Line-of-sight area identification method and line-of-sight area identification system
The method and system capture the entire face using a wide-angle imaging device and neural network to identify gaze areas, addressing the limitations of eye-focused technologies and enabling accurate gaze area identification for diverse users, including those with disabilities, without calibration or specialized equipment.
Patent Information
- Application Number
- JP2024015769
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2025-08-18
AI Technical Summary
Existing gaze area identification technologies rely heavily on high-resolution imaging of the eyes, which are variable in size and require complex calibration, making them difficult for individuals with severe disabilities to use, and are limited in application scope.
A method and system that captures the entire face using a wide-angle imaging device and employs a convolutional neural network to identify gaze areas by analyzing gaze coordinate points positioned at least 30 degrees to the left and right of the user's field of view, eliminating the need for high-resolution eye-focused imaging and calibration.
Enables accurate gaze area identification for a wide range of subjects, including those with severe disabilities, without the need for complex calibration or specialized equipment, and supports efficient gaze-based input and assistance.
Smart Images

Figure 2025120729000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a gaze area specifying method for specifying a gaze area of a user, and also to a gaze area specifying system. [Background technology]
[0002] Numerous methods for estimating gaze have been proposed (Non-Patent Document 1, Non-Patent Document 2), and the results are being used in a variety of fields. One example is a gaze area identification device (Non-Patent Document 3) that is designed to enable patients with cerebrovascular disease or amyotrophic lateral sclerosis, who have difficulty speaking, to express their intentions. This device has the ability to select icons on a screen with gaze or input characters to write and read sentences aloud, making it possible to communicate intentions and feelings, and is expected to improve quality of life.
[0003] Patent Document 1 discloses the following technology regarding a system for eye signals. A system and method are provided for identifying a device wearer's intention based primarily on eye movements. The system may be included in unobtrusive headwear that performs eye tracking and controls a screen display. The system may also utilize a remote eye tracking camera, a remote display, and / or other auxiliary inputs. The screen layout is optimized to facilitate the formation and reliable detection of high-speed eye signals. The detection of eye signals is based on tracking physiological eye movements that are under the voluntary control of the device wearer. The detection of eye signals results in operations compatible with wearable computing and a wide range of display devices.
[0004] Patent Document 2 discloses an information processing device for estimating a person's gaze direction, which includes an image acquisition unit that acquires an image including a person's face, an image extraction unit that extracts a partial image including the person's eyes from the image, and an estimation unit that acquires gaze information indicating the person's gaze direction from a trained learning device that has performed machine learning to estimate the gaze direction by inputting the partial image into the learning device.
[0005] Patent document 3 discloses a gaze recognition device that is characterized by comprising an eyeball displacement amount detection means that detects the amount of eyeball displacement in response to changes in the operator's gaze, a signal smoothing means that receives the displacement amount detection output from the eyeball displacement amount detection means and smooths the eyeball displacement amount for a predetermined period of time, and a neural network that receives the smoothed eyeball displacement amount output from the signal smoothing means and learns and recognizes the correspondence with coordinates on a display screen.
[0006] Patent Document 4 discloses a gaze area identification method, which includes an imaging step of capturing an image of the entire face of a user looking at a screen displaying input information to an electronic computer using imaging means fixed around a display unit having the screen to acquire image data, and an identification step of inputting the image data into a pre-created trained model of gaze area prediction to identify the gaze area of the user, wherein the pre-created trained model of the gaze area is created by using an image including the entire face having information on gaze coordinate points as training data, detecting a face area from the training data, and performing machine learning on a convolutional neural network of the face area. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Special Publication No. 2017-526078 [Patent Document 2] Japanese Patent Application Publication No. 2019-28843 [Patent Document 3] Japanese Patent Application Publication No. 5-46309 [Patent Document 4] Patent No. 7296069 Summary of the Invention [Problem to be solved by the invention]
[0008] As disclosed in Patent Documents 1 to 3, when identifying the gaze area and performing input / operation using it, processing is generally performed focusing on information obtained by capturing an image of the user's eyes. Relying on eye information requires capturing an image with high resolution of a small area. In particular, the size of the eyes varies from person to person, and even when focusing on the relative positions of the so-called whites and pupils of the eyes, the sizes of these also vary from person to person.
[0009] Furthermore, even in the black part of the eye, it is difficult to distinguish between the color tones of the pupil and the iris, so it is necessary to identify the gaze, necessitating higher-resolution images. When attempting to identify the gaze centered on such eyes, the camera specifications become important, and the image processing load of the captured data is also large. Furthermore, because there are individual differences in identifying the gaze area, calibration is also required before starting operation.
[0010] However, there are cases where the gaze area identification device is required to be used by children at special needs schools who have severe multiple disabilities that make it difficult for them to express their intentions or perform complex operations, and it may be difficult for such children to perform calibration processes that require repeated multiple advanced processes.
[0011] Furthermore, Patent Document 4 discloses a method for identifying a gaze area in which input for operating a computer is performed by the gaze of an operator without requiring calibration. However, the scope of application is limited because it uses an imaging means fixed around the periphery of a display unit having a screen, and there is a demand for a technology that can be used for a wider range of subjects.
[0012] Under such circumstances, an object of the present invention is to provide a method for identifying a user's line of sight that can be used for a wide range of subjects. [Means for solving the problem]
[0013] The present inventors have conducted extensive research to solve the above problems and have found that the following inventions meet the above objectives, thereby completing the present invention.
[0014] <1> an imaging step of capturing an image of the entire face of the user with an imaging means to acquire imaging data; and a specifying step of inputting the imaging data into a trained model for gaze area prediction created in advance to specify the gaze area of the user, the pre-created trained model of the gaze area is created by using an image including an entire face having information on gaze coordinate points as training data, detecting a face area from the training data, and performing machine learning on the face area using a convolutional neural network; the imaging means is disposed in front of the user, A method for identifying a gaze area, wherein the gaze coordinate point information of the learning data uses a plurality of gaze coordinate points arranged at least 30 degrees to the left and right of the user's field of view. <2> The plurality of gaze coordinate points are arranged in a matrix of at least five points in the horizontal direction and five points in the vertical direction, and the distance between adjacent gaze coordinate points is 15 cm or more. <1> The gaze area specifying method according to claim 1. <3> The gaze area of the user identified in the identifying step is identified by finding a weighted center of gravity for a plurality of gaze area candidates with top estimation probabilities. <1> or <2> The gaze area specifying method according to claim 1. <4> the imaging data in the imaging step is continuous imaging data including a plurality of imaging data obtained by imaging the entire face of the user multiple times consecutively at predetermined time intervals, The specifying step specifies a line-of-sight area for each piece of imaging data of the continuous imaging data; a time averaging process step of performing a moving average process on each of the gaze areas identified based on the continuous image capturing data, and setting the result as an average gaze area for the predetermined time period; <1> ~ <3> 10. The method for specifying a gaze area according to claim 9, wherein <5> an imaging means for imaging the entire face of a user and acquiring imaging data; an identification unit that inputs the imaging data into a trained model of gaze area prediction created in advance and identifies the gaze area of the user, the pre-created trained model of the gaze area is created by using an image including an entire face having information on a plurality of gaze coordinate points as training data, detecting a face area from the training data, and performing machine learning on the face area using a convolutional neural network; the imaging means is disposed in front of the user, A gaze area identification system in which the gaze coordinate point information of the learning data uses multiple gaze coordinate points arranged so that they are at least 30 degrees to the left and right of the user's field of view. [Effects of the Invention]
[0015] According to the present invention, the user's line of sight can be identified for a wide range of objects. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a schematic diagram of a gaze area specifying system according to an embodiment of the present invention. [Figure 2] 1 is a schematic diagram for explaining gaze coordinate points used in the gaze area method and the like of the present invention. FIG. [Figure 3] FIG. 1 is a flow diagram illustrating an example of a gaze region method of the present invention. [Figure 4] 1 is an example of an image used as learning data for creating a trained model of the gaze area of the present invention. [Figure 5] FIG. 10 is a diagram showing the accuracy when the gaze area is identified in a test example. [Figure 6] FIG. 10 is a diagram showing the accuracy when the gaze area is identified in a test example. [Figure 7] FIG. 10 is a diagram showing the accuracy when the gaze area is identified in a test example. [Figure 8] FIG. 10 is a diagram showing the accuracy when the gaze area is identified in a test example. [Figure 9]FIG. 10 is a diagram illustrating an outline of specifying a gaze area by a weighted center of gravity. [Figure 10] FIG. 10 is a diagram showing the accuracy when the gaze area is identified in a test example. DETAILED DESCRIPTION OF THE INVENTION
[0017] The following describes in detail an embodiment of the present invention, but the following description of the constituent elements is one example (typical example) of an embodiment of the present invention, and the present invention is not limited to the following content unless the gist of the present invention is changed. Note that when the expression "to" is used in this specification, it is used as an expression that includes the numerical values before and after it.
[0018] [Gaze area identification method of the present invention] The gaze area identification method of the present invention includes an imaging step of imaging the entire face of a user with an imaging means and acquiring imaging data, and an identification step of inputting the imaging data into a pre-created trained model for gaze area prediction to identify the gaze area of the user. Furthermore, in the gaze area identification method of the present invention, the pre-created trained model for gaze area is created by using an image including the entire face having information on gaze coordinate points as training data, detecting a face area from the training data, and performing machine learning on the face area using a convolutional neural network, the imaging means is disposed in front of the user, and the gaze coordinate point information of the training data uses a plurality of gaze coordinate points that are positioned so as to be at least 30 degrees to the left and right of the user's field of view.
[0019] [Gaze area identification system of the present invention] The gaze area identification system of the present invention includes an imaging means for imaging the entire face of a user and acquiring imaging data, and an identification unit for inputting the imaging data into a pre-created trained model for gaze area prediction to identify the gaze area of the user. Furthermore, the gaze area identification system of the present invention includes the pre-created trained model for gaze area created by using an image including the entire face having information on a plurality of gaze coordinate points as training data, detecting a face area from the training data, and performing machine learning on the face area using a convolutional neural network, wherein the imaging means is disposed in front of the user, and the gaze coordinate point information of the training data uses a plurality of gaze coordinate points that are positioned so as to be at least 30 degrees to the left and right of the user's field of view.
[0020] In the present application, the gaze area specifying method of the present invention can also be performed by the gaze area specifying system of the present invention, and the corresponding configurations in the present application can be mutually utilized.
[0021] The inventors have studied a method for identifying the user's line of sight. Patent Document 4 allows input by gaze without the need for calibration. Therefore, even people with severe and multiple disabilities who have intellectual disabilities and have difficulty moving their bodies as they wish can move their gaze according to calibration operation instructions, without the need for high-resolution cameras or dedicated attachment devices that are used in conventional methods to make judgments based on eye images.
[0022] However, there are cases where it is required to identify not only the gaze of the user relative to the content displayed in the image, but also the gaze of the user. For example, for a user who requires assistance and has difficulty speaking, the assistant can efficiently provide assistance by understanding what the user is looking at. Also, by looking in the direction the user wants to move, the assistant can obtain information about what the user is looking at and move in that direction. Furthermore, by utilizing the identification of the user's gaze, gaze-based input can be performed in various situations.
[0023] In response to such demands, the inventors have improved the imaging units such as cameras and the gaze coordinate points of learning data, resulting in the realization of methods and systems according to the present invention that can be used for a wide range of subjects in various environments.
[0024] This is thought to be because by using gaze coordinate points over a wide field of view, not only the positions of the pupils and whites of the eyes but also the orientation of the entire face, which is affected by factors such as posture of the neck, can be comprehensively analyzed as input data, and by targeting the entire face, it is easier to obtain useful data.
[0025] Fig. 1 is a schematic diagram of a gaze area specifying system according to an embodiment of the present invention. Fig. 2 is a schematic diagram for explaining gaze coordinate points used in the gaze area specifying method etc. of the present invention. Fig. 3 is a flow chart showing an example of the gaze area specifying method of the present invention.
[0026] [Gaze area identification system] Fig. 1 is a schematic diagram of a gaze area identification system 10. As shown in Fig. 1, the gaze area identification system 10 has an imaging means 21, a control unit 3 including a gaze area identification unit, and a memory unit 4. The gaze area identification system 10 may also have a display unit 5. The control unit 3, memory unit 4, and output unit 5 are built into a computer 6. The computer 6 and the imaging means 21 can input and output signals via wires or wirelessly.
[0027] [Flow for identifying gaze area] Fig. 3 is a flow diagram showing an example of a gaze area identification method. As shown in Fig. 3, in the gaze area identification method according to the present invention, step S11 involves capturing an image of a user's face looking at a screen displaying information to be input to a computer using an imaging means to obtain image data. Step S21 involves inputting the image data captured in step S11 into a trained model for gaze area prediction created in advance to identify the user's gaze area. Step S31 involves displaying the gaze area identified in step S21 on a display.
[0028] [Gaze Area Identification System 10] Gaze area identification system 10 is a system that identifies the gaze area of user 91. For example, the gaze area of user 91 is identified outdoors or indoors. The identified gaze area can be displayed on display unit 5 or used for operations based on the gaze of user 91.
[0029] [User 91] A user 91 of the gaze area identification system 10 is a person whose gaze area is to be identified. The user 91 is a person who identifies a gaze and expresses his or her intentions through the gaze, for example, in a situation where speech is difficult or due to physical reasons.
[0030] [Image capture] The imaging means 21 is a means for capturing an image of the entire face of the user 91. The imaging means 21 is disposed in front of the user. This imaging means 21 can capture an image of the user 91 and acquire image data. The imaging means 21 is disposed in a position where it can capture an image of the face of the user 91 when the user 91 is using the gaze area identification system 10, and has a number of pixels, an angle of view, and the like that can capture an image of the face of the user 91.
[0031] The imaging means 21 only needs to be able to capture an image of the entire face of the user 91, and can be placed in front of or diagonally in front of the user. The imaging means 21 can be placed at a distance of several tens of centimeters to several meters from the eyes of the user 91. For example, when the user 91 is used in a seated position, the imaging means 21 can be placed in front of the user 91 at a height similar to the user's eye level using a stand or the like. Alternatively, when the user 91 is in a wheelchair or the like, the imaging means 21 can be placed on the armrest of the wheelchair or the like. In the gaze area identification system 10 of FIG. 1, the imaging means 21 is placed on the armrest of a wheelchair-like chair 22.
[0032] The lower limit of the distance between the imaging means 21 and the eyes of the user 91 is preferably 30 cm or more, or 50 cm or more. The upper limit of the distance between the imaging means 21 and the eyes of the user 91 is preferably 5 m or less, 3 m or less, or 2 m or less. This distance makes it easy to photograph the face of the user 91, and is an appropriate range for the usage environment of the user 91 and the range for specifying the line of sight area.
[0033] The imaging means 21 may have a pixel count of about 1 MP (1 million pixels) or more and an angle of view of about 60 degrees or more, depending on factors such as the distance between the imaging means 2 and the user 91. If the number of pixels is too high, the analysis load will be heavy and processing to reduce the number of pixels may be necessary, so it may be about 12 MP or less. Also, if the angle of view is too wide, the area around the face may not be captured sufficiently, increasing the number of surrounding elements and potentially resulting in multiple people being captured, so it may be about 90 degrees or less. The pixel count may be about 0.8 MP to 12 MP or about 2 MP to 6 MP, and the angle of view may be about 70 to 90 degrees.
[0034] [Specify gaze area] The identification unit inputs the captured image data into a pre-created trained model for gaze area prediction to identify the gaze area of the user 91. The trained model is created by using an image including the entire face having information on gaze coordinate points as training data, detecting a face area from the training data, and performing machine learning on the face area using a convolutional neural network.
[0035] [Gaze coordinate point] Figure 2 is a schematic diagram for explaining an example of gaze coordinate points. Figure 2 shows an example of a user, imaging means, and gaze coordinate points for acquiring learning data. In this example, a person (user) is seated in a chair, and a web camera acting as an imaging means is placed in the user's line of sight, with a gaze coordinate point (the point in Figure 2) set on the wall beyond.
[0036] The gaze coordinate point information of the learning data uses a plurality of gaze coordinate points arranged at least 30 degrees to the left and right of the field of view of the user 91. By arranging a group of gaze coordinate points in such a wide area and acquiring learning data, it is possible to identify the line of sight corresponding to a wide field of view.
[0037] The positional relationship between each gaze coordinate point and its adjacent points depends on the user, the imaging means, and the location where the gaze coordinate points are set, but can be set to, for example, 15 cm to 50 cm, or 20 cm to 40 cm. The gaze coordinate points can be set, for example, by providing dots at predetermined intervals on a wall in a room. These dots can be pasted or written on the wall, or projected using optical means such as a projector. The distance between a person and the wall or other location where the gaze coordinate points are set can be, for example, approximately 80 cm to 2 m.
[0038] The viewing angle can be 30 degrees or more, 40 degrees or more, or 50 degrees or more. There is no need to set an upper limit, but if it is too wide, it may be difficult to set a gaze coordinate point or may be impractical, so it may be 70 degrees or less, 65 degrees or less, or 60 degrees or less.
[0039] [Pre-trained model] The pre-created trained model of the gaze area uses images containing information on multiple gaze coordinate points as training data. Furthermore, a face area is detected from the training data and used. The trained model is created by performing machine learning using a convolutional neural network on images from which a face area has been detected as training data, and extracting features related to the face and eyes. Since this trained model targets the face area, it extracts features related to the face and eyes.
[0040] To obtain a trained model, a large number of facial images with identified gaze points and regions are used. For example, first, training data of facial images is input. Next, facial regions are detected and extracted. Next, machine learning is performed using the information of the extracted facial regions as training data. Then, a trained model is obtained.
[0041] In the step of inputting face image learning data, a face image of a person looking at the gaze coordinate point is captured by an imaging means and input as learning data. A face region is then identified and extracted from the image. The face region is then extracted as input data. At this time, non-face parts that may become noise may be weighted less or not used in the learning data.
[0042] To extract a facial region, Haar-Like features can be used. Note that if an image does not allow a facial region to be identified, it may be excluded from the learning data.
[0043] Then, machine learning is performed using the face region as training data. For machine learning, images with identified relationships with gaze points, such as 100 or more, 1,000 or more, or 10,000 or more, are used as a training dataset. Note that, to prevent overlearning, training data of 100,000 or less, 50,000 or less, or 30,000 or less may be used. The training dataset may be obtained by acquiring data from a large number of subjects under roughly the same environmental conditions.
[0044] For machine learning, it is preferable to use a convolutional neural network (CNN). CNN models have layers such as convolutional layers and pooling layers, and various models exist depending on the number and combination of these layers, with the most appropriate one being adopted. For example, VGG, GoogLeNet (Inception), Xception, ResNet-18, etc. can be used. Specifically, a convolutional neural network based on the structure of VGG16, a type of VGG, can be used, with the final fully connected layer calculating the estimated probability of the gaze area.
[0045] References for VGG: Liu, S.; Deng, W. "Very deep convolutional neural network based image classification using small training sample size". Proceedings of 2015 3rd IAPR Asian Conference on Pattern Recognition (ACPR). Kuala, Lumpur, 2015-11-03 / 06, p.730-734, doi: 10.1109 / ACPR.2015.7486599.
[0046] References for GoogLeNet (Inception): Szegedy, C.; Liu, W.; et al. "Going deeper with convolutions". Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Boston, MA, 2015-07-07 / 12, p.1-9, doi: 10.1109 / CVPR.2015.7298594.
[0047] References for Xception: Chollet, F. "Xception: deep learning with depthwise separable convolutions". Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, HI, 2017-07-? / 26, p.1800-1807, doi: 10.1109 / CVPR.2017.195.
[0048] It should be noted that there seems to be an error in the original text where "2017-07-21 / 26" is miswritten as "2017-07-? / 26" in the translation of the Xception reference part. It has been corrected as much as possible while maintaining the format.ResNet-18 references: Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun, Deep Residual Learning for Image Recognition, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016), pp.770-778.
[0049] The acquired trained model is stored in the storage unit 4 of the electronic computer 6 shown in FIG. 1 etc., and is used for processing in the identification unit.
[0050] The storage unit 4 is a memory that stores the trained model, captured image data, extracted face regions, programs for processing these, and the like.
[0051] The identified gaze area can be used in various ways. For example, by looking at the location where the user wants to move, it can be used to input operations for the area the user is looking at. This can be used to move a wheelchair to that location, for example. Furthermore, it can be used to assign the role of a start / stop switch for a device to an object in a room, for example, to turn on an off light by gazing at it for a certain period of time, or to turn off an on light by gazing at it for a certain period of time. In addition, the user's gaze area can be displayed on a display unit. This supports and assists the user's gaze to express their intentions, making it possible to efficiently convey the user's gaze information to assistants 81 around the user, such as caregivers.
[0052] The display unit may be a monitor connected to a computer such as a personal computer, a monitor integrated with a computer such as a tablet terminal or laptop computer, or an image projected by a projector.
[0053] The display unit may be provided with an appropriate imaging means, such as a backup camera capable of capturing an image of the rear, opposite the user, and may display an image from the backup camera on the display unit 5, superimposing the identified gaze area 51 on the image. The gaze area identification system of FIG. 1 displays an image of the direction in which the chair 22 on which the user 91 is sitting is facing on the tablet terminal-like display unit 5 used by the assistant 81 of the user 91, and superimposes the gaze area 51 of the user 91. This makes it easier for the assistant 81 to identify the gaze of the user 91, thereby making it easier to confirm the user 91's intentions, etc.
[0054] [Electronic computer 6] The control unit 3 and output unit 5 can be used as application software installed in the electronic computer 6. The electronic computer 6 can also be a tablet terminal that is further integrated with the display unit 5 and the storage unit 4.
[0055] Weighted Center of Gravity Processing The gaze area identification can be performed by performing weighted center of gravity processing. This weighted center of gravity processing is performed by determining the weighted center of gravity of multiple gaze area candidates with the highest estimated probability, based on the gaze area of the user identified in the identification step. The weighted center of gravity processing can be performed by a weighted center of gravity processing unit.
[0056] When identifying the gaze area, an estimated probability is indicated according to the design of the gaze coordinates. The point with the highest estimated probability may be identified as the gaze area. If there are multiple points with a relatively high estimated probability, it is effective to calculate their weighted center of gravity. By performing weighted center of gravity processing, it is expected that the accuracy of identifying the gaze area and the error distance will be improved and the error distance will be reduced. The number of coordinates used to calculate the weighted center of gravity may be two or more, or three or more, depending on the design of the gaze coordinates and the trained model. Furthermore, it is possible to calculate the weighted center of gravity for all areas, but since areas with low probability have little impact and there is no risk of excessive variation, it is also possible to set an upper limit of 10 points or less, 8 points or less, or 6 points or less.
[0057] [Time averaging processing] The imaging data in the imaging process is continuous imaging data including multiple imaging data obtained by imaging the entire face of the user 91 multiple times consecutively at predetermined time intervals, and the identification process identifies a gaze area for each piece of imaging data in the continuous imaging data, and the process can include a time averaging process in which each gaze area identified based on the continuous imaging data is subjected to moving average processing to obtain a time-averaged gaze area for the predetermined time.
[0058] The time averaging processing unit is a part that performs averaging processing to obtain a time-averaged gaze area for a predetermined time. The time averaging processing step processes imaging data using continuous imaging data including multiple imaging data obtained by capturing images of the user's face multiple times consecutively at predetermined time intervals. In performing the time averaging processing, a gaze area is identified for each imaging data of the continuous imaging data by an identification step. Then, each gaze area identified based on the continuous imaging data is subjected to a moving average process to obtain the time-averaged gaze area for the predetermined time.
[0059] A person's gaze may fluctuate in a short period of time. Even if the person intends to look at the area where the options are displayed, they may look around to check, their gaze may wander within the area where the options are displayed, or their gaze may be processed incorrectly due to blinking, etc. In order to eliminate these fluctuations, it is preferable to perform an averaging process using a moving average.
[0060] In particular, the gaze area identification method of the present invention requires a low analysis load, allowing for the identification of a gaze area in a short time. Therefore, the gaze area can be identified in real time even for image data captured continuously at a constant frame rate. Images captured at a frame rate of approximately 15 to 60 fps can be used. Blinking noise is considered to be on the order of a few frames, and gaze-based intentions can be analyzed to indicate a high probability of intention within approximately 0.5 seconds. Therefore, for example, with a frame rate of 30 fps, a more reliable gaze area can be identified by performing a moving average process over approximately 10 to 20 frames.
[0061] According to the present invention, by using a general-purpose web camera or the like, it is possible to build a system that estimates a subject's spatial gaze area without the need for dedicated equipment. To achieve this, an appearance-based method using CNN can be used to create a new algorithm for spatial gaze area estimation. The present invention not only eliminates the burden and sense of restraint of wearing a device, but also has the advantage of not requiring calibration. It is expected that its introduction into various devices that utilize gaze will contribute to improving usability.
[0062] [Test example] The following test was conducted to determine the gaze area according to the present invention.
[0063] [Experiment 1 (Estimation at 40cm interval between fixation points)] First, we conducted an experiment to estimate the gaze area in space at a distance of 40 cm as follows. Figure 2 is a schematic diagram explaining the user, webcam, gaze coordinate points, etc. used in the experiment.
[0064] The webcam (Logicool HD Pro Webcam C910 (resolution: 1920 x 1080)) was placed at a height of 130 cm from the floor. The gaze coordinate points were set by placing points on the walls of the room. The measurement area was divided into 35 areas, 5 vertically and 7 horizontally, with the center being 130 cm from the floor, the same height as the webcam, and the interval between adjacent gaze coordinate points being 40 cm. Each area is represented alphabetically from A to I vertically, and numbers from -6 to +6 horizontally. The measurement range is 160 cm vertically and 240 cm horizontally. When the user was seated in a chair, the height of their eyes was 130 cm from the floor. The horizontal distance between the user and the webcam was 60 cm, and the horizontal distance between the webcam and the wall where the gaze point was located was 50 cm. The user's eyes, the webcam, and the center of the gaze coordinate point were positioned in a roughly straight line. In this case, the user's field of view was approximately 54 degrees to the left and right.
[0065] Setting up the training dataset The training dataset was collected from 16 participants, with five images per person, for a total of 2,800 sets of images. Figure 4 shows an example of an image in the training dataset for creating a trained model of the gaze area of the present invention.
[0066] Learning conditions The EarlyStopping class was used for learning, and the learning was terminated if the highest accuracy was not exceeded for 10 consecutive epochs. The batch size was set to 20, the learning rate was set to 0.001, and the coordinate with the highest estimation probability was used as the estimated coordinate.
[0067] Test results Figure 5 shows the estimation results for each coordinate using the trained model. The vertical axis (Esrimated coordinates) shows the estimation results when gazing at the coordinates (Gaze coordinates) shown on the horizontal axis. When gazing at a certain coordinate (horizontal axis), the vertical axis shows the likelihood of the coordinate estimated (identified) by the developed system; the darker the color, the more likely the coordinate is to be estimated. For example, when the gazed-at coordinate (horizontal axis) is "A+2," if the system's estimated (identified) coordinate (vertical axis) is also estimated to be "A+2," the darker the color of "A+2" on the vertical axis. This experiment revealed that an estimation accuracy of 82.24% was achieved. Therefore, it can be seen that there is a lot of data in the diagonal area of the figure, which indicates high accuracy.
[0068] Next, the distribution of estimation probabilities when gazing at coordinates is shown in the figure. The horizontal plane represents the estimated coordinate, and the height represents the estimation probability. Figure 9 shows the result when gazing at coordinate E0. The estimation probability for E0 was the highest at 89.88%, followed by the coordinate above it, C0, at 4.87%, and the coordinate below it, G0, at 2.63%. From the above, it was confirmed that the estimation probability for the correct coordinate was highest, and the other estimation probabilities were concentrated at adjacent coordinates and were relatively low, indicating no extreme bias in accuracy.
[0069] [Experiment 2 (Estimation at a 20cm gap between fixation areas)] Next, a test was conducted with a narrower interval between the gaze regions than in Experiment 1. Each region was 20cm x 20cm, and the measurement space was divided into a total of 117 regions, 9 vertically and 13 horizontally. In addition to the images from Experiment 1, new gaze images were collected, resulting in a total of 9,360 sets of data. Other learning conditions and estimation methods were the same as in Experiment 1.
[0070] Figure 7 shows the accuracy of each gaze area when identifying the gaze area in Experiment 2. Figure 8 shows the accuracy with which the top four points of the estimated gaze areas were the actual gaze areas. The average accuracy at the point with the highest probability in Figure 7 was 60.22%. There was a tendency for this to decrease particularly away from the center. On the other hand, near the center surrounded by the solid line, the accuracy was 77.33%, which shows that a higher accuracy than average was obtained.
[0071] [Table 1]
[0072] Table 1 shows the probability that a gaze coordinate is included within the top n estimated probabilities. As the number of coordinates n increases, the probability that a gaze coordinate is included increases, reaching 93.05% for the top four points.
[0073] In Experiment 1, it was shown that coordinates with high estimation probability tended to be clustered around the correct coordinate and its adjacent coordinates. This suggests that the accuracy did not drop significantly compared to Experiment 1 by increasing the number of coordinates, and that the model for estimating the gaze area in space was properly trained, resulting in this result due to the narrower coordinate interval.
[0074] Figure 8 shows the probability that gaze coordinates are included in the top four points for each coordinate. It is clear that estimation is possible without bias across the entire gaze area, and that a model has been created that is in line with the theme of this research. Therefore, it is believed that accuracy could be further improved by using information on the top n points with the highest estimation probability.
[0075] [Weighted center of gravity] When investigating the estimation of the gaze area using information on the top few points with the highest estimated probability, the weighted center of gravity was calculated as follows. First, the gaze coordinates were estimated using the learning model derived in Experiment 2. When the coordinates of the top n points with the highest estimated probability are (xi, yi) and their respective estimated probabilities are pi (i = 1, 2, , n), the weighted center of gravity coordinates (xW, yW) are given by the following formula. Then, discretization was performed to obtain the representative coordinates of the gaze area.
[0076]
number
[0077] FIG. 9 is a diagram illustrating an overview of identifying the gaze area using this weighted center of gravity. In the example of FIG. 9, the actual gaze area is E0. When estimating the gaze area using the weighted center of gravity, the top five points with the highest estimated probability are used. The point with the highest probability is E1 at 35%, followed by E0 at 35%, D0 at 15%, E-1 at 12%, and F0 at 8%. By calculating the weighted center of gravity coordinates of these top five areas, area E0 becomes the estimated area. This is an example where, if only the area with the highest specified probability were identified, what would have been erroneously estimated as an adjacent area, E+1, can be estimated as the actual gaze area, E0, by using the weighted center of gravity.
[0078] Figure 10 shows the sum of the error distances for the test data for each region using the weighted centroid of the top four points, based on the experimental results of the test example. Weighted centroid processing resulted in a 2.23% improvement in the discretized coordinate information. Furthermore, the total sum of the error distances for the entire region was 820.49 in Experiment 2 before weighted centroid processing, but after weighted centroid processing it was reduced to 640.87, a 21.9% reduction.
[0079] Next, we analyzed the histogram of the distance between the gaze coordinates and the identified estimated coordinates for all test data. By using weighted centroid processing, the number of data points between distances of 0 and 0.5 was roughly the same as when weighted centroid processing was not used (Experiment 2), while the number of data points between distances of 1 and 1.5 decreased. Overall, the distances became smaller, resulting in estimation results with less variance. This is because the estimated coordinates identified using weighted centroid processing are obtained as continuous values, which we can see contributes to improving internal accuracy. [Industrial Applicability]
[0080] INDUSTRIAL APPLICABILITY The present invention is industrially useful because it can identify the user's line of sight area and can be used for input by line of sight, display of line of sight, etc. [Explanation of symbols]
[0081] 10. Gaze Area Identification System 21 Imaging means 22 Chair 3. Control Unit 4 Storage section 5 Display section 51 Gaze area 6 Electronic computer 81 Auxiliary 91 User
Claims
1. an imaging step of capturing an image of the entire face of the user with an imaging means to acquire imaging data; and a specifying step of inputting the imaging data into a trained model for gaze area prediction created in advance to specify the gaze area of the user, the pre-created trained model of the gaze area is created by using an image including an entire face having information on gaze coordinate points as training data, detecting a face area from the training data, and performing machine learning on the face area using a convolutional neural network; the imaging means is disposed in front of the user, A method for identifying a gaze area, wherein the gaze coordinate point information of the learning data uses a plurality of gaze coordinate points arranged at least 30 degrees to the left and right of the user's field of view.
2. 2. The method for identifying a gaze area according to claim 1, wherein the plurality of gaze coordinate points form a matrix of at least five points in the horizontal direction and five points in the vertical direction, and are arranged so that the distance between adjacent gaze coordinate points is 15 cm or more.
3. The gaze area specifying method according to claim 1 , wherein the gaze area of the user specified in the specifying step is specified by finding a weighted center of gravity for a plurality of gaze area candidates with top estimated probabilities.
4. the imaging data in the imaging step is continuous imaging data including a plurality of imaging data obtained by imaging the entire face of the user multiple times consecutively at predetermined time intervals, The specifying step specifies a line-of-sight area for each piece of imaging data of the continuous imaging data; The gaze area specifying method according to claim 1, further comprising a time averaging process step of performing a moving average process on each of the gaze areas specified based on the continuous image capturing data, and setting the result as an average gaze area for the predetermined time period.
5. an imaging means for imaging the entire face of a user and acquiring imaging data; an identification unit that inputs the imaging data into a trained model of gaze area prediction created in advance and identifies the gaze area of the user, the pre-created trained model of the gaze area is created by using an image including an entire face having information on a plurality of gaze coordinate points as training data, detecting a face area from the training data, and performing machine learning on the face area using a convolutional neural network; the imaging means is disposed in front of the user, A gaze area identification system in which the gaze coordinate point information of the learning data uses a plurality of gaze coordinate points arranged at least 30 degrees to the left and right of the user's field of view.
Citation Information
Patent Citations
Sight line recognizing device
JP1993046309A
Systems and methods for biomechanics-based ocular signals for interacting with real and virtual objects
JP2017526078A
Information processing apparatus for estimating person's line of sight and estimation method, and learning device and learning method
JP2019028843A
Eye-gaze input device and eye-gaze input method
JP7296069B2