Control apparatus and estimation apparatus
The control device efficiently integrates gaze position estimation with pointing device operations to quickly reacquire and maneuver the pointer on large screens, addressing the issue of lost track by displaying it at the estimated gaze position when conditions are met.
Patent Information
- Application Number
- JP2024030847
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2025-09-10
AI Technical Summary
Users often lose track of the pointer on large screens, leading to increased time in searching and moving it to the desired position, and existing gaze-based pointer systems do not effectively integrate with pointing devices.
A control device that combines gaze position estimation with pointing device operations, displaying the pointer at the estimated gaze position when specific conditions are met, preventing continuous display control that interferes with manual manipulation.
Enables quick reacquisition and precise movement of the pointer to the target position without user inconvenience, even after losing sight, by integrating gaze estimation with pointing device signals.
Smart Images

Figure 2025132945000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a control device and an estimation device. [Background technology]
[0002] 2. Description of the Related Art Conventionally, control devices that control the position of a pointer displayed on a screen based on a position operation signal from a pointing device such as a mouse are widely known in personal computers and the like.
[0003] Furthermore, Patent Document 1 discloses a control device that uses a line-of-sight detection sensor to detect the line of sight of a user looking at the display unit of a digital camera, and performs AF (autofocus) to focus on a subject displayed at the detected line-of-sight position. This control device displays a line-of-sight pointer on the display unit that indicates the line-of-sight position, allowing the user to recognize the line-of-sight position on the display unit through the line-of-sight pointer. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2022-68749 Summary of the Invention [Problem to be solved by the invention]
[0005] With the recent increase in screen size, it is becoming increasingly common for users to lose track of the pointer displayed on the screen. In the past, users who lost track of the pointer would have to spend time searching for it, which meant it took a long time to move the pointer to the desired position.
[0006] The control device disclosed in Patent Document 1 does not use a pointing device, and the gaze pointer displayed on the display unit indicates the user's gaze position on the display unit, but is not a pointer whose position is controlled by a pointing device. Therefore, the control device disclosed in Patent Document 1 cannot be a means for solving the above problem. [Means for solving the problem]
[0007] In order to solve the above-mentioned problems, one aspect of the present invention is a control device that controls the position of a pointer displayed on a screen based on a position operation signal from a pointing device, and has a gaze position estimation means that estimates the user's gaze position on the screen, and when a predetermined pointer display condition is met, performs display control to display the pointer at the gaze position on the screen estimated by the gaze position estimation means, and does not perform the display control until the next time the predetermined pointer display condition is met. In the control device according to this aspect, when a predetermined pointer display condition is satisfied, a display control is performed in which a pointer whose position is operated by a pointing device is displayed at the user's gaze position on the screen estimated by the gaze position estimation means. Therefore, even if the user loses sight of the pointer displayed on the screen (or a pointer that has become hidden for some reason), the pointer is displayed in the user's line of sight by satisfying the predetermined pointer display condition (for example, by operating the pointing device), and the user can quickly find the pointer. If the display control were to continue even after the pointer appears in the user's line of sight, the pointer position would change in accordance with the user's line of sight, hindering subsequent pointer position manipulation using the pointing device. Therefore, the control device according to this aspect prevents the display control from being performed until the next time a predetermined pointer display condition is met. As a result, when the pointer appears in the user's line of sight and the user finds it, they can operate the pointing device to move the pointer from that position to the target position without any problems. According to this aspect, even if the user loses sight of the pointer displayed on the screen, the user can quickly find the pointer and move it to the target position without any trouble.
[0008] In the control device, the specified pointer display condition may include a condition that a position operation signal for changing the position of the pointer is received from the pointing device after a specified time has elapsed since the position of the pointer was last changed. While a user is operating a pointing device to change the position of the pointer on the screen, it is unlikely that the user will lose track of the pointer's position. However, if the user stops operating the pointing device for a while, for example, by performing another task or by moving away from the screen, the user will often lose track of the pointer's position when they try to operate the pointing device again to move the pointer to the target position. This is because, for example, the user forgets the last position of the pointer during the interruption, or the pointing device is operated without the user realizing it. According to this control device, when a position operation signal for changing the pointer position is received from the pointing device after a predetermined time has elapsed since the last change in the pointer position, the pointer is displayed at the user's gaze position on the screen estimated by the gaze position estimation means. Therefore, when a user who has stopped operating the pointing device operates the pointing device again to move the pointer to a target position, the pointer can be moved to the target position without the user being aware that they have lost sight of the pointer position.
[0009] In the control device, the predetermined pointer display condition may include a condition that a predetermined pointer display signal is received from the pointing device or another input device. According to this control device, when a predetermined pointer display signal is received from a pointing device or other input device, a pointer is displayed at the user's gaze position on the screen estimated by the gaze position estimation means. Therefore, if the user loses sight of the pointer, the pointer can be displayed immediately in the user's line of sight by performing a predetermined operation (an operation for displaying the pointer at the gaze position) on the pointing device or other input device, allowing the user to quickly find the pointer.
[0010] In the control device, the pointer display signal may be a position operation signal from the pointing device that repeatedly moves the position of the pointer back and forth on the screen. When users lose sight of the pointer position, many users operate the pointing device by repeatedly moving the pointer position back and forth (moving it left and right) to make it easier to recognize the pointer. With this control device, when users perform the same operation on the pointing device that many users do when they lose sight of the pointer position, the pointer is displayed in the user's line of sight, allowing the user to quickly find the pointer. Therefore, the user does not need to remember any special operations to display the pointer at the line of sight.
[0011] In the control device, the gaze position estimation means may include an acquisition unit that acquires imaging data of a user's facial image, and an estimation unit that estimates the gaze position of the user on the screen based on the imaging data acquired by the acquisition unit by having a computer execute a trained model that has been trained using a plurality of learning data including imaging data of the facial image of the learning subject at a viewing timing when the learning subject views a target mark displayed on the screen and position data of the target mark on the screen at the viewing timing. There are various methods for estimating a user's gaze position on a screen, but most of the conventional methods require equipment such as expensive sensors. On the other hand, there is a method that uses an inexpensive imaging device (camera) to capture an image of the user's face, etc., and estimates the user's gaze position on a screen from the captured image data. However, there are limitations to human programming when it comes to creating a computer program that can accurately estimate a user's gaze position from captured data such as a user's face image. In this control device, a program for estimating a user's gaze position uses a trained model trained using multiple training data sets, including image data of the subject's face at the timing when the subject gazes at a target mark displayed on the screen and position data of the target mark on the screen at the timing. To obtain this trained model, for example, the facial image of the subject is captured at the timing when the subject gazes at target marks displayed at various positions on the screen to obtain image data, and position data of the target mark at the timing. The trained model is then obtained by performing machine learning or the like using a training dataset containing multiple training data sets including image data of the face image and position data of the target mark. The trained model obtained in this manner makes it possible to accurately estimate the user's gaze position, not only when the subject is the user, but also when the user is a different user from the subject.
[0012] In the control device, the plurality of learning data may also include imaging data of an eye image including at least one eye of the learning subject extracted from imaging data of a facial image of the learning subject at the viewing timing. By changing the direction of their eyes, users can direct their gaze in a direction different from their facial direction (face direction), i.e., the direction of their face. Therefore, both facial direction (face direction) and eye direction (eye direction) are important characteristics when estimating the user's gaze position on the screen. In this control device, in addition to the image data of the face image of the training subject, image data of an eye image containing at least one of the training subject's eyes extracted from the image data is also included in the training data. In this way, by including image data of the eye image separately from the image data of the face image, a trained model can be obtained that produces estimation results in which the influence of the eye direction (eye direction) is emphasized. This makes it possible to estimate the user's gaze position with higher accuracy than when a trained model is obtained from training data that does not include image data of the eye image (in which the influence of eye direction on the estimation result is small). Furthermore, in this control device, the image data of the eye images is extracted from the image data of the face image of the person being trained, which is included in the training data, so a trained model can be obtained more cheaply than if the image data of the eye images were obtained through a separate imaging operation from the image data of the face image.
[0013] In the control device, the plurality of learning data may also include feature amount data generated from the imaging data. This makes it possible to obtain a trained model with improved estimation accuracy through feature engineering compared to a trained model obtained from training data that does not include such feature data (training data that includes only imaging data). In particular, when capturing facial images of a subject using an inexpensive imaging device such as a monocular camera used in webcams, the imaging data lacks depth information. Therefore, a trained model obtained from training data containing only imaging data obtained with such an inexpensive imaging device may not be able to estimate the user's gaze position on the screen with sufficient accuracy. With this control device, even when imaging data obtained with an inexpensive imaging device is used as training data, the lack of accuracy can be compensated for with feature data, resulting in a trained model that can estimate the user's gaze position with sufficient accuracy.
[0014] In the control device, the feature data may include at least one of face direction feature data for identifying the direction of the face of the learning subject, and eye direction feature data for identifying the direction of the eyeballs of the learning subject. As described above, a user can direct their gaze in a direction different from the direction of their face (face direction), i.e., the front direction of their face, by changing the direction of their eyeballs. Therefore, both the direction of their face (face direction) and the direction of their eyeballs (eye direction) are important features in estimating the user's gaze position. According to this control device, the trained model is obtained from training data including numerical data (face direction feature data, eye direction feature data) on at least one of the face direction and eye direction, which are important features for estimating the user's gaze position, and therefore the user's gaze position can be estimated with higher accuracy.
[0015] In the control device, the viewing timing may include the timing when a target mark moving on the screen is overlapped by a pointer whose position on the screen is controlled based on a position operation signal of a pointing device operated by the learner. In general, in order to obtain a trained model with high accuracy, it is desirable to collect a large amount of training data with various contents in a short time. Therefore, in order to obtain the trained model that estimates the gaze position of a user, a method is required for collecting a large amount of imaging data of various facial images of a training subject and position data of a target mark in a short time. One possible method for obtaining imaging data of various facial images of the student is for the student to repeatedly operate a pointing device to specify target marks displayed in sequence at different positions on the screen with a pointer. In this method, the student's facial image is captured at the timing when the student operates the pointing device and specifies the position with the pointer (for example, when the mouse is clicked). This makes it possible to obtain position data of the target mark specified with the pointer and imaging data of the student's facial image at that specified time, and by repeating this process, it is possible to collect a large amount of learning data. However, this method requires the user to perform a positioning operation (e.g., moving the mouse) to move the pointer to the position of the next target mark and a designation operation (e.g., clicking the mouse) to specify the target mark, which places a heavy burden on the user and is inconvenient. Moreover, the time it takes to move the pointer to the position of the next target mark is a blank period during which no imaging data can be acquired, so it takes a long time to obtain one piece of imaging data. In this control device, image data of the student's facial image and position data of the target mark are acquired each time a pointer, whose position is controlled based on a position operation signal from a pointing device operated by the student, overlaps with a target mark moving (continuously changing position) on the screen. This allows the student to continuously acquire image data of the facial image and position data of the target mark without having to perform a designation operation to specify the target mark, simply by operating the pointing device so that the pointer follows the target mark moving on the screen. This reduces the user's workload and the gaps in time during which image data cannot be acquired, making it possible to collect a large amount of image data of various facial images of the student and position data of the target mark in a short period of time. As a result, a highly accurate trained model can be easily obtained. Furthermore, by moving the target mark slowly across the screen, the control device can induce smooth pursuit movements in the eyes of the learner viewing the target mark, thereby improving the accuracy of the trained model's estimation of the gaze position when the user moves their gaze with smooth pursuit movements.
[0016] In the control device, the viewing timing may include the timing at which a pointer, whose position on the screen is controlled based on a position operation signal of a pointing device operated by the learner, overlaps with each target mark appearing at a different position on the screen. In this control device, image data of the student's facial image and position data of the target mark are acquired each time a pointer, whose position is controlled based on a position operation signal of a pointing device operated by the student, overlaps with each target mark appearing at a different position on the screen. This allows the student to acquire image data of the facial image and position data of the target mark simply by operating the pointing device to overlap the pointer with the target mark that appears (displayed) discretely on the screen, without having to perform a designation operation to specify the target mark. Therefore, a large amount of image data of various facial images of the student and position data of the target mark can be collected in a short time with little workload on the user. Furthermore, this control device can induce saccades in the eyes of the learner as they view target marks that appear at different positions on the screen, thereby improving the accuracy of the trained model's estimation of the gaze position when the user moves their gaze with saccades.
[0017] In the control device, the plurality of learning data may include image data of the user's face image acquired by the acquisition unit at a pointer designation timing when the user operates the pointing device to designate any position on the screen with the pointer, and position data of the pointer on the screen at the pointer designation timing. In this control device, an acquisition unit that acquires captured image data of a user's face image is used to acquire captured image data of a user (trainee)'s face image to be used as training data for the trained model. This allows training data to be collected while the user is performing some task while looking at the screen using a pointing device. Moreover, the captured image data of the user's face image included in this training data is that of the user himself / herself (not a training subject who is a different person from the user), and is captured in the user's usage environment (background, brightness, etc.). Therefore, the inclusion of such training data can improve the accuracy of estimating the gaze position of the user by the trained model.
[0018] Another aspect of the present invention is an estimation device that estimates a user's gaze position on a screen, comprising: an acquisition unit that acquires imaging data of a user's facial image; and an estimation unit that estimates the user's gaze position based on the imaging data acquired by the acquisition unit by having a computer execute a trained model that has been trained using a plurality of learning data including imaging data of the facial image of the learning subject at a viewing timing when the learning subject views a target mark displayed on the screen, feature data generated from the imaging data, and position data of the target mark on the screen at the viewing timing, wherein the feature data includes at least one of face direction feature data for identifying the direction of the learning subject's face and eye direction feature data for identifying the direction of the learning subject's eyes. There are various methods for estimating a user's gaze position, but most of the conventional methods require equipment such as expensive sensors. On the other hand, there is a method that uses an inexpensive imaging device (camera) to capture an image of the user's face, etc., and estimates the user's gaze position on the screen from the captured image data. However, there are limitations to human programming when it comes to creating a computer program that can accurately estimate a user's gaze position from captured image data such as the user's face image. In this aspect, a program for estimating a user's gaze position uses a trained model trained using a plurality of training data sets, including image data of the subject's face at the timing when the subject gazes at a target mark displayed on the screen and position data of the target mark on the screen at the timing. To obtain this trained model, for example, the facial image of the subject is captured at the timing when the subject gazes at target marks displayed at various positions on the screen to obtain image data, and position data of the target mark at the timing. The trained model is then obtained by performing machine learning or the like using a training dataset including a plurality of training data sets including image data of the facial image and position data of the target mark. The trained model obtained in this manner makes it possible to accurately estimate the user's gaze position, not only when the subject is the user, but also when the user is a different user from the subject. Furthermore, in this embodiment, the training data also includes feature data generated from imaging data of the face image of the training subject. This allows for a trained model with improved estimation accuracy through feature engineering compared to a trained model obtained from training data that does not include such feature data (trained data that includes only imaging data). In particular, when capturing facial images of a training subject using an inexpensive imaging device such as a monocular camera used in webcams, the imaging data lacks depth information. Therefore, a trained model obtained from training data containing only imaging data obtained using such an inexpensive imaging device may not be able to estimate the user's gaze position on the screen with sufficient accuracy. According to this aspect, even when imaging data obtained using an inexpensive imaging device is used as training data, the lack of accuracy can be compensated for by feature data, thereby obtaining a trained model that can estimate the user's gaze position with sufficient accuracy. Here, by changing the direction of the user's eyes, the user can direct their gaze in a direction different from the direction of their face (face direction), i.e., the direction of the face. Therefore, both the face direction (face direction) and the eye direction (eye direction) are important features for estimating the user's gaze position. According to this aspect, numerical data (face direction feature data, eye direction feature data) for at least one of the face direction and eye direction, which are important features for estimating the user's gaze position, is used as feature data. Therefore, the user's gaze position can be estimated with higher accuracy.
[0019] Yet another aspect of the present invention is an estimation device that estimates a user's gaze position on a screen, comprising: an acquisition unit that acquires image data of a user's facial image; and an estimation unit that estimates the user's gaze position on the screen based on the image data acquired by the acquisition unit by having a computer execute a trained model that has been trained using a plurality of learning data including image data of the face image of the learning subject at a viewing timing when the learning subject views a target mark displayed on the screen and position data of the target mark on the screen at the viewing timing, wherein the viewing timing includes a timing when a pointer whose position on the screen is controlled based on a position operation signal of a pointing device operated by the learning subject overlaps with the target mark moving on the screen. There are various methods for estimating a user's gaze position, but most of the conventional methods require equipment such as expensive sensors. On the other hand, there is a method that uses an inexpensive imaging device (camera) to capture an image of the user's face, etc., and estimates the user's gaze position on the screen from the captured image data. However, there are limitations to human programming when it comes to creating a computer program that can accurately estimate a user's gaze position from captured image data such as the user's face image. In this aspect, a program for estimating a user's gaze position uses a trained model trained using a plurality of training data sets, including image data of the subject's face at the timing when the subject gazes at a target mark displayed on the screen and position data of the target mark on the screen at the timing. To obtain this trained model, for example, the facial image of the subject is captured at the timing when the subject gazes at target marks displayed at various positions on the screen to obtain image data, and position data of the target mark at the timing. The trained model is then obtained by performing machine learning or the like using a training dataset including a plurality of training data sets including image data of the facial image and position data of the target mark. The trained model obtained in this manner makes it possible to accurately estimate the user's gaze position, not only when the subject is the user, but also when the user is a different user from the subject. Generally, in order to obtain a trained model with high accuracy, it is desirable to collect a large amount of training data with various contents in a short time. Therefore, in order to obtain the trained model that estimates the gaze position of a user, a method is required for collecting a large amount of imaging data of various facial images of the training subject and position data of target marks in a short time. One possible method for obtaining imaging data of various facial images of the student is for the student to repeatedly operate a pointing device to specify target marks displayed in sequence at different positions on the screen with a pointer. In this method, the student's facial image is captured at the timing when the student operates the pointing device and specifies the position with the pointer (for example, when the mouse is clicked). This makes it possible to obtain position data of the target mark specified with the pointer and imaging data of the student's facial image at that specified time, and by repeating this process, it is possible to collect a large amount of learning data. However, this method requires the user to perform a positioning operation (e.g., moving the mouse) to move the pointer to the position of the next target mark and a designation operation (e.g., clicking the mouse) to specify the target mark, which places a heavy burden on the user and is inconvenient. Moreover, the time it takes to move the pointer to the position of the next target mark is a blank period during which no imaging data can be acquired, so it takes a long time to obtain one piece of imaging data. In this mode, each time a pointer whose position is controlled based on a position operation signal from a pointing device operated by the learner overlaps with a target mark moving (continuously changing position) on the screen, captured image data of the learner's facial image and position data of the target mark are acquired. This allows the learner to continuously acquire captured image data of the facial image and position data of the target mark without having to perform a designation operation to specify the target mark, simply by operating the pointing device so that the pointer follows the target mark moving on the screen. This reduces the user's workload and the gaps in time during which image data cannot be acquired, making it possible to collect a large amount of captured image data of various facial images of the learner and position data of the target mark in a short period of time. As a result, a highly accurate trained model can be easily obtained. Furthermore, according to this aspect, by moving the target mark slowly on the screen, it is possible to induce smooth pursuit of the eyes of the learner who is viewing the target mark, thereby improving the accuracy of estimating the gaze position of the trained model when the user moves their gaze with smooth pursuit.
[0020] Yet another aspect of the present invention is an estimation device that estimates a user's gaze position on a screen, characterized by including: an acquisition unit that acquires imaging data of a user's face image; and an estimation unit that estimates the user's gaze position on the screen based on the imaging data acquired by the acquisition unit by having a computer execute a trained model that has been trained using a plurality of training data including imaging data of the user's face image acquired by the acquisition unit at a pointer designation timing when the user operates a pointing device to designate an arbitrary position on the screen with a pointer, and position data of the pointer on the screen at the pointer designation timing. There are various methods for estimating a user's gaze position, but most of the conventional methods require equipment such as expensive sensors. On the other hand, there is a method that uses an inexpensive imaging device (camera) to capture an image of the user's face, etc., and estimates the user's gaze position on the screen from the captured image data. However, there are limitations to human programming when it comes to creating a computer program that can accurately estimate a user's gaze position from captured image data such as the user's face image. In this aspect, a program for estimating a user's gaze position uses a trained model trained using a plurality of training data sets, including image data of the subject's face at the timing when the subject gazes at a target mark displayed on the screen and position data of the target mark on the screen at the timing. To obtain this trained model, for example, the facial image of the subject is captured at the timing when the subject gazes at target marks displayed at various positions on the screen to obtain image data, and position data of the target mark at the timing. The trained model is then obtained by performing machine learning or the like using a training dataset including a plurality of training data sets including image data of the facial image and position data of the target mark. The trained model obtained in this manner makes it possible to accurately estimate the user's gaze position, not only when the subject is the user, but also when the user is a different user from the subject. Moreover, in this embodiment, an acquisition unit that acquires captured image data of a user's face image is used to acquire captured image data of a user (student)'s face image to be used as training data for the trained model. This allows training data to be collected while the user is performing some task while looking at the screen using a pointing device. Moreover, the captured image data of the user's face image included in this training data is that of the user himself / herself (not a student other than the user), and is captured in the user's usage environment (background, brightness, etc.). Therefore, the inclusion of such training data can improve the accuracy of estimating the gaze position of the user by the trained model. [Effects of the Invention]
[0021] According to the present invention, even if the user loses sight of the pointer on the screen, the user can quickly find the pointer and move it to the target position without any trouble. [Brief explanation of the drawings]
[0022] [Figure 1] FIG. 1 is an external view schematically illustrating a personal computer as an information processing apparatus according to an embodiment. [Figure 2] A functional block diagram showing the main functions of the computer. [Figure 3] FIG. 2 is a block diagram showing the main hardware that makes up the main body of the personal computer. [Figure 4] 5 is a flowchart showing an example (control example 1) of the flow of pointer display control in the embodiment. [Figure 5] 10 is a flowchart showing another example (control example 2) of the flow of pointer display control in the embodiment. [Figure 6] 1A is an explanatory diagram showing an example of a screen displaying a target mark used when collecting learning data, and FIG. 1B is an explanatory diagram showing an example of a facial image of a learning subject viewing the target mark displayed on the screen. [Figure 7]An explanatory diagram of the learning phase of the trained model executed by the CPU in the device main body of the same computer. [Figure 8] (a) is an explanatory diagram showing an example of a face image of a training subject used as training data, (b) is an explanatory diagram showing an example of an eye image of the training subject's left eye used as training data, and (c) is an explanatory diagram showing an example of an eye image of the training subject's right eye used as training data. [Figure 9] 1A is an explanatory diagram showing an example of a face image of a learning subject when face orientation feature data is obtained, and FIG. 1B is an explanatory diagram showing position data of facial feature points obtained from the face image of FIG. [Figure 10] 1A is an explanatory diagram showing an example of an eye image of a learning subject when obtaining eyeball direction feature data, and FIG. 1B is an explanatory diagram of eye feature parameters obtained from the eye image of FIG. [Figure 11] An explanatory diagram showing an example of a machine learning model for creating a trained model in an embodiment. [Figure 12] FIG. 10 is an explanatory diagram showing an example of a method for obtaining imaging data of various facial images of a learning subject viewing a target mark and position data of the target mark. DETAILED DESCRIPTION OF THE INVENTION
[0023] Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings. In this embodiment, an example will be described in which a user uses an information processing device (personal computer: PC) consisting of a general desktop personal computer by operating a mouse as a pointing device and a keyboard as other input devices while looking at a display screen. However, the present invention is not limited to such PCs and can be applied to various devices as long as the device changes (moves) the position of a pointer displayed on the screen by operating a pointing device.
[0024] FIG. 1 is a schematic external view of a personal computer 100, which is an information processing apparatus according to this embodiment. FIG. 2 is a functional block diagram showing the main functions of the personal computer 100 in this embodiment. The personal computer 100 in this embodiment is a computer device such as a general-purpose personal computer or a server device. The personal computer 100 in this embodiment is mainly composed of a device main body 101 as a control device, a display 102 as a display means, input devices as operation reception means such as a keyboard 103 and a mouse 104 as a pointing device, and a web camera 105 as an imaging device.
[0025] The device main body 101 mainly includes a control unit 101a, a storage unit 101b, and a gaze position estimation unit 101c. The control unit 101a mainly performs input / output control such as control of the keyboard 103 and mouse 104, which are input devices, and control of the display 102, which is an output device. The storage unit 101b stores various computer programs for performing the functions of the personal computer 100. In particular, the storage unit 101b of this embodiment stores a trained model, which will be described later, as a computer program for realizing the function of the gaze position estimation unit 101c. The gaze position estimation unit 101c also functions as gaze position estimation means for estimating a position on the screen of the display 102 in the user's line of sight (user's gaze position).
[0026] FIG. 3 is a block diagram showing the main hardware constituting the device main body 101 of the personal computer 100 according to this embodiment. The device main body 101 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a camera controller 14, a mouse controller 15, a keyboard controller 16, a display controller 17, and a communication interface 18. The components constituting the device main body 101 are electrically connected to one another via bus lines 19 such as an address bus and a data bus.
[0027] The CPU 11 functions as a gaze position estimation unit 101c that estimates a position on the screen of the display 102 in the user's line of sight (user's gaze position) by executing a computer-readable program (such as an estimation program described below) stored in the ROM 12. The ROM 12 stores data, programs, etc. that the CPU 11 executes. The RAM 13 temporarily stores data, etc. when the CPU 11 executes a program.
[0028] The camera controller 14 is connected to a web camera 105, which serves as an imaging device for capturing an image of a user (a user viewing the display 102) using the personal computer 100. The camera controller 14 controls the operation of the web camera 105 and also functions as an acquisition unit for acquiring imaging data of the user's facial image captured by the web camera 105.
[0029] The mouse controller 15 is connected to the mouse 104, which is a pointing device, and controls the mouse 104 by receiving operation signals from the mouse 104 and sending them to the CPU 11. The mouse 104 sends at least a position operation signal for changing (moving) the position of a pointer displayed on the screen of the display 102, and a designation operation signal for designating a position on the screen of the display 102 with the pointer.
[0030] The keyboard controller 16 is connected to the keyboard 103 and controls the keyboard 103 by receiving operation signals from the keyboard 103 and sending them to the CPU 11 .
[0031] The display controller 17 is connected to the display 102 and performs display control on the screen of the display 102 in accordance with a display control signal from the CPU 11. In particular, the display controller 17 of this embodiment performs display control to change (move) the position of a pointer displayed on the screen of the display 102 based on a position operation signal from a mouse in accordance with a control signal from the CPU 11. The display controller 17 of this embodiment also performs display control to display a pointer at the user's gaze position estimated by the gaze position estimation unit 101c in accordance with a control signal from the CPU 11.
[0032] The communication interface 18 communicates with external devices and transmits and receives data.
[0033] A user of the personal computer 100 may lose sight of the pointer displayed on the screen of the display 102. In particular, in the case of a large-screen display 102 (including a multi-display), it is easy for the user to lose sight of the pointer.
[0034] Therefore, in this embodiment, a gaze position estimation unit 101c is provided that estimates the gaze position of the user on the screen of the display 102, and performs pointer display control to display the pointer at the estimated gaze position when predetermined pointer display conditions are met. As a result, even if the user loses sight of the pointer displayed on the screen (or a pointer that has been hidden for some reason), the pointer is displayed in the user's line of sight when the predetermined pointer display conditions are met, and the pointer can be found immediately.
[0035] If the above-described pointer display control were to continue even after the pointer is displayed in the user's line of sight, the pointer position would change following the movement of the user's line of sight, causing problems with subsequent pointer position manipulation using the mouse 104. For this reason, in this embodiment, the above-described pointer display control is not performed until the next time a predetermined pointer display condition is satisfied. As a result, when the pointer is displayed in the user's line of sight and the user finds the pointer, they can operate the mouse 104 to move the pointer from that position to the target position without any problems.
[0036] [Control example 1] FIG. 4 is a flowchart showing an example of the flow of pointer display control in this embodiment (hereinafter referred to as "control example 1"). When a position operation signal (a signal indicating the amount of movement of the mouse 104) is input from the mouse 104 (Yes in S1), the control unit 101a of the device main body 101 of the personal computer 100 performs display control to change (move) the position of the pointer displayed on the screen of the display 102 in accordance with the position operation signal (S2). This allows the user to move the position of the pointer on the screen in accordance with the movement of the mouse 104 by moving the mouse 104.
[0037] The predetermined pointer display condition in this control example 1 is that a predetermined time has elapsed since the pointer position was last changed (moved), and then a position operation signal for changing the pointer position is input (accepted) from mouse 104. That is, when a predetermined time has elapsed since the pointer position was last changed (Yes in S3), and a position operation signal is input from mouse 104 (Yes in S4), control is performed to display the pointer at a position on the screen of display 102 that is in the user's line of sight (line of sight position) (S5 to S7).
[0038] Specifically, when a predetermined pointer display condition is satisfied (Yes in S3, Yes in S4), first, control unit 101a controls webcam 105 to capture an image of the user's face with webcam 105 (S5). Then, control unit 101a acquires imaging data of the user's face image captured by webcam 105 and passes the imaging data to gaze position estimation unit 101c. Then, gaze position estimation unit 101c, which has acquired the imaging data of the user's face image, estimates the position of the user's gaze (gaze position) on the screen of display 102 based on the imaging data (S6), and the estimation result is sent to control unit 101a. Control unit 101a, which has received the estimation result of the user's gaze position, performs display control to display a pointer at the position on the screen indicated by the estimation result (S7).
[0039] If a user interrupts operation of the mouse 104 for a while, for example by performing another task or moving away from the screen, the user often loses track of the pointer position when attempting to operate the mouse 104 again to move the pointer to a target position. According to this control example 1, when a user who has interrupted operation of the mouse 104 for a predetermined period of time attempts to operate the mouse 104 again to move the pointer to a target position, the user can quickly find the pointer without being aware that he or she has lost track of the pointer position. Then, the user can operate the mouse 104 to move the found pointer to the target position.
[0040] [Control example 2] FIG. 5 is a flowchart showing another example of the flow of pointer display control in this embodiment (hereinafter referred to as "control example 2"). In this control example 2, when a position operation signal is input from the mouse 104 (Yes in S11), the control unit 101a performs display control to change (move) the position of the pointer displayed on the screen of the display 102 in accordance with the position operation signal (S12).
[0041] The predetermined pointer display condition in this control example 2 is a condition that a predetermined pointer display signal is input (accepted) from the mouse 104 or another input device such as the keyboard 103. That is, when the pointer display signal is input at any timing (Yes in S13), control is performed to display the pointer at a position on the screen of the display 102 that is in the line of sight of the user (line of sight position) (S14 to S16). The control performed in processing steps S14 to S16 is the same as in the above-mentioned control example 1.
[0042] An example of the pointer display signal is a key operation signal generated when a predetermined key operation (a key operation set to display the pointer at the line of sight) is performed on the keyboard 103. According to this, if the user loses sight of the pointer, the user can immediately find it by performing a predetermined key operation on the keyboard 103, which causes the pointer to be displayed immediately in the user's line of sight.
[0043] Furthermore, the pointer display signal may be, for example, a position operation signal from the mouse 104 when the mouse 104 is operated to repeatedly move the pointer position back and forth (for example, moving the pointer on the screen back and forth left and right). When losing sight of the pointer position, many users perform mouse operations to repeatedly move the pointer position back and forth (moving it left and right) so that the pointer is easier to recognize. Therefore, according to this example, when a mouse operation is performed, which many users perform when losing sight of the pointer position, the pointer is displayed in the user's line of sight, allowing the user to quickly find the pointer. In this case, the user does not need to remember any special operations to display the pointer at the line of sight.
[0044] Next, the gaze position estimation process in the gaze position estimation unit 101c of this embodiment will be described. The gaze position estimation unit 101c of this embodiment captures a user's facial image using a webcam 105 consisting of an inexpensive monocular camera, and estimates the user's gaze position on the screen from the captured image data. There are limitations to human programming in obtaining a computer program that can estimate the user's gaze position with sufficient accuracy from such captured image data. In particular, the captured image data of a user's facial image captured using an inexpensive webcam 105 contains little information about depth, making it difficult to estimate the user's gaze position with sufficient accuracy from such captured image data.
[0045] Therefore, in this embodiment, a trained model created by machine learning or the like is used as a program for estimating the user's gaze position. Specifically, the trained model is trained using a plurality of training data including image data of a face image A (see FIG. 6(b)) of a trainee (who may be the user or a person different from the user) at the timing when the trainee views a target mark M displayed on the screen as shown in FIG. 6(a) and position data of the target mark M on the screen at the timing when the target mark M is viewed (coordinates (x=2436, y=1032) in the example of FIG. 6).
[0046] In the stage of creating a trained model (learning phase) of this embodiment, a facial image A of the trainee is captured at each viewing timing when the trainee views a target mark M displayed at various positions on the screen. Then, imaging data of the facial image A of the trainee when viewing the target mark M at each position (viewing timing) is obtained, and position data of the target mark M at each position is also obtained as information indicating the trainee's gaze position. Then, machine learning or the like is performed using a training dataset that is a collection of multiple (large) training data including the imaging data of each facial image and the position data of the target mark M, to obtain a trained model. The trained model obtained in this way is an estimation program that, when estimating a user's gaze position (inference phase), uses imaging data of the user's facial image captured by the webcam 105 as input data and outputs the gaze position of the user to be estimated.
[0047] FIG. 7 is an explanatory diagram of the learning phase of the trained model 200 executed by the CPU 11 of the device main body 101 in the personal computer 100. The trained model 200 of this embodiment can be created by supervised learning (machine learning) using, for example, an external PC 202 or a cloud service that can generate the trained model 200. The created trained model 200 is a type of calculation algorithm, and is modularized and implemented as part of a control program (estimation program) executed by the CPU 11.
[0048] In one example of supervised learning, the learner operates a mouse so that the pointer overlaps target marks displayed at various positions on the screen of a learning personal computer (which may be the personal computer 100 of this embodiment). At this time, since the learner is usually visually recognizing the position of the target mark when the pointer overlaps the target mark, this timing is used as the visual recognition timing to capture an image of the learner's face. Then, image data of the learner's face image at this timing is acquired, and position data of the target mark at this timing is acquired as information indicating the learner's gaze position. In this way, learning data 201 can be obtained, which is a set of teacher data in which the position data of the target mark (the learner's gaze position) is used as correct answer data, along with the image data of the learner's face image.
[0049] In this embodiment, the trained model 200 is created by machine learning using supervised learning, but the trained model 200 may also be created by adopting other machine learning methods such as unsupervised learning or reinforcement learning.
[0050] Here, by changing the direction of the user's eyes, the user can direct their gaze in a direction different from the direction of their face (facial direction), i.e., the front direction of their face. Therefore, both the facial direction (facial direction) and the eye direction (eye direction) are important features in estimating the user's gaze position on the screen.
[0051] Therefore, in the learning phase of this embodiment, imaging data of the left and right eye images BL and BR of the learning subject, as shown in FIGS. 8(b) and 8(c), are extracted from imaging data of the learning subject's face image A, as shown in FIG. 8(a). The imaging data of the left eye image BL and the right eye image BR are then included in the learning data, along with the imaging data of the face image A from which the data was extracted. By including the imaging data of the eye images BL and BR in the learning data separately from the imaging data of the face image A in this way, a trained model can be obtained that outputs estimation results in which the influence of the eyeball direction (eyeball orientation) is emphasized. This makes it possible to estimate the user's gaze position with higher accuracy than when a trained model is obtained from training data that does not include imaging data of the eye images BL and BR (in which the eyeball direction has a smaller impact on the estimation results).
[0052] Furthermore, in this embodiment, the image data of the eye images BL and BR included in the training data is extracted from the image data of the face image A included in the training data, so a trained model can be obtained more cheaply than if the image data of the eye images BL and BR were obtained by a separate imaging operation from the image data of the face image A.
[0053] Furthermore, in the learning phase of this embodiment, feature data generated from the imaging data included in the learning data (imaging data of face image A and imaging data of eye images BL and BR) is also included in the learning data. This makes it possible to obtain a trained model with improved estimation accuracy through feature engineering, compared to a trained model obtained from training data that does not include such feature data (training data that includes only imaging data).
[0054] In particular, when facial image A of a training subject is captured using an inexpensive imaging device such as a monocular camera used in webcam 105, the imaging data contains little information about depth. Therefore, it is difficult for a trained model generated from training data including only imaging data obtained with such an inexpensive imaging device to estimate the user's gaze position with sufficient accuracy. By including feature data in the training data as in this embodiment, even when imaging data obtained with an inexpensive imaging device is used as training data, the lack of accuracy can be compensated for by the feature data, and a trained model that can estimate the user's gaze position with sufficient accuracy can be obtained.
[0055] Here, as the feature data, for example, face direction feature data for identifying the face direction of the learning subject, eye direction feature data for identifying the eye direction of the learning subject, etc. can be suitably used.
[0056] The face direction feature amount data includes, for example, position data (coordinate data) of facial feature points, and yaw angle (angle in the horizontal direction of the face), pitch angle (angle in the vertical direction of the face), and roll angle (angle around an axis extending in the forward direction of the face) calculated from the position data of the facial feature points. As an example, position data of facial feature points such as eye positions (e.g., inner corners of the left and right eyes), nose position (e.g., nostrils), mouth position (e.g., positions of the corners of the mouth), and facial contour position (e.g., tip of the chin) are extracted from imaging data of a face image such as that shown in FIG. 9(a), as shown in FIG. 9(b). The yaw angle, pitch angle, and roll angle can be calculated from these position data using a predetermined arithmetic expression.
[0057] The eyeball direction feature amount data can be, for example, a feature amount parameter that characterizes the eyeball direction, calculated from position data (coordinate data) of feature points of the eyeball. As an example of this feature amount parameter, for example, from the captured data of left and right eye images as shown in Fig. 10(a), the ratios Ll2 / Ll1 and Lr2 / Lr1 of the distances Ll2 and Lr2 from the inner canthus of each eye to the distances Ll1 and Lr1 from the inner canthus of each eye to the pupil center can be used, as shown in Fig. 10(b).
[0058] FIG. 11 is an explanatory diagram illustrating an example of a machine learning model for creating a trained model in this embodiment. In this example, the machine learning model inputs image data for a face image, a right-eye image, and a left-eye image into an image recognition model configured with separate convolutional neural networks (CNNs) to generate respective feature maps. Each feature map is fully connected via a concatenated layer, a flattened layer, and a dense layer to form a single feature map. Numerical feature inputs, such as face direction feature data (yaw angle, pitch angle, roll angle, left eye (inner canthus of the left eye) position coordinates (xL, yL), right eye (inner canthus of the right eye) position coordinates (xR, yR), and eye direction feature data (the ratio Ll2 / Ll1 for the left eye and the ratio Lr2 / Lr1 for the right eye), are then combined to generate a model in which full connection is performed and the output is two-dimensional (x, y) using a linear activation function.
[0059] FIG. 12 is an explanatory diagram showing an example of a method for obtaining image data of various facial images of a subject viewing the target mark M and position data of the target mark M. In FIG. In general, to obtain a trained model with high accuracy, it is desirable to collect a large amount of training data with various contents in a short time. To obtain a trained model that estimates the gaze position of a user as in this embodiment, it is desirable to collect a large amount of imaging data of various facial images of the training subject who views the target mark M and position data of the target mark M in a short time.
[0060] In this embodiment, as shown in Fig. 12, the student operates the mouse 104 so that the pointer P overlaps and follows the target mark M that is moving (continuously changing position) on the screen. Then, every time the pointer P overlaps the target mark M (continuously while the pointer P is overlapping the target mark M), captured data of the face image of the student and position data of the target mark are acquired.
[0061] According to this, the learner only needs to operate the mouse 104 so that the pointer P follows the target mark M moving on the screen, and the captured data of the face image and the position data of the target mark can be continuously acquired without performing a designation operation (i.e., an operation of clicking the mouse 104) to designate the target mark M. Therefore, the workload on the user is small, and the captured data can be continuously acquired (there is little blank time when the captured data cannot be acquired), and a large amount of captured data of various face images of the learner and the position data of the target mark can be collected in a short time.
[0062] In particular, in the example of FIG. 12, the target mark M moves on the screen along a movement path indicated by arrows (1) to (5). Specifically, the target mark M1 first moves slowly along the solid arrow (1), and then, as indicated by the dashed arrow (2), it momentarily moves to the position of the target mark M2 indicated by the symbol M2. That is, the target mark M1, which has moved to the tip of the solid arrow (1), disappears, and the target mark M2 appears at the base end position of the solid arrow (3). After that, the target mark M2 moves slowly along the solid arrow (3), and then, as indicated by the dashed arrow (4), it momentarily moves to the position of the target mark M3 indicated by the symbol M3. Then, the target mark M3 moves slowly along the solid arrow (5).
[0063] As shown by the solid arrows (1), (3), and (5), by slowly moving the target mark M, smooth pursuit can be induced in the eyes of the learner viewing the target mark M. This improves the accuracy of the trained model's estimation of the gaze position when the user moves their gaze with smooth pursuit.
[0064] In addition, as indicated by the dashed arrows (2) and (4), by instantaneously moving the target mark M to a different position on the screen, it is possible to induce a saccade in the eyes of the learner who is viewing the target mark M. This improves the accuracy of the trained model's estimation of the gaze position when the user moves their gaze with a saccade.
[0065] Alternatively, for example, the user may operate mouse 104 to designate an arbitrary position on the screen of display 102 (for example, the position of an icon on the screen) with pointer P, and the web camera 105 may capture an image of the user's face at the timing of the pointer designation, and the captured data of the face image and the position data of the pointer may be included in the learning data. In this way, learning data can be collected while the user is performing some task on personal computer 100 while looking at the screen of display 102 with mouse 104, making it easier to collect learning data.
[0066] Furthermore, the captured image data of the user's face image included in this training data is that of the user himself / herself, who is the subject of gaze position estimation (not a learning subject other than the user), and is captured in the user's usage environment (background, brightness, etc.). Therefore, the inclusion of such training data can improve the accuracy of the trained model in estimating the gaze position of the user.
[0067] It should be noted that many of the partial techniques used in the method for estimating the user's gaze position described in this embodiment can also be applied to estimating the user's gaze position for purposes other than this embodiment.
[0068] For example, a method of including at least one of facial direction feature data and eye direction feature data in addition to image data of the face image of the subject and position data of the target mark in the learning data used to generate a trained model that estimates the user's gaze position is useful not only for display control applications that display a pointer at the estimated gaze position on the screen when a specified pointer display condition is met, but also for other applications such as gaze input.
[0069] Furthermore, for example, a method of acquiring image data of the face image of the person being trained and position data of the target mark at the time when the pointer overlaps with a target mark moving on the screen as training data used to generate a trained model that estimates the user's gaze position is useful not only for display control applications that display a pointer at the estimated gaze position on the screen when specified pointer display conditions are met, but also for other applications such as gaze input.
[0070] Furthermore, for example, a method of using image data of a user's face captured by webcam 105 at the timing when the user operates mouse 104 to designate any position on the screen of display 102 (for example, the position of an icon on the screen) with pointer P, and position data of the pointer, as learning data used to generate a trained model that estimates the user's gaze position, is useful not only for display control applications that display a pointer at the estimated gaze position on the screen when specified pointer display conditions are met, but also for other applications such as gaze input. [Explanation of symbols]
[0071] 100: PC 101: Device main body 101a: control unit 101b: Storage section 101c: Gaze position estimation unit 102: Display 103: Keyboard 104: Mouse 105: Webcam A:Facial image BL: Left eye image BR: Right eye image M: Target mark P: pointer
Claims
1. A control device that controls the position of a pointer displayed on a screen based on a position operation signal from a pointing device, a gaze position estimation means for estimating a gaze position of a user on the screen, When a predetermined pointer display condition is satisfied, display control is performed to display the pointer at the gaze position on the screen estimated by the gaze position estimation means, and the display control is not performed until the next time the predetermined pointer display condition is satisfied.
2. 2. The control device according to claim 1, A control device characterized in that the specified pointer display condition includes a condition that a position operation signal for changing the position of the pointer is received from the pointing device after a specified time has elapsed since the position of the pointer was last changed.
3. 3. The control device according to claim 1 or 2, The control device according to claim 1, wherein the predetermined pointer display condition includes a condition that a predetermined pointer display signal is received from the pointing device or another input device.
4. 4. The control device according to claim 3, The control device, wherein the pointer display signal is a position operation signal from the pointing device that repeatedly moves the position of the pointer back and forth on the screen.
5. 2. The control device according to claim 1, The gaze position estimation means is a control device characterized in that it includes: an acquisition unit that acquires imaging data of a user's facial image; and an estimation unit that estimates the gaze position of the user on the screen based on the imaging data acquired by the acquisition unit by having a computer execute a trained model trained using a plurality of learning data including imaging data of the facial image of the learning subject at a viewing timing when the learning subject views a target mark displayed on the screen and position data of the target mark on the screen at the viewing timing.
6. 6. The control device according to claim 5, A control device characterized in that the multiple learning data also include image data of an eye image including at least one eye of the learning subject extracted from image data of the learning subject's face image at the viewing timing.
7. 7. The control device according to claim 5 or 6, The control device, wherein the plurality of learning data also includes feature amount data generated from the imaging data.
8. 8. The control device according to claim 7, The feature data includes at least one of face direction feature data for specifying the face direction of the learning subject, and eye direction feature data for specifying the eye direction of the learning subject.
9. 7. The control device according to claim 5 or 6, A control device characterized in that the visual recognition timing includes the timing when a pointer, whose position on the screen is controlled based on a position operation signal of a pointing device operated by the learner, overlaps with a target mark moving on the screen.
10. 7. The control device according to claim 5 or 6, A control device characterized in that the visual recognition timing includes the timing when a pointer, whose position on the screen is controlled based on a position operation signal of a pointing device operated by the learner, overlaps with each target mark appearing at a different position on the screen.
11. 7. The control device according to claim 5 or 6, The control device is characterized in that the plurality of learning data include image data of the user's face image acquired by the acquisition unit at a pointer designation timing when the user operates the pointing device to designate any position on the screen with the pointer, and position data of the pointer on the screen at the pointer designation timing.
12. An estimation device that estimates a user's gaze position on a screen, an acquisition unit that acquires captured data of a face image of a user; an estimation unit that estimates the gaze position of the user based on the imaging data acquired by the acquisition unit by causing a computer to execute a trained model that has been trained using a plurality of training data including imaging data of a face image of the learning subject at a viewing timing when the learning subject views a target mark displayed on a screen, feature amount data generated from the imaging data, and position data of the target mark on the screen at the viewing timing, the feature data includes at least one of face direction feature data for specifying the face direction of the learning subject, and eye direction feature data for specifying the eye direction of the learning subject.
13. An estimation device that estimates a user's gaze position on a screen, an acquisition unit that acquires captured data of a face image of a user; an estimation unit that estimates the gaze position of the user on the screen based on the imaging data acquired by the acquisition unit by causing a computer to execute a trained model that has been trained using a plurality of training data including imaging data of a face image of the trainee at a viewing timing when the trainee views a target mark displayed on the screen and position data of the target mark on the screen at the viewing timing, The estimation device is characterized in that the visual recognition timing includes the timing when a pointer, whose position on the screen is controlled based on a position operation signal of a pointing device operated by the learner, overlaps with a target mark moving on the screen.
14. An estimation device that estimates a user's gaze position on a screen, an acquisition unit that acquires captured data of a face image of a user; and an estimation unit that estimates the gaze position of the user on the screen based on the imaging data acquired by the acquisition unit by causing a computer to execute a trained model trained using a plurality of training data including imaging data of the user's face image acquired by the acquisition unit at a pointer designation timing when the user operates a pointing device to designate an arbitrary position on the screen with a pointer and position data of the pointer on the screen at the pointer designation timing.
Citation Information
Patent Citations
Electronic apparatus, control method thereof, program, and recording medium
JP2022068749A