Information processing device, information processing method, and program

The system collects data for machine learning during gaze input and switches to camera-based estimation, addressing the downtime issue in existing technologies, ensuring efficient and accurate gaze input without operational disruption.

JP2026061045APending Publication Date: 2026-04-09OKI ELECTRIC INDUSTRY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing hands-free input technologies using eye gaze require a significant amount of downtime and man-hours for data collection, which is impractical for environments like factories.

Method used

An information processing system that collects data for machine learning in parallel with gaze input processing, utilizing a gaze sensor to generate a gaze estimation model, and then switches to using a visible light camera for continuous operation without interrupting work.

Benefits of technology

Enables efficient data collection for machine learning without disrupting operations, allowing cost-effective implementation of gaze input across multiple devices with maintained accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026061045000001_ABST
    Figure 2026061045000001_ABST
Patent Text Reader

Abstract

This enables the collection of data for machine learning in parallel with eye-tracking input processing. [Solution] An information processing device comprising: an acquisition unit that acquires gaze information from a first sensor and an image captured from a second sensor; a gaze input processing unit that processes gaze input based on user gaze position information acquired from the gaze information; a processing unit that stores the user's gaze position information acquired from the gaze information and the user's face image obtained from the image captured in a storage unit, in association with a user ID; and a machine learning unit that performs machine learning to generate a gaze estimation model for estimating gaze position from the face image using the data stored in the storage unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, a hands-free input technology using eye gaze has been known.

[0003] In Patent Document 1 below, a gaze position estimation system that enables gaze estimation without calibration using a visible light camera and a gaze position estimation model generated by machine learning is disclosed.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the system disclosed in Patent Document 1, a preparation process for collecting data such as training eye images used for machine learning is required, and when introducing hands-free input for workers in a factory or the like, there is a problem that the work must be stopped once and a large number of man-hours must be allocated for data collection.

[0006] Therefore, an object of the present invention is to provide a novel and improved information processing apparatus, information processing method, and program capable of collecting data for machine learning in parallel with gaze input processing.

Means for Solving the Problems

[0007] To solve the above problems, according to one aspect of the present invention, an information processing device is provided, comprising: an acquisition unit that acquires gaze information from a first sensor and an image captured from a second sensor; a gaze input processing unit that processes gaze input based on user gaze position information acquired from the gaze information; a storage processing unit that stores the user's gaze position information acquired from the gaze information and the user's face image obtained from the image captured in association with a user ID in a storage unit; and a machine learning unit that performs machine learning to generate a gaze estimation model for estimating gaze position from the face image using the data stored in the storage unit.

[0008] The machine learning unit may further determine the estimation accuracy of the generated gaze estimation model.

[0009] The information processing device may further include a switching unit that switches the input information to the gaze input processing unit from gaze position information based on the gaze information to estimated gaze position information obtained using the gaze estimation model based on the face image.

[0010] The switching unit may perform the switch if it determines that the estimation accuracy of the gaze estimation model corresponding to the user is sufficient.

[0011] The information processing device may further include a notification control unit that performs control to notify the administrator terminal that it has been determined that the estimation accuracy of the gaze estimation model is sufficient.

[0012] The switching unit may perform the switching in accordance with instructions from the administrator terminal.

[0013] The information processing device may further include a similarity determination unit that determines the similarity of the user based on the user's facial image and classifies the user into a similarity cluster.

[0014] The storage processing unit may store information about the user's gaze position and the user's face image in a database of similarity clusters into which the user is classified, and the machine learning unit may perform machine learning to generate a gaze estimation model for estimating the gaze position from the face image using the data stored in the database.

[0015] Furthermore, in order to solve the above problems, according to another aspect of the present invention, an information processing method is provided in which a processor acquires gaze information from a first sensor and an image captured from a second sensor; processes gaze input based on user gaze position information acquired from the gaze information; stores the user gaze position information acquired from the gaze information and the user's face image obtained from the image captured in a storage unit in association with a user ID; and performs machine learning to generate a gaze estimation model for estimating the gaze position from the face image using the data stored in the storage unit.

[0016] Furthermore, in order to solve the above problems, according to another aspect of the present invention, a program is provided to cause a computer to function as: an acquisition unit that acquires gaze information from a first sensor and an image captured from a second sensor; a gaze input processing unit that processes gaze input based on user gaze position information acquired from the gaze information; a processing unit that stores the user's gaze position information acquired from the gaze information and the user's face image obtained from the image captured in a storage unit, in association with a user ID; and a machine learning unit that performs machine learning to generate a gaze estimation model for estimating gaze position from the face image using the data stored in the storage unit. [Effects of the Invention]

[0017] As described above, the present invention makes it possible to collect data for machine learning in parallel with eye-tracking input processing. [Brief explanation of the drawing]

[0018] [Figure 1]It is an explanatory diagram showing the configuration of the line-of-sight input system 1 according to an embodiment of the present invention. [Figure 2] It is a block diagram showing a configuration example of the server 20 according to an embodiment of the present invention. [Figure 3] It is a diagram showing an example of data stored in the user ID-specific DB 231 according to the present embodiment. [Figure 4] It is a flowchart showing an example of the flow of operation processing of the line-of-sight input system 1 according to the present embodiment. [Figure 5] It is a flowchart showing an example of the flow of operation processing when shifting to line-of-sight input using the estimated line-of-sight position based on camera data according to the present embodiment. [Figure 6] It is a block diagram showing a configuration example of the server 20x according to a modification example of the present embodiment. [Figure 7] It is a diagram showing an example of data stored in the similarity-specific DB 241 according to the modification example. [Figure 8] It is a flowchart showing an example of the flow of operation processing of a modification example of the line-of-sight input system 1 according to the present embodiment. [Figure 9] It is a diagram showing an example of the data configuration of each DB stored in the storage unit 240 according to the present modification example. [Figure 10] It is a flowchart showing an example of the flow of operation processing when shifting to line-of-sight input using the estimated line-of-sight position based on camera data according to the present modification example. [Figure 11] It is a diagram showing the hardware configuration of the information processing apparatus 900 as an example of the server 20 according to an embodiment of the present invention.

Embodiments for Carrying Out the Invention

[0019] [[ID=3S]]Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant description is omitted.

[0020] <1. Outline of the Line-of-Sight Input System> Embodiments of the present invention relate to an eye-tracking system (an example of an information processing system) that enables the collection of data for machine learning in parallel with eye-tracking processing.

[0021] Figure 1 is an explanatory diagram showing the configuration of the eye-tracking input system 1 according to an embodiment of the present invention. Here, we will explain the case in which the eye-tracking input system 1 according to this embodiment is applied to eye-tracking input by workers performing tasks in a factory or the like.

[0022] The eye-tracking input system 1 according to this embodiment includes a worker terminal 11, a camera 12, an eye-tracking sensor 13, and an ID acquisition unit 14, all located in the work environment 10, as well as a server 20 and an administrator terminal 30. Each device can communicate and transmit / receive data via a network 40.

[0023] The worker terminal 11 is an example of a display device having a display unit that displays various screens, and is fixed to a workbench 15, for example. Worker W performs work by looking at work instructions, etc., displayed on the worker terminal 11. At this time, it is assumed that worker W's hands are occupied with the work. Therefore, hands-free eye-tracking input is preferred for operating the worker terminal 11, such as page scrolling, and the eye-tracking input system 1 according to this embodiment is applied.

[0024] Camera 12 captures images of worker W and transmits the captured images to server 20. Camera 12 can be implemented using a visible light camera. For example, camera 12 is installed on or around the workbench 15 and captures images of worker W's face.

[0025] The gaze sensor 13 is an example of a sensor capable of detecting the gaze information of worker W. The gaze sensor 13 may be implemented, for example, by an infrared stereo camera. In the example shown in Figure 1, a bar-shaped gaze sensor 13 is detachably attached to the bottom of the worker terminal 11. The gaze sensor 13 transmits the detected sensing data to the server 20 as the gaze information of worker W. Note that the shape and installation position of the gaze sensor 13 are not limited to the example shown in Figure 1.

[0026] The ID acquisition unit 14 is an example of a reader that reads a user ID (identification information) from an ID card or the like owned by worker W. The ID acquisition unit 14 transmits the read user ID to the server 20. The ID acquisition unit 14 may also acquire biometric data as the user ID and transmit it to the server 20. Alternatively, the ID acquisition unit 14 may be installed in the worker terminal 11. In this case, the ID acquisition unit 14 may transmit, for example, a login ID (e.g., employee number, email address, etc.) entered in the worker terminal 11 as the user ID to the server 20.

[0027] Server 20 identifies worker W based on the user ID transmitted from ID acquisition unit 14. Server 20 also calculates the gaze position based on gaze information transmitted from gaze sensor 13 and transmits it to worker terminal 11 as gaze input information (coordinates of gaze position). The gaze position is acquired as a two-dimensional coordinate position on the display unit of worker terminal 11. In this embodiment, it is assumed that the gaze sensor 13 has been calibrated. The worker terminal 11 accepts hands-free input operations using gaze input on the screen displayed on the worker terminal 11, based on the gaze input information received from server 20. Server 20 may be an information processing device such as a PC located in the work environment 10. Furthermore, server 20 may be installed on a network owned and operated within the company, or it may be installed on the cloud.

[0028] The administrator terminal 30 is used by the administrator and can manage the entire work site, including the work progress of worker W. The administrator terminal 30 can also receive notifications from the server 20 as needed and present them to the administrator. The administrator can take various actions as necessary.

[0029] Here, gaze sensors used to detect gaze information are expensive, specialized sensors that are difficult to obtain and implement in multiple units, resulting in poor cost performance. On the other hand, visible light cameras are relatively inexpensive and easy to obtain and implement in multiple units, so it is desirable to be able to implement gaze input using visible light cameras. However, when performing gaze estimation using an estimation model obtained by machine learning with information acquired by a visible light camera, a preparatory process such as collecting data for machine learning is required beforehand, which may affect operations in factories, etc.

[0030] Therefore, the inventor of this invention, focusing on the above circumstances, has created an eye-tracking input system according to an embodiment of the present invention. The embodiment of the present invention makes it possible to collect data for machine learning in parallel with eye-tracking input processing.

[0031] Specifically, by performing gaze input using the gaze sensor 13 while collecting data necessary for machine learning (generating a dataset), and once a gaze estimation model is generated through machine learning, switching from gaze input using the gaze sensor 13 to gaze input based on the gaze estimation results using the camera 12, it is possible to generate a dataset and perform machine learning without interrupting operations in factories, etc.

[0032] Furthermore, when the server 20 notifies the administrator terminal 30 that an estimation model has been generated, the administrator may remove the gaze sensor 13 and attach it to another worker's terminal to generate gaze estimation models for other workers. In this way, the expensive dedicated gaze sensor 13 can be effectively utilized, and even with a small number of units, it is possible to handle situations requiring a large number of devices without reducing the accuracy of gaze input.

[0033] The configuration and operation of the eye-tracking input system according to this embodiment of the present invention will be described in detail below.

[0034] <2. Example configuration of Server 20> Figure 2 is a block diagram showing an example configuration of a server 20 according to an embodiment of the present invention. As shown in Figure 2, the server 20 includes a communication unit 210, a control unit 220, and a storage unit 230.

[0035] (Communications Section 210) The communication unit 210 is comprised of a communication interface and communicates with the worker terminal 11 and the administrator terminal 30 via the network 40.

[0036] (Control unit 220) The control unit 220 may include a CPU (Central Processing Unit) and its functions may be realized when the program stored in the storage unit 230 is loaded into RAM (Random Access Memory) by the CPU and executed. In this case, a computer-readable recording medium on which the program is recorded may also be provided. Alternatively, the control unit 220 may be composed of dedicated hardware or a combination of multiple hardware components.

[0037] As shown in Figure 2, the control unit 220 includes a gaze data acquisition unit 221, a camera data acquisition unit 222, a data storage processing unit 223, a machine learning unit 224, an administrator notification control unit 225, a switching unit 226, a gaze estimation unit 227, and a gaze input processing unit 228.

[0038] The gaze data acquisition unit 221 acquires gaze information from the gaze sensor 13 and further performs processing to calculate the gaze direction and gaze position of the worker W based on the gaze information (sensing data). The gaze position is calculated as a two-dimensional coordinate position P(xP, yP) on the display screen of the worker terminal 11. For example, if the update time unit for the gaze position is Δt, the two-dimensional position calculated at each time is expressed as Pn(xn, yn) as the coordinate position after Δt × n time has elapsed. In other words, if measurement is performed up to Δt × n time, the gaze data acquisition unit 221 can obtain an array of coordinate positions such as P1(x1, y1), P2(x2, y2), ...Pt(xt, yt) as the gaze position. The update time is a sufficient unit of time to calculate the gaze direction and gaze position.

[0039] The gaze position information calculated by the gaze data acquisition unit 221 is output to the data storage processing unit 223 and the gaze input processing unit 228.

[0040] The camera data acquisition unit 222 acquires camera data (captured images) from the camera 12, and further processes the acquired camera data to extract a face image including the eyes of the worker W. The above update time is set to a unit time sufficient to process the extraction of the face image.

[0041] The data storage processing unit 223 stores the gaze position information calculated by the gaze data acquisition unit 221 and the face image obtained by the camera data acquisition unit 222 in the storage unit 230, associating them with the user ID. In the storage unit 230, a database (user ID-specific DB 231) is prepared for each user ID. That is, user-specific (worker-specific) datasets used for machine learning can be generated. Meanwhile, the gaze position information is also output from the gaze data acquisition unit 221 to the gaze input processing unit 228 and used for gaze input processing. In this embodiment, it is possible to collect data for machine learning in parallel with gaze input processing.

[0042] The machine learning unit 224 performs machine learning using user-specific datasets stored in the memory unit 230 to generate user-specific gaze estimation models. As an example, the machine learning unit 224 performs supervised learning using data stored in a database (user ID-specific DB231). Specifically, the machine learning unit 224 is composed of, for example, a neural network and has weight parameters. In this case, learning is performed by inputting face images from the database as input data, calculating the error between the obtained output data, which is the two-dimensional coordinate position Q(xQ,yQ), and the ground truth data, which is the viewpoint position (two-dimensional coordinate position P) in the database, and updating the weight parameters of the neural network by backpropagation. Face images are input each time data is stored in the database (for example, every update time unit Δt), and learning is performed. However, the timing of learning is not limited to this.

[0043] The generated gaze estimation model is stored in the memory unit 230 in association with the user ID.

[0044] Furthermore, the machine learning unit 224 determines (verifies) whether the estimation accuracy of the gaze estimation model generated by machine learning has reached a level sufficient for operation. If it has reached a sufficient level, it outputs a message to the administrator notification control unit 225. An estimation accuracy level sufficient for operation means an accuracy level that allows gaze input to be performed without problems. For example, if the method of gaze input involves staring at a specific UI (input area) for a certain period of time, an accuracy level that can correctly detect the dwell of gaze within the UI is expected.

[0045] The administrator notification control unit 225 controls the notification of the estimation accuracy determination result of the gaze estimation model output from the machine learning unit 224 to the administrator terminal 30 via the network 40 from the communication unit 210.

[0046] The switching unit 226 switches the input information to the gaze input processing unit 228. The switching unit 226 disconnects the connection between the gaze data acquisition unit 221 and the gaze input processing unit 228, and connects the gaze estimation unit 227 and the gaze input processing unit 228, thereby switching the source of information input to the gaze input processing unit 228 from the gaze data acquisition unit 221 to the gaze estimation unit 227.

[0047] More specifically, the switching unit 226 instructs the gaze data acquisition unit 221 to stop outputting information (outputting gaze position information) to the gaze input processing unit 228, and instructs the gaze estimation unit 227 to perform gaze estimation and output the estimated gaze position information (estimation result) to the gaze input processing unit 228. Accordingly, the switching unit 226 instructs the camera data acquisition unit 222 to output the face image extracted from the camera data to the gaze estimation unit 227.

[0048] The switching unit 226 may switch input information in accordance with instructions from the administrator terminal 30. That is, when the administrator notification control unit 225 notifies the administrator terminal 30 that the estimation accuracy of the gaze estimation model has reached a level sufficient for operation, the administrator is expected to instruct the server 20 to switch to gaze input using the camera 12, which is a visible light camera. The switching unit 226 can perform the switch in accordance with such instructions. In addition, the administrator may, at the same time as instructing the switch, remove the gaze sensor 13 from the worker terminal 11 and attach it to another worker terminal so that gaze estimation models for other workers can be generated.

[0049] Furthermore, the switching unit 226 is not limited to instructions from the administrator, but may also be configured to automatically switch input information when the machine learning unit 224 determines that the estimation accuracy of the gaze estimation model has reached a level sufficient for operation. In this case, the switching unit 226 may notify the administrator that the switch has been made.

[0050] The gaze estimation unit 227 uses the face image output from the camera data acquisition unit 222 and the gaze estimation model generated by machine learning to estimate the gaze position, and outputs the estimation result (information on the gaze position; specifically, the two-dimensional coordinate position) to the gaze input processing unit 228.

[0051] The eye-tracking input processing unit 228 performs eye-tracking input processing based on eye-tracking position information (two-dimensional coordinate position). Specifically, the eye-tracking input processing unit 228 executes input according to the eye-tracking position and controls the display unit of the worker terminal 11 to display a feedback screen. It is assumed that information about the display screen is obtained from the worker terminal 11. Input could be, for example, if a work instruction sheet is displayed on the display unit, operations such as advancing or going back to the next page. As for specific methods of eye-tracking input, for example, an input area may be displayed on the display unit, and if the gaze remains in the input area for a certain period of time or longer, an input corresponding to the input area (advancing or going back to the next page, etc.) may be executed. Also, if the gaze moves to the right or left within the input area, input such as advancing or going back to the next page may be executed according to the direction of movement. Conventional eye-tracking input UIs and methods may be used for such specific methods of eye-tracking input.

[0052] Furthermore, some processing of the eye-tracking input processing unit 228 may be performed on the worker terminal 11. Specifically, the eye-tracking input processing unit 228 may transmit eye-tracking position information (two-dimensional coordinate position) to the worker terminal 11, and the worker terminal 11 may perform input execution and control the display of the feedback screen.

[0053] (Storage unit 230) The memory unit 230 is a storage device capable of storing programs and data for operating the control unit 220. The memory unit 230 can also temporarily store various data required during the operation of the control unit 220. For example, the storage device may be a non-volatile storage device.

[0054] The storage unit 230 according to this embodiment has a user ID-specific DB 231. The user ID-specific DB 231 is a database generated for each user ID. The user ID-specific DB 231 may store data used for machine learning and generated estimation models. For example, the user ID-specific DB 231 stores facial images, gaze position (position coordinates) information, and neural network weight parameters.

[0055] Figure 3 shows an example of data stored in DB231 by user ID according to this embodiment. The example shown in Figure 3 is an example of data stored in association with user ID 00001.

[0056] As shown in Figure 3, DB231 for each user ID stores a timestamp indicating the time of data acquisition, a facial image, and information on the gaze position (x, y coordinates).

[0057] The configuration of the server 20 according to an embodiment of the present invention has been described above. Note that the configuration shown in Figure 2 is just one example, and the configuration of the server 20 is not limited thereto. For example, each component of the server 20 may be implemented by a system consisting of multiple devices. Also, at least a part of the configuration of the server 20 may be implemented by the worker terminal 11 or the administrator terminal 30.

[0058] <3. Operation Processing> Figure 4 is a flowchart showing an example of the operation process flow of the eye-tracking input system 1 according to this embodiment.

[0059] As shown in Figure 4, first, the server 20 recognizes the user ID transmitted from the ID acquisition unit 14 (step S103). This allows the server 20 to determine which worker is using the service.

[0060] Next, the server 20 acquires gaze data (sensing data) transmitted from the gaze sensor 13 using the gaze data acquisition unit 221 (step S106) and calculates the gaze position (step S109).

[0061] Furthermore, the server 20 acquires camera data (captured images) transmitted from the camera 12 using the camera data acquisition unit 222 (step S112). In addition, the camera data acquisition unit 222 extracts the worker's face image from the camera data.

[0062] Next, the server 20, using the data storage processing unit 223, stores the gaze position information and facial image in the user ID-specific DB 231 corresponding to the user ID of worker W (step S115). This can generate a machine learning dataset for worker W.

[0063] Meanwhile, the server 20 performs gaze input based on gaze position information using the gaze input processing unit 228 (step S118).

[0064] Next, the server 20, using the machine learning unit 224, performs supervised learning using the data (dataset) stored in the user ID-specific DB 231 (step S121) to generate a gaze estimation model capable of estimating gaze using face images (step S124).

[0065] Next, the machine learning unit 224 determines whether the estimation accuracy of the gaze estimation model is sufficient. Specifically, the machine learning unit 224 inputs the face images stored in the user ID DB 231 into the gaze estimation model to estimate the gaze position (step S127), and evaluates the estimation accuracy of the gaze estimation model from the estimation results. Specifically, the machine learning unit 224 determines whether the estimation accuracy exceeds a threshold (step S130). For example, the evaluation of estimation accuracy may involve assigning an evaluation value corresponding to the degree of error between the gaze position Q estimated using the gaze estimation model and the correct answer (gaze position P).

[0066] If the estimation accuracy exceeds the threshold (step S130 / Yes), the server 20 notifies the administrator (specifically, the administrator terminal 30) via the administrator notification control unit 225 that the estimation accuracy of the generated gaze estimation model is sufficient (step S132).

[0067] If the estimation accuracy falls below the threshold (step S130 / No), the process returns to step S121, and machine learning is continued using the newly stored user ID-specific DB231.

[0068] Next, the server 20, using the switching unit 226, controls the connection to the gaze input processing unit 228 in accordance with instructions from the administrator (instruction to switch to gaze estimation based on captured images) (step S136). Specifically, the switching unit 226 stops the output of gaze position information from the gaze data acquisition unit 221 to the gaze input processing unit 228 and starts outputting the gaze position estimation results from the gaze estimation unit 227 to the gaze input processing unit 228.

[0069] Then, server 20 transitions to gaze input using the estimated gaze position based on camera data (step S139).

[0070] Figure 5 is a flowchart showing an example of the operation process flow when transitioning to gaze input using estimated gaze position based on camera data according to this embodiment.

[0071] As shown in Figure 5, first, the server 20 acquires camera data from the camera 12 (step S153). Specifically, the camera data acquisition unit 222 acquires the camera data, and then extracts a face image from the camera data. The extracted face image is output to the gaze estimation unit 227.

[0072] Next, the gaze estimation unit 227 inputs the extracted face image into the gaze estimation model corresponding to the user ID and estimates the gaze position (step S156). The estimation result (estimated gaze position) is output to the gaze input processing unit 228.

[0073] Next, the eye-tracking input processing unit 228 performs eye-tracking input based on the estimated gaze position (step S159).

[0074] The processes shown in steps S153 to S159 above are repeated until eye-tracking input is terminated (step S162).

[0075] The operation process according to this embodiment has been described above.

[0076] <4. Variation> Next, a modified version of this embodiment will be described. In the embodiment described above, a dataset for machine learning was generated for each user, but the embodiment is not limited to this, and the dataset may be shared and used among multiple users.

[0077] <<4-1. Example Configuration>> Figure 6 is a block diagram showing an example configuration of server 20x according to a modification of this embodiment. Below, we will mainly describe the parts of the server 20x configuration that differ from the server 20 configuration according to the above embodiment described with reference to Figure 2.

[0078] As shown in Figure 6, server 20x has a configuration that includes a similarity determination unit 229 in addition to the configuration of server 20 shown in Figure 2. Furthermore, server 20x has a storage unit 240 in addition to, or instead of, the storage unit 230 shown in Figure 2.

[0079] (Similarity determination unit 229) The similarity determination unit 229 determines the similarity of users (workers) based on the facial images output from the camera data acquisition unit 222. Specifically, the similarity determination unit 229 determines the similarity of users and classifies them into the corresponding similarity cluster. For such classification, an estimation model that has been previously trained by unsupervised learning may be used. For example, if the user group consists of workers in a factory, etc., facial photographs from a worker roster can be used as input data, and classification can be performed based on features such as eye shape, pupil diameter, and pupil color. The estimation model used for classification may be obtained by training with publicly available datasets, or it may be trained by other methods.

[0080] The similarity determination unit 229 outputs the determination result (information indicating which similarity cluster worker W is classified into) to the data storage processing unit 223 and the gaze estimation unit 227. The data storage processing unit 223 stores the gaze position information calculated by the gaze data acquisition unit 221 and the face image extracted by the camera data acquisition unit 222 into a database corresponding to the classified similarity cluster. The gaze estimation unit 227 estimates the gaze position from the face image using a gaze estimation model corresponding to the classified similarity cluster.

[0081] (Storage unit 240) The storage unit 240 has a similarity-based database 241. The similarity-based database 241 is generated for each similarity cluster. Figure 7 shows an example of data stored in the similarity-based database 241 in a modified example. The example shown in Figure 7 is an example of data collected and stored from each user classified into similarity cluster A.

[0082] As shown in Figure 7, the Similarity Database 241 stores a timestamp indicating the time of data acquisition, a user ID, a face image, and information on the gaze position (x, y coordinates). The gaze estimation model generated by the machine learning unit 224 using the data stored in the Similarity Database 241 is also stored in the Similarity Database 241.

[0083] In this modified version, by sharing and using the dataset among multiple users, it is expected that a gaze estimation model with sufficient estimation accuracy can be generated in a shorter time (shorter measurement time). Furthermore, if a gaze estimation model for the similarity cluster to which a new user is classified has already been generated and has sufficient estimation accuracy, gaze input using camera 12 can be implemented from the very beginning of the process without collecting data from that new user.

[0084] <<4-2. Operation Processing>> Figure 8 is a flowchart showing an example of the operation process flow of a modified eye-tracking input system 1 according to this embodiment.

[0085] The processes shown in steps S103 to S139 of the flow in Figure 8 include processes that are common to those explained with reference to Figure 4. The main differences will be explained below.

[0086] In this modified version, server 20x performs similarity determination (specifically, classification into similarity clusters) of workers based on the worker's face image extracted from camera data (step S200). As a result, in step S115, the data is stored in the database of the corresponding similarity cluster. Furthermore, in step S121, machine learning is performed using the data from the database of the corresponding similarity cluster.

[0087] Then, in step S124, gaze estimation models corresponding to similarity clusters are generated. The generated gaze estimation models are stored in the similarity-based DB241.

[0088] In this way, by sharing datasets among multiple similar users (users classified into the same similarity cluster), it is possible to generate gaze estimation models with sufficient estimation accuracy in a shorter measurement time.

[0089] As a result, the storage unit 240 in this modified configuration can store multiple similarity-based databases, as shown in Figure 9. Figure 9 shows an example of the data configuration of each database stored in the storage unit 240 in this modified configuration.

[0090] As shown in Figure 9, the memory unit 240 stores, for example, similarity-based DB-A, similarity-based DB-B, and similarity-based DB-C, each containing information on the dataset, estimation model, and estimation accuracy. When a new worker starts work, the server 20x determines whether the estimation accuracy of the gaze estimation model for the similarity cluster to which the worker is classified is sufficient. If the estimation accuracy is sufficient, the server 20x can switch (transition) to gaze input using the camera 12 without collecting data from the new worker.

[0091] Figure 10 is a flowchart showing an example of the operation process flow when transitioning to gaze input using estimated gaze position based on camera data according to this modification.

[0092] As shown in Figure 10, first, the server 20 acquires camera data of the worker from the camera 12 (step S213).

[0093] Next, the gaze estimation unit 227 uses a gaze estimation model (with sufficient estimation accuracy) of similarity clusters classified according to the worker's face to estimate the gaze position from the face image extracted from the camera data (step S216). The estimation result (estimated gaze position) is output to the gaze input processing unit 228.

[0094] Next, the eye-gaze input processing unit 228 performs eye-gaze input based on the estimated gaze position (step S219).

[0095] The processes shown in steps S213 to S219 above are repeated until eye-tracking input is terminated (step S222).

[0096] The operation process using this modified example has been explained above.

[0097] <5. Hardware Configuration> Next, we will describe an example of a hardware configuration that can be applied to the server 20 according to an embodiment of the present invention. In the following, we will describe an example of a hardware configuration of the information processing device 900 as an example of the server 20's hardware configuration.

[0098] The hardware configuration example of the information processing device 900 described below is merely one example of the hardware configuration of the server 20. Therefore, the hardware configuration of the server 20 may be modified by removing unnecessary components from the hardware configuration of the information processing device 900 described below, or by adding new components.

[0099] Figure 11 shows the hardware configuration of an information processing device 900 as an example of a server 20 according to an embodiment of the present invention. The information processing device 900 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, a RAM (Random Access Memory) 903, a host bus 904, a bridge 905, an external bus 906, an interface 907, an input device 908, an output device 909, a storage device 910, and a communication device 911.

[0100] The CPU 901 functions as both an arithmetic processing unit and a control unit, controlling the overall operation of the information processing unit 900 according to various programs. The CPU 901 may also be a microprocessor. The ROM 902 stores programs and arithmetic parameters used by the CPU 901. The RAM 903 temporarily stores programs used in the execution of the CPU 901 and parameters that change as needed during its execution. These are interconnected by a host bus 904, which consists of a CPU bus and other components.

[0101] The host bus 904 is connected to an external bus 906, such as a PCI (Peripheral Component Interconnect / Interface) bus, via a bridge 905. It is not always necessary to configure the host bus 904, bridge 905, and external bus 906 separately; these functions may be implemented on a single bus.

[0102] The input device 908 consists of input means for the user to input information, such as a mouse, keyboard, touch panel, buttons, microphone, switches, and levers, and an input control circuit that generates input signals based on the user's input and outputs them to the CPU 901. The user operating the information processing device 900 can input various types of data to the information processing device 900 or instruct it to perform processing operations by operating this input device 908.

[0103] The output device 909 includes, for example, display devices such as CRT (Cathode Ray Tube) display devices, liquid crystal display (LCD) devices, OLED (Organic Light Emitting Diode) devices, lamps, and audio output devices such as speakers.

[0104] The storage device 910 is a device for storing data. The storage device 910 may include a storage medium, a recording device for recording data on the storage medium, a reading device for reading data from the storage medium, and a deletion device for deleting data recorded on the storage medium. The storage device 910 is composed of, for example, an HDD (Hard Disk Drive). This storage device 910 drives the hard disk and stores programs executed by the CPU 901 and various data.

[0105] The communication device 911 is a communication interface composed of, for example, a communication device for connecting to a network. The communication device 911 may support either wireless or wired communication.

[0106] <6. Conclusion> Although preferred embodiments of the present invention have been described in detail above with reference to the attached drawings, the present invention is not limited to these examples. It is clear to any person with ordinary skill in the art to which the present invention belongs that various modifications or alterations can be conceived within the scope of the technical idea described in the claims, and these are also understood to fall within the technical scope of the present invention.

[0107] For example, the configuration of the eye-tracking system 1 is not limited to the example shown in Figure 1. For instance, at least one of the camera 12, eye-tracking sensor 13, and ID acquisition unit 14 may be connected to the worker terminal 11, and data may be transmitted to the server 20 via the worker terminal 11.

[0108] Furthermore, the display device in the work environment 10 is not limited to the worker terminal 11, which is implemented using a tablet device or the like; for example, it may be a projector.

[0109] Furthermore, each component of the server 20 may be provided on the worker terminal 11.

[0110] Each step in the operation processing of the eye-tracking system 1 described herein does not necessarily have to be processed chronologically in the order shown in the flowchart. For example, each step in the operation processing of the eye-tracking system 1 may be processed in an order different from that shown in the flowchart, or may be processed in parallel.

[0111] Furthermore, it is possible to create one or more computer programs to enable the server 20 to perform its functions using the hardware such as the CPU, ROM, and RAM built into the server 20. A computer-readable storage medium on which these one or more computer programs are stored is also provided. [Explanation of Symbols]

[0112] 1. Eye-tracking system 10. Work Environment 11. Worker terminal 12 cameras 13. Eye-tracking sensor 14 ID acquisition section 15 Workbench 20 servers 210 Communications Department 220 Control Unit 221 Eye-tracking data acquisition unit 222 Camera data acquisition unit 223 Data Storage Processing Unit 224 Machine Learning Department 225 Administrator Notification Control Unit 226 Switching section 227 Gaze estimation part 228 Eye-tracking input processing unit 229 Similarity determination unit 230, 240 storage section 30 Administrator terminals 40 Networks

Claims

1. An acquisition unit that acquires line-of-sight information from the first sensor and an image captured from the second sensor, A gaze input processing unit that processes gaze input based on the user's gaze position information obtained from the aforementioned gaze information, A storage processing unit stores the user's gaze position information obtained from the gaze information and the user's face image obtained from the captured image in a storage unit, in association with the user ID. A machine learning unit performs machine learning to generate a gaze estimation model for estimating gaze position from the face image using the data stored in the memory unit, An information processing device equipped with the following features.

2. The information processing apparatus according to claim 1, wherein the machine learning unit further determines the estimation accuracy of the generated gaze estimation model.

3. The aforementioned information processing device is The information processing apparatus according to claim 2, further comprising a switching unit that switches the input information to the gaze input processing unit from gaze position information based on the gaze information to estimated gaze position information obtained using the gaze estimation model based on the face image.

4. The information processing apparatus according to claim 3, wherein the switching unit performs a switch when it is determined that the estimation accuracy of the gaze estimation model corresponding to the user is sufficient.

5. The aforementioned information processing device is The information processing apparatus according to claim 3, further comprising a notification control unit that performs control to notify an administrator terminal that it has been determined that the estimation accuracy of the gaze estimation model is sufficient.

6. The switching unit performs the switching in accordance with instructions from the administrator terminal, as described in claim 5.

7. The aforementioned information processing device is The information processing apparatus according to claim 1, further comprising a similarity determination unit that determines the similarity of the user based on the user's facial image and classifies the user into a similarity cluster.

8. The storage processing unit stores the user's gaze position information and the user's facial image in the database of similarity clusters into which the user is classified. The machine learning unit performs machine learning to generate a gaze estimation model for estimating the gaze position from the face image using the data stored in the database. The information processing apparatus according to claim 7.

9. The processor, The process involves acquiring gaze information from the first sensor and an image captured from the second sensor. The process of processing eye-gaze input is performed based on the user's gaze position information obtained from the aforementioned gaze information. The user's gaze position information obtained from the gaze information and the user's face image obtained from the captured image are stored in the storage unit in association with the user ID. Using the data stored in the memory unit, machine learning is performed to generate a gaze estimation model for estimating the gaze position from the face image. Information processing methods, including those mentioned above.

10. Computers, An acquisition unit that acquires line-of-sight information from the first sensor and an image captured from the second sensor, A gaze input processing unit that processes gaze input based on the user's gaze position information obtained from the aforementioned gaze information, A processing unit that stores the user's gaze position information obtained from the gaze information and the user's face image obtained from the captured image in a storage unit, in association with the user ID. A machine learning unit performs machine learning to generate a gaze estimation model for estimating gaze position from the face image using the data stored in the memory unit, A program designed to function as such.

Citation Information

Patent Citations

  • Fixation point estimation system, fixation point estimation method, fixation point estimation program, and information recording medium for recording the same

    JP2020140630A