Eye tracking system, eye tracking method, and recording medium

By dynamically setting the guidance area and calculating the viewpoint coordinate difference in the eye-tracking system, the problem of complex correction process in the prior art is solved, and simple and efficient viewpoint coordinate correction is achieved.

CN115916032BActive Publication Date: 2026-05-08DOWANGO KK
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOWANGO KK
Filing Date
2021-09-01
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing technologies, the calibration process of eye-tracking systems is complex and requires prior preparation of calibration materials, making the calibration process inconvenient.

Method used

By dynamically setting a guide area in the eye-tracking system, calculating the difference between the user's viewpoint coordinates in the image and the guide area, and using this difference to correct the user's viewpoint coordinates, correction of any content can be achieved.

Benefits of technology

It simplifies the calibration process of eye-tracking systems, improves the convenience and accuracy of calibration, and enables effective calibration without the need for prior preparation of calibration content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115916032B_ABST
    Figure CN115916032B_ABST
Patent Text Reader

Abstract

An eye tracking system of an embodiment includes at least one processor. The at least one processor performs processing to dynamically set a partial region of a first content displayed on a screen as a guide region in order for a user to gaze at the partial region, determine a first viewpoint coordinate of the user in the screen based on movement of an eye of the user gazing at the guide region, calculate a difference between the determined first viewpoint coordinate and a region coordinate of the guide region in the screen, and correct a second viewpoint coordinate of a user viewing a second content displayed on the screen using the difference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of this disclosure relates to an eye-tracking system, an eye-tracking method, and a recording medium. Background Technology

[0002] Eye-tracking systems for calculating the position of a user's viewpoint are known. Patent Document 1 describes a correction method for a head-mounted eye-tracking device. In this method, while the wearer of the eye-tracking device is looking in a reference direction, eye data related to the wearer's eye position is acquired by the eye-tracking device, and this eye data is correlated with a gaze direction corresponding to the reference direction. The eye-tracking device includes a spectacle frame with ophthalmic lenses, and the gaze direction corresponding to the reference direction is determined taking into account the optical refraction function of the ophthalmic lenses. Other examples of correction methods are described in Patent Documents 2 and 3.

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent No. 6656156

[0006] Patent Document 2: Japanese Patent Application Publication No. 2010-259605

[0007] Patent Document 3: Japanese Patent Application Publication No. 2001-204692 Summary of the Invention

[0008] The problem that the invention aims to solve

[0009] The goal is to simply perform the corrections used for eye tracking.

[0010] Methods for solving problems

[0011] An eye-tracking system according to one aspect of this disclosure includes at least one processor. The at least one processor performs the following processes: dynamically setting a guide region to a portion of first content displayed on a screen so that a user gazes at it; determining the user's first viewpoint coordinates on the screen based on eye movements of the user gazing at the guide region; calculating the difference between the determined first viewpoint coordinates and the region coordinates of the guide region on the screen; and using the difference to correct the second viewpoint coordinates of the user viewing second content displayed on the screen.

[0012] In this respect, a guiding region for the user to gaze at is dynamically set for any content (the first content), and the difference for correction is calculated using this guiding region. Therefore, there is no need to prepare the content for correction in advance, thus making it easier to perform correction for eye tracking.

[0013] Invention Effects

[0014] According to one aspect of this disclosure, corrections for eye tracking can be readily performed. Attached Figure Description

[0015] Figure 1 This is a diagram illustrating an example of the application of an auxiliary system in an implementation method.

[0016] Figure 2 This is a diagram illustrating an example of the hardware structure associated with the auxiliary system of the implementation method.

[0017] Figure 3 This is a diagram illustrating an example of the functional structure associated with the auxiliary system of the implementation method.

[0018] Figure 4 This is a flowchart illustrating an example of the operation of an auxiliary system in an implementation method.

[0019] Figure 5 This is a flowchart illustrating an example of the operation of an eye-tracking system according to an implementation method.

[0020] Figure 6 This is an example diagram showing a guide area set for the first content.

[0021] Figure 7 This is an example diagram showing a guide area set for the first content.

[0022] Figure 8 This is a flowchart illustrating an example of the operation of an auxiliary system in an implementation method.

[0023] Figure 9 This is a flowchart illustrating an example of the operation of an auxiliary system in an implementation method.

[0024] Figure 10 This is a diagram representing an example of auxiliary information. Detailed Implementation

[0025] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the description of the drawings, the same or equivalent elements are labeled with the same reference numerals, and repeated descriptions are omitted.

[0026] [System Overview]

[0027] The assistive system of the implementation method is a computer system that assists users in visually recognizing content. Content refers to information provided by a computer or computer system that is identifiable by humans. Electronic data representing content is called content data. There are no limitations on the form of content presentation; for example, content can be presented as documents, images (e.g., photographs, videos, etc.), or a combination thereof. There are no limitations on the purpose and usage scenarios of content; for example, content can be used for various purposes such as education, news, speeches, business transactions, entertainment, medical care, games, and chat.

[0028] The assistive system provides content to the user by sending content data to the user terminal. The user is the person who wants to obtain information from the assistive system; that is, the viewer of the content. The user terminal can also be called a "viewer terminal." The assistive system can provide content data to the user terminal based on a request from the user, or based on instructions from a publisher different from the user. The publisher is the person who wants to transmit information to the user (viewer); that is, the sender of the content.

[0029] The assistive system provides users not only with content as needed, but also with supplementary information corresponding to the user's level of comprehension. User comprehension is an indicator of the degree to which the user understands the content. For example, if the content includes an article, user comprehension could also be an indicator of how much the user understands the article (e.g., whether the user understands the meaning of the words in the article, whether the user understands the grammar, etc.). Supplementary information is information used to facilitate the user's understanding of the content. For example, if the content includes an article, supplementary information could also be information indicating the meaning of the words in the article, the grammar, etc. In the following explanation, the user who will be the target for estimating comprehension (in other words, the user who will be the target for receiving supplementary information as needed) is called the target user, and the content visually confirmed by the target user is called the target content.

[0030] To output auxiliary information, the auxiliary system estimates the target user's comprehension level based on the movement of the target user's viewpoint on the screen displaying the target content. Specifically, the auxiliary system refers to correspondence data that represents the relationship between the movement of the user's viewpoint and the user's comprehension level. This correspondence data is electronic data generated through statistical processing of pre-acquired sample data. The sample data is electronic data representing the pairs between the movement of a user's viewpoint after visually recognizing the content and the user's comprehension level of the content. In the following description, the user providing the sample data used to generate the correspondence data is referred to as the sample user, and the content visually recognized by the sample user is referred to as the sample content.

[0031] The assistance system acquires data representing the movement of the target user's viewpoint from the target user's user terminal. This viewpoint movement data indicates how the user's viewpoint moves on the user terminal's screen and is also referred to as viewpoint data in this disclosure. Hereinafter, the data representing the movement of the target user's viewpoint (i.e., the target user's viewpoint data) will be referred to as target data. The assistance system uses the correspondence data and the target data to estimate the target user's level of understanding. Then, the assistance system outputs assistance information corresponding to the target user's level of understanding to the target user's user terminal as needed.

[0032] In this disclosure, when it is not necessary to distinguish between sample users and target users, they are sometimes referred to collectively as users.

[0033] Viewpoint data is acquired by an eye-tracking system. Based on the user's eye movements, the eye-tracking system determines the user's viewpoint coordinates at given time intervals, acquiring viewpoint data representing multiple viewpoint coordinates arranged along a time series. Viewpoint coordinates are the coordinates representing the position of the viewpoint on the user's screen. Viewpoint coordinates can also be represented using a two-dimensional coordinate system. The eye-tracking system can be installed on the user's terminal or on a different computer. Alternatively, the tracking system can also be implemented through collaboration between the user terminal and other computers.

[0034] The eye-tracking system performs a correction process to determine the user's viewpoint coordinates with higher accuracy. For example, firstly, the eye-tracking system designates a portion of the content displayed on the user's screen as a guiding area for the user to focus on. Hereinafter, the content with the designated guiding area is referred to as the first content. Then, based on the user's eye movements, the eye-tracking system determines the user's viewpoint coordinates as the first viewpoint coordinates and calculates the difference between these first viewpoint coordinates and the area coordinates of the guiding area. The area coordinates of the guiding area represent the position of the guiding area on the user's screen. Next, when the user visually confirms the content displayed on the user's screen (second content), the eye-tracking system determines the user's viewpoint coordinates as the second viewpoint coordinates based on the user's eye movements. Then, the eye-tracking system uses the pre-calculated difference to correct the determined second viewpoint coordinates. The second content is what the user is viewing during the correction of the second viewpoint coordinates.

[0035] As described above, the purpose and usage scenarios of the content are not limited. In this embodiment, educational content is shown as an example of content, and the assistance system assists students in visually recognizing the educational content. Therefore, the target content is "target content for education," and the sample content is "sample content for education." Educational content is content used to educate students, and may be, for example, tests such as exercises or exam questions, or textbooks. Educational content may also include articles, mathematical formulas, charts, or graphics. A student refers to someone who receives education in academics, skills, etc. A student is an example of a user (viewer). As described above, content may also be published to viewers based on the publisher's instructions. When the content is educational content, the publisher may also be a teacher. A teacher refers to someone who teaches students academics, skills, etc. A teacher may or may not have a teaching certificate. There are no limitations on the age or affiliation of the teacher and student. Therefore, the purpose and usage scenarios of the educational content are not limited. For example, educational content can be used in various schools, including nurseries, kindergartens, primary schools, middle schools, high schools, universities, graduate schools, vocational colleges, preparatory schools, and online schools, as well as in places or scenarios outside of schools. Relatedly, educational content can be used for various purposes such as early childhood education, compulsory education, higher education, and career education. Furthermore, educational content includes not only school education but also content used in seminars or training sessions in businesses and other settings.

[0036] [System Structure]

[0037] Figure 1 This diagram illustrates an example of the application of the auxiliary system 1 in this embodiment. In this embodiment, the auxiliary system 1 includes a server 10. The server 10 is communicatively connected to the user terminal 20 and the database 30 via a communication network N. The structure of the communication network N is not limited. For example, the communication network N may be configured to include the Internet or to include a local area network.

[0038] Server 10 is a computer that distributes content to user terminal 20 and provides auxiliary information to user terminal 20 as needed. Server 10 may also consist of one or more computers.

[0039] User terminal 20 is a computer used by a user. In this embodiment, the user is a student viewing educational content. In one example, user terminal 20 has the functions of accessing and receiving content data and auxiliary information from the auxiliary system 1 and displaying it, and sending viewpoint data to the auxiliary system 1. The type of user terminal 20 is not limited; for example, it can be a high-performance mobile phone (smartphone), tablet computer, wearable terminal (e.g., head-mounted display (HMD), smart glasses, etc.), laptop computer, mobile phone, or other portable terminal. Alternatively, user terminal 20 can also be a fixed terminal such as a desktop computer. Figure 1 The diagram shows three user terminals 20, but the number of user terminals 20 is not limited. In this embodiment, when distinguishing between the terminals of sample users and the terminals of target users, the terminals of sample users are labeled "user terminal 20A," and the terminals of target users are labeled "user terminal 20B." Users log in to the auxiliary system 1 by operating the user terminals 20 and can view content. In this embodiment, it is assumed that the user of the auxiliary system 1 is already logged in.

[0040] Database 30 is a non-temporary storage device that stores data used by auxiliary system 1. In this embodiment, database 30 stores content data, sample data, correspondence data, and auxiliary information. Database 30 can be a single database or a collection of multiple databases.

[0041] Figure 2 This is a diagram illustrating an example of the hardware structure associated with auxiliary system 1. Figure 2 This refers to the server computer 100, which functions as a server 10, and the terminal computer 200, which functions as a user terminal 20.

[0042] As an example, the server computer 100 includes a processor 101, a main storage unit 102, an auxiliary storage unit 103, and a communication unit 104 as hardware components.

[0043] Processor 101 is a computing device that executes operating systems and applications. Examples of processors include CPU (Central Processing Unit) and GPU (Graphics Processing Unit), but processor 101 is not limited to these types.

[0044] The main storage unit 102 is a device for storing programs used to implement the server 10, calculation results output from the processor 101, and the like. The main storage unit 102 is composed of at least one of ROM (Read Only Memory) and RAM (Random Access Memory).

[0045] The auxiliary storage unit 103 is generally a device capable of storing a larger amount of data than the main storage unit 102. The auxiliary storage unit 103 is, for example, composed of a non-volatile storage medium such as a hard disk or flash memory. The auxiliary storage unit 103 stores the server program P1, which enables the server computer 100 to function as a server 10, and various data. In this embodiment, an auxiliary program is installed as the server program P1.

[0046] The communication unit 104 is a device that performs data communication with other computers via a communication network N. The communication unit 104 may be composed of, for example, a network interface card (NIC) or a wireless communication module.

[0047] The various functional elements of server 10 are implemented by having processor 101 or main storage unit 102 read server program P1 and execute the program. Server program P1 contains code for implementing the various functional elements of server 10. Processor 101 activates communication unit 104 according to server program P1, performing data reading and writing in main storage unit 102 or auxiliary storage unit 103. The various functional elements of server 10 are implemented through such processing.

[0048] Server 10 can consist of one or more computers. When multiple computers are used, these computers are interconnected via a communication network N, thereby logically forming a server 10.

[0049] As an example, the terminal computer 200 includes a processor 201, a main storage unit 202, an auxiliary storage unit 203, a communication unit 204, an input interface 205, an output interface 206, and a camera unit 207 as hardware components.

[0050] Processor 201 is a computing device that executes an operating system and applications. Processor 201 may be, for example, a CPU or a GPU, but is not limited to these types.

[0051] The main storage unit 202 is a device for storing programs used to implement the user terminal 20, calculation results output from the processor 201, and the like. The main storage unit 202 is, for example, composed of at least one of ROM and RAM.

[0052] The auxiliary storage unit 203 is generally a device capable of storing a larger amount of data than the main storage unit 202. The auxiliary storage unit 203 is composed of, for example, a non-volatile storage medium such as a hard disk or flash memory. The auxiliary storage unit 203 stores the client program P2 used to enable the terminal computer 200 to function as a user terminal 20, and various other data.

[0053] The communication unit 204 is a device that performs data communication with other computers via a communication network N. The communication unit 204 may be composed of, for example, a network interface card (NIC) or a wireless communication module.

[0054] The input interface 205 is a device for receiving data based on user operations or actions. For example, the input interface 205 may consist of at least one of a keyboard, operation buttons, indicator devices, touch panel, microphone, sensor, and camera.

[0055] Output interface 206 is a device for outputting data processed by terminal computer 200. For example, output interface 206 may consist of at least one of a monitor, touch panel, HMD, and speaker.

[0056] The camera unit 207 is a device for capturing images of the real world, specifically a camera. The camera unit 207 can capture moving images (videos) or still images (photographs). The camera unit 207 can also function as an input interface 205.

[0057] The various functional elements of the user terminal 20 are implemented by having the processor 201 or the main storage unit 202 read the client program P2 and execute the program. The client program P2 contains code for implementing the various functional elements of the user terminal 20. The processor 201 operates the communication unit 204, the input interface 205, the output interface 206, or the camera unit 207 according to the client program P2, and reads and writes data from the main storage unit 202 or the auxiliary storage unit 203. Through this process, the various functional elements of the user terminal 20 are implemented.

[0058] At least one of the server program P1 and the client program P2 may also be provided on a tangible recording medium such as a CD-ROM, DVD-ROM, or semiconductor memory, which is not temporarily recorded thereon. Alternatively, at least one of these programs may also be provided as a data signal superimposed on a carrier wave via a communication network N. These programs may be provided individually or together.

[0059] Figure 3This diagram illustrates an example of the functional structure associated with the auxiliary system 1. The server 10 includes a content distribution unit 11, a statistical processing unit 12, an estimation unit 13, and an auxiliary unit 14 as functional elements. The statistical processing unit 12 is a functional element that generates correspondence data. The statistical processing unit 12 generates correspondence data by performing statistical processing on sample data stored in the database 30, and saves the correspondence data in the database 30. The estimation unit 13 is a functional element that estimates the target user's understanding of the target content. The estimation unit 13 obtains target data representing the movement of the target user's viewpoint from the user terminal 20B, and estimates the target user's understanding based on this target data and the correspondence data. The auxiliary unit 14 is a functional element that sends auxiliary information corresponding to the target user's understanding to the user terminal 20B.

[0060] The user terminal 20 includes a setting unit 21, a determining unit 22, a calculating unit 23, a tracking unit 24, and a display control unit 25 as functional elements. The setting unit 21 is a functional element that sets a portion of the first content displayed on the screen of the user terminal 20 as a guiding area. The recognizing unit 22 is a functional element that recognizes the user's first viewpoint coordinates based on the movement of the user's eyes while viewing the guiding area. The calculating unit 23 is a functional element that calculates the difference between the area coordinates of the guiding area set by the setting unit 21 and the first viewpoint coordinates determined by the determining unit 22. The tracking unit 24 is a functional element that generates viewpoint data by observing the movement of the user's eyes while viewing the content displayed on the screen of the user terminal 20. The tracking unit 24 uses the calculated difference to correct the second viewpoint coordinates of the user viewing the second content, generating viewpoint data representing the corrected second viewpoint coordinates. The display control unit 25 is a functional element that controls the display of the screen on the user terminal 20. In this embodiment, the eye-tracking system is composed of the setting unit 21, the determining unit 22, the calculating unit 23, the tracking unit 24, and the display control unit 25.

[0061] [System Actions]

[0062] Figure 4 This is a flowchart representing the actions of auxiliary system 1 as processing flow S1. (Refer to...) Figure 4 The overall processing of auxiliary system 1 will be explained.

[0063] In step S11, the statistical processing unit 12 of the server 10 performs statistical processing on multiple sample data to generate corresponding relationship data.

[0064] An example of collecting sample data, which is a prerequisite for step S11, will be described. First, the content distribution unit 11 distributes sample content to each of the multiple user terminals 20A. There is no time limit for distributing the sample content to each user terminal 20A. For example, the content distribution unit 11 can distribute the sample content to each user terminal 20A based on a request from that user terminal 20A, or it can distribute the sample content to two or more user terminals 20A simultaneously. In each user terminal 20A, the display control unit 25 receives and displays the sample content. Then, the tracking unit 24 of the user terminal 20A generates viewpoint data, which represents the movement of the viewpoint of the sample user who has visually confirmed the sample content. In one example, the sample user inputs their level of understanding of the sample content into the user terminal 20A by answering a questionnaire, and the user terminal 20A accepts this input data. Alternatively, the user terminal 20A or the server 10 can estimate the sample user's level of understanding based on the user's answers to the sample content (e.g., answers to questions). Input or estimated comprehension level, for example, indicates whether the user can understand the meaning of words in an article, or whether they can understand the grammar of the article. In one example, user terminal 20A generates sample data representing pairs between the generated viewpoint data and the input or estimated comprehension level, and sends this sample data to server 10. Alternatively, user terminal 20A can send viewpoint data to server 10, and server 10 can generate sample data representing pairs between the viewpoint data and the estimated comprehension level. In either case, server 10 stores the sample data in database 30. Server 10 stores multiple sample data obtained from multiple user terminals 20A regarding a specific sample content in database 30. Server 10 can store multiple sample data for each of the multiple sample contents. The auxiliary system 1 collects sample data through this series of processes.

[0065] The statistical processing unit 12 reads multiple sample data from the database 30, performs statistical processing on the multiple sample data, and generates corresponding relationship data. The method of statistical processing performed by the statistical processing unit 12 and the format of the generated corresponding relationship data are not limited.

[0066] As an example, the statistical processing unit 12 clusters multiple sample data based on the movement of the sample users' viewpoints and their understanding of the sample content, thereby generating correspondence data. The statistical processing unit 12 can also determine the similarity of viewpoint movements based on at least one of the following: viewpoint movement speed, the number of viewpoint reversals (the number of times the viewpoint's movement direction changes), and the area of ​​the viewpoint movement region. The statistical processing unit 12 can also determine the similarity of content comprehension based on at least one of the comprehension of word meanings and the comprehension of article grammar. The statistical processing unit 12 can also vectorize the features related to viewpoint movement and the features related to comprehension as feature vectors for each sample data, so that sample data with common or similar feature vectors belong to the same class. The statistical processing unit 12 derives the correspondence between the user's viewpoint movement and the user's comprehension based on the clustering results. More specifically, this correspondence can be described as a pair representing the user's viewpoint movement tendency and the corresponding comprehension level. The statistical processing unit 12 generates correspondence data representing this correspondence and stores this correspondence data in the database 30.

[0067] As another example, the statistical processing unit 12 can also generate corresponding relationship data by performing regression analysis. Specifically, the statistical processing unit 12 quantifies the movement of the sample users' viewpoints and the sample users' comprehension levels based on prescribed rules. The statistical processing unit 12 performs regression analysis on the quantified data to generate a regression equation with the sample users' comprehension levels as the objective variable and the movement of the sample users' viewpoints as the explanatory variable. At this time, the statistical processing unit 12 can also decompose the movement of the sample users' viewpoints into multiple factors such as the movement speed of the viewpoints and the number of viewpoint reversals, and set multiple explanatory variables corresponding to these multiple factors. For example, the statistical processing unit 12 can also quantify the movement speed of the viewpoints and the number of viewpoint reversals as independent explanatory variables and perform a multiple regression analysis using these multiple explanatory variables. The statistical processing unit 12 stores the regression equations generated by the regression analysis as corresponding relationship data in the database 30. The regression analysis method performed by the statistical processing unit 12 can also be partial least squares regression (PLS) or support vector regression (SVR). In summary, this corresponding relationship data also represents the pair between the user's viewpoint movement tendency and the comprehension level corresponding to that tendency.

[0068] As another example, the statistical processing unit 12 can also analyze the correspondence between the movement of a sample user's viewpoint and the sample user's level of understanding using machine learning, generating correspondence data. Machine learning can also be deep learning using neural networks. The statistical processing unit 12 uses a machine learning model configured to output data representing the user's level of understanding when data representing the movement of the user's viewpoint is input into the input layer, performing teacher-guided learning using the sample data as learning data, and adjusting the weighting parameters within the learning model. The statistical processing unit 12 stores the model with adjusted weighting parameters (the learned model) as correspondence data in the database 30. In the case of machine learning, the statistical processing unit 12 can also preprocess the sample data stored in the database 30, converting it into data suitable for machine learning.

[0069] The statistical processing unit 12 may also appropriately select sample data used in the statistical processing to generate various correspondence data. For example, the statistical processing unit 12 may generate correspondence data for each of multiple sample contents using sample data obtained from multiple sample users who viewed that sample content. In this case, correspondence data is generated for each content. Hereinafter, this correspondence data will be referred to as "content-specific correspondence data". Alternatively, the statistical processing unit 12 may also generate correspondence data using sample data from multiple sample contents (e.g., multiple sample contents belonging to the same category). In this case, common correspondence data is generated for multiple contents (e.g., multiple contents belonging to the same category). Hereinafter, this correspondence data will be referred to as "generalized correspondence data".

[0070] In step S12, the assistance unit 14 provides auxiliary information to the target user who is visually confirming the target content as needed. The assistance unit 14 estimates the target user's level of understanding of the target content and provides auxiliary information corresponding to that level of understanding as needed. The details of the processing of outputting auxiliary information will be described later. The correspondence between the user's level of understanding and the auxiliary information is predetermined, and the auxiliary information is pre-stored in the database 30 in a way that determines this correspondence. Alternatively, the user's level of understanding and the auxiliary information can be correlated in a way that supplements the target user's insufficient understanding of the target content. For example, for a user's level of understanding of a word contained in the text that they do not understand, the meaning of that word can also be used as auxiliary information.

[0071] Figure 5This is a flowchart representing the actions of the eye-tracking system as processing flow S2. The processing based on the eye-tracking system is roughly divided into the process of calculating the difference used in the correction of the viewpoint coordinates (steps S21 to S23), and the process of using the calculated difference to correct the user's viewpoint coordinates (steps S24 and S25).

[0072] In step S21, the setting unit 21 dynamically sets a portion of the first content displayed on the screen of the user terminal 20 as a guide area. The first content is any content distributed by the content distribution unit 11 and displayed by the display control unit 25. The first content can be educational content or content not intended for education. The guide area is an area designed to attract the user's attention and consists of a series of pixels arranged consecutively. Dynamically setting the guide area means setting the guide area within the first content in response to the display of the first content on the screen in an area not pre-defined for the user's attention. In one example, the guide area is set only while the first content is displayed on the screen. The position of the guide area within the first content displayed on the screen is not limited. For example, the setting unit 21 can set the guide area at any position, such as the center, top, bottom, or corner of the first content. In one example, after the setting unit 21 sets the guide area, the display control unit 25 displays the guide area within the first content based on that setting. The shape and area (number of pixels) of the guide area are also not limited. The guide area is the area that the user focuses on in order to correct the viewpoint coordinates. Therefore, typically, the setting unit 21 sets the area of ​​the guide area to be much smaller than the area of ​​the first content displayed on the screen (i.e., the area of ​​the display device).

[0073] There is no limitation on the method for dynamically setting the guide area. In one example, the setting unit 21 can also visually distinguish the guide area from the non-guide area by making the display mode of the guide area different from that of the area outside the guide area (hereinafter also referred to as the non-guide area). There is no limitation on the method for setting the display mode. As a specific example, the setting unit 21 can also distinguish the guide area from the non-guide area by relatively increasing the resolution of the guide area without changing the resolution of the guide area. As another specific example, the setting unit 21 can also distinguish the guide area from the non-guide area by blurring the non-guide area without changing the display mode of the guide area. For example, the setting unit 21 can also perform blurring by setting the color of a target pixel in the non-guide area to the average color of the colors of multiple pixels adjacent to that target pixel. The setting unit 21 can perform blurring while maintaining the resolution of the non-guide area, or it can perform blurring by reducing the resolution. As another specific example, the setting unit 21 can also distinguish the guide area from the non-guide area by surrounding the outer edge of the guide area with a specific color or a specific type of frame. The setting unit 21 can also distinguish the guide area from other areas by combining any two or more methods, including adjusting the resolution, blurring, and drawing the outline.

[0074] Alternatively, if the first content includes a selection object that can be chosen by the user, the setting unit 21 may also set the area displaying the selection object as a guide area. That is, the setting unit 21 may also define the selection object as a partial area and set it as a guide area. Typically, the selection object may also be a selection button or link displayed in the application's tutorial screen. Alternatively, when the user terminal 20 is performing a question exercise or test, the selection object may be a button for selecting a question or a button for starting the exercise or test. The setting unit 21 may reduce the resolution of the non-guided area while maintaining the resolution of the selection object set as a guide area. Based on or instead of this process, the setting unit 21 may blur the non-guided area or surround the outer edge of the selection object set as a guide area with a specific color or type of border.

[0075] The setting unit 21 sets the region coordinates of the guide area using any method. For example, the setting unit 21 may set the coordinates of the center or centroid of the guide area as the region coordinates. Alternatively, the setting unit 21 may set the position of any single pixel in the guide area as the region coordinates.

[0076] In step S22, the determining unit 22 determines the user's viewpoint coordinates in the gaze guidance area as the first viewpoint coordinates. The determining unit 22 determines the viewpoint coordinates based on the movement of the user's eyes. The method for determining the viewpoint coordinates is not limited. As an example, the determining unit 22 may also capture a peripheral image of the user's eyes using the camera unit 207 of the user terminal 20, and determine the viewpoint coordinates based on the position of the iris with the user's inner corner of the eye as a reference point. As another example, the determining unit 22 may also use the pre-contrast corneal reflection method (PCCR) to determine the user's viewpoint coordinates. When using the pre-contrast corneal reflection method, the user terminal 20 may also have an infrared emitting device and an infrared camera as hardware structures.

[0077] In step S23, the calculation unit 23 calculates the difference between the first viewpoint coordinates determined by the determination unit 22 and the area coordinates of the guide area set by the setting unit 21. For example, when the position on the screen of the user terminal 20 is represented by the XY coordinate system, if the first viewpoint coordinates are (105, 105) and the area coordinates are (100, 100), the difference is (105-100, 105-100) = (5, 5). The calculation unit 23 stores the calculated difference in any storage device such as the main storage unit 202 or the auxiliary storage unit 203.

[0078] To improve the accuracy of the correction, the user terminal 20 may also change the position of the guide area while repeatedly processing from step S21 to step S23. In this case, the calculation unit 23 may also set the statistical values ​​(e.g., average values) of the calculated multiple differences as the differences used in the subsequent correction process (step S25).

[0079] In step S24, the tracking unit 24 determines the viewpoint coordinates of the user viewing the second content as the second viewpoint coordinates. The second content is any content distributed by the content distribution unit 11 and displayed by the display control unit 25. For example, the second content can be sample content or target content. The tracking unit 24 can determine the second viewpoint coordinates using the same method as the determination unit 22 in determining the first viewpoint coordinates (i.e., the same method as in step S22). The second content can be different from or the same as the first content.

[0080] In step S25, the tracking unit 24 uses a difference to correct the second viewpoint coordinates. For example, when the second viewpoint coordinates determined in step S24 are (190, 155) and the difference calculated in step S23 is (5, 5), the tracking unit 24 corrects the second viewpoint coordinates so that (190-5, 155-5) = (185, 150).

[0081] The tracking unit 24 can also repeatedly perform steps S24 and S25 to obtain multiple corrected second viewpoint coordinates arranged in a time sequence and generate viewpoint data representing the movement of the user's viewpoint. Alternatively, the tracking unit 24 can also obtain multiple corrected second viewpoint coordinates, and the server 10 can generate viewpoint data based on these multiple second viewpoint coordinates.

[0082] Reference Figure 6 and Figure 7 The following explains the setting example of the guide area. Figure 6 and Figure 7 These are examples of diagrams showing the guide area set by the setting unit 21 for the first content.

[0083] exist Figure 6 In this example, the setting unit 21 sets the guiding area by reducing the resolution of the non-guiding area. In this example, the user terminal 20 displays first content C11 including a child, a lawn, and a ball, and calculates the difference while changing the position of the guiding area on the first content C11. As the position of the guiding area changes, the display changes in the order of screens D11, D12, and D13. Figure 6 In the middle, the non-guided area is represented by a dashed line.

[0084] First, the setting unit 21 sets the area of ​​the child's face as the guide area A11. The screen D11 corresponds to this setting. The setting unit 21 does not change the resolution of the guide area A11, but reduces the resolution of the area outside the guide area A11 (the non-guide area). For example, the setting unit 21 may reduce the resolution of the non-guide area so that the resolution of the guide area A11 is more than twice or four times that of the non-guide area. For example, when the resolution of the guide area A11 is 300 ppi, the resolution of the non-guide area may be less than 150 ppi or less than 75 ppi. With this resolution setting, the non-guide area is displayed more blurry than the guide area A11, so the user's gaze is usually directed towards the clearly displayed guide area A11. Thus, the viewpoint coordinates (first viewpoint coordinates) of the user looking at the guide area A11 can be determined. During the display of the screen D11, the determination unit 22 acquires the user's first viewpoint coordinates. Next, the calculation unit 23 calculates the difference between the first viewpoint coordinates and the area coordinates of the guide area A11.

[0085] Next, the setting unit 21 sets the portion of the sphere as the guide area A12. Screen D12 corresponds to this setting. The setting unit 21 restores the resolution of the guide area A12 to its original value and reduces the resolution of areas outside the guide area A12 (non-guide area). As a result, the user's gaze is usually directed towards the guide area A12. While screen D12 is being displayed, the determination unit 22 acquires the user's first viewpoint coordinates. Then, the calculation unit 23 calculates the difference between the first viewpoint coordinates and the area coordinates of the guide area A12.

[0086] Next, the setting unit 21 sets the lower right portion (the lawn area) of the first content C11 as the guide area A13. Screen D13 corresponds to this setting. The setting unit 21 restores the resolution of the guide area A13 to its original value and reduces the resolution of areas outside the guide area A13 (non-guide areas). As a result, the user's gaze is usually directed towards the guide area A13. While screen D13 is being displayed, the determination unit 22 acquires the user's first viewpoint coordinates. Then, the calculation unit 23 calculates the difference between the first viewpoint coordinates and the area coordinates of the guide area A13. The calculation unit 23 calculates a statistical value for the calculated multiple differences. This statistical value is used by the tracking unit 24 to correct the second viewpoint coordinates (step S25).

[0087] exist Figure 7 In this example, the setting unit 21 sets the selected object within the first content C21 as the guide area. In this example, the first content C21 is a tutorial for an online academic proficiency test. As the tutorial progresses, the display changes in the order of screens D11, D12, and D13.

[0088] Screen D21 contains the string "A question about Chinese." and an OK button. The OK button is the selection object. The setting unit 21 sets the area displaying the OK button as the guide area A21. Normally, the user looks at the selection object when operating on it. Therefore, the viewpoint coordinates (first viewpoint coordinates) of the user looking at the guide area A21 can be determined. In one example, when the user selects the OK button, the determination unit 22 obtains the user's first viewpoint coordinates. Then, the calculation unit 23 calculates the difference between the first viewpoint coordinates and the area coordinates of the guide area A21.

[0089] If the user presses the OK button, the display control unit 25 switches screen D21 to screen D22. Screen D22 includes a string such as "Please select the number of questions." and three selection buttons: "5", "10", and "15". These selection buttons represent the selection objects. The setting unit 21 designates the areas displaying the three selection buttons as guide areas A22, A23, and A24, respectively. In one example, when the user selects any one of the three selection buttons, the determining unit 22 determines the user's viewpoint coordinates (first viewpoint coordinates). Then, the calculation unit 23 calculates the difference between the first viewpoint coordinates and the coordinates of the area corresponding to the selected object (one of guide areas A22 to A24).

[0090] If the user selects a selection button, the display control unit 25 switches screen D22 to screen D23. Screen D23 includes the string "Start test?" and a start button. The start button is the selection object. The setting unit 21 sets the area displaying the start button as the guide area A25. In one example, when the user selects the start button, the determination unit 22 obtains the user's first viewpoint coordinates. Next, the calculation unit 23 calculates the difference between the first viewpoint coordinates and the area coordinates of the guide area A25. The calculation unit 23 calculates the statistical value of the calculated multiple differences. This statistical value is used by the tracking unit 24 to correct the second viewpoint coordinates (step S25).

[0091] Figure 8 This is a flowchart illustrating an example of the actions of the assistance system 1 as processing flow S3. Processing flow S3 represents the steps of providing assistance information to the target user viewing the target content. Processing flow S3 assumes that the target user has logged into assistance system 1. Furthermore, it assumes that the eye-tracking system has already calculated the difference used in the correction of the viewpoint coordinates.

[0092] In step S31, the display control unit 25 of the user terminal 20B displays the target content on the screen of the user terminal 20B. The display control unit 25 receives, for example, content data distributed from the content distribution unit 11 from the server 10, and displays the target content based on the content data.

[0093] In step S32, the tracking unit 24 of the user terminal 20B acquires the viewpoint coordinates (second viewpoint coordinates) of the target user visually confirming the target content. Specifically, the tracking unit 24 determines the viewpoint coordinates (viewpoint coordinates before correction) based on the movement of the target user's eyes while viewing the target content, and corrects the determined viewpoint coordinates using a pre-calculated difference. The tracking unit 24 may also acquire the corrected viewpoint coordinates at given time intervals, generating viewpoint data (i.e., target data representing the movement of the target user's viewpoint) arranged along a time sequence of these multiple viewpoint coordinates.

[0094] In step S33, the estimation unit 13 acquires target data. For example, the estimation unit 13 can receive target data from the tracking unit 24 of the user terminal 20B. Alternatively, the tracking unit 24 can send multiple corrected viewpoint coordinates to the server 10 in sequence, and the estimation unit 13 can generate viewpoint data (target data) arranged along a time series of these multiple viewpoint coordinates.

[0095] In step S34, the estimation unit 13 retrieves correspondence data from the database 30 and estimates the target user's understanding of the target content based on the target data and the correspondence data. For example, when the correspondence data is generated through clustering, the estimation unit 13 estimates the understanding represented by the class to which the target data belongs as the target user's understanding. For another example, when the correspondence data is generated through regression analysis, the estimation unit 13 applies the target data to the regression equation to estimate the target user's understanding. As yet another example, when the correspondence data is a fully learned model, the estimation unit 13 estimates the target user's understanding by inputting the target data into the fully learned model.

[0096] In step S35, the assistance unit 14 retrieves assistance information corresponding to the target user's level of understanding from the database 30 and sends the assistance information to the user terminal 20B. The display control unit 25 of the user terminal 20B displays the assistance information on the screen of the user terminal 20B. The timing of the assistance information output is not limited. For example, the display control unit 25 may output the assistance information after a predetermined time (e.g., 15 seconds) has elapsed since the target content was displayed on the screen of the user terminal 20B. Alternatively, the display control unit 25 may output the assistance information upon request from the user. The display control unit 25 may also adjust the display time of the assistance information according to the user's level of understanding. Alternatively, the display control unit 25 may only display the assistance information during a display time pre-set by the user, etc. Alternatively, the assistance unit 14 may display the assistance information until the display of the target content is switched, or it may display the assistance information until the user inputs information about the target content (e.g., an answer to a question). If the estimated level of understanding indicates that the target user's understanding of the target content is sufficient, the assistance unit 14 may terminate the process without outputting assistance information. The method of outputting the assistance information is not limited. When the auxiliary information includes audio data, the user terminal 20 can also output the audio data from the speaker.

[0097] As shown in step S36, the auxiliary system 1 repeatedly performs the processing from step S32 to step S35 while the target content is displayed on the user terminal 20B. For example, the auxiliary system 1 repeatedly performs this series of processes while the target content is displayed.

[0098] Figure 9This is a flowchart illustrating the actions of the assistive system 1 as an example of processing flow S4. Processing flow S4 also involves providing assistive information to the target user viewing the target content, but the specific steps differ from those in processing flow S3. Processing flow S4 also assumes that the target user has logged into assistive system 1 and that the eye-tracking system has already calculated the difference.

[0099] In step S41, the display control unit 25 of the user terminal 20B displays the target content on the screen of the user terminal 20B. In step S42, the tracking unit 24 of the user terminal 20B acquires the viewpoint coordinates (second viewpoint coordinates) of the target user who visually confirms the target content. In step S43, the estimation unit 13 acquires target data representing the movement of the target user's viewpoint. This series of processes is the same as steps S31 to S33.

[0100] In step S44, the estimation unit 13 obtains generalized correspondence data from the database 30, and estimates the target user's understanding of the target content (first understanding) based on the target data and the generalized correspondence data. The specific estimation method is the same as in step S34.

[0101] In step S45, the auxiliary unit 14 retrieves auxiliary information corresponding to the target user's first level of understanding from the database 30 and sends the auxiliary information to the user terminal 20B. The display control unit 25 of the user terminal 20B outputs the auxiliary information to the screen of the user terminal 20B.

[0102] In step S46, the assistance unit 14 determines whether to provide additional assistance to the target user, that is, whether to provide additional assistance information to the target user. If the assistance unit 14 determines that no additional assistance should be provided, the process moves to step S49. If the assistance unit 14 determines that additional assistance should be provided, the process moves to step S47. As an example, if user input regarding the target content (e.g., an answer to a question) is performed within a specified time, the assistance unit 14 may determine that no additional assistance should be provided; if no such user input is performed within the specified time, the assistance unit 14 may determine that additional assistance should be provided.

[0103] In step S47, the estimation unit 13 refers to the database 30 to obtain target content-specific correspondence data, and estimates the target user's understanding of the target content (second understanding) based on the target data and the content-specific correspondence data. This process assumes that the same content is used as both sample content and target content. The specific estimation method is the same as in step S34.

[0104] In step S48, the auxiliary unit 14 retrieves additional auxiliary information corresponding to the second level of understanding of the target user from the database 30 and sends the auxiliary information to the user terminal 20B. The display control unit 25 of the user terminal 20B outputs the additional auxiliary information to the screen of the user terminal 20B.

[0105] As shown in step S49, the auxiliary system 1 repeatedly performs the processing from step S42 to step S48 while the user terminal 20B is displaying the target content. For example, the auxiliary system 1 repeatedly performs this series of processes while the target content is being displayed.

[0106] Figure 10 This is a diagram illustrating an example of auxiliary information. In this example, suppose the target content Q11 is part of an English question, and the target user is a Japanese student. In this example, the auxiliary system 1 refers to correspondence data containing information related to comprehension level Ra (indicating insufficient vocabulary), comprehension level Rb (indicating insufficient grammatical ability), and comprehension level Rc (indicating insufficient understanding of the context of the text). For example, if the estimation unit 13 estimates that the target user's vocabulary is insufficient based on the target data and its correspondence data, the auxiliary unit 14 outputs auxiliary information B11 corresponding to that comprehension level. If the estimation unit 13 estimates that the target user's grammatical ability is insufficient, the auxiliary unit 14 outputs auxiliary information B12 corresponding to that comprehension level. If the estimation unit 13 estimates that the target user does not understand the context of the text, the auxiliary unit 14 outputs auxiliary information B13 corresponding to that comprehension level. The display control unit 25 of the user terminal 20B displays the output auxiliary information. The target user can refer to this auxiliary information to solve the problem.

[0107] [Effect]

[0108] As described above, one aspect of the eye-tracking system disclosed herein includes at least one processor. The at least one processor performs the following processes: dynamically setting a guide region to a portion of first content displayed on a screen so that a user gazes at it; determining the user's first viewpoint coordinates on the screen based on eye movements of the user gazing at the guide region; calculating the difference between the determined first viewpoint coordinates and the region coordinates of the guide region on the screen; and using the difference to correct the second viewpoint coordinates of the user viewing second content displayed on the screen.

[0109] One aspect of this disclosure relates to an eye-tracking method performed by an eye-tracking system having at least one processor. The eye-tracking method includes the following steps: dynamically defining a guide region as a portion of first content displayed on a screen to induce a user to gaze at it; determining the user's first viewpoint coordinates on the screen based on eye movements of the user gazing at the guide region; calculating the difference between the determined first viewpoint coordinates and the region coordinates of the guide region on the screen; and using the difference to correct the second viewpoint coordinates of a user viewing second content displayed on the screen.

[0110] An eye-tracking program according to one aspect of this disclosure causes a computer to perform the following steps: dynamically setting a portion of a first content displayed on a screen as a guide area in order to make a user gaze at the guide area; determining the user's first viewpoint coordinates on the screen based on the movement of the user's eyes while gazing at the guide area; calculating the difference between the determined first viewpoint coordinates and the area coordinates of the guide area on the screen; and using the difference to correct the second viewpoint coordinates of a user viewing a second content displayed on the screen.

[0111] In this respect, a guiding region for the user to gaze at is dynamically set for any content (the first content), and the difference for correction is calculated using this guiding region. Therefore, there is no need to prepare the content for correction in advance, thus making it easier to perform correction for eye tracking.

[0112] In other eye-tracking systems, at least one processor may designate a specific area as a guide area by adjusting the screen resolution so that the resolution of that area is higher than that of the areas outside that area. By relatively increasing the resolution of the guide area compared to the resolution of other areas, the user's gaze can naturally be directed towards the guide area when viewing the primary content.

[0113] In other eye-tracking systems, at least one processor may designate a portion of the area as a guiding region by blurring the area outside that region. In this case, the guiding region is displayed more clearly than the non-guiding region, thus allowing the user's gaze to naturally gravitate towards the guiding region when viewing the initial content.

[0114] In other eye-tracking systems, at least one processor may define a portion of the region as a guide region by surrounding the outer edge of the region with a frame. In this case, since the guide region is distinguished from other regions by the frame, the location of the guide region can be easily identified by the user.

[0115] In other eye-tracking systems, at least one processor may designate an area displaying a selectable object as a guide area. When selecting an object, the user typically gazes at it. Therefore, by designating the display area of ​​the selected object as the guide area, the user's gaze can be naturally directed towards the guide area.

[0116] In other eye-tracking systems, at least one processor may perform multiple difference calculations by changing the set position of the guide region in the first content. In this case, the calculated multiple differences can be used to correct the second viewpoint coordinates, thereby improving the correction accuracy.

[0117] [Variation Example]

[0118] The embodiments described above are based on the present disclosure. However, the present disclosure is not limited to the above embodiments. Various modifications can be made to the present disclosure without departing from its spirit.

[0119] In the above embodiment, the auxiliary system 1 is configured using server 10, but it can also be configured without server 10. In this case, the functional elements of server 10 can be installed on any user terminal 20, for example, on either the terminal used by the content publisher or the terminal used by the content viewer. Alternatively, the functional elements of server 10 can be installed separately on multiple user terminals 20, for example, separately on the terminal used by the publisher and the terminal used by the viewer. Relatedly, the auxiliary program can also be implemented as a client program. By having the functions of server 10 on the user terminal 20, the load on server 10 can be reduced. Furthermore, information related to viewers of content, such as students (e.g., data indicating viewpoint movement), is not sent outside the user terminal 20, thus providing more reliable protection for the viewer's privacy.

[0120] In the above embodiment, the eye-tracking system consists only of the user terminal 20, but the system can also be configured using the server 10. In this case, several functional elements of the user terminal 20 can also be installed on the server 10. For example, functional elements equivalent to the computing unit 23 can also be installed on the server 10.

[0121] In the above embodiments, auxiliary information is displayed separately from the target content, but auxiliary information can also be displayed as part of the target content. For example, if the target content includes an article, the auxiliary unit 14 can also emphasize a part of the article (e.g., a part important for understanding the article) as auxiliary information. That is, auxiliary information can also be a visual effect attached to the target content. In this case, the auxiliary unit 14 can also perform its emphasis display by making the color or font of the part of the article that is the object of auxiliary information different from other parts.

[0122] In the above embodiment, the assistance system 1 outputs assistance information corresponding to the target user's level of understanding. However, the assistance system 1 may also output assistance information without using this level of understanding. This variation will be described below.

[0123] Server 10 obtains viewpoint data representing the movement of the viewpoints of sample users who have visually confirmed the sample content, and sample data representing auxiliary information prompted to the sample users, from each user terminal 20A, and stores the sample data in database 30. In one example, the auxiliary information prompted to the sample users (i.e., the auxiliary information corresponding to the sample users) is determined through manual experiments or surveys, questionnaires for sample users, etc., and input into user terminal 20A. Statistical processing unit 12 performs statistical processing on the sample data in database 30, generates correspondence data representing the correspondence between the movement of the user's viewpoint and the auxiliary information of the content, and stores the correspondence data in database 30. Similar to the above embodiment, the statistical processing method and the form of the generated correspondence data are not limited. Therefore, statistical processing unit 12 can generate correspondence data using various methods such as clustering, regression analysis, and machine learning.

[0124] Server 10 outputs auxiliary information corresponding to the target data based on the target data and its corresponding relationship data received from user terminal 20B. In one example, estimation unit 13 retrieves the corresponding relationship data from database 30 and determines the auxiliary information corresponding to the target data. When the corresponding relationship data is generated through clustering, estimation unit 13 determines the auxiliary information represented by the class to which the target data belongs. As another example, when the corresponding relationship data is generated through regression analysis, estimation unit 13 applies the target data to the regression equation as auxiliary information. As yet another example, when the corresponding relationship data is a fully learned model, estimation unit 13 determines the auxiliary information by inputting the target data into the fully learned model. Auxiliary unit 14 retrieves the determined auxiliary information from database 30 and sends the auxiliary information to user terminal 20B.

[0125] In this disclosure, the expression "at least one processor executes a first process, executes a second process, ... executes an nth process" or its corresponding expression is a concept that includes the case where the executing entity (i.e., the processor) of the n processes from the first process to the nth process changes midway. That is, the expression includes both the case where all n processes are executed by the same processor and the case where the processor changes in any direction among the n processes.

[0126] The processing order of the method executed by at least one processor is not limited to the example in the above embodiments. For example, some of the above steps (processes) may be omitted, or the steps may be executed in a different order. In addition, any two or more of the above steps may be combined, or a part of the steps may be modified or deleted. Alternatively, other steps may be performed based on the above steps.

[0127] Symbol Explanation

[0128] 1…Auxiliary System, 10…Server, 11…Content Distribution Department, 12…Statistical Processing Department, 13…Estimation Department, 14…Auxiliary Department, 20, 20A, 20B…User Terminal, 21…Setting Department, 22…Determination Department, 23…Calculation Department, 24…Tracking Department, 25…Display Control Department, 30…Database, 100…Server Computer, 101…Processor, 102…Main Storage Department, 103…Secondary Storage Department, 104…Communication Department, 200…Terminal Computer, 201…Processor, 202…Main storage, 203…Auxiliary storage, 204…Communication, 205…Input interface, 206…Output interface, 207…Camera, A11, A12, A13, A21, A22, A23, A24, A25…Boot area, C11, C21…First content, D11, D12, D13, D21, D22, D23…Screen, N…Communication network, P1…Server program, P2…Client program.

Claims

1. An eye-tracking system comprising at least one processor executing an operating system and applications, wherein, The at least one processor performs the following processing: To draw the user's attention to a portion of the first content displayed on the screen, the following portions of the first content displayed on multiple screens after the application is started are dynamically set as their respective guide areas: these portions are areas in each of the multiple screens that display selectable objects that the user can choose. Based on the movement of the user's eyes in the guidance area when the selected object is selected in each of the multiple screens, the coordinates of the user's first viewpoint in each of the multiple screens are determined. For each of the plurality of frames, calculate the difference between the determined first viewpoint coordinates and the region coordinates of the guide area in the frame; The second viewpoint coordinates of the user viewing the second content displayed on the screen are corrected using statistical values ​​of the multiple differences corresponding to the multiple screens.

2. The eye-tracking system according to claim 1, wherein, The at least one processor sets the partial area as the guide area by adjusting the resolution of the screen in such a way that the resolution of the partial area becomes higher than the resolution of the areas outside the partial area.

3. The eye-tracking system according to claim 1 or 2, wherein, The at least one processor sets the partial region as the guiding region by blurring the region outside the partial region.

4. The eye-tracking system according to claim 1 or 2, wherein, The at least one processor defines the partial region as the guiding region by surrounding the outer edge of the partial region with a frame.

5. An eye-tracking method, wherein the eye-tracking method is executed by an eye-tracking system having at least one processor capable of executing an operating system and an application program. in, Includes the following steps: To draw the user's attention to a portion of the first content displayed on the screen, the following portions of the first content displayed on multiple screens after the application is started are dynamically set as their respective guide areas: these portions are areas in each of the multiple screens that display selectable objects that the user can choose. Based on the movement of the user's eyes in the guidance area when the selected object is selected in each of the multiple screens, the coordinates of the user's first viewpoint in each of the multiple screens are determined. For each of the plurality of frames, calculate the difference between the determined first viewpoint coordinates and the region coordinates of the guide area in the frame; The second viewpoint coordinates of the user viewing the second content displayed on the screen are corrected using statistical values ​​of the multiple differences corresponding to the multiple screens.

6. A recording medium having an eye-tracking program recorded thereon, the eye-tracking program causing a computer to perform the following steps: To encourage the user to focus on a portion of the first content displayed on the screen, the following portions of the first content displayed on multiple screens after the eye-tracking program has started are dynamically set as their respective guide areas: these portions are areas in each of the multiple screens that display selectable objects that the user can choose. Based on the movement of the user's eyes in the guidance area when the selected object is selected in each of the multiple screens, the first viewpoint coordinates of the user in each of the multiple screens are determined. For each of the plurality of frames, calculate the difference between the determined first viewpoint coordinates and the region coordinates of the guide area in the frame; The second viewpoint coordinates of the user viewing the second content displayed on the screen are corrected using statistical values ​​of the multiple differences corresponding to the multiple screens.

Citation Information

Patent Citations

  • Eye movement inspection device

    JP2001204692A

  • Visual line measuring device and visual line measuring program

    JP2010259605A

  • Program, device, and method for image processing

    JP2007087003A

  • Control device, method and surgery system

    JP2017176414A

  • Dynamic eye tracking calibration

    US20160139665A1