Calibration for gaze detection

By measuring head and eyeball rotation speeds and using implicit calibration, the system accurately performs gaze detection for users wearing head-mounted displays, overcoming the challenge of obscured eye visibility.

JP2026053370APending Publication Date: 2026-03-25FOVE INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Conventional gaze detection systems struggle to accurately perform calibration when users are wearing head-mounted displays, as the user's eye area is covered, making it impossible to confirm if they are looking at a specific indicator.

Method used

The system measures the rotational speed of the head and eyeballs in a certain direction and calibrates the gaze detection unit when these speeds are below a threshold, utilizing implicit calibration methods to perform gaze detection without explicit user interaction.

Benefits of technology

Enables accurate gaze detection for users wearing head-mounted displays by leveraging head and eyeball rotation speeds to calibrate the system, ensuring precise gaze direction estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053370000001_ABST
    Figure 2026053370000001_ABST
Patent Text Reader

Abstract

Precise calibration is performed to enable gaze detection of users wearing head-mounted displays. [Solution] The method comprises measuring the rotational speed of the head in a certain direction, measuring the rotational speed of the eyeballs in the same direction, and calibrating the gaze detection unit when the rotational speed of the head and the rotational speed of the eyeballs are below a threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video system, a video generation method, a video distribution method, a video generation program, and a video distribution program, particularly to a video system including a display attached to a head and a gaze detection device.

Background Art

[0002] Conventionally, when performing gaze detection to specify the point being viewed by a user, calibration is required. Here, calibration refers to having the user gaze at a specific indicator and specifying the positional relationship between the position where the specific indicator is displayed and the center of the user's cornea. A gaze detection system that performs calibration to execute gaze detection can identify the point being viewed by the user.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, the preparation for calibration is performed under the condition that the user is judged to be looking at a specific indicator. Therefore, there is a problem that when information is acquired without the user gazing at a specific indicator, actual gaze detection cannot be accurately performed. This problem is particularly prominent in the case of a head-mounted display where the user's eye area is covered by the device and the internal state cannot be seen, because it is impossible to confirm from the surroundings whether the user is actually looking at a specific indicator.

[0005] This invention has been made in consideration of the above-mentioned problems, and aims to provide a technology that can accurately perform calibration to realize gaze detection of a user wearing a head-mounted display. [Means for solving the problem]

[0006] To solve these problems, an aspect of the present invention is a method characterized by comprising the steps of: measuring the rotational speed of the head in a certain direction; measuring the rotational speed of the eyeballs in the same direction; and calibrating the gaze detection unit when the rotational speed of the head and the rotational speed of the eyeballs are lower than a threshold. [Effects of the Invention]

[0007] According to the present invention, it is possible to provide a technology for detecting the gaze direction of a user wearing a head-mounted display. [Brief explanation of the drawing]

[0008] [Figure 1] This is a schematic diagram of the video system 1 according to the first embodiment. [Figure 2] This is a block diagram showing the configuration of the video system 1 according to an embodiment. [Figure 3] This diagram shows the location of each component. [Figure 4] This is a flowchart of how to track eyes. [Figure 5] This shows the physical positions of the virtual camera and lens. [Figure 6] Camera images for lens shape are shown. [Figure 7] This shows a flowchart of the pupil prediction process based on a 3D model. [Figure 8] This figure shows an example of a scene image for calibration. [Figure 9] A flowchart of the hidden calibration process is shown. [Figure 10] A schematic diagram of the video system is shown. [Figure 11]Shows a flowchart of a process related to communication between a head-mounted display and a cloud server. [Figure 12] Shows a functional block diagram of a video system. [Figure 13] Shows another example of a functional block diagram of a video system. [Figure 14] Shows a graph indicating the rotation speeds of the head and eyes. [Figure 15] Shows the physical structure of the eyeball. [Figure 16] Shows an example of a calibration method for ACD. [Figure 17] Shows a refractive model for single-point calibration. [Figure 18] Shows a branch of implicit calibration. [Figure 19] Shows an overview of implicit calibration. [Figure 20] Shows a flowchart of implicit calibration.

Mode for Carrying Out the Invention

[0009] Hereinafter, each embodiment of the video system will be described with reference to the drawings. In the following description, the same components are denoted by the same reference numerals, and repeated descriptions are omitted.

[0010] Hereinafter, an overview of the first embodiment of the present invention will be described. FIG. 1 is a schematic diagram of a video system 1 according to the first embodiment. According to this embodiment, the video system 1 includes a head-mounted display 100 and a gaze detection device 200. As shown in FIG. 1, the head-mounted display 100 is used while being fixed to the head of the user 300.

[0011] The line-of-sight detection device 200 detects the line-of-sight direction of at least one of the right eye and the left eye of a user wearing the head-mounted display 100, and designates the focus of the user, that is, the point gazed at by the user within the three-dimensional image displayed on the head-mounted display. The line-of-sight detection device 200 also functions as a video generation device that generates a video to be displayed by the display 100 attached to the head. For example, the line-of-sight detection device 200 is a device that can play videos such as a stationary game machine, a portable game machine, a PC, a tablet, a smartphone, a phablet, a video player, a TV, etc., but the present invention is not limited thereto. The line-of-sight detection device 200 is connected to the head-mounted display 100 wirelessly or by wire. In the example shown in FIG. 1, the line-of-sight detection device 200 is wirelessly connected to the head-mounted display 100. The wireless connection between the line-of-sight detection device 200 and the head-mounted display 100 can be realized using known wireless communication technologies such as Wi-Fi (registered trademark) or Bluetooth (registered trademark). For example, the transfer of video between the head-mounted display 100 and the line-of-sight detection device 200 is executed according to standards such as Miracast (registered trademark), WiGig (registered trademark), WHDI (registered trademark). Other communication technologies can be used, for example, acoustic communication technology or optical transmission technology can be used.

[0012] The head-mounted display 100 comprises a housing 150, a mounting harness 160, and headphones 170. The housing 150 houses an image display system, such as an image display element for presenting video images to the user 300, and, although not shown in the figure, houses a Wi-Fi module, a Bluetooth® module, or another type of wireless communication module. The head-mounted display 100 is secured to the user 300's head with the mounting harness 160. The mounting harness 160 can be implemented, for example, with the help of a belt or elastic band. Once the user 300 secures the head-mounted display 100 with the mounting harness 160, the housing 150 is positioned to cover the user 300's eyes. Thus, when the user 300 wears the head-mounted display 100, the user 300's field of view is covered by the housing 150.

[0013] The headphones 170 output the audio of the video played back by the video generation device 200. The headphones 170 do not need to be fixed to the head-mounted display 100. Even if the head-mounted display 100 is fixed with the mounting harness 160, the user 300 can freely attach or detach the headphones 170.

[0014] Figure 2 is a block diagram showing the configuration of the video system 1 according to an embodiment.

[0015] The head-mounted display 100 includes a video display unit 110, an imaging unit 120, and a communication unit 130.

[0016] The video display unit 110 displays a video to the user 300. The video display unit 110 can be implemented, for example, as a liquid crystal monitor or an organic EL (electroluminescent) display.

[0017] The imaging unit 120 captures an image of the user's eye. The imaging unit 120 can be implemented as, for example, a CCD (charge-coupled device), CMOS (complementary metal-oxide-semiconductor) or other image sensor located within the housing 150.

[0018] The communication unit 130 provides a wireless or wired connection to the video generation device 200 for information transfer between the head-mounted display 100 and the video generation device 200. Specifically, the communication unit 130 transfers images captured by the imaging unit 120 to the video generation device 200 and receives video from the video generation device 200 for presentation by the video presentation unit 110. The communication unit 130 can be implemented, for example, as a Wi-Fi module, a Bluetooth® module, or another wireless communication module.

[0019] The gaze detection device 200 shown in Figure 2 is introduced. The gaze detection device 200 comprises a communication unit 210, a gaze detection unit 220, a calibration unit 230, and a storage unit 240.

[0020] The communication unit 210 provides a wireless or wired connection to the head-mounted display 100. The communication unit 210 receives images from the head-mounted display 100 captured by the imaging unit 120 and transmits video to the head-mounted display 100. The gaze detection unit 220 detects the gaze of the user viewing the image displayed on the display 100 and generates gaze data. The calibration unit 230 performs calibration of the gaze detection. The storage unit 240 stores the data for gaze detection and calibration.

[0021] <Eye-tracking using lens correction>

[0022] Eye-tracking using lens correction may include the following methods: The camera captures an image of the user's eyes. Find the reflected light from the image. Calculate the light rays from the camera to the reflected light. Light rays are transmitted as light rays that pass through a lens. The corneal center is located using transmitted light.

[0023] This method may further include the following: Find the pupil of the eye in the image. Calculate the second ray from the camera to the pupil. The second ray is transmitted as a second ray that passes through the lens. The position of the pupil is determined by the transmitted second ray.

[0024] Figure 3 shows a schematic diagram of eye-tracking with lens correction. Figure 3 shows a human eye, lens, virtual camera, and head-mounted display screen. Light rays from the camera pass through a standard lens or Fresnel lens and reach the human eye. The eye-tracking unit 220 uses the light rays to calculate eye tracking.

[0025] A standard lens or Fresnel lens is placed between the camera and the human eye. When detecting the direction of the eye's gaze, the gaze detection unit 220 uses the reflected light from the camera and the light rays to the pupil to detect the reflected light and pupil on the image of the human eye. In gaze tracking with lens correction, the light rays pass through the lens. Therefore, the gaze detection unit 220 must calculate such transmission.

[0026] The gaze detection unit 220 can calculate the rays from the camera to the position of the detected light (rays in front of the lens) using an internal matrix and an external matrix to assign a three-dimensional ray to any two-dimensional point (reflected light) on the camera image. To calculate the rays after the lens, the gaze detection unit 220 can apply Snell's law of ray tracing or use a pre-calculated transfer matrix. The gaze detection unit 220 uses these rays after the lens to calculate eye tracking (direction of gaze).

[0027] Lens correction can be performed using polynomial fitting. Let (x,y) represent a pixel on the camera image, (xp,yp) represent the xy position on the lens, and (xd,yd,zd) represent the xyz direction of the light ray from the lens. Next, for any pixel on the camera image, the gaze detection unit 220 can find the light ray after it has passed through the lens. JPEG2026053370000002.jpg32129JPEG2026053370000003.jpg33129JPEG2026053370 000004.jpg35141JPEG2026053370000005.jpg35141JPEG2026053370000006.jpg39152

[0028] Here, ai, bi, ci, di, ei, fi, gi, hi, pi, and qi are pre-calculated polynomial coefficients.

[0029] Note that (x,y) can be anything that can be directly derived from the pixel coordinates, such as the angle in spherical coordinates. Furthermore, (xd,yd,zd) can also have an alternative representation (e.g., spherical coordinates).

[0030] Figure 4 shows a flowchart of the eye-tracking method. The left side shows the conventional flow, and the right side shows eye-tracking with lens correction according to this embodiment.

[0031] First, the gaze detection unit 220 obtains an image of the eye from the camera. Then, the gaze detection unit 220 detects reflected light and the pupil by performing image processing. The gaze detection unit 220 uses internal and external matrices to obtain light rays from the camera for each light source.

[0032] In this eye-tracking system using lens correction, the eye-tracking detection unit 220 transmits light rays through the lens. The transmission is calculated using the matrix or polynomial fitting described above.

[0033] The gaze detection unit 220 solves the inverse problem to find the corneal center / radius.

[0034] Next, the gaze detection unit 220 uses internal and external matrices to obtain light rays from the camera to the pupil.

[0035] In our lens-corrected eye-tracking system, the eye-tracking unit 220 transmits this light ray through the lens.

[0036] The gaze detection unit 220 intersects this light ray with the sphere of the cornea.

[0037] The resulting intersection point is the 3D pupil position. The resulting optical axis is a vector from the corneal center to the 3D pupil position.

[0038] <Camera optimization through lens fitting>

[0039] Camera optimization through lens fitting may include the following methods: The camera captures an image of the user's eyes. The shape of the lens placed between the eye and the camera is detected. The camera's position and orientation are corrected so that the lens shape matches the expected lens shape.

[0040] Figure 5 shows the physical positions of the virtual camera and lens. When using such a lens, the expected position and orientation of the camera are of great importance in calculating the line of sight, as light rays from the camera to the user's eye are transmitted through the lens. Camera optimization through lens fitting adjusts the camera's position and orientation.

[0041] Figure 6 shows camera images for lens shape. The image on the left shows the expected camera image when the camera is oriented correctly. The image on the right shows the camera image when the camera is oriented incorrectly. As shown in the right image, the lens shape (white circle) is not in the center of the image.

[0042] The calibration unit 230 performs numerical optimization to correct the camera's position and orientation. As an optimization cost function, the calibration unit 230 attempts to fit the observed lens to the expected lens shape.

[0043] <Prediction of pupil, iris, and reflected light based on 3D models>

[0044] Predictions based on 3D models may include the following methods: The camera captures an image of the eye. The eye image is processed to obtain the position of the eye area. The eyeball model parameters are estimated based on the position of the eye area. The 3D line of sight direction is calculated based on the eyeball model parameters. Create a 3D eyeball model from eyeball model parameters. The following eyeball model parameters are estimated. The estimated eyeball model parameters are fed back into the image processing.

[0045] Figure 7 shows a flowchart of the pupil prediction process based on a 3D model. First, the eye-tracking system acquires an image of the eye via a camera. Next, it performs image processing based on the pupil and iris eccentricity, reflected light position, and the image of the eye from the camera. Then, it estimates eyeball model parameters such as the position and directional radius of the eyeball, pupil, and iris. It outputs a 3D gaze estimation. Next, it creates a 3D eyeball model from previous image frames and estimates the pupil and iris eccentricity and the reflected light position from the 3D model. Then, the pupil and iris eccentricity and the reflected light position are used in the next cycle of image processing.

[0046] <Hidden Calibration>

[0047] The calibration process requires additional effort from the user. Hidden calibration is performed while the user is viewing the content. Hidden calibration may include the following methods: Moving objects are displayed as engaging content in visual contexts to entertain users. The system uses a moving object as a calibration point to perform calibration of the user's line of sight. In hidden calibration, calibration is performed every time the scene changes.

[0048] Figure 8 shows an example of a scene image for calibration. The image on the left in Figure 8 shows a scream image of conventional calibration. In conventional calibration, moving dots are displayed on the screen before the content starts, and the user watches the dots. Furthermore, if recalibration is performed, the content must be stopped again to display the moving dots. However, stopping the content for calibration is stressful for the user. To address this problem, it is desirable to perform calibration without stopping the content.

[0049] For example, video content has a scene for a specific period of time that shows only moving objects on the screen, such as a logo, fireflies, and bright objects. During the displayed scene, the user looks at the moving objects, and the calibration unit can perform the calibration process. The right-hand diagram in Figure 8 shows an example of a scene displayed with fireflies.

[0050] If the video content contains multiple scenes, calibration can be performed multiple times within the content, gradually improving the accuracy of eye tracking.

[0051] Figure 9 shows a flowchart of the hidden calibration process. The application (such as a video player) draws a moving object on the screen without including any other content. Even if this calibration is not announced, the user's eyes are expected to track the moving object because only the object is displayed on the screen.

[0052] Next, the application sends the object's 3D position information (3D coordinates) to the eye-tracking unit.

[0053] Next, the eye-tracking unit performs calibration in real time using its location information. When the eye-tracking unit performs calibration, the application sends additional timestamp information along with the 3D location information.

[0054] <Foveal camera streaming>

[0055] Foveal camera streaming may include the following: acquiring an image to be displayed to the user; detecting the user's gaze direction; determining the user's region of interest on the image based on the gaze direction; compressing the region of interest of the image with a first compression ratio; compressing the area outside the region of interest of the image with a second compression ratio, the second being higher than the first; and transmitting the compressed region of interest and the compressed outer area. In this method, the resolution of the region of interest is higher than the resolution of the outer area.

[0056] In this method, the image may be a video, and in the step of encoding the region of interest, the outer region is compressed into a first video, and in the step of encoding the outer region, it is compressed into a second video, the frame rate of the first video is higher than the frame rate of the second video.

[0057] Foveal camera streaming may also be performed by methods including the following: Get the first image displayed to the user. Detects the user's gaze direction. Based on the direction of gaze, the user's region of interest on the first image is determined. Expand the region of interest to create a second image. Combine the first image and the second image. Combined image FA. Decode the combined image. The first and second images are separated from the combined image. Unenlarge the second image. The first and second images are processed.

[0058] Figure 10 shows a schematic diagram of the video system. In this embodiment, the video system comprises a head-mounted display 100, a gaze detection device 200, and a cloud server.

[0059] The head-mounted display 100 is further equipped with an external camera. The external camera is fixed in a housing 150 and positioned to record video images in the direction directly in front of the user's head. The external camera records video images of the entire world that it can record at full resolution. The video system has two image streams, one containing high-resolution images for the user's gaze area and another containing low-resolution images for other areas. Images, including the high-resolution and low-resolution images, are transmitted to a cloud server via a public communication network, either directly from the head-mounted display 100 or via the gaze detection device 200. In this technology, instead of transmitting full-resolution images of the entire world that the external camera can record, the video system 1 can reduce the bandwidth of video transmission by transmitting full-resolution images only to a limited area that the user sees (gaze area) and low-resolution images to other areas.

[0060] Based on the two types of image information received, the cloud server creates contextual information to be used in an AR (Augmented Reality) or MR (Mixed Reality) display. The cloud server aggregates information (e.g., object identification, face recognition, video images, etc.) to create the contextual information and transmits the contextual information to the head-mounted display 100.

[0061] Figure 11 shows a flowchart illustrating the communication process between the head-mounted display and the cloud server.

[0062] External cameras facing outwards from the head-mounted display capture images of the world (S1101).

[0063] Next, the control unit divides the video image into two streams based on the gaze tracking coordinates (S1102). In this step, the control unit detects the user's gaze point coordinates based on the gaze tracking coordinates and divides the video image into a region of interest and other regions. The region of interest can be obtained from the video image by dividing it into a region of a specific size that includes the gaze point.

[0064] Next, the two video image streams are sent to the cloud server via a communication network (e.g., a 5G network) (S1103). In this step, the image of the region of interest is sent to the server as a high-resolution image, while the image of the other region is sent to the server as a low-resolution image.

[0065] The cloud server then processes the image and adds contextual information (S1104).

[0066] Then, the image and context information are sent back to the head-mounted display, and the AR or MR image is displayed to the user (S1105).

[0067] Figure 12 shows the functional configuration diagram of the video system. The head-mounted display and gaze detection device include an external camera, a control unit, a gaze tracking unit, a sensing unit, a communication unit, and a display unit. The cloud server consists of a general recognition processing unit, a detailed processing unit, and an information aggregation unit.

[0068] The external camera acquires video images and inputs the resulting high-resolution raw video images to the control unit. The eye-tracking unit detects points (gaze coordinates) based on eye tracking and inputs gaze coordinate information to the control unit. The control unit determines the region of interest within each image based on the gaze coordinates. For example, the region of interest can be obtained from the video image by dividing it into a region of a specific size that includes the gaze point. The image data of the target region is compressed at a lower compression ratio and input to the communication unit. The communication unit also receives sensing data such as the tilt of the headset and other metadata obtained by the sensing unit. The sensing unit can be configured with a GPS or geomagnetic sensor. Image data of the region of interest is transmitted to the cloud server as a higher-resolution image. Image data outside the region of interest is compressed at a higher compression ratio and input to the communication unit. Image data outside the region of interest is transmitted to the cloud server as a lower-resolution image.

[0069] The general recognition processing unit of the cloud server receives low-resolution image data other than the "region of interest" (as well as headset tilt and metadata), and performs image processing to identify objects in the image (type, number, etc.).

[0070] The detailed processing unit on the cloud server receives high-resolution image data (and headset angle, metadata) of the region of interest and performs image processing to identify details such as face recognition and character recognition.

[0071] The information aggregation unit receives the identification results from the general recognition processing unit and the recognition results from the detailed processing unit. The information aggregation unit aggregates the received results to create a display image and transmits the display image to the head-mounted display via the communication network.

[0072] Figure 13 shows another example of a functional configuration diagram of the video system. In Figure 7-3, image data of the region of interest (high resolution) and image data outside the region of interest (low resolution) are sent separately to the cloud server. However, these image data can also be sent in a single video stream, as shown in Figure 7-4. After acquiring the region of interest, the control unit magnifies the image to reduce data outside the region of interest. The magnified image and sensing data are then sent from the sensing unit to the demagnetization unit in the cloud server. The demagnetization unit removes the magnification from the received image data and sends the demagnetized image data to the general recognition processing unit and the detail processing unit.

[0073] <Eye-tracking calibration using light response>

[0074] Eye-tracking calibration can be performed using the optical dynamic response. Specifically, the calibration method may include the following: measuring the head rotation speed in a certain direction; measuring the eyeball rotation speed in that direction; and performing calibration of the eye-tracking detection unit when the head rotation speed and eyeball rotation speed are below a threshold.

[0075] Ocular motility responses are eye movements that occur in response to the movement of an image on the retina. When looking at a point, the sum of the head rotation speed and the eye rotation speed is zero (0) during head rotation.

[0076] The calibration unit 230 can calibrate the gaze detection device 200 when the user fixates on a stable point that can be detected by detecting that the sum of the head rotation speed and eye rotation speed is zero. In other words, when the user rotates their head to the right, they should rotate their eyes to the left in order to fixate on a certain point.

[0077] Figure 14 shows a graph illustrating the rotational speeds of the head and eyes. The dotted line represents the eye rotational speed in the direction of rotation. The solid line represents the reverse head rotational speed (head rotational speed multiplied by -1). As shown in Figure 14, the reverse head rotational speed is approximately consistent with the eye rotational speed.

[0078] The head-mounted display 100 is equipped with an IMU. The IMU can measure the rotational speed of the user 300's head. The gaze detection unit can measure the rotational speed of the user's eyes. Eye rotational speed can be expressed as the speed of movement of the point of fixation. The calibration unit 230 can calculate the vertical and horizontal head rotational speeds from the values ​​measured by the IMU. The calibration unit 230 can also calculate the vertical and horizontal eye rotational speeds from the history of the point of fixation. The calibration unit 230 displays markers in a virtual space drawn on the display. The markers can be moved or stabilized. The calibration unit 230 calculates the horizontal and vertical head rotational speeds and the horizontal and vertical eye rotational speeds. The calibration unit 230 can perform calibration when the sum of the head rotational speed and eye rotational speed is lower than a predetermined threshold.

[0079] <Single-point calibration>

[0080] Single-point calibration may include the following methods: The pupil is imaged with a camera. The pupil position is corrected based on the depth of the anterior chamber. The direction of gaze is determined using the pupil correction position.

[0081] In the calibration method, the direction from the center of the cornea to the position of the pupil is determined as the line of sight.

[0082] In the calibration method, the direction from the center of the eyeball to the position of the pupil can be determined as the line of sight.

[0083] The calibration method may further include correcting the position of the pupil to an angle with respect to the direction from the camera to the pupil image.

[0084] Figure 15 shows the physical structure of the eyeball. The eyeball consists of several parts, including the pupil, cornea, and anterior chamber. The position of the pupil can be recognized by a camera image. In reality, there is anterior chamber depth (ACD) between the corneal surface and the pupil. Therefore, in order to improve the accuracy of gaze estimation, it is necessary to correct the pupil position by taking the ACD into account. The gaze direction is estimated using the corrected pupil position.

[0085] Figure 16 shows an example of a calibration method. In this case, the eye is looking at a calibration point known by the system, and the pupil is observed by a camera. P0 indicates the intersection of the light ray (the pupil observed by the camera) and the corneal sphere. P0 is the observed pupil on the camera. P0 is used for general gaze estimation.

[0086] However, in reality, the pupil is located at P1 within the corneal bulb according to the ACD. The direction from the center of the eyeball (or the center of the corneal bulb) to the center of the pupil is considered the line of sight of the eye. Calibration can be performed using the line of sight and known calibration points.

[0087] However, in reality, the pupil is located at P1 within the corneal bulb according to the ACD. The direction from the center of the eyeball (or the center of the corneal bulb) to the center of the pupil is considered the line of sight of the eye. Calibration can be performed using the line of sight and known calibration points.

[0088] <Refraction Model>

[0089] If the cornea refracts light rays, the correction is adjusted accordingly. Specifically, the pupil position is corrected by considering the anterior chamber depth (ACD) and the horizontal optical axis relative to the visual axis.

[0090] Figure 17 shows a refractive model for single-point calibration. P0 is obtained from the intersection of the light ray (the pupil observed by the camera) and the corneal sphere. We know the position of the cornea and the incident light ray. Therefore, we apply "Snell's Law" (or other refractive models). This changes the direction. The light ray is continued such that there is a constant distance (ACD) between the pupil P1 and the corneal sphere (P0). The direction from the corneal center (or ocular center) to P1 should be directed towards a known calibration point. If not, the calibration unit 230 optimizes the ACD. Therefore, the calibration unit 230 calibrates the ACD.

[0091] <Implicit Calibration>

[0092] The calibration process requires additional effort from the user. Through implicit calibration, the calibration is performed while the user is viewing the content.

[0093] In conventional (explicit) calibration, the following occurs: 1. The system displays a point target at a location known to the user. 2. Users need to verify their target for a certain period of time. 3. The system records the estimated user gaze during that period. 4. This system estimates eye-tracking parameters by combining recorded line-of-sight data with ground verification locations.

[0094] On the other hand, with implicit calibration, the following occurs: 1. There is no explicit point target. 2. Users do not need to take any specific actions, as they will be engaged in a normal VR / AR experience. 3. The system records estimated user gaze and head-mounted screen images during a normal VR / AR experience. 4. This system combines the estimated line of sight with the screen image to obtain the ground verification position. 5. This system combined recorded line-of-sight data with ground verification locations to estimate eye-tracking parameters.

[0095] Implicit calibration may be a method for calibrating gaze detection, including the following: Obtain an image of the user's eyes. The system detects the point of fixation based on the image of the eye. This process captures images of the user's field of view while they are viewing scenes that do not contain pre-identified content. The system detects edges in sub-images extracted from the field of view image and sub-images containing the point of fixation. Adjust the point of focus according to the detected edge.

[0096] In implicit calibration, the point at which the focus of attention is adjusted may be the point with the highest probability of an edge occurring.

[0097] Implicit calibration may further include the following: Accumulate the detected edges over a predetermined period. Calculate statistics from the edge distribution. After a predetermined period, the focus of attention is adjusted based on statistical data.

[0098] It is important to add that a screen image is required. This screen image is called the field of view. The field of view can cover both the screen image in VR and the external camera image in AR. It is also important to note that the user does not need to look at a specific target. The content of the scene can be arbitrary. Since there are two types of images, the eye image and the screen image, it is necessary to explicitly specify which image to refer to.

[0099] The correlation between the point of fixation and the field of view image can be defined as an assumption about human behavior. That is, in situation A, A and B can be automatically extracted from the field of view image, and it is more likely that the person will look at B. Currently, we are using the assumption that in all situations, people are more likely to look at the edge.

[0100] For example, there are other potentially usable assumptions. When a new image is presented, people are likely to see the faces of humans / animals / etc. first, or when presented with a video of an object moving against a stationary background, people are more likely to see the moving object. We do the following: 1. Calculate the average edge over time. 2. Find the vector to the point with the maximum average edge.

[0101] The general nature of the idea is to "automatically extract positive data from a field of view image." However, the positive data that underlies this is not a single point, but a probability distribution. If there is only one field of view image, the actual point of fixation cannot be predicted, but it can be predicted that the actual point of fixation will be located in a specific region with a certain probability.

[0102] When selecting a region of interest, the fixation point probability distribution is essentially transformed into a probability distribution of the eye-tracking parameters. After accumulating probability distributions from the field of view images at different time moments, the mean probability distribution can be calculated. This mean probability distribution gradually converges to a single point (a single value of the eye-tracking parameters). In other words, the larger the number of images, the smaller the standard deviation of this distribution becomes.

[0103] The general idea of ​​predicting gaze from images of the field of view is not new. There is a field of study called "saliency prediction" that investigates this topic. The hypothesis that "humans are more likely to look at the edges of objects" also stems from saliency prediction. Integrating saliency prediction into eye-tracking calibration is a new approach.

[0104] Figure 18 shows the branching of implicit calibration. "Bias" means that the difference between the estimated point of fixation and the actual point of fixation is constant. Humans tend to see points with high contrast, i.e., points with edges of objects. Therefore, in implicit calibration using edge accumulation, the following occurs: 1. Obtain estimated user gaze and images from the head-mounted display screen. 2. Select a small area (ROI; region of interest) of the screen image around the point of focus. 3. Detect edges on the ROI. 4. Accumulate statistics over time. 5. Find the point where the edge is maximized over time. 6. Estimate the bias as the difference between the ROI center and the maximum point.

[0105] Figure 19 shows an overview of implicit calibration. Circles represent gaze points from the eye-tracker (gaze detection unit 220). Stars represent actual gaze points. Rectangles represent the cumulative area of ​​the field of view. The authors used the point with the maximum number of edges as positive data for calibration. Therefore, it is not necessary to provide calibration points in the display field.

[0106] Figure 20 shows a flowchart of implicit calibration. The eye-tracker (eye-detection unit 220) provides an approximate gaze direction. The head-mounted display provides an image of the user's entire field of view. Using the approximate gaze direction and the image of the entire field of view, the calibration unit 230 calculates statistics over time, estimates the eye-tracking parameters, and feeds the parameters back to the eye-tracker. In this way, the gaze direction is gradually calibrated.

Claims

1. Measuring the speed of head rotation in a certain direction, Measuring the rotational speed of the eyeball in the aforementioned direction, Calibration of the gaze detection unit is performed when the head rotation speed and the eyeball rotation speed are below a threshold. A method for providing this.

2. Retrieving the image to display to the user, The detection of the user's gaze direction, Based on the aforementioned line of sight direction, the user's area of ​​interest in the image is determined, Compressing the region of interest in the image with a first compression ratio, Compressing the area outside the region of interest of the image with a second compression ratio higher than the first compression ratio, Transmitting the compressed region of interest and the compressed outer region, A method for providing this.

3. The method according to claim 2, wherein the resolution of the region of interest is higher than the resolution of the outer region.

4. The method according to claim 2, wherein in the step of compressing the region of interest, the outer region is encoded into a first video, and in the step of compressing the outer region, the outer region is encoded into a second video, wherein the frame rate of the first video is higher than the frame rate of the second video.

5. To obtain the first image displayed to the user, The detection of the user's gaze direction, Based on the aforementioned line of sight direction, the user's region of interest in the first image is determined, Stretching the aforementioned region of interest into a second image, The first and second images described above are combined, The aforementioned synthesized image is transmitted, Decoding the aforementioned synthesized image, The process of separating the first and second images from the synthesized image, To remove the stretching of the second image, Processing the first and second images, A method for providing this.

6. Obtaining an image of the user's eyes, Detecting the viewpoint from the aforementioned eye image, This involves acquiring an image of the user's field of view while viewing a scene that does not necessarily contain predetermined content, The process involves detecting edges in a partial image extracted from the image of the field of view, wherein the partial image includes the viewpoint, and the edge detection is performed accordingly. Adjusting the viewpoint according to the detected edge, A method for providing this.

7. The method according to claim 6, wherein the point on which the fixation point is adjusted is the point with the highest probability of an edge occurring.

8. Accumulating edges detected over a predetermined period, Calculating statistics from the distribution of the aforementioned edges, After the predetermined period, the perspective is adjusted according to the statistical quantity, The method according to claim 6, further comprising:

9. Taking a picture of the pupil with a camera, Correcting the position of the pupil based on the depth of the anterior chamber, The direction of gaze is determined using the corrected position of the pupil, A method for providing this.

10. The method according to claim 9, further comprising determining the direction from the center of the cornea to the position of the pupil as the line of sight.

11. The method according to claim 9, wherein in the step of determining the gaze direction, the direction from the center of the eyeball to the position of the pupil is determined as the gaze direction.

12. The method according to claim 9, further comprising the step of correcting the position of the pupil in the direction from the camera to the pupil image.

13. Displaying moving objects in a visual context as content that entertains the user, Using the aforementioned moving object as a calibration point, the user's line of sight is calibrated. A method for providing this.

14. The method according to claim 13, wherein the calibration is performed each time the scene changes.

15. Acquiring an image of the user's eyes from the camera, To detect the shape of the lens placed between the eye and the camera, To correct at least one of the position and orientation of the camera so that the shape of the lens conforms to the planned lens shape, A method for providing this.

16. Acquiring an image of the user's eyes from the camera, To find the light above the eye from the aforementioned image, Calculating the light ray from the camera to the light, Transmitting the aforementioned light ray as a light ray passing through a lens, Using the transmitted light ray, the corneal center is located, A method for providing this.

17. To locate the pupil of the eye in the aforementioned image, Calculating a second light ray from the camera to the pupil, The second ray is transmitted as the second ray that has passed through the lens, Using the aforementioned second technical college that was transmitted, find one of the aforementioned schools, The method according to claim 16, further comprising:

18. Acquiring an image of the eye from the camera, To obtain the position of the eye area, the image of the eye is processed using image processing, Estimating eyeball model parameters based on the aforementioned position of the eye portion, Calculating the 3D line of sight direction based on the aforementioned eyeball model parameters, Creating a 3D eyeball model from the aforementioned eyeball model parameters, To estimate the following eyeball model parameters, The estimated eyeball model parameters are fed back into the image processing, A method for providing this.

Citation Information

Patent Citations

  • Head-mounted display and program used therefor

    JP2012216123A