Calibration for gaze detection
By employing head and eye rotation speed measurements and advanced calibration techniques, the method addresses the inaccuracy of gaze detection in head-mounted displays, ensuring precise gaze tracking without additional user effort.
Patent Information
- Application Number
- JP2023546576
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-12
- Filing Date
- 2021-10-12
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-10-12
AI Technical Summary
Calibration for gaze detection in head-mounted displays is inaccurate due to the user's eyes being covered, making it impossible to confirm if they are looking at a specific index, especially in environments where the surroundings cannot be seen.
A method involving measuring head and eye rotation speeds, and calibrating the gaze detection unit when these speeds are below a threshold, along with techniques like lens correction, 3D modeling, covert calibration, foveated camera streaming, and implicit calibration to enhance accuracy.
Enables accurate gaze detection in head-mounted displays by utilizing head and eye rotation speeds, improving calibration methods, and reducing the need for user intervention, thereby enhancing the precision and efficiency of gaze tracking.
Smart Images

Figure 0007790749000006 
Figure 0007790749000007 
Figure 0007790749000008
Abstract
Description
[Technical Field]
[0001] The present invention relates to a video system, a video generation method, a video distribution method, a video generation program, and a video distribution program, particularly relating to a video system having a head-mounted display and a gaze detection device. [Background technology]
[0002] Conventionally, when performing gaze detection to specify a point where a user is looking, calibration is required. Here, calibration refers to having a user gaze at a specific indicator and specifying the positional relationship between the position where the specific indicator is displayed and the center of the user's cornea. A gaze detection system that performs calibration to perform gaze detection can identify the point where a user is looking. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-216123 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the calibration preparation is performed under the condition that it is determined that the user is looking at a specific index. Therefore, if information is acquired when the user is not gazing at a specific index, there is a problem that the actual gaze detection cannot be performed accurately. This problem is particularly noticeable in the case of a head-mounted display, where the area around the user's eyes is covered by a device and the internal state cannot be seen, because it is impossible to confirm from the surroundings whether the user is actually looking at a specific index.
[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a technology that can accurately perform calibration to detect the gaze of a user wearing a head-mounted display. [Means for solving the problem]
[0006] To solve such problems, an aspect of the present invention is a method characterized by comprising a step of measuring the head rotation speed in a certain direction, a step of measuring the eye rotation speed in said direction, and a step of calibrating a gaze detection unit when the head rotation speed and eye rotation speed are lower than a threshold value. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide a technique for detecting the gaze direction of a user wearing a head-mounted display. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a schematic diagram of a video system 1 according to a first embodiment. [Figure 2] 1 is a block diagram showing the configuration of a video system 1 according to an embodiment. [Figure 3] FIG. [Figure 4] 1 is a flowchart of a method for tracking eyes. [Figure 5] Indicates the physical location of the virtual camera and lens. [Figure 6] 1 shows a camera image for the lens shape. [Figure 7] 1 shows a flowchart of a process for pupil prediction based on a 3D model. [Figure 8] FIG. 10 is a diagram showing an example of a scene image for calibration. [Figure 9] 1 shows a flowchart of a process for hidden calibration. [Figure 10] 1 shows a schematic diagram of a video system. [Figure 11]1 shows a flowchart of a process for communication between a head-mounted display and a cloud server. [Figure 12] A functional configuration diagram of a video system is shown. [Figure 13] 10 shows another example of a functional configuration diagram of a video system. [Figure 14] 1 shows a graph illustrating head and eye rotation velocities. [Figure 15] Shows the physical structure of the eyeball. [Figure 16] An example of a method for calibrating an ACD is shown below. [Figure 17] 1 shows a refraction model for single point calibration. [Figure 18] 1 shows the branching of implicit calibration. [Figure 19] An overview of implicit calibration is given below. [Figure 20] 1 shows a flowchart of implicit calibration. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the video system will be described with reference to the accompanying drawings. In the following description, the same components are denoted by the same symbols, and repeated description will be omitted.
[0010] An overview of a first embodiment of the present invention will be described below. Fig. 1 is a schematic diagram of a video system 1 according to the first embodiment. According to this embodiment, the video system 1 includes a head-mounted display 100 and a gaze detection device 200. As shown in Fig. 1, the head-mounted display 100 is used while being fixed to the head of a user 300.
[0011] The gaze detection device 200 detects the gaze direction of at least one of the right and left eyes of a user wearing the head-mounted display 100 and specifies the user's focal point, i.e., the point at which the user gazes within the three-dimensional image displayed on the head-mounted display. The gaze detection device 200 also functions as a video generation device that generates video to be displayed by the head-mounted display 100. For example, the gaze detection device 200 may be a device capable of playing video, such as a stationary game console, a portable game console, a PC, a tablet, a smartphone, a phablet, a video player, or a television, but the present invention is not limited thereto. The gaze detection device 200 is connected to the head-mounted display 100 wirelessly or wirelessly. In the example shown in FIG. 1, the gaze detection device 200 is connected to the head-mounted display 100 wirelessly. The wireless connection between the gaze detection device 200 and the head-mounted display 100 can be realized using known wireless communication technologies such as Wi-Fi (registered trademark) or Bluetooth (registered trademark). For example, video transmission between the head mounted display 100 and the gaze detection device 200 is performed according to standards such as Miracast (registered trademark), WiGig (registered trademark), WHDI (registered trademark), etc. Other communication technologies may be used, such as acoustic communication technology or optical transmission technology.
[0012] The head-mounted display 100 includes a housing 150, a mounting harness 160, and headphones 170. The housing 150 houses an image display system, such as an image display element, for presenting a video image to the user 300. Although not shown, the housing 150 may also house a Wi-Fi module, a Bluetooth module, or another type of wireless communication module. The head-mounted display 100 is secured to the head of the user 300 with the mounting harness 160. The mounting harness 160 may be implemented with the aid of, for example, a belt or an elastic band. When the user 300 secures the head-mounted display 100 with the mounting harness 160, the housing 150 is positioned so that the user's 300's eyes are covered. Therefore, when the user 300 wears the head-mounted display 100, the user's field of vision is covered by the housing 150.
[0013] The headphones 170 output the audio of the video played by the video generation device 200. The headphones 170 do not need to be fixed to the head-mounted display 100. Even if the head-mounted display 100 is fixed with the attachment harness 160, the user 300 can freely attach or detach the headphones 170.
[0014] FIG. 2 is a block diagram showing the configuration of the video system 1 according to the embodiment.
[0015] The head mounted display 100 includes a video presentation unit 110, an imaging unit 120, and a communication unit .
[0016] The video presenting unit 110 presents a video to the user 300. The video presenting unit 110 can be implemented as, for example, a liquid crystal monitor or an organic EL (electroluminescence) display.
[0017] Imager 120 captures an image of the user's eye and may be implemented, for example, as a CCD (charge-coupled device), CMOS (complementary metal-oxide semiconductor), or other image sensor disposed within housing 150.
[0018] The communication unit 130 provides a wireless or wired connection to the video generation device 200 for information transfer between the head mounted display 100 and the video generation device 200. Specifically, the communication unit 130 transfers images captured by the imaging unit 120 to the video generation device 200 and receives video from the video generation device 200 for presentation by the video presentation unit 110. The communication unit 130 can be implemented, for example, as a Wi-Fi module, a Bluetooth (registered trademark) module, or other wireless communication module.
[0019] 2, the gaze detection device 200 includes a communication unit 210, a gaze detection unit 220, a calibration unit 230, and a storage unit 240.
[0020] The communication unit 210 provides a wireless or wired connection to the head mounted display 100. The communication unit 210 receives images of the head mounted display 100 captured by the imaging unit 120 and transmits video to the head mounted display 100. The gaze detection unit 220 detects the gaze of a user looking at an image displayed on the display 100 and generates gaze data. The calibration unit 230 calibrates gaze detection. The memory unit 240 stores data for gaze detection and calibration.
[0021] <Eye tracking with lens correction>
[0022] Eye tracking with lens correction may be a method that includes: An image of the user's eyes is acquired from the camera. Find the reflected light in the eyes from the image. Calculate the ray from the camera to the reflected light. The light beam is transmitted as a beam through a lens. The center of the cornea is found using transmitted light.
[0023] The method may further include: Find the eye pupil from the image. Calculate the second ray from the camera to the pupil. The second ray is transmitted as a second ray through the crystalline lens. The transmitted second ray finds the pupil position.
[0024] Figure 3 shows a schematic diagram of gaze tracking with lens correction. Figure 3 shows a human eye, a lens, a virtual camera, and a screen of a head-mounted display. A ray of light from the camera passes through a standard or Fresnel lens and reaches the human eye. The gaze detection unit 220 uses the ray of light to calculate eye tracking.
[0025] A standard lens or a Fresnel lens is provided between the camera and the human eye. When detecting the gaze direction of the eye, the gaze detection unit 220 detects the reflected light and pupil on the image of the human eye using light rays from the camera to each reflected light and pupil. In gaze tracking with lens correction, the light rays pass through the lens. Therefore, the gaze detection unit 220 must calculate such transmission.
[0026] The gaze detection unit 220 can calculate a ray from the camera (a ray before the lens) from the image to the detected light position using an intrinsic matrix and an extrinsic matrix to give a 3D ray for any 2D point (reflected light) on the camera image. The gaze detection unit 220 can apply Snell's law ray tracing or use a pre-calculated transfer matrix to calculate the ray after the lens. The gaze detection unit 220 uses this ray after the lens to calculate eye tracking (gaze direction).
[0027] Lens correction can be performed using polynomial fitting. Let (x,y) represent a pixel on the camera image, (xp,yp) represent the xy position on the lens, and (xd,yd,zd) represent the xyz direction of the ray from the lens. Then, for any pixel on the camera image, the gaze detection unit 220 can find the ray after passing through the lens. JPEG0007790749000001.jpg32129JPEG0007790749000002.jpg33129JPEG0007790749 000003.jpg35141JPEG0007790749000004.jpg35141JPEG0007790749000005.jpg39152
[0028] where ai, bi, ci, di, ei, fi, gi, hi, pi, qi are pre-computed polynomial coefficients.
[0029] Note that (x,y) can be anything that can be derived directly from pixel coordinates, such as an angle in spherical coordinates. Additionally, (xd,yd,zd) can also have an alternative representation (e.g., spherical coordinates).
[0030] 4 shows a flowchart of the gaze tracking method. The left side shows the conventional flow, and the right side shows the gaze tracking with lens correction according to this embodiment.
[0031] First, the gaze detection unit 220 obtains an image of the eye from the camera. Then, the gaze detection unit 220 detects the reflected light and pupil by processing the image. The gaze detection unit 220 uses internal and external matrices to obtain the light rays from the camera to each light.
[0032] Here, in lens-corrected gaze tracking, the gaze detection unit 220 transmits rays through the lens, which is calculated by matrix or polynomial fitting as described above.
[0033] The gaze detection unit 220 solves the inverse problem to find the corneal center / radius.
[0034] The gaze detector 220 then uses the intrinsic and extrinsic matrices to obtain the ray from the camera to the pupil.
[0035] In our lens-corrected gaze tracking, the gaze detector 220 transmits this ray through the lens.
[0036] The line of sight detection unit 220 causes this ray to intersect with the corneal sphere.
[0037] The resulting intersection point is the 3D pupil position. The resulting optical axis is the vector from the corneal center to the 3D pupil position.
[0038] <Camera optimization through lens fitting>
[0039] Camera optimization by lens fitting may be a method that includes: An image of the user's eyes is acquired from the camera. Detects the shape of the lens placed between the eye and the camera. At least one of the position and orientation of the camera is corrected so that the shape of the lens matches an expected lens shape.
[0040] Figure 5 shows the physical location of the virtual camera and lens. When using such a lens, the expected position and orientation of the camera are crucial for calculating the gaze direction, since the light rays from the camera to the user's eyes travel through the lens. Camera optimization through lens fitting involves adjusting the camera's position and orientation.
[0041] Figure 6 shows the camera images for the lens shape. The picture on the left shows the expected camera image when the camera orientation is correct. The picture on the right shows the camera image when the camera orientation is incorrect. As you can see in the right image, the lens shape (white circle) is not in the center of the image.
[0042] The calibration unit 230 performs a numerical optimization to correct the camera position and orientation. As an optimization cost function, the calibration unit 230 tries to match the observed lens to the expected lens shape.
[0043] <Pupil, iris, and reflected light prediction based on 3D models>
[0044] Prediction based on 3D models can be methods including: Acquire an image of the eye from the camera. The eye image is image processed to obtain the location of the eye parts. Eye model parameters are estimated based on the positions of the eye parts. Calculate the 3D gaze direction based on the eye model parameters. A 3D eye model is created from the eye model parameters. The following eye model parameters are estimated: The estimated next eye model parameters are fed back to image processing.
[0045] Figure 7 shows a flowchart of the process of pupil prediction based on a 3D model. First, the gaze tracking system acquires an image of the eye by a camera. Then, it performs image processing based on the pupil and iris eccentricity, the reflected light position, and the eye image from the camera. Then, it estimates eye model parameters such as the position and directional radius of the eyeball, pupil, and iris. It outputs a 3D gaze estimation. Next, it creates a 3D eye model from the previous image frame, and estimates the pupil and iris eccentricity and the reflected light position from the 3D model. Then, it uses the pupil and iris eccentricity and the reflected light position for the next cycle of image processing.
[0046] <Hidden Calibration>
[0047] The calibration process imposes additional effort on the user. With hidden calibration, the calibration is performed while the user is watching the content. Covert calibration may be a method that includes: To display moving objects in a visual scene as content that entertains users. A moving object is used as a calibration point to perform a calibration of the user's gaze direction. In hidden calibration, a calibration is performed every time the scene changes.
[0048] FIG. 8 shows an example of a scene image for calibration. The left image in FIG. 8 shows a scream image for conventional calibration. In conventional calibration, before content starts, moving dots are displayed on the screen so that the user can see the dots. Furthermore, when recalibrating, the content needs to be stopped to display the moving dots again. However, stopping the content for calibration causes stress to the user. To address this issue, it is desirable to perform calibration without stopping the content.
[0049] For example, a video content may have a scene for a specific time that shows only moving objects on the screen, such as a logo, a firefly, and a bright object. During the displayed scene, a user may view the moving objects, and the calibration unit may perform a calibration process. The right diagram of Figure 8 shows an example of a scene displayed with a firefly.
[0050] If the video content has multiple scenes, calibration can be performed multiple times during the content, gradually improving the accuracy of the eye-tracking.
[0051] Figure 9 shows a flowchart of the process of covert calibration. An application (such as a video player) draws a moving object on the screen without any other content. Even without this calibration being announced, the user's eyes are expected to track the moving object because only the object is visible on the screen.
[0052] Next, the application sends the object's 3D position information (3D coordinates) to the eye tracking unit.
[0053] The eye tracker then uses the position information to calibrate in real time. When the eye tracker calibrates, the application sends additional timestamp information along with the 3D position information.
[0054] <Foveated Camera Streaming>
[0055] Foveated camera streaming may be a method that includes: acquiring an image to display to a user; detecting a gaze direction of the user; determining a region of interest of the user on the image based on the gaze direction; compressing the region of interest of the image at a first compression rate; compressing an outer region of the image other than the region of interest at a second compression rate, the second compression rate being higher than the first compression rate; and transmitting the compressed region of interest and the compressed outer region, wherein in this method the resolution of the region of interest is higher than the resolution of the outer region.
[0056] In this method, the image may be a video, and in the step of encoding the region of interest, the external region is compressed into a first video, and in the step of encoding the external region, the image is compressed into a second video, and the frame rate of the first video is higher than the frame rate of the second video.
[0057] Foveated camera streaming may also be a method that includes the following. Gets the initial image that is displayed to the user. Detect the user's gaze direction. A region of interest of the user on the first image is determined based on the gaze direction. The region of interest is enlarged into a second image. The first image and the second image are combined. Send the combined image. Decode the combined image. Separate the first image and the second image from the combined image. Unzoom the second image. The first image and the second image are processed.
[0058] 10 shows a schematic diagram of a video system. In this embodiment, the video system includes a head-mounted display 100, a gaze detection device 200, and a cloud server.
[0059] The head-mounted display 100 further includes an external camera. The external camera is fixed to the housing 150 and positioned to record video images in a frontal direction of the user's head. The external camera records video images of the entire world that the external camera can record at full resolution. The video system has two image streams, including high-resolution images for the user's gaze area and low-resolution images for other areas. The images, including the high-resolution images and the low-resolution images, are transmitted from the head-mounted display 100 directly or via the gaze detection device 200 to a cloud server over a public communication network. In this technology, instead of transmitting full-resolution images of the entire world that the external camera can record, the video system 1 transmits full-resolution images only for a limited area (gaze area) that the user views and low-resolution images for other areas, thereby reducing the video transmission bandwidth.
[0060] Based on the received two types of image information, the cloud server creates context information to be used for an AR (Augmented Reality) or MR (Mixed Reality) display. The cloud server aggregates information (e.g., object identification, face recognition, video images, etc.) to create the context information and transmits the context information to the head-mounted display 100.
[0061] FIG. 11 shows a flowchart of the process for communication between the head mounted display and the cloud server.
[0062] An external camera facing outward from the head-mounted display captures an image of the world (S1101).
[0063] Next, the control unit divides the video image into two streams based on the eye tracking coordinates (S1102). In this step, the control unit detects the user's gaze point coordinates based on the eye tracking coordinates, and divides the video image into a region of interest and other regions. The region of interest can be obtained from the video image by dividing a region of a certain size that includes the gaze point.
[0064] Next, the two video image streams are transmitted to a cloud server via a communication network (e.g., a 5G network) (S1103). In this step, an image of the region of interest is sent to the server as a high-resolution image, while an image of the other region is sent to the server as a low-resolution image.
[0065] The cloud server then processes the image and adds context information (S1104).
[0066] The image and context information are then sent back to the head-mounted display, and the AR or MR image is displayed to the user (S1105).
[0067] Figure 12 shows the functional configuration of the video system. The head-mounted display and gaze detection device include an external camera, control unit, gaze tracking unit, sensing unit, communication unit, and display unit. The cloud server consists of a general recognition processing unit, detailed processing unit, and information aggregation unit.
[0068] The external camera captures video images and inputs the resulting high-resolution raw video images to the control unit. The gaze tracking unit detects points (gaze coordinates) based on gaze tracking and inputs the gaze coordinate information to the control unit. The control unit determines a region of interest within each image based on the gaze coordinates. For example, the region of interest can be obtained from the video image by dividing it into a region of a specific size that includes the gaze point. Image data of the region of interest is compressed at a lower compression ratio and input to the communication unit. The communication unit also receives sensing data, such as headset tilt and other metadata, obtained by the sensing unit. The sensing unit may be configured with a GPS or geomagnetic sensor. Image data of the region of interest is transmitted to a cloud server as a higher-resolution image. Image data outside the region of interest is compressed at a higher compression ratio and input to the communication unit. Image data outside the region of interest is transmitted to the cloud server as a lower-resolution image.
[0069] The general recognition processing unit of the cloud server receives low-resolution image data (as well as headset tilt and metadata) other than the "region of interest" and performs image processing to identify objects in the image (type, number, etc. of objects).
[0070] The detail processing section of the cloud server receives high-resolution image data of the region of interest (as well as headset angle and metadata) and performs image processing to identify details such as facial recognition and character recognition.
[0071] The information aggregation unit receives the identification result from the general recognition processing unit and the recognition result from the detailed processing unit, aggregates the received results to create a display image, and transmits the display image to the head-mounted display via a communication network.
[0072] Figure 13 shows another example of a functional configuration diagram of a video system. In Figure 7-3, image data of the region of interest (high resolution) and image data outside the region of interest (low resolution) are sent separately to the cloud server. However, these image data can also be sent in a single video stream, as shown in Figure 7-4. After acquiring the region of interest, the control unit enlarges the image to reduce the data outside the region of interest. The enlarged image and sensing data are then sent from the sensing unit to the de-enlargement unit in the cloud server. The de-enlargement unit de-enlarges the received image data and sends the de-enlarged image data to the general recognition processing unit and the detailed processing unit.
[0073] <Eye tracking calibration using light response>
[0074] Eye tracking calibration can be performed using photodynamic responses. That is, the calibration method includes the following: measuring the head rotation speed in a certain direction; measuring the eye rotation speed in the same direction; and calibrating the gaze detection unit if the head rotation speed or eye rotation speed is below a threshold.
[0075] Oculomotor responses are eye movements that occur in response to a moving image on the retina. When looking at a point, the sum of head rotation velocity and eye rotation velocity is zero (0) during head rotation.
[0076] The calibration unit 230 can calibrate the gaze detection device 200 when the user gazes at a stable point that can be detected by detecting that the sum of the head rotation speed and the eye rotation speed is zero. That is, when the user rotates his / her head to the right, he / she should rotate his / her eyes to the left to gaze at a certain point.
[0077] Figure 14 shows a graph illustrating head and eye rotation velocities. The dotted line shows the directional eye rotation velocity. The solid line shows the inverse head rotation velocity (head rotation velocity multiplied by -1). As shown in Figure 14, the inverse head rotation velocity is approximately aligned with the eye rotation velocity.
[0078] The head-mounted display 100 includes an IMU. The IMU can measure the rotation speed of the user's head. The gaze detection unit can measure the rotation speed of the user's eyes. The eye rotation speed can be expressed as the movement speed of the gaze point. The calibration unit 230 can calculate the head rotation speed in the up / down and left / right directions from the values measured by the IMU. The calibration unit 230 can also calculate the eye rotation speed in the up / down and left / right directions from the gaze point history. The calibration unit 230 displays a marker in the virtual space depicted on the display. The marker can be moved or stabilized. The calibration unit 230 calculates the head rotation speed in the left / right and up / down directions and the eye rotation speed in the left / right and up / down directions. The calibration unit 230 can perform calibration when the sum of the head rotation speed and the eye rotation speed is lower than a predetermined threshold.
[0079] <Single-point calibration>
[0080] The single point calibration may be a method that includes the following. The pupil is photographed with a camera. Correct pupil position based on anterior chamber depth. The corrected position of the pupil is used to determine the gaze direction.
[0081] In the calibration method, the direction from the center of the cornea to the position of the pupil is determined as the gaze direction.
[0082] In the calibration method, the direction from the center of the eyeball to the position of the pupil can be determined as the gaze direction.
[0083] The calibration method may further include correcting the pupil position to an angle relative to the direction from the camera to the pupil image.
[0084] Figure 15 shows the physical structure of the eyeball. The eyeball is composed of several parts, including the pupil, cornea, and anterior chamber. The position of the pupil can be recognized from a camera image. In reality, there is an anterior chamber depth (ACD) between the corneal surface and the pupil. Therefore, to improve the accuracy of gaze estimation, it is necessary to correct the pupil position taking the ACD into account. The corrected pupil position is used to estimate the gaze direction.
[0085] Figure 16 shows an example of a calibration method. In this case, the eye is looking at a calibration point known by the system, and the pupil is observed by the camera. P0 indicates the intersection of the ray (pupil observed by the camera) with the corneal sphere. P0 is the observed pupil on the camera. P0 is used for general gaze estimation.
[0086] However, in reality, the pupil is located at P1 within the corneal sphere by the ACD. The direction from the center of the eyeball (or the center of the corneal sphere) to the center of the pupil is considered the eye's gaze direction. Calibration can be performed using the gaze direction and a known calibration point.
[0087] However, in reality, the pupil is located at P1 within the corneal sphere by the ACD. The direction from the center of the eyeball (or the center of the corneal sphere) to the center of the pupil is considered the eye's gaze direction. Calibration can be performed using the gaze direction and a known calibration point.
[0088] <Refraction model>
[0089] Assuming the cornea refracts light rays, the correction is adjusted, i.e., correcting pupil position taking into account the anterior chamber depth (ACD) and the horizontal optical axis relative to the visual axis.
[0090] Figure 17 shows a refraction model for single-point calibration. P0 is obtained from the intersection of the ray (pupil observed by the camera) with the corneal sphere. We know the position of the cornea and the incident ray. Therefore, we apply "Snell's Law" (or other refraction model). This changes the direction. The ray continues so that there is a certain distance (ACD) between the pupil P1 and the corneal sphere (P0). The direction from the corneal center (or eye center) to P1 should point to a known calibration point. If not, the calibration unit 230 optimizes the ACD. Therefore, the calibration unit 230 calibrates the ACD.
[0091] Implicit Calibration
[0092] The calibration process imposes additional effort on the user. With implicit calibration, the calibration is performed while the user is viewing the content.
[0093] In traditional (explicit) calibration, 1. The system shows the user a point target at a known location. 2. The user must confirm the target for a certain period of time. 3. The system records an estimate of the user's gaze during that period. 4. The system combines the recorded gaze with ground truth positions to estimate gaze tracking parameters.
[0094] On the other hand, in implicit calibration, 1. There is no explicit point target. 2. The user does not need to take any specific action to engage in a normal VR / AR experience. 3. The system records an estimate of the user's gaze and the head-mounted screen image during a typical VR / AR experience. 4. The system combines the estimated line of sight with the screen image to obtain the ground truth position. 5. The system combined recorded gaze and ground truth positions to estimate eye-tracking parameters.
[0095] Implicit calibration may be a method for calibrating gaze detection that includes: An image of the user's eyes is captured. The gaze point is detected based on the image of the eye. An image of a user's field of view is captured while viewing a scene that does not necessarily include pre-specified content. The edges of the sub-image extracted from the view image and the sub-image containing the gaze point are detected. The gaze point is adjusted according to the detected edges.
[0096] In implicit calibration, the point at which the gaze point is adjusted may be the point that has the highest probability of an edge occurring.
[0097] Implicit calibration may further include: The detected edges are accumulated over a predetermined period. Calculate statistics from the edge distribution. After a predetermined period of time has passed, the gaze point is adjusted according to the statistics.
[0098] It is important to add that a screen image is required. The screen image is called the field of view. The field of view can cover both the screen image in VR and the external camera image in AR. It is also important to note that the user does not need to be looking at a specific target. The content of the scene can be arbitrary. There are two types of images: eye images and screen images, so you should explicitly specify which image you are referring to.
[0099] The correlation between gaze points and visual field images can be defined as an assumption about human behavior: in situation A, A and B can be automatically extracted from visual field images, and people are likely to look at B. Currently, we use the assumption that in all situations, people are likely to look at edges.
[0100] For example, there are other assumptions that could potentially be used: when presented with a new image, people are likely to see human / animal / etc. faces first, or when presented with a video with moving objects against a stationary background, people are likely to look at the moving objects. 1. Calculate the average of the edges over time. 2. Find the vector to the point with the maximum average edge.
[0101] The general idea is to automatically extract positive data from images of the visual field. However, the positive data is not a single point, but a probability distribution. While a single image of the visual field cannot predict the actual gaze point, it can predict that the actual gaze point will be located in a specific region with a certain probability.
[0102] When taking a region of interest, we essentially convert the gaze point probability distribution into a probability distribution of the eye tracking parameters. After accumulating the probability distributions at different moments in time from the visual field images, we can calculate the average probability distribution. This average probability distribution gradually converges to a single point (a single value of the eye tracking parameters). That is, the more images there are, the smaller the standard deviation of this distribution will be.
[0103] The general idea of predicting gaze from images of the visual field is not new. There is a field of research called "saliency prediction" that studies this topic. The hypothesis that "people are more likely to look at the edges of objects" also stems from saliency prediction. The method of integrating saliency prediction into eye-tracking calibration is new.
[0104] Figure 18 shows the branching of implicit calibration. "Bias" means that the difference between the estimated gaze point and the actual gaze point is constant. Humans tend to look at points with high contrast, i.e., the edges of objects. Therefore, in implicit calibration using edge accumulation: 1. Obtain an image of the user's estimated gaze and the head-mounted display screen. 2. Select a small region of the screen image (ROI; Region of Interest) around the fixation point. 3. Detect edges on the ROI. 4. Accumulate statistics over time. 5. Find the point where the edge is maximum over time. 6. Estimate the bias as the difference between the ROI center and the maximum point.
[0105] Figure 19 shows an overview of implicit calibration. The circle is the gaze point from the eye tracker (gaze detection unit 220). The star is the actual gaze point. The rectangle is the cumulative area of the visual field. We used the maximum amount of points on the edge as the positive data for calibration. Therefore, there is no need to provide calibration points in the display field.
[0106] Figure 20 shows a flowchart of implicit calibration. The gaze tracker (gaze detection unit 220) provides an approximate gaze direction. The head-mounted display provides an image of the full field of view that the user sees. Using the approximate gaze direction and the full-field of view image, the calibration unit 230 calculates statistics over time, estimates gaze tracking parameters, and feeds the parameters back to the gaze tracker. In this way, the gaze direction is gradually calibrated.
Claims
1. Measuring the speed of head rotation in a certain direction, measuring a rotational velocity of the eyeball in a direction opposite to said direction; calibrating a gaze detection unit when the sum of the head rotation speed and the eyeball rotation speed is less than a threshold value; A method for providing
2. The method described in claim 1, wherein the rotation speed of the head and the rotation speed of the eyeballs are measured while the marker is moved.
3. The method described in claim 2, wherein the marker is moved in virtual space.
4. A method as described in claim 1, wherein the rotation speed of the head in the up-down and left-right directions and the rotation speed of the eyeballs are measured.
5. A measuring device for measuring the speed of head rotation in a certain direction; a gaze detection unit that measures the rotation speed of the eyeball in a direction opposite to the direction; a calibration unit that calibrates the gaze detection unit when the sum of the head rotation speed and the eyeball rotation speed is less than a threshold value; A system comprising:
6. Measuring the speed of head rotation in a certain direction; measuring the rotational velocity of the eye in a direction opposite to said direction; calibrating a gaze detection unit when the sum of the head rotation speed and the eyeball rotation speed is less than a threshold; A program that causes a computer to execute the following.
Citation Information
Patent Citations
Instruction processor
JP1995064709A
Head-mounted display and program used therefor
JP2012216123A
Systems and methods for biomechanics-based ocular signals for interacting with real and virtual objects
JP2017526078A