Calibrating a gaze tracker
By adjusting the calibration parameters of the gaze tracker in real time, the problem of inaccurate calibration caused by user head movement or changes in device position was solved, thus improving the accuracy of gaze tracking and the usability of the device.
Patent Information
- Application Number
- CN202310304963.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-01
- Filing Date
- 2023-03-27
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-03-27
AI Technical Summary
In existing devices, gaze trackers are prone to inaccurate calibration during use due to user head movement or changes in device position, resulting in inaccurate gaze tracking.
By adjusting the calibration parameters of the gaze tracker in real time, the calibration parameters are automatically compensated to maintain accuracy based on the difference between the expected gaze position and the measured gaze position, reducing the need for recalibration operations that require user intervention.
It improves the accuracy of gaze tracking, reduces the need for recalibration with user intervention, and enhances device usability and operational smoothness.
Smart Images

Figure CN116823957B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 324,351, filed March 28, 2022, and U.S. Provisional Patent Application No. 63 / 409,293, filed September 23, 2022, which are incorporated by reference in their entirety. TECHNICAL FIELD
[0003] The present disclosure relates generally to calibrating a gaze tracker. BACKGROUND
[0004] Some devices include a display that presents visual content. Some devices manipulate the visual content based on input. Erroneous input can trigger the device to unexpectedly manipulate the visual content. Some devices perform various operations based on input. Erroneous input can trigger the device to perform unpredictable operations. BRIEF DESCRIPTION OF DRAWINGS
[0005] Accordingly, the present disclosure can be understood with reference to some example aspects of specific implementations, some of which are illustrated in the attached drawings.
[0006] Figures 1A-1J is an illustration of an example operating environment according to some implementations.
[0007] Figure 2 is a block diagram of a system to adjust calibration parameters of a gaze tracker according to some implementations.
[0008] Figure 3 is a flowchart representation of a method to adjust calibration parameters of a gaze tracker according to some implementations.
[0009] Figure 4 is a block diagram of a device to adjust calibration parameters of a gaze tracker according to some implementations.
[0010] According to common practice the various features illustrated in the drawings can not be drawn to scale. Accordingly, the dimensions of the various features can be arbitrarily expanded or reduced for the clarity of presentation. In addition, some of the drawings can not depict all of the components of a given system, method or device. Finally, like reference numerals can be used to denote like features throughout the specification and figures. SUMMARY
[0011] Various implementations disclosed herein include devices, systems, and methods for calibrating a gaze tracker. In various implementations, a device includes a display, an image sensor, a non-transitory memory, and one or more processors coupled with the display, image sensor, and non-transitory memory. In various implementations, a method includes displaying, on the display, a plurality of visual elements. In some implementations, the method includes determining, based on respective characteristic values of the plurality of visual elements, an expected gaze target indicating a first display area at which a user of the device intends to gaze while the plurality of visual elements are being displayed. In some implementations, the method includes obtaining, via the image sensor, an image including a set of pixels corresponding to a pupil of the user of the device. In some implementations, the method includes determining, by a gaze tracker, a measured gaze target based on the set of pixels corresponding to the pupil, the measured gaze target indicating a second display area at which the user is measurably gazing. In some implementations, the method includes adjusting a calibration parameter of the gaze tracker based on a difference between the first display area indicated by the expected gaze target and the second display area indicated by the measured gaze target.
[0012] According to some implementations, a device includes one or more processors, a non-transitory memory, and one or more programs. In some implementations, the one or more programs are stored in the non-transitory memory and executed by the one or more processors. In some implementations, the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some implementations, a non-transitory computer-readable storage medium has stored therein the instructions that, when executed by one or more processors of a device, cause the device to perform or cause to perform any of the methods described herein. According to some implementations, a device includes one or more processors, a non-transitory memory, and means for performing or causing to perform any of the methods described herein. DETAILED DESCRIPTION
[0013] Many details are described to provide a thorough understanding of example implementations shown in the drawings. However, the drawings merely show some example aspects of the present disclosure and should not be considered limiting. One of ordinary skill in the art would understand that other effective aspects and / or variants do not include all of the specific details described herein. Moreover, well-known systems, methods, components, devices and circuits have not been described in detail so as not to obscure the more relevant aspects of the example implementations described herein.
[0014] A physical environment refers to the physical world that people are able to sense and / or interact with without the aid of electronic devices. The physical environment can include physical features such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People are able to directly sense and / or interact with the physical environment such as through sight, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment can include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, etc. In the case of an XR system, a subset of a person’s physical motions, or representations thereof, are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR system are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect head movements and, in response, adjust graphical content and a sound field presented to the person in a manner similar to how such views and sounds would change in the physical environment. As another example, an XR system can detect movements of an electronic device (e.g., a mobile phone, a tablet, a laptop, etc.) that presents an XR environment and, in response, adjust graphical content and a sound field presented to the person in a manner similar to how such views and sounds would change in the physical environment. In some cases (e.g., for accessibility reasons), an XR system can adjust characteristics of graphical content in an XR environment in response to representations of physical motions (e.g., voice commands).
[0015] There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capability, windows with integrated display capability, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system can have an integrated opaque display and one or more speakers. Alternatively, a head-mounted system can be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system can incorporate one or more imaging sensors to capture images or video of a physical environment, and / or one or more microphones to capture audio of the physical environment. Rather than an opaque display, a head-mounted system can have a transparent or translucent display. A transparent or translucent display can have a medium through which light representative of an image is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, hologram medium, optical combiner, optical reflector, or any combination thereof. In some implementations, a transparent or translucent display can be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects graphical images onto a person's retina. Projection systems can also be configured to project virtual objects into a physical environment, for example, as a hologram or on a physical surface.
[0016] Some devices utilize gaze as input. Such devices include an image sensor and a gaze tracker. The image sensor captures a set of one or more images of a user of the device. The gaze tracker tracks the gaze of the user by identifying pixels that correspond to the user's pupils. The gaze tracker determines a gaze direction based on the pixels that correspond to the pupils. The gaze tracker is calibrated so that it accurately tracks the gaze of the user in a reliable manner. When using the device, it can be necessary to adjust the calibration so that the gaze tracker continues to accurately track the gaze. For example, a head-mounted device with an eye tracking camera can slide or move around on the user's head during use. In this example, the eye position relative to the eye tracking camera can change when the head-mounted device is in use. Too much change in the eye position relative to the eye tracking camera can cause the gaze tracking to be inaccurate.
[0017] The present disclosure provides methods, systems, and / or devices for adjusting the calibration of a gaze tracker while the gaze tracker is in use, so that the gaze tracker continues to accurately track the gaze of a user. The device adjusts the calibration of the gaze tracker as a background operation while the device is being used to perform other operations. The device adjusts the calibration of the gaze tracker based on the difference between an expected gaze location and a measured gaze location. If the difference between the expected gaze location and the measured gaze location is greater than an acceptable threshold, the device adjusts a calibration parameter of the gaze location. The device adjusts the calibration parameter so that the difference between a subsequent expected gaze location and a corresponding subsequent measured gaze location is within the acceptable threshold.
[0018] The device can determine the expected gaze location based on user input. As an example, the device expects the user to gaze at a button (e.g., a "send" button in a messaging application) when the user activates the button by pressing the button (e.g., via a physical input device such as a mouse, keyboard, touchpad, or touch screen), performing a gesture while gazing at the button, gazing at the button for a threshold length of time, or via voice input (e.g., by saying "send" or "send message"). In this example, if the gaze tracker indicates that the user is gazing 10 pixels away from the button or 10 pixels away from the expected portion of the button when the button is activated by pressing, gesturing, gazing at, or via voice input, the device determines that the gaze tracker can be generating an erroneous gaze target. As such, the device adjusts a calibration parameter of the gaze tracker based on the 10-pixel error. As another example, if the device is displaying a string of text that the user is expected to gaze at and the gaze tracker indicates that the user is gazing at a 15-pixel blank area away from the string of text, the device determines that the gaze tracker can need to be recalibrated and the device adjusts a calibration parameter of the gaze tracker based on the 15-pixel error.
[0019] The calibration parameter can vary with the position of the user's eye relative to the image sensor of the device. Since the gaze tracker utilizes the value of the calibration parameter to generate the measured gaze target, the measured gaze target varies with the position of the eye relative to the image. If the device is a head-mounted device and the head-mounted device moves while the head-mounted device is being worn on the user's head, the position of the eye relative to the image changes. Adjusting the calibration parameter compensates for the movement of the head-mounted device on the user's head. The calibration parameter can include a value that indicates the position of the eye relative to the image sensor. Adjusting the calibration parameter can include changing the value to a new value that indicates a new position of the eye relative to the image sensor.
[0020] Adjusting calibration parameters while using the device reduces the need for dedicated recalibration operations. For example, adjusting calibration parameters as a background operation reduces the need for recalibration performed as a foreground operation, where the device directs the user to perform certain operations in order to recalibrate the gaze tracker. As an example, adjusting calibration parameters during regular device use reduces the need for guided recalibration operations, where the device prompts the user to gaze at particular visual elements and recalibrates the gaze tracker based on the difference between the measured gaze position and the position of the particular visual element.
[0021] Figure 1A is a diagram illustrating an example physical environment 10 in accordance with some implementations. While relevant features are illustrated, one of ordinary skill in the art, in light of the disclosure, will appreciate that various other features are not illustrated for the sake of brevity and so as not to obscure more pertinent aspects of the example implementations disclosed herein. To that end, by way of non-limiting example, the physical environment 10 includes an electronic device 20 and a user 22 of the electronic device 20.
[0022] In some implementations, the electronic device 20 includes a handheld computing device that can be held by the user 22. For example, in some implementations, the electronic device 20 includes a smartphone, a tablet, a media player, a laptop computer, a desktop computer, etc. In some implementations, the electronic device 20 includes a wearable computing device that can be worn by the user 22. For example, in some implementations, the electronic device 20 includes a head-mountable device (HMD) or an electronic watch. In various implementations, the electronic device 20 includes an image sensor 24 that captures images of at least one eye of the user 22, a gaze tracker 26 that tracks a gaze of the user 22 based on images captured by the image sensor 24, and a display that presents a graphical environment 40 (e.g., a graphical user interface (GUI) having various GUI elements). In some implementations, the electronic device 20 includes or is connected to a physical input device, such as a mouse, a keyboard, a touch-sensitive surface (e.g., a touchpad), a clicker device, etc. In some implementations, the display includes a touch screen display that can detect user inputs (e.g., tap inputs, long-press inputs, drag inputs, etc.).
[0023] In some implementations, the electronic device 20 includes a smartphone or a tablet, and the image sensor 24 includes a front-facing camera that can capture images of the eyes of the user 22 while the user 22 is using the electronic device 20. In some implementations, the electronic device 20 includes an HMD, and the image sensor 24 includes a camera facing the user that captures images of the eyes of the user 22 while the HMD is worn on the head of the user 22. In some implementations, the electronic device 20 is a laptop that includes a touch-sensitive surface (e.g., a touchpad) for receiving user inputs from the user 22. In some implementations, the laptop is connected to a separate physical input device, such as a mouse, a touchpad, or a keyboard, for receiving user inputs from the user 22. In some implementations, the electronic device 20 is a desktop computer that is connected to a separate physical input device (e.g., a mouse, a touchpad, or a keyboard) for receiving user inputs from the user 22.
[0024] In some implementations, the gaze tracker 26 obtains a set of one or more images captured by the image sensor 24. The gaze tracker 26 identifies pixels that correspond to the eyes and / or pupils of the user 22. The gaze tracker 26 tracks the gaze of the user 22 based on the pixels that correspond to the eyes and / or pupils of the user 22. In some implementations, the gaze tracker 26 generates a gaze target that includes a gaze location value, a gaze intensity value, and a gaze duration value. The gaze location value indicates coordinates of a display region within the graphical environment 40 that the user 22 is gazing at. The gaze intensity value indicates a number of pixels that the user 22 is gazing at. The gaze duration value indicates a duration of time that the user 22 has been gazing at the display region indicated by the gaze location value.
[0025] In various implementations, the electronic device 20 calibrates the gaze tracker 26 so that the gaze tracker 26 accurately tracks the gaze of the user 22 in a reliable manner. In some implementations, the gaze tracker 26 is associated with a set of one or more calibration parameters 28 (referred to below as “calibration parameters 28”). In such implementations, calibrating the gaze tracker 26 includes setting values of the calibration parameters 28. In some implementations, the calibration parameters 28 include a first value 30 and a second value 32. Figure 1A In the example of FIG. 1, the calibration parameters 28 have a first value 30. In some implementations, the first value 30 includes a default value. In some implementations, the first value 30 varies with an expected position of the eyes of the user 22 relative to the image sensor 24. For example, the first value 30 can correspond to the eyes of the user 22 being aligned with the image sensor 24 (e.g., the eyes intersecting an axis at the center of a frustum of view of the image sensor 24). In other words, the first value 30 can represent the eyes being at the center of the field of view of the image sensor 24.
[0026] In some implementations, the graphical environment 40 includes a two-dimensional (2D) environment. In some implementations, the graphical environment 40 includes a three-dimensional (3D) environment, such as an XR environment. In Figure 1A In the example of FIG. 4, the graphical environment 40 includes various visual elements 50 (e.g., a first visual element 50a, a second visual element 50b, a third visual element 50c, and a fourth visual element 50d). In some implementations, the visual elements 50 include graphical objects (e.g., XR objects). In some implementations, the visual elements 50 include selectable affordances (e.g., buttons) that the user 22 can select by providing user input (e.g., touch input via a touchpad or touchscreen, mouse click input via a mouse, key press input via a keyboard, gaze input, voice input via a microphone, etc.). In some implementations, the visual elements 50 include text 52 (e.g., a brief description of the functionality of the first visual element 50a). In some implementations, the visual elements 50 include graphics 54 (e.g., images, such as a visual indication of the functionality of the second visual element 50b).
[0027] Referring to Figure 1B In some implementations, the visual elements 50 are associated with respective feature values 60. For example, the first visual element 50a is associated with a first feature value 60a, the second visual element 50b is associated with a second feature value 60b, the third visual element 50c is associated with a third feature value 60c, and the fourth visual element 50d is associated with a fourth feature value 60d.
[0028] As shown in Figure 1B In some implementations, the electronic device 20 determines the expected gaze target 70 based on the feature values 60. In Figure 1B In the example of FIG. 4, the expected gaze target 70 indicates an expected gaze location 72 corresponding to the third visual element 50c. The expected gaze target 70 indicates a display region in which the user 22 is expected to gaze. In Figure 1B In the example of FIG. 4, the user 22 is expected to gaze at the third visual element 50c because the expected gaze location 72 coincides with the location of the third visual element 50c.
[0029] In some implementations, the feature values 60 include respective saliency values for the corresponding visual elements 50, and the electronic device 20 determines the expected gaze target 70 based on the saliency values. For example, the electronic device 20 selects the location of the third visual element 50c as the expected gaze location 72 because the third visual element 50c has the greatest saliency value among the visual elements 50. In some implementations, the electronic device 20 obtains (e.g., generates or receives) a saliency map of the graphical environment 40, and the electronic device 20 retrieves the saliency values for the visual elements 50 from the saliency map.
[0030] In some implementations, the feature values 60 include respective position values for the corresponding visual elements 50, and the electronic device 20 determines the intended gaze target 70 based on the position values. For example, the electronic device 20 can select the position of a particular visual element 50 as the intended gaze location 72 because the particular visual element 50 has a position value that is within a threshold range of the position value at which the intended user 22 gazes (e.g., the particular visual element 50 is positioned near the center of the display area of the electronic device 20).
[0031] In some implementations, the feature values 60 include respective color values for the corresponding visual elements 50, and the electronic device 20 determines the intended gaze target 70 based on the color values. For example, the electronic device 20 can select the position of a particular visual element 50 as the intended gaze location 72 because the particular visual element 50 has a color value that matches a threshold color value at which the intended user 22 is expected to direct their gaze (e.g., the particular visual element 50 is red, while the rest of the display area of the electronic device 20 is black; or the particular visual element 50 is a color display, while the rest of the display area is a black and white display).
[0032] In some implementations, the feature values 60 include respective movement values for the corresponding visual elements 50, and the electronic device 20 determines the intended gaze target 70 based on the movement values. For example, the electronic device 20 can select the position of a particular visual element 50 as the intended gaze location 72 because the particular visual element 50 has a movement value that matches a threshold movement value at which the intended user 22 is expected to direct their gaze (e.g., the particular visual element 50 is moving, while the other visual elements 50 are stationary).
[0033] In some implementations, the feature values 60 include respective user interaction values for the corresponding visual elements 50, and the electronic device 20 determines the intended gaze target 70 based on the user interaction values. In some implementations, the user interaction values are based on current and / or historical user inputs provided by the user 22 to indicate respective levels of interaction with the visual elements 50. As an example, the electronic device 20 can select the position of a particular visual element 50 as the intended gaze location 72 because the particular visual element 50 has a user interaction value that matches a threshold interaction value and the user 22 is more likely to interact with (e.g., select) the particular visual element 50 (e.g., the user 22 is more likely to select the particular visual element 50 because the particular visual element 50 has been selected the most number of times among the visual elements 50).
[0034] In some implementations, the intended gaze target 70 is associated with a confidence value that indicates a confidence (e.g., a degree of certainty) associated with the intended gaze target 70. In some implementations, the confidence value varies with the feature values 60. In some implementations, the confidence value varies with (e.g., is proportional to) a variance of the feature values 60. As an example, if the third feature value 60c is the highest of the feature values 60, and the difference between the third feature value 60c and the second highest feature value 60 is greater than a threshold difference, the confidence value associated with the intended gaze target 70 can be set to a value greater than a threshold confidence value (e.g., the confidence value can be set to a value greater than 0.5, e.g., the confidence value can be set to “1”). In this example, if the difference between the highest feature value and the second highest feature value of the feature values 60 is less than the threshold difference, the confidence value associated with the intended gaze target 70 can be set to a value less than the threshold confidence value (e.g., the confidence value can be set to a value less than 0.5, e.g., the confidence value can be set to 0.2).
[0035] In some implementations, the intended gaze location 72 corresponds to a location indicated by user input received via a physical input device. In some implementations, the electronic device 20 detects user input at a particular location within the graphical environment 40 via a physical input device, and the electronic device 20 sets the particular location of the user input as the intended gaze location 72. In some implementations, in response to detecting a mouse click by a mouse, the electronic device 20 sets a cursor location of a cursor as the intended gaze location 72. For example, in response to detecting a mouse click while the cursor is positioned over the third visual element 50c, the electronic device 20 can set a location (e.g., a center) of the third visual element 50c as the intended gaze location 72. In some implementations, in response to detecting a tap or press via a touch-sensitive surface such as a touchpad or touch screen display, the electronic device 20 sets a cursor location of a cursor as the intended gaze location 72. For example, in response to detecting a tap or press via a touchpad or touch screen display, the electronic device 20 can set a location (e.g., a center) of the third visual element 50c as the intended gaze location 72.
[0036] In some implementations, in response to detecting a key press via the keyboard, the electronic device 20 sets the position of the focus element to the intended gaze position 72. For example, when the focus element is on the third visual element 50c, in response to detecting a press of the enter key, the electronic device 20 can set the position (e.g., center) of the third visual element 50c to the intended gaze position 72. In some implementations, in response to detecting a voice input corresponding to a select command, the electronic device 20 sets the position of the focus element to the intended gaze position 72. For example, when the focus element is on the third visual element 50c (e.g., when the user 22 says “select”), in response to detecting a select voice command, the electronic device 20 can set the position of the third visual element 50c to the intended gaze position 72.
[0037] Referring to Figure 1C , the image sensor 24 captures an image 78 of the user 22, and the gaze tracker 26 utilizes the image 78 to generate a measured gaze target 80. In some implementations, the measured gaze target 80 indicates a measured gaze position 82. The measured gaze position 82 represents a display area that the gaze tracker 26 has identified as corresponding to the gaze of the user 22. In some implementations, the measured gaze target 80 includes a measured gaze intensity that indicates a number of pixels that the gaze tracker 26 has identified as corresponding to the gaze of the user 22. In some implementations, the measured gaze target 80 includes a measured gaze duration that indicates a duration of time that the gaze tracker 26 has identified as corresponding to the gaze of the user 22. As Figure 1C can be seen, the measured gaze position 82 can be different from the intended gaze position 72.
[0038] Referring back to Figure 1B , in some implementations, the electronic device 20 selects the position of a particular visual element 50 as the intended gaze position 72 when the particular visual element 50 has a position value that is within a threshold range of position values from the position indicated by the measured gaze target 80. In some implementations, the electronic device 20 selects the position of a particular visual element 50 as the intended gaze position 72 when the visual element 50 has a position value that is closest to the position indicated by the measured gaze target 80. In some implementations, the electronic device 20 selects the position of a particular visual element 50 as the intended gaze position 72 when the measured gaze position 82 remains stationary (or below a threshold amount of movement) for a threshold length of time, when the visual element 50 has a position value that is closest to the position indicated by the measured gaze target 80. In some implementations, the electronic device 20 selects the position of a particular visual element 50 as the intended gaze position 72 when a select input is received, when the visual element 50 has a position value that is closest to the position indicated by the measured gaze target 80.
[0039] Referring to Figure 1DIn various specific implementations, electronic device 20 determines the expected gaze target 70 ( Figure 1B (shown in) and Figure 1C and Figure 1D The differences between the measured gaze targets 80 are shown in the figure. Figure 1D In the example, electronic device 20 identifies a difference 90 between the expected gaze position 72 and the measured gaze position 82. This difference 90 may be caused by movement of electronic device 20 relative to the user 22's eye. For example, difference 90 may be caused by the eye no longer being in the center of the image sensor 24's field of view. After identifying difference 90, electronic device 20 determines a new value 32 for calibration parameter 28 (e.g., different from...). Figure 1C (The first value 30 shown is a second value). In some specific implementations, the new value 32 varies with the difference 90. The new value 32 can compensate for the eye not being in the center of the field of view of the image sensor 24. The electronic device 20 replaces the first value 30 with the new value 32.
[0040] Figure 1E This corresponds to the time period that occurs after the electronic device 20 sets the calibration parameter 28 to a new value 32. After setting the calibration parameter 28 to the new value 32, the image sensor 24 captures another image 98, and the gaze tracker 26 generates another measured gaze target 100 based on this other image 98. The measured gaze target 100 indicates another measured gaze position 102 that coincides with the expected gaze position 72. Setting the calibration parameter 28 to the new value 32 reduces (e.g., makes it smaller or eliminates) the gaze target 100. Figure 1D The difference shown is 90, and the accuracy of gaze tracker 26 is improved. (As shown) Figures 1C-1E As shown, when electronic device 20 is performing non-calibration-related operations, electronic device 20 adjusts calibration parameter 28. In Figures 1C-1E In the example, the adjustment of calibration parameter 28 is performed as a background operation rather than a foreground operation in order to reduce disruption to the operability of electronic device 20.
[0041] Advantageously, electronic device 20 adjusts calibration parameters 28 without displaying prompts to user 22 to adjust the position of electronic device 20 above his / her head. For example, electronic device 20 does not request user 22 to move electronic device 20 so that user 22's eyes are centered in the field of view of image sensor 24. Furthermore, electronic device 20 adjusts calibration parameters 28 without performing guided calibration, which may include prompting user 22 to look at specific visual elements 50 in order to adjust calibration parameters 28. The aforementioned presentation of guided calibration reduces disruption to the operability of electronic device 20, thereby increasing the usability of electronic device.
[0042] Figure 1F and Figure 1GThe electronic device 20 is shown to determine the intended gaze target 112 based on user input 110 provided by user 22. Figure 1G The sequence is shown in the image. Figure 1F As shown, electronic device 20 detects user input 110 at a location corresponding to the fourth visual element 50d. In some embodiments, electronic device 20 includes a touchscreen display, and electronic device 20 detects user input 110 by detecting a tap on the touchscreen display. Alternatively, in some embodiments, electronic device 20 displays the graphical environment 40 as a virtual plane, and electronic device 20 detects user input 110 by detecting the intersection between a collider object representing user 22's finger (e.g., a finger) and the virtual plane of the graphical environment 40. For example, electronic device 20 detects user input 110 by detecting a 3D gesture performed by user 22. In some embodiments, electronic device 20 detects user input 110 by detecting voice input. For example, user 22 may utter a phrase corresponding to a request to select the fourth visual element 50d (e.g., user 22 may say "select the bottom right option"). In some embodiments, electronic device 20 detects user input 110 via a physical input device (e.g., a mouse, keyboard, touch-sensitive surface (e.g., touchpad), or clicker device). In some embodiments, the physical input device is connected to the electronic device 20 via a wire. Alternatively, in some embodiments, the physical input device provides instructions for user input 110 to the electronic device 20 via wireless communication.
[0043] In some implementations, feature value 60 indicates whether the corresponding visual element 50 has been selected. For example, feature value 60 may include a binary value, where a value of "0" indicates that the corresponding visual element 50 has not been selected, while a value of "1" indicates that the corresponding visual element 50 has been selected. Figure 1F In the example, the fourth feature value 60d may have a binary value "1" to indicate that the fourth visual element 50d has been selected, while the first feature value 60a, the second feature value 60b, and the third feature value 60c may have a binary value "0" to indicate that the first visual element 50a, the second visual element 50b, and the third visual element 50c have not been selected.
[0044] refer to Figure 1G In response to user input 110 detecting selection of fourth visual element 50d, electronic device 20 generates a anticipated gaze target 112, which indicates a anticipated gaze position 114 corresponding to the fourth visual element 50d. Figure 1G As can be seen, the expected gaze position 114 indicates that when user 22 is selecting the fourth visual element 50d, user 22 is expected to gaze at the fourth visual element 50d. Figure 1G In the example, if the difference between the measured gaze position and the expected gaze position 114 is greater than a threshold, the electronic device 20 adjusts...Figure 1A Calibration parameters 28 of the gaze tracker 26 are shown.
[0045] Figures 1H-1J The electronic device 20 is shown determining a sequence of expected gaze targets based on movement of a particular visual element 50. Figure 1H A fifth visual element 50e is shown moving in a direction indicated by arrow 120. For example, the fifth visual element 50e moves to the right of the graphical environment 40. In some implementations, the feature values 60 indicate movement of the corresponding visual elements 50. For example, the feature values 60 can indicate a respective speed at which the corresponding visual elements 50 are moving. In Figure 1H In the example shown, the visual elements 50a-d are stationary. As such, the feature values 60a-d can indicate a movement speed of 0. However, since the fifth visual element 50e is moving, the fifth feature value 60e can indicate a speed at which the fifth visual element 50e is moving. Alternatively, in some implementations, the feature values 60 include binary values, where a value of “0” indicates no movement and a value of “1” indicates movement.
[0046] As Figure 1H shown, the electronic device 20 determines a first expected gaze location 130a corresponding to the location of the fifth visual element 50e. In Figure 1H the example shown, the electronic device 20 selects the location of the fifth visual element 50e as the first expected gaze location 130a because the feature values 60 indicate that the fifth visual element 50e is moving while the remaining visual elements 50a-d are stationary. As Figure 1H shown, the electronic device 20 determines a first measured gaze location 140a that is offset from the first expected gaze location 130a by a difference 150. The difference 150 between the first expected gaze location 130a and the first measured gaze location 140a can be caused by a change in the position of the eyes of the user 22 relative to the image sensor 24. For example, in some implementations, the electronic device 20 includes a head-mounted device that the user 22 wears on his / her head, and the difference 150 can be caused by the electronic device 20 sliding on the head as the user 22 moves.
[0047] Referring back Figure 1I , as the fifth visual element 50e moves through the graphical environment 40 in the direction indicated by arrow 120, the electronic device 20 can generate additional expected gaze targets and additional measured gaze targets to determine whether to adjust the calibration parameters 28 of the gaze tracker 26. For example, the electronic device 20 determines a second expected gaze location 130b corresponding to a new location of the fifth visual element 50e. In Figure 1IIn the example, the previous position of the fifth visual element 50e is indicated by the dashed box 160. The electronic device 20 determines a second measured gaze position 140b, offset by a difference 150 from the second expected gaze position 130b. Due to the first measured gaze position 140a ( Figure 1H (as shown in the image) from the first expected gaze position 130a ( Figure 1H As shown in the figure, the difference 150 was offset and the second measured gaze position 140b was offset from the second expected gaze position 130b by a similar or identical difference 150, so the electronic device 20 determined to change the calibration parameter 28 of the gaze tracker 26 with a greater degree of certainty.
[0048] like Figure 1J As shown, the electronic device 20 adjusts the calibration parameter 28 by changing its value from a first value 30 to a third value 34. In some embodiments, the third value 34 varies with the difference 150 between the expected gaze positions 130a and 130b and the corresponding measured gaze positions 140a and 140b. For example, in some embodiments, the difference between the first value 30 and the third value 34 is related to... Figure 1H and Figure 1I The difference shown is proportional to 150.
[0049] Figure 1J The new position of the fifth visual element 50e, the third expected gaze position 130c, and the third measured gaze position 140c are shown. Figure 1J In the example, the previous position of the fifth visual element 50e is indicated by another dashed box 162. The electronic device 20 determines the third measured gaze position 140c after adjusting the calibration parameter 28. (As...) Figure 1J As can be seen, the third expected gaze position 130c and the third measured gaze position 140c are juxtaposed. In other words, the third measured gaze position 140c matches the third expected gaze position 130c. Since the gaze tracker 26 determines the third measured gaze position 140c after setting the calibration parameter 28 to the third value 34, the third measured gaze position 140c coincides with the third expected gaze position 130c. Therefore, the third measured gaze position 140c does not deviate from the third expected gaze position 130c.
[0050] Figure 2 It involves adjusting the calibration parameters of the gaze tracker according to certain specific implementations (e.g., Figure 1A The block diagram shows a system 200 for the calibration parameters 28 of the gaze tracker 26. In some embodiments, system 200 includes an expected gaze determiner 210, a measured gaze determiner 230, and a calibration parameter adjuster 250. In various embodiments, system 200 resides in... Figures 1A-1J The electronic device shown is located at 20 locations (e.g., implemented therein).
[0051] In various implementations, the expected gaze determiner 210 obtains (e.g., receives or determines) feature values 220 associated with corresponding visual elements being displayed on the display (e.g., Figure 1B As shown, the expected gaze determiner 210 determines an expected gaze target 212 based on the feature values 220. For example, the expected gaze determiner 210 determines Figure 2 As shown, the expected gaze determiner 210 determines an expected gaze target 212 based on the feature values 220. For example, the expected gaze determiner 210 determines Figure 1B As shown, the expected gaze determiner 210 determines an expected gaze target 212 based on the feature values 220. For example, the expected gaze determiner 210 determines Figure 1B As shown, the expected gaze determiner 210 determines an expected gaze target 212 based on the feature values 220. For example, the expected gaze determiner 210 determines
[0052] In some implementations, the expected gaze determiner 210 determines a confidence score associated with the expected gaze target 212. The confidence score indicates a degree of certainty in the expected gaze target 212. In some implementations, the confidence score varies with the feature values 220. For example, the confidence score can be based on a distribution of the feature values 220. As an example, if the feature values 220 have a large variance, the confidence score can be higher. Conversely, if the feature values 220 have a lower variance, the confidence score can be lower.
[0053] In some implementations, the feature values 220 include saliency values 220a that indicate respective saliency levels of the visual elements. In some implementations, the saliency values 220a are based on respective saliencies of the visual elements (e.g., more salient visual elements have larger saliency values 220a than less salient visual elements). In some implementations, the saliency values 220a are based on respective conspicuities of the visual elements (e.g., more conspicuous visual elements have larger saliency values 220a than less conspicuous visual elements). In some implementations, the expected gaze determiner 210 obtains (e.g., receives or generates) a saliency map that includes the saliency values 220a. In some implementations, the expected gaze determiner 210 determines the expected gaze target 212 such that the expected gaze location 212a corresponds to the visual element having the largest saliency value 220a.
[0054] In some implementations, the feature values 220 include a position value 220b that indicates a respective position of the visual element. In some implementations, the user is more likely to gaze at certain positions. For example, the user can be more likely to gaze at visual elements that are positioned toward the center of the display area. In this example, the expected gaze determiner 210 can generate the expected gaze target 212 such that the expected gaze position 212a points toward a visual element that is near the center of the display area. More generally, in various implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze position 212a corresponds to a visual element that is positioned within a portion of the display area that the user is more likely to gaze at. In some implementations, the expected gaze determiner 210 identifies the portion of the display area that the user is more likely to gaze at based on historical gaze tracking data. For example, if the historical gaze tracking data indicates that the user spends more time gazing at a particular portion of the display area (e.g., the center or the upper right), the expected gaze determiner 210 determines that the particular portion of the display area is the portion that the user is more likely to gaze at.
[0055] In some implementations, the feature values 220 include a color value 220c that indicates a respective color of the visual element. In some implementations, the user is more likely to gaze at color visual elements and less likely to gaze at black and white visual elements. More generally, in various implementations, the user is more likely to gaze at certain colors (e.g., bright colors such as red and blue) and less likely to gaze at other colors (e.g., dark colors such as gray). In some implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze position 212a points to a position of a visual element that has a color value 220c that matches a threshold color value (e.g., a preferred color, such as a bright color such as red or blue).
[0056] In some implementations, the feature values 220 include movement values 220d that indicate respective movement of the visual elements. In some implementations, the movement values 220d include binary values that indicate whether the corresponding visual element is moving (e.g., “0” indicates stationary and “1” indicates moving). In some implementations, the expected gaze determiner 210 determines that the user is more likely to gaze at moving visual elements and less likely to gaze at stationary visual elements. Accordingly, in some implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze location 212a is directed at a visual element that has a movement value 220d indicating movement (e.g., has a movement value 220d of “1”). In some implementations, the movement values 220d include movement speed. In some implementations, the expected gaze determiner 210 determines that the user is more likely to gaze at visual elements that are moving quickly and less likely to gaze at visual elements that are moving slowly. Accordingly, in some implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze location 212a is directed at a visual element that has the greatest movement value 220d.
[0057] In some implementations, the feature values 220 include interaction values 220e that indicate current and / or historical user interaction with the visual elements. In some implementations, the interaction values 220e include binary values that indicate whether the user is currently interacting with the corresponding visual element. For example, an interaction value 220e of “1” can indicate that the device has detected user input that selects the corresponding visual element, and an interaction value 220e of “0” can indicate that the device has not detected user input that selects the corresponding visual element. In some implementations, the expected gaze determiner 210 determines that the user is more likely to gaze at visual elements that the user is currently interacting with and less likely to gaze at visual elements that the user is not currently interacting with. Accordingly, in some implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze location 212a is directed at a visual element that the user has selected via user input (e.g., via a click input of a mouse, via a key press input of a keyboard, via a touch input of a touchpad or touchscreen, via a gesture input of a touchscreen or 3D gesture tracker, via a voice input of a microphone, etc.). As an example, with reference to Figure 1G , the expected gaze determiner 210 generates the expected gaze location 114 that is directed at the fourth visual element 50d in response to the user input 110 that selects the fourth visual element 50d.
[0058] In some implementations, the interaction values 220e indicate historical user interactions with the corresponding visual elements. In some implementations, the historical user interactions include historical user inputs for the visual elements. In some implementations, the expected gaze determiner 210 determines that a visual element with more historical user interactions is more likely to be gazed at than a visual element with fewer historical user interactions. Accordingly, in some implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze location 212a coincides with the location of the visual element with the greatest interaction value 220e.
[0059] In various implementations, the measured gaze determiner 230 obtains a set of one or more images 232 of the user’s eye, and the measured gaze determiner 230 determines a measured gaze target 234 based on the images 232. In some implementations, the measured gaze determiner 230 receives the images 232 from an image sensor (e.g., a camera facing the user, such as an eye tracking camera) that captures the images 232. In some implementations, the measured gaze determiner 230 implements a gaze tracker 26 as shown. Figure 1A In some implementations, the measured gaze determiner 230 includes a set of one or more calibration parameters 240 (e.g., calibration parameters 28 as shown) that are used to determine the measured gaze target 234. In some implementations, the measured gaze determiner 230 determines the measured gaze target 234 based on the images 232 and the set of one or more calibration parameters 240. Figure 1A In some implementations, the measured gaze determiner 230 includes a set of one or more calibration parameters 240 (e.g., calibration parameters 28 as shown) that are used to determine the measured gaze target 234. In some implementations, the measured gaze determiner 230 determines the measured gaze target 234 based on the images 232 and the set of one or more calibration parameters 240. Figure 2 In some implementations, the set of one or more calibration parameters 240 has an existing value 242. For example, as shown, the calibration parameters 28 have a first value 30. Figure 1A In some implementations, the set of one or more calibration parameters 240 has an existing value 242. For example, as shown, the calibration parameters 28 have a first value 30.
[0060] In various implementations, the measured gaze determiner 230 determines a measured gaze target 234 based on the images 232 of the user’s eye. In some implementations, the measured gaze target 234 includes a measured gaze location 234a (e.g., measured gaze location 82 as shown), a measured gaze intensity 234b, and / or a measured gaze duration 234c. In some implementations, the measured gaze target 234 varies with the existing value 242 of the calibration parameters 240. As such, changing the value of the calibration parameters 240 triggers a change in the measured gaze target 234. In some implementations, the measured gaze determiner 230 provides the measured gaze target 234 to the calibration parameter adjuster 250. Figure 1C
[0061] In some implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze location 212a points to a visual element proximate or closest to a location indicated by the measured gaze target 234. In some implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze location 212a points to a visual element proximate or closest to a location indicated by the measured gaze target 234 when the measured gaze target 234 remains stationary for a threshold length of time (or below a threshold amount of movement). In some implementations, the expected gaze determiner 210 generates the expected gaze target 212 such that the expected gaze location 212a points to a visual element proximate or closest to a location indicated by the measured gaze target 234 when a selection input (e.g., a touch input, a gesture input, a voice input, etc.) is received.
[0062] In some implementations, the calibration parameter adjuster 250 adjusts the calibration parameter 240 based on a comparison of the expected gaze target 212 and the measured gaze target 234. In some implementations, the calibration parameter adjuster 250 generates a new value 252 of the calibration parameter 240 when a difference between the expected gaze target 212 and the measured gaze target 234 exceeds a threshold difference. Referring to Figure 1D , the calibration parameter adjuster 250 changes a value of the calibration parameter 28 from a first value 30 to a new value 32 based on a difference 90 between the expected gaze location 72 and the measured gaze location 82. In some implementations, the calibration parameter adjuster 250 generates the new value 252 of the calibration parameter 240 so as to reduce a difference between a subsequent expected gaze target and a corresponding subsequent measured gaze target. In some implementations, the calibration parameter adjuster 250 produces the new value 252 as a background operation while the device continues to perform other operations (e.g., non-calibration related operations, such as displaying visual content).
[0063] In some implementations, the new value 252 varies with a difference between the expected gaze target 212 and the measured gaze target 234. In some implementations, a difference between the existing value 242 and the new value 252 is proportional to a difference between the expected gaze target 212 and the measured gaze target 234. For example, the greater the difference between the expected gaze target 212 and the measured gaze target 234, the greater the difference between the existing value 242 and the new value 252. In some implementations, the calibration parameter adjuster 250 uses a lookup table (LUT) to determine the new value 252. For example, the LUT can list various values representing varying differences between the expected gaze target 212 and the measured gaze target 234, and corresponding changes to make to the existing value 242.
[0064] In some implementations, the calibration parameter adjuster 250 obtains a confidence score associated with the expected gaze target 212. For example, the expected gaze determiner 210 provides the calibration parameter adjuster 250 with the expected gaze target 212 and the confidence score associated with it. In some implementations, in response to a confidence score greater than a threshold confidence score, the calibration parameter adjuster 250 generates a new value 252 for the calibration parameter 240. In some implementations, in response to a confidence score less than a threshold confidence score, the calibration parameter adjuster 250 abandons the generation of the new value 252 for the calibration parameter 240.
[0065] Figure 3 This is a flowchart illustrating a method 300 for adjusting calibration parameters of a gaze tracker. In various specific embodiments, method 300 comprises a device including a display, an image sensor, non-transitory memory, and one or more processors coupled to the display, image sensor, and non-transitory memory (e.g., Figures 1A-1J The electronic device 20 and / or shown Figure 2 The method 300 is executed by the system 200 shown. In some embodiments, the method 300 is executed by processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, the method 300 is executed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).
[0066] As shown in box 310, in various specific embodiments, method 300 includes displaying multiple visual elements on a display. For example, as Figure 1A As shown, electronic device 20 displays visual element 50. In some specific embodiments, the visual element includes text (e.g., Figure 1A The text 52 shown) and / or images (e.g., Figure 1A (See Figure 54). In some implementations, visual elements include GUI elements that are part of the GUI. For example, visual elements include selectable display representations (e.g., buttons) that can be selected by the user of the device.
[0067] As illustrated in box 320, in various specific embodiments, method 300 includes determining an expected gaze target for a first display area based on corresponding feature values of the plurality of visual elements, whereby a user of the device is expected to gaze at or intend to gaze at the first display area while the plurality of visual elements are being displayed. For example, as Figure 1B As shown, the electronic device 20 determines the expected gaze target 70 based on the feature value 60 of the visual element 50.
[0068] As shown in box 320a, in some embodiments, the corresponding feature values include corresponding saliency values. In such embodiments, method 300 includes determining the expected gaze target based on the corresponding saliency value. For example, asFigure 2 As shown, in some implementations, the feature values 220 include saliency values 220a, and the expected gaze determiner 210 determines the expected gaze target 212 based on the saliency values 220a. In some implementations, an expected user will gaze at the most salient visual element. As an example, if the device is displaying a virtual character, it can be expected that the user will gaze at the virtual character’s eyes rather than the virtual character’s knees. Thus, in this example, the virtual character’s eyes can have a greater saliency value than the virtual character’s knees.
[0069] In some implementations, the respective feature values include location values that indicate respective placements of the plurality of visual elements. In such implementations, the method 300 includes determining the expected gaze target based on the location values of the visual elements. For example, as shown in FIG. 2B, the expected gaze determiner 210 determines the expected gaze target 212 based on the location values 220b of the visual elements 202a-202d. In this example, the expected gaze target 212 is directed at the visual element 202b that is positioned near the center of the display area 204. Figure 2 As shown, in some implementations, the feature values 220 include location values 220b, and the expected gaze determiner 210 determines the expected gaze target 212 based on the location values 220b. As an example, it can be expected that a user will gaze at a visual element that is positioned near the center of a display area rather than a visual element that is positioned near an edge of the display area, which can be in the user’s peripheral vision. In this example, the expected gaze target is directed at the visual element that is near the center of the display area. In other implementations, it can be expected that a user will gaze at a visual element to select the visual element or that a user will gaze at a visual element while providing a separate selection input (e.g., a touch input, a gesture input, a voice input, etc.).
[0070] In some implementations, the respective feature values include color values that indicate respective colors of the plurality of visual elements. In such implementations, the method 300 includes determining the expected gaze target based on the color values of the visual elements. For example, as shown in FIG. 2C, the expected gaze determiner 210 determines the expected gaze target 212 based on the color values 220c of the visual elements 202a-202d. In this example, the expected gaze target 212 is directed at the visual element 202b that has the color that is most different from the other colors of the visual elements 202a-202d. Figure 2 As shown, in some implementations, the feature values 220 include color values 220c, and the expected gaze determiner 210 determines the expected gaze target 212 based on the color values 220c. As an example, it can be expected that a user will gaze at the most vibrant visual element. As another example, it can be expected that a user will gaze at a visual element that has a particular color (e.g., red, blue, etc.).
[0071] In some implementations, the respective feature values include movement values that indicate respective speeds at which the plurality of visual elements are moving. In such implementations, the method 300 includes determining the expected gaze target based on the movement values of the visual elements. For example, as shown in FIG. 2D, the expected gaze determiner 210 determines the expected gaze target 212 based on the movement values 220d of the visual elements 202a-202d. In this example, the expected gaze target 212 is directed at the visual element 202b that is moving at the fastest speed. Figure 2 As shown, in some implementations, the feature values 220 include movement values 220d, and the expected gaze determiner 210 determines the expected gaze target 212 based on the movement values 220d. As an example, it can be expected that a user will gaze at a moving visual element rather than a stationary visual element. For example, with reference to FIG. 2D, it can be expected that a user will gaze at the visual element 202b that is moving at the fastest speed rather than the visual element 202a that is stationary. Figures 1H-1JIt is expected that the user 22 will gaze at the fifth visual element 50e that is moving rather than the stationary visual elements 50a-d.
[0072] In some implementations, the respective feature values include an interaction value indicative of a respective level of interaction between the user and the plurality of visual elements. In such implementations, the method 300 includes determining the expected gaze target based on the interaction values associated with the visual elements. For example, as shown in Figure 2 the feature values 220 include an interaction value 220e, and the expected gaze determiner 210 determines the expected gaze target 212 based on the interaction value 220e. As another example, with reference to Figure 1G When the user input 110 is directed to the fourth visual element 50d, it is expected that the user 22 will gaze at the fourth visual element 50d.
[0073] In some implementations, determining the expected gaze target includes detecting, via a physical input device, an input directed to the first display region, and setting the first display region as the expected gaze target in response to detecting the input directed to the first display region. For example, with reference to Figure 1F and Figure 1G In some implementations, the electronic device 20 detects the user input 110 via a mouse, and sets the location indicated by the user input 110 as the expected gaze location 114. For example, the user 22 uses the mouse to move the cursor to the fourth visual element 50d, and clicks or double-clicks the fourth visual element 50d.
[0074] In some implementations, determining the expected gaze target includes detecting, via a mouse, a mouse click when the first display region corresponds to a cursor location, and setting the cursor location as the expected gaze target in response to detecting the mouse click. For example, with reference to Figure 1F and Figure 1G In some implementations, the electronic device 20 detects the user input 110 via a mouse. For example, the user 22 uses the mouse to move the cursor to the fourth visual element 50d, and clicks or double-clicks the fourth visual element 50d.
[0075] In some implementations, determining the expected gaze target includes detecting, via a touch-sensitive surface, a tap input when the first display region corresponds to a cursor location, and setting the cursor location as the expected gaze target in response to detecting the tap input. For example, with reference to Figure 1F and Figure 1G In some implementations, the electronic device 20 detects the user input 110 via a touchpad. For example, the user 22 uses the touchpad to move the cursor to the fourth visual element 50d, and clicks or double-clicks the fourth visual element 50d.
[0076] In some implementations, determining the expected gaze target includes detecting, via a keyboard, a key press when the first display region corresponds to a focus indicator, and setting the first display region as the expected gaze target in response to detecting the key press. For example, with reference toFigure 1F And Figure 1G In some implementations, the electronic device 20 detects the user input 110 via a keyboard. For example, the user 22 uses the directional keys on the keyboard to move the focus indicator to the fourth visual element 50d and presses an enter key to select the fourth visual element 50d.
[0077] In some implementations, determining the intended gaze target includes detecting, via an audio sensor, a voice input including a select command while the first display area corresponds to the focus indicator, and setting the first display area as the intended gaze target in response to detecting the voice input including the select command. For example, with reference to Figure 1F And Figure 1G In some implementations, the electronic device 20 detects the user input 110 by voice input received via a microphone. For example, if the fourth visual element 50d represents an icon of a messaging application, the user 22 can say “open the messaging application.”
[0078] As shown by block 330, in various implementations, the method 300 includes obtaining, via an image sensor, an image including a set of pixels corresponding to a pupil of a user of the device. For example, as shown by Figure 1C the image sensor 24 captures an image 78 of at least one eye of the user 22. In some implementations, the image sensor includes a user-facing camera that captures images after informed consent is obtained from the user. In some implementations, the image sensor includes an eye tracking camera that captures images after informed consent is obtained from the user.
[0079] As shown by block 340, in various implementations, the method 300 includes determining, by a gaze tracker, a measured gaze target based on the set of pixels corresponding to the pupil, the measured gaze target indicating a second display area that the user is currently measurably gazing at. For example, with reference to Figure 1C the gaze tracker 26 determines a measured gaze target 80 that indicates a measured gaze location 82 that the user 22 of the electronic device 20 is currently gazing at. As another example, with reference to Figure 2 the measured gaze determiner 230 determines a measured gaze target 234 based on the image 232. In some implementations, the second display area indicated by the measured gaze target is a distance from the first display area indicated by the intended gaze target. As an example, Figure 1D demonstrates a difference 90 between the intended gaze location 72 and the measured gaze location 82. As another example, the measured gaze target can indicate that the user is gazing at a blank area that is several pixels from the intended gaze location corresponding to the eye of the virtual character being displayed.
[0080] As shown in block 340a, in some implementations the plurality of visual elements includes a moving element at which the intended user is expected to gaze, the intended gaze target corresponds to a location of the moving element, and the measured gaze target indicates a location that is offset from the location of the moving element. For example, as shown in Figure 1H FIG. 6A, the visual elements 50 include a fifth visual element 50e that moves in the direction indicated by arrow 120, the first intended gaze location 130a overlaps the location of the fifth visual element 50e, and the first measured gaze location 140a is offset from the first intended gaze location 130a by a difference 150.
[0081] As shown in block 340b, in some implementations the plurality of visual elements includes a selectable affordance at which the intended user is expected to gaze when the selectable affordance is selected, the intended gaze target corresponds to a location of the selectable affordance, and the measured gaze target indicates a location that is offset from the location of the selectable affordance. For example, as shown in Figure 1G FIG. 6B, the user 22 selects the fourth visual element 50d, the intended gaze location 114 overlaps the location of the fourth visual element 50d, and the measured gaze location (not shown in FIG. 6B) does not overlap the fourth visual element 50d. Figure 1G
[0082] As shown in block 350, in various implementations the method 300 includes adjusting a calibration parameter of the gaze tracker based on a difference between a first display area indicated by the intended gaze target and a second display area indicated by the measured gaze target. For example, as shown in Figures 1C-1E FIG. 7, the electronic device 20 adjusts the calibration parameter 28 by changing a value of the calibration parameter 28 from the first value 30 to a new value 32 based on the difference 90 between the intended gaze location 72 and the measured gaze location 82. In various implementations, the device adjusts the calibration parameter such that subsequently captured images produce measured gaze locations that overlap the intended gaze locations. For example, as shown in Figure 1E FIG. 8, changing the value of the calibration parameter 28 from the first value 30 to the new value 32 causes the measured gaze location 102 to overlap the intended gaze location 72. As described herein, in various implementations the adjustment to the calibration parameter is performed as a background operation without prompting the user to view certain visual elements as part of a guided calibration operation. Adjusting the calibration parameter as a background operation reduces disruptions to device operability, thereby increasing the usability of the device.
[0083] As shown in block 350a, in some implementations adjusting the calibration parameter includes adjusting the calibration parameter when a distance between the first display area and the second display area is greater than a threshold distance. For example, with reference to Figures 1D-1E In response to the difference 90 between the expected gaze location 72 and the measured gaze location 82 being greater than the threshold distance, the electronic device 20 adjusts the calibration parameter 28. In some implementations, the method 300 includes forgoing adjustment of the calibration parameter when the distance between the first display region and the second display region is less than the threshold distance.
[0084] In some implementations, the adjustment to the calibration parameter is proportional to the distance between the first display region and the second display region. For example, the greater the distance between the first display region and the second display region, the greater the adjustment to the calibration parameter. As an example, reference is made to Figure 1D The adjustment to the calibration parameter 28 is proportional to the difference 90 between the expected gaze location 72 and the measured gaze location 82.
[0085] As shown in block 350b, in some implementations, adjusting the calibration parameter includes recording a change in position of the pupil from the first expected position to the second expected position. In some implementations, the device includes a head-mountable device that is expected to be worn in a particular way by the user. For example, the user’s eyes are expected to be at the center of the field of view of the eye tracking camera. However, as the user moves, the head-mountable device can shift, and the eyes can no longer be at the center of the field of view of the eye tracking camera. In some implementations, adjusting the calibration parameter compensates for movement of the head-mountable device on the user’s head. For example, the change in the calibration parameter compensates for the eyes not being at the center of the field of view of the eye tracking camera.
[0086] As shown in block 350c, in some implementations, the expected gaze target is associated with a confidence score, and adjusting the calibration parameter includes adjusting the calibration parameter in response to the confidence score being greater than a threshold confidence score and forgoing adjustment of the calibration parameter in response to the confidence score being less than the threshold confidence score. For example, as described with respect to Figure 2 In some implementations, when the expected gaze target 212 is associated with a confidence score that is greater than a threshold confidence score, the calibration parameter adjuster 250 generates a new value 252 of the calibration parameter 240, and when the expected gaze target 212 is associated with a confidence score that is less than the threshold confidence score, the calibration parameter adjuster 250 forgoes generating the new value 252.
[0087] In some implementations, the confidence score varies with a density of the plurality of visual elements. For example, the confidence score can vary with an amount of spacing between the plurality of visual elements. As an example, when the visual elements are positioned relatively close to one another, the confidence score of the expected gaze target can be low (e.g., below a threshold confidence score). In contrast, when the visual elements are positioned relatively far from one another, the confidence score of the expected gaze target can be high (e.g., greater than the threshold confidence score).
[0088] In some implementations, the confidence score varies with a distance between the first display region and the second display region. For example, referring to Figure 1D the confidence score of the expected gaze target 70 can vary with a difference 90 between the expected gaze location 72 and the measured gaze location 82. In some implementations, the confidence score is inversely proportional to a distance between the first display region and the second display region. For example, the greater the distance between the expected gaze location and the measured gaze location, the lower the confidence score of the expected gaze location.
[0089] As shown in block 350d, in some implementations, adjusting the calibration parameter includes adjusting the calibration parameter in response to device movement data indicating that the device has moved more than a threshold amount since a previous adjustment of the calibration parameter. As described herein, a head-mountable device can slide on a user’s head as the user moves while using the device. In some implementations, if sensor data from a visual-inertial odometry (VIO) system (e.g., accelerometer data from an accelerometer, gyroscope data from a gyroscope, and / or magnetometer data from a magnetometer) indicates that the head-mountable device and / or the user has moved more than a threshold amount since a previous adjustment of the calibration parameter, the device can adjust the calibration parameter in order to continue to provide accurate gaze tracking.
[0090] As shown in block 350e, in some implementations, adjusting the calibration parameter includes adjusting the calibration parameter in response to the second display region corresponding to a blank region. For example, referring to Figure 1C in response to the measured gaze location 82 pointing to a blank region in which no visual elements are being displayed, the electronic device 20 can adjust the calibration parameter 28. When the device displays visual elements, it is expected that the user will gaze at one or more of the displayed visual elements rather than the blank region. Thus, when gaze tracking indicates that the user is gazing at the blank region while the device is displaying visual elements, the gaze tracking is likely to generate false measured gaze targets, and can be improved by adjusting the calibration parameter.
[0091] As shown in block 350f, in some implementations, adjusting the calibration parameter includes adjusting the calibration parameter in response to the first display region having a first saliency value that is greater than a second saliency value of the second display region. In some implementations, it is expected that the user will gaze at the visual element having the greatest saliency value. For example, referring to Figure 1B in response to the saliency value of the second visual element 50b being greater than the saliency values of the remaining visual elements 50a, 50c, and 50d, the electronic device 20 can select the location corresponding to the second visual element 50b as the expected gaze location 72.
[0092] In some implementations, method 300 includes determining a desired gaze target based on a measured gaze target. In some implementations, the measured gaze target indicates that the user's gaze is directed to a portion of the object within a threshold distance of the object's center. For example, the measured gaze target might indicate that the user is gazing towards the edge of a button rather than its center. In various implementations, it is expected that the user will gaze at or towards the center of the object. For example, it is expected that the user will gaze at or near the center of a button. In such implementations, method 300 includes determining that the desired gaze location corresponds to the center of the object the user is measuringly gazing at.
[0093] Figure 4 This is a block diagram based on some specific implementations of device 400. In some specific implementations, device 400 implements... Figures 1A-1J The electronic device 20 and / or shown Figure 2 The system 200 is shown. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to avoid obscuring more relevant aspects of the specific embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, device 400 includes one or more processing units (CPUs) 401, a network interface 402, a programming interface 403, a memory 404, one or more input / output (I / O) devices 408, and one or more communication buses 405 for interconnecting these and various other components.
[0094] In some implementations, a network interface 402 is provided to establish and maintain a metadata tunnel between a cloud-hosted network management system and at least one private network including one or more compatible devices, among other uses. In some implementations, one or more communication buses 405 include circuitry for interconnecting and communicating between system components. Memory 404 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices, and may include non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 404 optionally includes one or more storage devices remotely located to one or more CPUs 401. Memory 304 includes a non-transitory computer-readable storage medium.
[0095] In some embodiments, memory 404 or a non-transitory computer-readable storage medium of memory 404 stores programs, modules, and data structures, or subsets thereof, including an optional operating system 406, an expected gaze determiner 210, a measured gaze determiner 230, and a calibration parameter adjuster 250. In various embodiments, device 400 executes... Figure 3 Method 300 is shown.
[0096] In some implementations, the expected gaze determiner 210 includes a tool for determining the expected gaze target (e.g., Figure 1B The expected gaze target 70 and / or Figure 2 The instructions 210a of the expected gaze target 212 shown, as well as heuristics and metadata 210b, are used. In some specific implementations, the expected gaze determiner 210 executes instructions 210a of the expected gaze target 212. Figure 3 The middle frame 320 represents at least some of the operations.
[0097] In some implementations, the measured gaze determiner 230 includes a method for determining the measured gaze target (e.g., based on the value of calibration parameter 240) Figure 1C The measured gaze target 80 and / or shown Figure 2 The instructions 230a of the measured gaze target 234 shown, as well as heuristics and metadata 230b, are used. In some specific implementations, the measured gaze determiner 230 executes instructions 230a of the measured gaze target 234. Figure 3 The middle frame 340 represents at least some of the operations.
[0098] In some specific implementations, the calibration parameter adjuster 250 includes a tool for adjusting the calibration parameter 240 (e.g., Figure 1A The calibration parameter 28) shown is indicated by instruction 250a, along with heuristics and metadata 250b. In some specific implementations, the calibration parameter adjuster 250 executes the instructions 250a, along with heuristics and metadata. Figure 3 Box 350 in the diagram represents at least some of the operations.
[0099] In some specific implementations, the one or more I / O devices 408 include features for obtaining input (e.g., Figure 1F The user input device 110 shown is an input device. In some embodiments, the input device includes a touchscreen (e.g., for detecting tap input), an image sensor (e.g., for detecting gesture input), and / or a microphone (e.g., for detecting voice input). In some embodiments, the one or more I / O devices 408 include an image for capturing the user's pupils (e.g., for capturing...). Figure 1C Image 78 shown Figure 1E The image 98 and / or shown Figure 2 The image sensor (e.g., a user-facing camera, such as an eye-tracking camera) shown in image 232. In some specific implementations, the one or more I / O devices 408 include a display for displaying visual elements.
[0100] In various implementations, the one or more I / O devices 408 include a video pass-through display that displays at least a portion of the physical environment surrounding the device 400 as an image captured by a camera. In various implementations, the one or more I / O devices 408 include an optical pass-through display that is at least partially transparent and passes light emitted or reflected by the physical environment.
[0101] It will be appreciated that Figure 4 The functional description of various features that can be present in particular implementations, differs from the structural illustrations of the implementations described herein. As will be appreciated by one of ordinary skill in the art, items shown separately could be combined and items shown separately could be separated. For example, Figure 4 Some of the functional blocks shown separately in the can be implemented as a single block, and various functions of a single functional block can be implemented through one or more functional blocks in various implementations. The actual number of blocks and the division of the particular functions between the blocks and how the features are allocated among the blocks will vary from one implementation to another and, in some implementations, depend in part on the particular combination of hardware, software, and / or firmware chosen for a particular implementation.
[0102] The implementations described herein contemplate the use of gaze information to present salient viewpoints and / or salient information. Implementers should consider the extent to which gaze information is collected, analyzed, disclosed, transmitted, and / or stored in order to respect established privacy policies and / or privacy practices. These considerations should include applying practices generally considered to meet or exceed industry and / or governmental requirements for maintaining user privacy. The present disclosure also contemplates that the use of a user’s gaze information can be limited to the extent necessary in achieving the implementations described. For example, in implementations where a user’s device provides processing capability, gaze information can be processed locally at the user’s device.
[0103] The various processes defined herein consider the option of obtaining and utilizing a user’s personal information. For example, such personal information can be utilized in order to provide improved privacy screens on electronic devices. However, to the extent that such personal information is collected, such information should be obtained with the user’s informed consent. As described herein, the user should be made aware of and have control over the use of their personal information.
[0104] Personal information will be used by appropriate parties for legitimate and reasonable purposes only. Parties utilizing such information will adhere to privacy policies and practices that at least comply with appropriate legal requirements. Furthermore, such policies should be comprehensive, user-accessible, and considered to meet or exceed governmental / industry standards. Furthermore, parties will not distribute, sell, or otherwise share such information except for any reasonable and legitimate purposes.
[0105] However, a user can limit the extent to which personal information is accessible or otherwise made available to various parties. For example, settings or other preferences can be adjusted to allow a user to decide whether and in what manner, if at all, personal information is to be accessed by various entities. Furthermore, while some features defined herein are described in the context of using personal information, aspects of these features can be implemented without requiring the use of such information. For example, if user preferences, account names, and / or location history are collected, this information can be obfuscated or otherwise generalized such that it does not identify the corresponding user.
[0106] While various aspects of implementations within the scope of the appended claims are described above, it should be apparent that the various features of the above described implementations can be embodied in a wide variety of forms, and that defined in the appended claims are intended to be used as illustrative only. As will be apparent to those skilled in the art from the disclosure herein, the aspects of the disclosure described herein can be implemented independently of any other aspects of the disclosure and that not all of these aspects need work or be realized in conjunction with any other aspects of the disclosure. For example, a device can be implemented or a method can be practiced using any number of the aspects set forth herein. In addition, to the extent that some aspects of this disclosure do not provide novelty over the prior art, those aspects, to the extent relevant to this disclosure, as well as those aspects considered obvious by the examiner or others skilled in the art, are intended to be expressly released.
Claims
1. A method of adjusting a calibration parameter of a gaze tracker, the method comprising: at a device comprising a display, an image sensor, a non-transitory memory, and one or more processors coupled with the display, the image sensor, and the non-transitory memory: displaying a graphical user interface (GUI) on the display, the graphical user interface comprising a plurality of selectable GUI elements; in response to a user input, determining an expected gaze target based on respective feature values of the plurality of selectable GUI elements, the expected gaze target indicating a first display area at which a user of the device intends to gaze while the plurality of selectable GUI elements are being displayed; obtaining, via the image sensor, an image comprising a set of pixels corresponding to a pupil of the user of the device; determining, by the gaze tracker, a measured gaze target based on the set of pixels corresponding to the pupil, the measured gaze target indicating a second display area at which the user is measurably gazing; and in response to the user input selecting a respective selectable GUI element from the displayed plurality of selectable GUI elements, adjusting the calibration parameter of the gaze tracker based on a difference between the first display area indicated by the expected gaze target and the second display area indicated by the measured gaze target without prompting a guided calibration.
2. The method of claim 1, wherein respective feature values corresponding to the user input comprise respective saliency values.
3. The method of claim 1 or 2, wherein respective feature values corresponding to the user input comprise location values indicating respective placements of the plurality of selectable GUI elements.
4. The method of claim 1 or 2, wherein respective feature values corresponding to the user input comprise color values indicating respective colors of the plurality of selectable GUI elements.
5. The method of claim 1 or 2, wherein respective feature values corresponding to the user input comprise movement values indicating respective speeds at which the plurality of selectable GUI elements are moving.
6. The method of claim 1 or 2, wherein respective feature values corresponding to the user input indicate respective levels of interaction between the user and the plurality of selectable GUI elements.
7. The method of claim 1 or 2, wherein adjusting the calibration parameter comprises adjusting the calibration parameter when a distance between the first display area and the second display area is greater than a threshold value.
8. The method of claim 1 or 2, wherein the adjustment to the calibration parameter is proportional to a distance between the first display area and the second display area.
9. The method of claim 1 or 2, wherein adjusting the calibration parameter comprises recording a change in a position of the pupil from a first expected position to a second expected position.
10. The method of claim 1 or 2, wherein the expected gaze target is associated with a confidence score, and wherein adjusting the calibration parameter comprises: adjusting the calibration parameter in response to the confidence score being greater than a threshold confidence score; and forfeit adjustment of the calibration parameter in response to the confidence score being less than the threshold confidence score.
11. The method of claim 10, wherein the confidence score varies with a density of the plurality of selectable GUI elements.
12. The method of claim 10, wherein the confidence score varies with a distance between the first display region and the second display region.
13. The method of any of claims 1, 2, 11, and 12, wherein adjusting the calibration parameter includes adjusting the calibration parameter in response to device movement data indicating that the device has moved more than a threshold amount since a previous adjustment of the calibration parameter.
14. The method of any of claims 1, 2, 11, and 12, wherein the plurality of selectable GUI elements includes a mobile selectable GUI element at which the user is expected to gaze; wherein the expected gaze target corresponds to a location of the mobile selectable GUI element; and wherein the measured gaze target indicates a location offset from the location of the mobile selectable GUI element.
15. The method of any of claims 1, 2, 11, and 12, wherein the plurality of selectable GUI elements includes a selectable affordance at which the user is expected to gaze when selecting the selectable affordance; wherein the expected gaze target corresponds to a location of the selectable affordance; and wherein the measured gaze target indicates a location offset from the location of the selectable affordance.
16. The method of any of claims 1, 2, 11, and 12, wherein adjusting the calibration parameter includes adjusting the calibration parameter in response to the second display region corresponding to a blank area.
17. The method of any of claims 1, 2, 11, and 12, wherein adjusting the calibration parameter includes adjusting the calibration parameter in response to the first display region having a first saliency value that is greater than a second saliency value of the second display region.
18. The method of any of claims 1, 2, 11, and 12, wherein the expected gaze target is determined based on the measured gaze target when a selection input is received.
19. An electronic device, the electronic device comprising: one or more processors; an image sensor; a display; a non-transitory memory; and one or more programs stored in the non-transitory memory, which, when executed by the one or more processors, cause the electronic device to: display, on the display, a graphical user interface (GUI) that includes a plurality of selectable GUI elements; determine, in response to a user input, an expected gaze target based on respective characteristic values of the plurality of selectable GUI elements, the expected gaze target indicating a first display region at which a user of the electronic device intends to gaze while the plurality of selectable GUI elements are being displayed; obtaining, via the image sensor, an image including a set of pixels corresponding to a pupil of the user of the electronic device; determining, by a gaze tracker, a measured gaze target based on the set of pixels corresponding to the pupil, the measured gaze target indicating a second display area at which the user is measurably gazing; and in response to the user input selecting a respective selectable GUI element from the plurality of selectable GUI elements displayed, adjusting, without prompting for guided calibration, a calibration parameter of the gaze tracker based on a difference between the first display area indicated by the expected gaze target and the second display area indicated by the measured gaze target.
20. A non-transitory memory storing one or more programs, which, when executed by one or more processors of a device including a display and an image sensor, cause the device to: display, on the display, a graphical user interface (GUI) including a plurality of selectable GUI elements; in response to a user input, determine, based on respective characteristic values of the plurality of selectable GUI elements, an expected gaze target indicating a first display area at which a user of the device intends to gaze while the plurality of selectable GUI elements are being displayed; obtain, via the image sensor, an image including a set of pixels corresponding to a pupil of the user of the device; determine, by a gaze tracker, a measured gaze target based on the set of pixels corresponding to the pupil, the measured gaze target indicating a second display area at which the user is measurably gazing; and in response to the user input selecting a respective selectable GUI element from the plurality of selectable GUI elements displayed, adjust, without prompting for guided calibration, a calibration parameter of the gaze tracker based on a difference between the first display area indicated by the expected gaze target and the second display area indicated by the measured gaze target.
Citation Information
Patent Citations
Cursor movement device
CN104641316A