Selecting real-world objects using wearable devices

The wearable device system improves object selection accuracy by using gaze tracking and reticle interaction to estimate user intent, enhancing user experience and optimizing resource usage.

JP7833555B2Active Publication Date: 2026-03-19GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing wearable devices struggle to accurately select virtual and real-world objects and attributes due to the inability to determine user intent effectively, leading to potential adverse user experiences.

Method used

A wearable device system that utilizes image sensors, gaze tracking, and reticle-based interaction to identify a subset of targets within the user's line of sight, estimating candidate targets through gaze direction and depth analysis, and adjusting the display accordingly.

Benefits of technology

Enhances user experience by minimizing frustration from incorrect selections and optimizing resource usage by accurately determining user intent and reducing unnecessary processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007833555000001
    Figure 0007833555000001
  • Figure 0007833555000002
    Figure 0007833555000002
  • Figure 0007833555000003
    Figure 0007833555000003
Patent Text Reader

Abstract

The method includes receiving an image from a sensor of the wearable device, rendering the image on a display of the wearable device, identifying a set of targets in the image, tracking a gaze direction associated with a user of the wearable device, rendering a gaze on the displayed image based on the tracked gaze direction, identifying a subset of targets based on the set of targets in a region of the image based on the gaze, triggering an action, and in response to the trigger, estimating candidate targets based on the subset of targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] ,

[0001] Cross - reference to Related Applications This application is a continuation of U.S. Non - Provisional Patent Application No. 17 / 651,209, entitled "SELECTION OF REAL - WORLD OBJECTS USING A WEARABLE DEVICE", filed on February 15, 2022, and claims priority therefrom. The disclosure of which is hereby incorporated by reference in its entirety.

[0002] Embodiments relate to the selection of objects and / or object attributes (e.g., text) using a wearable device.

Background Art

[0003] By using an Augmented Reality (AR) device (e.g., a wearable device), many operations can be performed to improve the user experience. One of these operations can be, for example, converting text. To perform some of these operations, the AR device selects virtual objects and / or real - world objects and / or object attributes (e.g., text) through user interaction as an input to the operation. The inability to accurately select virtual objects and / or real - world objects and / or object attributes can potentially have an adverse effect on the user experience.

Summary of the Invention

[0004] In a general embodiment, a device, a system, a non-temporary computer-readable medium (storing computer-executable program code that can be run on a computer system), and / or a method can perform a process by means of a method. The method includes receiving an image from a sensor of a wearable device, rendering the image on the display of the wearable device, identifying a set of targets in the image, tracking the gaze direction associated with the user of the wearable device, rendering the gaze on the displayed image based on the tracked gaze direction, identifying a subset of targets based on the set of targets in the area of ​​the image based on the gaze, triggering an action, and, in response to the trigger, estimating candidate targets based on the subset of targets.

[0005] In another common embodiment, the wearable device comprises an image sensor, a display, at least one processor, and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to cause the wearable device to perform, using at least one processor, the following: receive an image from the image sensor; render the image on the display; identify a set of targets in the image; track the gaze direction associated with the user of the wearable device; render the gaze on the displayed image based on the tracked gaze direction; identify a subset of targets based on the set of targets in the image region based on the gaze; trigger an action; and, in response to the trigger, estimate candidate targets based on the subset of targets.

[0006] Embodiments may include one or more of the following features. For example, a method (and / or computer program code) may further include identifying a subset of targets based on a region encompassing the line of sight, and estimating the depth associated with each target in the set of targets, wherein the estimation of candidate targets is based on the intersection of the line of sight at the depth contained in the region. A method (and / or computer program code) may further include detecting a change in line of sight direction, determining that the change is below a threshold, and re-rendering the image on the display of a wearable device. A method (and / or computer program code) may further include detecting a change in line of sight direction, determining that the change is below a threshold, and re-rendering the line of sight. A method (and / or computer program code) may further include detecting a change in line of sight direction, determining that the change is within the rendered image and closer to a subset of targets, and re-rendering the line of sight using a change in color. A method (and / or computer program code) may further include detecting a change in line of sight direction, determining that the change is greater than a threshold, and receiving another image from the sensor. The method (and / or computer program code) may further include rendering a reticle on a displayed image based on the position of a candidate target. The method may further include repositioning the reticle to a different position on the displayed image, and the candidate target is estimated based on the repositioned reticle. The method (and / or computer program code) may further include calibrating the wearable device based on the position of the wearable device's sensors and the center of the wearable device's display.

[0007] Exemplary embodiments will be better understood from the detailed description below and the accompanying drawings of this specification. Similar elements are indicated by similar reference numerals and are given for illustrative purposes only and are not limiting to the exemplary embodiments. [Brief explanation of the drawing]

[0008] [Figure 1A] An exemplary embodiment shows a side perspective view of a user observing multiple objects in a real-world scenario. [Figure 1B] An exemplary embodiment shows a front perspective view of a user looking at multiple objects in a real-world scene. [Figure 2] This illustrates a three-dimensional rendering of the image space using an exemplary embodiment. [Figure 3] This illustrates a three-dimensional rendering of the image space using an exemplary embodiment. [Figure 4] A block diagram of the calibration data flow according to an exemplary embodiment is shown. [Figure 5] A block diagram of eye-tracking using a target identification data flow, according to an exemplary embodiment, is shown. [Figure 6] A block diagram of a method for identifying a target according to an exemplary embodiment is shown. [Figure 7] A block diagram of the system according to an exemplary embodiment is shown. [Figure 8] Examples of computer devices and mobile computer devices are shown, according to at least one exemplary embodiment. [Modes for carrying out the invention]

[0009] It should be noted that these figures are intended to illustrate the general characteristics of the methods, structures, and / or materials used in specific exemplary embodiments and to supplement the following descriptions. However, these drawings are not to scale and may not accurately reflect the exact structural or performance characteristics of any given embodiment and should not be interpreted as defining or limiting the range of values ​​or characteristics encompassed by the exemplary embodiments. For example, the relative thickness and arrangement of molecules, layers, regions, and / or structural elements may be reduced or exaggerated for clarity. The use of similar or identical reference numerals in various drawings is intended to indicate the presence of similar or identical elements or features.

[0010] A user wearing a wearable device (e.g., smart glasses) may want to perform an action (e.g., convert specific text) based on a target (e.g., a road sign) or a part of a target (e.g., a line of a multi-line road sign) within their environment. However, there may be multiple candidate targets (multiple road signs) or parts of candidate targets (e.g., a multi-line road sign) on which to perform the action (e.g., convert).

[0011] The disclosed solution will enable users to communicate to a wearable device (e.g., smart glasses) which target or part of a target they intend to act upon. For example, the solution could provide functionality that a wearable device user can use to communicate to the wearable device their selection of real-world targets, objects, and / or areas of interest in the real world.

[0012] Technical solutions may include identifying a set of targets in an image captured using the image sensors of a wearable device. Then, the user's gaze on the wearable device can be used to determine a subset of targets based on their gaze direction. Next, if an action is triggered, candidate targets(s) can be estimated and / or selected from this subset. Various techniques can be used to reduce the set of targets to a subset and / or estimate candidate targets(s). For example, gaze, reticles, and / or several other visual tools can be used to narrow down the targets the user is likely to intend to perform an action on by focusing on or helping to focus the user's gaze.

[0013] The advantages of this solution include, for example, improving the user experience by minimizing frustration caused by receiving incorrect information. A technical advantage may be the reduction in the use of limited resources (processing, power, etc.) in the wearable device as a result of repeating an action due to receiving an inaccurate or incorrect response from that action.

[0014] Users wearing wearable devices can see many targets (or objects) in real-world situations. Wearable devices can make it difficult to determine which targets and what aspects of those targets the user is interested in. Figures 1A and 1B can be used to illustrate (and / or refer to) exemplary embodiments for determining which of many targets are candidate targets. Figures 1A and 1B illustrate how the embodiments described herein can address the difficulty of identifying which of many targets are candidate targets (or targets of interest).

[0015] Figure 1A shows a side perspective view of a user gazing at multiple targets in a real-world scene according to an exemplary embodiment. Figure 1A shows a user 105 wearing a wearable device 110 (e.g., an AR / VR device) and viewing a scene (e.g., a real-world scene including multiple targets 115, 120, 125, 130, 135). The multiple targets 115, 120, 125, 130, 135 are at depths D1, D2, D3, D4 (e.g., at a distance from the wearable device 110). Figure 1B shows a front perspective view of a user gazing at multiple targets in a real-world scene according to an exemplary embodiment. As shown in Figure 1B, the real-world scene 140 includes multiple targets 115, 120, 125, 130, 135. Multiple targets 115, 120, 125, 130, and 135 have associated texts: Text 1, Text 2, Text 3, Text 4, Text 5, Text 6, and Text 7. Figures 1A and 1B show the gaze directions GD1, GD2, and GD3 as the field of view (up / down, left / right) of a user 105 viewing a real-world scene 140.

[0016] Referring to Figure 1A, targets 115 and 120 are at depth D1, target 135 is at depth D2, target 125 is at depth D3, and target 130 is at depth D4. Referring to Figure 1B, targets 115 and 125 are on the left, targets 120 and 130 are on the right, and target 135 is between targets 115 and 120, overlapping them (as indicated by the dashed line, target 135 is behind targets 115 and 120).

[0017] Referring to Figure 1A, GD1 is typically oriented upward toward targets 120 and 115. Referring to Figure 1B, GD1 is typically oriented left toward targets 115 and 125. Based on the orientation of GD1, the maximum likelihood region of the image representing real-world scene 140 may be (or contain) target 115, where target 115 contains text 1, text 2, and text 3. Thus, target 115 may also be within the region of the image representing real-world scene 140, and a subset of targets may be identified as text 1, text 2, and text 3. In the exemplary embodiment, candidate targets can be estimated based on a subset of targets, or as one of text 1, text 2, and text 3. The techniques described below can be used to reduce the subset of targets and / or estimate candidate targets. In other words, using the techniques described below, one of Text 1, Text 2, and Text 3 can be selected or estimated as a candidate target (for example, a target of interest to the user 105 of the wearable device 110).

[0018] Referring to Figure 1A, GD2 is typically oriented slightly downward from directly towards targets 115, 120, 130, and 135. Referring to Figure 1B, GD2 is typically oriented slightly to the right from directly towards targets 120 and 135. Based on the orientation of GD2, the maximum likelihood region of the image representing real-world scene 140 may be (or include) targets 120 and 135, where target 120 includes texts 4 and 5. Thus, targets 120 and 135 may also be within the region of the image representing real-world scene 140, and a subset of targets can be identified as texts 4, 5, and 8 (text 8 being within target 135). In the exemplary embodiment, candidate targets can be estimated based on a subset of targets, or as one of texts 4, 5, and 8. The techniques described below can be used to reduce the subset of targets and / or estimate candidate targets. In other words, using the techniques described below, one of texts 4, 5, and 8 can be selected or estimated as a candidate target (for example, a target of interest to user 105 of wearable device 110). Since target 120 is at depth D1 and target 135 is at depth D2, one technique for estimating candidate targets can be depth-based.

[0019] Referring to FIG. 1A, GD3 is typically downward towards targets 130 and 135. Referring to FIG. 1B, GD3 is typically leftward towards targets 115 and 125. Based on the direction of GD3, the most likely region of the image representing the real-world scene 140 can be target 125 (or including target 125). Thus, target 125 may be within the region of the image representing the real-world scene 140, and a subset of the targets can be identified as text 6. In an exemplary embodiment, a candidate target can be estimated based on the subset of the targets or estimated as text 6. Using the techniques described below, the subset of the targets can be reduced and / or the candidate target can be estimated. Since the subset of the targets is a subset of 1, the most likely result may be to estimate or select text 6 as the candidate target (e.g., the target of interest to user 105 of wearable device 110).

[0020] In the foregoing techniques, details regarding the image space can be used to determine useful information regarding the line-of-sight direction, the center and offset of the camera, the depth of the object, and user input (e.g., via reticle, head movement, etc.). FIGS. 2-5 show various image space details used to determine useful information.

[0021] FIG. 2 shows a 3D rendering of an image space according to an exemplary embodiment. The 3D rendering can be based on the camera's view frustum and the view frustum of the lens or screen. The view frustum is a frustum of a pyramid that determines what is within the field of view. Only objects within the frustum of the view cone can be displayed on the screen and / or in the image. The camera (or eye) is at the tip of the pyramid. The pyramid extends in a direction away from the camera (or eye). The view frustum starts at the near plane and ends at the far plane. These planes are parallel, and their normal vectors are along the line of sight or the direction of the camera (e.g., the direction the eye or camera is looking). The length of the view frustum is determined by the distance from the camera (or eye) to the near plane and the distance from the camera to the far plane. In an exemplary embodiment, the division into far and near fields can result from the observation that the epipolar lines of the user's line of sight visible from the camera (e.g., the lines from the eye to the image plane) are concentrated in a small region of the image space. In an exemplary use case of a wearable device, a far-field object can be an object having a distance greater than 1 meter.

[0022] As shown in FIG. 2, there can be a view frustum 210 of a camera (e.g., camera 250), view frustums 245-1, 245-2 of the lens (of the wearable device), and view frustums 220-1, 220-3 of the screen (displayed on the lens of the wearable device). The view frustum 210 of the camera can have an associated epipolar line 230, and the view frustums 220-1, 220-3 of the screen can have associated lines of sight 235 (e.g., epipolar lines).

[0023] An image 220-^ of the real-world scene 205 can be located on the image plane associated with the view frustums 220-1, 220-3 of the screen. The image 220-2 can include objects 225-1, 225-2, 225-3.

[0024] In an exemplary embodiment, gazes 235-1 and 235-2 can be used to identify gaze directions (e.g., GD1, GD2, GD3). In Figure 2, gazes 235-1 and 235-2 refer to object 225-1. Furthermore, there is no indication that gazes 235-1 and 235-2 refer to any other objects (e.g., objects 225-2 and 225-3). Thus, in the example in Figure 2, object 225-1 may be a candidate target (e.g., a target of interest to user 105 of wearable device 110), or may contain a candidate target. If object 225-1 contains several targets (e.g., text), the exemplary embodiment may include the use of gazes, reticles, and / or some other visual tools to help the user indicate which of the several targets is the candidate target. The use of gazes and reticles can be illustrated using Figure 3.

[0025] Figure 3 shows a three-dimensional rendering of image space according to an exemplary embodiment. As shown in Figure 3, a screen 305 (e.g., a display on a lens, or as part of the lens of a wearable device) shows a line of sight 235 (representing either line of sight 235-1 or 235-2). The line of sight 235 can be of fixed size dimensions so as to be displayed on the screen 305. The line of sight 235 can be positioned on (within, together with, etc.) a rendered image, having a first end at the edge of the rendered image and a second end near the center of the rendered image. The line of sight 235 can have a somewhat triangular or conical shape. The line of sight 235 can have its first end at or near the outer edge of the rendered image (e.g., the far right edge). The first end of the line of sight 235 can be relatively longer than the second end. The second end of the line of sight 235 can reach a point tapering from the first end of the line of sight 235.

[0026] The line of sight 235 can be used to indicate the direction of the line of sight and can be used as a pointer to a potential target object. The line of sight 235 can be displayed in multiple parts 325, 330, and 335 of different colors. The parts 325, 330, and 335 can be positioned along the longitudinal axis of the line of sight 235. The length of each part 325, 330, and 335 can be shorter than the total length of the line of sight 235. The parts 325, 330, and 335 can taper along the longitudinal axis of the line of sight 235. The first part 325 can be positioned at the second edge of the line of sight 235 and can contain a point of the line of sight 235. The third part 335 can be positioned at the second end of the line of sight 235 at the edge of the rendered image. The third part 330 can be positioned between the first part 325 and the third part 335 along the longitudinal axis of the line of sight 235. The parts 325, 330, and 335 of different colors can be used to indicate proximity to an object. For example, there may be three colors. Color 1 (associated with part 325) can indicate that user 105 is close to the object (e.g., within 1 meter). Color 2 (associated with part 330) can indicate that user 105 is in the middle range to the object (e.g., between 1 and 3 meters). Color 3 (associated with part 335) can indicate that user 105 is in the far range to the object (e.g., more than 3 meters).

[0027] The screen 305 can be fixed in size so that objects that are far away (e.g., far field of view) may appear relatively smaller (e.g., compared to the same object that is closer). Furthermore, objects that are closer (e.g., near field of view) may appear relatively larger (e.g., compared to the same object that is far away). Additionally, as the user 105 approaches an object (or as the object approaches the user 105), the object may adapt and appear larger. In an exemplary embodiment, since the user 105 is relatively far from the object, the dominant color of the line of view 235 may be color 3 (associated with section 335). As the user 105 approaches the object, color 3 may become less dominant, color 2 may become the dominant color, color 1 may become less dominant, and color 3 may not be shown (or may be only a small streak), with color 1 (associated with section 325) and color 2 (associated with section 330) becoming more widespread. Next, as user 105 approaches the object, color 2 may progress in a less dominant manner until it becomes invisible (or only a small streak), and color 1 may become more dominant until it becomes virtually the only color in line of sight 235.

[0028] The gaze direction 235 can be used to identify a subset of targets based on the region encompassing the gaze direction. The depth associated with each target in the set of targets can be estimated. Candidate target estimation can be based on the intersection of the gaze directions at depths included in the region, and / or targets in the subset of targets. Changes in gaze direction can be detected. For example, head movements and / or eye movements can be detected. If the change in gaze direction is below a threshold, the image can be re-rendered on the wearable device's display (for example, because it causes minimal changes in gaze direction). The gaze can also be redrawn. If the change in gaze direction is above a threshold, another image can be received from the sensor and rendered on the wearable device's display. The gaze may or may not be redrawn.

[0029] As shown in Figure 3, screen 305 may include reticles 320-1 and 320-2. Reticles 320-1 and 320-2 can be used to identify targets (e.g., with minimal ambiguity). Reticles 320-1 and 320-2 are illustrated as rectangles. However, reticles 320-1 and 320-2 can be any shape (e.g., square, circular, elliptical, and / or similar). User 105 can move reticles 320-1 and 320-2 on screen 305 (e.g., by head and / or eye movements) to select (or assist in selection) candidate targets more accurately (e.g., reducing the likelihood of incorrect targeting). Alternatively (or additionally), user 105 can move the image along with reticles 320-1 and 320-2 that remain in place on screen 305. For example, reticle 320-1 is in a first position and reticle 320-2 is in a second position. In an exemplary embodiment, the second position may be a preferred position for selecting a target and / or candidate target. Thus, user 105 can move the position of reticle 320-1 to the position of reticle 320-2 on screen 305 (for example, by head and / or eye movements).

[0030] Estimating gaze can include processing the gaze of a user 105 who is in a fixed position relative to the wearable device 110. For example, the gaze direction can be collinear with the user 105's head in the case of a head-mounted wearable device 110. However, the gaze direction can also be based on the field of view direction and the head direction. In an exemplary embodiment, the reticle 320 can be forced to fix the field of view direction (e.g., the eyes focus on the position of the reticle 320). For example, the screen 305 can be a pass-through display such that the reticle 320 and the gaze 235 drawn on the screen 305 can intersect, and the gaze direction can be collinear with the user 105's head gaze direction when selecting an object as a candidate target.

[0031] Calibration can be used to align the screen 305 and the camera 250 (for example, to align the centers of the screen 305 and the camera 250). For example, circle 310 can represent the center vector of camera 250, and the point of line of sight 235 can represent the center vector of screen 305. The offset line 315 can represent the distance and direction of the offset between the center vector of camera 250 and the center vector of screen 305. Thus, calibration can be used to shift the image received from camera 250 and displayed on screen 305 based on the offset (represented by the offset line 315). Alternatively, calibration can be used to shift the line of sight 235 and / or reticle 320 displayed on screen 305 based on the offset (represented by the offset line 315). Calibration can include, for example, computer-aided design (CAD) calibration, factory calibration, in-field user calibration, and / or in-field automated calibration. Figure 4 shows a block diagram of the calibration data flow according to an exemplary embodiment.

[0032] CAD calibration 405 can be determined during the CAD design of the wearable device 110. For example, CAD calibration can be based on the orientation and positioning between the camera 250 and a given user's line of sight (e.g., the average user's line of sight) once the wearable device 110 is designed using, for example, CAD software tools. Factory calibration 410 can be a process or action that brings about adjustment of the calibration (e.g., CAD calibration) after a particular wearable device 110 has been manufactured. Factory calibration can take into account the actual positioning between the camera 250 and a given user's line of sight (e.g., the average user's line of sight), which is determined for that particular wearable device 110.

[0033] In-field user calibration 415 can be a process or operation that allows user 105 to further adjust the calibration by performing a series of predetermined steps. In-field user calibration can be performed before user 105's first use. For example, user 105 can use an object that is easily recognizable by a computer vision algorithm as a calibration marker. For example, the object can be printed on a product box, or the product box can be used as the object. User 105 can place the object in user 105's environment at a distance (e.g., more than 1 meter). User 105 can start the calibration software while fixating on the object. In-field user calibration 420 can be repeated multiple times to improve accuracy. In-field automatic calibration can be a process or operation that further adjusts the calibration during use. For example, a discrepancy between the center of the detected object and the line-of-sight estimate can be used as corrective feedback to the calibration process or operation. After calibration, target identification can be performed by eye-tracking (in-field automatic calibration while target identification is performed by eye-tracking).

[0034] Figure 5 shows a block diagram of eye-tracking using a target identification data flow according to an exemplary embodiment. As shown in Figure 5, the data flow includes a near-field block 505, a far-field block 510, a reticle block 515, an eye-line adjustment block 520, an eye-tracking block 525, and an identification target block 530.

[0035] The near field of view 505 can be configured to estimate the user's line of sight and / or direction of sight in the near field of view (e.g., within 1 meter of the wearable device). In the near field of view, the user's line of sight can extend over a significant portion of the image. Therefore, many objects can be within the user's field of view. Some objects may be further away than those considered to be in the near field of view. In other words, the user can see objects in the displayed image even if the objects are far away (e.g., more than 1 meter from the wearable device). For example, referring to Figures 1A and 1B, D1 may be 0.5 meters from the wearable device 110, and D3 may be 5 meters from the wearable device 110. Therefore, targets 115 and 120 may be in the near field of view, and target 125 may be in the far field of view. However, the user's line of sight can extend over a significant portion of the image displayed on the screen 305 (including the objects as targets). Therefore, targets 115, 120, and 125 may appear to be within the user's line of sight and possibly in the near field of view.

[0036] In an exemplary embodiment, the wearable device 110 may include a depth sensor (or another method for determining depth). Thus, the image displayed on the screen 305 may have an associated depth map (or other depth information). The depth map can be used to estimate the user's line of sight and / or direction of sight, as well as a subset of targets (e.g., objects) in the near field of view. For example, the near field of view may be predetermined as less than 1 meter (relative to the wearable device 110). Thus, the depth map can be used to exclude target 125 from the subset of targets.

[0037] The far field of view 510 can be configured to estimate the user's line of sight direction in the far field of view (e.g., more than 1 meter from the wearable device). In the far field of view, the user's line of sight can extend to a variable portion of the image. For example, as the user's line of sight moves away from the wearable device, the portion of the image that the user's line of sight extends to can become smaller and smaller. Therefore, the further away the user is fixated, the fewer objects can be in the user's field of view. Thus, estimating the user's line of sight and / or line of sight direction can be more accurate (compared to estimating in the near field of view). As described above, the wearable device 110 may include a depth sensor, and the image displayed on the screen 305 may have a corresponding depth map (or other depth information). Therefore, the line of sight 235 drawn on the screen 305 can use the depth map to improve the accuracy of target identification. Furthermore, the depth along the line of sight (e.g., metric depth) can be determined using one of the calibration techniques described above. Therefore, identifying a target that intersects with line of sight 235 may also include determining and / or estimating the depth of the identified target.

[0038] The reticle 515 can be configured to forcibly fix the field of view direction (for example, the eyes focus on the position of the reticle). For example, the screen 305 can be a pass-through display such that the reticle 320 and the line of sight 235 drawn on the screen 305 can intersect, and the line of sight direction can be collinear with the user 105's head line of sight when selecting an object as a candidate target. Estimating the line of sight direction may include determining that the user's line of sight is fixed on the wearable. For example, in the case of a head-mounted wearable device, the line of sight direction may be collinear with the user's head. However, humans tend to fixate with their heads as well as their eyes. Therefore, the reticle 516 can be configured to forcibly fix the line of sight direction.

[0039] The gaze adjustment 520 can be configured to fix the gaze direction based on the correlation between the wearable device's position, device trajectory, and the user's gaze. For example, when the user looks at a sign (e.g., upward), the gaze adjustment can be estimated by monitoring the user's eyes. A model takes the wearable's position in space up to a known range (3dof or 6dof) and past trajectories as input and generates a correction for the gaze estimate. The model can be a machine learning model (e.g., a trained neural network) or an algorithm used to compute an offset (based on the current gaze direction). The reticle 515 and the gaze adjustment 520 can be used together and / or separately.

[0040] Eye tracking 525 can be configured to track the user's rotational gaze (e.g., using head movements). For example, the initiation of rotational tracking (e.g., 3DoF movement) may involve the use of sensors associated with the wearable device. Sensors may include motion sensors (e.g., inertial measurement units (IMUs)) that provide linear acceleration and rotational velocity. Changes in the wearable's orientation can be translated into changes in the gaze position in image space, allowing the user to sequentially select from detected objects. Eye tracking 525 can be performed without capturing or rendering a new image. Rotational tracking can also be used to re-trigger capture and detection if the user's current gaze moves outside the currently displayed image. In exemplary embodiments, it may not be sensitive to small translational motions, where small is relative to the distance to the object of interest. However, in the case of large translational motions (e.g., movement exceeding a threshold), the current set of detected targets (e.g., objects) can be discarded, and the capture and rendering of a new image can be triggered along with target identification. Depth information can also be used to reproject objects as the wearable device moves without capturing a new image.

[0041] The identification target 530 can be configured to identify potential targets. The line of sight direction and / or line of sight can be used to identify a subset of targets based on the region encompassing the line of sight. The depth associated with each target in the set of targets can be estimated. Candidate target estimation can be based on the intersection of the line of sight at the depths contained in the region, and / or targets in the subset of targets. In exemplary embodiments, actions can be triggered (e.g., convert text, find the best price, read a map, get directions, identify an image, read a product label, identify a storefront, identify a restaurant, read a menu, identify a building, identify a product, identify something natural (e.g., a plant, flower, tree, and / or similar)), and / or, in response to the trigger, candidate targets can be estimated (or selected) based on a subset of targets. For example, if the action is to convert text, the text can be associated with candidate targets.

[0042] Figure 6 shows a block diagram of a method for identifying a target according to an exemplary embodiment. As shown in Figure 6, in step S605, an image is received from a sensor of the wearable device. For example, camera 250 can capture an image (or multiple images) representing a real-world scene. The image (or one of multiple images) can be rendered on screen 305.

[0043] In step S610, a set of targets for the image is identified. For example, an image may contain multiple objects. An object, or a subset of an object, can be selected as the set of targets. In step S615, the gaze direction associated with the user of the wearable device is tracked. For example, the gaze (e.g., gaze 235) can be drawn on the screen. The gaze can be used to track the user's gaze.

[0044] In step S620, a subset of targets is identified from a set of targets within the image region based on the line of sight. For example, a subset of targets can be identified based on the region encompassing the line of sight.

[0045] In step S625, the system receives a command that triggers an action. For example, an action could be triggered (e.g., convert text, find the best price, get directions, and / or something similar). For example, if the action is to convert text, the text can be associated with a candidate target. The command can be a voice command, a gesture, contact with a wearable device, and / or something similar.

[0046] In step S630, in response to an instruction that triggers an action, candidate targets are identified (determined, estimated, and / or similar) from a subset of targets. For example, in response to a trigger, candidate targets can be estimated (or selected) based on a subset of targets. The depth associated with each target in the set of targets can be estimated. The estimation of candidate targets can be based on line-of-sight intersections at depths included in the region. A candidate target can be one of a subset of targets selected based on line-of-sight intersections at depths included in the region.

[0047] Figure 7 shows a block diagram of a system according to an exemplary embodiment. In the example of Figure 7, it should be understood that the system (e.g., a wearable device) may include, or be associated with, a computing system or at least one computing device (e.g., a mobile computing device, a cell phone, a laptop computer, a tablet, etc.) and virtually represent any computing device configured to perform the techniques described herein. Thus, it should be understood that the system may include various components that can be used to implement the techniques described herein or different or future versions thereof. As an example, the system may include a processor 705 and memory 710 (e.g., non-temporary computer-readable memory). The processor 705 and memory 710 may be connected (e.g., communicateable) by a bus 715.

[0048] The processor 705 may be used to execute instructions stored in at least one memory 710. Therefore, the processor 705 can implement various features and functions described herein, or additional or alternative features and functions. The processor 705 and at least one memory 710 may be used for various other purposes. For example, at least one memory 710 may represent examples of various types of memory and associated hardware and software that can be used to implement any one of the modules described herein.

[0049] At least one memory 710 may be configured to store data and / or information associated with the device. At least one memory 710 may be a shared resource. Thus, at least one memory 710 may be configured to store data and / or information associated with other elements in a larger system (e.g., image / video processing or wired / wireless communication). The techniques described herein can be implemented by using the processor 705 and at least one memory 710 together. Thus, the techniques described herein can be implemented as code segments (e.g., software) stored in memory 710 and executed by the processor 705. Thus, memory 710 may include a calibration 400 block, a near-field 505 block, a far-field 510 block, a reticle 515 block, a gaze adjustment 520 block, a gaze tracking 525 block, and an identification target 530 block. In one or more exemplary embodiments, a subset of the components shown as being included in memory 710 may be used. For example, memory 710 may include a calibration 400 block without other components.

[0050] Figure 8 shows examples of computer devices 800 and mobile computer devices 850 that may be used in conjunction with the techniques described herein (for example, to implement wearable devices). Computing device 800 includes a processor 802, memory 804, a storage device 806, a high-speed interface 808 connected to memory 804 and a high-speed expansion port 810, and a low-speed bus 814 and a low-speed interface 812 connected to storage device 806. Each component 802, 804, 806, 808, 810, and 812 is interconnected using various buses and may be mounted on a common motherboard or in other configurations as needed. The processor 802 processes instructions for execution within computing device 800, including instructions stored in memory 804 or storage device 806, to display graphical information of a GUI on an external input / output device such as a display 816 connected to the high-speed interface 808. In other embodiments, multiple processors and / or multiple buses may be used as needed, along with multiple memories and memory types. Additionally, multiple computing devices 800 may be connected, with each device performing some of the necessary operations (for example, as a server bank, a group of blade servers, or a multiprocessor system).

[0051] Memory 804 stores information within the computing device 800. In one embodiment, memory 804 is a volatile memory unit; in another embodiment, memory 804 is a non-volatile memory unit. Memory 804 may also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0052] The storage device 806 can provide high-capacity storage to the computing device 800. In one embodiment, the storage device 806 may be or include a computer-readable medium such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, flash memory or other similar solid-state memory device, or a storage area network or other configuration device. A computer program product can be tangibly embodied in an information carrier. A computer program product may also include instructions that perform one or more of the above-described methods at runtime. The information carrier is a computer-readable or machine-readable medium such as memory 804, the storage device 806, or the memory of the processor 802.

[0053] The high-speed controller 808 manages bandwidth-intensive operations of the computing device 800, while the low-speed controller 812 manages bandwidth-intensive operations. Such function assignments are merely illustrative. In one embodiment, the high-speed controller 808 is connected to memory 804, a display 816 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 810 that can accept various expansion cards (not shown). In another embodiment, the low-speed controller 812 is connected to a storage device 806 and a low-speed expansion port 814. The low-speed expansion port may include various communication ports (e.g., USB, Bluetooth®, Ethernet®, Wireless Ethernet) and may be connected to one or more input / output devices, such as a keyboard, pointing device, scanner, or network devices such as switches or routers, for example, via a network adapter.

[0054] The computing device 800 may be implemented in many different forms, as shown in the figure. For example, it may be implemented as a standard server 820, or it may be implemented multiple times in a group of such servers. It may also be implemented as part of a rack server system 824. Additionally, it may be implemented in a personal computer such as a laptop computer 822. Alternatively, the components of computing device 800 may be combined with other components of a mobile device (not shown), such as device 850. Each of such devices may contain one or more computing devices 800, 850, and the entire system may consist of multiple computing devices 800, 850 communicating with each other.

[0055] The computing device 850 includes, among other components, a processor 852, memory 864, input / output devices such as a display 854, a communication interface 866, and a transceiver 868. The device 850 may also be provided with additional storage devices, such as a microdrive or other storage devices. The components 850, 852, 864, 854, 866, and 868 are interconnected using various buses, and some components may be mounted on a common motherboard, or in other configurations as needed.

[0056] The processor 852 can execute instructions within the computing device 850, including instructions stored in memory 864. The processor may be implemented as a chipset of chips including multiple separate analog and digital processors. The processor may also provide coordination of other components of the device 850, such as user interface control, applications run by the device 850, and wireless communication by the device 850.

[0057] The processor 852 may communicate with the user via a control interface 858 and a display interface 856 connected to the display 854. The display 854 may be, for example, a thin-film transistor liquid crystal display (TFT LCD), a light-emitting diode (LED) or organic light-emitting diode (OLED) display, or other suitable display technology. The display interface 856 may include appropriate circuitry for driving the display 854 to present graphic information and other information to the user. The control interface 858 may receive commands from the user and translate them for submission to the processor 852. Furthermore, an external interface 862 communicating with the processor 852 may be provided to enable short-range communication between device 850 and other devices. The external interface 862 may, for example, provide wired communication in some implementations, or wireless communication in other implementations, and may use multiple interfaces.

[0058] Memory 864 stores information within the computing device 850. Memory 864 can be implemented as one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Alternatively, an expansion memory 874 may be provided and connected to the device 850 via an expansion interface 872. The expansion interface 872 may include, for example, a single in-line memory module (SIMM) card interface. Such an expansion memory 874 may provide extra storage space to the device 850, or it may store applications or other information for the device 850. Specifically, the expansion memory 874 may include instructions for performing or supplementing the above operations, and may also include secure information. Therefore, for example, the expansion memory 874 may be provided as a security module for the device 850 and may be programmed with instructions that enable secure use of the device 850. Furthermore, a secure application may be provided via the SIMM card, along with additional information such as identification information placed on the SIMM card in a hack-proof manner.

[0059] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that perform one or more of the above-described methods at runtime. The information carrier is a computer-readable or machine-readable medium, such as memory 864, extended memory 874, or memory on processor 852, and this information carrier may be received, for example, via transceiver 868 or external interface 862.

[0060] Device 850 may perform wireless communication via a communication interface 866, which may include digital signal processing circuitry if necessary. The communication interface 866 can provide communication in various modes or protocols, including, among others, GSM® voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA®, CDMA2000, or GPRS. Such communication may be performed, for example, via a radio frequency transceiver 868. Furthermore, short-range communication may be performed using Bluetooth, WiFi, or other such transceivers (not shown). Additionally, a Global Positioning System (GPS) receiver module 870 may provide device 850 with additional navigation and location-related radio data, which may be used as needed by applications running on device 850.

[0061] Device 850 may also use audio codec 860 for voice communication. This audio codec 860 may receive voice information from the user and convert it into usable digital information. Similarly, audio codec 860 may also generate sounds audible to the user through a speaker (for example, in the handset of device 850). Such sounds may include sounds from voice telephone calls, recorded sounds (for example, voice messages, music files, etc.), and sounds generated by applications running on device 850.

[0062] The computing device 850 can be implemented in many different forms, as shown in the figure. For example, it may be implemented as a mobile phone 880. Alternatively, it may be implemented as part of a smartphone 882, a personal digital assistant, or other similar mobile device.

[0063] Various implementations of the systems and techniques described herein can be realized in digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor, at least one input device, and at least one output device, which may be specialized or general-purpose, connected to receive data and instructions from and transmit data and instructions to a storage system.

[0064] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic circuits (PLDs)) used to provide machine instructions and / or data to a programmable processor that contains a machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0065] To provide user interaction, the systems and techniques described herein can be implemented in a computer having a display device (LED (light-emitting diode), OLED (organic OLED), or LCD (liquid crystal display) monitor / screen) for displaying information to the user, and a keyboard and pointing device (e.g., mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide user interaction; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and the input from the user can be received in any form, such as acoustic, spoken language, or haptic input.

[0066] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., data servers), middleware components (e.g., application servers), or frontend components (e.g., a client computer having a graphical user interface or web browser through which a user can interact with the implementation of the systems and techniques described herein), or in a combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), and the Internet.

[0067] A computing system can include clients and servers. Clients and servers are generally far apart from each other and typically interact through a communication network. The client-server relationship arises from computer programs that run on each computer and have a client-server relationship with each other.

[0068] In some embodiments, the computing device shown in the figure may interface with an AR headset / HMD device 890 and include sensors that generate an augmented environment for viewing content inserted into physical space. For example, one or more sensors included in the computing device 850 shown in the figure or in other computing devices may provide input to the AR headset 890, or in general, provide input to the AR space. Sensors may include, but are not limited to, touchscreens, accelerometers, gyroscopes, pressure sensors, biometric sensors, temperature sensors, humidity sensors, and ambient light sensors. The computing device 850 may use the sensors to determine the absolute position and / or detected rotation of the computing device in the AR space, which can then be used as input to the AR space. For example, the computing device 850 may be incorporated into the AR space as a virtual object such as a controller, laser pointer, keyboard, or weapon. When incorporated into the AR space, the user's positioning of the computing device / virtual object allows the user to position the computing device to view the virtual object in a particular way in the AR space. For example, if the virtual object represents a laser pointer, the user can operate the computing device as if it were a real laser pointer. Users can move the computing device left and right, up and down, in circles, etc., and use the device in a similar way to using a laser pointer. In some embodiments, users can use a virtual laser pointer to point to a target location.

[0069] In some embodiments, one or more input devices included in or connected to the computing device 850 can be used as inputs to the AR space. Input devices may include, but are not limited to, a touchscreen, keyboard, one or more buttons, trackpad, pointing device, mouse, trackball, joystick, camera, microphone, earphones or buds with input capabilities, game controller, or other connectable input devices. When the computing device is integrated into the AR space, a user can interact with the input devices included in the computing device 850 to trigger specific actions in the AR space.

[0070] In some embodiments, the touchscreen of the computing device 850 can be rendered as a touchpad in the AR space. The user can interact with the touchscreen of the computing device 850. The interaction is rendered, for example, as movement on the rendered touchpad in the AR space in the AR headset 890. The rendered movement allows control of virtual objects in the AR space.

[0071] In some embodiments, one or more output devices included in the computing device 850 can provide output and / or feedback to the user of the AR headset 890 in the AR space. The output and / or feedback can be visual, tactical, or audible. The output and / or feedback may include, but is not limited to, vibration, turning one or more lights or strobes on and off, flashing and / or flashing, sounding an alarm, playing a chime, playing a song, and playing an audio file. The output devices may include, but are not limited to, a vibration motor, a vibration coil, a piezoelectric device, an electrostatic device, a light-emitting diode (LED), a strobe, and a speaker.

[0072] In some embodiments, the computing device 850 may be displayed as another object in a computer-generated 3D environment. User interactions with the computing device 850 (e.g., rotating, shaking, touching the touchscreen, swiping a finger on the touchscreen) can be interpreted as interactions with objects in the AR space. In the example of a laser pointer in the AR space, the computing device 850 appears as a virtual laser pointer in the computer-generated 3D environment. When the user interacts with the computing device 850, the user in the AR space sees the movement of the laser pointer. The user receives feedback from their interaction with the computing device 850 in the AR environment on the computing device 850 or the AR headset 890. The user's interaction with the computing device can be translated into an interaction with a user interface generated in the AR environment for the controllable device.

[0073] In some embodiments, the computing device 850 may include a touchscreen. For example, a user can interact with the touchscreen to interact with the user interface of the controllable device. For example, the touchscreen may include user interface elements such as sliders that can control the characteristics of the controllable device.

[0074] Computing device 800 is intended to represent a variety of digital computers and devices, including, but not limited to, laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 850 is intended to represent a variety of mobile devices, including personal digital assistants, mobile phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the implementation of the inventions described and / or claimed herein.

[0075] Many embodiments have been described. Needless to say, it should be understood that various modifications can be made without departing from the spirit and scope of this specification.

[0076] Furthermore, the logic flow shown in the figure does not require a specific sequence, i.e., a sequential order, to achieve the desired result. Additionally, other steps may be added to the described flow, or steps may be removed from the described flow, and other components may be added to the described system, or other components may be removed from the described system. Therefore, other embodiments are within the scope of the following claims.

[0077] In addition to the above description, the systems, programs, or functions described herein may provide the user with controls that allow the user to choose whether and when user information (e.g., information about the user's social networks, social actions, or activities, occupation, user preferences, or user's current location) may be collected, and whether user content or communications are sent from the server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, user identity may be processed so that personally identifiable information cannot be identified, or if location information is obtained (e.g., at the city, zip code, or state level), the user's geographical location may be generalized so that the user's specific location cannot be identified. Thus, users can control what information is collected about them, how that information is used, and what information is provided to them.

[0078] While specific features of the embodiments described herein have been illustrated, many modifications, substitutions, alterations, and equivalents will be conceivable to those skilled in the art. It should be understood that the appended claims are intended to encompass all such modifications and alterations that fall within the scope of the embodiments. These are presented only as examples and not as limitations, and various variations in form and detail may be made. Any part of the apparatus and / or method described herein may be combined in any combination except mutually exclusive combinations. The embodiments described herein may include various combinations and / or partial combinations of the functions, components, and / or features of the different embodiments described herein.

[0079] The exemplary embodiments may include a variety of modifications and alternative forms, which are shown as examples in the drawings and described in detail herein. However, it should be understood that the exemplary embodiments are not limited to any particular form disclosed, but rather encompass all modifications, equivalents, and alternatives included in the claims. Similar figures refer to similar components throughout the description of the drawings.

[0080] Some of the exemplary embodiments described above are explained as processes or methods shown as flowcharts. While flowcharts describe operations as sequential, many operations may occur in parallel, concurrently, or simultaneously. The order of operations may also be changed. A process may terminate when its operations are completed, but it may have additional steps not shown in the diagram. A process may correspond to a method, function, procedure, subroutine, subprogram, etc.

[0081] Some of these are illustrated by flowcharts. The methods described above can be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segment for performing the required task may be stored in a machine- or computer-readable medium such as a storage medium. A processor(s) may perform the required task.

[0082] The specific structural and functional details disclosed herein are for illustrative purposes only. However, the exemplary embodiments are embodied in many alternative forms and should not be construed as being limited only to the embodiments described herein.

[0083] Terms such as "first," "second," etc., may be used herein to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. As used herein, the terms and / or include any combination and all combinations of one or more of the items described relating to them.

[0084] When an element is said to connect to or join another element, it should be understood that there may be other elements that directly connect to or join to that element, or that intersect with it. In contrast, when an element is said to be directly connected to or joined to another element, there are no intervening elements. Other words used to indicate relationships between elements should be interpreted similarly (e.g., between and directly between, adjacent and directly adjacent).

[0085] The terms used herein are for the sole purpose of describing specific embodiments and are not intended to limit the exemplary embodiments. Where used herein, the singular forms a, an, and the are also intended to include the plural forms unless the context otherwise explicitly indicates. It should be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, identify the presence of the described feature, integer, step, action, element, and / or component, but do not exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof.

[0086] Furthermore, it should be noted that in some alternative embodiments, the functions / actions shown may occur in a different order than that shown in the diagrams. For example, two diagrams shown consecutively may actually be performed in reverse order, depending on the functions / actions that are performed simultaneously or involved.

[0087] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as those generally understood by those skilled in the art to which this invention belongs. Furthermore, terms defined in commonly used dictionaries, for example, should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and it will be further understood that, unless expressly defined herein, they should not be interpreted in an idealized or overly formal sense.

[0088] The above exemplary embodiments and corresponding detailed descriptions are presented with respect to software, or algorithms, and symbolic representations of operations on data bits in computer memory. These descriptions and representations are intended to effectively convey the nature of those operations to those skilled in the art. An algorithm, if the term is used herein and in general usage, is considered to be a set of self-consistent steps leading to a desired result. These steps require the physical manipulation of physical quantities. Usually, but not always, these quantities take the form of optical, electrical, or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated. For reasons of general use, it is sometimes convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.

[0089] In the exemplary embodiments described above, references to symbolic representations (e.g., in the form of flowcharts) of actions and behaviors that can be implemented as program modules or functional processes include routines, programs, objects, components, data structures, etc., which perform a particular type of task or a particular type of abstract data and can be described and / or implemented using existing hardware on existing structural elements. Such existing hardware may include one or more central processing units (CPUs), digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate array (FPGA) computers, etc.

[0090] However, it should be recognized that all these and similar terms should be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise specified, or as is evident from the description, terms such as processing, arithmetic, calculating, and determining displays refer to the actions and processes of a computer system or similar computing device that manipulate and convert data represented as physical quantities, electronic quantities, in the registers and memory of a computer system into other data similarly represented as physical quantities in computer system memory or registers, or other such information storage, transmission devices, or display devices.

[0091] It should also be noted that the software implementations of exemplary embodiments are typically encoded in some form of non-temporary program storage medium or implemented via some type of transmission medium. The program storage medium may be magnetic (e.g., floppy disk or hard drive) or optical (e.g., compact disk read-only memory, i.e., CD-ROM), and may be read-only or random-access. Similarly, the transmission medium may be twisted-pair wire, coaxial cable, optical fiber, or other suitable transmission medium known in the art. The exemplary embodiments are not limited to these aspects of any given embodiment.

[0092] Finally, while the attached claims describe specific combinations of the features described herein, it should be noted that the scope of this disclosure is not limited to the specific combinations claimed below, but rather extends to encompass any combination of the features or embodiments disclosed herein, regardless of whether such specific combinations are specifically enumerated in the attached claims at this time.

Claims

1. Receiving images from sensors in wearable devices, Identifying a set of targets in an image rendered on the display of the wearable device, To estimate the depth associated with each target in the set of targets, Based on the gaze direction associated with the user of the wearable device, rendering the gaze direction onto the rendered image, Based on the line of sight, identify a subset of targets based on the set of targets within the region of the image, Triggering an action, A method comprising: estimating candidate targets based on a subset of targets by excluding one or more targets from the subset of targets based on the depth associated with each target in response to the trigger.

2. Identifying a subset of the target based on the region encompassing the line of sight, It further includes, The method according to claim 1, wherein the estimation of the candidate target is based on the intersection of the lines of sight at the depth included in the region.

3. Detecting changes in line of sight, The determination that the aforementioned change is below the threshold, Re-rendering the image on the display of the wearable device, The method according to claim 1, further comprising:

4. Detecting changes in line of sight, The determination that the aforementioned change is below the threshold, Re-rendering the aforementioned line of sight, The method according to claim 1, further comprising:

5. Detecting changes in line of sight, The determination that the aforementioned change is present in the rendered image and is closer to the subset of the target, Re-rendering the aforementioned line of sight using color changes, The method according to claim 1, further comprising:

6. Detecting changes in line of sight, The determination that the aforementioned change is greater than the threshold, Receiving another image from the aforementioned sensor, The method according to claim 1, further comprising:

7. The method according to claim 1, further comprising rendering a reticle on the rendered image based on the position of the candidate target.

8. The method of claim 7, further comprising rearranging the reticle to a different position on the rendered image, wherein the candidate target is estimated based on the rearranged reticle.

9. The method according to claim 1, further comprising calibrating the wearable device based on the position of the sensor of the wearable device and the center of the display of the wearable device.

10. Image sensor and, The display and At least one processor, A wearable device comprising at least one memory containing computer program code, The at least one memory and the computer program code are transmitted to the wearable device using the at least one processor. Receiving an image from the aforementioned image sensor, Identifying a set of targets in the image rendered on the aforementioned display, To estimate the depth associated with each target in the set of targets, Based on the gaze direction associated with the user of the wearable device, rendering the gaze direction onto the rendered image, Based on the line of sight, identify a subset of targets based on the set of targets within the region of the image, Triggering an action, In response to the trigger, candidate targets are estimated based on the subset of targets by excluding one or more targets from the subset of targets based on the depth associated with each target, A wearable device configured to perform a certain action.

11. The aforementioned computer program code further provides the wearable device with: Identifying a subset of the target based on the region encompassing the line of sight, Have them do it, The wearable device according to claim 10, wherein the estimation of the candidate target is based on the intersection of the lines of sight at the depth included in the region.

12. The aforementioned computer program code further provides the wearable device with: Detecting changes in line of sight, The determination that the aforementioned change is below the threshold, Re-rendering the image on the display of the wearable device, A wearable device according to claim 10, which enables the following:

13. The aforementioned computer program code further provides the wearable device with: Detecting changes in line of sight, The determination that the aforementioned change is below the threshold, Re-rendering the aforementioned line of sight, A wearable device according to claim 10, which enables the following:

14. The aforementioned computer program code further provides the wearable device with: Detecting changes in line of sight, The determination that the aforementioned change is present in the rendered image and is closer to the subset of the target, Re-rendering the aforementioned line of sight using color changes, A wearable device according to claim 10, which enables the following:

15. The aforementioned computer program code further provides the wearable device with: Detecting changes in line of sight, The determination that the aforementioned change is greater than the threshold, Receiving another image from the aforementioned image sensor, A wearable device according to claim 10, which enables the following:

16. The aforementioned computer program code further provides the wearable device with: The wearable device according to claim 10, which causes the device to render a reticle on the rendered image based on the position of the candidate target.

17. The aforementioned computer program code further provides the wearable device with: The wearable device according to claim 16, wherein the reticle is rearranged to a different position on the rendered image, and the candidate target is estimated based on the rearranged reticle.

18. The aforementioned computer program code further provides the wearable device with: The wearable device according to claim 10, wherein the wearable device is calibrated based on the position of the image sensor of the wearable device and the center of the display of the wearable device.

19. A computer program including instructions, wherein when the instructions are executed by at least one processor, the at least one processor is configured to: Receiving images from sensors in wearable devices, Identifying a set of targets in an image rendered on the display of the wearable device, To estimate the depth associated with each target in the set of targets, Based on the gaze direction associated with the user of the wearable device, rendering the gaze direction onto the rendered image, Based on the line of sight, identify a subset of targets based on the set of targets within the region of the image, Triggering an action, In response to the trigger, candidate targets are estimated based on the subset of targets by excluding one or more targets from the subset of targets based on the depth associated with each target, A program that causes something to happen.

20. The instruction further provides to at least one processor: Identifying a subset of the target based on the region encompassing the line of sight, Have them do it, The program according to claim 19, wherein the estimation of the candidate target is based on the intersection of the lines of sight at the depth included in the region.

Citation Information

Patent Citations

  • Information processing method and information processor

    JP2008293357A

  • Program and image formation device

    JP2017058971A

  • Head-mounted display device, program, and method for controlling head-mounted display device

    JP2018200415A

  • Math operations in mixed or virtual reality

    US20180061132A1