Method and system for gaze-based control of mixed reality content

By using visual guideline technology in a mixed reality head-mounted display, and addressing the limitations of gaze control and hardware optics, the problems of cumbersome object control and difficult depth positioning in existing devices are solved, enabling fast and natural MR content enhancement display.

CN112262361BActive Publication Date: 2026-03-17INTERDIGITAL VC HOLDINGS INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-04-12
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing mixed reality head-mounted display (HMD) devices are cumbersome and unnatural to operate when controlling and displaying augmented reality/mixed reality objects, making it difficult to quickly and accurately locate the depth position of objects, and the limited focal plane of the optical devices causes eye fatigue for users.

Method used

Using visual guideline technology, user movements are guided in a 3D map through gaze control. The position and size of the guideline are determined by hardware optical limitations and real or MR objects in the user's view, allowing the augmented MR content to move along the line, achieving fast and natural object control.

Benefits of technology

It provides a fast and natural way to control the display of MR content-enhanced objects, utilizing precise focusing of hardware optical positioning to reduce user operation steps and improve the accuracy and comfort of object positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112262361B_ABST
    Figure CN112262361B_ABST
Patent Text Reader

Abstract

Systems and methods for discovering content and localizing it into an augmented reality space are presented. One method includes forming a three-dimensional (3D) map of a user's surrounding environment of an augmented reality (AR) head-mounted display (HMD), determining a depth directional position of a user's point of gaze based on an eye gaze direction and an eye vergence, determining a visual guide line path in the 3D map, guiding a user's action along the visual guide line path at one or more identified focal points, and rendering a mixed reality (MR) object along the visual guide line path at a location corresponding to the user's gaze direction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is a non-provisional application filed with U.S. Provisional Patent Application Serial No. 62 / 660,428 (filed April 20, 2018) entitled “Method and System for Gaze-Based Control of Mixed Reality Content”, and claims the benefit under 35 U.S. SC §119(e), the entire contents of which are incorporated herein by reference. Background Technology

[0003] In the realm of mixed reality (MR) head-mounted displays (HMDs) or goggles, there are interaction challenges that need to be addressed. Specifically, HMD users need to manipulate and control mixed reality / augmented reality (MR / AR) objects embedded in their environment. Enhanced control can involve providing additional information, as well as presenting information within the user's field of view, while optimizing the view based on the target device specifications, such as focusing on an accurate viewing plane and resolution.

[0004] Currently, the user interface activities required to view the details of AR / MR objects, such as selecting an object and moving it closer to obtain a detailed view, are typically accomplished through a combination of modality and pose, such as selecting head orientation and manipulating gestures. For example, to move an object closer to Microsoft... TM HoloLens TM The user must perform the following steps: 1. Turn his head toward the "Adjust" tool icon on the corner of the object to be moved; 2. Activate the adjustment tool by tapping with his finger; 3. Turn his head toward the "Drag to Move" icon that appears on the object; 4. Begin moving the object by performing a tap and hold finger gesture; 5. Move the object by moving his hand; 6. Release the tap and hold finger gesture to put the object down; 7. Turn his head toward the "Done" icon on the object; and 8. Perform a tapping gesture to activate the "Done" icon.

[0005] Although finger tapping and hand movement gestures can be replaced by using a handheld gyroscope-based "clicker" controller, this method is both laborious and noticeable, and difficult in public places.

[0006] The use of eye gaze for more natural interaction has been extensively studied, initially as an input method for people with disabilities. Gaze tracking systems are also being incorporated into AR / MR HMDs; for example, Google's Eyefluence. TM We are working with HMD manufacturers to build a gaze-based system.

[0007] Many interaction methods using gaze control have been investigated, such as gaze pointing, gaze pose, dwell-based selection, multimodal selection, selection by following moving objects, drag and drop, rotation control and sliders, window switching, image annotation, reading, and attention focus. An HMD-friendly approach for multimodal object control, titled "An Interaction Technique Using Gaze and Head Rotation," has been reported in the evaluation of head rotation. This was presented at the 9th NordiCHI 2016 conference and in the proceedings of enhanced gaze interaction using simple head poses. It was also presented at the ACM UbiComp 2012 conference, and in other interaction methods.

[0008] Of particular interest are the functions that focus on quickly inspecting, selecting, and manipulating (moving) objects. This function is similar to drag-and-drop using eye gaze and has been described in at least a few approaches, such as “Using Eye Movement in Human-Computer Interaction Technology: What You See Is What You Get” (ACM Journal of Information Systems, Vol. 9, p. 2 (ACM Trans. Inf. Syst. 9, 2), Jacob, April 1991) and “Gaze-Based Interaction for Virtual Environments” (J. UCS, 14(19), 3085-3098, Jimenez, Gutierrez, D., Latorre, 2008). However, the few reported studies focus only on 2D motion and fail to consider depth aspects, such as the movement of objects in 3D space that would be required on MR HMD devices. Furthermore, the approach requires something similar to the HoloLens described above. TM The example has multiple steps.

[0009] As MR content enhancement becomes more common, MR HMD users will face difficulties in controlling how and when that content is displayed. Such content enhancement can essentially be associated with any real-world or virtual content a user sees, such as streetlights, traffic signs, shop signs and advertisements in shop windows, public notices, people, vehicles, etc.

[0010] Another problem with current solutions is that some parts of the user's view are better suited for displaying content enhancements than others. For example, real and virtual objects of varying sizes occupy the user's view, and the user may be moving. Further limitations can stem from hardware; gaze recognition and optical display resolution can impose requirements on the localization needed to display enhanced content. It cannot be expected that the user will decide where to place content across the entire range of their view every time; current systems cannot identify the appropriate location of content and display it there, nor can they allow the user to quickly decide on the display area.

[0011] Another issue is the limited optical capabilities of HMD devices. Current devices, such as the Microsoft HoloLens, have limited optical capabilities. TM The current approach uses a single, fixed focal plane, which leads to a conflict in the convergence and divergence adaptation of MR objects that do not reside on that plane. This conflict causes eye strain and slows down the user's ability to determine the accurate depth position of objects. HMDs with multiple focal planes are expected to become commercially available in the near future. With such a device, users will benefit from the ability to control the position of content augmentation, ensuring it is precisely placed on the focal plane. Current solutions lack some form of visual guidance for focusing, so that the positioning (especially depth) of MR objects can be controlled solely through gaze.

[0012] There is a need to provide users with a quick and natural way to control the display of MR content-enhanced objects, which utilizes knowledge of the optically perfect (focus-precise) position of the hardware (HW) using only gaze.

[0013] The systems and methods described in this paper address these and other problems. Summary of the Invention

[0014] The system and method described in this paper provide an embodiment of using gaze control to bring augmented MR content related to distant objects closer to the user. The solution uses a visual guideline implemented as an MR object. This guideline contains points that help the user focus their gaze, placed at a depth equal to the focus-corrected viewing distance supported by the device hardware. The augmented MR content follows the user's gaze along this guideline, moving the content closer to or further away from the user. The position and size of this guideline are determined by the system based on HW constraints and existing real-world or MR objects in the user's view.

[0015] One or more embodiments relate to a method comprising: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); displaying a mixed reality (MR) object in the 3D map, the MR object including visual cues for content enhancement available to the user for the object; activating the content enhancements based on user input to the HMD regarding the visual cues; displaying a visual guidance line path in the 3D map; guiding the user's movements along the visual guidance line path at one or more identified focal points; and rendering the MR object along the visual guidance line path at a location corresponding to the user's gaze direction.

[0016] In one or more embodiments, activating content enhancement based on user input to the HMD includes user input of one or more of the following: gazing at the content enhancement, head pose, or gaze persistence.

[0017] In one or more embodiments, displaying a visual guideline path in a 3D graph includes displaying one or more identified focal points as multiple focal plane indicators at multiple depths within the 3D graph. In one or more embodiments, rendering an MR object along the visual guideline path at a location corresponding to the user's gaze direction includes moving the augmented object along the multiple focal plane indicators at multiple depths to magnify the MR object.

[0018] In one or more embodiments, guiding a user's actions along a visual guideline path at one or more identified focal points includes providing visual cues, wherein the visual cues include the next suggested action for the user.

[0019] In one or more embodiments, displaying a visual guideline path in a 3D diagram includes determining the depth direction of the user's gaze point based on eye gaze direction and eye convergence / divergence.

[0020] In one or more embodiments, a method includes: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); determining the depth orientation location of the user's gaze point based on eye gaze direction and eye convergence / divergence; determining a visual guide path in the 3D map; guiding the user's movements along the visual guide path at one or more identified focal points; and rendering a mixed reality (MR) object along the visual guide path at a location corresponding to the user's gaze direction.

[0021] In one embodiment, one or more focal points along the visual guide path are determined by tracking the user's gaze.

[0022] In one embodiment, one or more focal points along the visual guide path are determined based on one or more hardware limitations of the HMD.

[0023] In one embodiment, one or more focal points along the visual guide path are determined based on the distance to the MR object.

[0024] In one embodiment, one or more focal points along the visual guide path are determined by the user's movement.

[0025] In one embodiment, determining a visual guide path in a 3D image is based on the depth orientation of the user's gaze point and the available space in the 3D image. In another embodiment, determining a visual guide path in a 3D image includes forming a visual guide path to avoid one or more identified objects in the 3D image.

[0026] In one embodiment, determining the visual guide path in the 3D graph includes changing the visual guide path based on the user's movement, which includes one or more of head tilt, head pitch, head yaw, and posture.

[0027] In one embodiment, determining the visual guide path in a 3D drawing includes changing the visual guide path based on one or more rotation points determined by the user's gaze.

[0028] In one embodiment, the method further includes determining the number of points along the visual guide path based on the available focal plane in the 3D graph.

[0029] Another embodiment relates to a system including a processor and a non-transitory computer-readable storage medium storing instructions that, when executed on the processor, are operable to perform the following functions: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); determining the depth orientation location of the user's gaze point based on eye gaze direction and eye convergence / divergence; determining a visual guideline path in the 3D map; guiding the user's movements along the visual guideline path at one or more identified focal points; and rendering a mixed reality (MR) object along the visual guideline path at a location corresponding to the user's gaze direction.

[0030] In one or more embodiments of the system, one or more focal points along the visual guide path are determined by tracking the user's gaze.

[0031] In one or more embodiments of the system, one or more focal points along the visual guide path are determined according to one or more hardware limitations of the HMD.

[0032] In one or more embodiments of the system, one or more focal points along the visual guide path are determined based on the distance to the MR object.

[0033] In one or more embodiments of the system, one or more focal points along the visual guide path are determined by the user's movement.

[0034] In one or more embodiments of the system, the path of the visual guide line in the 3D graph is determined based on the depth direction position of the user's gaze point and the available space in the 3D graph.

[0035] In one or more embodiments of the system, determining a visual guide path in a 3D graph includes forming a visual guide path to avoid one or more identified objects in the 3D graph. In one or more embodiments of the system, determining a visual guide path in a 3D graph includes changing the visual guide path based on user movement, including one or more of head tilt, head pitch, head yaw, and posture.

[0036] In one or more embodiments of the system, determining the visual guide path in a 3D drawing includes changing the visual guide path based on a rotation point determined by the user's gaze.

[0037] Another embodiment of the system targets a non-transitory computer-readable storage medium for storing instructions that, when executed on a processor, are operable to perform additional functions, including determining the number of points along a visual guide path based on an available focal plane in a 3D graph.

[0038] Another embodiment relates to a system including a processor and a non-transitory computer-readable storage medium storing instructions operable, when executed on the processor, to perform the following functions: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); displaying a mixed reality (MR) object in the 3D map, the MR object including visual cues for content enhancement available to the user for the object; activating the content enhancements based on user input to the HMD regarding the visual cues; displaying a visual guide path in the 3D map; guiding the user's movements along the visual guide path at one or more recognized focal points; and rendering the MR object along the visual guide path at a location corresponding to the user's gaze direction.

[0039] In one or more embodiments of the system, activating content enhancement based on user input to the HMD includes user input of one or more of the following: gazing at the content enhancement, head pose, or gaze persistence.

[0040] In one or more embodiments of the system, displaying a visual guide path in a 3D graph includes displaying one or more identified focal points as multiple focal plane indicators at multiple depths within the 3D graph.

[0041] In one or more embodiments of the system, rendering the MR object along the visual guide path at a position corresponding to the user's gaze direction includes moving the augmented object along the plurality of focal plane indicators at the plurality of depths to magnify the MR object.

[0042] In one or more embodiments of the system, guiding a user's actions along a visual guide path at one or more recognized focal points includes providing visual cues, wherein the visual cues include the next suggested action for the user.

[0043] In one or more embodiments of the system, displaying the visual guide line path in a 3D graph includes determining the depth direction location of the user's gaze point based on eye gaze direction and eye convergence / divergence.

[0044] Another embodiment relates to a method for rendering a visual guide path, comprising: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); determining the depth direction location of the user's gaze point based on eye gaze direction and eye convergence / divergence; determining a visual guide path in the 3D map; and rendering one or more mixed reality (MR) objects along the visual guide path at a location corresponding to the user's gaze direction, while avoiding one or more pre-existing objects in the 3D map of the surrounding environment.

[0045] In one or more embodiments of the method, the visual guide line path is placed in a defined available space within a 3D map of the surrounding environment.

[0046] In one or more embodiments of the method, one or more pre-existing objects include one or more real-world objects and existing MR objects. Attached Figure Description

[0047] Figure 1 A solution architecture for gaze-based control of mixed reality (MR) content, according to an embodiment, is described.

[0048] Figure 2 The process for activating, controlling, and deactivating visual guides for controlling MR objects using gaze and head movement, according to an embodiment, is described.

[0049] Figure 3 A field of view according to an embodiment is depicted, showing the minimum required width for determining visual guide lines in a user's view.

[0050] Figure 4 A graph depicts the variation in gaze control accuracy, adapted from the conference paper “Facing Everyday Gaze Input: Accuracy and Precision of Eye Tracking and Its Implications for Design” (CHI 2017, May 6-11, 2017, p. 1125).

[0051] Figure 5 An example of a non-linear visual guide line, according to an embodiment, is depicted to prevent the line from colliding with an existing object.

[0052] Figure 6 An object in the foreground, according to an embodiment, is depicted with available MR content enhancements and indications of possible directions, where a user can pop up the content enhancement.

[0053] Figure 7 The activated content enhancement according to the embodiment is described.

[0054] Figure 8The illustration depicts a user, according to an embodiment, using their gaze to enhance the drawing of content to the nearest focus.

[0055] Figure 9 A view depicting a user focusing on an original object having points along a line indicating the next action, according to an embodiment.

[0056] Figure 10 Another view depicts, according to an embodiment, of a user focusing on an original object having points along a line indicating the next action.

[0057] Figure 11 Another view, according to an embodiment, depicts a user focusing on the original object with a magnified object.

[0058] Figure 12 A method for activating a bootstrap object according to an embodiment is described.

[0059] Figure 13 A schematic diagram of two steps of a method for selecting and activating objects according to an embodiment is shown.

[0060] Figure 14 Another view depicts a user focusing on an original object with an area highlighted along a visual guide path, according to an embodiment.

[0061] Figure 15 Another view of the user according to an embodiment is depicted.

[0062] Figure 16 Another view, according to an embodiment, depicts a user focusing on an original object having multiple objects shown at a distance.

[0063] Figure 17 Another view depicting a user focusing on an original object with multiple visual guide lines, as shown in an embodiment.

[0064] Figure 18A A schematic diagram depicts a method, according to an embodiment, for a user to change the direction of a visual guide line before the enhanced object begins to move along the visual guide line.

[0065] Figure 18B A method according to an embodiment is depicted, wherein the schematic diagram continues Figure 18B The diagram further illustrates the change in the direction of the visual guide lines.

[0066] Figure 19A This is a system diagram illustrating an example communication system that can implement one or more of the disclosed embodiments.

[0067] Figure 19B The embodiment illustrates that it can be performed... Figure 19AThe system diagram shown is of an example wireless transmit / receive unit (WTRU) used in the communication system. Detailed Implementation

[0068] Embodiments herein provide systems and methods that enable users of virtual and mixed reality devices to decide whether to display content enhancements, and if so, to very quickly and effortlessly control how much they view—that is, if the content enhancement looks interesting, the user can zoom in for a closer look, but can also reject the enhancement if it proves uninteresting upon closer inspection. In one embodiment, this distance is entirely user-controlled.

[0069] In some embodiments disclosed herein, little or no other interaction is required to display content enhancements besides the user's gaze. As will be understood, we are constantly scanning our surroundings for information anyway, so bringing content enhancements closer using additional input methods such as gestures would be cumbersome and potentially draw unwanted attention in crowds. References Figure 1 An overview of the system 100 according to an embodiment includes several components, including a Simultaneous Localization and Mapping (SLAM) / 3D mapping module 102, a gaze detection module 104, an enhanced view user interface module 110, a head pose detection module 120, a hardware information provider module 130, and an enhanced view content service module 140. The enhanced view user interface module 110 includes an enhanced view position determination module 112, a visualization module 114, and a control module 116.

[0070] The SLAM / 3D mapping module 102 maintains a 3D model of the user's surrounding environment and the user's location within the 3D model, including head positioning and orientation. Embodiments herein include the use of any suitable technology to maintain a 3D model of the user's surrounding environment, such as structured infrared patterns, stereo cameras, monocular visual ranging, time-of-flight cameras, etc.

[0071] The gaze detection module 104 identifies the direction of the user's gaze, including convergence and divergence information, to determine the depth at which the user is looking. The gaze detection module 104 also determines how long and how fully the user's gaze has lingered on an object. It detects the dwell time on raw, non-augmented MR, or real-world objects, as well as augmented MR objects.

[0072] The augmented view user interface module 110 includes an augmented view position determination module 112, which establishes potential positions in the 3D space surrounding the user for necessary control of the augmented view. Criteria may include, for example, existing real-world and MR objects, minimum eye movement and gaze detection resolution requirements, and user movement. User preferences, such as the preference to display the augmented view above rather than at or below the horizon, may also be considered. The augmented view user interface module 110 also includes a visualization module 114, which renders all mixed reality (MR) content and, in embodiments, provides functionality including directional visual cues, allowing the user to activate content augmentation features, render visual guide lines, guide user movement by highlighting available and recommended next actions, and render augmented content MR objects along the visual guide lines at positions corresponding to the user's gaze direction.

[0073] The control module 116 within the augmented view user interface (UI) module 110 provides functionality including obtaining a list of available augmentations near the user (e.g., by querying the optional augmented view content service 140). Among the available augmentations, the location determination module 112 requests which augmentations are likely to be displayed; and for possible augmentations, the augmentation initiation posture (gaze dwell and / or other methods known in the art, such as a combination of gaze plus head posture plus gesture) and the direction of indication for the augmented view are detected; the augmented view is activated and the position of the augmented object along the visual guide line is controlled; the suggested next action is determined and highlighted, for example, the next focus-correct snap point on the visual guide line can be highlighted when the user moves the augmented object closer; and using gaze dwell and “visual exhaustion” metrics from the gaze detection module 104, it is estimated whether the user has given sufficient attention to the augmented object so that it can be removed from the view if necessary.

[0074] In one embodiment, gestures for ending the augmented view may include, after a “visual exhaustion” metric has been met, having the user turn their gaze away from the augmented object; having the user turn their head; having the user use their gaze to move the augmented object back to its original position; and / or gestures.

[0075] In one embodiment, the hardware information provider provides hardware-based constraints related to the space required to calculate the display visual guide and the depth of the focused display plane, so that visual cues (“capture points”) can be rendered for the user at those planes along the guide.

[0076] The relevant hardware limitations are at least the number of precise focal planes supported by the device optics and the gaze detection resolution. For example, in one embodiment, if the hardware supports five focal planes, the guide line indicates five corresponding points. If the hardware further supports very precise gaze tracking, the points in the user's view may almost overlap. In embodiments with lower gaze tracking accuracy, these points are further apart in the x / y directions to allow for proper identification of each other. Therefore, in some embodiments, the line occupies more space in the user's view (in the left-right and / or up-down directions).

[0077] Figure 1 Head pose detection via head pose detection module 120 is also illustrated. In an embodiment, head pose detection continuously monitors the user's head movement and reports the identified head pose to UI control module 116. In an embodiment, control module 116 can use this information, for example, in conjunction with concurrent gaze direction / dwelling information, to determine activation, control, and deactivation events for view enhancement features.

[0078] In one embodiment, the enhanced view content service 140 is an external service that provides information on available enhanced MR content for nearby real-world and MR objects.

[0079] The enhanced view content service 140 provides a quick and easy way to obtain a detailed view of distant MR or real-world objects. Users can determine how close (and therefore how large) objects are allowed in their view and have complete control over events without automatic pop-ups. In one embodiment, to accommodate situations where hardware limitations would otherwise inconvenience the user, when the hardware only supports a limited number of focal planes, the service shows the user the locations of these planes and gives the user the option to place content there for optimal viewing.

[0080] The focal plane indicator is beneficial to users because the eye focuses on those planes most quickly and only takes a short time to grasp the relevant details.

[0081] Now for reference Figure 2 The flowchart illustrates a process 200 for gaze-based depth control of mixed reality (MR) objects. Specifically, the flowchart describes the process for activating, controlling, and deactivating MR content enhancements using gaze and head movement.

[0082] As shown in the figure Figure 2The process of interaction between the SLAM module 202, the augmented view control module 204, the augmented view position detection module 206, the gaze detection module 208, the visualization module 210, and the head pose detection module 212 is illustrated. Within the SLAM module 202, user localization, head motion, and 3D mapping 216 are performed. After the 3D mapping is established, head motion data 218 is transmitted to the head pose detection module 212, where constant pose recognition 220 is performed.

[0083] Next, the head positioning and orientation, and the 3D map 226 are provided to the augmented view control 204, which also receives augmentations 228 of all available content in the vicinity. The augmented view position detection module 206 receives hardware constraints 222 and then estimates the minimum required space 224 for the visual guides relative to any hardware constraints.

[0084] Within the enhanced view control module 204, a potential content enhancement location 230 is requested from the enhanced view location detection module 206, along with relevant information such as 3D graphics, head positioning and orientation, and the content to be displayed.

[0085] The enhanced view position detection module 206 estimates each content relative to the space available for display, and provides the enhanced view control module 204 (232) with those contents that may be displayed within the available space from the list of content elements 234.

[0086] Next, the enhanced view control module 204 provides the visualization module 210 (236) with a list of MR enhanced object locations and any pop-up directions, which then renders any visual cues 238.

[0087] Next, the gaze detection module 208 provides gaze and dwell data 240 to the augmented view control module 204, which also receives head pose data 242 from the head pose detection module 212. Within the augmented view control module 204, a determination 244 regarding a start event is made based on the data from the gaze detection module 208 and / or the head pose detection module 212. Furthermore, the augmented view control module 204 determines the capture point from any hardware focus attribute 246. Next, the augmented view control module 204 provides visual guides and augmented content 248, which are then provided to the visualization module 210. The visualization module 210 renders the guides and augmented content 250.

[0088] Next, the enhanced view control module 204 receives gaze and dwell data 252 from the gaze detection module 208, which is used to determine the positioning of the enhanced content based on the gaze direction and to identify and highlight potential next actions 254. Then, any updated content, positioning, and highlighting are provided to the visualization module 210 for rendering 258.

[0089] Gaze and dwell data 260 are repeatedly received from gaze detection module 208, as are head pose data 262 from head pose detection module 212.

[0090] Next, the enhanced view control module 204 determines the end event 264 and provides any enhanced content and animations 266 to the visualization module 210 for rendering 268.

[0091] According to the embodiments described herein, the system continuously performs background content enhancement scans. The system continuously monitors the user's position, head, and gaze direction to determine whether there are objects with MR content enhancements near the user that can be brought into the user's view using the system. This determination can be based on, for example, a geolocation-based search of a (remote) database containing the locations of enhanced objects.

[0092] The embodiments also involve determining the potential for display content enhancement. While acquiring information about enhanced objects near the user, the system continuously maintains information about whether and where enhanced content can be brought into the user's view. The 3D space around the user that could potentially be used for enhanced content placement may initially encompass all space visible to the user, or may be limited to specific viewing areas. For example, in some embodiments, areas directly above and below the user's eye level may be excluded.

[0093] According to some embodiments, the decision to bring augmented content to the user's view is made by considering different parameters that may reduce the available display area of ​​the augmented content. For example, the 3D space surrounding the user is considered. The system performs SLAM to determine the locations of real-world objects near the user. Locations with real-world obstacles are excluded as potential locations for displaying augmented content. Furthermore, identified real-world objects can be marked as objects that must not be occluded.

[0094] Another parameter considered includes MR objects in the user view. According to embodiments, 3D space in the user view already occupied by MR objects can be avoided. Typically, occlusion of existing MR objects is also avoided. However, metrics based on task, activity, or other priorities can be used to determine whether the enhanced content view is likely to occlude existing MR content, such as occlusion over a short period of time.

[0095] Another parameter to consider is eye movement requirements. When a user's eye movement (gaze) is used to control the position of the augmented object, the system needs to determine the minimum range of visual guide lines in the user's view so that gaze detection can distinguish visual control points. The minimum range can be determined by at least the following attributes: gaze tracking accuracy (which may be a hardware limitation), the number of points to be distinguished (which may correspond to the number of focal planes provided by the optics), user movement, distance to the augmented object, etc.

[0096] After considering space constraints, identify one or more potential paths for the visual guide lines. This identification can be based on user preferences (e.g., a user might prefer to use the top of their view for content enhancement), avoiding object occlusion, etc. The lines can be straight lines, splines, arcs, or any other form.

[0097] like Figure 2 As shown, the system also identifies visual cues suitable for content enhancement. Objects determined to be suitable for content enhancement according to this process are highlighted to the user. Highlights can be, for example, visual borders or icons. Specifically, highlighting can include indications of the available directions of visual guide lines.

[0098] The embodiments include various methods for content enhancement for user activation of objects in virtual or mixed reality. These methods include gaze fixation on an object or direction indicator, or by performing head gestures or postures while the gaze remains fixed on the object, as well as other methods known to those skilled in the art that benefit from this disclosure. The direction of head movement can be used to select one of several suggested directions for the guide line. For example, if a leftward direction is suggested, object activation occurs by turning the head to the left. Alternatively, after the guide line becomes visible, head tilting / pitching, such as while the gaze remains fixed on the object, can be used to fine-tune the position of the line.

[0099] In some embodiments, content enhancement activates the display of a selected visual guide line, thereby showing the marker at an optimal viewing distance. The enhanced object enters the view at or near the location of the original object. In some embodiments, the enhanced object is fixed from a corner to the visual guide line, ensuring that the object, line, and marker are always visible.

[0100] In one or more embodiments, a defined visual guide line may be displayed, but no additional markers or snap-to points may be displayed along the visual guide line. In such embodiments, even if the corresponding markers are not visually displayed, the snap-to point can be maintained internally as the user's gaze moves along the visual guide line, and the snap-to-point effect used for moving and displaying enhanced content can still be maintained.

[0101] As those skilled in the art who benefit from this disclosure will understand, visual guide lines may be optional. For example, in one or more embodiments, visual guide lines are not displayed; instead, additional markers or snap points are displayed, making the visual guide lines effectively invisible. Thus, the ability to move enhanced content along the visual guide lines from one point to another changes in response to the user's gaze, as if the visual guide lines were present but not displayed.

[0102] In other embodiments, neither the defined visual guideline nor markers or snap points are displayed. Instead, the enhanced content is moved from one point to another along the defined visual guideline in response to changes in the user's gaze, even if the user cannot see the guideline and associated markers or snap points. In one or more embodiments, a reduced set of markers or snap points may be displayed to provide the user with minimal visual cues for moving the enhanced content. For example, as the enhanced content is moved, markers or snap points adjacent to the current location of the enhanced content may be adaptively displayed, giving the user minimal visible indication of where the content will be moved next, using the offset of the user's gaze.

[0103] In some embodiments, the visual guide line is fixed relative to the pivot point of the user's head / neck. Therefore, if the user moves their head but not tilts / pitches, the guide line moves forward, with its origin fixed to the source (the object being enhanced).

[0104] One or more embodiments include depth control. In some embodiments, the process enables a user to control the depth orientation position of an augmented object using gaze in different ways. In one embodiment, the system identifies that the user's gaze is within a predefined boundary of a visual guide line and should therefore be used to control the positioning of the augmented object. The user's gaze direction and eye convergence / divergence are used to determine the position of the augmented object along the line. Optionally, in addition to determining the depth of focus related to what the user is looking at, the user's eye adaptation may also be used.

[0105] In one embodiment, the augmented object is moved to a corresponding position on a guide line. The next optimal viewing position is highlighted on that line to encourage the user to move the object to that position. At the optimal viewing positions, magnetic or snap effects can be used to hold the object in those positions, which may require additional eye movement to move past that point.

[0106] In one embodiment, depth control of an object stops when the gaze moves away from a predefined boundary, such as when the user is looking at the object instead of the line. If the user looks back at the line, depth control can continue.

[0107] In one embodiment, the system tracks how long and / or how intensively a user looks at the augmented content to determine whether the content can be discarded once the user turns their gaze away. Otherwise, normal, rapid head movements, such as glancing at a car honking, might unintentionally obscure the augmented content.

[0108] In one embodiment, ending the display of enhanced content can be performed via gestures, such as fixing the gaze on a point on the enhanced object and turning the head toward the far end of a visual guide line. In another embodiment, ending the display occurs by the user turning their head and / or gaze away from the enhanced object after the content has timed out.

[0109] Other methods to end the display of augmented content include the user using visual guides or gestures to move the augmented object back to the starting point using gaze.

[0110] According to one or more embodiments, Figure 2 The visualization module 210 renders enhanced content, including visual guide lines, which requires determining the space needed to display those lines. Figure 2 The enhanced view position detection module 206 shown determines whether necessary controls for MR content enhancement can be drawn in the user's view, which may appear continuously, constantly, or as required by the system. Therefore, in some embodiments, module 206 determines the minimum required size of a line in the user's view and any real or virtual objects that the line needs to avoid.

[0111] According to an embodiment, the minimum size of the visual guide line is determined by determining the number of guide points to be drawn. The number of guide points can be the total number of focus-correct planes supported by the device's hardware, a subset of the planes, or, if the number of planes is low, the list of points can include interpolated points to provide sufficient guide points. In one embodiment, interpolated points may be shown differently from the points corresponding to focus-correct planes.

[0112] According to one embodiment, the minimum interval between points is determined by the hardware required to detect each point based on the device's gaze tracking. Furthermore, the minimum interval can be affected by other factors, such as additional head movement caused by motion, or other environmental factors.

[0113] refer to Figure 3 This illustrates a method for determining the desired width of visual guide lines in a user view. As shown in the figure... Figure 3 An illustration shows the plane 306 with ten focal points in two different user views (302) and (304). As shown, the original position of the augmented object is at its furthest point from the user. Figure 3(302) shows a gaze detection resolution of 4 degrees, which means that the entire control line 310 must occupy approximately 40 degrees of eye movement and intervals along the control line in the user's view. Figure 3 (304) shows a 2-degree resolution 312, where control lines 314 cover half of the user view. Figure 3 In this system, movement is limited to the horizontal axis; by using both horizontal and vertical movement, extreme cases of eye movement by the user can be reduced.

[0114] Now for reference Figure 4 The graph illustrates the variation in gaze control accuracy. This graph is adapted from the conference paper "Facing Everyday Gaze Input: Accuracy and Precision of Eye Tracking and Its Implications for Design" (CHI 2017, May 6-11, 2017, p. 1125). As shown in the figure, the graph demonstrates that gaze control accuracy can vary due to various factors, such as viewing angle. The embodiments described herein are adapted to… Figure 4 The problems shown include, for example, the resolution of gaze control precision can vary in different parts of the field of view along x-coordinate 402 and y-coordinate 404. The original size 406, filter size 408, target point 410, screen size 412, and boundary line 420 are shown.

[0115] According to an embodiment of the method, after establishing the minimum length of the guide line, candidate paths for the line in the user's view are determined. The origin of the path (e.g., the far end) is located at the object to be enhanced, while the near end is fixed relative to the user's head. Path determination takes into account the 3D space occupied by real-world objects and existing MR objects in the user's view, ensuring that the guide line does not conflict with existing objects. Further criteria for selecting the path of the visual guide line may include occlusion, such as the line avoiding occlusion of some MR or real-world object, and user movement, such that the line is drawn in the direction the user is moving to maintain their gaze in that general direction.

[0116] Now for reference Figure 5 One embodiment involves bending the guide line 504 of the user 502 to prevent it from colliding with an existing object 510. For example, bending would be appropriate if a straight line from start to finish would conflict with the location of a real-world or MR object. Figure 5 The user view 500 is also shown, illustrating a 2-degree separation 520 between lines of sight, and how the guide line can occupy more lateral orientation area than is strictly necessary to make the view less crowded or avoid collisions with objects. An alternative is to draw the line entirely in the opposite direction, or to allow for resolution by drawing the line narrower in the user's view, so that the line and existing objects can be displayed side by side.

[0117] In some embodiments, user 502 can utilize gaze control to enhance the position of an object by focusing on a visible reference point. In one embodiment, a visible reference point exists regardless of whether the system has a finite or infinite number of focal planes. In one embodiment, markings along a visual guide line help focus on the next position, rather than requiring continuous up-and-down sliding along the line to focus.

[0118] In one embodiment, because the designed line covers a large distance in depth, all parts except the currently focused portion are more or less out of focus. To prevent the next position to be focused on from being difficult to find or requiring a longer focus than needed, the next position can be highlighted sequentially to provide visual stimulation in the user's peripheral vision.

[0119] Now for reference Figure 6 The method is illustrated by a user view 600, which includes visual cues for MR content enhancements available for object 602. The system has determined the potential direction of the content enhancement view pop-up indicated by the arrow. The user can activate the content enhancement in either direction by looking at one of the arrows, or by using some other gestures known in the art, such as a combination of gaze and head / hand gestures.

[0120] Figure 7 The next step in this method is illustrated; specifically, the user activates the enhanced feature on the right. The enhanced content "Call Card" 702 is shown, along with available focus points 704 and 706. To guide the user's action, a focus indicator is used to highlight the next suggested action. The enhanced content is shown at the furthest available focus point. The user can now move the enhanced content along the depth plane by looking at different points 704 and 706 along the visual guide. The next suggested action, which can be at different depths and thus out of focus, is highlighted to provide the user with peripheral visual cues of their location. Figure 7 The suggested action is to look at the next closer focal plane.

[0121] Figure 8 The illustration shows a user view after the user has drawn the enhanced object 800 to the nearest focus using their gaze. In this embodiment, the user can deactivate the enhanced content display 800 by turning their head, looking away after a short time, or by using their gaze to move the object back to its starting point.

[0122] Now for reference Figure 9 In one embodiment, an enhanced object 900 is shown. In some embodiments, when a user focuses on a particular depth, other areas of the scene may become out of focus, as indicated by the darkening of the user's view. Therefore, the user benefits from seeing a change in peripheral vision, highlighting the next focal point 901.

[0123] Now for reference Figure 10 Further examples of user views include displaying the next action of the next action when the user focuses on the enhanced object 1002. Figure 10 This shows that only the enhanced object 1002 can be focused to help the user identify where to look next, such as focus 1003. Therefore, Figure 10 This indicates that when a user focuses on the augmented object 1002, he / she naturally cannot see the surrounding environment very clearly because it may be out of focus. However, by showing the suggested next action 1003 (the next location to look at) clearly enough, the user should be able to recognize the next action, even if it is out of focus. Figure 11 Further eye movements from enhanced object 1100 to the next action 1101 are shown.

[0124] Now for reference Figure 12 The activation and control steps are illustrated in box 1210. As shown in box 1210, the real-world object 1204 or an existing MR object presents a visual cue 1206 to the user 1202, indicating that content enhancement (pop-up) is available for that object. Then, in box 1220, the user activates the enhancement by, for example, a combination of gazing at the cue and head posture, or a gaze hold. The enhanced object 1208 enters the view along with a visual guide 1210 that indicates a focal plane indicator 1212 at its corresponding depth. Then, in box 1230, as the user moves their gaze along the visual guide 1232, the enhanced object moves accordingly, effectively bringing the object closer for inspection. In boxes 1220 and 1230, Figure 12 Visual cues for the next suggested action are shown, including indicators of the next focal plane at different depths.

[0125] As shown in the figure, in box 1210, visual cues indicating the availability of augmented reality are presented to the user. The gaze determines which object to augment and the direction available for augmentation; this can be automatically determined and indicated by the visual cues. Next, in box 1220, activation occurs via gaze persistence or other user input, and guide lines with focal plane indicators and the augmented reality object appear. Next, in box 1230, the suggested next action is highlighted using gaze control, such as positioning the gaze, and the next action can be emphasized.

[0126] Now for reference Figure 13 An alternative activation method illustrates that in 1310, a user can lock their gaze 1302 onto an object for enhancement, and then, as shown in 1320, the user turns their head 1322 in the direction of a directional cue 1304, which triggers the generation and display of a guide line 1324. Thus, activation occurs through gaze locking and head rotation. Furthermore, if multiple directional cues are available (e.g., see...),... Figure 6If there are multiple directional prompts (602) in the prompt, the direction of the user's head rotation can be used to select one of the multiple directional prompts (e.g., the prompt that most closely corresponds to the direction of the user's head rotation).

[0127] Now for reference Figure 14 An example of a curved surface is shown. Instead of lines, the guide assistant can be displayed as a two-dimensional (2D) surface, such as the inner surface of a cone 1402 (plane, hyperbola, cone, ...), a series of concentric circles / triangles / rectangles, etc.

[0128] According to an embodiment, a guiding assistant is provided that allows the user greater flexibility in controlling the positioning of the augmented object using only their eyes. In one embodiment, focusing at an optimal depth can assist the user. Therefore, the user has control over the vertical and horizontal positioning of the object. The area available for placing the augmented object can be determined by the user through head movements such as tilting, pitching, and yaw. As shown, the area seen by the user appears to be viewed from inside cone 1402. A line is drawn along the optimal distance for focusing. A line is also shown as a guide for the next suggested action.

[0129] Therefore, once activated, the user can freely change the position of the visual guide line, either by head movement or by some other posture, or between available positions. For example, if the user activates the visual guide line to their right, tilting the head upwards sufficiently will move the visual guide line to another available position at the top of their field of vision. Through free movement, head pitch and yaw can move the guide line up / down and left / right, respectively.

[0130] In one embodiment, the system and method provide a way to select elements to enhance, for example, whether the number of potential objects is more than one or whether an object exists within a certain distance.

[0131] refer to Figure 15 The potential enhanced objects are shown in window 1500. The user can select which object to activate by looking at it or by other user input.

[0132] Figure 16 This shows the activation of object 1600 and bringing the object to the furthest focal plane.

[0133] Figure 17 Different visual guide lines 1704 and 1706 are shown for each object 1600. Thus, in one embodiment, the gaze pulls the object 1702 closer for closer inspection.

[0134] like Figure 18A and 18BAs shown, the user can freely change the direction of the guide line before the enhanced object begins to move along it. Therefore, in box 1802, user 1806 gazes at the defined object 1804. In box 1810, activation occurs, and guide line 1812 appears. (Reference) Figure 18B In box 1820, with the gaze fixed, rotating the user's head 1826 can rotate the guide line 1824 up or down or from left to right along a visible or invisible rotation point 1822, as shown in box 1830, where the user 1836 rotates the gaze to the left and changes the guide line 1832.

[0135] Therefore, gaze and head posture can modify the visual guide line before gaze-based depth control is activated. In one embodiment, after activation, head movements control the positioning of the visual guide line as long as the user maintains their gaze on the original object. Then, once the line is in the desired position, the line is fixed in place by moving the gaze from the original object to a point along the line.

[0136] Example network for implementation of the embodiments

[0137] Figure 19A This diagram illustrates an example communication system 1900 in which one or more of the disclosed embodiments may be implemented. The communication system 1900 may be a multi-access system providing content such as voice, data, video, messaging, and broadcasting to multiple wireless users. The communication system 1900 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 1900 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT-UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0138] like Figure 19AAs shown, the communication system 1900 may include wireless transmit / receive units (WTRUs) 1902a, 1902b, 1902c, 1902d, RAN 104 / 1913, CN 106 / 1915, public switched telephone network (PSTN) 1908, Internet 1910, and other networks 1912. However, it should be understood that any number of WTRUs, base stations, networks, and / or network elements are contemplated in the disclosed embodiments. Each of the WTRUs 1902a, 1902b, 1902c, and 1902d can be any type of device configured to operate and / or communicate in a wireless environment. As an example, WTRU 1902a, 1902b, 1902c, and 1902d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chains), consumer electronics devices, devices operating in commercial and / or industrial wireless network environments, etc. Any of WTRU 1902a, 1902b, 1902c, and 1902d may be interchangeably referred to as a UE.

[0139] The communication system 1900 may also include base station 1914a and / or base station 1914b. Each of base stations 1914a and 1914b may be any type of device configured to connect to at least one wireless interface of WTRUs 1902a, 1902b, 1902c, and 1902d to facilitate access to one or more communication networks, such as CN 1906 / 1915, the Internet 1910, and / or other networks 1912. As an example, base stations 1914a and 1914b may be base transceiver stations (BTS), node B, e-node B, home node B, home e-node B, gNB, NR node B, site controller, access point (AP), wireless router, etc. Although base stations 1914a and 1914b are each described as a single element, it should be understood that base stations 1914a and 1914b may include any number of interconnected base stations and / or network elements.

[0140] Base station 1914a may be part of RAN 1904 / 1913 and may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 1914a and / or base station 1914b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of radio services to a specific geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 1914a may be divided into three sectors. Thus, in one embodiment, base station 1914a may include three transceivers, i.e., one transceiver for each sector of the cell. In embodiments, base station 1914a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to send and / or receive signals in a desired spatial direction.

[0141] Base stations 1914a and 1902b can communicate with one or more of WTRUs 1902a, 1902b, 1902c, and 1902d via air interface 1916. This air interface can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 1916 can be established using any suitable radio access technology (RAT).

[0142] More specifically, as described above, the communication system 1900 can be a multi-access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 1914a and WTRUs 1902a, 1902b, and 1902c in RAN 1904 / 1913 can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish the air interface 1915 / 1916 / 1917. WCDMA can include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0143] In the embodiments, base stations 1914a and WTRUs 1902a, 1902b, 1902c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or LTE-A Advanced (LTE-A) and / or LTE-A Pro Advanced (LTE-A Pro) to establish an air interface 1916.

[0144] In the embodiments, base station 1914a and WTRUs 1902a, 1902b, 1902c can implement radio technologies such as NR radio access, which can use a new radio (NR) to establish an air interface 1916.

[0145] In the embodiments, base station 1914a and WTRUs 1902a, 1902b, and 1902c can implement multiple radio access technologies. For example, base station 1914a and WTRUs 1902a, 1902b, and 1902c can, for example, use a dual connectivity (DC) principle to implement both LTE and NR radio access together. Therefore, the air interface utilized by WTRUs 1902a, 1902b, and 1902c is characterized by multiple types of radio access technologies and / or transmissions to / from multiple types of base stations (e.g., eNBs and gNBs).

[0146] In other embodiments, base station 1914a and WTRUs 1902a, 1902b, and 1902c can implement radio technologies such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (i.e., Global Microwave Access Interoperability (WiMAX)), CDMA 2000, CDMA 2000 1X, CDMA 2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate Evolution of GSM (EDGE), and GSM EDGE (GERAN).

[0147] Figure 19ABase station 1914b can be, for example, a wireless router, home node B, home e node B, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area such as a business premises, home, vehicle, campus, industrial facility, air corridor (e.g., for drone use), road, etc. In one embodiment, base station 1914b and WTRUs 1902c, 1902d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, base station 1914b and WTRUs 1902c, 1902d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 1914b and WTRUs 1902c, 1902d can utilize cellular-based RATs (e.g., WCDMA, CDMA 2000, GSM, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. Figure 19A As shown, base station 1914b can have a direct connection to the Internet 1910. Therefore, base station 1914b can access the Internet 1910 without going through CN 1906 / 1915.

[0148] RAN 1904 / 1913 can communicate with CN 1906 / 1915, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU 1902a, 1902b, 1902c, and 1902d. Data may have varying Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 1906 / 1915 can provide call control, billing services, location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although in Figure 19A Although not shown, it should be understood that RAN 1904 / 1913 and / or CN1906 / 1915 can communicate directly or indirectly with other RANs that use the same RAT as or a different RAT than RAN 1904 / 1913. For example, in addition to being connected to RAN 1904 / 1913 which can utilize NR radio technology, CN 1906 / 1915 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0149] CN 1906 / 1915 can also serve as a gateway for WTRU 1902a, 1902b, 1902c, 1902d to access PSTN 1908, Internet 1910, and / or other networks 1912. PSTN 1908 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). Internet 1910 may include a global system of interconnected computer networks and devices using common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 1912 may include wired and / or wireless communication networks operated and / or managed by other service providers. For example, network 1912 may include another CN connected to one or more RANs, where the RAN may use the same RAT as RAN 1904 / 1913 or a different RAT.

[0150] Some or all of the WTRUs 1902a, 1902b, 1902c, and 1902d in the communication system 1900 may include multi-mode capability (e.g., WTRUs 1902a, 1902b, 1902c, and 1902d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 19A The WTRU 1902c shown can be configured to communicate with base station 1914a, which can use cellular-based radio technology, and with base station 1914b, which can use IEEE 802 radio technology.

[0151] Figure 19B This shows a system diagram of an example WTRU 1902. (See diagram below.) Figure 19B As shown, WTRU 1902 may include a processor 1918, a transceiver 1920, a transmitting / receiving element 1922, a speaker / microphone 1924, a keyboard 1926, a display / touchpad 1928, non-removable memory 1930, removable memory 1932, a power supply 1934, a Global Positioning System (GPS) chipset 1936 and / or other peripheral devices 1938, and other devices. It is understood that WTRU 1902 may include any sub-combination of the foregoing elements while remaining consistent with the implementation.

[0152] Processor 1918 can be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 1918 can perform signal decoding, data processing, power control, input / output processing, and / or any other function that enables WTRU 1902 to operate in a wireless environment. Processor 1918 can be coupled to transceiver 1920, which can be coupled to transmitting / receiving element 1922. Although Figure 19B While the processor 1918 and transceiver 1920 are described as separate components, it should be understood that the processor 1918 and transceiver 1920 may be integrated together in an electronic package or chip.

[0153] Transmitting / receiving element 1922 can be configured to transmit signals to or receive signals from a base station (e.g., base station 1914a) via air interface 1916. For example, in one embodiment, transmitting / receiving element 1922 can be an antenna configured to transmit and / or receive RF signals. In another embodiment, transmitting / receiving element 1922 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 1922 can be configured to transmit and / or receive both RF and optical signals. It will be appreciated that transmitting / receiving element 1922 can be configured to transmit and / or receive any combination of wireless signals.

[0154] Although the transmitter / receiver unit 1922 was in Figure 19B While described as a single element, the WTRU 1902 may include any number of transmit / receive units 1922. More specifically, the WTRU 1902 may use MIMO technology. Thus, in one embodiment, the WTRU 1902 may include two or more transmit / receive elements 1922 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 1916.

[0155] Transceiver 1920 can be configured to modulate signals to be transmitted by transmitting / receiving element 1922 and demodulate signals received by transmitting / receiving element 1922. As described above, WTRU 1902 can have multimode capability. Therefore, transceiver 1920 can include multiple transceivers to enable WTRU 1902 to communicate via multiple RATs, such as via NR and IEEE 802.11.

[0156] The processor 1918 of WTRU 1902 can be coupled to a speaker / microphone 1924, a keyboard 1926, and / or a display / touchpad 1928 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit), and can receive user input data from there. The processor 1918 can also output user data to the speaker / microphone 1924, keyboard 1926, and / or display / touchpad 1928. Additionally, the processor 1918 can access and store information from any type of suitable memory, such as non-removable memory 1930 and / or removable memory 1932. Non-removable memory 1930 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 1932 may include a subscriber identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, the processor 1918 may access information from memory and store data in memory that is not physically located on the WTRU 1902, for example on a server or home computer (not shown).

[0157] The processor 1918 can receive power from the power supply 1934 and can be configured to distribute and / or control power to other components in the WTRU 1902. The power supply 1934 can be any suitable device for powering the WTRU 1902. For example, the power supply 1934 may include one or more dry cell batteries (e.g., nickel-cadmium, nickel-zinc, nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0158] The processor 1918 may also be coupled to a GPS chipset 1936, which can be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 1902. In addition to, or alternatively to, information from the GPS chipset 1936, the WTRU 102 may receive location information from base stations (e.g., base stations 1914a, 1914b) via the air interface 1916, and / or determine its location based on the timing of signals received from two or more neighboring base stations. It should be understood that the WTRU 1902 may acquire location information using any suitable location determination method, while remaining consistent with the implementation method.

[0159] The processor 1918 may also be coupled to other peripheral devices 1938, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 1938 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or video), universal serial bus (USB) ports, vibration devices, television transceivers, hands-free headsets, modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 1938 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, attitude sensors, biometric sensors, and / or humidity sensors.

[0160] WTRU 1902 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., signals associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., a choke) or via signal processing by a processor (e.g., a separate processor (not shown) or via processor 1918). In embodiments, WTRU 1902 may include a half-duplex radio for which the transmission and reception of some or all signals (e.g., signals associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)) are concurrent and / or simultaneous.

[0161] As described above, the systems and methods described herein provide embodiments for using gaze control to bring augmented MR content related to distant objects closer to the user. The solution uses a visual guideline implemented as an MR object. This guideline contains points that help the user focus their gaze, placed at a depth equal to the focus-corrected viewing distance supported by the device hardware. The augmented MR content follows the user's gaze along this line, moving the content closer to or further away from the user. The position and size of this line are determined by the system based on HW constraints and existing real-world or MR objects in the user's view.

[0162] According to an embodiment, a method includes: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); determining the depth direction location of the user's gaze point based on eye gaze direction and eye convergence / divergence; determining a visual guide path in the 3D map; guiding the user's movements along the visual guide path at one or more identified focal points; and rendering a mixed reality (MR) object along the visual guide path at a location corresponding to the user's gaze direction.

[0163] In one embodiment, one or more focal points along the visual guide path are determined by tracking the user's gaze.

[0164] In one embodiment, one or more focal points along the visual guide path are determined based on one or more hardware limitations of the HMD.

[0165] In one or more embodiments, one or more focal points along the visual guide path are determined based on the number of visually correct focal planes that the HMD can display.

[0166] In one or more embodiments, one or more focal points along the visual guide path are determined based on the accuracy of gaze tracking available on the HMD.

[0167] In one or more embodiments, one or more focal points along the visual guide path are determined based on the number of visually correct focal planes that the HMD can display.

[0168] In one or more embodiments, one or more focal points along the visual guide path are determined based on the accuracy of gaze tracking available on the HMD.

[0169] In one embodiment, one or more focal points along the visual guide path are determined based on the distance to the MR object.

[0170] In one embodiment, one or more focal points along the visual guide path are determined by the user's movement.

[0171] In one embodiment, determining a visual guide path in a 3D image is based on the depth orientation of the user's gaze point and the available space in the 3D image. In another embodiment, determining a visual guide path in a 3D image includes forming a visual guide path to avoid one or more identified objects in the 3D image.

[0172] In one embodiment, determining the visual guide path in the 3D graph includes changing the visual guide path based on the user's movement, which includes one or more of head tilt, head pitch, head yaw, and posture.

[0173] In one embodiment, determining the visual guide path in a 3D drawing includes changing the visual guide path based on one or more rotation points determined by the user's gaze.

[0174] In one embodiment, the method further includes determining the number of points along the visual guide path based on the available focal plane in the 3D graph.

[0175] Another embodiment relates to a method comprising: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); displaying a mixed reality (MR) object in the 3D map, the MR object including visual cues for content enhancement available to the user for the object; activating the content enhancements based on user input to the HMD regarding the visual cues; displaying a visual guide path in the 3D map; guiding the user's movements along the visual guide path at one or more recognized focal points; and rendering the MR object along the visual guide path at a location corresponding to the user's gaze direction.

[0176] In one or more embodiments, activating content enhancement based on user input to the HMD includes user input of one or more of the following: gazing at the content enhancement, head pose, or gaze persistence.

[0177] In one or more embodiments, displaying a visual guideline path in a 3D graph includes displaying one or more identified focal points as multiple focal plane indicators at multiple depths within the 3D graph. In one or more embodiments, rendering an MR object along the visual guideline path at a location corresponding to the user's gaze direction includes moving the augmented object along the multiple focal plane indicators at multiple depths to magnify the MR object.

[0178] In one or more embodiments, guiding a user's actions along a visual guideline path at one or more identified focal points includes providing visual cues, wherein the visual cues include the next suggested action for the user.

[0179] In one or more embodiments, displaying a visual guideline path in a 3D diagram includes determining the depth direction of the user's gaze point based on eye gaze direction and eye convergence / divergence.

[0180] Another embodiment relates to a system including a processor and a non-transitory computer-readable storage medium storing instructions that, when executed on the processor, are operable to perform the following functions: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); determining the depth orientation location of the user's gaze point based on eye gaze direction and eye convergence / divergence; determining a visual guideline path in the 3D map; guiding the user's movements along the visual guideline path at one or more identified focal points; and rendering a mixed reality (MR) object along the visual guideline path at a location corresponding to the user's gaze direction.

[0181] In one or more embodiments of the system, one or more focal points along the visual guide path are determined by tracking the user's gaze.

[0182] In one or more embodiments of the system, one or more focal points along the visual guide path are determined according to one or more hardware limitations of the HMD.

[0183] In one or more embodiments of the system, one or more focal points along the visual guide path are determined based on the distance to the MR object.

[0184] In one or more embodiments of the system, one or more focal points along the visual guide path are determined by the user's movement.

[0185] In one or more embodiments of the system, the path of the visual guide line in the 3D graph is determined based on the depth direction position of the user's gaze point and the available space in the 3D graph.

[0186] In one or more embodiments of the system, determining a visual guide path in a 3D image includes forming a visual guide path to avoid one or more identified objects in the 3D image. In one or more embodiments of the system, determining a visual guide path in a 3D image includes changing the visual guide path based on user movement, including one or more of head tilt, head pitch, head yaw, and gestures.

[0187] In one or more embodiments of the system, determining the visual guide path in a 3D drawing includes changing the visual guide path based on a rotation point determined by the user's gaze.

[0188] Another embodiment of the system targets a non-transitory computer-readable storage medium for storing instructions that, when executed on a processor, are operable to perform additional functions, including determining the number of points along a visual guide path based on an available focal plane in a 3D graph.

[0189] Another embodiment relates to a system including a processor and a non-transitory computer-readable storage medium storing instructions operable, when executed on the processor, to perform the following functions: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); displaying a mixed reality (MR) object in the 3D map, the MR object including visual cues for content enhancement available to the user for the object; activating the content enhancements based on user input to the HMD regarding the visual cues; displaying a visual guide path in the 3D map; guiding the user's movements along the visual guide path at one or more recognized focal points; and rendering the MR object along the visual guide path at a location corresponding to the user's gaze direction.

[0190] In one or more embodiments of the system, activating content enhancement based on user input to the HMD includes user input of one or more of the following: gazing at the content enhancement, head pose, or gaze persistence.

[0191] In one or more embodiments of the system, displaying a visual guide path in a 3D graph includes displaying one or more identified focal points as multiple focal plane indicators at multiple depths within the 3D graph.

[0192] In one or more embodiments of the system, rendering the MR object along the visual guide path at a position corresponding to the user's gaze direction includes moving the augmented object along the plurality of focal plane indicators at the plurality of depths to magnify the MR object.

[0193] In one or more embodiments of the system, guiding a user's actions along a visual guide path at one or more recognized focal points includes providing visual cues, wherein the visual cues include the next suggested action for the user.

[0194] In one or more embodiments of the system, displaying the visual guide line path in a 3D graph includes determining the depth direction location of the user's gaze point based on eye gaze direction and eye convergence / divergence.

[0195] Another embodiment relates to a method for rendering a visual guide path, comprising: forming a three-dimensional (3D) map of the surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); determining the depth direction location of the user's gaze point based on eye gaze direction and eye convergence / divergence; determining a visual guide path in the 3D map; and rendering one or more mixed reality (MR) objects along the visual guide path at a location corresponding to the user's gaze direction, while avoiding one or more pre-existing objects in the 3D map of the surrounding environment.

[0196] In one or more embodiments of the method, the visual guide line path is placed in a defined available space within a 3D map of the surrounding environment.

[0197] In one or more embodiments of the method, one or more pre-existing objects include one or more real-world objects and existing MR objects.

Claims

1. A method for gaze-based control of mixed reality (MR) content, comprising: forming a three-dimensional (3D) map of a user's surroundings of an augmented reality (AR) head-mounted display (HMD); selecting an MR object to place or reposition in the 3D map; displaying a visual guide line path in the 3D map; detecting a gaze of the user at one or more positions along the visual guide line path; and rendering the MR object at a rendering position along the visual guide line path corresponding to the one or more detected positions of the gaze of the user.

2. The method of claim 1, wherein the visual guide line path is a straight line that occupies multiple depths relative to the user.

3. The method of any of claims 1-2, wherein the visual guide line path is a curved line displayed in 3D space.

4. The method of any of claims 1-2, wherein the visual guide line path is displayed with a plurality of point markers that identify the rendering positions at which the MR object can be rendered.

5. The method of any of claims 1-2, wherein a plurality of point markers correspond to discrete focal planes supported by an optical system of the HMD.

6. The method of any of claims 1-2, further comprising determining the visual guide line path to avoid objects in the 3D map.

7. The method of any of claims 1-2, further comprising determining the visual guide line path to avoid additional MR objects displayed by the HMD for the user.

8. The method of any of claims 1-2, wherein selecting the MR object comprises: identifying an object in the user's view for which a content augmentation is available; and selecting the available content augmentation object for the identified object as the selected MR object.

9. The method of any of claims 1-2, further comprising: detecting a gaze point of the user, wherein the MR object is selected based on the detected gaze point of the user.

10. The method of any of claims 1-2, wherein rendering the MR object at the rendering position along the visual guide line path changes a depth at which the MR object is rendered relative to the user.

11. The method of any of claims 1-2, further comprising: displaying a visual cue to the user that a content augmentation for an object corresponding to the MR object is available: wherein selecting the MR object comprises activating the content augmentation for the object.

12. The method of claim 11, further comprising: displaying the object corresponding to the MR object in the 3D map with the visual cue prior to selection.

13. The method of claim 11, wherein the object corresponding to the MR object comprises one of a virtual object, a pre-existing real world object, or an MR object.

14. The method of claim 11, wherein selecting the MR object further comprises: activating the content augmentation for the object corresponding to the MR object in accordance with user input to the AR HMD relative to the visual cue.

15. The method of any one of claims 1-2, wherein the user input to the AR HMD comprises one or more of a gaze, a head pose, or a gaze dwell on the visual cue.

16. The method of any one of claims 1-2, wherein displaying the visual guide line path in the 3D map comprises: displaying a plurality of focal plane indicators at a plurality of depths within the 3D map.

17. The method of claim 16, wherein rendering the MR object along the visual guide line path corresponding to the one or more detected positions of the gaze of the user comprises: moving the MR object along the plurality of focal plane indicators at the plurality of depths to zoom in on the MR object.

18. The method of any one of claims 1-2, wherein displaying the visual guide line path in the 3D map comprises determining a depth directional position of a gaze point of the user based on an eye gaze direction and an eye vergence.

19. The method of claim 18, further comprising: determining the visual guide line path in the 3D map; and determining the one or more positions as one or more identified focal points along the visual guide line path.

20. The method of claim 19, wherein determining the one or more identified focal points along the visual guide line path comprises any one of: determining the one or more identified focal points by gaze tracking of the user, determining the one or more identified focal points according to one or more hardware limitations of the AR HMD, determining one or more identified focal points based on a number of visually correct focal planes that can be displayed by the AR HMD, determining the one or more identified focal points based on a determined gaze tracking accuracy available on the AR HMD, determining the one or more identified focal points according to a distance of the MR object, or determining the one or more identified focal points by movement of the user.

21. The method of claim 19, wherein determining the visual guide line path in the 3D map comprises any one of: determining the visual guide line path based on a depth directional position of a gaze point of the user and available space in the 3D map, determining the visual guide line path by forming the visual guide line path to avoid one or more identified objects in the 3D map, determining the visual guide line path by altering the visual guide line path according to a point of rotation determined by user gaze, or determining the visual guide line path by altering the visual guide line path according to movement of the user, the movement of the user comprising one or more of a head tilt, a head pitch, a head yaw, and a pose.

22. The method of any one of claims 1-2, wherein rendering the MR object further comprises: ​ rendering one or more MR objects along the visual guide line path at locations corresponding to a direction of the gaze of the user, while avoiding one or more pre-existing objects in the 3D map of the surrounding environment, wherein the visual guide line path is placed in determined available space in the 3D map of the surrounding environment, and wherein the one or more pre-existing objects include at least one of a real world object or an existing MR object.

23. A system for gaze-based control of mixed reality (MR) content, comprising: a processor; and a non-transitory computer-readable storage medium storing instructions operable when executed by the processor to cause the system to: form a three-dimensional (3D) map of a surrounding environment of a user of an augmented reality (AR) head-mounted display (HMD); select an MR object to place or reposition in the 3D map; display a visual guide line path in the 3D map; detect a gaze of the user at one or more detected locations of the gaze of the user along the visual guide line path; and render the MR object at a rendering location along the visual guide line path corresponding to the one or more detected locations of the gaze of the user.

24. The system of claim 23, wherein the instructions are further operable when executed by the processor to cause the system to: display a visual cue to the user that a content augmentation for an object corresponding to the MR object is available, wherein selecting the MR object includes activating the content augmentation of the object.

25. The system of claim 24, wherein the instructions are further operable when executed by the processor to cause the system to: prior to selection, display the object corresponding to the MR object in the 3D map with the visual cue.

26. The system of claim 24, wherein the object corresponding to the MR object includes one of a virtual object, a pre-existing real world object, or an MR object. activate the content augmentation of the object corresponding to the MR object according to user input to the AR HMD relative to the visual cue.

27. The system of claim 24, wherein selecting the MR object further comprises:

28. The system of any of claims 23-24, wherein the user input to the AR HMD includes one or more of a gaze, a head pose, or a gaze dwell on the visual cue.

29. The system of any of claims 23-24, wherein displaying the visual guide line path in the 3D map includes: displaying a plurality of focal plane indicators at a plurality of depths within the 3D map.

30. The system of claim 29, wherein rendering the MR object along the visual guide line path corresponding to the one or more detected locations of the gaze of the user includes: moving the MR object along the plurality of focal plane indicators at the plurality of depths to zoom in on the MR object.

31. The method of claim 19, further comprising: ​ guiding the user's action along the visual guide line path at the one or more identified focal points comprises providing a visual cue, wherein the visual cue comprises a next suggested action for the user.

32. The system of any one of claims 23-24, wherein displaying the visual guide line path in the 3D map comprises determining a depth directional position of a fixation point of the user based on eye gaze direction and eye vergence.

Citation Information

Patent Citations

  • Visually guiding motion to be performed by a user

    US20130234926A1

  • Context-aware augmented reality object commands

    US20140237366A1