Pose and interaction with 3D virtual objects using multiple DOF controllers
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2026-08-14
AI Technical Summary
【0007】 本明細書に説明される主題の1つまたはそれを上回る実装の詳細が、付随の図面および以下の説明に記載される。他の特徴、側面、および利点は、説明、図面、および請求項から明白となるであろう。本概要または以下の発明を実施するための形態のいずれも、本発明の主題の範囲を定義または限定することを主張するものではない。 本発明は、例えば、以下を提供する。 (項目1) ウェアラブルデバイスのためのオブジェクトと相互作用するためのシステムであって、前記システムは、 3次元(3D)ビューをユーザに提示し、ユーザの動眼視野(FOR)内のオブジェクトとのユーザ相互作用を可能にするように構成される、ウェアラブルデバイスのディスプレイシステムであって、前記FORは、前記ディスプレイシステムを介して前記ユーザによって知覚可能な前記ユーザの周囲の環境の一部を含む、ウェアラブルデバイスのディスプレイシステムと、 前記ユーザの姿勢と関連付けられたデータを入手するように構成されるセンサと、 前記センサおよび前記ディスプレイシステムと通信するハードウェアプロセッサであって、前記ハードウェアプロセッサは、 前記センサによって入手されたデータに基づいて、前記ユーザの姿勢を判定することと、 前記FOR内のオブジェクトのグループ上への円錐投射を開始することであって、前記円錐投射は、少なくとも部分的に前記ユーザの姿勢に基づく方向に開口を伴う仮想円錐を投射することを含む、ことと、 前記ユーザの環境と関連付けられたコンテキスト情報を分析することと、 少なくとも部分的に前記コンテキスト情報に基づいて、前記仮想円錐の開口を更新することと、 前記円錐投射のための前記仮想円錐の視覚的表現をレンダリングすることと を行うようにプログラムされる、ハードウェアプロセッサと を備える、システム。 (項目2) 前記コンテキスト情報は、前記ユーザの視野(FOV)内のオブジェクトのサブグループのタイプ、レイアウト、場所、サイズ、または密度のうちの少なくとも1つを備え、前記FOVは、前記ディスプレイシステムを介して前記ユーザによって所与の時間に知覚可能な前記FORの一部を含む、項目1に記載のシステム。 (項目3) 前記ユーザのFOV内の前記オブジェクトのサブグループの密度は、 前記オブジェクトのサブグループ内のオブジェクトの数を計算すること、 前記オブジェクトのサブグループによって被覆される前記FOVのパーセンテージを計算すること、または 前記オブジェクトのサブグループ内のオブジェクトに関する等高線マップを計算すること のうちの少なくとも1つによって計算される、項目2に記載のシステム。 (項目4) 前記ハードウェアプロセッサはさらに、前記仮想円錐と前記FOR内のオブジェクトのグループの中の1つまたはそれを上回るオブジェクトとの間の衝突を検出するようにプログラムされ、前記衝突の検出に応答して、前記ハードウェアプロセッサはさらに、焦点インジケータを前記1つまたはそれを上回るオブジェクトに提示するようにプログラムされる、項目1-3のいずれか1項に記載のシステム。 (項目5) 前記ハードウェアプロセッサは、遮蔽曖昧性解消技法を前記仮想円錐と衝突する1つまたはそれを上回るオブジェクトに適用し、遮蔽されたオブジェクトを識別するようにプログラムされる、項目4に記載のシステム。 (項目6) 前記円錐は、中心光線を備え、前記開口は、前記中心光線を横断する、項目1-5のいずれか1項に記載のシステム。 (項目7) 前記仮想円錐は、近位端を備え、前記近位端は、前記ユーザの眼間の場所、ユーザの腕の一部上の場所、ユーザ入力デバイス上の場所、または前記ユーザの環境内の任意の他の場所のうちの少なくとも1つの場所にアンカリングされる、項目1-6のいずれか1項に記載のシステム。 (項目8) 前記ハードウェアプロセッサはさらに、ユーザ入力デバイスから、前記仮想円錐の深度を深度平面にアンカリングするインジケーションを受信するようにプログラムされ、円錐投射は、前記深度平面内の前記オブジェクトのグループ上に実施される、項目1-7のいずれか1項に記載のシステム。 (項目9) ウェアラブルデバイスのためのオブジェクトと相互作用するための方法であって、前記方法は、3次元(3D)空間内の第1の位置においてユーザに表示される標的仮想オブジェクトの選択を受信することと、 前記標的仮想オブジェクトに関する移動のインジケーションを受信することと、 前記標的仮想オブジェクトと関連付けられたコンテキスト情報を分析することと、 少なくとも部分的に前記コンテキスト情報に基づいて、前記標的仮想オブジェクトの移動に適用されるための乗数を計算することと、 前記標的仮想オブジェクトに関する移動量を計算することであって、前記移動量は、少なくとも部分的に前記移動のインジケーションおよび前記乗数に基づく、ことと、 前記ユーザに、第2の位置において前記標的仮想オブジェクトを表示することであって、前記第2の位置は、少なくとも部分的に前記第1の位置および前記移動量に基づく、ことと を含む、方法。 (項目10) 前記コンテキスト情報は、前記ユーザから前記標的仮想オブジェクトまでの距離を含む、項目9に記載の方法。 (項目11) 前記乗数は、前記距離の増加に伴って比例して増加する、項目10に記載の方法。 (項目12) 前記移動は、位置変化、速度、または加速のうちの1つまたはそれを上回るものを含む、項目9-11のいずれか1項に記載の方法。 (項目13) 前記移動のインジケーションは、前記ウェアラブルデバイスと関連付けられたユーザ入力デバイスの作動または前記ユーザの姿勢の変化のうちの少なくとも1つを含む、項目9-12のいずれか1項に記載の方法。 (項目14) 前記姿勢は、頭部姿勢、眼姿勢、または身体姿勢のうちの1つまたはそれを上回るものを含む、項目13に記載の方法。 (項目15) ウェアラブルデバイスのためのオブジェクトと相互作用するためのシステムであって、前記システムは、 3次元(3D)ビューをユーザに提示するように構成される、ウェアラブルデバイスのディスプレイシステムであって、前記3Dビューは、標的仮想オブジェクトを備える、ウェアラブルデバイスのディスプレイシステムと、 前記ディスプレイシステムと通信するハードウェアプロセッサであって、前記ハードウェアプロセッサは、 前記標的仮想オブジェクトに関する移動のインジケーションを受信することと、 前記標的仮想オブジェクトと関連付けられたコンテキスト情報を分析することと、 少なくとも部分的に前記コンテキスト情報に基づいて、前記標的仮想オブジェクトの移動に適用されるための乗数を計算することと、 前記標的仮想オブジェクトに関する移動量を計算することであって、前記移動量は、少なくとも部分的に前記移動のインジケーションおよび前記乗数に基づく、ことと、 前記ディスプレイシステムによって、第2の位置において前記標的仮想オブジェクトを表示すことであって、前記第2の位置は、少なくとも部分的に、前記第1の位置および前記移動量に基づく、ことと を行うようにプログラムされる、ハードウェアプロセッサと を備える、システム。 (項目16) 前記標的仮想オブジェクトの移動のインジケーションは、前記ウェアラブルデバイスのユーザの姿勢の変化または前記ウェアラブルデバイスと関連付けられたユーザ入力デバイスから受信された入力を含む、項目15に記載のシステム。 (項目17) 前記コンテキスト情報は、前記ユーザから前記標的仮想オブジェクトまでの距離を含む、項目15-16のいずれか1項に記載のシステム。 (項目18) 前記乗数は、前記距離が閾値距離未満であるとき、1に等しく、前記閾値距離は、前記ユーザの手が届く範囲と等しい、項目17に記載のシステム。 (項目19) 前記乗数は、前記距離の増加に伴って比例して増加する、項目17-18のいずれか1項に記載のシステム。 (項目20) 前記移動は、位置変化、速度、または加速のうちの1つまたはそれを上回るものを含む、項目15-19のいずれか1項に記載のシステム。
Smart Images

Figure 0007905502000004 
Figure 0007905502000005 
Figure 0007905502000006
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority under 35 U.S.C.§119(e) to U.S. Provisional Application No. 62 / 316,030, filed on March 31, 2016, with the title "CONE CASTING WITH DYNAMICALLY UPDATED APERTURE" and U.S. Provisional Application No. 62 / 325,679, filed on April 21, 2016, with the title "DYNAMIC MAPPING OF USER INPUT DEVICE", both of which are hereby incorporated by reference in their entirety.
[0002] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly, to interactions with virtual objects based on context information.
Background Art
[0003] Modern computing and display technologies are driving the development of systems for so-called “virtual reality,” “augmented reality,” or “mixed reality” experiences, in which digitally reproduced images or parts thereof are presented to the user in a manner that appears, or can be perceived, as real. Virtual reality or “VR” scenarios typically involve the presentation of digital or virtual image information without transparency to other real-world visual inputs. Augmented reality or “AR” scenarios typically involve the presentation of digital or virtual image information as an extension to the visualization of the real world around the user. Mixed reality or “MR” relates to the fusion of the real and virtual worlds to generate a new environment in which physical and virtual objects coexist and interact in real time. In conclusion, the human visual perception system is highly complex, making the generation of VR, AR, or MR technologies that facilitate a rich and comfortable, natural-feeling presentation of virtual image elements among other virtual or real-world image elements challenging. The systems and methods disclosed herein address various challenges related to VR, AR, and MR technologies. [Overview of the project] [Means for solving the problem]
[0004] In one embodiment, a system for interacting with objects for a wearable device is disclosed. The system comprises a display system for a wearable device configured to present a three-dimensional (3D) view to the user and enable user interaction with objects in the user's eye-moving field of view (FOR). FOR may include a portion of the user's surrounding environment that is perceptible to the user through the display system. The system may also comprise a sensor configured to acquire data associated with the user's posture, and a hardware processor communicating with the sensor and the display system. The hardware processor is programmed to determine the user's posture based on the data acquired by the sensor, to initiate a cone projection onto a group of objects in FOR, the cone projection comprising projecting a virtual cone with an opening in a direction at least partially based on the user's posture, to analyze contextual information associated with the user's environment, to update the opening of the virtual cone at least partially based on the contextual information, and to render a visual representation of the virtual cone for the cone projection.
[0005] In another embodiment, a method for interacting with an object for a wearable device is disclosed. The method includes receiving a selection of a target virtual object to be displayed to a user at a first position in three-dimensional (3D) space; receiving an indication of movement relating to the target virtual object; analyzing contextual information associated with the target virtual object; calculating a multiplier to be applied to the movement of the target virtual object, at least partially based on the contextual information; calculating a displacement amount relating to the target virtual object, the displacement amount being at least partially based on the indication of movement and the multiplier; and displaying the target virtual object to the user at a second position, the second position being at least partially based on the first position and the displacement amount.
[0006] In yet another embodiment, a system for interacting with an object for a wearable device is disclosed. The system comprises a display system for a wearable device configured to present a three-dimensional (3D) view to a user, the 3D view comprising a target virtual object. The system may also comprise a hardware processor that communicates with the display system. The hardware processor is programmed to receive indications of movement of the target virtual object, analyze contextual information associated with the target virtual object, calculate multipliers applied to the movement of the target virtual object, at least partially based on the contextual information, calculate a displacement of the target virtual object, the displacement being at least partially based on indications of movement and multipliers, and display the target virtual object at a second position, the second position being at least partially based on the first position and displacement.
[0007] Details of one or more implementations of the subject matter described herein are shown in the accompanying drawings and the following description. Other features, aspects, and advantages will be evident from the description, drawings, and claims. Neither this abstract nor any of the following embodiments for carrying out the invention shall claim to define or limit the scope of the subject matter of the invention. The present invention provides, for example, the following: (Item 1) A system for interacting with objects for a wearable device, wherein the system is A display system for a wearable device, configured to present a three-dimensional (3D) view to a user and enable user interaction with objects within the user's eye-moving field of view (FOR), wherein the FOR includes a portion of the environment surrounding the user that is perceptible to the user through the display system. A sensor configured to obtain data associated with the user's posture, A hardware processor that communicates with the sensor and the display system, wherein the hardware processor is Based on the data obtained by the aforementioned sensor, the user's posture is determined, Initiating a cone projection onto a group of objects within the FOR, the cone projection including projecting a virtual cone with an opening in a direction at least partially based on the user's posture, Analyzing contextual information associated with the user's environment, Updating the opening of the virtual cone based at least partially on the contextual information, Rendering a visual representation of the virtual cone for the aforementioned cone projection A hardware processor and A system that includes these features. (Item 2) The system according to item 1, wherein the context information comprises at least one of the types, layouts, locations, sizes, or densities of subgroups of objects within the user's field of view (FOV), and the FOV includes a portion of the FOR that is perceptible to the user at a given time via the display system. (Item 3) The density of the subgroups of the object within the user's FOV is: Calculating the number of objects within a subgroup of the aforementioned objects, Calculating the percentage of the FOV covered by the subgroup of the object, or Calculating contour maps for objects within a subgroup of the aforementioned objects. The system described in item 2, calculated by at least one of the following: (Item 4) The system according to any one of items 1-3, wherein the hardware processor is further programmed to detect collisions between the virtual cone and one or more objects in the group of objects within the FOR, and in response to the detection of the collision, the hardware processor is further programmed to present a focus indicator to the one or more objects. (Item 5) The system according to item 4, wherein the hardware processor is programmed to apply an occlusion ambiguity technique to one or more objects that collide with the virtual cone and to identify the occluded object. (Item 6) The system according to any one of items 1-5, wherein the cone comprises a central ray, and the aperture intersects the central ray. (Item 7) The system according to any one of items 1-6, wherein the virtual cone comprises a proximal end, the proximal end being anchored to at least one of the following locations: between the user's eyes, on part of the user's arm, on a user input device, or any other location in the user's environment. (Item 8) The hardware processor is further programmed to receive an indication from a user input device to anchor the depth of the virtual cone to a depth plane, and the cone projection is performed on the group of objects in the depth plane, according to any one of items 1-7 of the system. (Item 9) A method for interacting with an object for a wearable device, the method comprising receiving a selection of a target virtual object displayed to the user at a first position in a three-dimensional (3D) space, Receiving an indication of movement related to the target virtual object, Analyzing contextual information associated with the aforementioned target virtual object, Calculating a multiplier to be applied to the movement of the target virtual object, at least partially based on the context information, The calculation of the amount of movement of the target virtual object, wherein the amount of movement is at least partially based on the indication of the movement and the multiplier, The method involves displaying the target virtual object to the user at a second position, wherein the second position is at least partially based on the first position and the amount of movement. Methods that include... (Item 10) The context information includes the distance from the user to the target virtual object, as described in item 9. (Item 11) The method according to item 10, wherein the multiplier increases in proportion to the increase in distance. (Item 12) The motion described in any one of items 9-11 includes, or exceeds, a change in position, velocity, or acceleration. (Item 13) The method according to any one of items 9-12, wherein the indication of movement includes at least one of the activation of a user input device associated with the wearable device or a change in the user's posture. (Item 14) The method according to item 13, wherein the posture includes one or more of the head posture, eye posture, or body posture. (Item 15) A system for interacting with objects for a wearable device, wherein the system is A display system for a wearable device configured to present a three-dimensional (3D) view to a user, wherein the 3D view includes a target virtual object, A hardware processor that communicates with the display system, wherein the hardware processor is Receiving an indication of movement regarding the target virtual object; Analyzing context information associated with the target virtual object; Calculating a multiplier for application to the movement of the target virtual object, at least in part based on the context information; Calculating an amount of movement regarding the target virtual object, the amount of movement being at least in part based on the indication of movement and the multiplier; Causing the display system to display the target virtual object at a second position, the second position being at least in part based on the first position and the amount of movement; A hardware processor programmed to perform the above; A system comprising the above. (Item 16) The system according to item 15, wherein the indication of movement of the target virtual object includes a change in the posture of the user of the wearable device or an input received from a user input device associated with the wearable device. (Item 17) The system according to any one of items 15-16, wherein the context information includes the distance from the user to the target virtual object. (Item 18) The system according to item 17, wherein the multiplier is equal to 1 when the distance is less than a threshold distance, and the threshold distance is equal to the reach of the user's hand. (Item 19) A The system according to any one of items 17-18, wherein the multiplier increases proportionally with an increase in the distance. (Item 20) The system according to any one of items 15-19, wherein the movement includes one or more of a change in position, speed, or acceleration.
Brief Description of Drawings
[0008] [Figure 1] Figure 1 illustrates an example of a mixed reality scenario involving a virtual reality object and a physical object visible to a person.
[0009] [Figure 2] Figure 2 schematically illustrates an example of a wearable system.
[0010] [Figure 3] Figure 3 schematically illustrates aspects of an approach to simulating a 3D image using multiple depth planes.
[0011] [Figure 4] Figure 4 schematically illustrates an example of a waveguide stack for outputting image information to the user.
[0012] [Figure 5] Figure 5 shows an exemplary output beam that can be produced by a waveguide.
[0013] [Figure 6] Figure 6 is a schematic diagram showing an optical system that includes a waveguide apparatus, an optical coupler subsystem for optically coupling light to or from the waveguide apparatus, and a control subsystem used in the generation of a multifocal stereoscopic display, image, or light field.
[0014] [Figure 7] Figure 7 is a block diagram of an embodiment of the wearable system.
[0015] [Figure 8] Figure 8 is a process flow diagram of an example of how to render virtual content in relation to recognized objects.
[0016] [Figure 9] Figure 9 is a block diagram of another embodiment of the wearable system.
[0017] [Figure 10] Figure 10 is a process flow diagram of an embodiment of a method for determining user input to a wearable system.
[0018] [Figure 11] Figure 11 is a process flow diagram of an example of a method for interacting with a virtual user interface.
[0019] [Figure 12A] Figure 12A illustrates an example of conical projection with a non-negligible opening.
[0020] [Figure 12B] Figures 12B and 12C show an example of selecting a virtual object using conical projection with different dynamically adjusted apertures. [Figure 12C] Figures 12B and 12C show an example of selecting a virtual object using conical projection with different dynamically adjusted apertures.
[0021] [Figure 12D] Figures 12D, 12E, 12F, and 12G illustrate an example of dynamically adjusting the opening based on the density of an object. [Figure 12E] Figures 12D, 12E, 12F, and 12G illustrate an example of dynamically adjusting the opening based on the density of an object. [Figure 12F] Figures 12D, 12E, 12F, and 12G illustrate an example of dynamically adjusting the opening based on the density of an object. [Figure 12G] Figures 12D, 12E, 12F, and 12G illustrate an example of dynamically adjusting the opening based on the density of an object.
[0022] [Figure 13] Figures 13, 14, and 15 are flowcharts illustrating an exemplary process for selecting interactable objects using conical projection with dynamically adjustable apertures. [Figure 14] Figures 13, 14, and 15 are flowcharts illustrating an exemplary process for selecting interactable objects using conical projection with dynamically adjustable apertures. [Figure 15] Figures 13, 14, and 15 are flowcharts illustrating an exemplary process for selecting interactable objects using conical projection with dynamically adjustable apertures.
[0023] [Figure 16] Figure 16 schematically illustrates an example of moving a virtual object using a user input device.
[0024] [Figure 17] Figure 17 schematically illustrates an example of a multiplier as a function of distance.
[0025] [Figure 18] Figure 18 illustrates a flowchart of an exemplary process for moving a virtual object in response to movement of a user input device.
[0026] Throughout the drawings, reference numbers may be reused to indicate correspondences between the referenced elements. The drawings are provided to illustrate exemplary embodiments described herein and are not intended to limit the scope of this disclosure. [Modes for carrying out the invention]
[0027] overview A wearable system can be configured to display virtual content within an AR / VR / MR environment. The wearable system can enable a user to interact with physical or virtual objects in the user's environment. The user can interact with objects, for example, by selecting and moving them, by using posture, or by activating a user input device. For example, the user may move a user input device over a distance, and a virtual object will follow the user input device, moving the same amount of distance. Similarly, the wearable system may use cone projection to enable the user to select or target virtual objects in accordance with their posture. As the user moves their hand, the wearable system can accordingly target and select different virtual objects within the user's field of view.
[0028] These approaches can fatigue the user if the objects are relatively far apart. This is because, in order to move a virtual object to a desired location or to reach a desired object, the user needs to move their user input device or increase their physical movement (e.g., increasing the movement of their arms or head) over a similarly long distance. In addition, precise positioning for distant objects can be difficult because it can be hard to verify minute adjustments at a distance. On the other hand, when objects are closer together, the user may prefer more precise positioning in order to interact accurately with the desired object.
[0029] To reduce user fatigue and provide dynamic user interaction with the wearable system, the wearable system can automatically adjust its user interface behavior based on contextual information.
[0030] As an example of providing dynamic user interaction based on contextual information, a wearable system can automatically update the opening of a cone in a cone projection based on a contextual coefficient. For example, if a user turns their head towards a direction with a high density of objects, the wearable system may automatically reduce the cone opening so that there are few virtual selectable objects within the cone. Similarly, if a user turns their head towards a direction with a low density of objects, the wearable system may automatically increase the cone opening to include more objects within the cone, or reduce the amount of movement required to overlap the virtual objects with the cone opening.
[0031] In another embodiment, the wearable system can provide a multiplier that can convert the amount of movement of a user input device (and / or the user's movement) into a larger amount of movement of a virtual object. As a result, the user does not need to physically move long distances to move the virtual object to the desired location when the object is located far away. However, the multiplier may be set to 1 when the virtual object is close to the user (e.g., within the user's reach). Thus, the wearable system can provide a one-to-one operation between user movement and virtual object movement. This can enable the user to interact with nearby virtual objects with increased precision. Embodiments of context-based user interaction are described in detail below.
[0032] Examples of 3D displays for wearable systems A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present a user with 2D or 3D virtual images. The images may be still images, video frames, or videos in combination or equivalent. A wearable system may include wearable devices that can present a VR, AR, or MR environment, either alone or in combination, for user interaction. A wearable device may be a head-mounted device (HMD).
[0033] Figure 1 illustrates an example of a mixed reality scenario involving a virtual reality object and a physical object that are visible to a person. In Figure 1, MR scene 100 is depicted, and the user of the MR technology sees a real-world park-like setting 110 featuring people, trees, buildings in the background, and a concrete platform 120. In addition to these items, the user of the MR technology also perceives "seeing" a robotic figure 130 standing on the real-world platform 120 and a flying cartoonish avatar character 140 that appears to be a personification of a bumblebee, although these elements do not exist in the real world.
[0034] It may be desirable for a 3D display to generate a distance-accommodative response corresponding to the virtual depth of each point within the display's field of view, in order to produce a true sense of depth, more specifically, a simulated sense of surface depth. If the distance-accommodative response for a display point does not correspond to the virtual depth of that point as determined by the binocular depth cues for convergence and stereopsis, the human eye may experience distance-accommodative collision, which can result in unstable imaging, harmful eye strain, headaches, and, in the absence of distance-accommodative information, a near-complete loss of surface depth.
[0035] VR, AR, and MR experiences can be provided by a display system having a display that provides the viewer with images corresponding to multiple depth planes. The images may differ for each depth plane (e.g., providing slightly different presentations of scenes or objects) and can be individually focused by the viewer's eyes, thereby helping to provide the user with depth cues based on the eye's accommodation required to focus on different image features relating to scenes located on different depth planes, or based on observing different image features on different depth planes that are out of focus. As discussed elsewhere herein, such depth cues provide a reliable perception of depth.
[0036] Figure 2 illustrates an embodiment of the wearable system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems to support the functions of the display 220. The display 220 may be coupled to a frame 230, which is wearable by a user, wearer, or viewer 210. The display 220 can be positioned in front of the user 210's eyes. The display 220 can present AR / VR / MR content to the user. The display 220 may comprise a head-mounted display (HMD) that is worn on the user's head. In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the user's ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable acoustic control).
[0037] The wearable system 200 may include an outward-facing imaging system 464 (shown in Figure 4) that observes the world within the user's surrounding environment. The wearable system 200 may also include an inward-facing imaging system 462 (shown in Figure 4) that can track the user's eye movements. The inward-facing imaging system can track the movement of one eye or both eyes. The inward-facing imaging system 462 may be mounted on the frame 230 and may communicate with a processing module 260 or 270 that processes the image information acquired by the inward-facing imaging system and can determine, for example, the pupil diameter or orientation of the user's eyes, eye movements, or eye posture.
[0038] As an example, the wearable system 200 can acquire images of the user's posture using an outward-facing imaging system 464 or an inward-facing imaging system 462. The images may be still images, video frames or videos, a combination thereof, or equivalent.
[0039] The display 220 is operably coupled to a local data processing module 260 (250), which can be mounted in various configurations, such as being fixedly attached to the frame 230 by wired or wireless connections, fixed to a helmet or hat worn by the user, built into headphones, or otherwise detachably attached to the user 210 (for example, in a backpack configuration or a belt-connected configuration).
[0040] The local processing and data module 260 may include a hardware processor and digital memory such as non-volatile memory (e.g., flash memory), both of which may be used to assist in data processing, caching, and storage. The data may include (a) data captured from sensors (e.g., operably coupled to frame 230 or otherwise attached to user 210) such as image acquisition devices (e.g., cameras in inward-facing imaging systems and / or outward-facing imaging systems), microphones, inertial measuring units (IMUs), accelerometers, compasses, global positioning system (GPS) units, wireless devices, or gyroscopes, or (b) data acquired or processed using the remote processing module 270 or remote data repository 280 for transmission to the display 220 after such processing or reading. The local processing and data modules 260 may be operably coupled to the remote processing module 270 or the remote data repository 280 by communication links 262 or 264, such as via wired or wireless communication links, so that these remote modules are available as resources to the local processing and data modules 260. In addition, the remote processing module 280 and the remote data repository 280 may be operably coupled to each other.
[0041] In some embodiments, the remote processing module 270 may comprise one or more processors configured to analyze and process data and / or image information. In some embodiments, the remote data repository 280 may comprise a digital data storage facility, which may be available through the internet or other networking configurations in a “cloud” resource configuration. In some embodiments, all data is stored, and all calculations are performed in the local processing and data module, enabling fully autonomous use from the remote module.
[0042] The human visual system is complex and struggles to provide a realistic perception of depth. While not limited by theory, it is believed that an object viewer may perceive an object as three-dimensional due to a combination of vergence and accommodation. The vergence and divergence of two eyes relative to each other (i.e., the rolling of pupils toward or away from each other to converge and fix the gaze on an object) is closely related to the focusing (or "accommodation") of the eye's lens. Under normal conditions, changing the focus of the eye's lens, or accommodation, to shift focus from one object to another at a different distance, will automatically produce a consistent change in vergence and divergence at the same distance, under a relationship known as the "accommodation-vergence-divergence reflex." Similarly, a change in vergence and divergence will, under normal conditions, induce a consistent change in accommodation. A display system that provides better coordination between distance accommodation and convergence / divergence motion can create a more realistic and comfortable simulation of three-dimensional images.
[0043] Figure 3 illustrates aspects of an approach to simulating a three-dimensional image using multiple depth planes. Referring to Figure 3, objects at various distances from eyes 302 and 304 on the z-axis are accommodated by eyes 302 and 304 so that those objects are in focus. Eyes 302 and 304 take on specific accommodated states, focusing objects at different distances along the z-axis. As a result, a specific accommodated state can be associated with one of the depth planes 306 having an associated focal length so that an object or part of an object in a particular depth plane is in focus when the eye is accommodated to that depth plane. In some embodiments, the three-dimensional image may be simulated by providing a different presentation of the image for each of eyes 302 and 304, and by providing a different presentation of the image corresponding to each of the depth planes. For the sake of clarity in the illustration, it should be understood that the fields of view of eyes 302 and 304 may overlap, for example, as the distance along the z-axis increases, although they are shown as separate. Furthermore, for the sake of illustration, it should be understood that, although shown as flat, the outline of the depth plane can be curved in physical space so that all features within the depth plane are in focus with the eye in a particular state of perspective accommodation. While not limited by theory, it is thought that the human eye can interpret a finite number of depth planes and typically provide depth perception. Consequently, a highly realistic simulation of perceived depth can be achieved by providing the eye with different presentations of images corresponding to each of these limited number of depth planes.
[0044] Waveguide stack assembly Figure 4 illustrates an embodiment of a waveguide stack for outputting image information to a user. The wearable system 400 includes a waveguide stack or stacked waveguide assembly 480, which may be used to provide three-dimensional perception to the eyes / brain using a plurality of waveguides 432b, 434b, 436b, 438b, and 4400b. In some embodiments, the wearable system 400 may correspond to the wearable system 200 of Figure 2, and Figure 4 schematically shows some parts of the wearable system 200 in more detail. For example, in some embodiments, the waveguide assembly 480 may be integrated into the display 220 of Figure 2.
[0045] Continuing with Figure 4, the waveguide assembly 480 may also include several features 458, 456, 454, and 452 between the waveguides. In some embodiments, features 458, 456, 454, and 452 may be lenses. In other embodiments, features 458, 456, 454, and 452 may not be lenses. Rather, they may simply be spacers (e.g., cladding layers or structures for forming air gaps).
[0046] Waveguides 432b, 434b, 436b, 438b, 440b or multiple lenses 458, 456, 454, 452 may be configured to transmit image information to the eye using varying levels of wavefront curvature or ray divergence. Each waveguide level may be associated with a specific depth plane and configured to output image information corresponding to that depth plane. Image input devices 420, 422, 424, 426, 428 may be used to input image information into waveguides 440b, 438b, 436b, 434b, 432b, each of which may be configured to disperse incident light across each individual waveguide for output toward the eye 410. Light exits from the output surfaces of image input devices 420, 422, 424, 426, and 428 and is fed into the corresponding input edges of waveguides 440b, 438b, 436b, 434b, and 432b. In some embodiments, a single beam of light (e.g., a collimated beam) may be fed into each waveguide and output an entire field of cloned collimated beams, which are directed toward the eye 410 at a specific angle (and divergence) corresponding to a depth plane associated with a particular waveguide.
[0047] In some embodiments, the image input devices 420, 422, 424, 426, and 428 are discrete displays that generate image information for input into their respective corresponding waveguides 440b, 438b, 436b, 434b, and 432b, respectively. In some other embodiments, the image input devices 420, 422, 424, 426, and 428 are output terminals of a single multiplexed display that can send image information to each of the image input devices 420, 422, 424, 426, and 428 via, for example, one or more optical conduits (such as optical fiber cables).
[0048] The controller 460 controls the operation of the stacked waveguide assembly 480 and the image input devices 420, 422, 424, 426, and 428. The controller 460 includes programming (e.g., instructions in a non-transient computer-readable medium) to coordinate the timing and delivery of image information to the waveguides 440b, 438b, 436b, 434b, and 432b. In some embodiments, the controller 460 may be a single integrated device or a distributed system connected by wired or wireless communication channels. In some embodiments, the controller 460 may be part of a processing module 260 or 270 (illustrated in Figure 2).
[0049] Waveguides 440b, 438b, 436b, 434b, and 432b may be configured to propagate light within each individual waveguide by total internal reflection (TIR). Waveguides 440b, 438b, 436b, 434b, and 432b may each be planar or have another shape (e.g., curved), with major upper and lower surfaces and edges extending between their major upper and lower surfaces. In the illustrated configuration, waveguides 440b, 438b, 436b, 434b, and 432b may each include light extraction optical elements 440a, 438a, 436a, 434a, and 432a, respectively, configured to extract light from the waveguides by redirecting the light, propagating it within each individual waveguide, and outputting image information from the waveguides to the eye 410. The extracted light may also be referred to as externally coupled light, and the light extraction optical elements may also be referred to as externally coupled optical elements. The beam of extracted light is output by the waveguide to the location where the light propagating within the waveguide strikes the light redirection element. The light extraction optical elements (440a, 438a, 436a, 434a, 432a) may be, for example, reflective or diffracting optical features. For ease of explanation and clarity of the drawings, they are shown positioned on the bottom main surface of waveguides 440b, 438b, 436b, 434b, 432b, but in some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be positioned on the top or bottom main surface, or directly within the volume of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, and 432a may be mounted on a transparent substrate and formed within a layer of material that forms the waveguides 440b, 438b, 436b, 434b, and 432b. In some other embodiments, the waveguides 440b, 438b, 436b, 434b, and 432b may be monolithic components of the material, and the light extraction optical elements 440a, 438a, 436a, 434a, and 432a may be formed on and / or inside the surface of that component of the material.
[0050] Continuing with Figure 4, as discussed herein, each waveguide 440b, 438b, 436b, 434b, and 432b is configured to emit light and form an image corresponding to a particular depth plane. For example, the waveguide 432b closest to the eye may be configured to deliver collimated light to the eye 410 as it is introduced into such waveguide 432b. The collimated light may represent the optical infinity focal plane. The next waveguide 434b may be configured to emit collimated light that passes through a first lens 452 (e.g., a negative lens) before reaching the eye 410. The first lens 452 may generate a slight convex wavefront curvature so that the eye / brain interprets the light originating from the next upper waveguide 434b as originating from a first focal plane closer inward from optical infinity toward the eye 410. Similarly, the third upper waveguide 436b passes its output light through both the first lens 452 and the second lens 454 before reaching the eye 410. The combined refractive power of the first and second lenses 452 and 454 may be configured to generate a different, incremental wavefront curvature so that the eye / brain interprets the light emanating from the third waveguide 436b as emanating from a second focal plane that is even closer inward toward the person from optical infinity than the light from the next upper waveguide 434b.
[0051] Other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are configured similarly, with the highest waveguide 440b in the stack being used to transmit its output through all the lenses between it and the eye for a concentrated focusing force representing the focal plane closest to the person. When viewing / interpreting light originating from the other side world 470 of the stacked waveguide assembly 480, a compensating lens layer 430 may be positioned on top of the stack to compensate for the stack of lenses 458, 456, 454, 452, and to compensate for the concentrated force of the lower lens stacks 458, 456, 454, 452. Such a configuration provides the same number of perceived focal planes as there are available waveguide / lens pairs. Both the light-extracting optical elements of the waveguides and the focusing sides of the lenses may be static (e.g., not dynamic or electroactive). In some alternative embodiments, one or both may be dynamic using electroactive features.
[0052] Continuing with Figure 4, the light extraction optical elements 440a, 438a, 436a, 434a, and 432a may be configured to redirect light from their respective waveguides and output the light with an appropriate amount of divergence or collimation for a particular depth plane associated with the waveguide. As a result, waveguides having different associated depth planes may have different configurations of light extraction optical elements that output light with different amounts of divergence depending on the associated depth plane. In some embodiments, as discussed herein, the light extraction optical elements 440a, 438a, 436a, 434a, and 432a may be three-dimensional or surface features that can be configured to output light at specific angles. For example, the light extraction optical elements 440a, 438a, 436a, 434a, and 432a may be volume holograms, surface holograms, and / or diffraction gratings. Optical elements for light extraction, such as diffraction gratings, are described in U.S. Patent Publication No. 2015 / 0178939, published on June 25, 2015 (which is incorporated herein by reference as a whole).
[0053] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, and 432a are diffraction features that form a diffraction pattern, i.e., “diffractive optical elements” (also referred to herein as “DOEs”). Preferably, the DOEs have relatively low diffraction efficiency such that only a portion of the beam light is deflected toward the eye 410 using each intersection of the DOEs, while the remainder continues to travel through the waveguide via total internal reflection. The light carrying the image information is therefore split into several associated emission beams that exit the waveguide at multiple locations, resulting in a very uniform pattern of emission toward the eye 304 with respect to this particular collimated beam bouncing within the waveguide.
[0054] In some embodiments, one or more DOEs may be switchable between an "on" state in which they actively diffract and an "off" state in which they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer-dispersed liquid crystal, in which microdroplets have a diffraction pattern in the host medium, and the refractive index of the microdroplets can be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract incident light), or the microdroplets can be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts incident light).
[0055] In some embodiments, the number and distribution of depth planes or depth of field may vary dynamically based on the pupil size or orientation of the viewer's eye. The depth of field may change inversely with the viewer's pupil size. As a result, as the pupil size of the viewer's eye decreases, the depth of field increases so that one plane that is indistinguishable because its location is beyond the eye's depth of focus becomes discernible and appears more in focus with the decrease in pupil size and the corresponding increase in depth of field. Similarly, the number of spaced-out depth planes used to present different images to the viewer may decrease with the decreased pupil size. For example, it may not be possible for a viewer to clearly perceive the details of both the first and second depth planes at one pupil size without adjusting the eye's accommodation from one depth plane to the other. However, these two depth planes may simultaneously be sufficient to focus on the user at a different pupil size without changing accommodation.
[0056] In some embodiments, the display system may vary the number of waveguides receiving image information based on a determination of pupil size and / or orientation, or in response to the reception of an electrical signal indicating a particular pupil size and / or orientation. For example, if the user's eye is unable to distinguish between two depth planes associated with two waveguides, the controller 460 may be configured or programmed to stop providing image information to one of these waveguides. Advantageously, this can reduce the processing load on the system and thereby increase the system's responsiveness. In embodiments where the DOE for a waveguide is switchable between on and off states, the DOE may be switched off when the waveguide receives image information.
[0057] In some embodiments, it may be desirable to satisfy the condition that the emitted beam has a diameter less than the diameter of the viewer's eye. However, satisfying this condition may be difficult in light of the variability of the viewer's pupil size. In some embodiments, this condition is satisfied over a wide range of pupil sizes by varying the size of the emitted beam in response to the determination of the viewer's pupil size. For example, as the pupil size decreases, the size of the emitted beam may also decrease. In some embodiments, the size of the emitted beam may be varied using a variable aperture.
[0058] The wearable system 400 may include an outward-facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 may be referred to as the field of view (FOV), and the imaging system 464 may sometimes be referred to as an FOV camera. The entire area available for viewing or imaging by the viewer may be referred to as the eye-moving field of view (FOR). FOR may include a solid angle of 4π steradians surrounding the wearable system 400 so that the wearer moves their body, head, or eyes to perceive substantially any direction in space. In other circumstances, the wearer's movement may be more restrained, and accordingly, the wearer's FOR may correspond to a smaller solid angle. Images obtained from the outward-facing imaging system 464 can be used, for example, to track gestures made by the user (e.g., hand or finger gestures) and to detect objects in the world 470 in front of the user.
[0059] The wearable system 400 may also include an inward-facing imaging system 466 (e.g., a digital camera) that observes user movements such as eye and face movements. The inward-facing imaging system 466 may be used to capture images of the eyes 410 and to determine the size or orientation of the pupils of the eyes 304. The inward-facing imaging system 466 may be used to determine the direction the user is looking (e.g., eye posture) or to obtain images for the user's biometric identification (e.g., via iris recognition). In some embodiments, at least one camera may be used for each eye, independently, to determine the pupil size or eye posture of each eye separately, thereby allowing the presentation of image information to each eye to be dynamically adjusted for that eye. In some other embodiments, the pupil diameter or orientation of only one eye 410 (e.g., using only one camera per pair of eyes) is determined and assumed to be similar with respect to both of the user's eyes. Images obtained by the inward-facing imaging system 466 may be used by the wearable system 400 to determine the user's eye posture or mood, or to determine the audio or visual content to be presented to the user. The wearable system 400 may also use sensors such as an IMU, accelerometer, and gyroscope to determine head posture (e.g., head position or head orientation).
[0060] The wearable system 400 may include a user input device 466 that allows the user to input commands into a controller 460 and interact with the wearable system 400. For example, the user input device 466 may include a trackpad, touchscreen, joystick, multi-degree-of-freedom (DOF) controller, capacitive sensing device, game controller, keyboard, mouse, directional pad (D-pad), wand, tactile device, totem (e.g., functioning as a virtual user input device), etc. A multi-DOF controller may sense user input in translation (e.g., left / right, forward / backward, or up / down) or rotation (e.g., yaw, pitch, or roll), which can be some or all of the controller's possible movements. A multi-DOF controller that supports translation may be referred to as 3DOF, while a multi-DOF controller that supports both translation and rotation may be referred to as 6DOF. In some cases, the user may use a finger (e.g., thumb) to press or swipe over a touch sensor input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input device 466 may be held in the user's hand while using the wearable system 400. The user input device 466 can be connected to the wearable system 400 via wired or wireless communication.
[0061] Figure 5 shows an embodiment of an outgoing beam output by a waveguide. Although one waveguide is shown, other waveguides within the waveguide assembly 480 may function similarly, and it should be understood that the waveguide assembly 480 includes multiple waveguides. Light 520 is injected into waveguide 432b at the input edge 432c of waveguide 432b and propagates through waveguide 432b by TIR. At the point where light 520 collides with DOE 432a, a portion of the light exits the waveguide as an outgoing beam 510. The outgoing beams 510 are shown as substantially parallel, but they may also be redirected to propagate towards the eye 410 at a certain angle depending on the depth plane associated with waveguide 432b (e.g., forming a divergent outgoing beam). It should be understood that a nearly parallel emitted beam may represent a waveguide with an optical element that externally couples the light and forms an image that appears to be set in the depth plane at long distances from the eye 410 (e.g., optical infinity). Other waveguides or other sets of optical elements may output a more divergent emitted beam pattern, which would require the eye 410 to adjust to a closer distance and focus on the retina, and would be interpreted by the brain as light from a distance closer to the eye 410 than optical infinity.
[0062] Figure 6 is a schematic diagram showing an optical system that includes a waveguide apparatus, an optical coupler subsystem for optically coupling light to or from the waveguide apparatus, and a control subsystem used in the generation of a multifocal stereoscopic display, image, or light field. The optical system may include a waveguide apparatus, an optical coupler subsystem for optically coupling light to or from the waveguide apparatus, and a control subsystem. The optical system may be used to generate a multifocal stereoscopic display, image, or light field. The optical system may include one or more primary plane waveguides 632a (only one is shown in Figure 6) and one or more DOEs 632b associated with at least one of each of the primary waveguides 632a. The plane waveguides 632b may be analogous to waveguides 432b, 434b, 436b, 438b, and 440b discussed with reference to Figure 4. The optical system may employ a dispersed waveguide apparatus to relay light along a first axis (vertical or Y-axis in the diagram of Figure 6) and expand the effective exit pupil of light along the first axis (e.g., Y-axis). The dispersed waveguide apparatus may include, for example, a dispersed plane waveguide 622b and at least one DOE 622a (illustrated by a double dashed line) associated with the dispersed plane waveguide 622b. The dispersed plane waveguide 622b may be similar to or the same as a primary plane waveguide 632b having a different orientation in at least some respects. Similarly, at least one DOE 622a may be similar to or the same as a DOE 632a in at least some respects. For example, the dispersed plane waveguide 622b or DOE 622a may be made of the same material as the primary plane waveguide 632b or DOE 632a, respectively. An embodiment of the optical display system 600 shown in Figure 6 can be integrated into the wearable system 200 shown in Figure 2.
[0063] The relayed and dilated light can be optically coupled from the dispersed waveguide apparatus into one or more primary plane waveguides 632b. The primary plane waveguides 632b can relay light along a second axis (e.g., horizontal or X-axis in the diagram of Figure 6) perpendicular to the first axis. It should be noted that the second axis can be a non-orthogonal axis to the first axis. The primary plane waveguides 632b dilate the effective exit pupil of the light along their second axis (e.g., X-axis). For example, a dispersed plane waveguide 622b can relay and dilate light along the vertical or Y-axis, and can pass its light into a primary plane waveguide 632b which can relay and dilate light along the horizontal or X-axis.
[0064] The optical system may include one or more colored light sources (e.g., red, green, and blue laser light) 610 that can be optically coupled into the proximal end of a single-mode optical fiber 640. The distal end of the optical fiber 640 may be screwed or received through a hollow tube 642 made of piezoelectric material. The distal end protrudes from the tube 642 as an unfixed, flexible cantilever 644. The piezoelectric tube 642 can be associated with four quadrant electrodes (not shown). The electrodes may be plated, for example, on the outside, outer surface, outer periphery, or diameter of the tube 642. A core electrode (not shown) may also be located in the core, center, inner periphery, or inner diameter of the tube 642.
[0065] For example, a drive electronic device 650, electrically coupled via wire 660, drives a pair of opposing electrodes to independently bend the piezoelectric tube 642 along two axes. The protruding distal tip of the optical fiber 644 has a mechanical resonance mode. The resonance frequency may depend on the diameter, length, and material properties of the optical fiber 644. By vibrating the piezoelectric tube 642 near the first mechanical resonance mode of the fiber cantilever 644, the fiber cantilever 644 can be vibrated and swept through a large deflection.
[0066] By stimulating resonant vibrations in two axes, the tip of the fiber cantilever 644 is scanned in two axes within an area that fills a two-dimensional (2-D) scan. By modulating the intensity of the light source 610 in synchronization with the scanning of the fiber cantilever 644, the light emitted from the fiber cantilever 644 can form an image. A description of such a setup is provided in U.S. Patent Publication No. 2014 / 0003762, which is incorporated herein by reference in its entirety.
[0067] Components of the optical coupler subsystem can collimate light emitted from the scanning fiber cantilever 644. The collimated light can be reflected by the mirrored surface 648 into a narrow-dispersion planar waveguide 622b containing at least one diffractive optical element (DOE) 622a. The collimated light propagates perpendicularly along the dispersive planar waveguide 622b (with respect to the diagram in Figure 6) by TIR, thereby repeatedly intersecting with the DOE 622a. The DOE 622a preferably has a low diffraction efficiency. This allows a portion of the light (e.g., 10%) to be diffracted toward the edge of the larger primary planar waveguide 632b at each intersection with the DOE 622a, while a portion of the light can be continued along its original trajectory along the length of the dispersive planar waveguide 622b via TIR.
[0068] At each intersection with DOE622a, additional light can be diffracted toward the entrance of the primary waveguide 632b. By splitting the incident light into multiple external coupling sets, the exit pupil of the light can be vertically extended by DOE4 in the dispersed plane waveguide 622b. This vertically extended light, externally coupled from the dispersed plane waveguide 622b, can enter the edge of the primary plane waveguide 632b.
[0069] Light entering the primary waveguide 632b can propagate horizontally along the primary waveguide 632b (relative to the diagram in Figure 6) via TIR. As the light intersects with DOE 632a at multiple points, it propagates horizontally along at least a portion of the length of the primary waveguide 632b via TIR. DOE 632a may be advantageously designed or configured to have a phase profile which is the sum of linear and radially symmetric diffraction patterns, and to produce both deflection and focusing of light. DOE 632a may advantageously have a low diffraction efficiency (e.g., 10%) such that only a portion of the beam of light is deflected towards the viewer's eye at each intersection of DOE 632a, while the rest of the light continues to propagate through the primary waveguide 632b via TIR.
[0070] At each intersection point between the propagating light and the DOE632a, a portion of the light is diffracted toward the adjacent surface of the primary waveguide 632b, allowing the light to escape from the TIR and be emitted from the surface of the primary waveguide 632b. In some embodiments, the radially symmetric diffraction pattern of the DOE632a also imparts a certain focal level to the diffracted light, shaping (e.g., imparting curvature) the wavefronts of the individual beams and steering the beams to an angle that matches the designed focal level.
[0071] Therefore, these different paths can be used to couple light outside the primary plane waveguide 632b by resulting in different filling patterns in the DOE 632a at different angles, focal levels, and / or in the exit pupil. Different filling patterns in the exit pupil can be advantageously used to generate a light field display with multiple depth planes. Each layer in the waveguide assembly or a set of layers in a stack (e.g., three layers) may be employed to generate individual colors (e.g., red, blue, and green). Thus, for example, a first set of three adjacent layers may be employed to generate red, blue, and green light at a first depth of focus, respectively. A second set of three adjacent layers may be employed to generate red, blue, and green light at a second depth of focus, respectively. Multiple sets may be employed to generate a full 3D or 4D color image light field with various depths of focus.
[0072] (Other components of the wearable system) In many implementations, the wearable system may include other components in addition to, or as alternatives to, the components of the wearable system described above. The wearable system may include, for example, one or more tactile devices or components. The tactile devices or components may be operable to provide a sense of touch to the user. For example, the tactile devices or components may provide a sense of pressure and / or texture when touching virtual content (e.g., virtual objects, virtual tools, other virtual structures). The tactile sensation may replicate the sensation of a physical object represented by a virtual object, or the sensation of an imaginary object or character represented by virtual content (e.g., a dragon). In some implementations, the tactile devices or components may be worn by the user (e.g., user-wearable gloves). In some implementations, the tactile devices or components may be held by the user.
[0073] A wearable system may include, for example, one or more physical objects that are operable by the user and enable input to or interaction with the wearable system. These physical objects may be referred to herein as totems. Some totems may take the form of inanimate objects, such as a piece of metal or plastic, a wall, or the surface of a table. In some implementations, a totem may not actually have any physical input structures (e.g., keys, triggers, joysticks, trackballs, rocker switches). Instead, a totem may simply provide a physical surface, and the wearable system may render a user interface so that it appears to the user as being on one or more of the totems. For example, a wearable system may render images of a computer keyboard and trackpad so that they appear to reside on one or more of the totems. For example, a wearable system may render a virtual computer keyboard and virtual trackpad so that they appear to be on the surface of a thin rectangular aluminum plate that acts as a totem. The rectangular plate itself does not have any physical keys or trackpads or sensors. However, the wearable system may detect user operation or interaction or touch using the rectangular plate as a selection or input made via a virtual keyboard or virtual trackpad. The user input device 466 (shown in Figure 4) may be an embodiment of the totem, which may include a trackpad, touchpad, trigger, joystick, trackball, rocker or virtual switch, mouse, keyboard, multi-degree-of-freedom controller, or another physical input device. The user may use the totem alone or in combination with posture to interact with the wearable system and / or other users.
[0074] Examples of wearable devices, HMDs, and display systems and usable tactile devices and totems of the present disclosure are described in U.S. Patent Publication No. 2015 / 0016777 (which is incorporated herein in whole by reference).
[0075] (Examples of wearable systems, environments, and interfaces) Wearable systems may employ various mapping-related techniques to achieve high depth of field within the rendered light field. When mapping a virtual world, it is advantageous to capture all features and points in the real world and accurately depict virtual objects in relation to the real world. To achieve this objective, FOV images captured by the user of the wearable system can be added to the world model by including new images that convey information about various points and features of the real world. For example, a wearable system can collect a set of map points (2D points or 3D points, etc.), find new map points, and render a more accurate version of the world model. The world model of the first user can be communicated to a second user (e.g., via a network such as a cloud network) so that the second user can experience the world surrounding the first user.
[0076] Figure 7 is a block diagram of an embodiment of the MR environment 700. The MR environment 700 may be configured to receive inputs (e.g., visual input 702 from the user's wearable system, steady input 704 such as an indoor camera, sensor input 706 from various sensors, gestures, totems, eye tracking, user input, etc. from a user input device 466) from one or more of the user's wearable systems (e.g., wearable system 200 or display system 220) or steady indoor systems (e.g., indoor camera, etc.). The wearable system can use various sensors (e.g., accelerometer, gyroscope, temperature sensor, motion sensor, depth sensor, GPS sensor, inward-facing imaging system, outward-facing imaging system, etc.) to determine the location of the user's environment and various other attributes. This information may be further supplemented with information from a steady camera in the room, which may provide images from different viewpoints or various cues. Image data acquired by the camera (e.g., indoor camera or camera of an outward-facing imaging system) may be reduced to a set of mapping points.
[0077] One or more object recognition devices 708 can crawl through received data (e.g., a collection of points), recognize or map the points, tag images, and link semantic information to objects using a map database 710. The map database 710 may contain various points and their corresponding objects collected over time. The various devices and the map database can be interconnected through a network (e.g., LAN, WAN, etc.) and can access the cloud.
[0078] Based on this information and the point set in the map database, object recognition devices 708a-708n may recognize objects in the environment. For example, an object recognition device can recognize faces, people, windows, walls, user input devices, televisions, other objects in the user's environment, etc. One or more object recognition devices may be specialized for objects with certain characteristics. For example, object recognition device 708a may be used to recognize faces, while another object recognition device may be used to recognize totems.
[0079] Object recognition may be performed using various computer vision techniques. For example, a wearable system can analyze images acquired by an outward-facing imaging system 464 (shown in Figure 4) to perform scene reconstruction, event detection, video tracking, object recognition, object pose estimation, learning, indexing, motion estimation, or image restoration, etc. One or more computer vision algorithms may be used to perform these tasks. Non-restrictive examples of computer vision algorithms include scale-invariant feature transformation (SIFT), speed-up robust features (SURF), orientation FAST and rotation BRIEF (ORB), binary robust invariant scalable keypoint (BRISK), fast retinal keypoint (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous localization and mapping (vSLAM) techniques, sequential Bayesian estimators (e.g., Kalman filter, extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), iterative nearest neighbor (ICP), semi-global matching (SGM), semi-global block matching (SGBM), feature point histograms, various machine learning algorithms (e.g., support vector machines, k-nearest neighbor algorithm, Naive Bayes, neural networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.).
[0080] Object recognition can be performed, in addition to or alternatively, by various machine learning algorithms. Once trained, machine learning algorithms can be stored by the HMD. Some embodiments of machine learning algorithms may include supervised or unsupervised machine learning algorithms, and include regression algorithms (e.g., ordinary least-squares regression), instance-based algorithms (e.g., learning vector quantization), decision tree algorithms (e.g., classification and regression trees), Bayesian algorithms (e.g., Naive Bayes), clustering algorithms (e.g., k-means clustering), association rule learning algorithms (e.g., a priori algorithms), artificial neural network algorithms (e.g., Perceptron), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural networks), dimensionality reduction algorithms (e.g., principal component analysis), ensemble algorithms (e.g., Stacked Generalization), and / or other machine learning algorithms. In some embodiments, individual models can be customized for individual datasets. For example, a wearable device can generate or store a base model. The base model is used as a starting point and may generate additional models specific to data types (e.g., a specific user in a telepresence session), datasets (e.g., a set of additional images acquired from a user in a telepresence session), conditional situations, or other modifications. In some embodiments, the wearable HMD can be configured to utilize multiple techniques to generate a model for analyzing aggregated data. Other techniques may include using predefined thresholds or data values.
[0081] Based on the main information and set of points in the map database, object recognition devices 708a-708n may recognize objects, supplement them with semantic information, and give them life. For example, if an object recognition device recognizes that a set of points is a door, the system may combine some semantic information (e.g., the door has a hinge and moves 90 degrees around the hinge). If an object recognition device recognizes that a set of points is a mirror, the system may combine semantic information that the mirror has a reflective surface that can reflect images of objects in the room. Over time, the map database grows as the system (which may reside locally or be accessible via a wireless network) accumulates more data from around the world. Once an object is recognized, the information may be transmitted to one or more wearable systems. For example, the MR environment 700 may contain information about the scene being generated in California. The environment 700 may be transmitted to one or more users in New York. Based on data received from the FOV camera and other inputs, the object recognition device and other software components can map points collected from various images and recognize objects, etc., so that the scene can be accurately "passed" to a second user who may be located in a different part of the world. Environment 700 may also use a topology map for location identification purposes.
[0082] Figure 8 is a process flow diagram of an embodiment of Method 800 for rendering virtual content in relation to a recognized object. Method 800 describes a way in which a virtual scene may be represented to a user of a wearable system. The user may be geographically distant from the scene. For example, the user may be in New York but may want to view a scene currently happening in California, or may want to go for a walk with a friend who is in California.
[0083] In block 810, the wearable system may receive input about the user's environment from the user and other users. This may be achieved through various input devices and knowledge already held in a map database. The user's FOV camera, sensors, GPS, eye tracking, etc., transmit information to the system in block 810. In block 820, the system may determine sparse points based on this information. Sparse points may be used to determine posture data (e.g., head posture, eye posture, body posture, or hand gestures) which can be used to display and understand the orientation and position of various objects around the user. In block 830, object recognition devices 708a-708n may crawl through these collected points and recognize one or more objects using the map database. This information may then be transmitted to the user's individual wearable systems in block 840, and a desired virtual scene may be displayed to the user in block 850, as appropriate. For example, a desired virtual scene (e.g., a user in CA) may be displayed in an appropriate orientation, position, etc., in relation to the user's various objects and other surroundings in New York.
[0084] Figure 9 is a block diagram of another embodiment of a wearable system. In this embodiment, the wearable system 900 includes a map, which may include map data about the world. The map may reside partially locally on the wearable system, or partially in a networked storage location (e.g., within a cloud system) accessible by a wired or wireless network. An attitude process 910 may run on a wearable computing architecture (e.g., a processing module 260 or a controller 460) and utilize data from the map to determine the position and orientation of the wearable computing hardware or the user. The attitude data may be calculated from data collected on the fly as the user experiences the system and operates within its world. The data may include images of objects in a real or virtual environment, data from sensors (generally including accelerometer and gyroscope components, such as an inertial measurement unit), and surface information.
[0085] Sparse point representations may be the output of a simultaneous location identification and mapping (SLAM or V-SLAM) process (referring to configurations where the input is image / visual only). The system can be configured to discover not only the locations of various components within the world, but also what the world is made of. Pose can be building blocks that achieve many goals, including filling in maps and using data from maps.
[0086] In one embodiment, sparse point locations may not be entirely accurate in themselves, and further information may be required to generate a multi-focus AR, VR, or MR experience. Generally, dense representations, referring to depth map information, may be used, at least partially, to fill these gaps. Such information may be calculated from a process referred to as stereoscopic viewing 940, where depth information is determined using techniques such as triangulation or time-of-flight sensing. Image information and active patterns (such as infrared patterns generated using an active projector) may serve as inputs to the stereoscopic viewing process 940. A significant amount of depth map information may be fused together, some of which may be summarized using surface representations. For example, mathematically definable surfaces may be efficient (e.g., compared to large point clouds) and applicable inputs to other processing devices such as game engines. Thus, the output of the stereoscopic viewing process (e.g., depth map) 940 may be combined in the fusion process 930. The orientation may also be an input to the fusion process 930, and the output of the fusion 930 becomes an input to fill the map process 920. Subsurfaces may connect with each other in topographic mapping, etc., to form a larger surface, and the map becomes a large-scale hybrid of points and surfaces.
[0087] Various inputs may be used to resolve various aspects in the mixed reality process 960. For example, in the embodiment depicted in Figure 9, game parameters may be inputs for determining that the system's user is playing a monster battle game with one or more monsters in various locations, that the monsters are dead or have escaped under various conditions (e.g., when the user shoots the monsters), and for determining walls or other objects and equivalents in various locations. A world map may include information about the locations where such objects exist relative to each other, which is another useful input to mixed reality. Attitudes to the world are also inputs and play an important role for almost any interactive system.
[0088] User control or input is another input to the wearable system 900. As described herein, user input can include visual input, gestures, totems, audio input, sensor input, etc. For example, to move around or play a game, the user may need to command the wearable system 900 about what they want to do. There are various forms of user control that can be utilized, not just moving around in space on their own. In one embodiment, an object such as a totem (e.g., a user input device) or a toy gun may be held by the user and tracked by the system. The system would preferably be configured to know that the user is holding the item and to understand the type of interaction the user is having with the item (for example, if the totem or object is a gun, the system may be configured to understand not only its location and orientation, but also whether the user is clicking a trigger or other sensing button or element, which may be equipped with sensors (such as an IMU) that can help determine what is happening even when such activity is not within the field of view of any camera).
[0089] Hand gesture tracking or recognition may also provide input information. The wearable system 900 may be configured to track and interpret hand gestures for button presses, left or right gestures, stop gestures, grips, holds, etc. For example, in one configuration, the user may want to flip through email or a calendar in a non-gaming environment, or to "fist bump" with another person or performer. The wearable system 900 may be configured to take advantage of a minimum amount of hand gestures, which may or may not be dynamic. For example, the gestures may be simple static gestures, such as spreading the hand to indicate stop, raising the thumb to indicate OK, lowering the thumb to indicate not OK, or flipping the hand left or right or up or down to indicate a directional command.
[0090] Eye tracking is another input (for example, tracking where the user is looking, controlling display technology, and rendering at a specific depth or range). In one embodiment, eye convergence and divergence may be determined using triangulation, and then accommodation may be determined using a convergence / divergence / accommodation model developed for that particular person.
[0091] Regarding the camera system, the exemplary wearable system 900 shown in Figure 9 may include three pairs of cameras, namely a relative wide-field-of-view (FOV) or passive SLAM pair of cameras arranged on either side of the user's face, and a different pair of cameras oriented in front of the user to handle the stereoscopic imaging process 940 and to capture hand gestures and totem / object trajectories in front of the user's face. The FOV camera and the pair of cameras for the stereoscopic process 940 may be part of an outward-facing imaging system 464 (shown in Figure 4). The wearable system 900 may also include an eye-tracking camera oriented toward the user's eye (which may be part of an inward-facing imaging system 462 (shown in Figure 4)) to triangulate eye vectors and other information. The wearable system 900 may also include one or more textured light projectors (such as infrared (IR) projectors) to bring texture into the scene.
[0092] Figure 10 is a process flow diagram of an embodiment of method 1000 for determining user input to a wearable system. In this embodiment, the user may interact with a totem. The user may have multiple totems. For example, the user may have one designated totem for a social media application, another totem for playing a game, etc. In block 1010, the wearable system may detect the movement of the totem. The movement of the totem may be perceived through an outward-facing system or detected through sensors (e.g., tactile gloves, image sensors, hand tracking devices, eye tracking cameras, head posture sensors, etc.).
[0093] Based at least partially on detected gestures, eye postures, head postures, or inputs through the totem, the wearable system detects the position, orientation, and / or movement of the totem (or the user's eyes or head or gestures) relative to a reference frame in block 1020. The reference frame may be a set of map points on which the wearable system translates the movement of the totem (or user) into actions or commands. In block 1030, the user's interaction with the totem is mapped. Based on the mapping of the user interaction to the reference frame 1020, the system determines the user input in block 1040.
[0094] For example, the user may move a totem or physical object back and forth, turn a virtual page, move to the next page, or move from one user interface (UI) display screen to another UI screen. In another embodiment, the user may move their head or eyes to view different real or virtual objects within the user's FOR. If the user gazes at a particular real or virtual object for longer than a threshold time, that real or virtual object may be selected as user input. In some implementations, the convergence and divergence movements of the user's eyes can be tracked, and a near / far accommodation / convergence / divergence movement model can be used to determine the near / far accommodation state of the user's eyes, providing information about the depth plane in which the user is focused. In some implementations, the wearable system can use raycasting techniques to determine real or virtual objects aligned with the direction of the user's head or eye posture. In various implementations, raycasting techniques may include projecting a narrow beam of light with virtually no width, or projecting a beam of light with substantial width (e.g., a cone or frustum).
[0095] The user interface may be projected by a display system as described herein (e.g., display 220 in Figure 2). It may also be displayed using various other techniques, such as one or more projectors. The projectors may project images onto a physical object, such as a canvas or a sphere. Interaction with the user interface may be tracked using one or more cameras outside or within the system (e.g., using an inward-facing imaging system 462 or an outward-facing imaging system 464).
[0096] Figure 11 is a process flow diagram of an embodiment of Method 1100 for interacting with a virtual user interface. Method 1100 may be performed by a wearable system as described herein.
[0097] In block 1110, the wearable system may identify a specific UI. The type of UI may be given by the user. The wearable system may identify that a particular UI needs to be populated based on user input (e.g., gestures, visual data, audio data, sensory data, direct commands, etc.). In block 1120, the wearable system may generate data for a virtual UI. For example, data associated with the UI's boundaries, general structure, shape, etc., may be generated. In addition, the wearable system may determine the map coordinates of the user's physical location so that the wearable system can display the UI in relation to the user's physical location. For example, if the UI is body-centered, the wearable system may determine the coordinates of the user's physical standing position, head posture, or eye posture so that a ring UI can be displayed around the user, or a planar UI can be displayed on a wall or in front of the user. If the UI is hand-centered, the map coordinates of the user's hand may be determined. These map points may be derived through an FOV camera, data received through sensor inputs, or any other type of collected data.
[0098] In block 1130, the wearable system may send data from the cloud to the display, or the data may be sent from a local database to the display component. In block 1140, the UI is displayed to the user based on the transmitted data. For example, a light field display may project the virtual UI into one or both of the user's eyes. Once the virtual UI is generated, in block 1150, the wearable system may simply wait for a command from the user and generate more virtual content on the virtual UI. For example, the UI may be a body-centered ring around the user's body. The wearable system may then wait for a command (gesture, head or eye movement, input from a user input device, etc.), and if recognized (block 1160), the virtual content associated with the command may be displayed to the user (block 1170). In one embodiment, the wearable system may wait for a hand gesture from the user before mixing multiple stem tracks.
[0099] Additional embodiments of wearable systems, UI, and user experience (UX) are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein in whole by reference.
[0100] (Overview of user interactions based on context information) Wearable systems can support various user interactions with objects within FOR based on contextual information. For example, a wearable system can adjust the size of the cone projection's opening as the user interacts with an object using the cone projection. In another embodiment, a wearable system can adjust the amount of movement of a virtual object associated with the operation of a user input device based on contextual information. Detailed embodiments of these interactions are provided below.
[0101] (Example object) The user's FOR may contain a group of objects that can be perceived by the user through the wearable system. The objects within the user's FOR 1200 may be virtual and / or physical objects. Virtual objects may include, for example, operating system objects such as a trash can for deleted files, a terminal for entering commands, a file manager for accessing files or directories, icons, menus, applications for audio or video streaming, and notifications from the operating system. Virtual objects may also include objects within applications, such as avatars, virtual objects in games, graphics, or images. Some virtual objects may be both operating system objects and objects within applications. In some embodiments, the wearable system may add virtual elements to existing physical objects. For example, the wearable system may add a virtual menu associated with a television in a room, which may give the user options to turn on the television or change channels using the wearable system.
[0102] Objects within the user's FOR can be part of a world map, as illustrated with reference to Figure 9. Data associated with objects (e.g., location, semantic information, properties, etc.) can be stored in various data structures, such as arrays, lists, trees, hashes, and graphs. The index of each stored object may be determined, where applicable, by the object's location. For example, the data structure may index objects by a single coordinate, such as the distance of the object from a reference point (e.g., distance to the left (or right) of the reference point, distance from the top (or bottom) of the reference point, or depth from the reference point). In some implementations, the wearable system includes a light field display capable of displaying virtual objects to the user in different depth planes. Interactive objects can be organized into multiple arrays located in different fixed depth planes.
[0103] A user can interact with a subset of objects within their FOR. This subset of objects may sometimes be referred to as interactable objects. A user can interact with objects using various techniques, such as by selecting an object, moving an object, opening a menu or toolbar associated with an object, or selecting a new set of interactable objects. A user may also interact with interactable objects by using hand gestures to activate a user input device (see, for example, user input device 466 in Figure 4), such as by clicking a mouse, tapping a touchpad, swiping a touchscreen, placing a hand over or touching a capacitive button, pressing a key on a keyboard or game controller (e.g., a 5-way d-pad), pointing a joystick, wand, or totem towards an object, pressing a button on a remote control, or other interactions with a user input device. A user may also interact with interactable objects using head, eye, or body posture, such as by gazing at or pointing at an object for a period of time. These hand gestures and user postures can trigger selection events in the wearable system, for example, by performing user interface actions (such as displaying a menu associated with a target interactable object, or performing game actions on an in-game avatar).
[0104] (Example of cone projection) As described herein, a user can interact with objects in their environment using their posture. For example, a user looking around a room might see a table, a chair, a wall, and a virtual television display on one of the walls. To determine the object the user is looking at, a wearable system may use a cone projection technique, which, generally speaking, projects an invisible cone in the direction the user is looking and identifies any object that intersects the cone. Cone projection may involve projecting a single ray, which has no lateral thickness, from the HMD (of the wearable system) toward a physical or virtual object. Cone projection with a single ray may also be called ray projection.
[0105] Ray projection can use collision detection agents to trace along the ray and identify whether any object intersects the ray and where. A wearable system can track the user's posture (e.g., body, head, or eye direction) using inertial measurement units (e.g., accelerometers), eye-tracking cameras, etc., to determine the direction the user is looking. The wearable system can use the user's posture to determine the direction from which the ray should be projected. Ray projection techniques can also be used in conjunction with user input devices such as handheld multi-degree-of-freedom (DOF) input devices. For example, a user can activate a multi-DOF input device to anchor the size and / or length of the ray while the user moves around. In another embodiment, instead of projecting the ray from an HMD, the wearable system can project the ray from a user input device.
[0106] In one embodiment, instead of projecting a ray with negligible thickness, the wearable system can project a cone with a non-negligible aperture (crossing the central ray 1224). Figure 12A illustrates an embodiment of cone projection with a non-negligible aperture. Cone projection can project a conical (or other shape) volume 1220 with an adjustable aperture. The cone 1220 can be a geometric cone, which has a proximal end 1228a and a distal end 1228b. The size of the aperture can correspond to the size of the distal end 1228b of the cone. For example, a large aperture may correspond to a large surface area at the distal end 1228b of the cone (e.g., the end away from the HMD, user, or user input device). In another embodiment, a large opening can correspond to a large diameter 1226 on the distal end 1228b of the cone 1220, while a small opening can correspond to a small diameter 1226 on the distal end 1228b of the cone 1220. As further illustrated with reference to Figure 12A, the proximal end 1228a of the cone 1220 may have its origin at various positions, for example, the center of the user's ARD (e.g., between the user's eyes), a point on one of the user's limbs (e.g., the hand, such as the fingers), or a user input device or totem (e.g., a toy weapon) held or manipulated by the user.
[0107] The central ray 1224 can represent the direction of the cone. The direction of the cone can correspond to the user's body posture (head posture, hand gestures, etc.) or the user's gaze direction (also referred to as eye posture). Embodiment 1206 in Figure 12A illustrates a cone projection with posture, and the wearable system can determine the direction of the cone 1224 using the user's head posture or eye posture. This embodiment also illustrates a coordinate system for head posture. The head 1250 may have multiple degrees of freedom. As the head 1250 moves toward different directions, the head posture will change with respect to the natural resting direction 1260. The coordinate system in Figure 12A shows three angular degrees of freedom (e.g., yaw, pitch, and roll) which can be used to measure the head posture relative to the natural resting state 1260 of the head. As shown in Figure 12A, the head 1250 can tilt forward and backward (e.g., pitch), change direction left and right (e.g., yaw), and tilt laterally (e.g., roll). Other implementations may also use other techniques or angular representations for measuring head posture, such as any other type of Euler angle system. The wearable system may use an IMU to determine the user's head posture. An inward-facing imaging system 462 (shown in Figure 4) may be used to determine the user's eye posture.
[0108] Example 1204 illustrates another embodiment of cone projection with orientation, in which a wearable system can determine the orientation 1224 of the cone based on the user's hand gestures. In this embodiment, the proximal end 1228a of the cone 1220 is at the tip of the user's finger 1214. As the user points their finger in a different direction, the position of the cone 1220 (and central ray 1224) can be moved accordingly.
[0109] The direction of the cone can also correspond to the position or orientation of the user input device or the operation of the user input device. For example, the direction of the cone may be based on the user drawing trajectory on the touch surface of the user input device. The user can move their finger forward on the touch surface to indicate that the direction of the cone is forward. Example 1202 illustrates another cone projection using a user input device. In this example, the proximal end 1228a is located at the tip of a weapon-shaped user input device 1212. As the user input device 1212 is moved, the cone 1220 and the central ray 1224 can also move with the user input device 1212.
[0110] The direction of the cone can further be based on the position or orientation of the HMD. For example, the cone may be projected in a first direction when the HMD is tilted, and in a second direction when the HMD is not tilted.
[0111] (Start of cone projection) The wearable system can initiate cone projection when user 1210 activates user input device 466, for example, by clicking a mouse, tapping a touchpad, swiping a touchscreen, placing or touching a capacitive button, pressing a key on a keyboard or game controller (e.g., a 5-way d-pad), pointing a joystick, wand, or totem towards an object, pressing a button on a remote control, or through other interaction with a user input device.
[0112] The wearable system may also initiate cone projection based on the user's posture 1210, such as prolonged gazing in one direction or a hand gesture (e.g., waving a hand in front of an outward-facing imaging system 464). In some implementations, the wearable system can automatically initiate cone projection events based on contextual information. For example, the wearable system may automatically initiate cone projection when the user is on the main page of the AR display. In another embodiment, the wearable system can determine the relative positions of objects in the user's gaze direction. If the wearable system determines that the objects are located relatively far apart from each other, the wearable system may automatically initiate cone projection, and thus the user does not need to move with precision to select an object within a group of sparsely located objects.
[0113] (Exemplary properties of a cone) The cone 1220 may have various properties, such as size, shape, or color. These properties may be displayed to the user so that the cone is perceptible to the user. In some cases, a part of the cone 1220 may be displayed (e.g., the end of the cone, the surface of the cone, the central ray of the cone, etc.). In other embodiments, the cone 1220 may be a cuboid, polyhedron, pyramid, frustum of a cone, etc. The distal end 1228b of the cone may have any cross-section, such as circular, oval, polygonal, or irregular.
[0114] In Figures 12A, 12B, and 12C, the cone 1220 may have a proximal end 1228a and a distal end 1228b. The proximal end 1228a (also referred to as the zero point of the central ray 1224) can be associated with the location from which the cone projection occurs. The proximal end 1228a may be anchored to a location in 3D space from which the virtual cone appears to emanate. The location may be a position on the user's head (e.g., between the user's eyes), a user input device that functions as a pointer (e.g., a 6DOF or 3DOF handheld controller), a fingertip (which can be detected by gesture recognition), etc. With respect to handheld controllers, the location to which the proximal end 1228a is anchored may depend on the shape factor of the device. For example, in a weapon-shaped controller 1212 (for use in a shooting game), the proximal end 1228a may be at the tip of the muzzle of the controller 1212. In this embodiment, the proximal end 1228a of the cone originates from the center of the barrel, and the cone 1220 (or the central ray 1224 of the cone 1220) can be projected forward such that the center of the cone projection will be concentric with the barrel of the weapon-shaped controller 1212. In various embodiments, the proximal end 1228a of the cone can be anchored at any location in the user's environment.
[0115] Once the proximal end 1228a of the cone 1220 is anchored to a location, the orientation and movement of the cone 1220 may be based on the movement of an object associated with that location. For example, as described with reference to Example 1206, if the cone is anchored to the user's head, the cone 1220 can move based on the user's head posture. In another embodiment, in Example 1202, if the cone 1220 is anchored to a user input device, the cone 1220 can move based on the operation of the user input device, for example, based on a change in the position or orientation of the user input device.
[0116] The distal end 1228b of the cone can extend until it reaches a termination threshold. The termination threshold may involve a collision between the cone and a virtual or physical object in the environment (e.g., a wall). The termination threshold may also be based on a threshold distance. For example, the distal end 1228b can continue to extend away from the proximal end 1228a until the cone collides with an object, or until the distance between the distal end 1228b and the proximal end 1228a reaches a threshold distance (e.g., 20 centimeters, 1 meter, 2 meters, 10 meters, etc.). In some embodiments, the cone can extend beyond an object, even if a collision may occur between the cone and the object. For example, the distal end 1228b can extend through a real-world object (a table, chair, wall, etc.) and terminate when it reaches a termination threshold. Assuming the termination threshold is the wall of a virtual room located outside the user's current room, the wearable system can allow the cone to extend beyond the current room until it reaches the surface of the virtual room. In one embodiment, a world mesh can be used to define the extent of one or more rooms. The wearable system can detect the presence of a termination threshold by determining whether a virtual cone intersects with a portion of the world mesh. Advantageously, in some embodiments, when the cone extends through real-world objects, the user can easily target virtual objects. As an example, the HMD can present a virtual hole on a physical wall, through which the user can remotely interact with virtual content in other rooms, even if the user is not physically present in those rooms. The HMD can determine objects in other rooms based on a world map as illustrated in Figure 9.
[0117] The cone 1220 may have a depth. The depth of the cone 1220 may be represented by the distance between the proximal end 1228a and the distal end 1228b of the cone 1220. The depth of the cone can be automatically adjusted by the wearable system, the user, or a combination thereof. For example, if the wearable system determines that an object is located away from the user, the wearable system may increase the depth of the cone. In some implementations, the depth of the cone may be anchored to a depth plane. For example, the user may choose to anchor the depth of the cone to a depth plane within 1 meter of the user. As a result, during cone projection, the wearable system will not capture objects outside the 1-meter boundary. In some embodiments, if the depth of the cone is anchored to a depth plane, the cone projection will capture only objects in that depth plane. Therefore, the cone projection will not capture objects that are closer to or further away from the user than the anchored depth plane. In addition to, or as an alternative to, setting the depth of the cone 1220, the wearable system may set the distal end 1228b to the depth plane so that the cone projection can enable user interaction with objects in or below the depth plane.
[0118] A wearable system can anchor the depth, proximal end 1228a, or distal end 1228b of a cone in response to detection of a hand gesture, body posture, gaze direction, activation of a user input device, voice command, or other technique. In addition to or alternative to the embodiments described herein, the anchoring location of the proximal end 1228a, distal end 1228b, or anchored depth may be based on contextual information, such as the type of user interaction or the function of the object to which the cone is anchored. For example, the proximal end 1228a may be anchored to the center of the user's head due to user utility and operability. In another embodiment, when a user points to an object using a hand gesture or a user input device, the proximal end 1228a can be anchored to the user's fingertip or the tip of the user input device to increase the accuracy of the direction the user is pointing.
[0119] The cone 1220 may have a color. The color of the cone 1220 may depend on the user's preferences, the user's environment (virtual or physical), etc. For example, when the user is in a virtual jungle, which is a collection of trees with lush green leaves, the wearable system may provide a dark gray cone and increase the contrast between the cone and objects in the user's environment so that the user can have better visibility of the cone's location.
[0120] A wearable system can generate a visual representation of at least a portion of a cone for display to the user. The properties of the cone 1220 may be reflected in the visual representation of the cone 1220. The visual representation of the cone 1220 may correspond to at least a portion of the cone, such as the cone's opening, surface, or central ray. For example, if the virtual cone is a geometric cone, the visual representation of the virtual cone may include a gray geometric cone extending from the user's interpupillary position. In another embodiment, the visual representation may include a portion of the cone that interacts with real or virtual content. Assuming the virtual cone is a geometric cone, the visual representation may include a circular pattern representing the base of the geometric cone, since the base of the geometric cone may be used to target and select virtual objects. In one embodiment, the visual representation is triggered based on user interface behavior. In an embodiment, the visual representation may be associated with the state of an object. The wearable system may present a visual representation when the object changes from a stationary state or a hovering state (where the object can be moved or selected). Furthermore, wearable systems can conceal their visual representation when an object changes from a hovering state to a selected state. In some implementations, when an object is in a hovering state, the wearable system can receive input from a user input device (in addition to or alternative to cone projection), allowing the user to select the virtual object using the user input device while the object is in a hovering state.
[0121] In one embodiment, the cone 1220 may be invisible to the user. The wearable system may assign the focus indicator to one or more objects to indicate the direction and / or location of the cone. For example, the wearable system may assign the focus indicator to an object that is in front of the user and intersects the user's gaze direction. The focus indicator may have a halo, color, perceived size or depth changes (e.g., to make the target object appear closer and / or larger when selected), shape changes of the cursor sprite graphic (e.g., the cursor changes from a circle to an arrow), or other audible, tactile, or visual effects to attract the user's attention.
[0122] The cone 1220 may have an aperture that intersects the central ray 1224. In some embodiments, the central ray 1224 is invisible to the user 1210. The size of the aperture may correspond to the size of the distal end 1228b of the cone. For example, a large aperture may correspond to a large diameter 1226 at the distal end 1228b of the cone 1220, while a small aperture may correspond to a small diameter 1226 at the distal end 1228b of the cone 1220.
[0123] As further illustrated with reference to Figures 12B and 12C, the aperture can be adjusted by the user, the wearable system, or a combination thereof. For example, the user may adjust the aperture through user interface actions, such as selecting an aperture option shown on the AR display. The user may also adjust the aperture by activating a user input device, for example, by scrolling the user input device or by pressing a button to anchor the aperture size. In addition to or in lieu of user input, the wearable system may update the aperture size based on one or more contextual factors described below.
[0124] (An example of cone projection with dynamically updated aperture) Cone projection can be used to increase accuracy when interacting with objects in the user's environment, especially when those objects are located at a distance where a small movement from the user can be converted into a large movement of light rays. Cone projection can also be used to reduce the amount of movement required from the user in order to overlap the cone with one or more virtual objects. In some implementations, the user can improve the speed and accuracy of selecting target objects by manually updating the cone's aperture, for example, by using a narrower cone when many objects are present and a wider cone when fewer objects are present. In other implementations, the wearable system can determine a contextual factor associated with the objects in the user's environment, in addition to or as an alternative to manual updates, enabling automatic cone updates, which is advantageous as it requires little user input and thus makes it easier for the user to interact with objects in the environment.
[0125] Figures 12B and 12C provide an example of cone projection onto a group 1230 of objects (e.g., 1230a, 1230b, 1230c, 1230d, 1230e) within the user's FOR 1200 (at least some of these objects are within the user's FOV). The objects may be virtual and / or physical objects. During cone projection, the wearable system can project a cone (visible or invisible to the user) 1220 in a certain direction and identify any objects that intersect the cone 1220. For example, in Figure 12B, object 1230a (shown by a thick line) intersects the cone 1220. In Figure 12C, objects 1230d and 1230e (shown by thick lines) intersect the cone 1220. Objects 1230b and 1230c (shown in gray) are outside cone 1220 and do not intersect with cone 1220.
[0126] A wearable system can automatically update its aperture based on contextual information. Contextual information may include information related to the user's environment (e.g., lighting conditions of the user's virtual or physical environment), user preferences, user physical conditions (e.g., whether the user is nearsighted), type of object in the user's environment (e.g., physical or virtual), information associated with objects in the user's environment, or the layout of objects (e.g., object density, object location and size), characteristics of objects the user interacts with (e.g., object function, type of user interface behavior supported by the object), combinations thereof, or equivalents. Density can be measured in various ways, such as the number of objects per projected area or the number of objects per solid angle. Density may also be expressed in other ways, such as the spacing between neighboring objects (smaller spacing reflects increased density). The wearable system can use object location information to determine the layout and density of objects within a region. As shown in Figure 12B, the wearable system may determine that the density of object group 1230 is high. The wearable system may therefore use a cone 1220 with a smaller opening. In Figure 12C, since objects 1230d and 1230c are located relatively far apart from each other, the wearable system may use a cone 1220 with a larger opening (compared to the cone in Figure 12B). Additional details regarding the calculation of object density and the adjustment of the opening size based on density are further explained in Figures 12D-12G.
[0127] The wearable system can dynamically update its opening (e.g., size or shape) based on the user's posture. For example, the user might initially be looking at the group of objects 1230 in Figure 12B, but as the user turns their head, they might end up looking at the group of objects in Figure 12C (where the objects are sparsely spaced relative to each other). As a result, the wearable system may increase the size of the opening (e.g., as shown by the change in the cone opening between Figure 12B and Figure 12C). Similarly, if the user turns their head back and looks at the group of objects 1230 in Figure 12B, the wearable system may decrease the size of the opening.
[0128] In addition, or alternatively, the wearable system can update the opening size based on user preferences. For example, if the user prefers to select a large group of items simultaneously, the wearable system may increase the opening size.
[0129] As another embodiment of dynamically updating the aperture based on contextual information, if the user is in a dark environment or is nearsighted, the wearable system may increase the size of the aperture to make it easier for the user to capture objects. In one implementation, the first cone projection may capture multiple objects. The wearable system may perform a second cone projection to further select a target object among the captured objects. The wearable system may also allow the user to select a target object from the captured objects using body posture or a user input device. The object selection process may be a recursive process, with one, two, three, or more cone projections being performed to select a target object.
[0130] (An example of dynamic aperture updating based on object density) As illustrated with reference to Figures 12B and 12C, the cone opening can be dynamically updated during cone projection based on the density of objects in the user's FOR. Figures 12D, 12E, 12F, and 12G illustrate an example of dynamically adjusting the opening based on object density. Figure 12D illustrates a contour map associated with the density of objects in the user's FOR 1208. Virtual objects 1271 are represented by small textured dots. The density of virtual objects is reflected by the amount of contour lines in a given region. For example, contour lines are close together in region 1272, indicating a high density of objects in region 1272. In another example, the contour lines in region 1278 are relatively sparse. Therefore, the density of objects in region 1278 is low.
[0131] The visual presentation of the opening 1270 is illustrated as a shaded circle in Figure 12D. The visual representation in this embodiment can correspond to the distal end 1228b of the virtual cone 1220. The opening size can vary based on the density of objects in a given region. For example, the opening size can depend on the density of objects where the center of the circle is located. As illustrated in Figure 12D, when the opening is in region 1272, the size of the opening 1270 can decrease (as indicated by the relatively small size of the opening circle). However, when the user starts from region 1276 in FOR 1208, the size of the opening 1270 is slightly larger in region 1272. When the user further changes their head position and looks at region 1274, the size of the opening is larger than in region 1276 because the density of objects in region 1274 is lower than in region 1276. In yet another embodiment, in region 1278, the size of the opening 1270 would increase because there are virtually no arbitrary objects within region 1278 of FOR 1208. Density is illustrated in these embodiments using contour maps, but density can also be determined using heat maps, surface plots, or other graphical or numerical representations. In general, the term “contour map” also includes these other types of density representations (in 1D, 2D, or 3D). Furthermore, contour maps may generally not be presented to the user but may be calculated and used by the ARD processor to dynamically determine the properties of the cone. Contour maps may be dynamically updated as physical or virtual objects move within the user’s FOV or FOR.
[0132] Various techniques can be employed to calculate object density. In one embodiment, density can be calculated by counting all virtual objects within the user's FOV. The number of virtual objects may be used as input to a function that defines the size of the aperture based on the number of virtual objects within the FOV. Image 1282a in Figure 12E shows an FOV with three virtual objects represented by a circle, an ellipse, and a triangle, and a virtual representation of the aperture 1280 illustrated using a textured circle. However, if the number of virtual objects decreases from three (in Image 1282a) to two (in Image 1282b), the size of the aperture 1280 can increase accordingly. A wearable system can calculate the increase using function 1288 in Figure 12F. In this figure, the size of the aperture is represented by the y-axis 1286b, while the number (or density) of virtual objects within the FOV is represented by the x-axis 1286a. As illustrated, as the number of virtual objects increases (e.g., density increases), the aperture size decreases according to function 1288. In one embodiment, the minimum aperture size is zero, which reduces the cone to a single ray. Function 1288 is illustrated as a linear function, but any other type of function, such as one or more power-law functions, may also be used. In some embodiments, function 1288 may include one or more threshold conditions. For example, when the density of objects reaches a certain low threshold, the size of aperture 1280 will no longer increase, even if the density of objects may decrease further. On the other hand, when the density of objects reaches a certain high threshold, the size of aperture 1280 will no longer decrease, even if the density of objects may increase further. However, when the density is between the low and high thresholds, the aperture size may decrease according to, for example, an exponential function.
[0133] Figure 12G illustrates another exemplary technique for calculating density. For example, in addition to, or as an alternative to, calculating the number of virtual objects in the FOV, the wearable system can calculate the percentage of the FOV covered by virtual objects. Images 1292a and 1292b illustrate adjustment of aperture size based on the number of objects in the FOV. As illustrated in this embodiment, the percentage of the FOV covered by the virtual image differs between images 1292a and 1292b (objects in image 1292a are more sparsely located), but the size of aperture 1280 does not change in these two images because the number of objects (e.g., three virtual objects) is the same across images 1292a and 1292b. In contrast, images 1294a and 1294b illustrate adjustment of aperture size based on the percentage of the FOV covered by virtual objects. As shown in image 1294a, aperture 1280 will increase in size because a lower percentage of the FOV will be covered by virtual objects (in contrast to the rest of the same in image 1292a).
[0134] (Examples of collisions) A wearable system can determine whether one or more objects collide with a cone during its projection. The wearable system may use a collision detection agent to detect collisions. For example, a collision detection agent can identify objects that intersect the surface of the cone and / or objects that are inside the cone. Such identification can be made based on the volume and location of the cone and the location information of the objects (stored in a world map, as illustrated with reference to Figure 9). Objects in the user's environment may be associated with a mesh (also referred to as a world mesh). A collision detection agent can determine whether a portion of the object's mesh overlaps with the portion relating to the cone and detect a collision. In some implementations, the wearable system may be configured to detect collisions only between the cone and objects on a certain depth plane.
[0135] The wearable system may provide a focus indicator for objects that intersect with the cone. For example, in Figures 12B and 12C, the focus indicator may be a red highlight around all or part of the object. Thus, in Figure 12B, when the wearable system determines that object 1230a intersects with cone 1220, the wearable system can display a red highlight around object 1230a to the user 1210. Similarly, in Figure 12C, the wearable system identifies objects 1230e and 1230d as objects that intersect with cone 1220. The wearable system can provide red highlights around objects 1230d and 1230e.
[0136] When a collision involves multiple objects, a wearable system may present user interface elements to help the user select one or more objects from among them. For example, a wearable system could provide a focus indicator that shows the target object the user is currently interacting with. The user can then use hand gestures to activate a user input device and move the focus indicator to another target object.
[0137] In some embodiments, an object may be behind another object in the user's 3D environment (for example, a nearby object may, at least partially, occlude a more distant object). Advantageously, the wearable system may apply disambiguation techniques during the cone projection (e.g., to determine occluded objects, determine the depth order or position between occluded objects, etc.) to capture both objects in front and behind. For example, a paper shredder may be behind a computer in the user's room. The user may not be able to see the shredder (because it is obscured by the computer), but the wearable system can project a cone towards the computer and detect a collision with both the shredder and the computer (because both the shredder and the computer are within the wearable system's world map). The wearable system may display a pop-up menu offering the user the option to select either the shredder or the computer, or the wearable system may use contextual information to determine which object to select (for example, if the user is attempting to delete a document, the system may select the paper shredder). In one implementation, the wearable system may be configured to capture only objects directly in front of it. In this embodiment, the wearable system would detect only the collision between the cone and the shredder.
[0138] In response to collision detection, the wearable system may enable the user to interact with the interactable object in various ways, such as selecting the object, moving the object, opening a menu or toolbar associated with the object, or performing game actions on an avatar in a game. The user may interact with the interactable object through posture (e.g., head, body posture), hand gestures, input from a user input device, a combination thereof, or equivalents. For example, when a cone collides with multiple interactable objects, the user may activate a user input device to select from among the multiple interactable objects.
[0139] (An illustrative process for dynamically updating the opening) Figure 13 is a flowchart illustrating an exemplary process for selecting an object using conical projection with a dynamically adjustable aperture. This process 1300 can be performed by a wearable system (shown in Figures 2 and 4).
[0140] In block 1310, the wearable system can initiate cone projection. Cone projection can be triggered by the user's posture or a hand gesture on a user input device. For example, cone projection may be triggered by a click on a user input device and / or by the user looking in a certain direction for an extended period of time. As shown in block 1320, the wearable system can analyze prominent features of the user's environment, such as the type of object, the layout of the object (physical or virtual), the location of the object, the size of the object, the density of the object, and the distance between the object and the user. For example, the wearable system can calculate the density of objects in the user's gaze direction by determining the number and size of objects in front of the user. Prominent features of the environment may be part of the contextual information described herein.
[0141] In block 1330, the wearable system can adjust the size of the opening based on contextual information. As discussed with reference to Figures 12B and 12C, the wearable system can increase the opening size when objects are sparsely located and / or when no obstacles are present. A large opening size can correspond to a large diameter 1226 on the distal end 1228b of the cone 1220. As the user moves around and / or the environment changes, the wearable system may update the size of the opening based on contextual information. Contextual information can be combined with other information such as user preferences, user posture, and characteristics of the cone (e.g., depth, color, location, etc.) to determine and update the opening.
[0142] The wearable system can render a conical projection visualization in block 1340. The conical projection visualization may include a cone with a non-negligible aperture. As illustrated with reference to Figures 12A, 12B, and 12C, the cone may have various sizes, shapes, or colors.
[0143] In block 1350, the wearable system can transform the cone projection and scan for collisions. For example, the wearable system can transform the amount of movement of the cone using the technique described with reference to Figures 16-18. The wearable system can also determine whether the cone is colliding with one or more objects by calculating the position of the cone relative to the positions of objects in the user's environment. As discussed with reference to Figures 12A, 12B, and 12C, one or more objects may intersect the surface of the cone or be located within the cone.
[0144] If the wearable system does not detect a collision, in block 1360, the wearable system repeats block 1320, and the wearable system can analyze the user's environment and update the aperture based on the user's environment (as shown in block 1330). If the wearable system detects a collision, the wearable system can indicate the collision, for example, by placing a focus indicator on the object that was hit. If the cone collides with multiple interactable objects, the wearable system can use ambiguity resolution techniques to capture one or more occluded objects.
[0145] In block 1380, the user can interact with the collided object in various ways, as described with reference to Figures 12A, 12B, and 12C. For example, the user can select the object, open a menu associated with the object, or move the object.
[0146] Figure 14 is another flowchart illustrating an exemplary process for selecting objects using conical projection with a dynamically adjustable aperture. This process 1400 can be performed by a wearable system (shown in Figures 2 and 4). In block 1410, the wearable system determines the group of objects in the user's FOR.
[0147] In block 1420, the wearable system can initiate a cone projection onto a group of objects within the user's FOR. The wearable system can initiate the cone projection based on input from a user input device (e.g., a swing of a wand) or posture (e.g., a hand gesture). The wearable system can also automatically trigger the cone projection based on certain conditions. For example, the wearable system may automatically initiate the cone projection when the user is in front of the wearable system's main display. The cone projection may use a virtual cone, which may have a central ray and an aperture that crosses the central ray. The central ray may be based on the user's gaze direction.
[0148] In block 1430, the wearable system can determine the user's posture. The user's posture may be head, eye, or body posture, either alone or in combination. Based on the user's posture, the wearable system can determine the user's field of view (FOV). The FOV may include a portion of the FOR perceived by the user at a given time.
[0149] Based on the user's field of view (FOV), block 1440 allows the wearable system to determine subgroups of objects within the user's FOV. As the user's FOV changes, the objects within the user's FOV may also change. The wearable system may be configured to analyze contextual information about the objects within the user's FOV. For example, the wearable system may determine the density of objects based on their size and location within the FOV.
[0150] In block 1450, the wearable system can determine the size of the aperture with respect to a cone projection event. The size of the aperture may be determined based on contextual information. For example, if the wearable system determines that the object density is high, it may use a cone with a small aperture to increase the accuracy of user interaction. In some embodiments, the wearable system can also adjust the depth of the cone. For example, if the wearable system determines that all objects are located away from the user, it may extend the cone to the depth plane containing these objects. Similarly, if the wearable system determines that objects are located close to the user, it may reduce the depth of the cone.
[0151] The wearable system can generate a visual representation of a cone projection in block 1460. The visual representation of the cone projection can incorporate the properties of a cone, as illustrated with reference to Figures 12B and 12C. For example, the wearable system can display a virtual cone with color, shape, and depth. The location of the virtual cone may be associated with the user's head posture, body posture, or gaze direction. The cone may be a geometric cone, a cuboid, a polyhedron, a pyramid, a frustum of a cone, or any other three-dimensional shape, which may or may not be a regular shape.
[0152] As the user moves around, the cone can also move with the user. As illustrated further with reference to Figure 15-18, as the user moves around, the amount of cone movement corresponding to the user's movement can also be calculated based on contextual information. For example, if the density of objects in the FOV is low, a small movement of the user can result in a large movement of the cone. On the other hand, if the density is high, the same movement may result in a smaller movement of the cone, thereby enabling more refined interactions with objects.
[0153] Figure 15 shows an exemplary process 1500 for conical projection with a dynamically adjustable aperture. Process 1500 in Figure 15 can be performed by a wearable system (shown in Figures 2 and 4). In block 1510, the wearable system can determine contextual information within the user's environment. Contextual information may include information about the user's environment and / or information associated with objects, such as the layout of objects, the density of objects, and the distance between objects and the user.
[0154] In block 1520, the wearable system can project a cone with a dynamically adjustable aperture based on contextual information. For example, when the object density is low, the aperture may be large.
[0155] In block 1530, the wearable system can detect collisions between an object and a cone. In some embodiments, the wearable system can detect collisions based on the location of the object and the location of the cone. A collision is detected if at least a portion of the object overlaps with the cone. In some embodiments, the cone may collide with multiple objects. The wearable system can apply ambiguity resolution techniques to capture one or more occluded objects. As a result, the wearable system can detect collisions between the cone and occluded objects.
[0156] Upon collision detection, the wearable system may assign a focus indicator to the object colliding with the cone. The wearable system may also provide user interface options, such as selecting an object from the collided objects. In block 1540, the wearable system may be configured to receive user interaction with the collided object. For example, the user may move the object, open a menu associated with the object, or select the object.
[0157] (Overview of context-based movement transformations) During conical projection, in addition to adjusting the cone's opening, or as an alternative, contextual information can also be used to translate movements associated with a user input device or a part of the user's body (e.g., changes in the user's posture) into user interface actions, such as moving a virtual object.
[0158] Users can move virtual objects or shift focus indicators by activating user input devices and / or by using posture such as head, eye, or body position. As is evident in the AR / VR / MR world, the movement of virtual objects does not refer to the actual physical movement of virtual objects, since virtual objects are computer-generated images and not physical objects. The movement of virtual objects refers to the apparent movement of virtual objects as they are displayed to the user by the AR or VR system.
[0159] Figure 16 schematically illustrates an embodiment of moving a virtual object using a user input device. For example, a user may hold and move a virtual object by selecting it using the user input device, or by physically moving the user input device 466. The user input device 466 may initially be at a first position 1610a. A user 1210 may select a target virtual object 1640 located at the first position 1610b by activating the user input device 466 (for example, by activating a touch sensor pad on the device). The target virtual object 1640 can be any type of virtual object that can be displayed and moved by a wearable system. For example, the virtual object may be an avatar, a user interface element (e.g., a virtual display), or any type of graphical element displayed by a wearable system (e.g., a focus indicator). User 1210 can move the target virtual object from a first position 1610b to a second position 1620b by moving the user input device 466 along the trajectory 1650b. However, since the target virtual object may be far from the user, the user may need to move the user input device a long distance before the target virtual object reaches its desired location, which may require the user to use large hand and arm movements and ultimately lead to user fatigue.
[0160] Embodiments of wearable systems may provide a technique for rapidly and efficiently moving distant virtual objects by moving the virtual object using a multiplier that tends to increase with distance to the virtual object and a quantity based on the movement of the controller. Such embodiments may advantageously allow the user to move distant virtual objects using shorter hand and arm movements, thereby reducing user fatigue.
[0161] A wearable system can calculate a multiplier to map the movement of a user input device to the movement of a target virtual object. The movement of the target virtual object may be based on the movement of the input controller and the multiplier. For example, the amount of movement of the target virtual object may be equal to the amount of movement of the input controller multiplied by the multiplier. This can reduce the amount that the user needs to move before the target virtual object reaches the desired location. For example, as shown in Figure 16, a wearable system may determine a multiplier that allows the user to move the user input device along a trajectory 1650a (shorter than trajectory 1650b) to move the virtual object from position 1620b to position 1610b.
[0162] In addition, or alternatively, the user 1210 can move a virtual object using head pose. For example, as shown in Figure 16, the head may have multiple degrees of freedom. As the head moves in different directions, the head pose will change with respect to the natural resting direction 1260. The exemplary coordinate system in Figure 16 shows three angular degrees of freedom (e.g., yaw, pitch, and roll) that can be used to measure the head pose with respect to the head's natural resting state 1260. As illustrated in Figure 16, the head can tilt forward and backward (e.g., pitch), change direction left and right (e.g., yaw), and tilt laterally (e.g., roll). Other implementations may also use other techniques or angular representations for measuring head pose, such as any other type of Euler angular system. A wearable system (see, for example, wearable system 200 in Figure 2 and wearable system 400 in Figure 4) may be used to determine the user's head pose, for example, using an accelerometer, an inertial measurement unit, etc., as discussed herein. The wearable system may also move virtual objects based on eye posture (e.g., as measured by an eye-tracking camera) and head posture. For example, a user may select a virtual object by gazing at it for an extended period and then use head posture to move the selected object. The techniques for mapping the movement of user input devices described herein can also be applied to changes in the user's head, eyes, and / or body posture, i.e., the amount of movement of the virtual object is multiplied by the amount of physical movement of the user's body (e.g., eyes, head, hands, etc.).
[0163] (Example of a distance-based multiplier) As described above, the wearable system can calculate a multiplier for mapping the movement of a user input device to the movement of a target virtual object. The multiplier may be calculated based on contextual information, such as the distance between the user and the target virtual object. For example, as shown in Figure 16, the multiplier may be calculated using the distance between the position of the user's head 1210 and the position of the virtual object 1640.
[0164] Figure 17 schematically illustrates an example of the multiplier as a function of distance. As shown in Figure 17, axis 1704 indicates the magnitude of the multiplier. Axis 1702 illustrates various distances (e.g., in feet or meters) between two endpoints. The endpoints may be determined in various ways. For example, one endpoint may be the user's location (e.g., measured from the user's ARD) or the location of a user input device. The other endpoint may be the location of a target virtual object.
[0165] The distance between the user and the virtual object may change as the endpoints used to calculate the distance change. For example, the user and / or the virtual object may move around. User 1210 may activate a user input device to pull the virtual object closer. During this process, the multiplier may change based on various coefficients described herein. For example, the multiplier may decrease as the virtual object moves closer to the user, or increase as the virtual object moves further away from the user.
[0166] Curves 1710, 1720, and 1730 illustrate embodiments of the relationship between the multiplier and distance. As shown by curve 1710, the multiplier can be equal to 1 when the distance is less than the threshold 1752. Curve 1710 shows a linear relationship between distance and the multiplier between thresholds 1752 and 1754. As illustrated with reference to Figure 16, this proportional-linear relationship allows a wearable system to map small changes in the position of the user input device to large changes in the position of objects located further away (up to threshold 1754). Curve 1710 reaches its maximum value at threshold 1754, and therefore any further increase in distance will not change the magnitude of the multiplier. This can prevent very distant virtual objects from moving very large distances in response to small movements of the user input device.
[0167] Thresholding of the multiplier in curve 1710 is optional (at either or both thresholds 1752 and 1754). The wearable system may generate the multiplier without using multiple thresholds.
[0168] To enable more precise one-to-one operation, an exemplary threshold may be the user's reach. The user's reach may be an adjustable parameter that can be set by the user or the HMD (to accommodate users with different reach ranges). In various implementations, the user's reach may be in the range of approximately 10 cm to approximately 1.5 m. Referring to Figure 16, for example, if the target virtual object is within reach, as user 1210 moves the user input device 466 along the trajectory 1650a from position 1610a to position 1620a, the target virtual object may also move along the trajectory 1650a. If the target virtual object 1640 is beyond reach, the multiplier may increase. For example, in Figure 16, if the target virtual object 1640 is initially at position 1610b, as the user input device 466 moves from position 1610a to position 1620a, the target virtual object 1640 moves from position 1610b to position 1620b, thereby moving a greater distance than that of the user input device 466.
[0169] The relationship between distance and multiplier is not limited to a linear relationship. Rather, it may be determined based on various algorithms and / or coefficients. For example, as shown in Figure 17, curve 1720 may be generated using one or more power laws between distance and multiplier, for example, the multiplier being proportional to a power of the distance. The power may be 0.5, 1.5, or 2. Similarly, curve 1730 may be generated based on user preference, where the multiplier is equal to 1 when the object is within a user-adjustable threshold distance.
[0170] As an example, the movement of a virtual object (e.g., angular movement) may be represented by the variable delta_object, and the movement of a user input device may be represented by the variable delta_input. The delta is associated with a multiplier. [ka]
[0171] A sensor within the user input device or an outward-facing camera on the ARD may be used to measure delta_input. The multiplier as a function of distance d can be determined from a lookup table, a functional form (e.g., a power law), or a curve (e.g., see the example in Figure 17). In some implementations, the distance may be normalized by the distance from the user to the input device. For example, distance d may be determined as follows: [ka] In equation (2), the normalized distance is dimensionless and equal to 1 when the object is at the distance of the input device. As previously mentioned, the multiplier may be set to 1 for objects within reach (e.g., within the distance between the camera and the input device). Thus, equation (2) allows the wearable system to dynamically adjust the hand length distance based on where the user holds the input device. The exemplary power law multiplier can be as follows: [ka] In the formula, the power p is, for example, 1 (linear), 2 (quadratic), or any other integer or real number.
[0172] (Other exemplary multipliers) The multiplier can also be calculated using other factors, such as contextual information about the user's physical and / or virtual environment. For example, if a virtual object is located within a high-density cluster of objects, the wearable system may use a smaller multiplier to increase the accuracy of placing the object. Contextual information may also include the properties of the virtual object. For example, in a driving game, the wearable system may provide a larger multiplier for a good car and a smaller multiplier for an average car.
[0173] The multiplier may depend on the direction of movement. For example, in the xyz coordinate system shown in Figure 6, the multiplier for the x-axis may differ from the multiplier for the z-axis. Referring to Figure 16, instead of moving the virtual object 1640 from 1610b to 1620b, user 1210 may want to pull the virtual object 1640 closer to themselves. In this situation, the wearable system may use a multiplier smaller than the multiplier used to move the virtual object 1640 from 1610b to 1620b. Thus, the virtual object 1640 may not suddenly appear very close to the user.
[0174] A wearable system can allow the user to construct a multiplier. For example, the wearable system may give the user several options for selecting a multiplier. A user who prefers slow movement may select a multiplier with a small magnitude. The user may also provide certain coefficients and / or the importance of the coefficients that the wearable system will use to automatically determine the multiplier. For example, the user may set the weight of distance higher than the weight associated with the properties of the virtual object. Thus, distance will have a greater influence on the magnitude of the multiplier than the properties of the virtual object. Furthermore, as illustrated with reference to Figure 17, the multiplier may have one or more thresholds. One or more of the thresholds may be calculated based on the values of a set of coefficients (such as coefficients determined from contextual information). In one embodiment, one threshold may be calculated based on one set of coefficients, while another threshold may be calculated based on another set of coefficients (which may not overlap with the first set of coefficients).
[0175] (Example application of multipliers) As illustrated with reference to Figures 16 and 17, a wearable system can apply multipliers to map the movement of a user input device to the movement of a virtual object. Movement may include velocity, acceleration, or changes in position (such as rotation or movement from one location to another). For example, a wearable system may be configured to move a virtual object faster when it is located further away.
[0176] In another embodiment, the multiplier may also be used to determine the acceleration of a virtual object. When the virtual object is away from the user, it may have a large initial acceleration when the user activates a user input device and moves the virtual object. In some embodiments, the multiplier for acceleration may peak or decrease after a certain threshold. For example, to avoid moving an object too fast, the wearable system may decrease the multiplier for acceleration when the virtual object reaches the midpoint of its trajectory or when the velocity of the virtual object reaches a threshold.
[0177] In some implementations, the wearable system may use a focus indicator to indicate the current position of the user input device and / or the user's posture (e.g., head, body, eye posture). A multiplier may be applied to indicate changes in the position of the focus indicator. For example, the wearable system may display a virtual cone during cone projection (see the explanation of cone projection in Figure 12-15). When the depth of the cone is set to a distant location, the wearable system may apply a large multiplier. Thus, as the user moves around, the virtual cone may travel a long distance.
[0178] In addition, or alternatively, a wearable system can map the movement of a user input device to the movement of multiple virtual objects. For example, in a virtual game, a player can move a group of virtual soldiers together by activating a user input device. The wearable system can translate the movement of the user input device into the movement of the group of virtual soldiers by applying a multiplier to the group of virtual soldiers together, and / or by applying a multiplier to each of the virtual soldiers within the group.
[0179] (An illustrative process for moving a virtual object) Figure 18 illustrates a flowchart of an exemplary process for moving a virtual object in response to movement of a user input device. Process 1800 can be performed by the wearable system shown in Figures 2 and 4.
[0180] In block 1810, the wearable system receives a selection of a target virtual object. The virtual object may be displayed by the wearable system at a first position in 3D space. The user can select the target virtual object by activating a user input device. In addition, or alternatively, the wearable system may be configured to support the user in moving the target virtual object using various body, head, or eye poses. For example, the user may select the target virtual object by pointing their finger at it, or move the target virtual object by moving their arm.
[0181] In block 1820, the wearable system can receive indications of movement relative to a target virtual object. The wearable system may receive such indications from a user input device. The wearable system may also receive such indications from a sensor (e.g., an outward-facing imaging system 464) which can determine changes in the user's posture. The indications may be a trajectory of movement or a change in the position of part of the user's body or user input device.
[0182] In block 1830, the wearable system determines the value of a multiplier that will be applied based on the contextual information described herein. For example, the wearable system may calculate the multiplier based on the distance between the object and the user input device, and the multiplier may increase with increasing distance to the target virtual object (at least over a range of distances from the user input device; see, for example, the embodiment in equation (3)). In some embodiments, the multiplier is a non-decreasing function of the distance between the object and the user input device.
[0183] As shown in block 1840, this multiplier may be used to calculate the amount of movement relative to a target virtual object. For example, if the multiplier is calculated using the distance between the object and the user input device, the multiplier may be large for distant target virtual objects. The wearable system may use equation (3) to relate the amount of movement of the input device and the multiplier to determine the amount of movement of the target virtual object. The trajectory of the movement of the target virtual object may be calculated using other coefficients along with the multiplier. For example, the wearable system may calculate the trajectory based on the user's environment. When another object is present along the path of the target virtual object, the wearable system may be configured to move the target virtual object to avoid collisions with the other object.
[0184] In block 1850, the wearable system can display the movement of a target virtual object based on a calculated trajectory or multiplier. For example, the wearable system can calculate a second position in 3D space based on the amount of movement calculated in block 1840. The wearable system can therefore display the target virtual object at the second position. As discussed with reference to Figure 16, the wearable system may also be configured to display the movement of a visible focus indicator using a multiplier.
[0185] (Additional embodiment) In a first aspect, a method for selecting a virtual object located in a three-dimensional (3D) space, the method being performed under the control of an augmented reality (AR) system comprising computer hardware, the AR system being configured to enable user interaction with objects in the user's eye-moving field of view (FOR), the FOR comprising a portion of the user's surrounding environment perceptible to the user via the AR system, the method comprising: determining a group of objects in the user's FOR; determining the user's posture; initiating a cone projection onto the group of objects, the cone projection comprising projecting a virtual cone with an opening in a direction at least partially based on the user's posture; analyzing contextual information associated with subgroups of objects within the group of objects; updating the opening with respect to the cone projection event at least partially based on the contextual information; and rendering a visual representation of the cone projection.
[0186] The second aspect is the method according to aspect 1, wherein the subgroup of objects is within the user's field of view (FOV), and the FOV includes a portion of FOR that is perceptible to the user via the AR system at a given time.
[0187] In the third aspect, the method as described in aspect 1 or 2, the contextual information includes one or more of the types, layouts, locations, sizes, or densities of objects within a subgroup of objects.
[0188] In the fourth aspect, contextual information further includes user preferences, as described in aspect 3.
[0189] The method according to any one of the five aspects of Aspects 1-4, further comprising detecting a collision between a cone and one or more objects.
[0190] In the sixth aspect, the method of aspect 5, wherein one or more objects include interactable objects.
[0191] In the seventh aspect, the method according to aspect 6, further comprising performing an action on the interactable object in response to detecting a collision with the interactable object.
[0192] The eighth aspect is the method of aspect 7, wherein the action includes one or more of selecting an interactable object, moving an interactable object, or opening a menu associated with an interactable object.
[0193] The method according to aspect 5 or 6, further comprising, in a ninth aspect, applying an occlusion ambiguity technique to one or more objects that collide with the cone.
[0194] The method according to any one of the sides of 1-9, further comprising updating the opening of the cone based at least partially on a change in the user's posture.
[0195] On the eleventh side, the cone has a certain shape, as described in any one of the sides 1-10.
[0196] The method according to the 12th aspect, wherein the shape includes one or more of a geometric cone projection, a cuboid, a polyhedron, a pyramid, or a frustum of a cone.
[0197] In the 13th side, the cone has a central ray, as described in any one of sides 1-12.
[0198] In the 14th aspect, the central ray is determined at least partially based on the user's posture, as described in aspect 13.
[0199] On the 15th side, the aperture is the method described on side 13 or 14, which crosses the central ray.
[0200] The method described in any one of the 16th aspect, further comprising resolving ambiguity of an object colliding with a cone.
[0201] Aspect 17 is an augmented reality system configured to implement the method described in any one of Aspects 1-16.
[0202] In its 18th aspect, a method for transforming a virtual object located in a three-dimensional (3D) space, the method being performed under the control of an augmented reality (AR) system comprising computer hardware and a user input device, the AR system being configured to enable user interaction with a virtual object in the user's eye-tracking field of view (FOR), the FOR including a portion of the user's surrounding environment perceptible to the user via the AR system, the virtual object being presented via the AR system for display to the user, the method comprising: determining a group of virtual objects in the user's FOR; receiving a selection of a target virtual object within the group of virtual objects in the user's FOR; calculating the distance to the target virtual object; determining a multiplier, at least partially based on the distance to the target virtual object; receiving a first movement of the user input device; and calculating a second movement of the target virtual object, the second movement being at least partially based on the first movement and the multiplier; and moving the target virtual object by an amount at least partially based on the second movement.
[0203] Aspect 19 is the method of aspect 18, wherein calculating the distance to a virtual object includes calculating the distance between the virtual object and a user input device, the distance between the virtual object and a sensor on the AR system, or the distance between the user input device and a sensor on the AR system.
[0204] In the 20th aspect, the method according to aspect 18, the second movement is equal to the first movement multiplied by the multiplier.
[0205] In the 21st aspect, the multiplier increases with increasing distance over a first distance range, as described in aspect 18.
[0206] The method according to aspect 21, wherein in aspect 22, the multiplier increases linearly with increasing distance over a first range.
[0207] In the 23rd aspect, the multiplier increases as a power of the distance over a first range, as described in aspect 21.
[0208] In the 24th aspect, the multiplier is equal to the first threshold when the distance is less than the first distance, as described in aspect 18.
[0209] In the 25th aspect, the first distance is equal to the range of the user's hand, as described in aspect 24.
[0210] In the 26th aspect, the first threshold is equal to 1, as described in aspect 24.
[0211] In the 27th aspect, the method described in any one of the sections of aspects 18-26, wherein the first movement or the second movement each includes a first speed or a second speed.
[0212] In aspect 28, the method described in any one of aspects 18-26, wherein the first movement and the second movement each include a first acceleration and a second acceleration, respectively.
[0213] In aspect 29, the AR system is the method described in any one of aspects 18-28, comprising a head-mounted display.
[0214] In Aspect 30, the target virtual object is interactable, as described in any one of Aspects 18-29.
[0215] In its 31st aspect, a method for moving a virtual object located in a three-dimensional (3D) space, the method being performed under the control of an augmented reality (AR) system comprising computer hardware and a user input device, the AR system being configured to present the virtual object in the 3D space for display to a user, the method comprising: receiving a selection of a target virtual object to be displayed to the user at a first position in the 3D space; receiving an indication of movement relating to the target virtual object; determining a multiplier to be applied to the movement of the target virtual object; calculating a displacement amount relating to the target virtual object, the displacement amount being at least partially based on the indication of movement and the multiplier; and displaying the target virtual object to the user at a second position, the second position being at least partially based on the first position and the displacement amount.
[0216] Aspect 32 is the method according to aspect 31, wherein determining a multiplier to be applied to the movement of a target virtual object includes calculating the distance to the target virtual object.
[0217] In the 33rd aspect, the method according to aspect 32, wherein the distance is between the target virtual object and the user input device, between the target virtual object and a sensor on the AR system, or between the user input device and a sensor on the AR system.
[0218] In the 34th aspect, the multiplier increases as the distance increases, as described in the method for aspect 32.
[0219] In Aspect 35, the multiplier is the method described in any one of Aspects 31-34, which is at least partially based on user preferences.
[0220] In Aspect 36, the motion is the method described in any one of Aspects 31–35, including a change of position, velocity, or acceleration, or more.
[0221] In Aspect 37, the target virtual object is the method described in any one of Aspects 31-36, which includes a group of virtual objects.
[0222] In Aspect 38, the target virtual object is interactable, as described in any one of Aspects 31-37.
[0223] Aspect 39, the method described in any one of Aspects 31-38, wherein receiving a movement indication includes receiving a movement indication from a user input device.
[0224] Aspect 40 is the method described in any one of Aspects 31-38, wherein receiving an indication of movement includes receiving an indication of a change in the user's posture.
[0225] In the 41st aspect, the method according to aspect 40, wherein the user's posture includes one or more of the head posture, eye posture, or body posture.
[0226] In the 42nd aspect, an augmented reality (AR) system for transforming virtual objects located in a three-dimensional (3D) space, the system comprising a system display system, a user input device, and a computer processor, the computer processor being configured to communicate with the display system and the user input device and to determine a group of virtual objects in the user's FOR, receive a selection of a target virtual object in the group of virtual objects in the user's FOR, calculate the distance to the target virtual object, determine a multiplier at least partially based on the distance to the target virtual object, receive a first movement of the user input device, calculate a second movement of the target virtual object, wherein the second movement is at least partially based on the first movement and the multiplier, and move the target virtual object by an amount at least partially based on the second movement.
[0227] In aspect 43, the system according to aspect 42 includes calculating the distance to a target virtual object, which includes calculating the distance between the target virtual object and a user input device, the distance between the virtual object and a sensor on the AR system, or the distance between the user input device and a sensor on the AR system.
[0228] In the 44th aspect, the system described in aspect 42 is such that the second movement is equal to the first movement multiplied by the multiplier.
[0229] In aspect 45, the multiplier increases with increasing distance over a first distance range, as described in aspect 42.
[0230] In aspect 46, the multiplier increases linearly with increasing distance over a first range, as described in aspect 45.
[0231] In the 47th aspect, the multiplier increases as a power of the distance over a first range, as described in the system of the 45th aspect.
[0232] In aspect 48, the system as described in aspect 42, wherein the multiplier is equal to a first threshold when the distance is less than a first distance.
[0233] In the 49th aspect, the first distance is equal to the user's reach, as described in the system in the 48th aspect.
[0234] In aspect 50, the first threshold is equal to 1, as described in aspect 48 of the system.
[0235] In aspect 51, the system according to any one of aspects 42-50, wherein the first movement or the second movement each includes a first speed or a second speed.
[0236] In aspect 52, the system described in any one of aspects 42-50, wherein the first movement and the second movement each include a first acceleration and a second acceleration, respectively.
[0237] In Aspect 53, the AR system is the system described in any one of Aspects 42-52, comprising a head-mounted display.
[0238] In Aspect 54, the target virtual object is an interactable system as described in any one of Aspects 42-53.
[0239] In the 55th aspect, an augmented reality (AR) system for moving a virtual object located in a three-dimensional (3D) space, the system comprising: a system display system; a user input device; and a computer processor, the computer processor being configured to communicate with the display system and the user input device and to receive a selection of a target virtual object to be displayed to the user at a first position in the 3D space; to receive an indication of movement relating to the target virtual object; to determine a multiplier to be applied to the movement of the target virtual object; to calculate a movement amount relating to the target virtual object, the movement amount being at least partially based on the indication of movement and the multiplier; and to display the target virtual object to the user at a second position, the second position being at least partially based on the first position and the movement amount.
[0240] In aspect 56, the system described in aspect 55 includes determining a multiplier to be applied to the movement of a target virtual object, which involves calculating the distance to the target virtual object.
[0241] In aspect 57, the system according to aspect 56, wherein the distance is between a virtual object and a user input device, between a virtual object and a sensor on the AR system, or between a user input device and a sensor on the AR system.
[0242] In aspect 58, the multiplier increases as the distance increases, as described in aspect 56.
[0243] In the 59th aspect, the multiplier is the system according to any one of aspects 55 - 58, which is at least partially based on user preference.
[0244] In the 60th aspect, the movement includes one or more of a change in position, speed, or acceleration, and is the system according to any one of aspects 55 - 59.
[0245] In the 61st aspect, the target virtual object includes a group of virtual objects, and is the system according to any one of aspects 55 - 60.
[0246] In the 62nd aspect, the target virtual object is interactable, and is the system according to any one of aspects 55 - 61.
[0247] In the 63rd aspect, receiving an indication of movement includes receiving an indication of movement from a user input device, and is the system according to any one of aspects 55 - 62.
[0248] In the 64th aspect, receiving an indication of movement includes receiving an indication of a change in the user's posture, and is the system according to any one of aspects 55 - 63.
[0249] In the 65th aspect, the user's posture includes one or more of a head posture, an eye posture, or a body posture, and is the system according to aspect 64.
[0250] In its 66th aspect, a system for interacting with objects for a wearable device, the system being a display system for a wearable device configured to present a three-dimensional (3D) view to a user and enable user interaction with objects in the user's eye-moving field of view (FOR), wherein the FOR includes a portion of the user's surrounding environment that is perceptible to the user via the display system, and comprising: a sensor configured to acquire data associated with the user's posture, and a hardware processor communicating with the sensor and the display system, wherein the hardware processor is programmed to determine the user's posture based on the data acquired by the sensor, to initiate a cone projection onto a group of objects in the FOR, the cone projection comprising projecting a virtual cone with an opening in a direction at least partially based on the user's posture, to analyze contextual information associated with the user's environment, to update the opening of the virtual cone at least partially based on the contextual information, and to render a visual representation of the virtual cone for the cone projection.
[0251] In Aspect 67, the context information includes at least one of the types, layouts, locations, sizes, or densities of subgroups of objects within the user's field of view (FOV), the FOV including a portion of FOR that is perceptible to the user at a given time via the display system, as described in Aspect 66.
[0252] In aspect 68, the system according to aspect 67, wherein the density of subgroups of objects within the user's FOV is calculated by at least one of the following: calculating the number of objects within a subgroup of objects; calculating the percentage of the FOV covered by the subgroup of objects; or calculating a contour map of the objects within a subgroup of objects.
[0253] In aspect 69, the system according to any one of aspects 66-68, wherein the hardware processor is further programmed to detect collisions between a virtual cone and one or more objects in a group of objects within the FOR, and in response to collision detection, the hardware processor is further programmed to present a focus indicator to one or more objects.
[0254] In aspect 70, the system described in aspect 69, the hardware processor is programmed to apply an occlusion deambiguation technique to one or more objects that collide with a virtual cone and to identify the occluded object.
[0255] In the 71st aspect, the cone comprises a central ray, and the aperture intersects the central ray, as described in any one of aspects 66-70.
[0256] In the 72nd aspect, the system according to any one of aspects 66-71, wherein the virtual cone comprises a proximal end, which is anchored to at least one of the following locations: between the user's eyes, on part of the user's arm, on a user input device, or any other location in the user's environment.
[0257] In Aspect 73, the system described in any one of Aspects 66-72, wherein the hardware processor is further programmed to receive an indication from a user input device to anchor the depth of a virtual cone to a depth plane, and the cone projection is performed on a group of objects in the depth plane.
[0258] In aspect 74, the system according to any one of aspects 66-73 comprises at least one of an inertial measurement unit or an outward-facing imaging system.
[0259] In aspect 75, the virtual cone is a system according to any one of aspects 66-74, comprising at least one of a geometric cone projection, a cuboid, a polyhedron, a pyramid, or a frustum of a cone.
[0260] In its 76th aspect, a method for interacting with an object for a wearable device, the method comprising: receiving a selection of a target virtual object to be displayed to a user at a first position in three-dimensional (3D) space; receiving an indication of movement relating to the target virtual object; analyzing contextual information associated with the target virtual object; calculating a multiplier to be applied to the movement of the target virtual object, at least in part based on the contextual information; calculating a displacement relating to the target virtual object, the displacement being at least in part based on the indication of movement and the multiplier; and displaying the target virtual object to the user at a second position, the second position being at least in part based on the first position and the displacement.
[0261] In Aspect 77, contextual information includes the distance from the user to the target virtual object, as described in Aspect 76.
[0262] In aspect 78, the multiplier increases proportionally with increasing distance, as described in aspect 77.
[0263] In Aspect 79, the motion is a method described in any one of Aspects 76–78, including a change of position, velocity, or acceleration, or more.
[0264] In Aspect 80, the method of any one of Aspects 76-79, wherein the indication of movement includes at least one of the activation of a user input device associated with a wearable device or a change in the user's posture.
[0265] In the 81st aspect, the posture includes one or more than one of the head posture, eye posture, or body posture, and is the method described in aspect 80.
[0266] In the 82nd aspect, a system for interacting with an object for a wearable device, the system comprising: a display system of a wearable device configured to present a three-dimensional (3D) view to a user, the 3D view comprising a target virtual object; and a hardware processor communicating with the display system, the hardware processor being configured to: receive an indication of movement regarding the target virtual object; analyze context information associated with the target virtual object; calculate a multiplier to be applied to the movement of the target virtual object based at least in part on the context information; calculate an amount of movement regarding the target virtual object, the amount of movement being based at least in part on the indication of movement and the multiplier; and cause the display system to display the target virtual object at a second position, the second position being based at least in part on the first position and the amount of movement.
[0267] In the 83rd aspect, the indication of movement of the target virtual object includes a change in the posture of the user of the wearable device or an input received from a user input device associated with the wearable device, and is the system described in aspect 82.
[0268] In the 84th aspect, the context information includes the distance from the user to the target virtual object, and is the system described in any one of aspects 82-83.
[0269] In the 85th aspect, the multiplier is equal to 1 when the distance is less than a threshold distance, and the threshold distance is equal to the reach of the user's hand, and is the system described in aspect 84.
[0270] In aspect 86, the multiplier increases proportionally with increasing distance, as described in any one of aspects 84-85.
[0271] In Aspect 87, the movement includes one or more of a change in position, velocity, or acceleration, as described in any one of Aspects 82-86.
[0272] (Conclusion) The processes, methods, and algorithms described herein and / or depicted in accompanying diagrams may be embodied in code modules executed by one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to perform specific computer instructions, thereby being fully or partially automated. For example, a computing system may include a general-purpose computer (e.g., a server) or a dedicated computer, dedicated circuit, etc., programmed with specific computer instructions. The code modules may be written in a programming language that can be compiled and linked into an executable program, installed in a dynamic link library, or interpreted. In some implementations, specific operations and methods may be performed by circuits specific to a given function.
[0273] Furthermore, functional implementations of the present disclosure are sufficiently mathematical, computational, or technically complex that application-specific hardware (utilizing appropriate specialized executable instructions) or one or more physical computing devices may need to implement the functionality, for example, due to the volume or complexity of the computations involved or to provide results substantially in real time. For example, video may contain many frames, each frame may have millions of pixels, and specifically programmed computer hardware needs to process the video data to provide a desired image processing task or application in a commercially reasonable amount of time.
[0274] Code modules or any type of data may be stored on any type of non-transient computer-readable medium, such as physical computer storage devices, including hard drives, solid-state memory, random-access memory (RAM), read-only memory (ROM), optical discs, volatile or non-volatile storage devices, combinations thereof, and / or equivalents. The method and modules (or data) may also be transmitted as data signals generated on various computer-readable transmission media, including wireless-based and wired / cable-based media (e.g., as part of a carrier wave or other analog or digital propagation signal), and may take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed process or process steps may be stored persistently or otherwise in any type of non-transient tangible computer storage device, or communicated via computer-readable transmission media.
[0275] Any process, block, state, step, or functionality in the flowcharts described herein and / or depicted in the accompanying diagrams should be understood as potentially representing a code module, segment, or portion of code containing one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in the process. Various processes, blocks, states, steps, or functionality may be combined, rearranged, added, deleted, modified, or otherwise changed from the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may implement some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the blocks, steps, or states associated therewith may be implemented in other appropriate sequences, for example, sequentially, in parallel, or in some other manner. Tasks or events may be added to or removed from the illustrative embodiments disclosed. Furthermore, the isolation of various system components in the implementations described herein is for illustrative purposes only and should not be understood as requiring such isolation in all implementations. It should be understood that the program components, methods, and systems described can generally be integrated together in a single computer product or packaged across multiple computer products. Many implementation variations are possible.
[0276] This process, method, and system may be implemented in a network (or distributed) computing environment. Network environments include enterprise-wide computer networks, intranets, local area networks (LANs), wide area networks (WANs), personal area networks (PANs), cloud computing networks, crowdsourced computing networks, the Internet, and the World Wide Web. The network may be a wired or wireless network or any other type of communication network.
[0277] Each system and method of this disclosure has several innovative aspects, none of which alone contribute to or are required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of each other or in various combinations. All possible and secondary combinations are intended to fall within the scope of this disclosure. Various modifications of the implementations described herein may be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations without departing from the spirit or scope of this disclosure. Accordingly, the claims are not intended to be limited to the implementations shown herein, but should be given the broadest scope consistent with the disclosure, principles, and novel features disclosed herein.
[0278] Some features described herein in the context of separate implementations may also be implemented in combinations within a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately in multiple implementations or in any preferred secondary combination. Furthermore, features described above as acting in a combination and further, which may be initially claimed as such, may in some cases be removed from the combination, and the claimed combination may be subject to secondary combinations or variations of secondary combinations. No single feature or group of features is required or essential in any embodiment.
[0279] In particular, conditional statements used herein, such as “can,” “could,” “might,” “may,” “e.g.,” and equivalents, are generally intended to convey that one embodiment includes certain features, elements, and / or steps, while other embodiments do not, unless otherwise specifically stated or understood in the context in which they are used. Therefore, such conditional statements are generally not intended to suggest that features, elements, and / or steps are required in any way for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps are included or should be implemented in any particular embodiment, with or without input or prompting by the author. The terms “equipment,” “includes,” “have,” and equivalents are synonyms and are used in a non-restrictive manner to encompass additional elements, features, actions, behaviors, etc. Furthermore, the term "or" is used in its inclusive sense (and not in its exclusive sense), and therefore, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. In addition, the articles "a," "an," and "the," as used in this application and the attached claims, should be interpreted as meaning "one or more" or "at least one," unless otherwise specified.
[0280] As used herein, the phrase “at least one of” a list of items refers to any combination of those items that includes a single element. In one embodiment, “at least one of A, B, or C” is intended to encompass A, B, C, A and B, A and C, B and C, and A, B, and C. Connecting phrases such as “at least one of X, Y, and Z” are generally understood differently in contexts such as those used to convey that an item, term, etc., may be at least one of X, Y, or Z, unless otherwise specifically stated. Thus, such connecting phrases are generally not intended to suggest that one embodiment requires the presence of at least one of X, at least one of Y, and at least one of Z, respectively.
[0281] Similarly, while actions may be depicted in a diagram in a specific order, it should be recognized that this does not mean that such actions must be performed in a specific order or sequential order, or that all illustrated actions must be performed, in order to achieve the desired result. Furthermore, a diagram may graphically depict one or more exemplary processes in the form of a flowchart. However, other actions not depicted may also be incorporated into the graphically illustrated exemplary methods and processes. For example, one or more additional actions may be performed before, after, simultaneously with, or in between any of the illustrated actions. In addition, actions may be rearranged or rearranged in other implementations. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. In addition, other implementations are also within the scope of the following claims. In some cases, the actions enumerated in the claims may be performed in a different order, and the desired result can still be achieved.
Claims
1. A system, wherein the system is A display system for a wearable device, wherein the wearable device is configured to provide a three-dimensional (3D) view to the user and to enable user interaction with objects within the user's field of view (FOV), the FOV including a portion of the environment around the user that is perceptible to the user through the display system. A hardware processor that communicates with the aforementioned display system Equipped with, The aforementioned hardware processor is Starting the cone projection of the virtual cone, Determining the context information within the FOV of the user, Anchoring the proximal end of the virtual cone to an anchoring location associated with a physical object specified by the user via user input, wherein the user is enabled to anchor the proximal end to a pointing device, the user is enabled to anchor the proximal end to the user, and movement of the anchoring location associated with the physical object specified by the user causes a corresponding movement of the virtual cone, thereby causing the proximal end of the virtual cone to remain at the anchoring location associated with the physical object, the virtual cone further including a distal end or depth, and the hardware processor is programmed to anchor the distal end and / or depth of the virtual cone to the physical object based on at least one of user input, body gestures, body posture, gaze direction, or voice commands, Determining one or more objects within the virtual cone anchored to the physical object, Performing an action on one or more of the objects within the virtual cone anchored to the physical object A system that is programmed to do so.
2. The system according to claim 1, wherein the hardware processor is further programmed to determine the context information relating to the user's environment, the context information comprising at least one of the types, layouts, locations, sizes, distances, or densities of objects in the user's FOV.
3. In order to determine the density of the object in the user's FOV, the hardware processor: Calculate the number of objects within the aforementioned FOV. Calculating the portion of the FOV covered by the aforementioned object, or Calculating a contour map for the aforementioned object. The system according to claim 2, which is programmed to perform the following:
4. The system according to claim 2, wherein the context information includes at least one of the user's preferences, physical conditions associated with the user, or information associated with the environment.
5. The system according to claim 1, wherein the opening is dynamically adjustable, and in order to update the dynamically adjustable opening, the hardware processor is programmed to select the opening between a minimum size and a maximum size.
6. The system according to claim 5, wherein the minimum size is zero.
7. The system according to claim 1, wherein the action includes rendering focus indicators associated with one or more objects.
8. The hardware processor is further programmed to detect collisions between the virtual cone and one or more objects in the environment, The system according to claim 5, wherein, in order to detect the collision between the virtual cone and the one or more objects, the hardware processor is programmed to determine that the one or more objects intersect the virtual surface of the virtual cone, or that the one or more objects are within the opening of the virtual cone.
9. The system according to claim 1, wherein the one or more objects include a plurality of collision objects, and the hardware processor is programmed to identify occluded objects or determine the depth order or position between occluded objects by applying an occluding ambiguity technique to the plurality of collision objects.
10. The system according to claim 1, wherein the one or more objects include a plurality of collision objects, and the hardware processor is programmed to select one or more of the plurality of collision objects by presenting a user interface element.
11. The action with one or more objects is: Selection of one or more objects, The movement of one or more of the aforementioned objects, Display of menus or toolbars associated with one or more of the aforementioned objects, or Game actions on virtual avatars within the game The system according to claim 1, comprising one or more of the following.
12. The system according to claim 1, wherein the virtual cone includes a central ray, the opening of the virtual cone intersects the central ray, and the direction of the central ray is based on the user's posture.
13. The hardware processor is Determining the density of objects within the user's FOV, The opening of the virtual cone is updated based on the density of the objects within the user's FOV. The system according to claim 1, configured to perform the following:
14. The system according to claim 1, wherein the hardware processor is further programmed to detect collisions between the virtual cone and one or more objects in the environment, the virtual cone includes a distal end, and in order to detect the collisions, the hardware processor is programmed to scan for collisions with the one or more objects within the distal end of the virtual cone.
15. The system according to claim 1, wherein the virtual cone further includes a distal end or depth, and the hardware processor is programmed to anchor the distal end or depth of the virtual cone based in part on the context information.
16. The system further comprises a posture sensor configured to obtain posture data associated with the user's posture, The system according to claim 1, wherein the hardware processor is programmed to translate the virtual cone based on the posture data associated with the user's posture.
17. The system according to claim 1, wherein the hardware processor is further programmed to resize the opening of the virtual cone based at least in part on the context information in the user's FOV.
18. The system according to claim 1, wherein, in order to initiate the cone projection, the hardware processor is programmed to extend the distal end of the virtual cone until the distal end reaches a termination threshold.
19. The system according to claim 18, wherein the termination threshold includes a threshold distance or the boundary of the environment.
20. The system according to claim 1, wherein the hardware processor is further programmed to render a visual representation of at least a portion of the virtual cone to the user via the display system.
Citation Information
Patent Citations
Information processing system, control method of the same, and program, and information processing apparatus, control method of the same, and program
JP2016035742A
Handheld synthetic vision device
US20090293012A1
Tools for Use within a Three Dimensional Scene
US20120013613A1
Method and system for providing consistency between a virtual representation and corresponding physical spaces
US20130219302A1
System and Method for Displaying Data Having Spatial Coordinates
US20130300740A1