Rendering of location-specific virtual content at any location

The mixed reality system addresses the challenges of AR and MR by anchoring virtual content to a scene node, allowing flexible saving and reopening across locations, enhancing user experience and computational efficiency.

JP7711266B2Active Publication Date: 2025-07-22MAGIC LEAP INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024093652
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-10-05
Filing Date
2024-06-10
Publication Date
2025-07-22
Estimated Expiration
2039-10-04

AI Technical Summary

Technical Problem

Existing augmented reality (AR) and mixed reality (MR) technologies struggle to provide a comfortable, natural, and rich presentation of virtual image elements in relation to the physical world, often leading to issues like unstable imaging, eye strain, and limited flexibility in saving and reopening virtual content across different locations.

Method used

A mixed reality system that allows users to select, save, and reopen virtual content by anchoring it to a scene anchor node, enabling flexible rendering and interaction with virtual objects across various environments using a unified user interface, regardless of location changes.

Benefits of technology

Enables seamless saving and reopening of virtual content at any location, improving user experience by reducing computational load, battery usage, and providing a realistic perception of depth through accurate alignment of virtual content with the physical world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711266000001
    Figure 0007711266000001
  • Figure 0007711266000002
    Figure 0007711266000002
  • Figure 0007711266000003
    Figure 0007711266000003
Patent Text Reader

Abstract

To provide rendering location-specific virtual content in any appropriate location.SOLUTION: Augmented reality systems and methods are provided for creating, saving, and rendering designs comprising multiple items of virtual content in a three-dimensional (3D) environment of a user. The designs may be saved as a scene, which is built by a user from pre-built sub-components, built components, and / or previously saved scenes. Location information, expressed as a saved scene anchor and position relative to the saved scene anchor for each item of virtual content, may also be saved. Upon opening the scene, the saved scene anchor node may be correlated to a location within the mixed reality environment of the user for whom the scene is opened. The virtual items of the scene may be positioned with the same relationship to that location as they have to the saved scene anchor node. That location may be selected automatically and / or by user input.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit under 35 U.S.C. § 119 of U.S. Provisional Patent Application No. 62 / 742,061, filed on October 5, 2018, and titled "RENDERING LOCATION SPECIFIC VIRTUAL CONTENT IN ANY LOCATION", which is incorporated herein by reference in its entirety.

[0002] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly to automatically repositioning virtual objects within a three - dimensional (3D) space.

Background Art

[0003] Modern computing and display technologies have facilitated the development of systems for so - called "virtual reality", "augmented reality", or "mixed reality" experiences, in which digitally reproduced images or portions thereof are presented to a user in a manner that appears or can be perceived as being real. Virtual reality, i.e., the "VR" scenario, typically involves the presentation of digital or virtual image information without transparency to other actual real - world visual inputs. Augmented reality, i.e., the "AR" scenario, typically involves the presentation of digital or virtual image information as an augmentation to the visualization of the actual world around the user. Mixed reality, i.e., "MR", is related to the fusion of the real world and the virtual world to produce a new environment in which physical and virtual objects co - exist and interact in real time. In conclusion, the human visual perception system is very complex, and it is difficult to produce VR, AR, or MR technologies that facilitate a comfortable, natural, and rich presentation of virtual image elements among other virtual or real - world image elements. The systems and methods disclosed herein address various issues related to VR, AR, and MR technologies.

Summary of the Invention

Means for Solving the Problem

[0004] Various embodiments of an augmented reality system for rendering virtual content at any location are described.

[0005] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, the drawings, and the claims. Neither this summary nor the following detailed description purports to define or limit the scope of the subject matter of the invention. The present invention provides, for example, the following. (Item 1) A method of operating a mixed reality system of a type that maintains an environment for a user comprising virtual content configured to be rendered to appear to the user in relation to the physical world, the method using at least one processor to select virtual content within the environment; store the stored scene data in a non-volatile computer storage medium, the stored scene data comprising data representing the selected virtual content; and position information indicating the position of the selected virtual content relative to the stored scene anchor node and including. (Item 2) The mixed reality system identifies one or more coordinate frames based on objects within the physical world, and the method further includes installing the stored scene anchor node in one of the one or more identified coordinate frames. A method of operating the mixed reality system according to item 1. (Item 3) Selecting the virtual content within the environment includes receiving user input to select a camera icon having a display area; Rendering a representation of virtual content within a part of the environment in the display area; Generating saved scene data representing at least virtual content within a part of the environment; A method of operating the composite reality system according to item 1, comprising: (Item 4) Changing the position or orientation of the camera icon within the environment based on user input; Dynamically updating the virtual content rendered within the display area based on the position and orientation of the camera icon; A method of operating the composite reality system according to item 3, further comprising: (Item 5) The method further comprises capturing an image representing the virtual content rendered within the display based on user input; Selecting the virtual content includes selecting the virtual content represented in the captured image; A method of operating the composite reality system according to item 4; (Item 6) Further comprising generating an icon associated with the saved scene data in a menu of saved scenes available for opening, the icon comprising the captured image, a method of operating the composite reality system according to item 5; (Item 7) Generating the saved scene data representing at least virtual content within a part of the environment is triggered based on user input, a method of operating the composite reality system according to item 3; (Item 8) Generating the saved scene data representing at least virtual content within a part of the environment is automatically triggered based on detecting that the scene is framed within the display area, a method of operating the composite reality system according to item 3; (Item 9) A method of operating a composite reality system according to item 1, wherein selecting virtual content within the environment includes selecting a virtual object within the user's field of view. (Item 10) A method of operating a composite reality system according to item 1, wherein selecting virtual content within the environment includes selecting a virtual object within the user's oculomotor field of view. (Item 11) The method further includes creating a scene within the environment by receiving user input indicating a plurality of pre-constructed sub-components for inclusion within the environment, wherein selecting virtual content within the environment includes receiving user input indicating at least a portion of the scene, A method of operating a composite reality system according to item 1. (Item 12) Selecting virtual content within the environment includes receiving user input indicating at least a portion of the environment and calculating a virtual representation of one or more physical objects within the environment and includes a method of operating a composite reality system according to item 1. (Item 13) A composite reality system configured to maintain an environment for a user with virtual content and render the content on a display device so as to appear to the user in relation to the physical world, the system comprising: at least one processor; a non-volatile computer storage medium; a non-transitory computer-readable medium encoded with computer-executable instructions, which when executed by the at least one processor, select virtual content within the environment and store stored scene data in the non-volatile computer storage medium, the stored scene data Data representing the selected virtual content, Position information indicating the position of the selected virtual content relative to the saved scene anchor node Comprising, A non-transitory computer-readable medium that performs Comprising, a system. (Item 14) The composite reality system further comprises one or more sensors configured to obtain information about the physical world, The computer-executable instructions further Based on the obtained information, identifying one or more coordinate frames, Placing the saved scene anchor node in one of the one or more identified coordinate frames The composite reality system according to item 13, configured to perform (Item 15) Selecting virtual content within the environment Receiving user input to select a camera icon having a display area, Rendering a representation of virtual content within a portion of the environment within the display area, Generating saved scene data representing at least the virtual content within a portion of the environment The composite reality system according to item 13, including (Item 16) The computer-executable instructions further Based on user input, changing the position or orientation of the camera icon within the environment, Based on the position and orientation of the camera icon, dynamically updating the virtual content rendered within the display area The composite reality system according to item 13, configured to perform (Item 17) The computer-executable instructions are further configured to capture an image representing virtual content rendered within the display based on user input, Selecting virtual content includes selecting the virtual content represented in the captured image, The mixed reality system according to item 16. (Item 18) The computer-executable instructions are further configured to generate an icon associated with the stored scene data within a menu of stored scenes available for opening, the icon comprising the captured image, the mixed reality system according to item 17. (Item 19) Generating stored scene data representing at least virtual content within a portion of the environment is triggered based on user input, the mixed reality system according to item 15. (Item 20) Generating stored scene data representing at least virtual content within a portion of the environment is automatically triggered based on detecting that the scene is framed within the display area, the mixed reality system according to item 15. (Item 21) A mixed reality system, A display configured to render virtual content to a user viewing the physical world, At least one processor, A computer memory storing computer-executable instructions, the computer-executable instructions, when executed by the at least one processor, Receiving an input from the user to select a stored scene from a stored scene library, each stored scene comprising virtual content, Determining a location of the virtual content relative to an object in the physical world within the user's field of view, Controlling the display to render the virtual content at the determined location And a computer memory configured to perform A mixed reality system comprising. (Item 22) The system includes a network interface, The computer-executable instructions are configured to read the virtual content from a remote server via the network interface. The composite reality system according to item 21. (Item 23) The system includes a network interface, The computer-executable instructions are configured to control the display and render a menu including a plurality of icons representing stored scenes in the stored scene library. Receiving an input from the user to select the stored scene includes a user selection of an icon among the plurality of icons. The composite reality system according to item 21. (Item 24) The computer-executable instructions are further configured to render a virtual user interface when executed by the at least one processor. The computer-executable instructions configured to receive an input from the user are configured to receive user input via the virtual user interface. The composite reality system according to item 21. (Item 25) The computer-executable instructions are further configured to render a virtual user interface including a menu of a plurality of icons representing stored scenes in the library when executed by the at least one processor. Receiving an input from the user to select a stored scene in the stored scene library includes receiving user input via the virtual interface, and the user input selects and moves an icon within the menu. The composite reality system according to item 21. (Item 26) The computer-executable instructions are configured to move a visual anchor node and determine a location of the virtual content relative to an object in the physical world within a user's field of view based on a user input that displays a preview representation of the scene at a location indicated by the position of the visual anchor node. The computer-executable instructions are further configured to replace the preview representation of the scene with a virtual object corresponding to an object within the preview representation but having a different appearance and nature in response to a user input. The mixed reality system according to item 21.

Brief Description of the Drawings

[0006]

Figure 1

[0007]

Figure 2

[0008]

Figure 3

[0009]

Figure 4

[0010]

Figure 5

[0011]

Figure 6

[0012]

Figure 7

[0013]

Figure 8

[0014]

Figure 9

[0015]

Figure 10

[0016]

Figure 11

[0017]

Figure 12

[0018]

Figure 13A

[0019]

Figure 13B

[0020]

Figure 14

[0021]

Figure 15

[0022]

Figure 16

[0023]

Figure 17

[0024]

Figure 18

[0025]

Figure 19

[0026]

Figure 20

[0027]

Figure 21A

[0028]

Figure 21B

[0029]

Figure 21C

[0030] Throughout the drawings, reference numerals may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate the exemplary embodiments described herein and are not intended to limit the scope of the present disclosure.

DETAILED DESCRIPTION

[0031] (Overview) In an AR / MR environment, a user may desire to create, construct, and / or design new virtual objects. The user can be an engineer who needs to create a prototype for a work project, or the user can be a high school student in their teens who enjoys building and creating through play, such as someone who can build using physical elements in complex LEGO (registered trademark) kits and puzzles. In some situations, the user may need to build virtual objects with complex structures that can take some time to build, for example, over several days, months, or even years. In some embodiments, virtual objects with complex structures may comprise repeating components that are arranged or used in different ways. As a result, there are situations where the user may sometimes desire to build components from pre-built sub-components, save one or more of the built components as separate saved scenes, and then use one or more of the pre-saved scenes to build various designs.

[0032] For example, a user may desire to create a gardening design using an AR / MR wearable device. The wearable device may download an application that stores or provides access to pre-built sub-components such as various types of trees (e.g., pine trees, oak trees), flowers (e.g., gladiolus, sunflowers, etc.), and various other plants (e.g., shrubs, vines, etc.). A gardening designer may recognize through experience that a certain combination of plants is favorable. The gardening designer may utilize an application on the AR / MR wearable device to create constructed components that incorporate these known favorable combinations of plants. As an example, a constructed component may include, for example, a raspberry planter, a tulip planter, and clover, or any other suitable mixed planting arrangement. After saving one or more scenes comprising one or more pre-processed sub-components combined to form one or more constructed components, the gardening designer may then desire to create a complete gardening design for their home. The gardening design may comprise one or more saved scenes, and / or one or more constructed components, and / or one or more pre-built sub-components. The gardening design can be designed more quickly and easily by utilizing the saved scenes than if the gardening designer had started with only pre-built sub-components.

[0033] The saved scenes can also enable more flexible design options. For example, some applications may only allow the user to select from pre-built sub-components that are more or less complex than what the designer requires. The ability to save constructed components can allow for fewer pre-built sub-components to be required, potentially making the size of the originally downloaded application smaller than an application that does not allow for the saving of constructed components.

[0034] In some embodiments, the application may allow a user to create a design in one location and then reopen the design saved at exactly the same location at a later time. In some embodiments, the application may allow a user to create a design in one location and then reopen the design saved at any other location within the real world. For example, this may allow a user to create a design in their office and then reopen the design for a presentation during a meeting in a conference room.

[0035] However, some applications may only allow reopening of designs saved at a specific location within the real world, which can be a problem for users who can create designs in their office but need to share the design in a conference room. In a system that only allows reopening of designs saved at a specific location, if the real-world location is no longer available (e.g., the user's office building burns down), the saved design may no longer be accessible because it is dependent on that specific location (e.g., it may not exist at a second location or may include objects that are digitally tethered or installed to real-world objects with different characteristics). Additionally, a system that only allows a user to reopen a saved design, for example, once per session or once per room, may not meet the user's needs. For example, a user may desire to present the user's design during a meeting from multiple perspectives simultaneously and thus may desire to load multiple copies of the saved design into the conference room.

[0036] The systems and methods of the present application solve these problems. Such a system may, for example, enable a user to define virtual content based on what the user perceives in an augmented reality environment. The system may then save a digital representation of that virtual content as a scene. At a later time, the user may instruct the same or possibly a different augmented reality system to open the saved scene. As part of opening the saved scene, the augmented reality system may incorporate the virtual content of the saved scene into a composite reality environment for the user of the augmented reality system such that the user for whom the scene is opened can then perceive the virtual content.

[0037] In some embodiments, a scene may be assembled from a plurality of constructed components that may be constructed by defining a combination of pre-built sub-components. The constructed components can be of any complexity and may, for example, even be pre-saved scenes. Further, the constructed components need not be assembled from pre-built sub-components. The components may also be constructed using tools provided by the augmented reality system. As an example, the system may process data about a physical object collected using sensors of the augmented reality system and form a digital representation of the physical object. This digital representation may be used to render a representation of the physical object and thus serve as a virtual object. Further, it should be understood that the scene to be saved need not even have a plurality of constructed components or a plurality of pre-built sub-components. The scene may have a single component. Thus, the description of saving or opening a scene refers to the manipulation of virtual content at any level of complexity and from any source.

[0038] A saved scene may comprise at least one saved scene anchor node for the saved scene (e.g., a parent node within a hierarchical data structure that may represent the saved scene in a coordinate system within space). The virtual content of the scene may have a spatial relationship established with respect to the saved scene anchor node such that once the location of the saved scene anchor node is established within the environment of a user of the augmented reality system, the location of the virtual content of the scene within that environment can be determined by the system. Using this location information, the system may render the virtual content of the scene to the user. The location of the virtual content of the saved scene within the environment of a user of the augmented reality system may be determined, for example, by user input that positions a visual anchor that represents a location within the user's environment. The user may manipulate the location of the visual anchor through the virtual user interface of the augmented reality system. The virtual content of the saved scene may be rendered with the saved scene anchor node that is aligned with the visual anchor.

[0039] In some scenarios, the saved scene anchor node may correspond to a fixed location within the physical world, and the saved scene may be reopened with the virtual content of the scene having the same position relative to the location it had when the scene was saved. In that case, the user may experience the scene when the scene is opened if that fixed location within the physical world is within the user's environment. In such scenarios, the saved scene anchor node may comprise a persistent coordinate frame (PCF) that is derived from an object existing in the real world that changes little, not at all, or infrequently over time. The location associated with the saved scene anchor node may be represented by the saved PCF. The saved PCF may be utilized to reopen the saved scene such that the saved scene is rendered at the exact same location within the space where it was rendered when the saved scene was saved.

[0040] Alternatively, the location within the physical world where the augmented reality system can render the content of the saved scene to the user may not be fixed. The location may depend on user input or on the user's surroundings when the scene is opened. For example, if the user needs to open scenes saved at different locations, the user can do so by appropriately utilizing one or more adjustable visual anchor nodes.

[0041] Alternatively or in addition, when the saved scene is reopened for the user while a fixed location is within its environment or its field of view, the system may conditionally position the virtual content of the scene saved at that fixed location. If not applicable, the system may provide a user interface through which the user can provide an input indicating the location where the virtual content of the saved scene will be located. The system can accomplish this by identifying the PCF (current PCF) closest to the user. If the current PCF matches the saved PCF, the system may reopen the saved scene with the virtual content of the scene saved at the exact same location as when it was saved, by placing the virtual objects of the scene saved in a spatial configuration fixed to the PCF. If the current PCF does not match the saved PCF, the system may preview-place the saved scene at a default location or at a location selected by the user. Based on the preview-placement, the user or the system may move the entire saved scene to the desired location and tell the system to instantiate the scene. In some embodiments, the saved scene may be rendered at the default location in a preview format that lacks some details or functionality of the saved scene. The step of instantiating the scene may include the step of rendering to the user the complete saved scene, including all visual, physical, and other saved scene data.

[0042] One skilled in the art can address the problem of saving scenes by creating a process for saving the constructed subcomponents and a separate process for saving the scenes. An exemplary system as described in this application may merge the two processes and provide a single user interaction. This has the benefits of computer operation. For example, writing and managing a single process instead of two can improve reliability. Additionally, since the processor only needs to access one program instead of accessing and switching between two or more processes, the processor may be able to operate faster. Additionally, there may be a benefit in terms of usability for the user in that the user only needs to learn one interaction instead of two or more interactions.

[0043] If a user desires to view several virtual objects in one room, e.g., 20 different virtual objects in the user's office, without saving the scene, the system would need to track 20 different object locations relative to the real world. One benefit of the systems and methods of this application is that all 20 virtual objects can be saved within a single scene such that only a single saved scene anchor node (e.g., a PCF) needs to be tracked (the 20 objects are placed relative to the PCF rather than the world and may require less computation). This can have the benefit of computer operation in terms of less computation, which can lead to less battery usage, or less heat generated during processing, and / or enable the application to launch on a smaller processor.

[0044] The disclosed system and method may enable a simple and easy unified user experience for saving and / or reopening scenes within an application by creating a single user interaction in each of a plurality of different situations. The step of reopening a scene may include the step of loading and / or instantiating the scene. Three examples of situations in which a user may reopen a scene using the same interaction are as follows. In situation 1, the user may desire to reopen the saved scene such that the user appears in the exact same location and environment where the scene was saved (e.g., the saved PCF matches the current PCF). In situation 2, the user may desire to reopen the saved scene such that the user appears in the exact same location where the scene was saved, but the environment is different (e.g., the saved PCF matches the current PCF, but the digital mesh describing the physical environment has changed). For example, the user may save and reopen a scene in the user's office during work, but the user may have added an extra table to the office. In situation 3, the user may desire to reopen the saved scene such that the user appears in a different location with an environment different from where the scene was saved (e.g., the saved PCF does not match the current PCF and the saved world mesh does not match the current world mesh).

[0045] Regardless of the situation, the user may interact with the system through the same user interface using controls that are available in each of the plurality of situations. The user interface may be, for example, a graphical user interface in which the control is associated with an icon visible to the user, and the control is activated by the user taking an action indicating selection of the icon. To save a scene, for example, the user may select a camera icon (1900, FIG. 19), frame the scene, and then capture an image.

[0046] Regardless of the situation, to load a scene, the user may, for example, select a saved scene icon such as icon 1910 (Figure 19). The user may pull out a saved scene icon from the user menu (which can create a preview of the saved scene), and an example thereof is illustrated in Figures 19 and 20. The user may then release the saved scene icon (which can set a visual anchor node for the saved scene object). Optionally, the user may move the visual anchor node to move the saved scene relative to the real world, and then instantiate the scene (which can send the complete saved scene data into the rendering pipeline).

[0047] One of ordinary skill in the art may address three situations as three different problems and create three separate solutions (e.g., programs) and user interactions to solve those problems. The systems and methods of the present application solve these three problems using a single program. This has the benefits of computer operation. For example, writing and managing a single process instead of two processes improves reliability. Additionally, since the processor only needs to access one program instead of switching between multiple processes, the processor can operate faster. An additional benefit of computer operation is provided because the system only needs to track a single point (e.g., an anchor node) within the real world. (Example of 3D Display of Wearable System)

[0048] A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present a 2D or 3D virtual image to a user. The image may be a still image, a frame of a video, or a video in a combination or equivalent. The wearable system can include wearable devices that can present a VR, AR, or MR environment, alone or in combination, for user interaction. The wearable device can be a head-mounted device (HMD) that is used synonymously with an AR device (ARD). Further, for the purposes of this disclosure, the term "AR" is used synonymously with the term "MR".

[0049] FIG. 1 depicts an illustration of a composite reality scenario as viewed by a person using an MR system with a virtual reality object and a physical object. FIG. 1 depicts an MR scene 100 in which the user of the MR technology sees a real-world park-like setting 110 featuring people, trees, buildings in the background, and a concrete platform 120. In addition to these items, the user of the MR technology also "sees" a robotic image 130 standing on the real-world platform 120 and an avatar character 140 in the form of a flying cartoon that appears like an anthropomorphic bumblebee, although these elements do not exist in the real world.

[0050] It may be desirable for a 3D display to generate a perspective response corresponding to its virtual depth for each point within the display's field of view in order to produce a true sense of depth, more specifically, a simulated sense of surface depth. If the perspective response for a display point does not correspond to the virtual depth of that point such that it is determined by both binocular depth cues of convergence and stereopsis, the human eye will experience a vergence conflict, resulting in unstable imaging, harmful eye strain, headaches, and in the absence of vergence information, a near-complete lack of surface depth.

[0051] VR, AR, and MR experiences can be provided by a display system having a display that provides an image corresponding to a plurality of depth planes to a viewer. The images may be different for each depth plane (e.g., providing a somewhat different presentation of a scene or object), and are separately focused by the viewer's eyes, thereby based on the eye accommodation required to focus on different image features of a scene located on different depth planes, or based on observing different image features on different depth planes that are out of focus, which can help provide depth cues to the user. As discussed anywhere in this specification, such depth cues provide a reliable perception of depth.

[0052] FIG. 2 illustrates an embodiment of a wearable system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems for supporting the functions of the display 220. The display 220 may be coupled to a frame 230, which can be worn by a user, wearer, or viewer 210. The display 220 can be positioned in front of the eyes of the user 210. The display 220 can present AR / VR / MR content to the user. The display 220 can comprise a head-mounted display that is worn on the user's head. In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the user's external ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other external ear canal to provide stereo / spatial sound control). The display 220 can include an audio sensor (e.g., a microphone) for detecting an audio stream from an environment in which voice recognition is to be performed.

[0053] The wearable system 200 can include an outward-facing imaging system 464 (shown in FIG. 4) that observes the world within the environment around the user. The wearable system 200 can also include an inward-facing imaging system 462 (shown in FIG. 4) that can track the user's eye movements. The inward-facing imaging system may track the movement of one eye or both eyes. The inward-facing imaging system 462 may be attached to the frame 230 and communicate electrically with a processing module 260 or 270 that processes the image information obtained by the inward-facing imaging system and can determine, for example, the pupil diameter or orientation, eye movement, or eye pose of the user 210.

[0054] As an example, the wearable system 200 can obtain an image that reveals the user's pose using the outward-facing imaging system 464 or the inward-facing imaging system 462. The image may be a still image, a frame of video, or video, or any combination of such information sources or other similar information sources.

[0055] The display 220 can be operably coupled (250) to a local data processing module 260 and can be mounted in various configurations, such as fixed to the frame 230 by a wired conductor or wireless connectivity, fixed to a helmet or hat worn by the user, built into headphones, or removably attached to the user 210 in another manner (e.g., in a backpack configuration, in a belt attachment configuration).

[0056] The local processing and data module 260 may comprise a digital memory such as a hardware processor and a non-volatile memory (e.g., flash memory), both of which may be utilized to assist in the processing, caching, and storage of data. The data may include (a) data captured from sensors such as an image capture device (e.g., a camera within an inward-facing imaging system and / or an outward-facing imaging system), an audio sensor (e.g., a microphone), an inertial measurement unit (IMU), an accelerometer, a compass, a global positioning system (GPS) unit, a wireless device, or a gyroscope (e.g., operatively coupled to the frame 230 or otherwise attachable to the user 210), or (b) data that may be obtained or processed using the remote processing module 270 or the remote data repository 280 for possible passage to the display 220 after processing or reading. The local processing and data module 260 may be operatively coupled to the remote processing module 270 or the remote data repository 280 via a communication link 262 or 264, such as a wired or wireless communication link, such that these remote modules are available as resources to the local processing and data module 260. Additionally, the remote processing module 280 and the remote data repository 280 may be operatively coupled to each other.

[0057] In some embodiments, the remote processing module 270 may comprise one or more processors configured to analyze and process data or image information. In some embodiments, the remote data repository 280 may comprise a digital data storage facility, which may be available through other networking configurations in an Internet or "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, enabling complete autonomy from the remote modules.

[0058] The human visual system is complex and it is difficult to provide a realistic perception of depth. Although not limited by any particular theory, it is thought that an object viewer can perceive an object in three dimensions due to a combination of convergence / divergence motion and accommodation. The convergence / divergence motion of two eyes relative to each other (i.e., the rotational movement of the pupils towards or away from each other to converge the lines of sight of the eyes and fixate on an object) is closely associated with the focusing (or "accommodation") of the eye's lens. Under normal conditions, a change in the focus of the eye's lens or the eye's accommodation to change the focus from one object to another at a different distance will automatically cause a corresponding change in the convergence / divergence motion to the same distance under a relationship known as the "accommodation-vergence reflex". Similarly, a change in the convergence / divergence motion will, under normal conditions, induce a corresponding change in accommodation. A display system that provides a better match between accommodation and convergence / divergence motion can form a more realistic and comfortable simulation of a three-dimensional image.

[0059] Figure 3 illustrates a side view of an approach for simulating a 3D image using multiple depth planes. Referring to Figure 3, objects at various distances from eyes 302 and 304 on the z-axis are focused by eyes 302 and 304 such that those objects are in focus. Eyes 302 and 304 take on a particular focused state and focus the objects at different distances along the z-axis. As a result, a particular focused state can be said to be associated with a particular one of depth planes 306 having an associated focal length such that an object or a portion of an object in a particular depth plane is in focus when the eyes are in the focused state with respect to that depth plane. In some embodiments, the 3D image may be simulated by providing different presentations of the image for each of eyes 302 and 304 and also by providing different presentations of the image corresponding to each of the depth planes. Although shown as being separate for purposes of clarity, it should be understood that the fields of view of eyes 302 and 304 may overlap, for example, as the distance along the z-axis increases. Additionally, although shown as being flat for purposes of facilitating illustration, it should be understood that the contours of the depth planes may be curved in physical space such that all features within the depth plane are in focus with the eyes in a particular focused state. Without being limited by theory, it is believed that the human eye typically interprets a finite number of depth planes and can provide depth perception. As a result, a highly realistic simulation of the perceived depth can be achieved by providing different presentations of the image corresponding to each of these limited number of depth planes to the eyes. (Waveguide stack assembly)

[0060] FIG. 4 illustrates an example of a waveguide stack for outputting image information to a user. Wearable system 400 includes a stack of waveguides or a stacked waveguide assembly 480 that can be utilized to provide three-dimensional perception to the eye / brain using a plurality of waveguides 432b, 434b, 436b, 438b, 4400b. In some embodiments, wearable system 400 may correspond to wearable system 200 of FIG. 2, and FIG. 4 schematically shows some parts of wearable system 200 in more detail. For example, in some embodiments, waveguide assembly 480 may be integrated within display 220 of FIG. 2.

[0061] Continuing to refer to FIG. 4, waveguide assembly 480 may also include a plurality of features 458, 456, 454, 452 between the waveguides. In some embodiments, features 458, 456, 454, 452 may be lenses. In other embodiments, features 458, 456, 454, 452 may not be lenses. Rather, they may simply be spacers (e.g., a cladding layer or structure for forming an air gap).

[0062] Waveguides 432b, 434b, 436b, 438b, 440b, or a plurality of lenses 458, 456, 454, 452 may be configured to transmit image information to the eye using various levels of wavefront curvature or ray divergence. Each waveguide level may be associated with a particular depth plane and may be configured to output image information corresponding to that depth plane. Image input devices 420, 422, 424, 426, 428 may be utilized to input image information into waveguides 440b, 438b, 436b, 434b, 432b, respectively, which may be configured to disperse incident light across each individual waveguide for output toward the eye 410. Light exits from the output surfaces of image input devices 420, 422, 424, 426, 428 and is input into the corresponding input edges of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, a single beam of light (e.g., a collimated beam) may be input into each waveguide and output an entire field of cloned collimated beams directed toward the eye 410 at a particular angle (and divergence amount) corresponding to a depth plane associated with a particular waveguide.

[0063] In some embodiments, image input devices 420, 422, 424, 426, 428 are discrete displays that each produce image information for input into the corresponding waveguides 440b, 438b, 436b, 434b, 432b, respectively. In some other embodiments, image input devices 420, 422, 424, 426, 428 are the output ends of a single multiplexed display that can send image information to each of image input devices 420, 422, 424, 426, 428, e.g., via one or more optical waveguides (such as an optical fiber cable).

[0064] Controller 460 controls the operation of the stacked waveguide assemblies 480 and the image input devices 420, 422, 424, 426, 428. Controller 460 includes programming (e.g., instructions in a non-transitory computer-readable medium) that adjusts the timing and provides the image information to waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, controller 460 may be a single integrated device or a distributed system connected by wired or wireless communication channels. Controller 460 may be part of processing module 260 or 270 (illustrated in FIG. 2) in some embodiments.

[0065] Waveguides 440b, 438b, 436b, 434b, 432b may be configured to propagate light within each individual waveguide by total internal reflection (TIR). The waveguides 440b, 438b, 436b, 434b, 432b may each be planar, or have another shape (e.g., curved), with a major top surface and a major bottom surface and an edge extending between their major top and bottom surfaces. In the illustrated configuration, the waveguides 440b, 438b, 436b, 434b, 432b each include light extraction optical elements 440a, 438a, 436a, 434a, 432a configured to extract light out of the waveguide by redirecting the light propagating within each individual waveguide out of the waveguide and outputting the image information to the eye 410. The extracted light may also be referred to as external coupled light, and the light extraction optical elements may also be referred to as external coupling optical elements. The beam of the extracted light is output by the waveguide at the location where the light propagating within the waveguide impinges on the light redirecting element. The light extraction optical elements (440a, 438a, 436a, 434a, 432a) may be, for example, reflective or diffractive optical features. For ease of explanation and clarity of the drawings, they are shown disposed on the bottom major surface of the waveguides 440b, 438b, 436b, 434b, 432b, but in some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be disposed on the top major surface or the bottom major surface, or may be disposed directly within the volume of the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be attached to a transparent substrate and formed within a layer of the material forming the waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, the waveguides 440b, 438b, 436b, 434b, 432b may be a monolithic piece of material, and the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed on and / or within the surface of that piece of material.

[0066] Continuing to refer to FIG. 4, as discussed herein, each waveguide 440b, 438b, 436b, 434b, 432b is configured to output light and form an image corresponding to a particular depth plane. For example, the waveguide 432b closest to the eye may be configured to deliver collimated light to the eye 410 as it is input into such waveguide 432b. The collimated light may represent an optically infinite focal plane. The next upper waveguide 434b may be configured to output collimated light that passes through a first lens 452 (e.g., a negative lens) before reaching the eye 410. The first lens 452 may be configured to create a slightly convex wavefront curvature such that the eye / brain interprets the light originating from its next upper waveguide 434b as originating from a first focal plane that is closer inwardly toward the eye 410 from the optically infinite. Similarly, the third upper waveguide 436b passes its output light through both the first lens 452 and the second lens 454 before reaching the eye 410. The combined refractive power of the first and second lenses 452 and 454 may be configured to create another incremental amount of wavefront curvature such that the eye / brain interprets the light originating from the third waveguide 436b as originating from a second focal plane that is even closer inwardly toward the person from the optically infinite than the light from the next upper waveguide 434b was interpreted as originating from.

[0067] Other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are similarly configured, and the top waveguide 440b in the stack sends its output through all of the lenses between it and the eye to represent the aggregated focusing power of the focal plane closest to the person. When viewing / interpretating light originating from the world 470 on the other side of the stacked waveguide assembly 480, a compensating lens layer 430 may be disposed on top of the stack to compensate for the stack of lenses 458, 456, 454, 452 in order to compensate for the aggregating power of the lower lens stack 458, 456, 454, 452. Such a configuration provides the same number of perceived focal planes as there are available waveguide / lens pairs. Both the light extraction optical elements of the waveguides and the focusing sides of the lenses may be static (e.g., not dynamic or electroactive). In some alternative embodiments, one or both may be dynamic using electroactive features.

[0068] Continuing to refer to FIG. 4, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be configured to redirect light from their respective waveguides for a particular depth plane associated with the waveguide and output the light with an appropriate amount of divergence or collimation. As a result, waveguides having different associated depth planes may have differently configured light extraction optical elements that output light with different amounts of divergence depending on the associated depth plane. In some embodiments, as discussed herein, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be three-dimensional or surface features configured to output light at specific angles. For example, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume holograms, surface holograms, and / or diffraction gratings. Light extraction optical elements such as diffraction gratings are described in U.S. Patent Publication No. 2015 / 0178939, published Jun. 25, 2015, which is incorporated herein by reference in its entirety.

[0069] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a are diffraction features or “diffractive optical elements” (also referred to herein as “DOEs”) that form a diffraction pattern. Preferably, the DOE has a relatively low diffraction efficiency such that only a portion of the light of the beam is deflected toward the eye 410 at each intersection of the DOE, while the remainder continues to travel through the waveguide via total internal reflection. The light carrying the image information can thus be split into a plurality of associated output beams that exit the waveguide at multiple locations, resulting in a very uniform pattern of output emission toward the eye 304 with respect to this particular collimated beam that bounces within the waveguide.

[0070] In some embodiments, one or more DOEs may be switchable between an “on” state in which they actively diffract and an “off” state in which they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer dispersed liquid crystal in which microdroplets comprise a diffraction pattern, and the refractive index of the microdroplets can be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract the incident light), or the microdroplets can be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts the incident light).

[0071] In some embodiments, the number and distribution of depth planes or depth of field may vary dynamically based on the pupil size or orientation of the viewer's eye. The depth of field may vary inversely with the pupil size of the viewer. As a result, as the pupil size of the viewer's eye decreases, the depth of field increases such that a plane that was indistinguishable because its location was beyond the depth of focus of the eye becomes distinguishable, and may appear more in focus with the reduction in pupil size and corresponding increase in depth of field. Similarly, the number of separated depth planes used to present different images to the viewer may be reduced with a reduced pupil size. For example, a viewer may not be able to clearly perceive the details of both a first depth plane and a second depth plane at one pupil size without adjusting the eye's focusing from one depth plane to the other. However, these two depth planes may be simultaneously sufficiently in focus for the user at another pupil size without changing the focusing adjustment.

[0072] In some embodiments, the display system may vary the number of waveguides that receive image information based on a determination of pupil size and / or orientation, or in response to receiving an electrical signal indicating a particular pupil size or orientation. For example, if the user's eye is indistinguishable between two depth planes associated with two waveguides, the controller 460 (which may be an embodiment of the local processing and data module 260) can be configured or programmed to stop providing image information to one of these waveguides. Advantageously, this can reduce the processing burden on the system, thereby increasing the responsiveness of the system. In embodiments where the DOE for the waveguide is switchable between on and off states, the DOE may be switched to the off state when the waveguide receives image information.

[0073] In some embodiments, it may be desirable to satisfy the condition that the emitted beam has a diameter less than the diameter of the viewer's eye. However, satisfying this condition can be difficult in light of the variability of the viewer's pupil size. In some embodiments, this condition is satisfied over a wide range of pupil sizes by varying the size of the emitted beam in response to a determination of the viewer's pupil size. For example, as the pupil size decreases, the size of the emitted beam may also decrease. In some embodiments, the emitted beam size may be varied using a variable aperture.

[0074] The wearable system 400 can include an outward-facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 can be referred to as the field of view (FOV) of the world camera, and the imaging system 464 is sometimes also referred to as the FOV camera. The entire area available for viewing or imaging by the viewer can be referred to as the field of regard (FOR). The FOR may include a solid angle of 4π steradians surrounding the wearable system 400 so that the wearer can move their body, head, or eyes and perceive substantially any direction in space. In other contexts, the movement of the wearer may be more restricted, and accordingly, the wearer's FOR may contact a smaller solid angle. Images obtained from the outward-facing imaging system 464 can be used, for example, to track gestures made by the user (e.g., hand or finger gestures) or to detect objects within the world 470 in front of the user.

[0075] The wearable system 400 can also include an inward-facing imaging system 466 (e.g., a digital camera) that observes user movements such as eye movements and face movements. The inward-facing imaging system 466 can capture an image of the eye 410 and may be used to determine the size or orientation of the pupil of the eye 304. The inward-facing imaging system 466 can be used to determine the direction in which the user is looking (e.g., eye posture), or to acquire an image for user biometric identification (e.g., via iris identification). In some embodiments, at least one camera can be used to separately determine the pupil size or eye posture of each eye independently for each eye, thereby enabling the presentation of image information to each eye to be dynamically adjusted with respect to that eye. In some other embodiments, only the pupil diameter or orientation of a single eye 410 (e.g., using only a single camera per pair of eyes) is determined and assumed to be similar for both eyes of the user. Images acquired by the inward-facing imaging system 466 may be analyzed to determine the user's eye posture or mood, which can be used by the wearable system 400 to determine the audio or visual content to be presented to the user. The wearable system 400 can also use sensors such as an IMU, accelerometer, gyroscope, etc. to determine the head posture (e.g., head position or head orientation).

[0076] The wearable system 400 can include a user input device 466 through which a user can input commands to the controller 460 and interact with the wearable system 400. For example, the user input device 466 can include a trackpad, a touch screen, a joystick, a multi-degree-of-freedom (DOF) controller, a capacitance sensing device, a game controller, a keyboard, a mouse, a directional pad (D-pad), a wand, a haptic device, a totem, a component that senses the movement of the user recognized by the system as an input (e.g., a virtual user input device), and the like. A multi-DOF controller can sense user input in some or all of the possible translational (e.g., left / right, forward / backward, or up / down) or rotational (e.g., yaw, pitch, or roll) movements of the controller. A multi-DOF controller that supports translational movement can be referred to as 3DOF, while a multi-DOF controller that supports both translational and rotational movement can be referred to as 6DOF. In some cases, the user may use a finger (e.g., the thumb) to press or swipe on a touch sensor-based input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input device 466 may be held by the user's hand during use of the wearable system 400. The user input device 466 can communicate with the wearable system 400 either wired or wirelessly.

[0077] FIG. 5 shows an embodiment of an output beam output by a waveguide. One waveguide is shown, but it should be understood that other waveguides within waveguide assembly 480 may function similarly, and waveguide assembly 480 includes a plurality of waveguides. Light 520 is input into waveguide 432b at input edge 432c of waveguide 432b and propagates within waveguide 432b by TIR. At the point where light 520 impinges on DOE 432a, a portion of the light exits the waveguide as output beam 510. Output beams 510 are shown as being substantially parallel, but they may also be redirected to propagate to eye 410 at an angle (e.g., to form a diverging output beam) depending on the depth plane associated with waveguide 432b. It should be understood that a waveguide with an optical extraction element that externally couples light to form an image that appears to be set in a depth plane at a long distance (e.g., optical infinity) from eye 410 may be shown for the substantially parallel output beams. Other waveguides or other sets of optical extraction elements may output a more diverging output beam pattern, which would require eye 410 to focus at a closer distance and would be interpreted by the brain as light from a distance closer to eye 410 than optical infinity.

[0078] FIG. 6 is a schematic diagram showing an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in the generation of a multi-focus stereo display, an image, or a light field. The optical system can include a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem. The optical system can be used to generate a multi-focus stereo, an image, or a light field. The optical system can include one or more primary planar waveguides 632a (only one is shown in FIG. 6) and one or more DOEs 632b associated with at least some of the primary waveguides 632a. The planar waveguide 632b can be similar to the waveguides 432b, 434b, 436b, 438b, 440b discussed with reference to FIG. 4. The optical system can employ a diffractive waveguide device to relay light along a first axis (vertical or Y-axis in the figure of FIG. 6) and expand the effective exit pupil of the light along the first axis (e.g., the Y-axis). The diffractive waveguide device can include, for example, a diffractive planar waveguide 622b and at least one DOE 622a (illustrated by the dashed line) associated with the diffractive planar waveguide 622b. The diffractive planar waveguide 622b can be similar to or the same as the primary planar waveguide 632b having a different orientation at at least some points. Similarly, at least one DOE 622a can be similar to or the same as the DOE 632a at at least some points. For example, the diffractive planar waveguide 622b or the DOE 622a can each be made of the same material as the primary planar waveguide 632b or the DOE 632a. The embodiment of the optical display system 600 shown in FIG. 6 can be integrated into the wearable system 200 shown in FIG. 2.

[0079] The relayed light with an expanded exit pupil can be optically coupled into one or more primary planar waveguides 632b from the diffractive waveguide device. The primary planar waveguide 632b can preferably relay light along a second axis (e.g., horizontal or X-axis in the figure of FIG. 6) that is orthogonal to the first axis. It should be noted that the second axis can be a non-orthogonal axis with respect to the first axis. The primary planar waveguide 632b expands the effective exit pupil of the light along its second axis (e.g., X-axis). For example, the diffractive planar waveguide 622b can pass the light through the primary planar waveguide 632b that can relay and expand the light along the vertical or Y-axis and relay and expand the light along the horizontal or X-axis.

[0080] The optical system may include one or more colored light sources (e.g., red, green, and blue laser light) 610 that can be optically coupled into the proximal end of the single-mode optical fiber 640. The distal end of the optical fiber 640 may be screwed or received through the hollow tube 642 of the piezoelectric material. The distal end projects from the tube 642 as a flexible cantilever 644 that is not fixed. The piezoelectric tube 642 can be associated with four quadrant electrodes (not shown). The electrodes may be deposited, for example, on the outside, outer surface, or outer circumference, or diameter of the tube 642. The core electrode (not shown) may also be located at the core, center, inner circumference, or inner diameter of the tube 642.

[0081] For example, the drive electronics 650, which are electrically coupled via the wire 660, drive a pair of opposing electrodes and independently bend the piezoelectric tube 642 about two axes. The protruding distal tip of the optical fiber 644 has a mechanical resonance mode. The resonance frequency can depend on the diameter, length, and material properties of the optical fiber 644. By vibrating the piezoelectric tube 642 near the first mechanical resonance mode of the fiber cantilever 644, the fiber cantilever 644 can be vibrated and swept through large deflections.

[0082] By stimulating resonant vibrations along two axes, the tip of the fiber cantilever 644 is scanned in two axes within the area filling the 2-D scan. By modulating the intensity of the light source 610 in synchronization with the scan of the fiber cantilever 644, light emitted from the fiber cantilever 644 can form an image. An explanation of such a setup is provided in U.S. Patent Publication No. 2014 / 0003762, which is incorporated herein by reference in its entirety.

[0083] Components of the optical coupler subsystem can collimate the light emitted from the scanning fiber cantilever 644. The collimated light can be reflected by the mirror surface 648 into a narrow-dispersion planar waveguide 622b containing at least one diffractive optical element (DOE) 622a. The collimated light can propagate perpendicular (with respect to the view in the figure of FIG. 6) along the dispersion planar waveguide 622b by TIR, and thereby can repeatedly intersect the DOE 622a. The DOE 622a preferably has a low diffraction efficiency. This can diffract a portion of the light (e.g., 10%) towards the edge of the larger primary planar waveguide 632b at each point of intersection with the DOE 622a and can continue a portion of the light on its original trajectory along the length of the dispersion planar waveguide 622b via TIR.

[0084] At each point of intersection with the DOE 622a, additional light can be diffracted towards the entrance of the primary waveguide 632b. By splitting the incident light into a plurality of external coupling sets, the exit pupil of the light can be vertically expanded by the DOE 622a within the dispersion planar waveguide 622b. The vertically expanded light that is externally coupled from the dispersion planar waveguide 622b can be incident on the edge of the primary planar waveguide 632b.

[0085] Light incident on the first waveguide 632b can propagate horizontally along the first waveguide 632b (with respect to the figure in FIG. 6) via TIR. As the light intersects the DOE 632a at multiple points, it propagates horizontally along at least a portion of the length of the first waveguide 632b via TIR. The DOE 632a advantageously has a phase profile that is a sum of a linear diffraction pattern and a radially symmetric diffraction pattern and may be designed or configured to produce both deflection and focusing of the light. The DOE 632a advantageously has a low diffraction efficiency (e.g., 10%) such that only a portion of the light of the beam is deflected towards the viewer's eye at each intersection of the DOE 632a, while the remainder of the light continues to propagate through the first waveguide 632b via TIR.

[0086] At each point of intersection between the propagating light and the DOE 632a, a portion of the light is diffracted towards the adjacent surface of the first waveguide 632b, allowing the light to escape from TIR and be emitted from the surface of the first waveguide 632b. In some embodiments, the radially symmetric diffraction pattern of the DOE 632a additionally imparts a certain focal level to the diffracted light, shaping (e.g., imparting curvature to) the wavefronts of the individual beams and steering the beams to an angle that matches the designed focal level.

[0087] Therefore, these different paths can couple light out of the primary planar waveguide 632b by resulting in different multiplicity of the DOE632a, focal levels, or filling patterns at different angles in the exit pupil. Different filling patterns in the exit pupil can be beneficially used to create a light field display with multiple depth planes. Each layer within the waveguide assembly or a set of layers within a stack (e.g., three layers) may be employed to generate an individual color (e.g., red, blue, green). Thus, for example, a first set of three adjacent layers may be employed to produce red, blue, and green light, respectively, at a first focal depth. A second set of three adjacent layers may be employed to produce red, blue, and green light, respectively, at a second focal depth. Multiple sets may be employed to generate a full 3D or 4D color image light field with various focal depths. (Other components of the wearable system)

[0088] In many implementations, a wearable system may include other components in addition to or instead of the components of the wearable system described above. The wearable system may include, for example, one or more haptic devices or components. The haptic device or component may be operable to provide a haptic sensation to the user. For example, the haptic device or component may provide a haptic sensation of pressure or texture when touching virtual content (e.g., virtual objects, virtual tools, other virtual structures). The haptic sensation may reproduce the sensation of a physical object represented by the virtual object or may reproduce the sensation of an imaginary object or character (e.g., a dragon) represented by the virtual content. In some implementations, the haptic device or component may be worn by the user (e.g., a user-wearable glove). In some implementations, the haptic device or component may be held by the user.

[0089] A wearable system may include, for example, one or more physical objects that are operable by a user and enable input to or interaction with the wearable system. These physical objects may be referred to herein as totems. Some totems may take the form of inanimate objects, such as pieces of metal or plastic, walls, table surfaces, etc. In some implementations, a totem may not actually have any physical input structures (e.g., keys, triggers, joysticks, trackballs, rocker switches). Instead, a totem may simply provide a physical surface, and the wearable system may render a user interface so as to appear to the user to be on one or more surfaces of the totem. For example, the wearable system may render an image of a computer keyboard and trackpad so as to appear to be resident on one or more surfaces of the totem. For example, the wearable system may render a virtual computer keyboard and virtual trackpad so as to appear on the surface of a thin rectangular plate of aluminum that serves as a totem. The rectangular plate itself does not have any physical keys or trackpads or sensors. However, the wearable system may detect user operations or interactions or touches using the rectangular plate as selections or inputs made via the virtual keyboard or virtual trackpad. The user input device 466 (shown in FIG. 4) may be an embodiment of a totem that may include a trackpad, touchpad, trigger, joystick, trackball, rocker or virtual switch, mouse, keyboard, multi-degree-of-freedom controller, or another physical input device. The user may use the totem alone or in combination with a gesture to interact with the wearable system and / or other users.

[0090] Examples of haptic devices and totems that can be used with the wearable devices, HMDS, and display systems of the present disclosure are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. (Exemplary Wearable Systems, Environments, and Interfaces)

[0091] The wearable system may employ various mapping-related techniques to achieve within a light field rendered with a high depth of field. When mapping the virtual world, it is advantageous to capture all features and points within the real world and accurately depict virtual objects in relation to the real world. To achieve this goal, the FOV images captured from the user of the wearable system can be added to the world model by including new photographs that convey information about various points and features of the real world. For example, the wearable system can collect a set of map points (such as 2D points or 3D points), find new map points, and render a more accurate version of the world model. The world model of the first user can be communicated to the second user (e.g., via a network such as a cloud network) so that the second user can experience the world surrounding the first user.

[0092] FIG. 7 is a block diagram of an example of an MR system 700 operable to process data related to an MR environment such as a room. The MR environment 700 may be configured to receive inputs (e.g., visual input 702 from a user's wearable system, stationary input 704 such as an indoor camera, sensory input 706 from various sensors, gestures, totems, eye tracking, user input, etc. from user input device 466) from one or more user wearable systems (e.g., wearable system 200 or display system 220) or stationary indoor systems (e.g., indoor cameras, etc.). The wearable system may use various sensors (e.g., accelerometers, gyroscopes, temperature sensors, motion sensors, depth sensors, GPS sensors, inward-facing imaging systems, outward-facing imaging systems, etc.) to determine the location and various other attributes of the user's environment. This information may be further supplemented with information from stationary cameras in the room that may provide images or various cues from different viewpoints. Image data obtained by the cameras (such as indoor cameras and / or cameras of outward-facing imaging systems) may be transformed into a set of mapping points.

[0093] One or more object recognition devices 708 may crawl through the received data (e.g., a set of points), recognize or map the points, tag the images, and attach semantic information to the objects using map database 710. The map database 710 may comprise various points collected over time and their corresponding objects. The various devices and the map database may be interconnected with each other through a network (e.g., LAN, WAN, etc.) and may access the cloud.

[0094] Based on the present information and the set of points in the map database, object recognition devices 708a…708n (only object recognition devices 708a and 708n of which are shown for simplicity) may recognize objects in the environment. For example, the object recognition device may recognize a face, a person, a window, a wall, a user input device, a television, a document (e.g., a travel document, a driver's license, a passport as described in the security examples of the present specification), other objects in the user's environment, and the like. One or more object recognition devices may be specialized for objects with certain characteristics. For example, object recognition device 708a may be used to recognize a face, while another object recognition device may be used to recognize a document.

[0095] Object recognition may be performed using a variety of computer vision techniques. For example, a wearable system may analyze an image obtained by an outward-facing imaging system 464 (shown in FIG. 4) and perform scene reconstruction, event detection, video tracking, object recognition (e.g., of a person or document), object pose estimation, face recognition (e.g., from an image of a person in the environment or on a document), learning, indexing, motion estimation, or image analysis (e.g., identifying marks in a document such as a photo, signature, identification information, travel information, etc.). One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), Oriented FAST and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoints (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, Visual Simultaneous Localization and Mapping (vSLAM) techniques, Sequential Bayesian estimators (e.g., Kalman filter, Extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), Iterative Closest Point (ICP), Semi-Global Matching (SGM), Semi-Global Block Matching (SGBM), feature point histogram, various machine learning algorithms (e.g., support vector machine, k-nearest neighbor algorithm, naive Bayes, neural network (including convolutional or deep neural network), or other supervised / unsupervised models, etc.).

[0096] Object recognition can additionally or alternatively be performed by various machine learning algorithms. Once trained, the machine learning algorithms can be stored by the HMD. Some examples of machine learning algorithms include supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., Apriori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural network, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., Stacked Generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models can be customized for individual datasets. For example, a wearable device can generate or store a base model. The base model is used as a starting point and may generate additional models specific to a data type (e.g., a particular user within a telepresence session), a dataset (e.g., a set of additional images of the user within a telepresence session), a conditional situation, or other variations. In some embodiments, the wearable HMD can be configured to generate models for the analysis of aggregated data using multiple techniques. Other techniques may include using predefined thresholds or data values.

[0097] Based on the book information and the set of points in the map database, the object recognition devices 708a - 708n may recognize an object, complement the object with semantic information, and endow the object with operability. For example, when the object recognition device recognizes that a set of points is a door, the system may attach certain semantic information (e.g., the door has hinges and has a 90 - degree movement around the hinges). When the object recognition device recognizes that a set of points is a mirror, the system may attach semantic information that the mirror has a reflective surface that can reflect images of objects in the room. The semantic information can include, as described herein, the affordances of the object. For example, the semantic information may include the normal of the object. The system can assign a vector, the direction of which indicates the normal of the object. Over time, the map database grows as the system (which may be resident locally or accessible through a wireless network) accumulates more data from the world. Once an object is recognized, the information may be transmitted to one or more wearable systems. For example, the MR system 700 may include information about a scene occurring in California. The information about the scene may be transmitted to one or more users in New York. Based on data received from the FOV camera and other inputs, the object recognition device and other software components can map the points collected from various images so that the scene can be accurately "passed" to a second user who may be in a different part of the world, and can recognize objects, etc. The environment 700 may also use a topological map for location - determination purposes.

[0098] FIG. 8 is a process flow diagram of an example of a method 800 for rendering virtual content in relation to recognized objects. Method 800 describes a way in which a virtual scene can be presented to a user of a wearable system. The user can be geographically remote from the scene. For example, the user can be present in New York but may desire to view a scene currently taking place in California, or may desire to take a walk with a friend who lives in California.

[0099] In block 810, the wearable system may receive input regarding the user's environment from the user and other users. This can be accomplished through various input devices and knowledge already held within a map database. The user's FOV camera, sensors, GPS, eye tracking, etc. communicate information to the system in block 810. The system may determine sparse points based on this information in block 820. The sparse points may be used when displaying and understanding the orientation and position of various objects around the user, and may also be used when determining pose data (e.g., head pose, eye pose, body pose, or hand gesture). Object recognition devices 708a - 708n may crawl through these collected points and use the map database to recognize one or more objects in block 830. This information may then be communicated to the user's individual wearable system in block 840, and a desired virtual scene may be displayed to the user accordingly in block 850. For example, a desired virtual scene (e.g., for a user in CA) may be displayed in appropriate orientation, position, etc. in relation to the various objects and surroundings of a user in New York.

[0100] FIG. 9 is a block diagram of another embodiment of a wearable system. In this embodiment, the wearable system 900 comprises a map that may include map data regarding the world. The map may reside, in part, locally on the wearable system and, in part, on a networked storage location (e.g., within a cloud system) accessible by a wired or wireless network. A pose process 910 is executed on a wearable computing architecture (e.g., processing module 260 or controller 460) and may utilize data from the map to determine the position and orientation of the wearable computing hardware or the user. The pose data may be calculated from data collected on-the-fly as the user experiences the system and operates within the world. The data may comprise images regarding objects in a physical or virtual environment, data from sensors (e.g., an inertial measurement unit generally comprising accelerometer and gyroscope components), and surface information.

[0101] The sparse point representation may be the output of a simultaneous localization and mapping (e.g., SLAM or vSLAM, which refers to configurations where the input is only image / vision) process. The system can be configured to find not only the locations of various components within the world, but also what the world is composed of. Pose can be a building block that achieves many goals, including using data taken into and from the map.

[0102] In one embodiment, the sparse point locations may not be entirely proper in themselves, and additional information may be required to produce a multi-focus AR, VR, or MR experience. Generally, a dense representation, which refers to depth map information, may be utilized to at least partially fill in this gap. Such information may be calculated from a process referred to as stereopsis 940, and the depth information is determined using techniques such as triangulation or time-of-flight sensing. Image information and active patterns (such as infrared patterns created using an active projector) may serve as inputs to the stereopsis process 940. A significant amount of depth map information may be fused together, and some of this may be summarized using surface representations. For example, a mathematically definable surface may be an efficient (e.g., for large point clouds) and summary-friendly input to other processing devices such as a game engine. Thus, the output of the stereopsis process (e.g., depth map) 940 may be combined in the fusion process 930. The pose 950 may similarly be an input to this fusion process 930, and the output of the fusion 930 is an input to the map ingestion process 920. Sub-surfaces may connect to each other in topographic mapping and form larger surfaces, and the map becomes a large hybrid of points and surfaces.

[0103] To resolve various aspects in the composite reality process 960, various inputs may be utilized. For example, in the embodiment depicted in FIG. 9, the game parameters may be inputs for determining that the user of the system is playing a monster battle game with one or more monsters in various locations, that the monster is dead, is fleeing under various conditions (such as when the user shoots the monster), walls or other objects in various locations, and the like. The world map may include information regarding where such objects exist relative to each other so as to be another useful input to the composite reality. The pose with respect to the world is similarly an input and plays an important role for almost any two-way system.

[0104] Control or input from the user is another input to the wearable system 900. As described herein, user input can include visual input, gestures, totems, audio input, sensory input, and the like. For example, in order to move around or play a game, the user may need to command the wearable system 900 with respect to a desired target. There are various forms of user control that can be utilized, not just moving oneself within a space. In one embodiment, a totem (e.g., a user input device) or an object such as a toy gun may be held by the user and tracked by the system. The system is preferably configured to ascertain that the user is holding the item and understand the type of interaction the user is having with the item (e.g., if the totem or object is a gun, the system may be equipped with sensors such as an IMU to assist in determining the situation that is occurring, even when such activity is not within the field of view of any of the cameras, and may be configured to understand whether the user is clicking a trigger or other sensing button or element).

[0105] Hand gesture tracking or recognition may also provide input information. The wearable system 900 may be configured to track and interpret hand gestures for button presses, gestures such as left or right, stop, grip, hold, etc. For example, in one configuration, the user may desire to scroll through an email or calendar in a non-game environment or perform a "fist bump" with another person or player. The wearable system 900 may be configured to utilize a minimal amount of hand gestures, which may be dynamic or not. For example, the gestures may be simple static gestures such as spreading the hand to indicate stop, raising the thumb to indicate OK, lowering the thumb to indicate not OK, or flipping the hand left or right or up or down to indicate a directional command.

[0106] Eye tracking is another input (e.g., tracking where the user is looking, controlling display technology, and rendering to a specific depth or range). In one embodiment, the convergence / divergence movement of the eyes may be determined using triangulation, and then the accommodation may be determined using a convergence / divergence movement / accommodation model developed for that particular person. Eye tracking may be performed by an eye camera to determine the eye gaze (e.g., the direction or orientation of one or both eyes). Other techniques, such as measurement of electrical potential by electrodes placed near the eyes (e.g., electrooculogram recording), may also be used for eye tracking.

[0107] Voice recognition can be another input that can be used alone or in combination with other inputs (e.g., totem tracking, eye tracking, gesture tracking, etc.). System 900 can include an audio sensor (e.g., a microphone) that receives an audio stream from the environment. The received audio stream can be processed (e.g., by processing modules 260, 270 or central server 1650) to recognize the user's voice from other voices or background audio and extract commands, parameters, etc. from the audio stream. For example, system 900 can identify that the phrase "Show me your ID" has been uttered from the audio stream, identify that the phrase has been uttered by the wearer of system 900 (e.g., a security examiner rather than another person within the examiner's environment), and from the context of the phrase and the situation (e.g., a security checkpoint), extract the executable command to be performed (e.g., computer vision analysis of things within the wearer's FOV) and the object on which the command is to be performed ("your ID"). System 900 can incorporate speaker recognition technology to determine the person speaking (e.g., whether the speech is from the wearer of the ARD or another person or voice (e.g., recorded speech transmitted by a loudspeaker within the environment)) and speech recognition technology to determine the content being uttered. Voice recognition techniques can include frequency estimation, hidden Markov models, Gaussian mixture models, pattern matching algorithms, neural networks, matrix representations, vector quantization, speaker diarization, decision trees, and dynamic time warping (DTW) techniques. Voice recognition techniques can also include anti-speaker techniques such as cohort models and world models. Spectral features may be used in representing speaker characteristics.

[0108] Regarding the camera system, the exemplary wearable system 900 shown in FIG. 9 can include three pairs of cameras, namely, a relatively wide FOV or passive SLAM pair of cameras arranged on both sides of the user's face, and a pair of cameras oriented in front of the user that handle the stereoscopic imaging process 940 and also capture the gestures of the hands and the trajectories of totems / objects in front of the user's face. The pair of cameras for the FOV cameras and the stereoscopic process 940 may be part of the outward-facing imaging system 464 (shown in FIG. 4). The wearable system 900 can include an eye-tracking camera (which can be part of the inward-facing imaging system 462 shown in FIG. 4) oriented towards the user's eyes to triangulate eye vectors and other information. The wearable system 900 may also include one or more textured light projectors (such as an infrared (IR) projector) to project textures into the scene.

[0109] FIG. 10 is a process flow diagram of an embodiment of a method 1000 for determining user input to a wearable system. In this embodiment, the user may interact with a totem. The user may have multiple totems. For example, the user may have one designated totem for a social media application, another totem for playing a game, etc. In block 1010, the wearable system may detect the movement of the totem. The movement of the totem may be recognized through the outward-facing imaging system or detected through sensors (such as a touch glove, an image sensor, a hand-tracking device, an eye-tracking camera, a head pose sensor, etc.).

[0110] At least in part, based on the detected gesture, eye gesture, head gesture, or input through the totem, the wearable system detects, at block 1020, the position, orientation, or movement of the totem (or the user's eye or head or gesture) relative to a reference frame. The reference frame may be a set of map points based on which the wearable system converts the movement of the totem (or the user) into an action or command. At block 1030, the user's interaction with the totem is mapped. Based on the mapping of the user interaction relative to the reference frame 1020, the system determines, at block 1040, the user input.

[0111] For example, the user may move the totem or physical object back and forth to scroll a virtual page, move to the next page, or move from one user interface (UI) display screen to another UI screen. As another example, the user may move their head or eyes to view different real or virtual objects within the user's FOR. If the user's fixation on a particular real or virtual object is longer than a threshold time, the real or virtual object may be selected as the user input. In some implementations, the user's eye convergence / divergence movement can be tracked, and a focus adjustment / convergence / divergence movement model can be used to determine the user's eye focus adjustment state, which provides information about the depth plane on which the user is in focus. In some implementations, the wearable system can use ray casting techniques to determine real or virtual objects along the direction of the user's head or eye gesture. In various implementations, the ray casting techniques can include steps of casting a thin beam of light with substantially little lateral width, or a beam of light with a substantial lateral width (e.g., a cone or frustum of a cone).

[0112] The user interface may be projected by a display system (such as display 220 in FIG. 2) as described herein. It may also be displayed using various other techniques such as one or more projectors. The projector may project an image onto a physical object such as a canvas or a sphere. Interaction with the user interface may be tracked using one or more cameras external to the system or part of the system (e.g., using an inward-facing imaging system 462 or an outward-facing imaging system 464).

[0113] FIG. 11 is a process flow diagram of an example of a method 1100 for interacting with a virtual user interface. The method 1100 may be implemented by a wearable system as described herein. Embodiments of the method 1100 may be used by a wearable system to detect a person or a document within the FOV of the wearable system.

[0114] In block 1110, the wearable system may identify a specific UI. The type of UI may be pre-determined by the user. The wearable system may identify that a specific UI needs to be captured based on user input (e.g., gestures, visual data, audio data, sensory data, direct commands, etc.). The UI may be specific to a security scenario, for example, where the wearer of the system observes a user presenting a document to the wearer (e.g., at a passenger checkpoint). In block 1120, the wearable system may generate data for a virtual UI. For example, data associated with the boundaries, general structure, shape, etc. of the UI may be generated. Additionally, the wearable system may determine the map coordinates of the user's physical location so that the wearable system can display the UI in relation to the user's physical location. For example, if the UI is centered on the body, the wearable system may determine the coordinates of the user's physical standing position, head pose, or eye pose so that a ring UI can be displayed around the user or a flat UI can be displayed on a wall or in front of the user. In the security context described herein, the UI may be displayed as if it were surrounding the traveler presenting the document to the wearer of the system so that the wearer can easily view the UI while looking at the traveler and the traveler's document. If the UI is centered on the hand, the map coordinates of the user's hand may be determined. These map points may be derived through a FOV camera, data received through sensory input, or any other type of collected data.

[0115] In block 1130, the wearable system may send data from the cloud to the display, or the data may be sent from a local database to the display component. In block 1140, the UI is presented to the user based on the sent data. For example, a light field display can project a virtual UI into one or both of the user's eyes. Once the virtual UI is created, the wearable system may, in block 1150, simply wait for commands from the user and generate more virtual content on the virtual UI. For example, the UI may be a body-centered ring around the user's body or the body of a person (e.g., a traveler) within the user's environment. The wearable system may then wait for commands (e.g., gestures, head or eye movements, voice commands, inputs from a user input device, etc.), and if recognized (block 1160), the virtual content associated with the command may be presented to the user (block 1170).

[0116] Additional embodiments of wearable systems, UIs, and user experiences (UXs) are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. (Persistent Coordinate Frame)

[0117] In some embodiments, the wearable system may store one or more persistent coordinate frames (PCFs) in a map database such as map database 710. The PCF may be constructed around a point in space in the real world (e.g., the user's physical environment) that does not change over time. In some embodiments, the PCF may be constructed around a point in space that does not change frequently over time or is unlikely to change over time. For example, most buildings are designed to stay in one place, while cars are designed to move people and things from one place to another, so a point on a building is less likely to change location over time than a point on a car.

[0118] The PCF may provide a mechanism for defining a location. In some embodiments, the PCF may be represented as a point with a coordinate system. The PCF coordinate system may be fixed in the real world and may not change for each session (i.e., when the user turns the system off and then on again).

[0119] In some embodiments, the PCF coordinate system may be an arbitrary point in space selected by the system at the start of a user session such that the PCF persists only for the duration of the session and may be consistent with a local coordinate frame. Such a PCF may be used for user pose determination.

[0120] The PCF may be recorded in a map database such as map database 710. In some embodiments, the system determines one or more PCFs and stores the PCFs in a map such as a digital map of the real world (which may be implemented in the system described herein as a "world mesh"). In some embodiments, the system may select the PCF by looking for features, points, and / or objects that are invariant over time. In some embodiments, the system may select the PCF by looking for features, points, objects, etc. that do not change during a user session on the system. The system may optionally utilize one or more computer vision algorithms in combination with other rule-based algorithms that look for one or more of the features described above. Examples of computer vision algorithms are described above in the context of object recognition. In some embodiments, the PCF may be a system-level determination that can be utilized by multiple applications or other system processes. In some embodiments, the PCF may be determined at the application level. Thus, it should be understood that the PCF can be represented in any of a plurality of ways, such as a set of one or more points or features recognizable in sensor data that represent the physical world around the wearable system. (world mesh)

[0121] 3D reconstruction is a 3D computer vision technique that takes as input an image (e.g., a colored / grayscale image, a depth image, or the like) and (e.g., automatically) generates a 3D mesh representing the user's environment and / or an observed scene such as the real world. In some embodiments, the 3D mesh representing the observed scene may be referred to as a world mesh. 3D reconstruction has many applications in virtual reality, mapping, robotics, gaming, movie production, and the like.

[0122] As an example, a 3D reconstruction algorithm can receive an input image (e.g., a colored / grayscale image, a colored / grayscale image + depth image, or depth only), and process the input image as appropriate to form a captured depth map. For example, a passive depth map can be generated from a colored image using a multi-view stereopsis algorithm, and an active depth map can be obtained using an active sensing technique such as a structured light depth sensor. Although the following examples are illustrated, embodiments of the present application may utilize a world mesh that can be generated from any suitable world mesh creation method. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0123] FIG. 17 is a simplified flowchart illustrating a method for creating a 3D mesh of a scene using a plurality of frames of a captured depth map. Referring to FIG. 17, a method for creating a 3D model of a scene, e.g., a 3D triangular mesh representing a 3D surface associated with the scene, from a plurality of frames of a captured depth map is illustrated. Method 1700 includes receiving a set of captured depth maps (step 1702). The captured depth maps are depth images having associated depth values where each pixel represents the depth from the pixel to the camera from which the depth image was acquired. Compared to a colored image that may have more than three channels per pixel (e.g., an RGB image with red, green, and blue components), the depth map can have a single channel per pixel (i.e., the pixel distance from the camera). The process of receiving a set of captured depth maps can include processing an input image, e.g., an RGB image, to produce one or more captured depth maps, also referred to as frames of the captured depth map. In other embodiments, the captured depth maps are acquired using a time-of-flight camera, LIDAR, a stereo camera, or the like, and are thus received by the system.

[0124] The set of captured depth maps includes depth maps from different camera angles and / or positions. As an example, a depth map stream can be provided by a moving depth camera. As the moving depth camera pans and / or moves, depth maps are produced as a stream of depth images. As another example, a stationary depth camera can be used to collect a plurality of depth maps of part or all of the scene from different angles and / or different positions, or combinations thereof.

[0125] The method also includes steps of aligning (1704) the camera poses associated with a set of depth maps captured within a reference frame, and overlaying (1706) the set of depth maps captured within the reference frame. In certain embodiments, the process of pose estimation is utilized to align depth points from all cameras and create a locally and globally consistent point cloud in 3D world coordinates. Depth points from the same position in world coordinates should be aligned as closely to each other as possible. However, due to inaccuracies present in the depth maps, pose estimation is usually not perfect, especially on structural features such as wall corners, wall ends, door frames within indoor scenes, and the like, and when present in the generated mesh, causes artifacts on these structural features. Further, these inaccuracies can be exacerbated when mesh boundaries are considered to be occluders (i.e., objects that occlude background objects) as the artifacts will be much more noticeable to the user.

[0126] To align the camera poses, which indicate the position and orientation of the camera associated with each depth image, the depth maps are overlaid such that differences in the positions of adjacent and / or overlapping pixels are reduced or minimized. Once the positions of the pixels within the reference frame are adjusted, the camera poses are adjusted and / or updated to align with the adjusted pixel positions. Thus, the camera poses are aligned (1706) within the reference frame. In other words, the rendered depth map can be created by projecting the depth points of all depth maps onto a reference frame (e.g., a 3D world coordinate system) based on the estimated camera pose.

[0127] The method further includes performing volumetric fusion (1708) and forming a reconstructed 3D mesh (1710). The volumetric fusion process can include fusing a plurality of captured depth maps into a volumetric representation as a discretized version of the signed distance function of the observed scene. 3D mesh generation can include the use of a marching cubes algorithm or other suitable method to extract a polygonal mesh from the volumetric representation in 3D space.

[0128] Further details explaining a method and system for creating a 3D mesh of a real-world environment (e.g., a world mesh) are provided in U.S. Non-Provisional Patent Application No. 15 / 274,823, entitled "Methods and Systems for Detecting and Combining Structural Features in 3D Reconstruction", which is hereby incorporated by reference in its entirety. (User operation process)

[0129] Figure 12 illustrates an exemplary process 1200 of user interaction using the systems and methods described herein, where a user can create and save a scene and then later open that scene. For example, a user can play a game on an AR / VR / MR wearable system such as the wearable system 200 and / or 900 described above. The game can enable the user to build a virtual structure using virtual blocks. The user may desire to build an elaborate structure, such as a replica of the user's house, spend the day in it, and save the structure for later use. The user may ultimately desire to build the user's neighborhood or the entire city. When the user saves the user's house, the user can, for example, open the house again the next day and continue working on the neighborhood. Since many neighborhoods reuse the house design, the user can build and save separately only five basic designs and can load one or more than once of those designs into a single scene to build the neighborhood. The neighborhood scene can then be saved as an additional scene. If the user desires to continue building, in some embodiments, the user can load one or more of the neighborhood scenes in combination with one or more of the five basic house designs and continue building the entire city.

[0130] The user may elect to save any combination of block designs (i.e., a single wall, a single house, an entire street of houses, a neighborhood, etc.) for subsequent reuse in future games / designs.

[0131] In this embodiment, the user starts the process of memorizing a scene by selecting an icon that serves as a control for starting the capture process. In step 1202, the user may select a camera icon. In some embodiments, the icon may not be a camera, but rather a different visual representation (e.g., text, image, etc.) of a computer program or code on the system that can create a virtual camera on the system. For example, the visual representation may be words such as "camera", "start camera", or "take a photo", or the visual representation may be an image such as an image of a camera, an image of a photo, an image of a flower, or an image of a person. Any suitable visual representation may be used.

[0132] When the user selects the camera icon or the visual representation of the virtual camera, the virtual camera may appear in the user's environment. The virtual camera may be interacted with by the user. For example, the user may look through the virtual camera and view real and virtual world content through the FOV of the camera, the user may hold the virtual camera and move it around, and / or the user may operate the camera through a user menu (e.g., to disable the virtual camera, to take a photo, etc.). In some embodiments, the user interaction is as described above in FIGS. 9-11. In some embodiments, the virtual camera may provide the same FOV as the user's FOV through a wearable device. In some embodiments, the virtual camera may provide the same FOV as the user's right eye, the user's left eye, or both of the user's eyes.

[0133] In step 1204, the virtual camera may be operated until the scene is framed. In some embodiments, the camera may frame the scene using a technique based on user preferences that can identify the manner in which the scene is selected. For example, the frame may initially be the default view through the virtual camera viewfinder (i.e., the device or part of the camera that shows the field of view of the lens and / or camera system) when the virtual camera is first presented to the user. The viewfinder may appear to the user as a preview image displayed on the virtual camera, similar to the way many real-world digital cameras have a display on the back of the real-world camera to preview an image before it is captured. In some embodiments, the user may operate the camera and change the frame of the scene. The changed frame may change the preview image displayed on the virtual camera. For example, the user may select the virtual camera by pressing a button on a multi-DOF controller such as the totem described above and move the totem to move the virtual camera. Once the user has the desired view through the virtual camera, the user may release the button and stop the movement of the virtual camera. The virtual camera may be moved in some or all of the possible translations (e.g., left / right, forward / backward, or up / down) or rotations (e.g., yaw, pitch, or roll). Alternative user interactions such as clicking and releasing the button for virtual camera selection and a second click and release of the button to release the virtual camera may be used. Other interactions may also be used to frame the scene through the virtual camera.

[0134] The virtual camera may display a preview image to the user during step 1204. The preview image may contain virtual content within the virtual camera FOV. Alternatively, or in addition, in some embodiments, the preview image may comprise a visual representation of the world mesh data. For example, to the user, a mesh version of a real-world bench within the virtual camera FOV may be visible in the same location as the actual bench. The virtual content and the mesh data may be spatially correct. For example, if a virtual avatar sits on a real-world bench, the virtual avatar will appear to sit on the mesh bench at the same location, orientation, and / or position. In some embodiments, only the virtual content is displayed within the preview image. In some embodiments, the virtual content is previewed in the same spatial arrangement as where the virtual content is installed in the real world. In some embodiments, the virtual content is previewed in different spatial arrangements such as clusters, rows, circles, or other suitable arrangements.

[0135] In some embodiments, the system may automatically frame the scene through the virtual camera such that, for example, through the viewfinder, it includes the maximum number of virtual objects possible. Alternative methods of automatic scene framing, such as framing the virtual scene so that the user's FOV matches the virtual camera FOV, may be used. In some embodiments, the system may automatically frame the scene using the object hierarchy such that higher-priority objects are within the frame. For example, living things such as people, dogs, cats, birds, etc. may have a higher priority than inanimate objects such as tables, chairs, cups, etc. Other suitable methods may be used to automatically frame the scene or to create a priority system for automatic framing.

[0136] Regardless of how the scene is framed, the system may capture the saved scene data. The saved scene data may be used by the augmented reality system to render the virtual content of the scene when the saved scene is opened later. In some embodiments, the saved scene may comprise scene objects saved in locations that are spatially fixed relative to each other.

[0137] In some embodiments, the saved scene data may comprise data that fully represents the saved scene such that the scene can be re-rendered at a later time and / or at a location different from where the scene was saved and when it was saved. In some embodiments, the saved scene data may comprise data required by the system to render and display the saved scene to the user. In some embodiments, the saved scene data comprises a saved PCF, an image of the saved scene, and / or saved scene objects.

[0138] In some embodiments, the saved PCF may be the PCF that was closest to the user when the saved scene was saved. In some embodiments, the saved PCF may be the PCF within the scene that was framed (step 1204). In some embodiments, the saved PCF may be the PCF within the user's FOV and / or FOR when the scene was saved. In some embodiments, there may be more than one PCF available for saving. In this case, the system may automatically select the most reliable PCF (e.g., the PCF that is least likely to change over time as described for the PCF above). In some embodiments, the saved scene data may have more than one PCF associated with the saved scene. The saved scene data may specify a primary PCF and one or more backup or secondary PCFs.

[0139] In some embodiments, the virtual objects that will be saved as part of a scene may be determined based on the objects within the framed scene when the scene is saved. The saved scene object may be, for example, a virtual object (e.g., digital content) that appears to be located in the user's real-world environment at the time the scene was saved. In some embodiments, the saved scene object may exclude one or more (up to all) user menus within the saved scene. In some embodiments, the saved scene object includes all virtual objects within the FOV of the virtual camera. In some embodiments, the saved scene object includes all virtual objects that the user can perceive in the real world within the user's FOR. In some embodiments, the saved scene object includes all virtual objects that the user can perceive in the real world within the user's FOV. In some embodiments, the saved scene object includes all virtual content in the user's environment regardless of whether the virtual content is within the FOV of the virtual camera, the user's FOV, and / or the user's FOR. In some embodiments, the saved scene object may include any subset of the virtual objects within the user's environment. The subset may be based on criteria such as the type of virtual object. For example, one subset may relate to building blocks, and a different subset may relate to the landscaping (e.g., plants) around a building. An exemplary process for saving a scene is described below in connection with FIG. 13A.

[0140] In addition to storing virtual content and location information, the system in step 1206 may capture an image of the scene framed. In some embodiments, the image is captured when the user provides a user interaction such as pressing a button on the controller, through a gesture, through the user's head pose, through the user's line of sight, and / or any other suitable user interaction. In some embodiments, the system may automatically capture the image. The system may automatically capture the image when the system has finished automatically framing the scene, as in some embodiments of step 1204. In some embodiments, the system may automatically capture the image, for example, by using a timer, after the user has framed the scene in step 1204 (e.g., if 5 seconds have elapsed since the user last moved the camera, the system will automatically capture the image). Other suitable methods of automatic image capture may also be used. In some embodiments, the storage of the scene may be initiated in response to the same event that triggers the capture of the image, but the two actions may be controlled independently in some embodiments.

[0141] In some embodiments, the captured image may be stored in a permanent memory such as a hard drive. In some embodiments, the permanent memory may be the local processing and data module 260 as described above. In some embodiments, the system may save the scene when the image is captured. In some embodiments, the captured image may comprise virtual objects, world meshes, real-world objects, and / or any other content with renderable data. In some embodiments, the user may desire to save more than one scene, and thus the process may loop back to step 1202. Further details of a method and system related to capturing an image comprising virtual objects and real-world objects are provided in U.S. Non-Provisional Patent Application No. 15 / 924,144, now published as US2018 / 0268611, entitled "Technique for recording augmented reality data", which is hereby expressly incorporated by reference in its entirety.

[0142] In step 1208, the saved scene icon may be selected. The saved scene icon may be the scene that was placed in the frame captured in step 1206. In some embodiments, the saved scene icon may be any suitable visual representation of the saved scene. In some embodiments, the saved scene icon may comprise the image captured in step 1206. For example, the saved scene icon may appear as a 3D box comprising an image of one side of the box being captured. In some embodiments, the saved scene icon may be one of the one or more saved scene icons. The saved scene icon may be presented within a user menu designed for saved scene selection.

[0143] Once scenes are saved, they may be opened by a user such that the virtual content of the saved scene appears in the user's augmented reality environment in which the scene is opened. In an exemplary user interface to the augmented reality system, a scene may be opened by selecting a saved scene icon in a manner that indicates selection of an icon to trigger opening the scene. The indication may be via user-initiation of a command such as a load icon, or may be inferred from context, for example. In some embodiments, the user may select a saved scene icon. The user may use any suitable user interaction such as a button click, gesture, and / or voice command to select the saved scene icon. In some embodiments, the system may automatically select a saved scene icon. For example, the system may automatically select a saved scene icon based on the user location. The system may automatically select a saved scene icon corresponding to a saved scene that was previously saved within the room where the user is currently located. In some embodiments, the system may automatically select a saved scene icon based on context. For example, if the user saved a scene at school, the system may automatically select the saved scene icon when the user is in any educational setting.

[0144] FIG. 12 includes steps for selecting and opening a saved scene. In step 1210, a saved scene icon is moved from its default location. In some embodiments, the saved scene icon is located within a saved scene user menu that may contain one or more saved scene icons representing one or more saved scenes. In some embodiments, the saved scene icon is not located within the saved scene user menu and instead is an isolated icon. For example, the saved scene icon may be automatically placed within the saved scene. As a specific example, the saved scene icon may be automatically placed at the saved PCF location or at the location where the user was when the scene was saved.

[0145] In step 1212, the saved scene icon may be installed in the user's environment. In some embodiments, the user may select the saved scene icon from the saved scene user menu 1208, for example, by pressing a button on the multi-DOF controller. By moving the controller, the user can then pull out the saved scene icon from the user menu 1210 and then install the saved scene icon 1212 by releasing the button on the multi-DOF controller. In some embodiments, the user may select the saved scene icon from its default saved PCF location 1208, for example, by pressing a button on the multi-DOF controller. The user can then move the saved scene icon out from the default saved PCF location 1210 and then install the saved scene icon 1212, for example, closer to the user's current location, by releasing the button on the multi-DOF controller. In some embodiments, the system may automatically install the saved scene icon 1212. For example, the system may automatically move the saved scene icon to a fixed distance from the user, or may automatically install the saved scene icon on the multi-DOF controller at the tip of a totem or the like. In some embodiments, the system may automatically install the saved scene icon at a fixed location relative to the user's hand when a specific gesture, such as a pinching or pointing gesture, is performed. Regardless of whether the installation is performed automatically by the system or by the user, any other suitable method may be used to install the saved scene icon 1212.

[0146] After step 1212, the saved scene icon may be instantiated into a copy of the original scene 1220, or the saved scene icon may be further manipulated by steps 1214 - 1218.

[0147] In step 1214, the user may select the position of the visual anchor node that indicates the location where the saved scene will be opened. Once the visual anchor node is selected, when the saved scene is opened, the system renders the virtual content of the saved scene with a saved scene anchor node that is aligned with the visual anchor node such that the virtual content has the same spatial relationship to the visual anchor node as it has to the saved scene anchor node. In some embodiments, the saved scene icon may indicate the location of the visual anchor node. In some embodiments, the saved scene icon may change its visual representation after being placed in 1212 and include a visual anchor node that was previously not visible to the user. In some embodiments, the saved scene icon may be replaced with a different visual representation of the visual anchor node. In some embodiments, the saved scene icon and the visual anchor node are the same. In some embodiments, the visual anchor node may be a separate icon from the saved scene icon that provides a visual representation of the saved scene anchor node.

[0148] The saved scene anchor node may be a root node relative to which all saved scene objects are positioned in order to maintain a consistent spatial relativity among the saved scene objects within the saved scene. In some embodiments, the saved scene anchor node is the top node within the saved scene hierarchy of nodes that represents at least a portion of the saved scene data. In some embodiments, the saved scene anchor node is an anchor node within a hierarchical structure that represents at least a portion of the saved scene data. In some embodiments, the saved scene anchor node may act as a reference point for positioning saved scene objects relative to each other. In some embodiments, the saved scene anchor node may represent a scene graph for the saved scene objects.

[0149] Regardless of how the visual anchor node appears to the user, the user may instruct the system to set the location of the visual anchor node through the user interface. In step 1216, the user may move the visual anchor node. In some embodiments, when the visual anchor node is moved, virtual scene objects move with the visual anchor node. In some embodiments, the visual anchor node is a visual representation of a saved scene anchor node. In some embodiments, the visual anchor node may provide a point and a coordinate system for manipulating the location and orientation of the saved scene anchor node using it. In some embodiments, moving the visual anchor node 1216 may mean translation, rotation, and / or 6DOF movement.

[0150] In step 1218, the user may place a visual anchor node. In some embodiments, the user may select the visual anchor 1214, for example, by pressing a button on the totem. The user may then move the visual anchor 1216 within the user's real-world environment and then place the visual anchor 1218, for example, by releasing a button on the totem. Moving the visual anchor moves the entire saved scene relative to the user's real world. Steps 1214 - 1218 may function to move all of the saved scene objects within the user's real world. Once the saved scene arrives at the desired location, the saved scene may be instantiated 1220. In step 1220, instantiating the saved scene may mean that a complete copy of the saved scene is presented to the user. However, it should be understood that the virtual content of the saved scene may be presented with the same physical and other characteristics as other virtual content rendered by the augmented reality system. Virtual content that is blocked by physical objects where the saved scene is opened may not be visible to the user. Similarly, when the position of the virtual content, as determined by its position relative to the visual scene anchor, may be outside the user's FOV, it may also not be visible when the scene is opened. The augmented reality system may still have information about this virtual content available for rendering it, which may occur when the user's pose or environment changes such that the virtual content becomes visible to the user.

[0151] After step 1220, the process may repeat starting from step 1208. This loop may enable the user to load more than one saved scene into the user's environment at a time.

[0152] In one exemplary embodiment, process 1200 begins with a user who has already assembled one or more component virtual object parts in a scene, such as constructing a replica of the user's house from component (e.g., pre-designed, pre-loaded) virtual building blocks provided as manipulable objects from an application. The user then selects the camera icon 1202, frames the scene 1204, which helps the user recall what the scene contains, and then the user presses a button to capture an image 1206 of the replica of the user's house. At this point, the replica of the user's house is saved in the system, and a saved scene icon corresponding to the replica of the user's house may be displayed to the user within the saved scene user menu. The user may turn the system off, go to bed that night, and then resume construction the next day. The user may select the saved scene icon 1208 corresponding to the replica of the user's house from the saved scene user menu by clicking a button on the totem, pull out the saved scene from the saved scene user menu 1210, then drag the saved scene icon to a desired location and release the button on the totem to place the saved scene icon 1212 in front of the user. A preview of the saved scene object may automatically appear to the user with a visual anchor node centered on the saved scene object. The preview may appear as a white-painted spatially correct visual-only copy of the saved scene. The user may decide to change the location where the saved scene is located based on the preview, and thus, by clicking a button on the totem, select the visual anchor node 1214 and move the totem to move the visual anchor node 1216 (which can move all of the saved scene object along with the visual anchor node and maintain the relative positioning inside between the saved scene objects with respect to the visual anchor node), and then release the button on the totem to place the visual anchor node 1218 at a different location within the user's environment.The user may then select the "Load Scene" virtual button from the user menu, which instantiates the saved scene 1220 and thus can fully load the virtual scene by rendering the complete saved scene data (e.g., an exact copy of the original scene potentially excluding different locations from where it was saved). (Process for saving a scene)

[0153] FIG. 13A illustrates an exemplary process 1300a for saving a scene using the systems and methods described herein. Process 1300a may begin with an application that is already open and that may be launched on an AR / VR / MR wearable system such as wearable system 200 and / or 900 described above, which may enable a user to place one or more pre-designed virtual objects into the user's real-world environment. The user's environment may already be meshed (e.g., a world mesh available to the application has already been created). In some embodiments, the world mesh may be an input to a map database, such as map database 710 from FIGS. 7 and / or 8. The world mesh may be combined with other world meshes from other users, or the same user from different sessions, and / or over time, to form a larger world mesh that is stored within the map database. The larger world mesh may be referred to as a traversable world and may include mesh data from the real world, in addition to object tags, location tags, and the like. In some embodiments, the meshed real-world environment available to the application may be of any size and may be determined by the processing capabilities of the wearable system (not exceeding the maximum allocated computational resources allocated to the application). In some embodiments, the meshed real-world environment (world mesh) available to the application may have an occupied area of 15 feet by 15 feet with a height of 10 feet. In some embodiments, the world mesh available to the application may be the size of the room or building in which the user is located. In some embodiments, the world mesh size and shape available to the application may be the first 300 square feet that is meshed by the wearable system when the application is first opened. In some embodiments, the mesh available to the application is the first area and volume that is meshed until a maximum threshold is met.Any other surface area or volume measurements may also be used in any quantity, as long as the wearable system has the computing resources available for it. The shape of the world mesh available to the application may be of any shape. In some embodiments, the surface area may be defined with respect to the world mesh available to the application, and may be in the shape of a square, rectangle, circle, polygon, etc., and may be a single continuous area, or (e.g., as long as the sum does not exceed a maximum surface area threshold) two or more discontinuous areas. In some embodiments, the volume may be defined with respect to the world mesh available to the application, and may be a cube, sphere, torus, cylinder, rectangular prism, cone, pyramid, prism, etc., and may be a continuous volume, or (e.g., as long as the sum does not exceed a maximum volume threshold) two or more discontinuous volumes.

[0154] Exemplary process 1300a may start at step 1202 when a camera icon is selected, as described with respect to FIG. 12. Step 1202 may cause the system to place a virtual rendering camera within virtual rendering scene 1304. The virtual rendering scene may be a digital representation of all renderable virtual content available to the user, created by manipulation of data on one or more processors (e.g., processing modules 260 or 270, or remote data repository 280, or wearable system 200, etc.). One or more virtual rendering cameras may be placed within the virtual rendering scene. The virtual rendering camera may function similarly to a real-world camera in that the virtual rendering camera has a location and orientation within the (virtual rendering world) space and can capture a 2D image of the 3D scene from that location and orientation.

[0155] The virtual rendering camera may act as an input for the rendering pipeline for the wearable system, and the location and orientation of the virtual rendering camera define that portion of the environment that includes virtual content existing in a part of the user's environment that will be rendered to the user as an indication of what the rendering camera is pointing at. The rendering pipeline may include one or more processes required to convert the virtual content for which the wearable system is ready to display digital data to the user. In some embodiments, the rendering for the wearable system may occur within a rendering engine that may be located within a graphics processing unit (GPU) that is part of and / or connected to processing module 160 and / or 270. In some embodiments, the rendering engine may be a software module within the GPU that may provide an image to display 220. In some embodiments, the virtual rendering camera may be placed within the virtual rendering scene in a position and orientation such that the virtual rendering camera has the same viewpoint and / or FOV as the default view through the virtual camera viewfinder as described in the context of FIG. 12.

[0156] In some embodiments, sometimes referred to as the "rendering camera", the "pinhole perspective camera" (or simply the "perspective camera"), or the "virtual pinhole camera", the "virtual rendering camera" is a simulated camera that may be used to render virtual image content from a database of objects within a virtual world. The objects may have locations and orientations relative to the user or wearer and potentially relative to actual objects in the environment surrounding the user or wearer. In other words, the rendering camera may represent a viewing point within the rendering space from which the user or wearer would view 3D virtual content (e.g., virtual objects) in the virtual rendering world space. The user may view the rendering camera's viewing point by viewing a 2D image captured from the perspective of the virtual rendering camera. The rendering camera may be managed by a rendering engine to render a virtual image based on a database of virtual objects to be presented to the eye. The virtual image may be rendered as if taken from the perspective of the user or wearer, from the perspective of a virtual camera that would frame the scene in step 1204, or from any other desired perspective. For example, the virtual image may be rendered as if captured by a pinhole camera (corresponding to the "rendering camera") having a specific set of intrinsic parameters (e.g., focal length, camera pixel size, principal point coordinates, skew / distortion parameters, etc.) and a specific set of extrinsic parameters (e.g., translation and rotation components relative to the virtual world). The virtual image is taken from the perspective of such a camera having the position and orientation of the rendering camera (e.g., the extrinsic parameters of the rendering camera).

[0157] The system may define and / or adjust intrinsic and extrinsic rendering camera parameters. For example, the system may define a particular set of extrinsic rendering camera parameters such that the virtual image is rendered as if captured from the perspective of a camera having a specific location with respect to the user's or wearer's eye, so as to appear as an image from the user's or wearer's perspective. The system may later dynamically adjust the extrinsic rendering camera parameters on-the-fly to maintain alignment with the specific location, for example, as used for eye tracking. Similarly, the intrinsic rendering camera parameters may be defined and dynamically adjusted over time. In some implementations, the image is rendered as if captured from the perspective of a camera having an aperture (e.g., a pinhole) at a specific location (such as the center of the viewpoint or the center of rotation, or other locations, etc.) with respect to the user's or wearer's eye.

[0158] Further details describing methods and systems related to rendering pipelines and rendering cameras are provided in U.S. Non-Provisional Patent Application No. 15 / 274,823, entitled "Methods and Systems for Detecting and Combining Structural Features in 3D Reconstruction", and U.S. Non-Provisional Patent Application No. 15 / 683,677, entitled "Virtual, augmented, and mixed reality systems and methods", which are hereby expressly incorporated by reference in their entirety.

[0159] In step 1306, the system may render the scene 1306 that is framed in step 1204, such that it is framed in the frame. In some embodiments, the system may render the scene framed in the frame at the normal refresh rate defined by the system when the location and / or orientation of the virtual camera changes during scene framing 1204, or at any other suitable time. In some embodiments, as the virtual camera moves during step 1204, the virtual rendering camera moves with a corresponding movement within the virtual rendering scene and maintains a viewpoint corresponding to the virtual camera viewpoint. In some embodiments, the virtual rendering camera may require that all renderable data available to the virtual rendering camera be sent to the rendering pipeline. The renderable data may be a set of data required by the wearable system to display virtual content to the user. The renderable data may be a set of data required by the wearable system to render virtual content. The renderable data may represent virtual content within the field of view of the virtual rendering camera. Alternatively, or in addition, the renderable data may represent objects in the physical world and may include data that is extracted from an image of the physical world obtained using the sensors of the wearable system and converted into a form that can be rendered.

[0160] For example, raw world mesh data that can be generated from images collected using a camera of a wearable augmented reality system may not be renderable. The raw world mesh data may be a set of vertices and thus may not be data that can be presented to a user as a 3D object or surface. The raw world mesh data in raw text format may include vertices with three points (x, y, and z) relative to the location of a virtual camera (and thus a virtual rendering camera since their viewpoints are synchronized). The raw world mesh data in raw text format may also include data representing other nodes or vertices to which each vertex is connected. Each vertex may be connected to one or more other vertices. In some embodiments, the world mesh data may be considered to be locations and data in space that define other nodes to which vertices are connected. The vertices and connected vertex data within the raw world mesh data may subsequently be used to construct polygons, surfaces, and thus a mesh, but additional data will be required to make the raw world mesh data renderable. For example, a shader, or any other program capable of performing the same function, may be used to visualize the raw world mesh data. The program may automatically follow a set of rules for computationally visualizing the raw world mesh data. The shader may be programmed to draw dots at each vertex location and then draw lines between each vertex and the set of vertices connected to that vertex. The shader may be programmed to visualize the data in one or more colors, patterns, etc. (e.g., blue, green, rainbow, checkerboard, etc.). In some embodiments, the shader is a program designed to illustrate the points and connections of the world mesh data. In some embodiments, the shader may connect the lines and dots to create a surface to which a texture or other visual representation may be added. In some embodiments, the renderable world mesh data may include UV data such as UV coordinates.In some embodiments, the renderable world mesh data may include depth checks to determine overlaps or obstacles between vertices in the raw world mesh data, or the depth checks may alternatively be performed within the rendering pipeline as a separate process. In some embodiments, applying shaders and / or other processes to the raw world mesh data enables the user to visualize and subsequently view what is typically non-visual data. In some embodiments, other raw data or conventionally non-visual data may be converted to renderable data using the processes described with respect to the raw world mesh data. For example, anything having a location in space may be visualized, e.g., by applying a shader. In some embodiments, a PCF may be visualized by providing data that defines an icon that can be positioned and oriented like a PCF.

[0161] Another example of renderable data (e.g., renderable 3D data, 3D renderable digital objects) is data related to 3D virtual objects such as characters in a video game, virtual avatars, or building blocks used to construct a replica of a user's house as described with respect to FIG. 12. Renderable data related to 3D virtual objects may include mesh data and mesh renderer data. In some embodiments, the mesh data may include one or more of vertex data, normal data, UV data, and / or triangle index data. In some embodiments, the mesh renderer data may include one or more texture data sets and one or more properties (e.g., material properties such as gloss, specular level, roughness, scatter color, ambient color, specular color, etc.).

[0162] In some embodiments, the virtual rendering camera may have settings that determine a subset of the renderable data to be rendered. For example, the virtual rendering camera may be capable of rendering three sets of renderable data, namely virtual objects, world meshes, and PCFs. The virtual camera may be set to render only virtual objects, only world meshes, only PCFs, or any combination of those three sets of renderable data. In some embodiments, any number of subsets of renderable data may exist. In some embodiments, setting the virtual rendering camera and using virtual object data to render world mesh data may enable a user to view conventional visual data (e.g., 3D virtual objects) overlaid on conventional non-visual data (e.g., raw world mesh data).

[0163] FIG. 13B illustrates an exemplary process 1300b for rendering a framed scene using the systems and methods described herein. Process 1300b may describe step 1306 in more detail. The process 1300b for rendering a framed scene may begin with step 1318 of requesting renderable data. In some embodiments, a virtual rendering camera, which may correspond to the virtual camera of FIG. 12, requests renderable data for all renderable objects located within the FOV of the virtual rendering camera within the virtual rendering scene. In some embodiments, the virtual rendering camera may request only the renderable data for an object that is programmed to request as the virtual rendering camera requests. For example, the virtual rendering camera may be programmed to request only the renderable data for a 3D virtual object. In some embodiments, the virtual rendering camera may request the renderable data for a 3D virtual object and the renderable data for world mesh data.

[0164] In step 1320, renderable data is sent to a rendering pipeline 1320. In some embodiments, the rendering pipeline may be a rendering pipeline for a wearable system such as the wearable system 200. In other embodiments, the rendering pipeline may be located on different devices, on different computers, or implemented remotely. In step 1322, a rendered scene is created. In some embodiments, the rendered scene may be the output of the rendering pipeline. In some embodiments, the rendered scene may be a 2D image of a 3D scene. In some embodiments, the rendered scene may be displayed to the user. In some embodiments, the rendered scene may comprise scene data that is ready to be displayed but is not actually being displayed.

[0165] In some embodiments, the system renders a scene framed and displays the rendered scene to the user through a virtual camera viewfinder. This may enable previewing a 2D image that would be captured if step 1206 of the user viewing and capturing an image of the framed scene were performed at that time, even as the virtual camera is moving. In some embodiments, the rendered scene may comprise visual and non-visual renderable data.

[0166] After step 1204, the user may select to capture an image 1206, or select cancel 1302 and cancel the scene saving process 1300a. In some embodiments, step 1302 may be performed by the user. In some embodiments, step 1302 may be automatically performed by the system. For example, if the system cannot automatically frame the scene in step 1204 as programmed (e.g., to capture all virtual objects in the frame), the system may automatically cancel the scene saving process 1300a in step 1302. If cancel is selected in step 1302, the system removes the virtual rendering camera corresponding to the virtual camera from the process 1200, e.g., from the virtual rendering scene 1308.

[0167] When an image is captured at step 1206, the system captures a 2D image 1310. In some embodiments, the system captures the 2D image by taking a photograph (storing data representing the rendered 2D image) using a virtual rendering camera within a virtual rendering scene. In some embodiments, the virtual rendering camera may function similarly to a real-world camera. The virtual rendering camera may convert 3D scene data into a 2D image by capturing the projection of the 3D scene onto the 2D image plane from the viewpoint of the virtual rendering camera. In some embodiments, the 3D scene data is captured as pixel information. In some embodiments, while the virtual rendering camera may be programmed to capture at any location between one subset and all subsets of the renderable data, since the real-world camera captures everything present in the view, the virtual rendering camera is not similar to the real-world camera. The 2D image captured at step 1310 may comprise only a subset of the renderable data that the camera is programmed to capture. The subset of the renderable data may comprise conventional visual data such as 3D objects and / or conventional non-visual data such as world meshes or PCFs.

[0168] In steps 1312 and 1314, the system saves the scene. The step of saving the scene may involve saving data representing the virtual content being rendered by the virtual rendering camera, and in some embodiments, any or all of the types of saved scene data as described above, such as location information. In step 1312, the system associates the scene with the nearest PCF. The application may send a request for the PCF ID to a lower-level system operation that manages a list of PCFs and their corresponding locations. The lower-level system operation may manage a map database that may include PCF data. In some embodiments, the PCFs are managed by a separate PCF application. The PCF associated with the scene may be referred to as the saved PCF.

[0169] In step 1314, the system writes the saved scene data to the persistent memory of the wearable system. In some embodiments, the persistent memory may be a hard drive. In some embodiments, the persistent memory may be the local processing and data module 260 as described above. In some embodiments, the saved scene data may comprise data that fully represents the saved scene. In some embodiments, the saved scene data may comprise data required by the system to render and display the saved scene to the user. In some embodiments, the saved scene data comprises the saved PCF, an image of the saved scene, and / or saved scene objects. The saved scene objects may be represented by saved scene object data. In some embodiments, the saved scene object data may comprise tags regarding the type of object. In embodiments where the saved scene object is derived by modifying a pre-designed base object, the saved scene object data may also indicate the difference between the saved scene object and the pre-designed base object. In some embodiments, the saved scene object data may comprise the pre-designed base object name with additional properties and / or status data added. In some embodiments, the saved scene object data may comprise a renderable mesh with added state, physical, and other properties that the saved scene object may be required to reload as a copy of the way it was saved.

[0170] In step 1316, the system may add the saved scene icon to the user menu. The saved scene icon may be a visual representation for the saved scene and optionally may comprise the 2D image captured in step 1310. The saved scene icon may be placed within a user menu, such as a saved scene user menu. The saved scene user menu may contain one or more saved scenes for the application. (Process for loading a saved scene)

[0171] Figure 14 illustrates an exemplary process 1400 for loading a scene using the systems and methods described herein. Process 1400 may begin at step 1402 where the user opens a user menu. The user menu may include one or more saved scene icons that may represent one or more saved scenes. In some embodiments, the user menu may be a saved scene user menu. In response to step 1402, the system may perform a PCF check 1422. The PCF check may include one or more processes to determine the PCF (current PCF) closest to the user's current location. In some embodiments, the application may determine the user's location. In some embodiments, the location may be based on the user's head pose location.

[0172] If the saved PCF matches the current PCF, the saved scene may be placed within the current PCF section of the user menu 1424. If the saved PCF does not match the current PCF, the saved scene may be placed within another PCF section within the user menu. In some embodiments, steps 1424 and 1426 may be combined, for example, if the user menu does not categorize saved scenes based on the user's current location or if the user's current PCF cannot be determined. At this point in process 1400, the user may view the saved scene user menu. In some embodiments, the user menu is separated into two sections: one section for saved scenes having a saved PCF that matches the current PCF and a second section for saved scenes having a saved PCF that does not match the current PCF. In some embodiments, one of the two sections may be empty.

[0173] In step 1404, the user may select a saved scene icon from the user menu. The user menu may include a saved scene user menu. In some embodiments, the user may select the saved scene icon through user interaction such as clicking a button on a totem or other user controller.

[0174] In step 1406, the user may take an action indicating that the content of the virtual content of the selected saved scene is to be loaded into the user's environment where the saved scene will be opened. For example, the user may remove the saved scene icon from the user menu 1406. In some embodiments, step 1406 may occur as the user continues to press the button used to select the saved scene icon. Step 1428 may result from step 1406. When the saved scene icon is removed from the user menu 1406, the system may load the saved scene (or saved scene data) from the hard drive or other persistent memory 1428 into volatile memory.

[0175] The location within the environment for the saved scene content may also be determined. In the illustrated embodiment, when the saved scene is opened at the same location where it was saved, the system may display the visual content of the saved scene at the same location that was present at the time the scene was saved. Alternatively, if the saved scene is opened at a different location, alternative approaches such as receiving user input as described below in connection with steps 416, 1418, and 1420 may be used. To support opening a saved scene with objects at the same location as when the scene was remembered, at step 1430, the system may perform a PCF check. The PCF check may include one or more processes that determine the PCF closest to the user's current location (the current PCF). In some embodiments or process 1400, either the PCF check 1422 or the PCF check 1430 may be performed instead of both. In some embodiments, additional PCF checks may be added to process 1400. The PCF check may occur at fixed time intervals (e.g., once per minute, once per second, once every 5 minutes, etc.) or may be based on a change in the user's location (e.g., if user movement is detected, the system may add an additional PCF check).

[0176] In step 1432, if the saved PCF matches the current PCF, the saved scene object is preview-installed with respect to the PCF. In some embodiments, preview-installing may include the step of rendering only the visual data associated with the saved scene data. In some embodiments, preview-installing may include the step of rendering the visual data associated with the saved scene data in combination with one or more shaders for modifying the appearance of the visual data and / or additional visual data. An example of additional visual data may be one or more lines extending from the visual anchor node to each of the saved scene objects. In some embodiments, the user may participate in performing step 1432, such as by providing an input indicating the location of the saved scene object. In some embodiments, the system may automatically perform step 1432. For example, the system may automatically perform step 1432 by calculating the location at the center of the saved scene object and then installing a visual anchor node at the center.

[0177] In step 1434, if the saved PCF does not match the current PCF, the saved scene object is preview installed relative to the visual anchor node. In some embodiments, the relative installation may be the step of installing the saved scene object such that the visual anchor node is at the center of the saved scene object. Alternatively, the saved scene objects may be positioned such that they have the same spatial relationship to the visual anchor node as they had to the scene anchor node where they were saved. In some embodiments, the installation may be determined by installing at a fixed distance away from the user (e.g., 2 feet away from the user in the z - direction at eye height) relative to the visual anchor node. In some embodiments, the user may be involved in performing step 1434, such as by providing an input indicating the location of the visual anchor node. In some embodiments, the system may perform step 1434 automatically. For example, the system may perform step 1434 automatically by automatically installing the visual anchor node at a fixed distance from the user menu, or may select the location of the visual anchor node relative to the location of physical or virtual objects in the user's environment where the saved scene is to be loaded.

[0178] In step 1436, the system may display a user prompt to the user and either cancel or instantiate the saved scene. The user prompt may have any suitable visual appearance and may function to enable at least one user interaction to cause the system to either cancel (e.g., prevent the process from proceeding to 1440) and / or instantiate the scene again. In some embodiments, the user prompt may display one or more interactive virtual objects, such as buttons labeled, for example, "Cancel" or "Load Scene".

[0179] In some embodiments, the system may display a saved scene preview as soon as the saved scene icon is removed from the user menu at step 1406, such that the user can view the preview in response to moving the icon. At step 1408, the user may release the saved scene icon and place the saved scene icon within the user's real-world environment. At step 1410, the user can view the saved scene preview and user prompt and can either cancel the loading of the saved scene or instantiate the scene. In some embodiments, the saved scene preview may be associated with the saved scene and optionally comprise only visual data with the visual data modified. In some embodiments, the visual data may appear painted white. In some embodiments, the saved scene preview may appear as a ghost preview of the visual data corresponding to the saved scene data. In some embodiments, the saved scene preview may appear as an empty copy of the data that looks recognizably similar to the saved scene but may not have the same functionality or audio. In some embodiments, the saved scene preview may be visual data corresponding to the saved scene data to which state data or physics has not been applied.

[0180] At step 1412, the user may select to cancel. Step 1412 may cause the system to stop displaying the content and remove the saved scene from the volatile memory 1438. In some embodiments, the system may only stop displaying the saved scene content but may keep the saved scene in the volatile memory. In some embodiments, the content may comprise all or part of the saved scene data, user menu, user prompt, or any other virtual content that may be specific to the saved scene selected at step 1404.

[0181] In step 1414, the user may choose to instantiate the saved scene. In some embodiments, this may be the same as step 1220. In some embodiments, step 1414 may fully load the virtual scene by rendering the complete saved scene data (e.g., an exact copy of the original scene, except potentially in a location different from where it was saved). In some embodiments, instantiation may include the step of applying the physical and state data to the visual preview.

[0182] In step 1416, the user may select a visual anchor node. In step 1418, the user may move the visual anchor node. This may cause the system to move the saved scene parent node location to match the visual anchor node location 1442. In some embodiments, the visual anchor node location may be moved without modifying any of the others in the saved scene data. This may be achieved by placing the saved scene objects relative to the visual anchor node in a manner that preserves the spatial relationship between the saved scene objects and the saved scene anchor, regardless of where the visual anchor may be located. In some embodiments, the process may loop back to step 1436 after step 1442.

[0183] In step 1420, the user may release the visual anchor node indicating the location of the visual anchor node. The process may loop back to step 1410 after step 1420, where the user again has the option to cancel the loading of the saved scene 1412, instantiate the saved scene 1414, or move the visual anchor node (and thus the entire saved scene objects that remain spatially coincident with each other) 1416 - 1420.

[0184] In some embodiments, the system may perform one or more of steps 1402 - 1420, described as involving user interaction with the system, either partially or fully automatically. In an exemplary embodiment of a system that performs steps 1402 - 1420 automatically, the system may automatically open user menu 1402 when the current PCF matches the saved PCF. The system may initiate the PCF check process for the entire time the application is running, or the system may refresh the PCF check at fixed intervals (e.g., every minute), or the system may initiate the PCF check when the system detects a change (e.g., user movement). In step 1402, the system may automatically select the saved scene if there is only one saved scene with a saved PCF that matches the current PCF. In step 1406, the system may automatically remove the saved scene icon from the user menu if the current PCF matches the saved PCF and has been saved for a threshold time period (e.g., 5 minutes). In step 1408, the system may automatically release the saved scene icon at a fixed distance from the user (e.g., 1 foot to the right of the user in the x - direction). In step 1412, the system may automatically disable process 1400 when the user exits the room, thus making the saved PCF no longer match the current PCF. In step 1414, the system may automatically instantiate the saved scene if the current PCF matches the saved PCF and has been saved for a threshold time period. The system may automatically perform steps 1416 - 1420 by moving the visual anchor node while the user is moving, maintaining a fixed relative spatial relationship to the user (e.g., the visual anchor node is fixed 2 feet in front of the user in the z - direction). (Process for Loading Saved Scenes - Shared Path)

[0185] Figure 15 illustrates an exemplary process 1500 for loading a scene using the systems and methods described herein. At step 1502, a saved scene icon may be selected. In some embodiments, the user may select a saved scene icon. In some embodiments, the saved scene icon may be selected from a user menu or as a stand-alone icon placed within the user's real-world environment. In some embodiments, the saved scene icon may be automatically selected by a wearable system and / or application. For example, the system may automatically select the saved scene icon that is closest to the user, or the system may automatically select the saved scene icon that is most frequently used regardless of location.

[0186] At step 1504, the saved scene icon is moved from its default location. In some embodiments, the user may remove the saved scene icon from the user menu as described in step 1406 in process 1400. In some embodiments, the system may automatically move the saved scene icon from its default position (e.g., within the user menu or at the installation location within the user's environment). For example, the system may automatically move the saved scene icon and maintain a fixed distance from the user.

[0187] The system may determine a location for a saved scene object within the user's environment where the saved scene is open. In some embodiments, the system may display a visual anchor node to the user and enable the user to input a command to move the location of the visual anchor node and / or influence the location where the saved scene object is placed by positioning the saved scene object relative to the visual anchor node. In some embodiments, each of the saved scene objects may have a position relative to a saved scene anchor node, and the saved scene objects may have the same relative positions with respect to the visual anchor node and thus may be positioned to position the saved scene anchor node relative to the visual anchor node.

[0188] In other embodiments, the user input may define a spatial relationship between one or more saved scene objects and the visual anchor node. In step 1506, the visual anchor node may be placed relative to the saved scene object. The step of placing the visual anchor node relative to the saved scene object may function to connect a particular location to the saved scene anchor. The choice of visual anchor location may affect, for example, the relative level of ease or difficulty with which the saved scene can be further manipulated, such as to later place a scene object in the environment. In some embodiments, the visual anchor node may be placed relative to the saved scene object by releasing a button on the totem (when the user presses a button, for example, to select and move a saved scene icon). In some embodiments, the user may select a location to place the visual anchor node in relation to the saved scene object. For example, the user may select to place the visual anchor node in close proximity to a particular saved scene object. The user may select to do this if the user is only concerned with the location where a particular object will be placed. In some embodiments, the user may desire to place the visual anchor node in a particular location that facilitates further manipulation of the visual anchor node.

[0189] In step 1508, the saved scene object may be preview placed with respect to the real world. In some embodiments, the saved scene object may be preview placed with respect to the real world by moving the visual anchor node. The step of moving the visual anchor node may move all of the saved scene objects with respect to the visual anchor node, and thus, maintain a fixed relative spatial configuration among the saved scene objects within the saved scene. In some embodiments, the step of moving the visual anchor node may change the anchor node location to match the current visual anchor node location. In some embodiments, the user may operate the visual anchor node to place the saved scene at a desired location with respect to the user's environment. In some embodiments, the system may automatically preview place the saved scene object with respect to the real world. For example, the objects of the saved scene may be positioned with respect to a surface at the time the saved scene was stored. In response to opening the saved scene, to position the saved scene object, the system may find the nearest surface having attributes and / or affordances similar to the surface on which the virtual object was positioned when saved. Examples of affordances may include surface orientation (e.g., vertical surface, horizontal surface), object type (e.g., table, bench, etc.), or relative height from the ground (e.g., low, medium, high height categories). Additional types of attributes or affordances may be used. In some embodiments, the system may automatically preview place the object in the real world on the next nearest meshed surface.

[0190] Affordance may have a relationship between an object and the object's environment that can provide an opportunity for an action or use associated with the object. Affordance may be determined, for example, based on the function, orientation, type, location, shape, or size of a virtual object or destination object. Affordance may also be based on the environment in which the virtual object or destination object is located. The affordance of a virtual object may be programmed as part of the virtual object and stored in the remote data repository 280. For example, a virtual object may be programmed to include a vector indicating the normal of the virtual object.

[0191] For example, the affordance of a virtual display screen (e.g., a virtual TV) is that the display screen can be viewed from a direction indicated by the normal to the screen. The affordance of a vertical wall is that an object can be installed on the wall (e.g., suspended on the wall) with their surfaces parallel to the normal to the wall. Additional affordances of the virtual display and the wall can be that each has a top and a bottom. The affordance associated with an object can help ensure more realistic interactions such as the object automatically suspending a virtual TV with the right side up on a vertical surface.

[0192] The automatic installation of virtual objects by a wearable system that utilizes affordance is described in U.S. Patent Publication No. 2018 / 0045963, published on February 15, 2018, which is incorporated herein by reference in its entirety.

[0193] In step 1510, the saved scene object for which preview is set may be moved relative to the real world. In some embodiments, this may be accomplished by moving the visual anchor node. For example, the user may select the visual anchor node by performing a pinching gesture, and may move the visual anchor node while maintaining the pinching gesture until the desired location is reached, and then the user may stop the movement and release the pinching gesture. In some embodiments, the system may automatically move the saved scene relative to the real world.

[0194] In step 1512, the saved scene may be instantiated. The step of instantiating the scene may include applying the complete saved scene data to the saved scene object, as opposed to only applying the preview version of the saved scene. For example, the preview of the saved scene may involve only displaying visual data or a modified version of the visual data. Step 1512 may instead display visual data such that, except for the saved scene potentially having a new anchor location, the rest of the data (e.g., physical) is added and saved to the saved scene data. (Process for loading a saved scene - Split path)

[0195] FIG. 16 illustrates an exemplary process 1600 for loading a scene using the systems and methods described herein.

[0196] In step 1602, the saved scene icon may be selected. In some embodiments, step 1602 may be the same as step 1502 and / or 1404. In some embodiments, the saved scene icon may be a visual representation of the saved scene and / or the saved scene data. In some embodiments, the saved scene icon may be the same visual representation as the visual anchor node. In some embodiments, the saved scene icon may include a visual anchor node.

[0197] In step 1604, the saved scene icon may be moved from its default location. In some embodiments, step 1604 may be the same as step 1504 and / or 1406. In some embodiments, step 1604 may include moving the totem around while continuously pressing a button on the totem. In some embodiments, step 1602 may include pressing and releasing a button on the totem to select an object, and step 1604 may include moving the totem around to cause a corresponding movement as a totem in the saved scene.

[0198] In some embodiments, the saved scene may be loaded at the same location where it was first saved (e.g., the saved scene objects are at the same real-world location where they were when the scene was first saved). For example, if the user creates a scene within the user's kitchen, the user may load the scene in the user's kitchen. In some embodiments, the saved scene may be loaded at a location different from where it was first saved. For example, if the user saves a scene at their friend's house but desires to continue interacting with the scene at home, the user may load the saved scene at the user's home. In some embodiments, this split process 1600 may correspond to the split between 1432 and 1434 in process 1400.

[0199] In step 1608, a visual anchor node may be placed with respect to the saved scene object. In some embodiments, this may occur when the saved scene PCF matches the user's current PCF (i.e., the scene is loaded at the same location where it was saved). For example, the application may be programmed to automatically place (e.g., preview place) the scene objects saved at the same real-world location where they were saved. In this case, the initial placement of the saved scene icon from its default location functions to place a visual anchor node with respect to the saved scene object that is already placed. In some embodiments, when the saved PCF matches the current PCF, the system may automatically place, e.g., preview place, the scene objects saved within the user's environment, and step 1608 may determine a location for the saved scene anchor. In some embodiments, once the visual anchor node is placed in 1608, the relative spatial location between the visual anchor node and the saved scene object may be fixed.

[0200] In step 1606, a visual anchor node may be placed with respect to the real world. In some embodiments, user input may be obtained at the location of the visual anchor node when the saved scene PCF does not match the user's current PCF and / or when the user's current PCF cannot be obtained. Alternatively, or in addition, the process may proceed to step 1608 where the system can receive an input and position a visual anchor node with respect to the saved scene object. For example, the application may be programmed to automatically place a visual anchor node with respect to the saved scene object. In some embodiments, the visual anchor node may be automatically placed at the center of the saved scene object (as its placement with respect to the saved scene object). Other suitable relative placement methods may also be used.

[0201] In step 1610, the saved scene may be moved with respect to the real world. By this step in process 1600, visual anchor nodes are placed with respect to the saved scene objects (step 1608 or 1606), and the saved scene objects are preview placed at the user's initial location within the real world (with respect to path 1608, the saved scene objects are automatically placed at the same real world location where they were saved, and with respect to path 1606, the saved scene objects are placed in a defined spatial configuration with respect to the visual anchor nodes). The saved scene may optionally be moved in step 1610. In some embodiments, step 1610 may be step 1510, 1416 - 1420, and / or 1210. In some embodiments, step 1610 is not performed and process 1600 proceeds directly to step 1612.

[0202] In step 1612, the saved scene may be instantiated. In some embodiments, step 1612 may be step 1512, 1414, and / or 1212. In some embodiments, since the application may already have loaded the scene saved during step 1602 into volatile memory, in step 1612, the scene may feed complete saved scene data into the rendering pipeline. The saved scene may then optionally be displayed to the user through a wearable device, such as wearable device 200, as an exact copy of the saved scene (e.g., the same relative spatial relationships among the saved scene objects), except at a different location within the user's real world.

[0203] In some embodiments, the wearable system and the application may be used synonymously. The application may be downloaded onto the wearable system and thus become part of the wearable system.

[0204] While opening the saved scene, an example of what can be seen by a user of the augmented reality system is provided by FIGS. 21A-C. In this example, a user interface is shown that allows the user to activate control and move the visual anchor node 2110. FIG. 21A illustrates the step of the user moving the visual anchor node 2110. In this example, the saved scene consists of a cube object that is visible in preview mode in FIG. 21A. In this example, the saved scene has a saved scene anchor node that coincides with the visual anchor node. The saved scene object has, in this example, a predetermined relationship to the saved scene anchor node and thus to the visual anchor node. That relationship is shown by the dotted line visible in FIG. 21A.

[0205] FIG. 21B illustrates the step of the user loading the saved scene with the visual anchor node to a desired location by activating the load icon 2120. In this example, the user selection is indicated by a line 2122 to the selected icon that mimics a laser pointer. In the augmented reality environment, that line can be manipulated by the user moving, pointing at, or otherwise in any other suitable way the totem. In this example, activating a control such as the load control can result from the user providing some other input such as pressing a button on the totem while the icon associated with the control is selected.

[0206] FIG. 21C illustrates virtual content 2130, here shown as a block, that is instantiated with respect to the saved scene anchor node that is aligned with the defined visual anchor node. In contrast to the preview mode of FIG. 21A, the saved scene object may be rendered with the full color, physical, and other attributes of the virtual object. (Exemplary Embodiment)

[0207] Concepts such as discussed in this specification, when executed by at least one processor, may be embodied as a non-transitory computer-readable medium encoded with computer-executable instructions to operate a type of mixed reality system that maintains an environment for a user, comprising virtual content configured to render to appear to the user in relation to the physical world to select a saved scene. Each saved scene comprises virtual content and a position relative to a saved scene anchor node, and determines whether a saved scene anchor node associated with a selected scene is associated with a location within the physical world. When a saved scene anchor node is associated with a location within the physical world, the mixed reality system may add virtual content to the environment at the location indicated by the saved scene anchor node. When a saved scene anchor node is not associated with a location within the physical world, the mixed reality system may determine a location within the environment and add virtual content to the environment at the determined location.

[0208] In some embodiments, when a saved scene anchor node is not associated with a location within the physical world, the step of determining a location within the environment includes rendering a visual anchor to the user and receiving user input indicating the position of the visual anchor.

[0209] In some embodiments, when a saved scene anchor node is not associated with a location within the physical world, the step of determining a location within the environment includes identifying a surface within the physical world based on an affordance and / or similarity of attributes to a physical surface associated with the saved scene to be selected, and determining a location relative to the identified surface.

[0210] In some embodiments, when a saved scene anchor node is not associated with a location within the physical world, the step of determining a location within the environment includes determining a location for the user.

[0211] In some embodiments, when the saved scene anchor node is not associated with a location in the physical world, the step of determining a location within the environment includes determining a location relative to the location of a virtual object within the environment.

[0212] In some embodiments, computer-executable instructions configured to select a saved scene may be configured to automatically select the saved scene based on a user's location in the physical world relative to the saved scene anchor node.

[0213] In some embodiments, computer-executable instructions configured to select a saved scene may be configured to select the saved scene based on user input.

[0214] Alternatively, or in addition, concepts as discussed herein, when executed by at least one processor, select a first pre-built virtual sub-component and a second pre-built virtual sub-component from a library, receive user input that defines the relative positions of the first pre-built virtual sub-component and the second pre-built virtual sub-component, and store saved scene data comprising the first pre-built virtual sub-component and the second pre-built virtual sub-component, and data that identifies the relative positions of the first pre-built virtual sub-component and the second pre-built virtual sub-component, store virtual content comprising at least the first pre-built virtual sub-component and the second pre-built virtual sub-component as a scene, render an icon representing the scene stored in the virtual user's menu that includes icons for a plurality of saved scenes, and may be embodied as a non-transitory computer-readable medium encoded with computer-executable instructions for operating a type of mixed reality system that maintains an environment for a user and is configured to render for the user to appear.

[0215] In some embodiments, the virtual content further comprises at least one constructed component.

[0216] In some embodiments, the virtual content further comprises at least one pre - saved scene. (Other considerations)

[0217] The processes, methods, and algorithms described herein and / or depicted in the accompanying figures are each embodied in one or more physical computing systems, hardware computer processors, application - specific circuits, and / or electronic hardware configured to execute specific and particular computer instructions, thereby being fully or partially automated. For example, a computing system can include a general - purpose computer (e.g., a server) or a special - purpose computer, a dedicated circuit, etc., programmed with specific computer instructions. The code modules can be installed in a dynamic - link library that can be compiled and linked into an executable program, or can be written in an interpreted - type programming language. In some implementations, certain operations and methods can be implemented by circuits specific to a given function.

[0218] Furthermore, the functional implementations of the present disclosure are sufficiently mathematically, computationally, or technically complex that a special - purpose hardware or one or more physical computing devices (utilizing appropriate specialized executable instructions) may be required to implement the functionality, for example, due to the amount or complexity of the calculations involved or to provide the results substantially in real - time. For example, a video can include many frames, each frame can have millions of pixels, and specifically programmed computer hardware is required to process the video data to provide the desired image - processing tasks or applications within a commercially reasonable amount of time.

[0219] A code module or any type of data can be stored on any type of non-transitory computer-readable medium, such as a physical computer storage device including a hard drive, solid state memory, random access memory (RAM), read only memory (ROM), optical disk, volatile or non-volatile storage device, combinations of the same, and / or equivalents. The methods and modules (or data) can also be transmitted as data signals generated on various computer-readable transmission media, including wireless-based and wire / cable-based media (e.g., as part of a carrier wave or other analog or digital propagated signal), and can take various forms (e.g., as part of a single or multiplexed analog signal or as a plurality of discrete digital packets or frames). The results of the disclosed process or process steps can be persistently or otherwise stored within any type of non-transitory tangible computer storage device or communicated via a computer-readable transmission medium.

[0220] Any process, block, state, step, or functionality in a flowchart described herein and / or depicted in the accompanying figures is to be understood as potentially representing a code module, segment, or portion of code that includes one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in a process. The various processes, blocks, states, steps, or functionality can be combined, rearranged, added to, removed from, modified, or otherwise changed from the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules can implement some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the associated blocks, steps, or states can be performed in other suitable sequences, e.g., sequentially, in parallel, or in some other manner. Tasks or events can be added to or removed from the disclosed illustrative embodiments. Further, the separation of the various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems can generally be integrated together in a single computer product or packaged into multiple computer products. Many implementation variations are possible.

[0221] The present process, method, and system may be implemented in a network (or distributed) computing environment. The network environment may include an enterprise-wide computer network, an intranet, a local area network (LAN), a wide area network (WAN), a personal area network (PAN), a cloud computing network, a cloud source computing network, the Internet, and the World Wide Web. The network may be a wired or wireless network or any other type of communication network.

[0222] The systems and methods of the present disclosure each have several innovative aspects, none of which alone contribute to or are required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure. Various modifications to the implementations described in this disclosure may be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations without departing from the spirit or scope of the present disclosure. Accordingly, the claims are not intended to be limited to the implementations shown herein but are to be accorded the widest scope consistent with the disclosure, principles, and novel features disclosed herein.

[0223] In the context of separate implementations, certain features described herein can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented separately in multiple implementations or in any suitable sub-combination. Further, features may be described above as acting in a certain combination and, further, may be initially claimed as such, but one or more features from the claimed combination can, in some cases, be deleted from the combination and the claimed combination can be directed to a sub-combination or a variation of a sub-combination. No single feature or group of features is necessary or essential to every embodiment.

[0224] In particular, conditional clauses used herein such as "can", "could", "might", "may", "e.g.", and equivalents, generally convey that one embodiment includes certain features, elements, and / or steps while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context in which they are used. Thus, such conditional clauses are not generally intended to imply that the features, elements, and / or steps are required in any way for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps should be included or implemented in any particular embodiment, regardless of the author's input or prompting. The terms "comprising", "including", "having", and equivalents are synonyms and are used inclusively in a non-limiting manner, without excluding additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not in its exclusive sense), and thus, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Additionally, the articles "a", "an", and "the" as used in this application and the appended claims should be construed to mean "one or more" or "at least one" unless otherwise defined.

[0225] As used herein, the phrase referring to a list of items "at least one of" refers to any combination of those items, including a single element. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Connective phrases such as "at least one of X, Y, and Z" are generally understood in a context such that, unless otherwise specifically stated, they are used to convey that an item, term, etc. can be at least one of X, Y, or Z. Thus, such connective phrases are generally not intended to suggest that an embodiment requires that at least one of X, at least one of Y, and at least one of Z each be present.

[0226] Similarly, operations may be depicted in the drawings in a particular order, but it should be recognized that this is not required for achieving the desired result, and that such operations may be performed in the particular order shown, or in a sequential order, or that all of the illustrated operations need not be performed. Additionally, the drawings may schematically depict one or more exemplary processes in the form of flowcharts. However, other operations not depicted may also be incorporated into the exemplary methods and processes schematically illustrated. For example, one or more additional operations may be performed before, after, simultaneously with, or during any of the illustrated operations. In addition, operations may be rearranged or reordered in other implementations. In some situations, multitasking and parallel processing may be advantageous. Further, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result.

Claims

1. A method for operating a mixed reality system, the method comprising using at least one processor to receive an input from a user selecting a stored scene from a stored scene library, the stored scene comprising virtual content and position information indicating the position of the virtual content relative to a stored scene anchor node in a first coordinate frame representing a first location within the physical world, determine that the input from the user selecting the stored scene is received at a second location different from the first location, the determining comprising determining a second coordinate frame based on the user's current environment, comparing the second coordinate frame with the first coordinate frame representing the first location, and determining that the second coordinate frame is different from the first coordinate frame when it is determined that the second coordinate frame does not match the first coordinate frame including, in response to determining that the input from the user selecting the stored scene is received at the second location different from the first location, determining a location for the stored scene anchor node relative to the second location, controlling a display to render the virtual content at the determined location for the stored scene anchor node including performing, a method.

2. using the at least one processor to determine that the input from the user selecting the stored scene is received at the first location, and controlling the display to render the virtual content at the first location for the stored scene anchor node further including performing, the method according to claim 1.

3. The mixed reality system identifies one or more coordinate frames based on objects within the physical world, the method further including using the at least one processor to place the stored scene anchor node in one of the one or more identified coordinate frames, the method according to claim 1.

4. The method according to claim 1, further comprising using the at least one processor to read the virtual content from a remote server via a network interface. **Claim 5** The method according to claim 1, further comprising using the at least one processor to control the display and render a menu comprising a plurality of icons representing saved scenes in the saved scene library, wherein receiving the input from the user to select the saved scene comprises user selection of an icon among the plurality of icons. **Claim 6** The method according to claim 1, further comprising using the at least one processor to render a virtual user interface, wherein receiving the input from the user comprises receiving user input via the virtual user interface. **Claim 7** The method according to claim 1, further comprising using the at least one processor to render a virtual user interface comprising a menu of a plurality of icons representing saved scenes in the saved scene library, wherein receiving the input from the user to select a saved scene in the saved scene library comprises receiving user input via the virtual user interface, and the user input comprises selecting and moving an icon within the menu. **Claim 8** A mixed reality system, the mixed reality system comprising a display, at least one sensor configured to obtain information about the environment of the mixed reality system, at least one processor, and a non-transitory computer-readable medium encoded with computer-executable instructions wherein the computer-executable instructions, when executed by the at least one processor, receive an input from a user to select a saved scene in a saved scene library, the saved scene comprising virtual content and location information, the location information indicating a location of the virtual content relative to a saved scene anchor node at a first location in the physical world, virtual content and location information, and a set of one or more points in the physical world representing the first location in the physical world and ​ Determining that the input from the user selecting the saved scene was received at a second location different from the first location, the determining comprising: Obtaining data from the at least one sensor, and Using the data from the at least one sensor to determine that the one or more points representing the first location in the physical world do not exist in the user's current environment Including, In response to determining that the input from the user selecting the saved scene was received at the second location different from the first location, Determining a location for the saved scene anchor node for the second location, and Controlling the display to render the virtual content for the determined location for the saved scene anchor node Performing A composite reality system that performs.

9. The computer-executable instructions further comprise: Determining that the input from the user selecting the saved scene was received at the first location, and Controlling the display to render the virtual content for the first location for the saved scene anchor node The composite reality system according to claim 8, further configured to perform.

10. The computer-executable instructions further comprise: Identifying one or more coordinate frames based on the obtained information about the environment of the composite reality system, and Installing the saved scene anchor node in one of the one or more identified coordinate frames The composite reality system according to claim 8, further configured to perform.

11. The system comprises a network interface, The computer-executable instructions are configured to read the virtual content from a remote server via the network interface, the composite reality system according to claim 8.

12. The system comprises a network interface, The computer-executable instructions are configured to control the display to render a menu comprising a plurality of icons representing saved scenes in the saved scene library. Receiving input from the user to select the saved scene includes user selection of an icon among the plurality of icons, the composite reality system according to claim 8.

13. The computer-executable instructions are further configured to render a virtual user interface when executed by the at least one processor, The computer-executable instructions configured to receive input from the user are configured to receive user input via the virtual user interface, the composite reality system according to claim 8.

14. The computer-executable instructions are further configured to render a virtual user interface comprising a menu of a plurality of icons representing saved scenes in the saved scene library when executed by the at least one processor, Receiving input from the user to select a saved scene in the saved scene library includes receiving user input via the virtual user interface, and the user input selects and moves an icon within the menu, the composite reality system according to claim 8.

15. A non-transitory computer-readable medium encoded with computer-executable instructions, wherein the computer-executable instructions, when executed by at least one processor, execute a method of operating a composite reality system, and the method uses the at least one processor, Receiving input from a user to select a saved scene in a saved scene library, where the saved scene is Virtual content and position information, where the position information indicates the position of the virtual content relative to a saved scene anchor node at a first location in the physical world, virtual content and position information, and A saved three-dimensional (3D) mesh representing the first location in the physical world, where the saved 3D mesh is generated from a depth map of the first location, the saved 3D mesh Comprising, Determining that the input from the user to select the saved scene is received at a second location different from the first location, and the determining is Accessing the stored 3D mesh representing the stored scene, and Determining that the stored 3D mesh is different from a 3D mesh generated from a depth map of the user's current environment within the physical world Including, In response to determining that the input from the user selecting the stored scene was received at a second location different from the first location, Determining a location for the stored scene anchor node for the second location, and Controlling a display to render the virtual content at the determined location for the stored scene anchor node Performing Including performing, a non-transitory computer-readable medium.

16. The computer-executable instructions Determining that the input from the user selecting the stored scene was received at the first location, and Controlling the display to render the virtual content at the first location for the stored scene anchor node The non-transitory computer-readable medium according to claim 15, further configured to perform.

17. The mixed reality system includes at least one sensor configured to obtain information about the environment of the mixed reality system, The method Determining one or more coordinate frames based on the obtained information about the environment of the mixed reality system, and Installing the stored scene anchor node in one of the one or more identified coordinate frames The non-transitory computer-readable medium according to claim 15, further comprising.

18. The system includes a network interface, The computer-executable instructions are configured to control the display to render a menu comprising a plurality of icons representing stored scenes in the stored scene library, Receiving input from the user selecting the stored scene includes user selection of an icon among the plurality of icons, the non-transitory computer-readable medium according to claim 15.

19. The computer-executable instructions are further configured to render a virtual user interface when executed by the at least one processor, The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions configured to receive input from the user are configured to receive user input via the virtual user interface.

20. When executed by the at least one processor, the computer-executable instructions are further configured to render a virtual user interface comprising a menu of a plurality of icons representing saved scenes in the saved scene library. Receiving input from the user to select a saved scene in the saved scene library includes receiving user input via the virtual user interface, and the user input selects and moves an icon within the menu. The non-transitory computer-readable medium of claim 15.

21. The first coordinate frame comprises one or more feature points in the physical world at the first location, and the second coordinate frame comprises one or more feature points in the user's current environment. Comparing the second coordinate frame with the first coordinate frame includes determining whether the one or more feature points of the first coordinate frame represent the same location in the physical world as the one or more feature points of the second coordinate frame. The method of claim 1.

22. In response to determining that the input from the user to select the saved scene was received at a second location different from the first location. The method of claim 1, further comprising rendering a preview of the virtual content at a default location.

23. Generating a menu within the display, the menu comprising: A first section for indicating saved scenes that match the second coordinate frame; A second section for indicating saved scenes that do not match the second coordinate frame, the second section including the saved scenes; and Including; Receiving the input from the user through the menu; The method of claim 1, further comprising.

Citation Information

Patent Citations

  • Systems and methods for augmented and virtual reality

    JP2016522463A

  • Automatic placement of a virtual object in a three-dimensional space

    WO2018031621A1

  • Automatic control of wearable display device based on external conditions

    WO2018125428A1