Automatic placement of virtual objects in 3D space

The AR system addresses the challenge of repositioning virtual objects in AR and MR by simulating physical forces and environmental interactions, enhancing user comfort and efficiency in object placement and alignment.

JP7777572B2Active Publication Date: 2025-11-28MAGIC LEAP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023212872
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-08-11
Filing Date
2023-12-18
Publication Date
2025-11-28
Estimated Expiration
2037-08-09

AI Technical Summary

Technical Problem

Existing augmented reality (AR) and mixed reality (MR) systems face challenges in automatically repositioning virtual objects in three-dimensional space, leading to user confusion and fatigue due to optical illusions and the difficulty in aligning virtual objects with destination objects, which requires precise and time-consuming user interactions.

Method used

An AR system that automatically repositions virtual objects by calculating trajectories, simulating physical forces like gravity and magnetism to align objects, and adjusting orientations based on environmental affordances, allowing users to intuitively place and unbind virtual objects using user inputs and postures.

Benefits of technology

Enhances user comfort by reducing cognitive and physical fatigue through seamless object placement and alignment, providing a more natural and efficient interaction experience in AR and MR environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007777572000001
    Figure 0007777572000001
  • Figure 0007777572000002
    Figure 0007777572000002
  • Figure 0007777572000003
    Figure 0007777572000003
Patent Text Reader

Abstract

To provide an augmented reality system and a method for automatically repositioning a virtual object in a three-dimensional (3D) space with respect to a destination object in a three-dimensional environment of a user.SOLUTION: A method automatically attaches a target virtual object to a destination object and re-orient the target virtual object based on affordance of the virtual object or the destination object. The method also tracks movement of a user and detaches the virtual object from the destination object when the movement of the user exceeds a threshold condition.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 62 / 373,693, filed August 11, 2016, entitled "AUTOMATIC PLACEMENT OF VIRTUAL OBJECTS IN A 3D ENVIRONMENT," and U.S. Provisional Application No. 62 / 373,692, filed August 11, 2016, entitled "VIRTUAL OBJECT USER INTERFACE WITH GRAVITY," the disclosures of which are incorporated herein by reference in their entireties.

[0002] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly to automatically repositioning virtual objects in three-dimensional (3D) space. [Background technology]

[0003] Modern computing and display technologies have facilitated the development of systems for so-called “virtual reality,” “augmented reality,” or “mixed reality” experiences in which digitally reproduced images, or portions thereof, are presented to a user in a manner that appears or can be perceived as real. Virtual reality or “VR” scenarios typically involve the presentation of digital or virtual image information without transparency to other actual real-world visual inputs. Augmented reality or “AR” scenarios typically involve the presentation of digital or virtual image information as an augmentation to the visualization of the real world around the user. Mixed reality or “MR” relates to the merging of real and virtual worlds to generate new environments in which physical and virtual objects coexist and interact in real time. Consequently, the human visual perception system is highly complex, making it challenging to produce VR, AR, or MR technologies that facilitate a comfortable, natural-feeling, and rich presentation of virtual image elements among other virtual or real-world image elements. The systems and methods disclosed herein address various challenges associated with VR, AR, and MR technologies. Summary of the Invention [Means for solving the problem]

[0004] Various embodiments of an augmented reality system for automatically repositioning virtual objects are described.

[0005] In one exemplary embodiment, an augmented reality (AR) system for automatically repositioning virtual objects within a three-dimensional (3D) environment is disclosed. The AR system includes an AR display configured to present virtual content within a 3D view and a hardware processor in communication with the AR display. The hardware processor is programmed to: identify a target virtual object within the user's 3D environment, the target virtual object being assigned a vector representing a first location and a first orientation; receive an indication to link the target virtual object to a destination object, the destination object being assigned at least one vector representing a second location and a second orientation; calculate a trajectory between the target virtual object and the destination object based at least in part on the first location and the second location; move the target virtual object toward the destination object along the trajectory; track a current location of the target virtual object; calculate a distance between the target virtual object and the destination object based at least in part on the current location and the second location of the target virtual object; determine whether the distance between the target virtual object and the destination virtual object is less than a threshold distance;

[0006] In another exemplary embodiment, a method for automatically repositioning a virtual object within a three-dimensional (3D) environment is disclosed. The method may be implemented under the control of an augmented reality (AR) system including computer hardware, the AR system configured to enable user interaction with the object within the 3D environment. The method includes identifying a target virtual object within a user's 3D environment, the target virtual object having a first position and a first orientation; receiving an indication to reposition the target virtual object relative to a destination object; identifying parameters for repositioning the target virtual object; analyzing affordances associated with at least one of the 3D environment, the target virtual object, and the destination object; calculating values ​​of the parameters for repositioning the target virtual object based on the affordances; determining a second position and second orientation for the target virtual object and a movement of the target virtual object based on the values ​​of the parameters for repositioning the target virtual object; and rendering the target virtual object at the second position and second orientation and the movement of the target virtual object to reach the second position and second orientation from the first position and first orientation.

[0007] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Neither this summary nor the following detailed description purports to define or limit the scope of the inventive subject matter. The present invention provides, for example, the following. (Item 1) 1. An augmented reality (AR) system for automatically repositioning a virtual object within a three-dimensional (3D) environment, the AR system comprising: an AR display configured to present virtual content; a hardware processor in communication with the AR display, the hardware processor comprising: identifying a target virtual object within a user's 3D environment, the target virtual object being assigned a vector representing a first location and a first orientation; receiving an indication to link the target virtual object to a destination object, the destination object being assigned at least one vector representing a second location and a second orientation; calculating a trajectory between the target virtual object and the destination object based at least in part on the first location and the second location; moving the target virtual object along the trajectory toward the destination object; tracking a current location of the target virtual object; calculating a distance between the target virtual object and the destination object based at least in part on the current location of the target virtual object and the second location; determining whether a distance between the target virtual object and a destination virtual object is less than a threshold distance; automatically linking the target virtual object to the destination object and orienting the target virtual object in the second orientation in response to comparing the distance to be less than or equal to the threshold distance; rendering, by the AR display, the target virtual object at the second location with the second orientation, wherein the target virtual object is overlaid on the destination object; and a hardware processor programmed to perform the An AR system equipped with: (Item 2) The hardware processor further comprises: and programmed to analyze affordances of at least one of the target virtual object, the destination object, or the environment; To automatically orient the target virtual object, the hardware processor is programmed to rotate the target virtual object to align a first normal of the target virtual object with a second normal of the destination object. Item 1. The AR system according to item 1. (Item 3) 3. The AR system of claim 2, wherein the affordance includes at least one of function, orientation, type, location, shape, or size. (Item 4) Item 10. The AR system of item 1, wherein to automatically bind the target virtual object, the hardware processor is programmed to simulate an attractive force between the target virtual object and the destination object, the attractive force including at least one of gravity, elastic force, adhesive force, or magnetic force. (Item 5) Item 1. The AR system of item 1, wherein, to calculate the distance, the hardware processor is programmed to calculate a displacement between a current location of the target virtual object and the second location associated with the destination object. (Item 6) Item 6. The AR system of item 5, wherein the threshold distance is zero. (Item 7) Item 1. The AR system of item 1, wherein the indication for linking the target virtual object is determined from at least one of an actuation of a user input device or a user posture. (Item 8) Item 8. The AR system of item 7, wherein the hardware processor is further programmed to assign a focus indicator to a current position of the user, the current position of the user being determined at least in part based on a posture of the user or a position associated with the user input device. (Item 9) The hardware processor further comprises: receiving an indication to uncouple the target virtual object from the destination object, the indication being associated with a change in a user's current location; determining, based at least in part on the received indication, whether a threshold condition for disengaging the target virtual object is met; and In response to determining that the threshold condition is met, unbinding the target virtual object from the destination object; moving the target virtual object from the second location associated with the destination object to a third location; rendering the target virtual object at the third location; and Item 1. The AR system according to item 1, programmed to: (Item 10) Item 10. The AR system of item 9, wherein in response to determining that the threshold condition is satisfied, the hardware processor is further programmed to move the target virtual object to the third location while retaining the second orientation for the target virtual object. (Item 11) 10. The AR system of claim 9, wherein the third location corresponds to a position of the focus indicator that corresponds to a current position of the user. (Item 12) The threshold condition for disassociating the target virtual object from the other object is: a second distance between the second location where the target virtual object is tied to the destination object and the position of the focus indicator is greater than or equal to a second threshold distance; a speed for moving the focus indicator from the second location to the position of the focus indicator is equal to or greater than a threshold speed; the acceleration causing movement away from the second location is greater than or equal to a threshold acceleration; or a jerk for moving away from the second location that is equal to or greater than a threshold jerk; Item 12. The AR system according to item 11, comprising at least one of: (Item 13) Item 14. The AR system of item 9, wherein the hardware processor is programmed to simulate physical forces when coupling and uncoupling the target virtual object from the destination object, the physical forces including at least one of friction or elasticity. 1. A method for automatically repositioning a virtual object within a three-dimensional (3D) environment, the method comprising: an augmented reality (AR) system configured to enable user interaction with objects in a 3D environment under control of the AR system, the AR system comprising computer hardware; identifying a target virtual object within a user's 3D environment, the target virtual object having a first position and a first orientation; receiving an indication to reposition the target virtual object relative to a destination object; identifying parameters for repositioning the target virtual object; analyzing affordances associated with at least one of the 3D environment, the target virtual object, and the destination object; calculating a value of a parameter for repositioning the target virtual object based on the affordance; determining a second position and a second orientation for the target virtual object and a movement of the target virtual object based on values ​​of parameters for repositioning the target virtual object; rendering the target virtual object at the second position and the second orientation and movement of the target virtual object from the first position and the first orientation to reach the second position and the second orientation; A method comprising: (Item 15) Item 15. The method of item 14, wherein repositioning the target object includes at least one of linking the target object to the destination object, reorienting the target object, or unlinking the target object from the destination object. (Item 16) Item 15. The method of item 14, wherein the destination object is a physical object. (Item 17) Item 15. The method of item 14, further comprising: determining whether the indication to reposition the target virtual object satisfies a threshold condition; and performing the calculating, determining, and rendering in response to determining that the indication satisfies the threshold condition. (Item 18) Item 18. The method of item 17, wherein the threshold condition includes a distance between the target virtual object and the destination object. (Item 19) Item 15. The method of item 14, wherein one or more physical attributes are assigned to the target virtual object, and movement of the target virtual object is determined by simulating an interaction of the target virtual object, the destination object, and the environment based on the physical attributes of the target virtual object. (Item 20) 20. The method of claim 19, wherein the one or more physical attributes assigned to the target virtual object include at least one of mass, size, density, topology, hardness, elasticity, or electromagnetic attributes. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 depicts an illustration of a mixed reality scenario with a virtual reality object and a physical object viewed by a person. [Figure 2] FIG. 2 illustrates diagrammatically an example of a wearable system. [Figure 3]FIG. 3 diagrammatically illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes. [Figure 4] FIG. 4 illustrates diagrammatically an embodiment of a waveguide stack for outputting image information to a user. [Figure 5] FIG. 5 shows an exemplary output beam that may be output by a waveguide. [Figure 6] FIG. 6 is a schematic diagram showing an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal stereoscopic display, image, or light field. [Figure 7] FIG. 7 is a block diagram of an embodiment of a wearable system. [Figure 8] FIG. 8 is a process flow diagram of an embodiment of a method for rendering virtual content in relation to recognized objects. [Figure 9] FIG. 9 is a block diagram of another embodiment of a wearable system. [Figure 10] FIG. 10 is a process flow diagram of an example method for determining user input to a wearable system. [Figure 11] FIG. 11 is a process flow diagram of an embodiment of a method for interacting with a virtual user interface. [Figure 12] 12A and 12B illustrate an example of automatically binding a virtual object to a table. [Figure 13] 13A, 13B, 13C, and 13D illustrate an example of automatically orienting a virtual object when a portion of the virtual object touches a wall. [Figure 14A] 14A, 14B, 14C, and 14D illustrate an example of moving a virtual object from a table to a wall. [Figure 14B] 14A, 14B, 14C, and 14D illustrate an example of moving a virtual object from a table to a wall. [Figure 14C]14A, 14B, 14C, and 14D illustrate an example of moving a virtual object from a table to a wall. [Figure 14D] 14A, 14B, 14C, and 14D illustrate an example of moving a virtual object from a table to a wall. [Figure 15A] 15A, 15B, and 15C illustrate an example of connecting and orienting virtual objects from a side view. [Figure 15B] 15A, 15B, and 15C illustrate an example of connecting and orienting virtual objects from a side view. [Figure 15C] 15A, 15B, and 15C illustrate an example of connecting and orienting virtual objects from a side view. [Figure 15D] 15D and 15E illustrate examples of tying and untying a virtual object from a wall. [Figure 15E] 15D and 15E illustrate examples of tying and untying a virtual object from a wall. [Figure 15F] 15F, 15G, and 15H illustrate additional examples of connecting and orienting virtual objects from a side view. [Figure 15G] 15F, 15G, and 15H illustrate additional examples of connecting and orienting virtual objects from a side view. [Figure 15H] 15F, 15G, and 15H illustrate additional examples of connecting and orienting virtual objects from a side view. [Figure 16] FIG. 16 is an exemplary method for connecting and orienting virtual objects. [Figure 17] FIG. 17 is an exemplary method for associating and unassociating a virtual object with another object in a user's environment.

[0009] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010] (overview) In an AR / MR environment, a user may wish to reposition a virtual object by changing its position or orientation. As an example, a user can move a virtual object in three-dimensional (3D) space and connect the virtual object to a physical object in the user's environment. The virtual object may be a two-dimensional (2D) or 3D object. For example, the virtual object may be a flat surface, a 2D television display, or a 3D virtual coffee pot. The user can move the virtual object along a trajectory and connect the virtual object to a physical object by using a user input device (e.g., a totem, etc.) and / or by changing the user's posture. For example, the user may move a virtual television (TV) screen from a table to a wall by moving the user input device. Similarly, an AR system may allow a user to select a virtual object and move it with their head posture. As the user moves their head, the virtual object also moves and is positioned and oriented accordingly.

[0011] However, moving a virtual object in 3D space can sometimes be problematic for a user because the movement can create optical illusions that can confuse the user about its current location. For example, a user may be confused about whether an object is moving away from or toward them. These optical illusions can cause cognitive fatigue when a user interacts with an AR system.

[0012] Furthermore, when a user attempts to place a virtual object on the surface of or inside a destination object, the user often needs to perform refined movements to orient and position the virtual object in multiple directions in 3D space and align the virtual object with the destination object. For example, when a user moves a virtual TV screen from a table to a wall, the user may need to orient the virtual screen so that the surface normal of the TV screen faces the user (e.g., the content displayed by the TV screen faces the user instead of the wall). The user may also orient the virtual screen so that the user does not need to turn their head when viewing the virtual TV screen. In addition, to make the virtual TV screen appear on the wall (rather than appearing to be built into the wall), the user may need to make fine adjustments to the position of the virtual TV screen. These operations may be time-consuming and difficult for a user to perform with precision and may cause physical fatigue to the user.

[0013] To solve some or all of these problems, an AR system can be configured to automatically reposition a target virtual object by changing the position or orientation of the target virtual object. As an example, the AR system can orient the target virtual object and connect the target virtual object to the destination object when the distance between the virtual object and the target object is less than a threshold distance. The AR system can also automatically reposition the target virtual object by moving the virtual object as if it were subject to a physical force (e.g., a spring force such as Hooke's Law, gravity, adhesive force, electromagnetic force, etc.). For example, when the virtual object and the target object are within a threshold distance, the AR system may automatically "snap" the virtual object onto the target object as if the virtual object and the target object were attracted together due to an attractive force (e.g., mimicking a magnetic force or gravity). Thus, the AR system may apply a virtual force between the objects, which simulates or acts like a physical force between the objects. In many cases, the virtual (or simulated physical) force may be an attractive force, but this is not a limitation, and in other cases, the virtual (or simulated physical) force may be a reactive force that tends to move objects away from one another. A repulsive virtual force may be advantageous when placing a target virtual object such that other nearby virtual objects are repelled (at least slightly) from the target object, thereby moving slightly and providing room between the other nearby objects for the placement of the target virtual object.

[0014] The AR system may further orient the virtual object and align the virtual object's surface normal with the user's gaze direction. As an example, the virtual object may initially float within the user's environment. The user may indicate an intent to move the virtual object onto a horizontal surface, such as the surface of a table or a floor (e.g., via a body gesture or activation of a user input device). The AR system may simulate the effects of gravity and automatically drop the virtual object onto the horizontal surface without additional user effort once the virtual object is sufficiently close to the horizontal surface.

[0015] In some situations, a user may wish to unbind a virtual object from an object to which it is bound. The AR system may simulate an attractive force between the virtual object and the object (e.g., simulating how a magnet may stick to a magnetic surface such as a refrigerator, or how a book rests on a horizontal table) so that the user may not be able to immediately unbind the virtual object from the object unless the user provides sufficient indication that the virtual object should be unbindable. For example, the user may "grasp" the virtual object with their hand or a virtual indicator and "jerk" the object (e.g., by a sufficiently rapid change in the position of the user's hand or the position of the virtual indicator). The indication to unbind the virtual object may be indicated by movement exceeding a threshold condition (such as when movement exceeds a threshold distance, a threshold speed, a threshold acceleration, or a threshold rate of change in acceleration, a combination thereof, or the like). This may be advantageous, particularly because it reduces the likelihood that the user will accidentally unbind the virtual object while interacting with it. As an example, while a user is playing a game using a virtual screen tied to a wall, the user may need to move their totem to find or interact with friends or foes. This type of game movement may correspond to the type of movement for untethering a virtual object from a wall. By untethering a virtual object from a wall only if the user's movement is sufficiently above a suitable threshold, the virtual screen will not be inadvertently untethered during game play. In addition, users typically cannot keep their posture or user input device still for an extended period of time. As a result, a virtual object may be accidentally untethered by a slight movement of the user when the user does not intend to untether the virtual object.Thus, by only untying a virtual object if the user's movement is sufficiently above a suitable threshold, slight movements or jerks by the user will not inadvertently untying the virtual object from its intended location or orientation. (Example of a 3D display for a wearable system)

[0016] A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present 2D or 3D virtual images to a user. The images may be still images, frames of video, or videos, in combination or the like. A wearable system can include a wearable device that can present a VR, AR, or MR environment, alone or in combination, for user interaction. The wearable device can be a head-mounted device (HMD), which is used synonymously with AR device (ARD). Additionally, for purposes of this disclosure, the term "AR" is used synonymously with the term "MR."

[0017] Figure 1 depicts an illustration of a mixed reality scenario involving a virtual reality object and a physical object viewed by a person. In Figure 1, an MR scene 100 is depicted in which a user of the MR technology sees a real-world park-like setting 110 featuring people, trees, a building in the background, and a concrete platform 120. In addition to these items, the user of the MR technology also perceives as "seeing" a robotic figure 130 standing on the real-world platform 120 and a flying, cartoon-like avatar character 140 that appears to be an anthropomorphic bumblebee, although these elements do not exist in the real world.

[0018] In order for a 3D display to produce a true depth sensation, and more specifically, a simulated sensation of surface depth, it may be desirable for the display to generate, for each point in its field of view, an accommodation response that corresponds to that point's virtual depth. If the accommodation response to a display point does not correspond to that point's virtual depth as determined by convergence and stereoscopic binocular depth cues, the human eye may experience accommodation conflict, resulting in unstable imaging, adverse eye strain, headaches, and, in the absence of accommodative information, a near-complete lack of surface depth.

[0019] VR, AR, and MR experiences can be provided by a display system having a display that provides a viewer with images corresponding to multiple depth planes. The images may be different for each depth plane (e.g., providing slightly different presentations of a scene or object) and may be focused separately by the viewer's eyes, thereby serving to provide depth cues to the user based on the ocular accommodation required to focus on different image features of a scene located on different depth planes, or based on observing different image features on different depth planes that are out of focus. As discussed elsewhere herein, such depth cues provide a believable perception of depth.

[0020] FIG. 2 illustrates an example of a wearable system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems to support the functionality of the display 220. The display 220 may be coupled to a frame 230, which is wearable by a user, wearer, or viewer 210. The display 220 can be positioned directly in front of the eyes of the user 210. The display 220 can present AR / VR / MR content to the user. The display 220 can comprise a head-mounted display worn on the user's head. In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the user's ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable sound control). The display 220 can include an audio sensor (e.g., a microphone) to detect an audio stream from the environment in which speech recognition is to be performed.

[0021] The wearable system 200 may include an outward-facing imaging system 464 (shown in FIG. 4 ) that observes the world in the user's surrounding environment. The wearable system 200 may also include an inward-facing imaging system 462 (shown in FIG. 4 ) that can track the user's eye movements. The inward-facing imaging system may track either one eye's movements or both eyes' movements. The inward-facing imaging system 462 may be mounted to the frame 230 and may be in electrical communication with a processing module 260 or 270 that may process image information obtained by the inward-facing imaging system and determine, for example, the pupil diameter or orientation of the user's 210 eyes, eye movements, or eye posture.

[0022] As an example, the wearable system 200 can obtain an image of the user's posture using an outward-facing imaging system 464 or an inward-facing imaging system 462. The image may be a still image, a frame of video or video, a combination thereof, or the like.

[0023] The display 220 is operably coupled (250) to a local data processing module 260, which may be mounted in a variety of configurations, such as fixedly attached to the frame 230, by wired or wireless connection, fixedly attached to a helmet or hat worn by the user, built into headphones, or otherwise removably attached to the user 210 (e.g., in a backpack-style configuration, in a belt-coupled configuration).

[0024] The local processing and data module 260 may comprise a hardware processor and digital memory such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storing data. The data may include data (a) captured from sensors (e.g., which may be operatively coupled to the frame 230 or otherwise attached to the user 210), such as an image capture device (e.g., a camera in an inward-facing and / or outward-facing imaging system), audio sensors (e.g., a microphone), an inertial measurement unit (IMU), an accelerometer, a compass, a global positioning system (GPS) unit, a wireless device, or a gyroscope, or data (b) obtained or processed using the remote processing module 270 or remote data repository 280, possibly for processing or readout and subsequent passage to the display 220. The local processing and data module 260 may be operably coupled to a remote processing module 270 or a remote data repository 280 over a communication link 262 or 264, such as via a wired or wireless communication link, such that these remote modules are available as resources to the local processing and data module 260. In addition, the remote processing module 280 and the remote data repository 280 may be operably coupled to each other.

[0025] In some embodiments, remote processing module 270 may comprise one or more processors configured to analyze and process data or image information. In some embodiments, remote data repository 280 may comprise a digital data storage facility, which may be available through the Internet or other networking configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, allowing for fully autonomous use from the remote module.

[0026] The human visual system is complex and difficult to provide a realistic perception of depth. Without being limited by theory, it is believed that viewers of an object may perceive the object as three-dimensional due to a combination of vergence and accommodation. The vergence of the two eyes relative to one another (i.e., the rotation of the pupils so that they move toward or away from one another, converging the eyes' lines of sight, and fixating on an object) is closely linked to the focusing of the eye's lenses (or "accommodation"). Under normal conditions, a change in the focus of the eye's lenses or the eye's accommodation to change focus from one object to another at a different distance will automatically produce a change in vergence at the same distance, a relationship known as the "accommodation-vergence reflex." Similarly, a change in vergence will induce a change in accommodation under normal conditions. Display systems that provide better matching between accommodation and convergence-divergence movements may produce more realistic and comfortable simulations of three-dimensional images.

[0027] FIG. 3 illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes. With reference to FIG. 3 , objects at various distances from the eyes 302 and 304 on the z-axis are accommodated by the eyes 302 and 304 such that the objects are in focus. The eyes 302 and 304 assume particular accommodated states, focusing objects at different distances along the z-axis. As a result, a particular accommodated state may be said to be associated with a particular one of the depth planes 306 having an associated focal length such that an object or portion of an object at a particular depth plane is in focus when the eye is in an accommodated state relative to that depth plane. In some embodiments, a three-dimensional image may be simulated by providing different representations of an image for each of the eyes 302 and 304, and by providing different representations of an image corresponding to each of the depth planes. While shown as separate for clarity of illustration, it should be understood that the fields of view of the eyes 302 and 304 may overlap, for example, as the distance along the z-axis increases. Additionally, while shown as flat for ease of illustration, it should be understood that the contours of a depth plane may be curved in physical space so that all features within the depth plane are in focus with the eye in a particular state of accommodation. Without being limited by theory, it is believed that the human eye is typically capable of interpreting a finite number of depth planes to provide depth perception. As a result, a highly realistic simulation of perceived depth may be achieved by providing the eye with different presentations of images corresponding to each of these limited number of depth planes. (Waveguide stack assembly)

[0028] FIG. 4 illustrates an example of a waveguide stack for outputting image information to a user. Wearable system 400 includes a stack of waveguides or stacked waveguide assembly 480 that can be utilized to provide three-dimensional perception to the eye / brain using multiple waveguides 432b, 434b, 436b, 438b, 4400b. In some embodiments, wearable system 400 may correspond to wearable system 200 of FIG. 2, and FIG. 4 schematically illustrates several portions of wearable system 200 in more detail. For example, in some embodiments, waveguide assembly 480 may be integrated into display 220 of FIG. 2.

[0029] 4, the waveguide assembly 480 may also include multiple features 458, 456, 454, 452 between the waveguides. In some embodiments, the features 458, 456, 454, 452 may be lenses. In other embodiments, the features 458, 456, 454, 452 may not be lenses. Rather, they may simply be spacers (e.g., cladding layers or structures to form air gaps).

[0030] Waveguides 432b, 434b, 436b, 438b, 440b or multiple lenses 458, 456, 454, 452 may be configured to transmit image information to the eye using various levels of wavefront curvature or ray divergence. Each waveguide level may be associated with a particular depth plane and configured to output image information corresponding to that depth plane. Image injection devices 420, 422, 424, 426, 428 may be utilized to inject image information into waveguides 440b, 438b, 436b, 434b, 432b, respectively, which may be configured to disperse incident light across each individual waveguide for output toward the eye 410. Light exits the output surfaces of image injection devices 420, 422, 424, 426, 428 and is injected into the corresponding input edges of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, a single beam of light (e.g., a collimated beam) may be injected into each waveguide, outputting an entire field of cloned collimated beams directed toward eye 410 at a particular angle (and divergence) corresponding to the depth plane associated with the particular waveguide.

[0031] In some embodiments, image input devices 420, 422, 424, 426, 428 are discrete displays that generate image information for input into each corresponding waveguide 440b, 438b, 436b, 434b, 432b, respectively. In some other embodiments, image input devices 420, 422, 424, 426, 428 are outputs of a single multiplexed display that may, for example, send image information to each of image input devices 420, 422, 424, 426, 428 via one or more optical conduits (such as fiber optic cables).

[0032] A controller 460 controls the operation of stacked waveguide assembly 480 and image injection devices 420, 422, 424, 426, 428. Controller 460 includes programming (e.g., instructions in a non-transitory computer-readable medium) that coordinates the timing and provision of image information to waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, controller 460 may be a single integrated device or a distributed system connected by a wired or wireless communication channel. Controller 460 may, in some embodiments, be part of processing module 260 or 270 (shown in FIG. 2).

[0033] Waveguides 440b, 438b, 436b, 434b, 432b may be configured to propagate light within each individual waveguide by total internal reflection (TIR). Waveguides 440b, 438b, 436b, 434b, 432b may each be planar or have another shape (e.g., curved) with major top and bottom surfaces and edges extending between their major top and bottom surfaces. In the illustrated configuration, waveguides 440b, 438b, 436b, 434b, 432b may each include light extraction optical elements 440a, 438a, 436a, 434a, 432a configured to extract light from the waveguides by redirecting the light to propagate within each individual waveguide and outputting image information from the waveguides to the eye 410. The extracted light may also be referred to as out-coupled light, and the light extraction optical element may also be referred to as out-coupling optical element. The extracted light beam is output by the waveguide where the light propagating within the waveguide strikes the light redirecting element. The light extraction optical element (440a, 438a, 436a, 434a, 432a) may be, for example, a reflective or diffractive optical feature. While shown disposed on the bottom major surfaces of the waveguides 440b, 438b, 436b, 434b, 432b for ease of explanation and clarity of the drawings, in some embodiments, the light extraction optical element 440a, 438a, 436a, 434a, 432a may be disposed on the top or bottom major surfaces or directly within the volume of the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed in a layer of material that is attached to a transparent substrate and forms the waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, the waveguides 440b, 438b, 436b, 434b, 432b may be a monolithic piece of material, and the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed on and / or within that piece of material.

[0034] Continuing with reference to FIG. 4, as discussed herein, each waveguide 440b, 438b, 436b, 434b, 432b is configured to output light and form an image corresponding to a particular depth plane. For example, the waveguide 432b closest to the eye may be configured to deliver collimated light to the eye 410 as it is launched into such waveguide 432b. The collimated light may represent an optical infinity focal plane. The next upper waveguide 434b may be configured to send collimated light that passes through a first lens 452 (e.g., a negative lens) before reaching the eye 410. The first lens 452 may be configured to create a slight convex wavefront curvature so that the eye / brain interprets light emerging from the next upper waveguide 434b as emerging from a first focal plane closer inward from optical infinity toward the eye 410. Similarly, the third upper waveguide 436b passes its output light through both the first lens 452 and the second lens 454 before reaching the eye 410. The combined refractive power of the first and second lenses 452 and 454 may be configured to produce another, increasing amount of wavefront curvature so that the eye / brain interprets the light emerging from the third waveguide 436b as emerging from a second focal plane even closer inward toward the person from optical infinity, which was the light from the next upper waveguide 434b.

[0035] Other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are similarly configured, with the highest waveguide 440b in the stack sending its output through all of the lenses between it and the eye for a collective focal power representing the focal plane closest to the person. To compensate for the stack of lenses 458, 456, 454, 452 when viewing / interpreting light originating from the world 470 on the other side of the stacked waveguide assembly 480, a compensating lens layer 430 may be placed on top of the stack to compensate for the collective power of the lower lens stacks 458, 456, 454, 452. Such a configuration provides as many perceived focal planes as there are available waveguide / lens pairs. Both the light extraction optical elements of the waveguides and the focusing sides of the lenses may be static (e.g., not dynamic or electro-active). In some alternative embodiments, one or both may be dynamic using electro-active features.

[0036] Continuing with reference to FIG. 4, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be configured to both redirect light from its respective waveguide and output this light with the appropriate amount of divergence or collimation for the particular depth plane associated with the waveguide. As a result, waveguides with different associated depth planes may have differently configured light extraction optical elements that output light with different amounts of divergence depending on the associated depth plane. In some embodiments, as discussed herein, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be solid or surface features that can be configured to output light at specific angles. For example, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume holograms, surface holograms, and / or diffraction gratings. Light extraction optical elements such as diffraction gratings are described in U.S. Patent Publication No. 2015 / 0178939, published June 25, 2015, which is incorporated herein by reference in its entirety.

[0037] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a are diffractive features, i.e., "diffractive optical elements" (also referred to herein as "DOEs"), that form a diffraction pattern. Preferably, the DOEs have a relatively low diffraction efficiency so that only a portion of the light in the beam is deflected toward the eye 410 with each intersection point of the DOE, while the remainder continues traveling through the waveguide via total internal reflection. The light carrying the image information is thus split into several related output beams that exit the waveguide at multiple locations, which can result in a very uniform pattern of output emission toward the eye 304 for this particular collimated beam bouncing within the waveguide.

[0038] In some embodiments, one or more DOEs may be switchable between an "on" state in which they actively diffract and an "off" state in which they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer-dispersed liquid crystal in which microdroplets comprise a diffractive pattern in a host medium, and the refractive index of the microdroplets can be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract incident light), or the microdroplets can be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts incident light).

[0039] In some embodiments, the number and distribution of depth planes or depths of field may be dynamically varied based on the size or orientation of the viewer's pupil. The depth of field may vary inversely with the viewer's pupil size. As a result, as the size of the viewer's pupil decreases, the depth of field increases so that a plane that is indistinguishable because its location exceeds the eye's depth of focus may become distinguishable and appear more focused with a corresponding decrease in pupil size and an increase in depth of field. Similarly, the number of spaced depth planes used to present different images to the viewer may be reduced with a decreased pupil size. For example, a viewer may not be able to clearly perceive details in both a first depth plane and a second depth plane at one pupil size without adjusting their eye's accommodation from one depth plane to the other. However, these two depth planes may be sufficient to simultaneously focus on the user at another pupil size without changing accommodation.

[0040] In some embodiments, the display system may vary the number of waveguides receiving image information based on a determination of pupil size and / or orientation, or in response to receiving an electrical signal indicating a particular pupil size and / or orientation. For example, if a user's eye is unable to distinguish between two depth planes associated with two waveguides, controller 460 (which may be an embodiment of local processing and data module 260) may be configured or programmed to stop providing image information to one of those waveguides. Advantageously, this may reduce the processing burden on the system, thereby increasing system responsiveness. In embodiments in which the DOE for a waveguide is switchable between on and off states, the DOE may be switched to the off state when the waveguide receives image information.

[0041] In some embodiments, it may be desirable to have the output beam satisfy the condition of having a diameter less than the diameter of the viewer's eye. However, meeting this condition may be difficult in light of the variability in the size of the viewer's pupil. In some embodiments, this condition is met over a wide range of pupil sizes by varying the size of the output beam in response to a determination of the size of the viewer's pupil. For example, as the pupil size decreases, the size of the output beam may also decrease. In some embodiments, the output beam size may be varied using a variable aperture.

[0042] The wearable system 400 may include an outward-facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 may be referred to as the world camera's field of view (FOV), and the imaging system 464 is sometimes referred to as an FOV camera. The entire area available for viewing or imaging by a viewer may be referred to as the ocular field of view (FOR). The FOR may include a solid angle of 4π steradians surrounding the wearable system 400 as the wearer moves their body, head, or eyes to perceive virtually any direction in space. In other contexts, the wearer's movement may be more constrained, and accordingly, the wearer's FOR may subtend a smaller solid angle. Images obtained from the outward-facing imaging system 464 can be used to track gestures (e.g., hand or finger gestures) made by the user, detect objects in the world 470 in front of the user, etc.

[0043] The wearable system 400 may also include an inward-facing imaging system 466 (e.g., a digital camera) that observes user movements, such as eye and facial movements. The inward-facing imaging system 466 may be used to capture images of the eyes 410 and determine the size or orientation of the pupils of the eyes 304. The inward-facing imaging system 466 may be used to obtain images for use in determining the direction the user is looking (e.g., eye pose) or for biometric identification of the user (e.g., via iris identification). In some embodiments, at least one camera may be utilized for each eye independently to separately determine the pupil size or eye pose of each eye, thereby allowing the presentation of image information to each eye to be dynamically adjusted for that eye. In some other embodiments, the pupil diameter or orientation of only a single eye 410 (e.g., using only a single camera per pair of eyes) is determined and assumed to be similar for both eyes of the user. Images obtained by inward-facing imaging system 466 may be analyzed to determine the user's eye posture or mood, which may be used by wearable system 400 to determine audio or visual content to be presented to the user. Wearable system 400 may also determine head pose (e.g., head position or head orientation) using sensors such as an IMU, accelerometer, gyroscope, etc.

[0044] The wearable system 400 may include a user input device 466 through which a user may input commands into the controller 460 and interact with the wearable system 400. For example, the user input device 466 may include a trackpad, touchscreen, joystick, multi-degree-of-freedom (DOF) controller, capacitive sensing device, game controller, keyboard, mouse, directional pad (D-pad), wand, tactile device, totem (e.g., functioning as a virtual user input device), etc. A multi-DOF controller may sense user input in possible translation (e.g., left / right, forward / backward, or up / down) or rotation (e.g., yaw, pitch, or roll) of some or all of the controller. A multi-DOF controller that supports translation may be referred to as 3DOF, while a multi-DOF controller that supports translation and rotation may be referred to as 6DOF. In some cases, a user may use a finger (e.g., a thumb) to press or swipe across a touch-sensitive input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input device 466 may be held by the user's hand during use of the wearable system 400. The user input device 466 may communicate with the wearable system 400 via wired or wireless communication.

[0045] 5 shows an example of an output beam output by a waveguide. While one waveguide is shown, it should be understood that other waveguides in waveguide assembly 480 may function similarly, and that waveguide assembly 480 includes multiple waveguides. Light 520 is launched into waveguide 432b at input edge 432c of waveguide 432b and propagates within waveguide 432b by TIR. At the point where light 520 impinges on DOE 432a, a portion of the light exits the waveguide as output beam 510. While output beams 510 are shown as approximately parallel, they may also be redirected to propagate to eye 410 at an angle (e.g., divergent output beam formation) depending on the depth plane associated with waveguide 432b. It should be understood that a nearly collimated exit beam may refer to a waveguide with light extraction optics that outcouples light and forms an image that appears to be set at a depth plane at a long distance (e.g., optical infinity) from the eye 410. Other waveguides or other sets of light extraction optics may output a more divergent exit beam pattern, which would require the eye 410 to accommodate to a closer distance and focus on the retina, and would be interpreted by the brain as light from a distance closer to the eye 410 than optical infinity.

[0046] FIG. 6 is a schematic diagram illustrating an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal volumetric display, image, or light field. The optical system can include a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem. The optical system can be used to generate a multifocal volumetric display, image, or light field. The optical system can include one or more primary planar waveguides 632a (only one is shown in FIG. 6) and one or more DOEs 632b associated with each of at least some of the primary waveguides 632a. The planar waveguides 632b can be similar to the waveguides 432b, 434b, 436b, 438b, and 440b discussed with reference to FIG. 4. The optical system may employ a dispersive waveguide device to relay light along a first axis (the vertical or Y-axis in the illustration of FIG. 6 ) and expand the effective exit pupil of the light along the first axis (e.g., the Y-axis). The dispersive waveguide device may include, for example, a dispersive planar waveguide 622 b and at least one DOE 622 a (illustrated by a double-dashed line) associated with the dispersive planar waveguide 622 b. The dispersive planar waveguide 622 b may be similar or identical in at least some respects to a primary planar waveguide 632 b having a different orientation therefrom. Similarly, the at least one DOE 622 a may be similar or identical in at least some respects to the DOE 632 a. For example, the dispersive planar waveguide 622 b or the DOE 622 a may be made of the same material as the primary planar waveguide 632 b or the DOE 632 a, respectively. The embodiment of the optical display system 600 shown in FIG. 6 can be integrated into the wearable system 200 shown in FIG.

[0047] The relayed, exit-pupil-expanded light can be optically coupled from the dispersive waveguide device into one or more primary planar waveguides 632b. The primary planar waveguides 632b can relay the light along a second axis, preferably orthogonal to the first axis (e.g., the horizontal or X-axis in the diagram of FIG. 6). Notably, the second axis can be non-orthogonal to the first axis. The primary planar waveguides 632b expand the effective exit pupil of the light along that second axis (e.g., the X-axis). For example, the dispersive planar waveguide 622b can relay and expand the light along the vertical or Y-axis and pass the light to a primary planar waveguide 632b, which can relay and expand the light along the horizontal or X-axis.

[0048] The optical system may include one or more colored light sources (e.g., red, green, and blue laser light) 610, which may be optically coupled into the proximal end of a single-mode optical fiber 640. The distal end of the optical fiber 640 may be threaded or received through a hollow tube 642 of piezoelectric material. The distal end protrudes from the tube 642 as a free-standing, flexible cantilever 644. The piezoelectric tube 642 may be associated with four quadrant electrodes (not shown). The electrodes may be plated, for example, on the outside, outer surface or outer periphery, or diameter of the tube 642. A core electrode (not shown) may also be located in the core, center, inner periphery, or inner diameter of the tube 642.

[0049] For example, drive electronics 650, electrically coupled via wires 660, drive opposing pairs of electrodes to bend piezoelectric tube 642 independently in two axes. The protruding distal tip of optical fiber 644 has a mechanical resonant mode. The frequency of the resonance may depend on the diameter, length, and material properties of optical fiber 644. By oscillating piezoelectric tube 642 near the first mechanical resonant mode of fiber cantilever 644, fiber cantilever 644 may be caused to oscillate and sweep through a large deflection.

[0050] By stimulating resonant vibrations in two axes, the tip of fiber cantilever 644 is scanned biaxially within an area filling a two-dimensional (2-D) scan. By modulating the intensity of light source 610 synchronously with the scanning of fiber cantilever 644, light emitted from fiber cantilever 644 can form an image. A description of such a setup is provided in U.S. Patent Publication No. 2014 / 0003762, which is incorporated herein by reference in its entirety.

[0051] Components of the optical coupler subsystem can collimate light emitted from the scanning fiber cantilever 644. The collimated light can be reflected by a mirrored surface 648 into a narrow dispersive planar waveguide 622b containing at least one diffractive optical element (DOE) 622a. The collimated light can propagate perpendicularly (with respect to the view of FIG. 6) along the dispersive planar waveguide 622b via TIR, and in doing so, repeatedly intersect with the DOE 622a. The DOE 622a preferably has a low diffraction efficiency. This causes a portion of the light (e.g., 10%) to diffract toward the edge of the larger primary planar waveguide 632b at each point of intersection with the DOE 622a, allowing a portion of the light to continue on its original trajectory down the length of the dispersive planar waveguide 622b via TIR.

[0052] At each point of intersection with the DOE 622a, additional light can be diffracted toward the entrance of the primary waveguide 632b. By splitting the incident light into multiple outcoupled sets, the exit pupil of the light can be vertically expanded by the DOE 622a within the dispersive planar waveguide 622b. This vertically expanded light outcoupled from the dispersive planar waveguide 622b can enter the edge of the primary planar waveguide 632b.

[0053] Light entering the primary waveguide 632b can propagate horizontally (with respect to the illustration of FIG. 6) along the primary waveguide 632b via TIR. The light propagates horizontally along at least a portion of the length of the primary waveguide 632b via TIR as it intersects the DOE 632a at multiple points. The DOE 632a advantageously has a phase profile that is the sum of a linear diffraction pattern and a radially symmetric diffraction pattern, and may be designed or configured to produce both deflection and focusing of the light. The DOE 632a advantageously may have a low diffraction efficiency (e.g., 10%) so that only a portion of the light in the beam is deflected toward the viewer's eye at each intersection of the DOE 632a, while the remainder of the light continues to propagate through the primary waveguide 632b via TIR.

[0054] At each point of intersection between the propagating light and the DOE 632a, a portion of the light is diffracted toward the adjacent face of the primary waveguide 632b, allowing the light to escape the TIR and emerge from the face of the primary waveguide 632b. In some embodiments, the radially symmetric diffraction pattern of the DOE 632a additionally imparts a level of focus to the diffracted light, both shaping (e.g., imparting curvature) the optical wavefronts of the individual beams and steering the beams to angles that match the designed level of focus.

[0055] Thus, these different paths can couple light out of the primary planar waveguide 632b by resulting in different fill patterns at the DOE 632a's multiplicity, focal level, or exit pupil at different angles. Different fill patterns at the exit pupil can be advantageously used to generate light field displays with multiple depth planes. Each layer in the waveguide assembly or set of layers (e.g., three layers) in the stack may be employed to generate distinct colors (e.g., red, blue, and green). Thus, for example, a first set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a first focal depth. A second set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a second focal depth. Multiple sets may be employed to generate full 3D or 4D color image light fields with various focal depths. (Other components of the wearable system)

[0056] In many implementations, the wearable system may include other components in addition to or as an alternative to the components of the wearable system described above. The wearable system may include, for example, one or more tactile devices or components. The tactile device or component may be operable to provide a haptic sensation to the user. For example, the tactile device or component may provide a sensation of pressure and / or texture upon touching virtual content (e.g., a virtual object, virtual tool, other virtual structure). The haptic sensation may replicate the sensation of a physical object represented by the virtual object, or may replicate the sensation of an imaginary object or character (e.g., a dragon) represented by the virtual content. In some implementations, the tactile device or component may be worn by the user (e.g., a user-wearable glove). In some implementations, the tactile device or component may be held by the user.

[0057] A wearable system may include, for example, one or more physical objects that can be manipulated by a user to enable input to or interaction with the wearable system. These physical objects may be referred to herein as totems. Some totems may take the form of inanimate objects, such as, for example, a piece of metal or plastic, a wall, the surface of a table, etc. In some implementations, a totem may not actually have any physical input structures (e.g., keys, triggers, joysticks, trackballs, rocker switches). Instead, the totem may simply provide a physical surface, and the wearable system may render a user interface to appear to the user on one or more surfaces of the totem. For example, the wearable system may render an image of a computer keyboard and trackpad to appear to reside on one or more surfaces of the totem. For example, the wearable system may render a virtual computer keyboard and virtual trackpad to appear on the surface of a thin rectangular plate of aluminum that serves as the totem. The rectangular plate itself does not have any physical keys or trackpads or sensors. However, the wearable system may detect user manipulation or interaction or touch with the rectangular plate as a selection or input made via a virtual keyboard or virtual trackpad. User input device 466 (shown in FIG. 4) may be an embodiment of a totem, which may include a trackpad, touchpad, trigger, joystick, trackball, rocker or virtual switch, mouse, keyboard, multi-degree-of-freedom controller, or another physical input device. A user may use the totem alone or in combination with posture to interact with the wearable system and / or other users.

[0058] Examples of tactile devices and totems usable with the wearable devices, HMDs, and display systems of the present disclosure are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. Exemplary Wearable Systems, Environments, and Interfaces

[0059] The wearable system may employ various mapping-related techniques to achieve a high depth of field within the rendered light field. When mapping a virtual world, it is advantageous to capture all features and points in the real world and accurately depict virtual objects in relation to the real world. To achieve this goal, FOV images captured from a user of the wearable system can be added to the world model by including new photos that convey information about various points and features in the real world. For example, the wearable system can collect a set of map points (such as 2D or 3D points), find new map points, and render a more accurate version of the world model. The world model of a first user can be communicated to a second user (e.g., via a network such as a cloud network) so that the second user can experience the world surrounding the first user.

[0060] 7 is a block diagram of an example MR environment 700. The MR environment 700 may be configured to receive inputs (e.g., visual input 702 from a user's wearable system, stationary input 704 such as a room camera, sensory input 706 from various sensors, gestures, totems, eye tracking, user input, etc. from user input device 466) from one or more user-wearable systems (e.g., wearable system 200 or display system 220) or stationary room systems (e.g., room cameras, etc.). The wearable systems can determine the location and various other attributes of the user's environment using various sensors (e.g., accelerometers, gyroscopes, temperature sensors, movement sensors, depth sensors, GPS sensors, inward-facing imaging systems, outward-facing imaging systems, etc.). This information may be further supplemented with information from stationary cameras in the room, which may provide images from different perspectives or various cues. Image data acquired by cameras (e.g., room cameras or outward-facing imaging system cameras) may be reduced to a set of mapping points.

[0061] One or more object recognizers 708 can crawl through the received data (e.g., a collection of points), recognize or map the points, tag the images, and associate semantic information with the objects using a map database 710. The map database 710 may comprise various points and their corresponding objects collected over time. The various devices and the map database may be interconnected through a network (e.g., a LAN, a WAN, etc.) and accessible to the cloud.

[0062] Based on this information and the set of points in the map database, the object recognizers 708a-708n may recognize objects in the environment. For example, the object recognizers may recognize faces, people, windows, walls, user input devices, televisions, documents (e.g., travel documents, driver's licenses, passports as described in the security embodiments herein), other objects in the user's environment, etc. One or more object recognizers may be specialized for objects with certain characteristics. For example, object recognizer 708a may be used to recognize faces, while another object recognizer may be used to recognize documents.

[0063] Object recognition may be performed using various computer vision techniques. For example, the wearable system may analyze images acquired by the outward-facing imaging system 464 (shown in FIG. 4) and perform scene reconstruction, event detection, video tracking, object recognition (e.g., people or documents), object pose estimation, face recognition (e.g., from images of people in the environment or on documents), learning, indexing, motion estimation, or image analysis (e.g., identifying indicia in documents such as photographs, signatures, identification information, travel information, etc.). One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale Invariant Feature Transform (SIFT), Speed-Up Robust Features (SURF), Orientation FAST and Rotation BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retinal Keypoints (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous localization and mapping (vSLAM) techniques, sequential Bayes estimators (e.g., Kalman filter, extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), iterative nearest neighbor (ICP), semi-global matching (SGM), semi-global block matching (SGBM), feature point histograms, various machine learning algorithms (e.g., support vector machines, k-nearest neighbor algorithms, naive Bayes, neural networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), etc.

[0064] Object recognition can additionally or alternatively be performed by various machine learning algorithms. Once trained, the machine learning algorithms can be stored by the HMD. Some examples of machine learning algorithms can include supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., a priori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural network, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., stacked generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models can be customized for individual datasets. For example, the wearable device can generate or store a base model. The base model may be used as a starting point to generate additional models specific to a data type (e.g., a particular user in a telepresence session), a data set (e.g., a set of additional images acquired of a user in a telepresence session), a conditional situation, or other variations. In some embodiments, the wearable HMD can be configured to generate models for analysis of aggregated data using multiple techniques. Other techniques may include using predefined thresholds or data values.

[0065] Based on this information and the set of points in the map database, the object recognizer 708a-708n may recognize the object, complement the object with semantic information, and bring it to life. For example, if the object recognizer recognizes that a set of points is a door, the system may associate some semantic information (e.g., a door has a hinge and 90-degree movement around the hinge). If the object recognizer recognizes that a set of points is a mirror, the system may associate semantic information that a mirror has a reflective surface that can reflect images of objects in a room. The semantic information may include the affordances of the object, as described herein. For example, the semantic information may include the object's normal. The system can assign a vector whose direction indicates the object's normal. Over time, the map database grows as the system (which may reside locally or be accessible over a wireless network) accumulates more data from the world. Once the object is recognized, the information may be transmitted to one or more wearable systems. For example, MR environment 700 may contain information about a scene being generated in California. Environment 700 may be transmitted to one or more users in New York. Based on data received from the FOV camera and other inputs, object recognizers and other software components can map points collected from various images, recognize objects, etc., so that the scene can be accurately "passed" to a second user who may be in a different part of the world. Environment 700 may also use a topology map for localization purposes.

[0066] 8 is a process flow diagram of an example method 800 for rendering virtual content in relation to recognized objects. Method 800 describes how a virtual scene can be presented to a user of a wearable system. The user may be geographically remote from the scene. For example, a user may be in New York but may want to view a scene currently occurring in California, or may want to go for a walk with a friend who is in California.

[0067] In block 810, the wearable system may receive input from the user and other users regarding the user's environment. This may be accomplished through various input devices and knowledge already held in a map database. The user's FOV camera, sensors, GPS, eye tracking, etc., communicate information to the system in block 810. The system may determine sparse points based on this information in block 820. The sparse points may be used to determine pose data (e.g., head pose, eye pose, body pose, or hand gestures) that can be used in displaying and understanding the orientation and position of various objects in the user's surroundings. The object recognizer 708a, 708n may crawl through these collected points and recognize one or more objects using the map database in block 830. This information may then be communicated to the user's respective wearable system in block 840, and the desired virtual scene may be displayed to the user appropriately in block 850. For example, a desired virtual scene (eg, a user in CA) may be displayed in the proper orientation, position, etc., relative to various objects and other surroundings of the user in New York.

[0068] FIG. 9 is a block diagram of another example of a wearable system. In this example, the wearable system 900 includes a map, which may include map data about the world. The map may reside partially locally on the wearable system and partially in a networked storage location (e.g., in a cloud system) accessible by a wired or wireless network. An attitude process 910 may run on the wearable computing architecture (e.g., processing module 260 or controller 460) and utilize data from the map to determine the position and orientation of the wearable computing hardware or the user. The attitude data may be calculated from data collected on the fly as the user experiences the system and moves within its world. The data may include images of objects in the real or virtual environment, data from sensors (such as inertial measurement units, which generally include accelerometer and gyroscope components), and surface information.

[0069] The sparse point representation may be the output of a simultaneous localization and mapping (e.g., SLAM or vSLAM, which refers to configurations where the input is image / vision only) process. The system can be configured to find not only the location of various components in the world, but also what the world is made of. Poses can be building blocks that accomplish many goals, including capturing in and using data from maps.

[0070] In one embodiment, the sparse point locations may not be entirely adequate by themselves, and additional information may be required to generate a multifocal AR, VR, or MR experience. A dense representation, generally referring to depth map information, may be utilized to fill in this gap, at least in part. Such information may be calculated from a process referred to as stereoscopic vision 940, where depth information is determined using techniques such as triangulation or time-of-flight sensing. Image information and active patterns (such as infrared patterns generated using an active projector) may serve as inputs to the stereoscopic vision process 940. A significant amount of depth map information may be fused together, and some of this may be summarized using a surface representation. For example, mathematically definable surfaces may be an efficient (e.g., for large point clouds) and easy-to-summarize input to other processing devices, such as a game engine. Thus, the outputs of the stereoscopic vision process (e.g., depth maps) 940 may be combined in a fusion process 930. Pose 950 may also be input to this fusion process 930, the output of which is input to map capture process 920. Sub-surfaces may interconnect to form larger surfaces, such as in topographic mapping, and the map becomes a large-scale hybrid of points and surfaces.

[0071] Various inputs may be utilized to resolve various aspects of the mixed reality process 960. For example, in the embodiment depicted in Figure 9, game parameters may be inputs for determining that a user of the system is playing a monster battle game with one or more monsters in various locations, whether a monster is dead or fleeing under various conditions (such as when the user shoots the monster), walls or other objects in various locations, and the like. A world map may contain information about where such objects are located relative to one another, which is another useful input for mixed reality. Attitude relative to the world is likewise an input and plays an important role for nearly any interactive system.

[0072] Controls or inputs from the user are another input to the wearable system 900. As described herein, user inputs can include visual inputs, gestures, totems, audio inputs, sensory inputs, etc. To move around or play games, for example, the user may need to command the wearable system 900 with respect to a desired object. There are various forms of user control that can be utilized beyond just moving around in space. In one embodiment, a totem (e.g., a user input device) or an object such as a toy gun may be held by the user and tracked by the system. The system would preferably be configured to know that the user is holding an item and understand the type of interaction the user is having with the item (e.g., if the totem or object is a gun, the system may be configured to understand not only the location and orientation, but also whether the user is clicking a trigger or other sensitive button or element, which may be equipped with sensors such as an IMU, which can help determine the situation occurring even when such activity is not within the field of view of any of the cameras).

[0073] Hand gesture tracking or recognition may also provide input information. The wearable system 900 may be configured to track and interpret hand gestures to gesture for button presses, left or right, stop, grasp, hold, etc. For example, in one configuration, a user may wish to flip through email or calendar in a non-gaming environment or perform a “fist bump” with another person or player. The wearable system 900 may be configured to utilize a minimal amount of hand gestures, which may or may not be dynamic. For example, gestures may be simple static gestures, such as extending the hand to indicate stop, thumbs up to indicate OK, thumbs down to indicate not OK, or flipping the hand left and right or up and down to indicate a directional command.

[0074] Eye tracking is another input (e.g., tracking where the user is looking and controlling display technology to render at a specific depth or range). In one embodiment, eye vergence may be determined using triangulation, and then accommodation may be determined using a vergence / accommodation model developed for that particular person. Eye tracking is performed by an eye camera and can determine eye gaze (e.g., direction or orientation of one or both eyes). Other techniques can also be used for eye tracking, such as measuring electrical potentials with electrodes placed near the eyes (e.g., electro-oculography).

[0075] Voice recognition may be another input that may be used alone or in combination with other inputs (e.g., totem tracking, eye tracking, gesture tracking, etc.). System 900 may include an audio sensor (e.g., a microphone) that receives an audio stream from the environment. The received audio stream may be processed (e.g., by processing modules 260, 270 or central server 1650) to recognize the user's voice (from other voices or background audio) and extract commands, parameters, etc. from the audio stream. For example, system 900 may identify from the audio stream that the phrase "Show me your ID" was uttered, identify that this phrase was uttered by the wearer of system 900 (e.g., a security screener, rather than another person in the screener's environment), and derive from the phrase and situational context (e.g., a security checkpoint) an executable command to be performed (e.g., computer vision analysis of things within the wearer's FOV) and the presence of an object ("your ID") on which the command should be performed. System 900 can incorporate speaker recognition techniques to determine who is speaking (e.g., whether the speech is from the ARD wearer or another person or voice (e.g., recorded speech transmitted by loudspeakers in the environment)) and speech recognition techniques to determine what is being said. Speech recognition techniques can include frequency estimation, hidden Markov models, Gaussian mixture models, pattern matching algorithms, neural networks, matrix representations, vector quantization, speaker diarization, decision trees, and dynamic time warping (DTW) techniques. Speech recognition techniques can also include anti-speaker techniques such as cohort models and world models. Spectral features can be used to represent speaker characteristics.

[0076] With respect to the camera system, the exemplary wearable system 900 shown in FIG. 9 may include three pairs of cameras: a pair of relatively wide-FOV or passive SLAM cameras arranged on either side of the user's face, and a different pair of cameras oriented in front of the user to handle the stereoscopic imaging process 940 and capture hand gestures and totem / object trajectories in front of the user's face. The FOV cameras and pair of cameras for the stereo process 940 may be part of the outward-facing imaging system 464 (shown in FIG. 4). The wearable system 900 may include an eye-tracking camera (which may be part of the inward-facing imaging system 462 shown in FIG. 4) oriented toward the user's eyes to triangulate eye vectors and other information. The wearable system 900 may also include one or more textured light projectors (such as infrared (IR) projectors) to inject texture into the scene.

[0077] 10 is a process flow diagram of an example embodiment of a method 1000 for determining user input to a wearable system. In this example, a user may interact with a totem. A user may have multiple totems. For example, a user may have one totem designated for social media applications, another totem for playing games, etc. In block 1010, the wearable system may detect movement of the totem. Movement of the totem may be recognized through an outward-facing imaging system or may be detected through sensors (e.g., tactile gloves, image sensors, hand tracking devices, eye tracking cameras, head pose sensors, etc.).

[0078] Based at least in part on the detected gestures, eye poses, head poses, or inputs through the totem, the wearable system detects the position, orientation, or movement of the totem (or the user's eyes or head or gestures) relative to a frame of reference in block 1020. The frame of reference may be a set of map points based on which the wearable system translates the totem's (or the user's) movements into actions or commands. In block 1030, the user's interactions with the totem are mapped. Based on the mapping of the user interactions to the frame of reference 1020, the system determines the user input in block 1040.

[0079] For example, a user may move a totem or physical object back and forth, turn a virtual page, move to the next page, or move from one user interface (UI) display screen to another. As another example, a user may move their head or eyes to view different real or virtual objects within the user's FOR. If the user's gaze at a particular real or virtual object is longer than a threshold time, that real or virtual object may be selected as user input. In some implementations, the user's eye vergence-divergence can be tracked, and an accommodation / vergence-divergence model can be used to determine the user's eye accommodation state, which provides information about the depth plane the user is focusing on. In some implementations, the wearable system can use ray-casting techniques to determine real or virtual objects that are aligned with the user's head or eye pose. In various implementations, ray casting techniques can include casting a thin bundle of rays with substantially little lateral width, or casting rays with substantial lateral width (e.g., a cone or truncated cone).

[0080] The user interface may be projected by a display system as described herein (such as display 220 in FIG. 2 ). It may also be displayed using a variety of other techniques, such as one or more projectors. A projector may project an image onto a physical object, such as a canvas or a sphere. Interactions with the user interface may be tracked using one or more cameras outside or part of the system (e.g., using inward-facing imaging system 462 or outward-facing imaging system 464).

[0081] 11 is a process flow diagram of an example method 1100 for interacting with a virtual user interface. Method 1100 may be performed by a wearable system described herein. An embodiment of method 1100 can be used by a wearable system to detect a person or document within the FOV of the wearable system.

[0082] In block 1110, the wearable system may identify a specific UI. The type of UI may be provided by the user. The wearable system may identify that a specific UI needs to be captured based on user input (e.g., gestures, visual data, audio data, sensory data, direct commands, etc.). The UI can be specific to a security scenario, where the wearer of the system observes a user presenting a document to the wearer (e.g., at a passenger checkpoint). In block 1120, the wearable system may generate data for a virtual UI. For example, data associated with the UI's boundaries, general structure, shape, etc. may be generated. Additionally, the wearable system may determine map coordinates of the user's physical location so that the wearable system can display the UI in relation to the user's physical location. For example, if the UI is body-centered, the wearable system may determine coordinates of the user's physical position, head pose, or eye pose so that a ring UI can be displayed around the user or a planar UI can be displayed on a wall or in front of the user. In the security context described herein, the UI may be displayed as if it were surrounding the traveler presenting documents to the wearer of the system, so that the wearer can easily view the UI while viewing the traveler and their documents. If the UI is hand-centered, map coordinates of the user's hand may be determined. These map points may be derived through an FOV camera, data received through sensory input, or any other type of collected data.

[0083] In block 1130, the wearable system may send data from the cloud to the display, or data may be sent from a local database to the display component. In block 1140, a UI is displayed to the user based on the sent data. For example, a light field display can project the virtual UI into one or both of the user's eyes. Once the virtual UI is generated, the wearable system may simply wait for commands from the user and generate more virtual content on the virtual UI in block 1150. For example, the UI may be a body-centered ring around the user's body or the body of a person (e.g., a traveler) in the user's environment. The wearable system may then wait for a command (gesture, head or eye movement, voice command, input from a user input device, etc.) and, if recognized (block 1160), virtual content associated with the command may be displayed to the user (block 1170).

[0084] Additional examples of wearable systems, UIs, and user experiences (UX) are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. (Example objects in the user's environment)

[0085] As described with reference to FIG. 4, a user of an augmented reality device (ARD) can have a field of view (FOR), which constitutes a portion of the user's surrounding environment that can be perceived by the user via the AR system. For a head-mounted ARD, the FOR can include substantially all of the 4π steradian solid angle surrounding the wearer, as the wearer can move their body, head, or eyes to perceive virtually any direction in space. In other contexts, the user's movement can be more constricted, and therefore the user's FOR can cover a smaller solid angle.

[0086] A FOR can contain a group of objects, which can be perceived by a user through an ARD. The objects may be virtual and / or physical objects. Virtual objects may include, for example, operating system objects, such as a trash can for deleted files, a terminal for entering commands, a file manager for accessing files or directories, icons, menus, applications for audio or video streaming, notifications from the operating system, etc. Virtual objects may also include objects within an application, such as, for example, avatars, widgets (e.g., virtual representations of clocks), virtual objects, graphics, or images within a game. Some virtual objects can be both operating system objects and objects within an application.

[0087] A virtual object may be a three-dimensional (3D), two-dimensional (2D), or one-dimensional (1D) object. For example, a virtual object may be a 3D coffee mug (which may represent virtual controls for a physical coffee maker). A virtual object may also be a 2D graphical representation of a clock, which displays the current time to the user. In some implementations, one or more virtual objects may be displayed within (or associated with) another virtual object. For example, a virtual coffee mug may be shown inside a user interface plane, but the virtual coffee mug may appear to be 3D while the user interface plane may appear to be 2D.

[0088] In some embodiments, virtual objects may be associated with physical objects. For example, as shown in FIG. 12B, a virtual book 1220 may appear to be resting on a table 1242. A user can interact with the virtual book 1220 (read the book, turn its pages, etc.) as if the physical book were resting on the table 1242. As another example, a virtual wardrobe application may be associated with a mirror in a user's FOR. When the user is near the mirror, the user may be able to interact with the virtual wardrobe application, which allows the user to simulate different clothing looks using the ARD.

[0089] Objects in a user's FOR can be part of a world model, as described with reference to FIG. 9. Data associated with objects (e.g., location, semantic information, properties, etc.) can be stored in various data structures, such as arrays, lists, trees, hashes, graphs, etc. The index of each stored object may be determined, for example, by the object's location, if applicable. For example, the data structure may index objects by a single coordinate, such as the object's distance from a reference position (e.g., distance to the left (or right) of the reference position, distance from the top (or bottom) of the reference position, or depth from the reference position). In situations where the ARD includes a light field display capable of displaying virtual objects in different depth planes to the user, the virtual objects can be organized into multiple arrays located at different fixed depth planes. In some implementations, objects in the environment may be represented in vector form, which can be used to calculate the position and movement of the virtual object. For example, an object may have an origin, a downward vector in the direction of gravity, and a forward vector in the direction of the object's surface normal. For example, the surface normal of a display (e.g., a virtual TV) may indicate the direction in which the displayed image may be viewed (rather than the direction in which the back of the display may be viewed). The AR system may calculate the difference between the components of two vectors, thereby determining the spatial relationship between the objects associated with the two vectors. The AR system may also use this difference to calculate the amount of movement required to place one object on top of (or inside) another object. (Example of moving a virtual object)

[0090] A user can interact with a subset of the objects in the user's FOR. This subset of objects may sometimes be referred to as interactable objects. A user can interact with the interactable objects by performing user interface actions, such as selecting or moving the interactable object, activating a menu associated with the interactable object, selecting an action to be performed using the interactable object, etc. As is evident in the AR / VR / MR world, movement of a virtual object does not refer to actual physical movement of the virtual object, as virtual objects are computer-generated images and not physical objects.

[0091] A user may perform various user interface operations using, alone or in combination, head postures, eye postures, body postures, voice commands, or hand gestures on a user input device. For example, a user may interact with an interactable object by using hand gestures, e.g., by actuating a user input device (e.g., see user input device 466 in FIG. 4 ), such as, alone or in combination, clicking on a mouse, tapping on a touchpad, swiping on a touchscreen, hovering over or touching a capacitive button, pressing a key on a keyboard or game controller (e.g., a five-way d-pad), pointing towards an object with a joystick, wand, or totem, pressing a button on a remote control, or other interaction with a user input device. A user may also interact with an interactable object using head, eye, or body postures, e.g., by gazing or pointing at an object for a period of time, tilting their head, waving their hand towards an object, etc.

[0092] In some implementations, the AR system may provide a focus indicator (such as focus indicator 1450 shown in FIGS. 14A-14D and 15A-15H) that indicates the location of a target object (see, e.g., focus indicator in FIG. 14A). The focus indicator may also be used to indicate the current position of a user input device or the user's posture (see, e.g., focus indicator in FIG. 15D). In addition to, or as an alternative to, providing an indication of location, the focus indicator can also provide an indication of the target object, the orientation of the user input device, or the user's posture. For example, the focus indicator can comprise a halo, color, a change in perceived size or depth (e.g., causing the target object to appear closer and / or larger when selected), a graphical representation of a cursor (such as a reticle), or other audible, tactile, or visual effect that attracts the user's attention. The focus indicator can appear as a 1D, 2D, or 3D image, which may include a still-frame image or a moving image.

[0093] As an example of presenting a focus indicator by an AR system, when a user gazes at a blank wall, the AR system may project a virtual cone or ray onto the wall to indicate the user's current gaze direction. As another example, a user may activate a user input device to indicate a desire to interact with an object in the environment. The AR system may assign a focus indicator to the object so that the user may perceive the object more easily. As the user changes their posture or activates a user input device, the AR system may move the focus indicator from one location to another. Example of snapping and orienting virtual objects

[0094] As described herein, because a user may move and rotate a virtual object in multiple directions, the user may sometimes find it difficult to precisely position and orient the virtual object. To reduce user fatigue and provide an improved AR device (ARD) with which the user interacts, the ARD can automatically reposition the virtual object relative to a destination object in the environment. For example, the ARD can automatically connect (also referred to as "snapping") the virtual object to another virtual or physical object in the environment when the virtual object is within a threshold distance from the virtual or physical object. In addition to, or as an alternative to, snapping the virtual object to the destination object, the ARD can automatically change the position or orientation of the virtual object when the virtual object approaches the destination object. For example, the ARD can rotate the virtual object so that the virtual object's normal faces the user. As another example, the ARD can align the boundaries of the virtual image with those of a physical book so that the virtual object may appear to be part of the physical book.

[0095] The AR system can reposition a virtual object to an appropriate location or orientation based on the affordances of the virtual object or target object. Affordances include relationships between an object and its environment that provide opportunities for actions or uses associated with the object. Affordances may be determined, for example, based on the function, orientation, type, location, shape, or size of the virtual object or destination object. Affordances may also be based on the environment in which the virtual object or destination object is located. Affordances of a virtual object may be programmed as part of the virtual object and stored in the remote data repository 280. For example, a virtual object may be programmed to include a vector indicating the virtual object's normal.

[0096] For example, the affordance of a virtual display screen (e.g., a virtual TV) is that the display screen can be viewed from a direction indicated by the normal to the screen. The affordance of a vertical wall is that an object can be placed on a wall (e.g., "hung" on the wall) with its surface normal parallel to the normal to the wall. A user can use the AR system to move the virtual display close to the wall, and when close enough to the wall, the AR system can automatically snap the virtual display onto the wall with the display normal parallel to the wall normal without further user input. An additional affordance of the virtual display and the wall can be that they each have a top or a bottom. When the virtual display is snapped onto the wall, the AR system can automatically orient the virtual display so that the bottom of the virtual display is oriented toward the bottom of the wall (or the top of the display is oriented toward the top of the wall), thereby ensuring that the virtual display does not present an upside-down image.

[0097] In some situations, to reposition the virtual object, the ARD may also change other characteristics of the virtual object. For example, the size or shape of the virtual image may not be identical to that of the physical book. As a result, the ARD may change the size or shape of the virtual image to match that of the physical book, aligning the boundaries of the virtual image with those of the physical book. The ARD may (in addition or alternatively) reposition other nearby virtual objects to provide sufficient space for the repositioned virtual object.

[0098] In effect, the AR system respects the affordances of physical and virtual objects and positions or orients virtual objects relative to other physical or virtual objects based, at least in part, on their respective affordances. Further details regarding these features are described below. (Example of Automatically Binding Virtual Objects to Physical Objects)

[0099] 12A and 12B illustrate an example of binding a virtual object to a table. As shown in FIG. 12A, a virtual book 1220 is initially floating above a table 1242 in a room 1200a. A user may have previously moved the book 1220 from its initial position to a position above the table (as shown in FIG. 12A). The user can provide an indication to place the virtual book 1220 on the table 1242 by activating a user input device. For example, a user can select the virtual book 1220 by clicking its totem and pointing the totem to the table 1242, indicating that the virtual book 1220 should be placed on the table 1242. The AR system can move the virtual book 1220 to the table 1242 without requiring the user to drag the book onto the table 1220. As another example, a user may point their eyes at a destination position on a table where the book should be placed, and the AR system can determine the destination position based on head pose or eye gaze; the AR system can use an IMU to obtain data regarding head pose or an eye-tracking camera to determine the user's gaze direction. When the user activates the totem, the AR system can automatically place the virtual book 1220 at the destination position on the table 1242. In some implementations, the AR system can simulate a gravitational force acting to pull the virtual object toward the target object, so the user does not need to indicate a destination object, such as the table 1242 or the wall 1210, to which the target object will be moved. For example, when the user selects the virtual book 1220, the AR system can simulate the effect of gravity on the virtual book 1220, causing the virtual book 1220 to move in a downward direction 1232. The AR system can automatically identify the table 1242 as the destination object because it is the first object on the path of the virtual book 1220 in its downward movement (as indicated by the arrow 1232). Thus, a virtual book can appear to be dropped onto the table 1242 just as if it were a physical book.

[0100] The AR system can determine parameters and calculate values ​​for the parameters for repositioning the virtual book 1220 based on the affordances of the virtual book, the environment (e.g., room 1200a), or the destination object. Some example parameters of movement can include the amount of movement (e.g., distance traveled or trajectory), the speed of movement, the acceleration of movement, or other physics parameters.

[0101] The AR system can calculate the amount of movement for the virtual book 1220 based on the position of the virtual book 1220 and the position of the table 1242. For example, the AR system may associate vectors with physical and virtual objects in a room. The vectors may include location and direction information for the physical and virtual objects. The vector for the virtual book 1220 may have a component in the direction of gravity (e.g., direction 1232). Similarly, the vector for the surface of the table 1242 may also have a component in the direction of gravity indicating its current location. The AR system can obtain the difference between the position of the surface of the table 1242 and the position of the virtual book 1220 in the direction of gravity. The AR system can use this difference to determine the distance the virtual book 1220 should be moved downward. The vector can also include information about the magnitude of gravity. The magnitude of gravity (e.g., gravitational acceleration) may vary based on the user's physical or virtual environment. For example, when a user is at home, gravity may be at an Earth value of 9.8 m / s. 2 (1 "g"). However, a user may play a game using an AR system, which may present a virtual environment within the game. As an example, if the virtual environment is a room, the magnitude of gravitational acceleration may be 1 / 6 "g", and if the virtual environment is Jupiter, the magnitude of gravitational acceleration may be 2.5 "g".

[0102] To provide an improved user experience using an AR system, the AR system can simulate the movement of the virtual book 1220 in FIGS. 12A-12B using various laws of physics as if the virtual book 1220 were a physical object. For example, the downward movement of the virtual book 1220 (as indicated by arrow 1232) can be based on a free-fall motion. The free-fall motion can be combined with other forces in the user's environment (such as air resistance) to provide a realistic user experience. As another example, the room 1200a may have an open window. When a gust of wind blows into the room 1200a, the AR system can simulate the effect of the wind by automatically turning the pages of the virtual book 1220.

[0103] 12A and 12B , the affordance of the table is that it can support objects on its surface, so the virtual book 1220 falls onto the table 1242. Therefore, when simulating the effects of gravity, the AR system will not display the virtual book 1220 on the floor 1230 because the AR system respects the affordance of the table that objects do not pass through the table. The table 1242 will prevent the virtual book 1220 from continuing to move in the direction of gravity as if the virtual book 1242 were a physical book.

[0104] In some situations, only a portion of the table 1242 is on the path of the downward movement of the virtual book 1220. As a result, a portion of the virtual book 1220 may extend beyond the surface of the table 1242. The AR system can determine whether the center of gravity of the virtual book 1220 is on the surface of the table 1242. If the center of gravity of the virtual book 1220 rests on the surface of the table 1242, the AR system can display the virtual book 1220 on the table 1242. If the center of gravity of the virtual book 1220 is outside the surface of the table 1242, the AR system can determine that the virtual book 1220 will not remain on the table 1242 and can instead display the virtual book 1220 on the floor.

[0105] As another example, the virtual object may be a virtual tennis ball whose affordances include bouncing off hard surfaces. Thus, when the virtual tennis ball hits table 1242, the AR system may show the virtual tennis ball bouncing off table 1242 and landing on floor 1230.

[0106] An additional affordance of virtual objects (e.g., books) and tables is that the normal to the object should be parallel to the normal to the table (e.g., the virtual book is flat on the table). The AR system can automatically properly orient the virtual book 1220 so that it appears to the user as flat on the table 1242, as shown in Figure 12B.

[0107] In addition to moving virtual objects in the direction of gravity, the AR system can also move virtual objects in other directions. For example, a user may wish to move a note onto a wall. The user may point (e.g., with a hand or totem) in the direction of wall 1210. Based on the direction indicated by the user, the AR system may use an outward-facing imaging system and / or world model to identify the surface of wall 1210 and automatically "fly" the note onto the wall (see, e.g., example scenes 1300a, 1300b, 1300c, and 1300d shown in Figures 13A-13D). In other implementations, objects on the wall may have affordances that may attract notes. For example, objects on the wall may represent surfaces to which notes are typically attached (e.g., a magnetic memo board to which magnetic notes can be attached or a cork memo board to which notes can be pinned). The AR system can identify that the wall object has a "sticky" affordance (e.g., a magnetic or cork memo board) and that the virtual note has a corresponding sticky affordance (e.g., the note is magnetic or pinnable). The AR system can automatically attach the virtual note to the wall object.

[0108] As described with reference to FIG. 7 , semantic information can be associated with physical objects, virtual objects, the physical environment, and the virtual environment. The semantic information can include affordances. For example, the AR system can assign physical attributes to virtual objects, such as the following non-exclusive illustrative attributes: mass, density, diameter, hardness (or softness), elasticity, viscosity, electromagnetic attributes (e.g., charge, conductivity, magnetic attributes), and topology (e.g., solid, liquid, or gas). The AR system can also assign physical attributes to the virtual environment, such as gravity and air resistance. The values ​​of the assigned attributes can be used to simulate interactions with the virtual objects using various laws of physics. For example, the movement of a virtual object may be based on a force applied to the virtual object. Referring to FIGS. 12A and 12B , the movement of a virtual book 1220 may be determined based on a force calculated using the gravity and air resistance of a room 1200 a and the mass of the virtual book 1220. (Virtual copy and paste example)

[0109] In some embodiments, rather than repositioning the virtual object to a target destination object, the AR system may duplicate the virtual object and move the duplicated virtual object to the destination object or a location within the environment. As an example, the AR system may present a virtual menu to the user, whereby the user may select one or more items from the virtual menu (e.g., using a totem or head or eye pose). The AR system may automatically place the selected item from the virtual menu in an appropriate location based on the affordances of the selected item and objects in the user's environment. For example, the virtual menu may present items to the user, including a virtual TV, a virtual audio player, a virtual game, etc. The user may select a virtual TV. The AR system may copy the virtual TV from the menu into a clipboard. The user may look around the environment and find a target location where the user wants the AR system to place the virtual TV. Once the user finds the target location, the user may actuate a user input device (e.g., a totem) to confirm the selection of the target location, and the AR system may automatically display the virtual TV in the desired location. The AR system may display the virtual TV based on the affordances of the TV and the target location. For example, if the desired location is on a vertical wall, the AR system can display the virtual TV as if it were hanging flat against the wall. As another example, if the target location is on a horizontal table, the AR system can display the virtual TV as if it were lying flat on the table. In both of these examples, the AR system orients the virtual TV (which has a normal indicating the direction in which the TV can be viewed) so that the normal of the virtual TV is aligned with the normal to the target location, i.e., the wall normal or the table normal.

[0110] In some situations, a virtual object may not be able to be placed at a target location selected by a user. For example, a user may select a table surface to place a virtual note. However, the table surface may already be covered by another document, and therefore the virtual note cannot be placed on the table surface. The AR system may simulate a reaction force, such as a repulsive spring force, which would prevent the virtual note from being placed on the table surface.

[0111] Thus, such an embodiment of an AR system can copy a virtual object and paste it where desired with minimal user interaction because the AR system knows the affordances of the virtual object as well as other objects in the user's environment. The AR system can use these affordances to place the virtual object in its natural position and / or orientation within the user's environment. (Example of automatically pivoting a virtual object)

[0112] 13A-13D show automatically adjusting the position and orientation of virtual object 1320 when virtual object 1320 approaches wall 1210. Virtual object 1320 has four corners 1326, 1328, 1322, and 1324. Once the AR system determines that virtual object 1320 is touching wall 1210 (e.g., corner 1322 is touching the wall, as shown in FIG. 13B ), the AR system can reorient virtual object 1320 so that virtual object 1320 appears aligned with the orientation of wall 1210, e.g., object normal 1355 is parallel to wall normal 1350, as shown in FIG. 13D . The movement to reorient virtual object 1320 may be along multiple axes in 3D space. For example, as shown in Figures 13B and 13C, the AR system can pivot the virtual object 1320 in the direction 1330a, so that corner 1324 also touches the wall 1210, as shown in Figure 13C. To further align the virtual object 1320 with the wall 1210, the AR system can pivot the virtual object 1320 in the direction 1330b, as shown in Figure 13C. Thus, the top corners 1326 and 1328 of the virtual object 1320 also touch the wall, as shown in Figure 13D. In some cases, the virtual object 1320 may also have an affordance whose natural orientation on the wall 1210 is to hang horizontally (e.g., a virtual painting). The AR system may use this affordance to properly orient the virtual object 1320 on the wall 1210 (e.g., the virtual painting appears to hang properly from the wall, rather than at an angle).

[0113] The AR system can reorient the virtual object 1320 by making its surface normal 1355 parallel to the surface normal 1350 of the wall 1210, rather than anti-parallel to the normal 1350. For example, only one side of a virtual TV screen may be configured to display a video, or only one side of a virtual painting may be configured to display a painting. Thus, the AR system may need to flip the virtual object so that the side with the content faces the user (instead of facing the wall). Example of snapping and reorienting when a virtual object is within a threshold distance of another object

[0114] A user can move a virtual object to another location, for example, by dragging the virtual object using user input device 466. When the virtual object is in proximity to a destination object, the AR system may automatically snap and orient the virtual object so that the user does not have to make small adjustments to align the virtual object with the destination object.

[0115] 14A-14D, a user of the ARD moves a virtual object from a table to a wall. The virtual object 1430 may be a virtual TV screen. In FIG. 14A, the virtual TV screen 1430 in room 1200b is initially on table 1242 (as indicated by the dashed line on table 1242) and is being moved to wall 1210. The user may select virtual screen 1430 and move virtual screen 1430 in direction 1440a (e.g., toward a desired destination location on wall 1210). The AR system may show a visible focus indicator 1450 on virtual screen 1430 to indicate the current location of virtual screen 1430 and to indicate that the user has selected screen 1430.

[0116] The AR system can monitor the position of the virtual screen 1430 as the user moves it. The AR system can begin automatically snapping and orienting the virtual screen 1430 when the distance between the virtual screen 1430 and the wall 1210 is less than a threshold distance. The AR system may set a threshold distance so that the AR system may begin automatically attaching and orienting the virtual screen 1430 when at least a portion of the virtual screen touches the wall 1210, as described with reference to FIGS. 13A-13D . In some implementations, the threshold distance may be large enough so that the AR system begins automatically attaching and / or orienting the virtual screen 1430 to the wall 1210 before any portion of the virtual screen 1430 touches the wall. In these implementations, the AR system may simulate a magnetic effect between the wall and the virtual object. For example, when the virtual screen 1430 gets close enough (e.g., less than a threshold distance), the object is automatically attracted to the wall without any further effort from the user.

[0117] The distance between the virtual screen 1430 and the wall 1210 may be measured in various ways. For example, it may be calculated based on the displacement between the center of gravity of the virtual screen 1430 and the surface of the wall 1210. In some embodiments, when an object is associated with a vector that describes the object's position, the AR system can calculate the Euclidean distance using the vector with respect to the virtual screen 1430 and the vector with respect to the wall 1210. The distance may also be calculated based on the components of the vector. For example, the ARD may calculate the position difference between the vector on the horizontal axis and the virtual screen 1430 as the distance.

[0118] The threshold distance can depend on the affordances associated with the virtual object. For example, the virtual screen 1430 is associated with a size (e.g., horizontal size, vertical size, thickness, diagonal size, etc.), and the threshold distance can be a percentage of the size. As an example, if the threshold distance is approximately equal to the vertical size of the screen, the AR system may begin orienting the screen toward its destination location and orientation once the screen is within the vertical size distance from a wall. As another example, if the threshold distance is smaller than the vertical size of the screen, the AR system may not begin orienting the virtual screen 1430 until it is much closer to the wall. In various implementations, the threshold distance can be set by the user or can be set to a default value (e.g., the size of the object, etc.). The threshold distance may vary based on the user's experience with the AR system. For example, if a user finds that a small threshold distance leads to rapid reorientation of the virtual object and is distracting, the user may reset the threshold distance to be larger so that the reorientation occurs more gradually over a larger distance.

[0119] When the virtual screen 1430 is not parallel to the wall, the AR system may use the portion of the virtual screen 1430 closest to the wall (such as the bottom portion of the virtual screen 1430 shown in FIG. 14B ) as the end point when calculating the distance. Using this method, when the distance becomes small enough or becomes zero, the AR system can detect that at least a portion of the virtual screen 1430 has touched the wall 1210. However, if the distance is calculated based on the displacement between the wall 1210 and the center of gravity of the virtual screen 1430, the distance may be greater than zero when the AR system detects a collision (see FIG. 14B ). This is because when the virtual screen 1430 is not parallel to the wall 1210, a portion of the virtual screen 1430 may reach the wall 1210 before the center of gravity of the virtual screen 1430 reaches the wall 1210.

[0120] In various implementations, a virtual object may have the affordance that only one surface of the virtual object displays content (e.g., a TV screen or a painting) or has a texture designed to be visible to the user. For example, with respect to a virtual TV object, only one surface of the object may act as a screen and display content to the user. As a result, the AR system may orient the virtual object so that the surface with the content faces the user instead of the wall. For example, with reference to FIGS. 14A, 14B, and 14C, the surface normal of virtual screen 1430 may initially face the ceiling (instead of the table) so that the user can see the content of virtual screen 1430 when standing in front of the table. However, if the user simply lifts virtual screen 1430 and attaches it to wall 1210, the surface normal may face the wall instead of the user. As a result, the user may not see the content and may see the back surface of virtual screen 1430. To ensure that the user can still see the content when virtual screen 1430 is moved to wall 1210, the AR system can rotate virtual screen 1430 180 degrees about an axis so that the surface with the content faces the user (and not the wall). With this rotation, side 1434, which was the bottom side of virtual screen 1430 in Figure 14B, becomes the top side in Figure 14C, while side 1432, which was the top side of virtual screen 1430 in Figure 14B, becomes the bottom side in Figure 14C.

[0121] In addition to, or as an alternative to, flipping the virtual object along an axis as described with reference to FIG. 14C , the AR system may also rotate the virtual object around other axes so as to retain the same orientation as before the virtual object was moved to the destination location. Referring to FIGS. 14A and 14D , the virtual TV screen 1430 (and its contents) may initially be in a portrait orientation on the table 1242. However, after the virtual TV screen 1430 is moved to the wall 1210, as shown in FIGS. 14B and 14C , the virtual TV screen is oriented in a landscape orientation. To retain the same user experience as when the virtual screen 1430 was on the table 1242, the AR system may rotate the virtual TV screen 1430 90 degrees so that the virtual TV screen 1430 appears in a portrait orientation (as shown in FIG. 14D ). Using this rotation, side 1434, which was shown as the top of screen 1430 in FIG. 14C, becomes the left side of screen 1430 in FIG. 14D, while side 1432, which was shown as the bottom side of screen 1430 in FIG. 14C, becomes the right side of screen 1430 in FIG. 14D.

[0122] 15A, 15B, and 15C illustrate side views of an example of connecting and orienting a virtual object. In this example, the virtual object is depicted as a planar object (e.g., a virtual screen) for illustrative purposes, but virtual objects are not limited to planar shapes. Similar to FIG. 14A, the virtual screen 1430 in FIG. 15A moves in a direction 1440a toward the wall 1210. The AR system can display a focus indicator 1450 to indicate a position associated with the user (e.g., the user's gaze direction or the position of the user's user input device 466, etc.). In this example, the focus indicator 1450 is on the virtual screen 1430, which can indicate selection of the virtual screen 1430. The virtual screen 1430 is at an angle 1552 with the wall 1210 (shown in FIG. 15B). When the virtual screen 1430 touches the surface 1510a of the wall 1210, as shown in Figure 15B, the AR system can rotate the virtual screen 1430 in an angular direction 1440b. As a result, the angle formed between the virtual screen 1430 and the surface 1510a of the wall 1210 is reduced from angle 1552 to angle 1554, as shown in Figure 15C, and in Figure 15D, the angle is reduced to zero because the screen 1430 is flat against the wall 1210. In this example, the screen 1430 remains on the wall (due to the "stickiness" of the wall) even after the user moves slightly away from the wall 1210 (the focus indicator 1450 is no longer on the virtual object 1430).

[0123] While the exemplary illustrations described herein show that the reorientation of a virtual object occurs after the virtual object touches a wall, it should be noted that the examples herein are for illustrative purposes and are not intended to be limiting. The reorientation can also occur before the virtual object touches a physical object. For example, the AR system can calculate the surface normals of the wall and the virtual object and orient the virtual object while the virtual object moves toward the wall. Additionally, it should be noted that while the examples provided herein show the bottom portion of the virtual object touching the wall first, any other portion of the object may also touch the wall first. For example, when a virtual object is parallel to a wall, the entire virtual object may collide with the wall simultaneously. As another example, when the object is a 3D cup, the handle of the cup may collide with the wall before any other portion of the cup. (Simulated gravitational effects between virtual and physical objects)

[0124] As described herein, the AR system can simulate the effect of an attractive force between a virtual object and a physical or virtual object, such as a wall or table (e.g., gravity in FIGS. 12A-12B or magnetic force or "stickiness" as shown in FIGS. 13A-13D or 14A-14D). As the distance between the virtual object and the physical object falls below a threshold, the AR system can automatically attach the virtual object to the physical object as if the two objects were attracted together due to an attractive force (e.g., the opposite polarity attractive force of a magnet or the downward attractive force of gravity).

[0125] In some cases, the AR system may utilize multiple attractive forces, which may more accurately represent the trajectory of a physical object. For example, referring to FIGS. 14A-14B , a user wishing to move virtual screen 1430 from table 1242 toward wall 1210 may perform a gesture to throw virtual screen 1430 onto wall 1210. The AR system may utilize magnetic attractive forces, as described above, to attach screen 1430 to wall 1210. In addition, the AR system may utilize downward gravity to represent an arcing trajectory of screen 1430 as it moves toward wall 1210. The AR system's use of one or more attractive forces may cause virtual objects to appear to move and interact with other objects in the user's environment in a more natural manner, since the virtual objects behave in much the same way as physical objects move. This may advantageously lead to a more natural and realistic user experience.

[0126] In other situations, a user may want to move a virtual object away from the object to which it is currently tied. However, sometimes, when a virtual object is tied to a wall, the AR system may be unable to distinguish between a user's movement indicating an intention to unbind the virtual object and a user's interaction with the virtual object. As an example, while a user is playing a game using a virtual screen tied to a wall, the user may need to move the totem to find or interact with a friend or foe. This type of game movement may match the type of movement that unbinds a virtual object from a wall. By unbind- ing the virtual screen only if the user's movement sufficiently exceeds a suitable threshold, the virtual screen will not be inadvertently unbind- ed during gameplay. In addition, users typically cannot maintain their posture or user input device still for an extended period of time. As a result, a virtual object may be accidentally unbind- ed by a slight user movement when the user does not intend to unbind the virtual object. Thus, by only untying a virtual object if the user's movement is sufficiently above a suitable threshold, slight movements or jerks by the user will not inadvertently untying the virtual object from its intended location or orientation.

[0127] To solve these problems and improve AR systems, the AR system can simulate gravitational forces between virtual and physical objects so that a user cannot immediately untether a virtual object from another object unless the user's change in position exceeds a threshold. For example, as shown in FIG. 15D , a screen 1430 is tethered to a wall 1210. The user can move their user input device so that the focus indicator 1450 is moved away from the wall. However, the virtual screen 1430 may still be tethered to the surface 1510 a due to the effect of the simulated gravitational force between the screen and the wall. As the focus indicator moves further away and exceeds a threshold distance between the screen 1430 and the focus indicator 1450, the AR system may untether the virtual screen 1430 from the wall 1210, as shown in FIG. 15E . This interaction provided by the AR system acts as if an invisible virtual tether exists between the focus indicator 1450 and the screen 1430. When the distance between the focus indicator 1450 and the screen 1430 exceeds the length of the virtual string, the virtual string becomes taut and unties the screen from the wall 1210. The length of the virtual string represents the threshold distance.

[0128] In addition to or as an alternative to the threshold distance, the AR system can also use other factors to determine whether to unbind a virtual object. For example, the AR system may measure the acceleration or velocity of the user's movement. The AR system may measure the acceleration and velocity using an IMU, as described with reference to FIGS. 2 and 4 . If the acceleration or velocity exceeds a threshold, the AR system may unbind the virtual object from the physical object. In some cases, the rate of change of acceleration (known as a "jerk") can be measured, and if the jerk provided by the user exceeds a jerk threshold, the virtual object is unbind. An implementation of the AR system that utilizes an acceleration or jerk threshold may more naturally represent a user unbind an object from another object. For example, to remove a physical object stuck to a wall, the user may grasp a part of the physical object and jerk it away from the wall. The representation of such a jerk in the virtual world can be modeled using an acceleration and / or jerk threshold.

[0129] In another implementation, the AR system may simulate other physical forces, such as friction or elasticity. For example, the AR system may simulate the interaction between the focus indicator 1450 and the screen 1430 as if there were a virtual rubber band connecting them. As the distance between the focus indicator 1450 and the screen 1430 increases, the virtual pulling force of the virtual rubber band increases, and when the virtual pulling force exceeds a force threshold representing the stickiness of the wall 1210, the screen 1430 becomes detached from the wall 1210. In another example, when a virtual object appears on a horizontal surface of a physical object, the AR system may simulate the interaction between the focus indicator and the virtual object as if virtual friction exists between the virtual object and the horizontal surface. While a user may drag the focus indicator along the horizontal surface, the AR system may begin moving the virtual object only when the force applied by the user is sufficient to overcome the virtual friction.

[0130] Thus, in various implementations, an AR system can utilize distance, speed, acceleration, and / or jerk measurements, along with corresponding thresholds, to determine whether to tie or untie a virtual object from another object. Similarly, an AR system can utilize thresholds for the attractive forces (representing stickiness, gravity, or magnetism) between objects to determine the strength with which a virtual object is tied to another object. For example, a virtual object intended to be immovably placed on another object may be associated with a very high attractive force threshold so that it may be very difficult for a user to tie or untie the virtual object. Unlike the physical world, where the properties of physical objects are determined by factors such as their weight, in the virtual world, the properties of virtual objects can be changed. For example, if a user desires to intentionally move an "immovable" virtual object, the user may instruct the AR system to temporarily change the virtual object's settings so that its associated threshold is much lower. After the virtual object has been moved to a new location, the user can instruct the AR system to reset the virtual object's threshold so that it is immovable again.

[0131] When a virtual object is detached from a physical object, the orientation of the virtual object can remain the same as when the virtual object is attached to a physical object. For example, in Figure 15D, when virtual screen 1430 is attached to a wall, virtual screen 1430 is parallel to wall surface 1510a. Thus, as virtual screen 1430 moves away from the wall in Figure 15E, virtual screen 1430 can remain parallel to wall 1430.

[0132] In some embodiments, when a virtual object is detached from a wall, the AR system may change the orientation of the virtual object back to its original orientation before the virtual object was detached from the wall. For example, as shown in Figure 15F, when virtual screen 1430 is detached from wall 1210, the AR system may return the orientation of virtual screen 1430 to the same orientation as shown in Figures 15A and 15B.

[0133] Once virtual screen 1430 is detached from wall 1210, the user can move object 1430 and reattach it to the wall (or another object). As shown in Figures 15G and 15H, when the user moves virtual screen 1430 back toward wall 1210, virtual object 1430 can be oriented in direction 1440c and reattached to surface 1510b of wall 1210. Thus, as shown in Figure 15D, the user can reattach object 1430 to the left side of wall 1210 (e.g., compare the position of virtual screen 1430 in Figure 15D with its position in Figure 15H).

[0134] 14A-14D and 15A-15H are described with reference to automatically binding a virtual object to a wall, but the same techniques can also be applied to automatically binding a virtual object to another object, such as a table. For example, the AR system can simulate the effect of downward gravity when a virtual book approaches a table and automatically orient and display the virtual book on the table without any additional user effort. As another example, a user can use hand gestures to move a virtual painting to the floor and drop it onto the floor (via virtual gravity). The user's current position may be indicated by a focus indicator. As the user moves the painting to the floor, the focus indicator can follow the user's position. However, when the virtual painting approaches a table in the user's environment, the virtual painting may be accidentally bound to the table because the table may be on the floor. The user can use their hand, head, or eye gestures to continue moving the focus indicator toward the floor. When the distance between the location of the focus indicator and the table is large enough, the AR system may detach the virtual picture from the table and move it to the location of the focus indicator. The user may then continue to move the virtual picture until it is near the floor, and the AR system may drop it into place under virtual gravity.

[0135] Additionally or alternatively, the techniques described herein can also be used to push a virtual object inside another object. For example, a user may have a virtual box that contains a photo of the user. When the AR system receives an indication to place the photo inside the virtual box, the AR system can automatically move the photo inside the virtual box and align the photo with the virtual box. (Example Method for Automatically Snapping and Orienting Virtual Objects)

[0136] 16 is an exemplary method for connecting and orienting virtual objects. The process 1600 shown in FIG. 16 may be implemented by the AR system 200 described with reference to FIG.

[0137] In block 1610, the AR system may identify a virtual object in the user's environment with which the user desires to interact. This virtual object may also be referred to as a target virtual object. The AR system may identify the target virtual object based on the user's pose. For example, the AR system may select a virtual object as the target virtual object when it intersects with the user's line of sight. The AR system may also identify the target virtual object when the user actuates a user input device. For example, when the user clicks the user input device, the AR system may automatically designate the target virtual object based on the current location of the user input device.

[0138] The target virtual object may have a first location and a first orientation. In some embodiments, the location and orientation of the target virtual object may be represented in vector form. As the target virtual object moves around, the values ​​in the vector may be updated accordingly.

[0139] In block 1620, the AR system may receive an indication to move the target virtual object to a destination object. For example, a user may point to a table and actuate a user input device to indicate an intent to move the target virtual object to the table. As another example, a user may drag a virtual object toward a wall to place the virtual object on the wall. The destination object may have a location and an orientation. In some implementations, like the target virtual object, the location and orientation of the destination object may also be represented in vector form.

[0140] In block 1630, the AR system may calculate a distance between the target virtual object and the destination object based on the location of the target virtual object and the location of the destination object. The AR system may compare the calculated distance with a threshold distance. If the distance is less than the threshold distance, as shown in block 1640, the AR system may automatically orient and connect the target virtual object to the destination object.

[0141] In some implementations, the AR system may include multiple threshold distances (or speeds, accelerations, or jerkiness), with each threshold distance associated with a type of action. For example, the AR system may set a first threshold distance, and the AR system may automatically rotate the virtual object when the distance between the target virtual object and the destination object is less than or equal to the threshold distance. The AR system may also set a second threshold distance, and the AR system may automatically link the target virtual object to the destination object when the distance is less than the second threshold distance. In this example, if the first threshold distance is equal to the second threshold distance, the AR system may start linking and orienting the target virtual object simultaneously once the threshold distance is met. If the first threshold distance is greater than the second threshold distance, the AR system may start orienting the target virtual object before the AR system starts automatically linking the target virtual object to the destination object. On the other hand, if the first threshold distance is less than the second threshold distance, the AR system may first attach the target virtual object to the destination object and then orient the target virtual object (such as when a portion of the target virtual object is already attached to the destination object). As another example, the AR system may detach the virtual object if the user movement exceeds a corresponding velocity, acceleration, or jerk threshold.

[0142] When orienting a target virtual object, the AR system can orient the target virtual object along multiple axes. For example, the AR system can rotate the virtual object so that the surface normal of the virtual object faces the user (instead of facing the surface of the destination object). The AR system can also orient the virtual object so that it appears in the same orientation as the destination object. The AR system can further adjust the orientation of the virtual object so that the user does not view the contents of the virtual object from an uncomfortable angle. In some embodiments, the AR system may simulate the effects of magnetic and / or gravitational forces when connecting two objects together. For example, the AR system can exhibit an attracting effect on the wall when the virtual TV screen is close to a wall. As another example, the AR system may simulate a free-fall motion when a user indicates an intention to place a virtual book on a table.

[0143] In addition to the effects of magnetism, adhesion, and gravity, the AR system may also simulate other physical effects as if the virtual object were a physical object. For example, the AR system may assign mass to the virtual object. When two virtual objects collide, the AR system may simulate momentum effects so that the two virtual objects may move together over a distance after the collision. Exemplary Methods for Attaching and Unattaching Virtual Objects

[0144] 17 is an exemplary method for associating and unassociating a virtual object with another object in a user's environment. The process 1700 shown in FIG. 17 may be implemented by the AR system described with reference to FIG. 2.

[0145] In block 1710, the AR system may identify a target virtual object within a user's field of view (FOR). The FOR constitutes a portion of the user's surrounding environment that can be perceived by the user through the AR system. The target virtual object may be linked to another object within the FOR. The AR system may assign a first location to the target object once linked to the other object.

[0146] In block 1720, the AR system may receive an indication to move the target virtual object to a second position within the user's FOR. The indication may be a change in the user's posture (such as moving their hand), a movement of a user input device (such as moving a totem), or a hand gesture on a user input device (such as moving along a trajectory on a touchpad).

[0147] In block 1730, the AR system may determine whether the indication of movement satisfies a threshold condition for dissociating the target virtual object from another object. As shown in block 1742, if the threshold condition is met, the AR system may dissociate the virtual object and move it to the user's position (e.g., as indicated by the focus indicator). The threshold condition may be based on the speed / acceleration of movement and / or the position change. For example, if a user desires to dissociate a virtual object from a wall, the AR system may calculate the distance the user has moved the totem from the wall. If the AR system determines that the distance between the totem and the wall meets the threshold distance, the AR system may dissociate the virtual object. As an example, if the user moves the totem fast enough to meet the threshold speed and / or threshold acceleration, the AR system may also dissociate the virtual object from the wall.

[0148] If the AR system determines that the threshold condition is not met, the AR system may not disassociate from the virtual object, as shown in block 1744. In some embodiments, the AR system may provide a focus indicator that indicates the user's current location. For example, when the threshold condition is not met, the AR system may show the virtual object as still being associated with other objects in the environment while showing the focus indicator at the user's current location.

[0149] Note that while the examples described herein refer to moving a single virtual object, these examples are not limiting. For example, an AR system may use the techniques described herein to automatically orient, connect, and disconnect a group of virtual objects from other objects in the environment. (Additional Embodiments)

[0150] In a first aspect, a method for automatically snapping a target virtual object to a destination object in a user's three-dimensional (3D) environment, under control of an augmented reality (AR) system comprising computer hardware, the AR system configured to enable user interaction with objects in the user's 3D environment, the AR system comprising a user input device, the method including: identifying a target virtual object and a destination object in the user's 3D environment, the target virtual object being associated with a first orientation and a first location, and the destination object being associated with a second orientation and a second location; calculating a distance between the target virtual object and the destination object based at least in part on the first location and the second location; comparing the distance to a threshold distance; automatically attaching the target virtual object to a surface of the destination object in response to comparing the distance being less than or equal to the threshold distance; and automatically orienting the target virtual object to align the target virtual object with the destination object based at least in part on the first orientation and the second orientation.

[0151] In a second aspect, the step of identifying the target virtual object and the destination object is based on at least one of head pose, eye pose, body pose, or hand gesture.

[0152] In a third aspect, the method of any one of aspects 1-2, wherein the destination object includes at least one of a physical object or a virtual object.

[0153] In a fourth aspect, the method of aspect 3, wherein the destination object comprises a wall or a table.

[0154] In a fifth aspect, the method of any one of aspects 1-4, wherein the first orientation or the second orientation includes at least one of a vertical orientation or a horizontal orientation.

[0155] In a sixth aspect, the method of any one of aspects 1-5, wherein calculating the distance includes calculating a displacement between the target virtual object and the destination object.

[0156] In a seventh aspect, the method of any one of aspects 1-6, wherein the threshold distance is zero.

[0157] In an eighth aspect, the method of any one of aspects 1-7, wherein the step of automatically orienting the target virtual object and aligning the target virtual object with the destination object includes at least one of the steps of automatically orienting the target virtual object so that a surface normal of the target virtual object faces the AR system, automatically orienting the target virtual object so that a surface normal of the target virtual object is perpendicular to the surface of the destination object, or automatically orienting the target virtual object so that a surface normal of the target virtual object is parallel to the normal of the destination object.

[0158] In a ninth aspect, the method of any one of aspects 1-8 further comprises assigning a focus indicator to at least one of the target virtual object or the destination object.

[0159] In a tenth aspect, an augmented reality system includes computer hardware and a user input device, the augmented reality system configured to implement a method according to any one of aspects 1-9.

[0160] In an eleventh aspect, a method for automatically snapping a target virtual object to a destination object in a user's three-dimensional (3D) environment, under control of an augmented reality (AR) system, comprising computer hardware, the AR system configured to enable user interaction with an object in the user's 3D environment, the AR system comprising a user input device and an orientation sensor configured to measure a user orientation, the method comprising: identifying a target virtual object in the user's 3D environment, the target virtual object being associated with a first location and a first orientation; and receiving, using the orientation sensor, an indication from the user to move the target virtual object to the destination object, the destination object being associated with a second location and a second orientation. calculating a trajectory between the target virtual object and the destination object based at least in part on the first location and the second location, and moving the target virtual object along the trajectory toward the destination object; calculating a distance between the target virtual object and the destination object based at least in part on the current location and the second location of the target virtual object; comparing the distance to a threshold distance; and automatically associating the target virtual object with the destination object in response to comparing the distance to be less than or equal to the threshold distance; and automatically orienting the target virtual object to align the target virtual object with the destination object based at least in part on the first orientation and the second orientation.

[0161] In a twelfth aspect, the method of aspect 11, wherein the attitude sensor includes at least one of an outward-facing imaging system, an inertial measurement unit, or an inward-facing imaging system.

[0162] In a thirteenth aspect, the method of aspect 12 further includes a step of assigning a focus indicator to a user's current position, the user's current position being determined at least in part based on the user's posture or a position associated with a user input device.

[0163] In a fourteenth aspect, the method of aspect 13, wherein the pose includes at least one of a head pose, an eye pose, or a body pose.

[0164] In a fifteenth aspect, the method of aspect 14, wherein receiving an indication from the user to move the target virtual object to the destination object includes identifying a change in the user's posture using a posture sensor, identifying the destination object based at least in part on the user's posture, and receiving a confirmation from the user to move the target virtual object to the destination object.

[0165] In a sixteenth aspect, the method of aspect 13, wherein receiving an indication from the user that the target virtual object is to be moved to the destination object includes at least one of receiving an indication of the destination object from a user input device or receiving a confirmation from the user that the target virtual object is to be moved to the destination object.

[0166] In a seventeenth aspect, the method of any one of aspects 15-16, wherein the confirmation includes at least one of a change in the user's posture or a hand gesture on a user input device.

[0167] In an eighteenth aspect, an augmented reality system includes computer hardware, a user input device, and an orientation sensor, and is configured to implement the method of any one of aspects 11-17.

[0168] In a nineteenth aspect, a method for snapping a target virtual object to a destination object in a user's three-dimensional (3D) environment, the method comprising: under control of an augmented reality (AR) system having computer hardware, the AR system configured to enable user interaction with objects in the user's 3D environment; identifying a target virtual object and a destination object in the user's 3D environment; receiving an indication to link the target virtual object to the destination object; determining an affordance associated with at least one of the target virtual object or the destination object; automatically orienting the target virtual object based, at least in part, on the affordance; and automatically linking the target virtual object to the destination object.

[0169] In a twentieth aspect, the method of aspect 19, wherein the step of identifying the target virtual object and the destination object is based on at least one of head pose, eye pose, body pose, hand gesture, or input from a user input device.

[0170] In a 21st aspect, the method of any one of aspects 19-20, wherein the destination object includes at least one of a physical object or a virtual object.

[0171] In a twenty-second aspect, the method of aspect 21, wherein the destination object includes a vertical or horizontal surface.

[0172] In a 23rd aspect, the method of any one of aspects 19-22, wherein the step of receiving an indication to link the target virtual object to the destination object includes one or more of the steps of detecting a change in the user's posture or receiving an indication of the destination object from a user input device.

[0173] In a 24th aspect, the method of aspect 23, wherein the pose includes at least one of a head pose, an eye pose, or a body pose.

[0174] In a 25th aspect, the method of any one of aspects 19-24, wherein the affordance is determined based on one or more of the function, orientation, type, location, shape, size, or environment of the target virtual object or destination object.

[0175] In a 26th aspect, the method of any one of aspects 19-25, wherein the step of automatically orienting the target virtual object includes one or more of the following steps: automatically orienting the target virtual object so that a surface normal of the target virtual object faces the AR system; automatically orienting the target virtual object so that a surface normal of the target virtual object is perpendicular to the surface of the destination object; or automatically orienting the target virtual object so that a surface normal of the target virtual object is parallel to the normal of the destination object.

[0176] In a 27th aspect, the method of any one of aspects 19-26, wherein the step of automatically linking the target virtual object to the destination object is performed by simulating an attractive force between the target virtual object and the destination object.

[0177] In a twenty-eighth aspect, the method of aspect 27, wherein the attractive force comprises one or more of a gravitational force or a magnetic force.

[0178] In a 29th aspect, an augmented reality system includes computer hardware configured to implement a method according to any one of aspects 19-28.

[0179] In a 30th aspect, a method for detaching a target virtual object from another object in a three-dimensional (3D) environment, under control of an augmented reality (AR) system comprising computer hardware, the AR system configured to enable user interaction with an object within a user's field of view (FOR), the FOR constituting a part of an environment surrounding the user that can be perceived by the user through the AR system, the method including: receiving a selection of the target virtual object by the user, the target virtual object being associated with a first position within the user's FOR; displaying a focus indicator associated with the target virtual object to the user; receiving an indication from the user to move the target virtual object; displaying the focus indicator to the user at an updated position based, at least in part, on the indication; calculating a distance between the first position of the target virtual object and the updated position of the focus indicator; comparing the distance to a threshold distance; and, in response to the comparison that the distance is greater than or equal to the threshold distance, moving the target virtual object from the first position to a second position associated with the updated position of the focus indicator; and displaying the target virtual object to the user at the second position.

[0180] In a thirty-first aspect, the method described in aspect 30, wherein the step of receiving a selection of a target virtual object by a user includes at least one of detecting a change in the user's posture or receiving input from a user input device.

[0181] In a 32nd aspect, the method of any one of aspects 30-31, wherein the other object includes at least one of a physical object or a virtual object.

[0182] In a thirty-third aspect, the method of aspect 32 is described, wherein the other object includes a wall or a table.

[0183] In a 34th aspect, the method of any one of aspects 30-33, wherein the step of receiving an indication to move the target virtual object includes at least one of detecting movement of a user input device, detecting a hand gesture on the user input device, or detecting a change in the user's posture.

[0184] In a 35th aspect, the method of any one of aspects 31-34, wherein the user's posture includes head posture, eye posture, or body posture.

[0185] In a 36th aspect, the method of any one of aspects 30-35, wherein the step of calculating the distance includes a step of calculating a displacement between the first position and the updated position.

[0186] In a 37th aspect, the method of any one of aspects 30-36, wherein the threshold distance is user-defined.

[0187] In aspect 38, the method of any one of aspects 30-37, wherein the step of displaying the target virtual object at a second position to the user includes determining an orientation associated with the target virtual object before the target virtual object touches the other object, and displaying the target virtual object to the user with the orientation at the second position.

[0188] In a thirty-ninth aspect, a method according to any one of aspects 1-38, wherein the step of displaying the target virtual object at a second position to the user includes the steps of determining an orientation associated with the target virtual object at the first position, and displaying the target virtual object to the user with the orientation at the second position.

[0189] In a fortieth aspect, a method for associating and unassociating a target virtual object from another object in a three-dimensional (3D) environment, under control of an augmented reality (AR) system comprising computer hardware, the AR system configured to enable user interaction with objects within a user's field of view (FOR), the FOR constituting a portion of an environment surrounding the user that can be perceived by the user via the AR system, includes receiving a selection of the target virtual object, the target virtual object associated with a first position within the user's FOR, and at least a portion of the target virtual object touching the other object; and notifying the user of a focus indicator associated with the target virtual object. receiving, from a user, an indication to disassociate the target virtual object from the other object; displaying, to the user, a focus indicator at an updated position based at least in part on the indication; determining whether the indication satisfies a threshold condition for disassociating the target virtual object from the other object; in response to determining that the threshold condition is met, displaying, to the user, the target virtual object at a second position associated with the updated position of the focus indicator; and in response to determining that the threshold condition is not met, displaying, to the user, the target virtual object at the first position.

[0190] In a forty-first aspect, the method described in aspect 40, wherein the step of receiving the selection of the target virtual object includes at least one of detecting a change in the user's posture or receiving input from a user input device.

[0191] In a forty-second aspect, the method of any one of aspects 40-41, wherein the other object includes at least one of a physical object or a virtual object.

[0192] In a forty-third aspect, the method of aspect 42 is described, wherein the other object includes a wall or a table.

[0193] In a 44th aspect, the method of any one of aspects 40-43, wherein the step of receiving an indication to tie or untie the target virtual object includes at least one of detecting movement of a user input device, detecting a hand gesture on the user input device, or detecting a change in the user's posture.

[0194] In a forty-fifth aspect, the method of any one of aspects 41-44, wherein the user's posture includes head posture, eye posture, or body posture.

[0195] In a forty-sixth aspect, the method of any one of aspects 41-45, wherein the threshold conditions for disassociating the target virtual object from the other objects include at least one of: a distance between the first position and the updated position being greater than or equal to a threshold distance; a velocity for moving from the first position to the updated position being greater than or equal to a threshold velocity; an acceleration for moving away from the first position being greater than or equal to a threshold acceleration; or a jerk for moving away from the first position being greater than or equal to a threshold jerk.

[0196] In a forty-seventh aspect, the method of any one of aspects 41-46, wherein the step of displaying the target virtual object at a second position to the user includes the steps of determining an orientation associated with the target virtual object before the target virtual object touches the other object, and displaying the target virtual object to the user with the orientation at the second position.

[0197] In aspect 48, a method according to any one of aspects 41-47, wherein the step of displaying the target virtual object at a second position to the user includes the steps of determining an orientation associated with the target virtual object at the first position, and displaying the target virtual object to the user with the orientation at the second position.

[0198] In a forty-ninth aspect, a method for disassociating a target virtual object from another object in a user's three-dimensional (3D) environment, under control of an augmented reality (AR) system comprising computer hardware, the AR system configured to enable user interaction with objects in the user's 3D environment, the method including: receiving a selection of the target virtual object, the target virtual object being associated with another object in the 3D environment at an initial position; displaying to the user a focus indicator associated with the target virtual object at the initial position; receiving from the user an indication to disassociate the target virtual object from the other object; displaying to the user the focus indicator at an updated position based, at least in part, on the indication; determining whether the indication satisfies a threshold condition for disassociating the target virtual object from the other object; and in response to determining that the threshold condition is met, disassociating the target virtual object based, at least in part, on the indication.

[0199] In a 50th aspect, the method described in aspect 49, wherein the step of receiving the selection of the target virtual object includes at least one of detecting a change in the user's posture or receiving input from a user input device.

[0200] In a 51st aspect, the method of any one of aspects 49-50, wherein the other object includes at least one of a physical object or a virtual object.

[0201] In a 52nd aspect, the method described in aspect 51, wherein the other object includes a vertical surface or a horizontal surface.

[0202] In a 53rd aspect, the method of any one of aspects 49-52, wherein the step of receiving an indication to associate or unassociate the target virtual object from another object includes at least one of detecting movement of a user input device, detecting a hand gesture on a user input device, or detecting a change in the user's posture.

[0203] In a 54th aspect, the method of any one of aspects 49-53, wherein the threshold conditions for disassociating the target virtual object from other objects include at least one of: a distance between the initial position and the updated position being equal to or greater than a threshold distance; a velocity for moving from the initial position to the updated position being equal to or greater than a threshold velocity; an acceleration for moving away from the initial position being equal to or greater than a threshold acceleration; or a jerk for moving away from the initial position being equal to or greater than a threshold jerk.

[0204] In a 55th aspect, the method of any one of aspects 49-54, wherein the step of attaching and detaching the target virtual object is performed by simulating a physical force.

[0205] In a 56th aspect, the method of aspect 55, wherein the physical force includes at least one of gravity, magnetic force, friction, or elasticity.

[0206] In a fifty-seventh aspect, an augmented reality system includes computer hardware configured to implement a method according to any one of aspects 30-56.

[0207] In a fifty-eighth aspect, an augmented reality (AR) system for automatically repositioning a virtual object within a three-dimensional (3D) environment includes an AR display configured to present virtual content within a 3D view; and a hardware processor in communication with the AR display, the system including: identifying a target virtual object within a user's 3D environment, the target virtual object being assigned a vector representing a first location and a first orientation; receiving an indication to link the target virtual object to a destination object; the destination object being assigned at least one vector representing a second location and a second orientation; calculating a trajectory between the target virtual object and the destination object based, at least in part, on the first location and the second location; and linking the target virtual object to the destination object. and a hardware processor programmed to: move the target virtual object along a trajectory toward a destination object; track a current location of the target virtual object; calculate a distance between the target virtual object and the destination object based at least in part on the current location and a second location of the target virtual object; determine whether the distance between the target virtual object and the destination virtual object is less than a threshold distance; and, in response to comparing the distance to be less than or equal to the threshold distance, automatically link the target virtual object to the destination object; orient the target virtual object in a second orientation; and render the target virtual object at the second location with the second orientation by an AR display, wherein the target virtual object is overlaid on the destination object.

[0208] In a fifty-ninth aspect, the hardware processor is further programmed to analyze affordances of at least one of the target virtual object, the destination object, or the environment, and to automatically orient the target virtual object, the hardware processor is programmed to rotate the target virtual object and align a first normal of the target virtual object with a second normal of the destination object, in the AR system described in aspect 58.

[0209] In a 60th aspect, the AR system described in aspect 59, wherein the affordance includes at least one of function, orientation, type, location, shape, or size.

[0210] In aspect 61, an AR system described in any one of aspects 58-60, wherein to automatically connect the target virtual object, the hardware processor is programmed to simulate an attractive force between the target virtual object and the destination object, the attractive force including at least one of gravity, elastic force, adhesive force, or magnetic force.

[0211] In aspect 62, an AR system described in any one of aspects 58-61, wherein to calculate the distance, the hardware processor is programmed to calculate the displacement between the current location of the target virtual object and a second location associated with the destination object.

[0212] In a 63rd aspect, the AR system described in aspect 62, wherein the threshold distance is zero.

[0213] In aspect 64, an AR system described in any one of aspects 58-63, wherein the indication for linking the target virtual object is determined from at least one of an actuation of a user input device or a user posture.

[0214] In aspect 65, the hardware processor is further programmed to assign a focus indicator to a user's current position, the user's current position being determined, at least in part, based on the user's posture or a position associated with a user input device, in an AR system as described in aspect 64.

[0215] In aspect 66, the hardware processor is further programmed to receive an indication to disassociate the target virtual object from the destination object, the indication being associated with a change in the user's current position, and in response to determining that the threshold condition is met, determine whether the threshold condition for disassociating the target virtual object is met based at least in part on the received indication, disassociate the target virtual object from the destination object, move the target virtual object from a second location associated with the destination object to a third location, and render the target virtual object at the third location. An AR system described in any one of aspects 58-65.

[0216] In aspect 67, the AR system described in aspect 66, wherein in response to determining that the threshold condition is satisfied, the hardware processor is further programmed to move the target virtual object to a third location while retaining the second orientation for the target virtual object.

[0217] In aspect 68, the AR system described in aspect 67, wherein the third location corresponds to the position of the focus indicator, which corresponds to the user's current position.

[0218] In a sixty-ninth aspect, the AR system of aspect 68, wherein the threshold conditions for untying the target virtual object from another object include at least one of: a second distance between a second location where the target virtual object is tied to the destination object and the position of the focus indicator being equal to or greater than a second threshold distance; a velocity for moving from the second location to the position of the focus indicator being equal to or greater than a threshold velocity; an acceleration for moving away from the second location being equal to or greater than a threshold acceleration; or a jerk for moving away from the second location being equal to or greater than a threshold jerk.

[0219] In a seventieth aspect, the hardware processor is programmed to simulate physical forces when attaching and detaching the target virtual object from the destination object, the physical forces including at least one of friction or elasticity, in an AR system as described in aspect 69.

[0220] In a seventy-first aspect, a method for automatically repositioning a virtual object within a three-dimensional (3D) environment, under control of an augmented reality (AR) system comprising computer hardware, the AR system configured to enable user interaction with objects within the 3D environment, the method including: identifying a target virtual object within the user's 3D environment, the target virtual object having a first position and a first orientation; receiving an indication to reposition the target virtual object relative to a destination object; identifying parameters for repositioning the target virtual object; analyzing affordances associated with at least one of the 3D environment, the target virtual object, and the destination object; calculating values ​​of the parameters for repositioning the target virtual object based on the affordances; determining a second position and second orientation for the target virtual object and a movement of the target virtual object based on the values ​​of the parameters for repositioning the target virtual object; and rendering the target virtual object at the second position and second orientation and the movement of the target virtual object to reach the second position and second orientation from the first position and first orientation.

[0221] In aspect 72, the method described in aspect 71, wherein the step of repositioning the target object includes at least one of the steps of linking the target object to the destination object, reorienting the target object, or unlinking the target object from the destination object.

[0222] In a 73rd aspect, the method of any one of aspects 71-72, wherein the destination object is a physical object.

[0223] In aspect 74, a method according to any one of aspects 71-73 further comprising: determining whether an indication to reposition the target virtual object satisfies a threshold condition; and performing the calculation, determination, and rendering in response to determining that the indication satisfies the threshold condition.

[0224] In a seventy-fifth aspect, the method of any one of aspects 74, wherein the threshold condition includes a distance between the target virtual object and the destination object.

[0225] In a 76th aspect, a method according to any one of aspects 71-75, wherein one or more physical attributes are assigned to the target virtual object, and movement of the target virtual object is determined by simulating an interaction between the target virtual object, the destination object, and the environment based on the physical attributes of the target virtual object.

[0226] In a seventy-seventh aspect, the method described in aspect 76, wherein the one or more physical attributes assigned to the target virtual object include at least one of mass, size, density, topology, hardness, elasticity, or electromagnetic attributes. (Other considerations)

[0227] Each of the processes, methods, and algorithms described herein and / or depicted in the accompanying figures may be embodied in code modules executed by one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to execute specific and particular computer instructions, and thereby may be fully or partially automated. For example, a computing system may include a general-purpose computer (e.g., a server) or a special-purpose computer programmed with specific computer instructions, special-purpose circuitry, etc. Code modules may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language. In some implementations, particular operations and methods may be performed by circuitry specific to a given function.

[0228] Furthermore, certain implementations of the functionality of the present disclosure may be sufficiently mathematically, computationally, or technically complex that special-purpose hardware (utilizing appropriate specialized executable instructions) or one or more physical computing devices may be required to implement the functionality, e.g., due to the amount or complexity of the calculations involved, or to provide results in substantially real time. For example, a video may contain many frames, each frame may have millions of pixels, and specifically programmed computer hardware may be required to process the video data to provide the desired image processing task or application in a commercially reasonable amount of time.

[0229] Code modules or any type of data may be stored on any type of non-transitory computer-readable medium, such as physical computer storage devices, including hard drives, solid-state memory, random-access memory (RAM), read-only memory (ROM), optical disks, volatile or non-volatile storage devices, combinations of the same, and / or the like. The methods and modules (or data) may also be transmitted as data signals (e.g., as part of a carrier wave or other analog or digital propagated signal) generated over various computer-readable transmission media, including wireless-based and wired / cable-based media, and may take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed processes or process steps may be stored, persistently or otherwise, in any type of non-transitory tangible computer storage device or communicated via a computer-readable transmission medium.

[0230] Any process, block, state, step, or functionality in the flow diagrams described herein and / or depicted in the accompanying figures should be understood as potentially representing a code module, segment, or portion of code, comprising one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in the process. Various processes, blocks, states, steps, or functionality can be combined, rearranged, added, deleted, modified, or otherwise changed from the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may perform some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the blocks, steps, or states associated therewith can be performed in other suitable sequences, e.g., serially, in parallel, or in some other manner. Tasks or events may be added to or removed from the disclosed exemplary embodiments. Furthermore, the separation of various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems may generally be integrated together in a single computer product or packaged in multiple computer products. Many implementation variations are possible.

[0231] The processes, methods, and systems can be implemented in a network (or distributed) computing environment. Network environments include enterprise-wide computer networks, intranets, local area networks (LANs), wide area networks (WANs), personal area networks (PANs), cloud computing networks, crowdsourced computing networks, the Internet, and the World Wide Web. The network can be a wired or wireless network or any other type of communication network.

[0232] The systems and methods of the present disclosure each have several innovative aspects, none of which is solely responsible for or required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure. Various modifications of the implementations described in the present disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations without departing from the spirit or scope of the present disclosure. Therefore, the claims are not intended to be limited to the implementations shown herein, but are to be accorded the widest scope consistent with the present disclosure, the principles, and novel features disclosed herein.

[0233] Certain features described herein in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented separately in multiple implementations or in any suitable subcombination. Furthermore, while features may be described above as operative in a combination and may even be initially claimed as such, one or more features from the claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination. No single feature or group of features is required or essential to every embodiment.

[0234] Conditional statements used herein, such as "can," "could," "might," "may," "eg," and the like, among others, are intended to generally convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context as used. Thus, such conditional statements are not generally intended to imply that features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps should be included or performed in any particular embodiment, with or without authorial input or prompting. The terms "comprise," "include," "have," and the like are synonymous and used inclusively in a non-limiting manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not its exclusive sense), so, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Additionally, the articles "a," "an," and "the," as used in this application and the appended claims, should be interpreted to mean "one or more" or "at least one," unless otherwise specified.

[0235] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single elements. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Transitional phrases such as "at least one of X, Y, and Z," unless specifically stated otherwise, are generally understood differently in the context in which they are used to convey that an item, term, etc. may be at least one of X, Y, or Z. Thus, such transitional phrases generally are not intended to suggest that an embodiment requires that at least one of X, at least one of Y, and at least one of Z, respectively, be present.

[0236] Similarly, while operations may be depicted in the figures in a particular order, it should be recognized that such operations need not be performed in the particular order shown, or in sequential order, or that all of the depicted operations need not be performed to achieve desirable results. Furthermore, the figures may diagrammatically depict one or more example processes in the form of a flowchart. However, other operations not depicted may be incorporated within the diagrammatically depicted example methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any of the depicted operations. Additionally, operations may be rearranged or reordered in other implementations. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results.

Claims

1. 1. An augmented reality (AR) system for automatically repositioning a virtual object within a three-dimensional (3D) environment, the AR system comprising: an AR display configured to present virtual content within a 3D view; a hardware processor in communication with the AR display; Equipped with The hardware processor includes: identifying a target virtual object within the user's 3D environment, the target virtual object being assigned a vector representing a first location and a first orientation; receiving an indication to link the target virtual object to a destination object, the destination object being assigned at least one vector representing a second location and a second orientation; calculating a trajectory between the target virtual object and the destination object based at least in part on the first location and the second location; moving the target virtual object along the trajectory toward the destination object; tracking a current location of the target virtual object; calculating a distance between the target virtual object and the destination object based at least in part on the current location and the second location of the target virtual object; calculating at least one affordance associated with at least one of the 3D environment, the target virtual object, or the destination object, the affordance including attributes used to simulate an interaction with one or more of the 3D environment, the target virtual object, or the destination object using laws of physics; determining whether the distance between the target virtual object and the destination object is less than a threshold distance, the threshold distance varying based on the at least one affordance associated with the at least one of the 3D environment, the target virtual object, or the destination object; automatically linking the target virtual object to the destination object and orienting the target virtual object in the second orientation in response to comparing the distance less than or equal to the threshold distance, wherein to automatically link the target virtual object, the hardware processor is programmed to simulate an attractive force between the target virtual object and the destination object based on the affordance, the simulated attractive force comprising at least one of a simulated gravity force, a simulated elastic force, or a simulated adhesive force in accordance with the law of physics associated with the at least one affordance; rendering, by the AR display, the target virtual object at the second location with the second orientation, the target virtual object being overlaid on the destination object; and The AR system is programmed to perform the following:

2. The hardware processor further comprises:

2. The AR system of claim 1, wherein the AR system is programmed to automatically orient the target virtual object, and wherein the hardware processor is programmed to rotate the target virtual object to align a first normal of the target virtual object with a second normal of the destination object.

3. The AR system of claim 2 , wherein the affordances include at least one of a function, an orientation, a type, a location, a shape, or a size.

4. 2. The AR system of claim 1, wherein to calculate the distance, the hardware processor is programmed to calculate a displacement between the current location of the target virtual object and the second location associated with the destination object.

5. The AR system of claim 4 , wherein the threshold distance is zero.

6. The AR system of claim 1 , wherein the indication to bind the target virtual object is determined from at least one of an actuation of a user input device or a user pose.

7. the hardware processor is further programmed to assign a focus indicator to the user's current position; The AR system of claim 6 , wherein the current position of the user is determined based at least in part on the pose of the user or a position associated with the user input device.

8. The hardware processor further comprises: receiving an indication to uncouple the target virtual object from the destination object, the indication being associated with a change in the user's current location; determining whether a threshold condition for disassociating the target virtual object is met based at least in part on the received indication; and in response to determining that the threshold condition is satisfied, disassociating the target virtual object from the destination object, moving the target virtual object from the second location associated with the destination object to a third location, and rendering the target virtual object at the third location. The AR system of claim 1 , programmed to execute the following:

9. 9. The AR system of claim 8, wherein in response to determining that the threshold condition is met, the hardware processor is further programmed to: move the target virtual object to the third location while retaining the second orientation for the target virtual object.

10. The AR system of claim 9 , wherein the third location corresponds to a position of a focus indicator that corresponds to the current position of the user.

11. The threshold condition for disassociating the target virtual object from other objects includes: a second distance between the second location where the target virtual object is bound to the destination object and the position of the focus indicator is greater than or equal to a second threshold distance; a speed for moving the focus indicator from the second location to the position is equal to or greater than a threshold speed; the acceleration to move away from the second location is greater than or equal to a threshold acceleration; or a jerk for moving away from the second location equal to or greater than a threshold jerk; The AR system of claim 10 , comprising at least one of:

12. 12. The AR system of claim 11, wherein the hardware processor is programmed to simulate physical forces when attaching and detaching the target virtual object from the destination object, the physical forces including at least one of friction or elasticity.

Citation Information

Patent Citations

  • Real object interference display device

    JP2008307091A

  • Game apparatus and game program

    JP2010279742A

  • A system and method for haptic messaging based on physical laws.

    JP2011528476A

  • Method and system for generating augmented reality using display in motor vehicle

    JP2013131222A

  • Interactions of virtual objects with surfaces

    US20140333666A1