Context-aware user interface menus
The wearable system addresses VR/AR/MR challenges by analyzing user context and environment to present a relevant subset of virtual objects, enhancing depth perception and comfort through a 3D display with multiple depth planes.
Patent Information
- Application Number
- JP2023201574
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-08-29
- Filing Date
- 2023-11-29
- Publication Date
- 2025-12-01
- Estimated Expiration
- 2037-05-18
AI Technical Summary
Modern VR, AR, and MR technologies face challenges in providing comfortable and natural-feeling presentations of virtual image elements due to the complexity of the human visual perception system, leading to issues like unstable imaging, eye strain, and difficulty in quickly identifying relevant virtual objects amidst a large number of options.
A wearable system that analyzes user posture, environment, and physiological parameters to identify contextual information, allowing for the presentation of a relevant subset of virtual objects based on the user's surroundings and psychological state, using a 3D display with multiple depth planes to enhance depth perception.
The system provides a more comfortable and realistic VR/AR/MR experience by presenting a tailored set of virtual objects and enhancing depth perception, reducing eye strain and improving object recognition.
Smart Images

Figure 0007778127000001 
Figure 0007778127000002 
Figure 0007778127000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 62 / 339,572, filed May 20, 2016, and entitled "CONTEXTUAL AWARENESS OF USER INTERFACE MENUS," and U.S. Provisional Application No. 62 / 380,869, filed August 29, 2016, and entitled "AUGMENTED COGNITION USER INTERFACE," the disclosures of which are incorporated herein by reference in their entireties.
[0002] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly to presenting and selecting virtual objects based on contextual information. [Background technology]
[0003] Modern computing and display technologies have facilitated the development of systems for so-called “virtual reality,” “augmented reality,” or “mixed reality” experiences, in which digitally reproduced images, or portions thereof, are presented to a user in a manner that appears or can be perceived as real. Virtual reality or “VR” scenarios typically involve the presentation of digital or virtual image information without transparency to other actual real-world visual inputs. Augmented reality or “AR” scenarios typically involve the presentation of digital or virtual image information as an augmentation to the visualization of the real world around the user. Mixed reality or “MR” relates to the merging of real and virtual worlds to create new environments in which physical and virtual objects coexist and interact in real time. Consequently, the human visual perception system is highly complex, making it challenging to create VR, AR, or MR technologies that facilitate comfortable, natural-feeling, and rich presentations of virtual image elements among other virtual or real-world image elements. The systems and methods disclosed herein address various challenges associated with VR, AR, and MR technologies. Summary of the Invention [Means for solving the problem]
[0004] In one embodiment, a wearable system for generating virtual content within a user's three-dimensional (3D) environment is disclosed. The wearable system may include an augmented reality display configured to present the virtual content to a user in a 3D view, a posture sensor configured to obtain user position or orientation data, analyze the position or orientation data, and identify the user's posture, and a hardware processor in communication with the posture sensor and the display. The hardware processor may be programmed to: identify a physical object in the user's environment within the 3D environment based at least in part on the user's posture; receive an indication to initiate an interaction with the physical object; identify a set of virtual objects in the user's environment associated with the physical object; determine contextual information associated with the physical object; filter the set of virtual objects; identify a subset of the virtual objects from the set of virtual objects based on the contextual information; generate a virtual menu including the subset of virtual objects; determine a spatial location within the 3D environment for presenting the virtual menu based at least in part on the determined contextual information; and present the virtual menu at the spatial location via the augmented reality display.
[0005] In another embodiment, a method for generating virtual content within a user's three-dimensional (3D) environment is disclosed. The method can include analyzing data obtained from a posture sensor to identify a posture of the user, identifying an interactable object within the user's 3D environment based at least in part on the posture, receiving an indication to initiate an interaction with the interactable object, determining context information associated with the interactable object, selecting a subset of user interface actions from a set of user interface actions available on the interactable object based on the context information, and generating instructions for presenting the subset of user interface actions to the user in a 3D view.
[0006] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, drawings, and claims. Neither this summary nor the following detailed description purports to define or limit the scope of the inventive subject matter. The present specification also provides, for example, the following items: (Item 1) 1. A method for generating virtual content within a user's three-dimensional (3D) environment, the method comprising: analyzing data obtained from the posture sensor to identify a posture of the user; identifying walls within the user's 3D environment based at least in part on the pose; and receiving an indication to initiate an interaction with the wall; determining context information associated with the wall; selecting a subset of user interface actions from a set of user interface actions available on the wall based on the context information; generating instructions for presenting the subset of user interface actions to the user in a 3D view; A method comprising: (Item 2) Item 10. The method of item 1, wherein the pose includes at least one of eye gaze, head pose, or gesture. (Item 3) Item 10. The method of item 1, wherein identifying the walls includes performing a cone projection based on the user's head pose. (Item 4) Item 10. The method of item 1, wherein the context information includes an orientation of a surface of the wall, and selecting the subset of user interface actions includes identifying user interface actions that can be performed on a surface having the orientation. (Item 5) Item 10. The method of item 1, wherein the indication to initiate an interaction with the wall includes actuation of a user input device or a change in the user's posture. (Item 6) 2. The method of claim 1, further comprising receiving physiological parameters of the user and determining a psychological state of the user, the psychological state being part of the context information for selecting the subset of user interactions. (Item 7) 7. The method of claim 6, wherein the physiological parameter relates to at least one of heart rate, mydriasis, galvanic skin response, blood pressure, electroencephalogram, respiratory rate, or eye movement. (Item 8) Item 10. The method of claim 1, wherein generating instructions for presenting the subset of user interactions to the user in a 3D view includes generating a virtual menu including the subset of user interface actions, determining a spatial location of the virtual menu based on characteristics of the wall, and generating display instructions for presentation of the virtual menu at the spatial location within the user's 3D environment. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 depicts an illustration of a mixed reality scenario with a virtual reality object and a physical object viewed by a person. [Figure 2] FIG. 2 illustrates diagrammatically an example of a wearable system. [Figure 3] FIG. 3 diagrammatically illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes. [Figure 4] FIG. 4 illustrates diagrammatically an embodiment of a waveguide stack for outputting image information to a user. [Figure 5] FIG. 5 shows an exemplary output beam that may be output by a waveguide. [Figure 6]FIG. 6 is a schematic diagram showing an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal stereoscopic display, image, or light field. [Figure 7] FIG. 7 is a block diagram of an embodiment of a wearable system. [Figure 8] FIG. 8 is a process flow diagram of an embodiment of a method for rendering virtual content in relation to recognized objects. [Figure 9] FIG. 9 is a block diagram of another embodiment of a wearable system. [Figure 10] FIG. 10 is a process flow diagram of an example method for determining user input to a wearable system. [Figure 11] FIG. 11 is a process flow diagram of an embodiment of a method for interacting with a virtual user interface. [Figure 12] FIG. 12 illustrates an example of user interaction with a virtual user interface in an office environment. [Figure 13] 13 and 14 illustrate an example of user interaction with a virtual user interface within a living room environment. [Figure 14] 13 and 14 illustrate an example of user interaction with a virtual user interface within a living room environment. [Figure 15] FIG. 15 illustrates an example of user interaction with a virtual user interface within a bedroom environment. [Figure 16] FIG. 16 is a flowchart of an exemplary method for generating a virtual menu based on context information. [Figure 17] FIG. 17 is a flowchart of an exemplary method for selecting virtual content based, at least in part, on a user's physiological and / or psychological state. DETAILED DESCRIPTION OF THE INVENTION
[0008] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure. Additionally, the figures in this disclosure are for illustrative purposes and are not to scale. (overview)
[0009] Modern computer interfaces support a wide range of functionality. However, users can become overwhelmed by the number of options and be unable to quickly identify objects of interest. In AR / VR / MR environments, the user's field of view (FOV), as perceived through the AR / VR / MR display of a wearable device, can be smaller than the user's natural FOV, making it more difficult to present a relevant set of virtual objects than in typical computing environments.
[0010] The wearable system described herein can alleviate this problem by analyzing the user's environment and providing a smaller and more relevant subset of functionality on the user interface. The wearable system may provide this subset of functionality based on contextual information, such as the user's environment or objects within the user's environment. The wearable system can recognize physical objects (such as tables and walls) and their relationship to the environment. For example, the wearable system can recognize that a cup should be placed on a table (instead of a wall) and a painting should be placed on a vertical wall (instead of a table). Based on this relationship, the wearable system can project a virtual cup onto a table in the user's room and a virtual painting onto the vertical wall.
[0011] In addition to the relationship between an object and its environment, other factors such as the object's orientation (e.g., horizontal or vertical), the nature of the user's environment, and previous usage patterns (time, location, etc.) can also be used to determine the virtual object to be displayed by the wearable system. The nature of the user's environment can include whether it is a private environment (e.g., in the user's home or office) where the user can interact with the wearable device in a relatively secure and private manner, or a public environment where there may be others nearby (where the user may not wish their interaction with the device to be seen or heard). The distinction between private and public environments is not exclusive. For example, a park is a public environment if many people are nearby, but can be a private environment if the user is alone or no others are nearby. The distinction may be made, at least in part, based on the number of people nearby, their proximity to the user, their relationship to the user (e.g., whether they are friends, family, or strangers), etc. Additionally or alternatively, the wearable system may identify a label (such as a fiducial mark) associated with the object. The label may contain information about a virtual object (such as an item on a virtual menu) that should be displayed for the object associated with the reference marker.
[0012] In some embodiments, the wearable system may also include various physiological sensors. These sensors may measure or estimate the user's physiological parameters, such as heart rate, respiratory rate, galvanic skin response, blood pressure, and brainwave state. These sensors may be used in conjunction with an inward-facing imaging system to determine the user's eye movement and pupil dilation, which may also be reflective of the user's physiological or psychological state. Data obtained by the physiological sensors or the inward-facing imaging system may be analyzed by the wearable system to determine the user's psychological state, such as mood and interests. The wearable system may use the user's psychological state as part of the contextual information and present a set of virtual objects based, at least in part, on the user's psychological state. (Example of a 3D display for a wearable system)
[0013] A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present 2D or 3D virtual images to a user. The images may be still images, frames of video, or videos, in combination or the like. A wearable system can include a wearable device that, alone or in combination, can present a VR, AR, or MR environment for user interaction. The wearable device can be a head-mounted device (HMD), which is used synonymously with AR device (ARD). The wearable device may be in the form of a helmet, glasses, a headset, or any other wearable configuration.
[0014] Figure 1 depicts an illustration of a mixed reality scenario involving certain virtual reality objects and certain physical objects viewed by a person. In Figure 1, an MR scene 100 is depicted in which a user of the MR technology sees a real-world park-like setting 110 featuring people, trees, buildings in the background, and a concrete platform 120. In addition to these items, the user of the MR technology also perceives as "seeing" a robotic figure 130 standing on the real-world platform 120 and a flying, cartoon-like avatar character 140 that appears to be an anthropomorphic bumblebee, although these elements do not exist in the real world.
[0015] In order for a 3D display to produce a true depth sensation, or more specifically, a simulated sensation of surface depth, it may be desirable for the display to generate, for each point in its field of view, an accommodation response that corresponds to that point's virtual depth. If the accommodation response to a display point does not correspond to that point's virtual depth as determined by convergence and stereoscopic binocular depth cues, the human eye may experience accommodation conflict, resulting in unstable imaging, adverse eye strain, headaches, and, in the absence of accommodative information, a near-complete lack of surface depth.
[0016] VR, AR, and MR experiences can be provided by a display system having a display that provides a viewer with images corresponding to multiple depth planes. The images may be different for each depth plane (e.g., providing slightly different presentations of a scene or object) and may be focused separately by the viewer's eyes, thereby serving to provide depth cues to the user based on the ocular accommodation required to focus on different image features of a scene located on different depth planes, or based on observing different image features on different depth planes that are out of focus. As discussed elsewhere herein, such depth cues provide a believable perception of depth.
[0017] FIG. 2 illustrates an example of a wearable system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems to support the functionality of the display 220. The display 220 may be coupled to a frame 230, which is wearable by a user, wearer, or viewer 210. The display 220 can be positioned directly in front of the eyes of the user 210. The display 220 can present AR / VR / MR content to the user. The display 220 can include a head-mounted display (HMD) worn on the user's head. In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the user's ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable sound control).
[0018] The wearable system 200 may include an outward-facing imaging system 464 (shown in FIG. 4 ) that observes the world in the user's surrounding environment. The wearable system 200 may also include an inward-facing imaging system 462 (shown in FIG. 4 ) that can track the user's eye movements. The inward-facing imaging system may track either one eye's movements or both eyes' movements. The inward-facing imaging system 462 may be mounted to the frame 230 and may be in electrical communication with a processing module 260 or 270 that may process image information acquired by the inward-facing imaging system and determine, for example, the pupil diameter or orientation of the user's 210 eyes, eye movements, or eye posture.
[0019] As an example, the wearable system 200 can capture an image of the user's posture using the outward-facing imaging system 464 or the inward-facing imaging system 462. The image may be a still image, a frame of video, or video, a combination thereof, or the like.
[0020] In some embodiments, the wearable system 200 can include one or more physiological sensors 232. Examples of such sensors include sensors configured for ophthalmic testing, such as confocal microscopy sensors, electronystagmography (ENG) sensors, electrooculography (EOG) sensors, electroretinogram (ERG) sensors, laser Doppler flowmeter (LDF) sensors, photoacoustic imaging and pressure reading sensors, two-photon excitation microscopy sensors, and / or ultrasound sensors. Other examples of sensors include sensors configured for other electrodiagnostic techniques, such as electrocardiogram (ECG) sensors, electroencephalogram (EEG) sensors, electromyography (EMG) sensors, electrophysiology testing (EP) sensors, event-related potential (ERP) sensors, near-infrared functional neuroimaging (fNIR) sensors, low-resolution electromagnetic tomography of the brain (LORETA) sensors, and / or optical coherence tomography (OCT) sensors. Still other examples of sensors 232 include physiological sensors such as blood glucose sensors, blood pressure sensors, electrodermal sensors, photoplethysmography devices, sensing devices for computer-assisted auscultation, galvanic skin response sensors, and / or temperature sensors. Sensors 232 may also include CO2 monitoring sensors, respiratory rate sensors, end-tidal CO2 sensors, and / or breathalyzers.
[0021] An example of sensor 232 is shown diagrammatically as connected to frame 230. This connection may take the form of physical attachment to frame 230 and may be anywhere on frame 230. As an example, sensor 232 may be mounted on frame 230 so as to be positioned adjacent the user's temple or at the contact point between frame 230 and the user's nose. As another example, sensor 232 may be positioned on a portion of frame 230 that extends over the user's ear. In some other embodiments, sensor 232 may extend from frame 230 and contact user 210. For example, sensor 232 may touch a part of the user's body (such as the user's arm) and connect to frame 230 via a wired connection. In other embodiments, sensor 232 may not be physically attached to frame 230. Rather, sensor 232 may communicate with wearable system 200 via a wireless connection. In some embodiments, the wearable system 200 may have the form of a helmet, and the sensor 232 may be positioned towards the top or sides of the user's head.
[0022] In some implementations, the sensor 232 provides a direct measurement of a physiological parameter that is used as context information by the wearable system. For example, a heart rate sensor may directly measure the user's heart rate. In other implementations, the sensor 232 (or a group of sensors) may provide a measurement that is used to estimate another physiological parameter. For example, stress may be estimated as a combination of heart rate measurements and galvanic skin response measurements. Statistical techniques can be applied to the sensor data to estimate a physiological (or psychological) state. As an example, the sensor data can be combined using machine learning techniques (e.g., decision trees, neural networks, support vector machines, Bayesian techniques) to estimate the user's state. The state estimate may provide a binary state (e.g., stress or baseline), multiple states (e.g., stress, baseline, or relaxed), or a probabilistic measure (e.g., the probability that the user is under stress). The physiological or psychological state may reflect any emotional state of the user, such as, for example, anxiety, stress, anger, love, boredom, despair or disappointment, happiness, sadness, loneliness, shock, or surprise.
[0023] The display 220 is operably coupled (250) to a local data processing module 260, which may be mounted in a variety of configurations, such as fixedly attached to the frame 230, fixedly attached to a helmet or hat worn by the user, built into headphones, or otherwise removably attached to the user 210 (e.g., in a backpack configuration, in a belt-connected configuration), such as by wired or wireless connection.
[0024] The local processing and data module 260 may comprise a hardware processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storing data. The data may include a) data captured from sensors (e.g., that may be operatively coupled to the frame 230 or otherwise attached to the user 210), such as an image capture device (e.g., a camera in an inward-facing or outward-facing imaging system), a microphone, an inertial measurement unit (IMU), an accelerometer, a compass, a global positioning system (GPS) unit, a wireless device, or a gyroscope, or b) data acquired or processed using the remote processing module 270 or the remote data repository 280, possibly for transmission to the display 220 after processing or retrieval. The local processing and data module 260 may be operably coupled to a remote processing module 270 or a remote data repository 280 by a communication link 262 or 264, such as via a wired or wireless communication link, so that these remote modules are available as resources to the local processing and data module 260. In addition, the remote processing module 280 and the remote data repository 280 may be operably coupled to each other.
[0025] In some embodiments, remote processing module 270 may comprise one or more hardware processors configured to analyze and process data and / or image information. In some embodiments, remote data repository 280 may comprise a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, allowing for fully autonomous use from the remote module.
[0026] The human visual system is complex and difficult to provide a realistic perception of depth. Without being limited by theory, it is believed that viewers of an object may perceive the object as three-dimensional due to a combination of vergence and accommodation. The vergence of the two eyes relative to one another (i.e., the rolling of the pupils toward or away from one another to converge and fixate on an object) is closely coupled to the focusing of the eye's lenses (or "accommodation"). Under normal conditions, changing the focus of the eye's lenses, or accommodating the eyes and shifting focus from one object to another at a different distance, will automatically produce a corresponding change in vergence to the same distance, a relationship known as the "accommodation-vergence reflex." Similarly, a change in vergence will trigger a corresponding change in accommodation under normal conditions. Display systems that provide better matching between accommodation and convergence-divergence movements may produce more realistic and comfortable simulations of three-dimensional images.
[0027] FIG. 3 illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes. With reference to FIG. 3 , objects at various distances from the eyes 302 and 304 on the z-axis are accommodated by the eyes 302 and 304 such that the objects are in focus. The eyes 302 and 304 assume particular accommodated states, focusing objects at different distances along the z-axis. As a result, a particular accommodated state may be said to be associated with a particular one of the depth planes 306, having an associated focal length, such that an object or portion of an object at a particular depth plane is in focus when the eye is in an accommodated state relative to that depth plane. In some embodiments, a three-dimensional image may be simulated by providing a different representation of an image to each eye 302 and 304, and by providing a different representation of an image corresponding to each of the depth planes. While shown as separate for clarity of illustration, it should be understood that the fields of view of the eyes 302 and 304 may overlap, for example, as the distance along the z-axis increases. Additionally, while shown as flat for ease of illustration, it should be understood that the contours of a depth plane may be curved in physical space so that all features within the depth plane are in focus with the eye in a particular state of accommodation. Without being limited by theory, it is believed that the human eye can typically interpret a finite number of depth planes to provide depth perception. As a result, a highly realistic simulation of perceived depth can be achieved by providing the eye with a different representation of an image corresponding to each of these limited number of depth planes. (Waveguide stack assembly)
[0028] FIG. 4 illustrates an example of a waveguide stack for outputting image information to a user. Wearable system 400 includes a stack of waveguides or stacked waveguide assembly 480 that can be utilized to provide three-dimensional perception to the eye / brain using multiple waveguides 432b, 434b, 436b, 438b, 4400b. In some embodiments, wearable system 400 may correspond to wearable system 200 of FIG. 2, and FIG. 4 diagrammatically illustrates several portions of wearable system 200 in more detail. For example, in some embodiments, waveguide assembly 480 may be integrated into display 220 of FIG. 2.
[0029] 4, the waveguide assembly 480 may also include multiple features 458, 456, 454, 452 between the waveguides. In some embodiments, the features 458, 456, 454, 452 may be lenses. In other embodiments, the features 458, 456, 454, 452 may not be lenses. Rather, they may simply be spacers (e.g., cladding layers or structures to form air gaps).
[0030] Waveguides 432b, 434b, 436b, 438b, 440b or multiple lenses 458, 456, 454, 452 may be configured to transmit image information to the eye with various levels of wavefront curvature or ray divergence. Each waveguide level may be associated with a particular depth plane and configured to output image information corresponding to that depth plane. Image injection devices 420, 422, 424, 426, 428 may be utilized to inject image information into waveguides 440b, 438b, 436b, 434b, 432b, respectively, which may be configured to disperse incident light across each individual waveguide for output toward the eye 410. Light exits the output surfaces of image injection devices 420, 422, 424, 426, 428 and is injected into the corresponding input edges of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, a single beam of light (e.g., a collimated beam) may be injected into each waveguide, outputting an entire field of cloned collimated beams directed towards eye 410 at a particular angle (and divergence) corresponding to the depth plane associated with the particular waveguide.
[0031] In some embodiments, image input devices 420, 422, 424, 426, 428 are each discrete displays that generate image information for input into corresponding waveguides 440b, 438b, 436b, 434b, 432b, respectively. In some other embodiments, image input devices 420, 422, 424, 426, 428 are outputs of a single multiplexed display that may, for example, send image information to each of image input devices 420, 422, 424, 426, 428 via one or more optical conduits (such as fiber optic cables).
[0032] A controller 460 controls the operation of stacked waveguide assembly 480 and image injection devices 420, 422, 424, 426, 428. Controller 460 includes programming (e.g., instructions in a non-transitory computer-readable medium) that coordinates the timing and provision of image information to waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, controller 460 may be a single, integrated device or a distributed system connected by a wired or wireless communication channel. Controller 460 may, in some embodiments, be part of processing module 260 or 270 (shown in FIG. 2).
[0033] Waveguides 440b, 438b, 436b, 434b, 432b may be configured to propagate light within each respective waveguide by total internal reflection (TIR). Waveguides 440b, 438b, 436b, 434b, 432b may each be planar or have another shape (e.g., curved) with major top and bottom surfaces and edges extending between the major top and bottom surfaces. In the illustrated configuration, waveguides 440b, 438b, 436b, 434b, 432b may each include light extraction optical elements 440a, 438a, 436a, 434a, 432a configured to extract light from the waveguides by redirecting the light to propagate within each respective waveguide and outputting image information from the waveguides to the eye 410. The extracted light may also be referred to as out-coupled light, and the light extraction optical element may also be referred to as out-coupling optical element. The extracted light beam is output by the waveguide where the light propagating within the waveguide strikes the light redirecting element. The light extraction optical element (440a, 438a, 436a, 434a, 432a) may be, for example, a reflective or diffractive optical feature. While shown disposed on the bottom major surfaces of the waveguides 440b, 438b, 436b, 434b, 432b for ease of explanation and clarity of drawing, in some embodiments, the light extraction optical element 440a, 438a, 436a, 434a, 432a may be disposed on the top or bottom major surfaces or directly within the volume of the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed in a layer of material that is attached to a transparent substrate and forms the waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, the waveguides 440b, 438b, 436b, 434b, 432b may be a monolithic piece of material, and the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed on and / or within that piece of material.
[0034] Continuing with reference to FIG. 4, as discussed herein, each waveguide 440b, 438b, 436b, 434b, 432b is configured to output light and form an image corresponding to a particular depth plane. For example, the waveguide 432b closest to the eye may be configured to deliver collimated light to the eye 410 as it is launched into such waveguide 432b. The collimated light may represent an optical infinity focal plane. The next upper waveguide 434b may be configured to send collimated light that passes through a first lens 452 (e.g., a negative lens) before reaching the eye 410. The first lens 452 may be configured to create a slight convex wavefront curvature so that the eye / brain interprets light emerging from the next upper waveguide 434b as emerging from a first focal plane closer inward from optical infinity toward the eye 410. Similarly, the third upper waveguide 436b passes its output light through both the first lens 452 and the second lens 454 before reaching the eye 410. The combined refractive power of the first and second lenses 452 and 454 may be configured to produce another, increasing amount of wavefront curvature such that the eye / brain interprets the light emerging from the third upper waveguide 436b as originating from a second focal plane that is closer inward from optical infinity towards the person than was the light from the next upper waveguide 434b.
[0035] Other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are similarly configured, with the highest waveguide 440b in the stack sending its output through all of the lenses between it and the eye for a collective focal power representing the focal plane closest to the person. To compensate for the stack of lenses 458, 456, 454, 452 when viewing / interpreting light originating from the world 470 on the other side of the stacked waveguide assembly 480, a compensating lens layer 430 may be placed on top of the stack to compensate for the collective power of the lower lens stacks 458, 456, 454, 452. Such a configuration provides as many perceived focal planes as there are available waveguide / lens pairs. Both the light extraction optical elements of the waveguides and the focusing sides of the lenses may be static (e.g., not dynamic or electro-active). In some alternative embodiments, either or both may be dynamic using electro-active features.
[0036] Continuing with reference to FIG. 4 , light extraction optical elements 440a, 438a, 436a, 434a, 432a may be configured to both redirect light from its respective waveguide and output the light with an appropriate amount of divergence or collimation for a particular depth plane associated with that waveguide. As a result, waveguides with different associated depth planes may have different configurations of light extraction optical elements, which output light with different amounts of divergence depending on the associated depth plane. In some embodiments, as discussed herein, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume or surface features, which may be configured to output light at specific angles. For example, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume holograms, surface holograms, and / or diffraction gratings. Light extraction optical elements such as diffraction gratings are described in U.S. Patent Publication No. 2015 / 0178939, published June 25, 2015, which is incorporated herein by reference in its entirety.
[0037] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a are diffractive features, i.e., "diffractive optical elements" (also referred to herein as "DOEs"), that form a diffraction pattern. Preferably, the DOEs have a relatively low diffraction efficiency so that only a portion of the light in the beam is deflected toward the eye 410 at each intersection of the DOE, while the remainder continues traveling through the waveguide via total internal reflection. The light carrying the image information is thus split into several related output beams that exit the waveguide at multiple locations, which can result in a very uniform pattern of output emission toward the eye 304 for this particular collimated beam bouncing within the waveguide.
[0038] In some embodiments, one or more DOEs may be switchable between an "on" state in which they actively diffract and an "off" state in which they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer-dispersed liquid crystal in which microdroplets comprise a diffractive pattern in a host medium, and the refractive index of the microdroplets may be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract incident light), or the microdroplets may be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts incident light).
[0039] In some embodiments, the number and distribution of depth planes or depths of field may be dynamically varied based on the size or orientation of the viewer's pupil. The depth of field may vary inversely with the viewer's pupil size. As a result, as the size of the viewer's pupil decreases, the depth of field increases so that a plane that is indistinguishable because its location exceeds the eye's depth of focus may become distinguishable and appear more focused with a corresponding decrease in pupil size and an increase in depth of field. Similarly, the number of spaced depth planes used to present different images to the viewer may be reduced with a decreased pupil size. For example, a viewer may not be able to clearly perceive details in both a first depth plane and a second depth plane at one pupil size without adjusting their eye's accommodation from one depth plane to the other. However, these two depth planes may be sufficient to simultaneously focus on the user at another pupil size without changing accommodation.
[0040] In some embodiments, the display system may vary the number of waveguides receiving image information based on a determination of pupil size and / or orientation or in response to receiving an electrical signal indicating a particular pupil size and / or orientation. For example, if a user's eye is unable to distinguish between two depth planes associated with two waveguides, controller 460 may be configured or programmed to stop providing image information to one of those waveguides. Advantageously, this may reduce the processing burden on the system, thereby increasing system responsiveness. In embodiments in which the DOE for a waveguide is switchable between on and off states, the DOE may be switched to the off state when the waveguide receives image information.
[0041] In some embodiments, it may be desirable to have the output beam satisfy the condition of having a diameter less than the diameter of the viewer's eye. However, meeting this condition may be difficult in light of the variability in the size of the viewer's pupil. In some embodiments, this condition is met over a wide range of pupil sizes by varying the size of the output beam in response to a determination of the size of the viewer's pupil. For example, as the pupil size decreases, the size of the output beam may also decrease. In some embodiments, the output beam size may be varied using a variable aperture.
[0042] The wearable system 400 may include an outward-facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 may be referred to as the world camera's field of view (FOV), and the imaging system 464 is sometimes referred to as an FOV camera. The entire area available for viewing or imaging by a viewer may be referred to as the ocular field of view (FOR). The FOR may include a solid angle of 4π steradians surrounding the wearable system 400 as the wearer moves their body, head, or eyes to perceive virtually any direction in space. In other situations, the wearer's movement may be more constrained, and accordingly, the wearer's FOR may subtend a smaller solid angle. Images obtained from the outward-facing imaging system 464 can be used to track gestures (e.g., hand or finger gestures) made by the user, detect objects in the world 470 in front of the user, etc.
[0043] The wearable system 400 may also include an inward-facing imaging system 466 (e.g., a digital camera) that observes user movements, such as eye and facial movements. The inward-facing imaging system 466 may be used to capture images of the eyes 410 and determine the size and / or orientation of the pupils of the eyes 304. The inward-facing imaging system 466 may be used to obtain images for use in determining the direction the user is looking (e.g., eye pose) or for biometric identification of the user (e.g., via iris identification). In some embodiments, at least one camera may be utilized for each eye independently to separately determine the pupil size or eye pose of each eye, thereby allowing the presentation of image information to each eye to be dynamically adjusted for that eye. In some other embodiments, the pupil diameter or orientation of only a single eye 410 (e.g., using only a single camera per pair of eyes) is determined and assumed to be similar for both eyes of the user. Images obtained by inward-facing imaging system 466 may be analyzed to determine the user's eye posture or mood, which may be used by wearable system 400 to determine audio or visual content to be presented to the user. Wearable system 400 may also determine head pose (e.g., head position or head orientation) using sensors such as an IMU, accelerometer, gyroscope, etc.
[0044] The wearable system 400 may include a user input device 466 through which a user may input commands into the controller 460 and interact with the wearable system 400. For example, the user input device 466 may include a trackpad, touchscreen, joystick, multi-degree-of-freedom (DOF) controller, capacitive sensing device, game controller, keyboard, mouse, directional pad (D-pad), wand, tactile device, totem (e.g., functioning as a virtual user input device), etc. A multi-DOF controller may sense user input in possible translation (e.g., left / right, forward / backward, or up / down) or rotation (e.g., yaw, pitch, or roll) of some or all of the controller. A multi-DOF controller that supports translation may be referred to as 3DOF, while a multi-DOF controller that supports translation and rotation may be referred to as 6DOF. In some cases, a user may use a finger (e.g., a thumb) to press or swipe across a touch-sensitive input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input device 466 may be held by the user's hand during use of the wearable system 400. The user input device 466 may communicate with the wearable system 400 via wired or wireless communication.
[0045] The wearable system 400 may also include a physiological sensor 468 (which may be an exemplary embodiment of the sensor 232 in FIG. 2 ) configured to measure physiological parameters of the user, such as heart rate, galvanic skin response, respiration rate, etc. The physiological sensor may communicate acquired data to the controller 460. The controller 460 can use the data acquired by the physiological sensor, alone or in combination with data acquired by other sensors, to determine the user's physiological and / or psychological state. For example, the controller 460 can combine heart rate data acquired by the physiological sensor 468 with pupil dilation information acquired by the inward-facing imaging system 462 to determine whether the user is happy or angry. As described further below, the wearable system can selectively present virtual content to the user based on the user's physiological and / or psychological state.
[0046] FIG. 5 shows an example of an output beam output by a waveguide. While one waveguide is illustrated, it should be understood that other waveguides in waveguide assembly 480 may function similarly, and that waveguide assembly 480 may include multiple waveguides. Light 520 is launched into waveguide 432b at input edge 432c of waveguide 432b and propagates within waveguide 432b by TIR. At the point where light 520 impinges on DOE 432a, a portion of the light exits the waveguide as output beam 510. While output beams 510 are illustrated as being approximately parallel, they may also be redirected to propagate to eye 410 at an angle (e.g., divergent output beam formation) depending on the depth plane associated with waveguide 432b. It should be understood that a nearly collimated exit beam may refer to a waveguide with light-extracting optics that outcouples light to form an image that appears to be set at a depth plane at a large distance (e.g., optical infinity) from the eye 410. Other waveguides or other sets of light-extracting optics may output a more divergent exit beam pattern, which would require the eye 410 to accommodate to a closer distance and focus on the retina, and would be interpreted by the brain as light from a distance closer to the eye 410 than optical infinity.
[0047] FIG. 6 is a schematic diagram illustrating an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal volumetric display, image, or light field. The optical system can include a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem. The optical system can be used to generate a multifocal volumetric display, image, or light field. The optical system can include one or more primary planar waveguides 632a (only one is shown in FIG. 6) and one or more DOEs 632b associated with each of at least some of the primary waveguides 632a. The planar waveguides 632b can be similar to the waveguides 432b, 434b, 436b, 438b, and 440b discussed with reference to FIG. 4. The optical system may employ a dispersive waveguide device to relay light along a first axis (the vertical or Y-axis in the illustration of FIG. 6 ) and expand the effective exit pupil of the light along the first axis (e.g., the Y-axis). The dispersive waveguide device may include, for example, a dispersive planar waveguide 622 b and at least one DOE 622 a (illustrated by a double-dashed line) associated with the dispersive planar waveguide 622 b. The dispersive planar waveguide 622 b may be similar or identical in at least some respects to a primary planar waveguide 632 b having a different orientation therefrom. Similarly, the at least one DOE 622 a may be similar or identical in at least some respects to the DOE 632 a. For example, the dispersive planar waveguide 622 b or the DOE 622 a may be made of the same material as the primary planar waveguide 632 b or the DOE 632 a, respectively. The embodiment of the optical display system 600 shown in FIG. 6 can be integrated into the wearable system 200 shown in FIG.
[0048] The relayed, exit-pupil-expanded light can be optically coupled from the dispersive waveguide device into one or more primary planar waveguides 632b. The primary planar waveguides 632b can relay the light along a second axis, preferably orthogonal to the first axis (e.g., the horizontal or X-axis in the diagram of FIG. 6). Notably, the second axis can be non-orthogonal to the first axis. The primary planar waveguides 632b expand the effective exit pupil of the light along that second axis (e.g., the X-axis). For example, the dispersive planar waveguide 622b can relay and expand the light along the vertical or Y-axis and pass the light to a primary planar waveguide 632b, which can relay and expand the light along the horizontal or X-axis.
[0049] The optical system may include one or more colored light sources (e.g., red, green, and blue laser light) 610, which may be optically coupled into the proximal end of a single-mode optical fiber 640. The distal end of the optical fiber 640 may be threaded or received through a hollow tube 642 of piezoelectric material. The distal end protrudes from the tube 642 as a free-standing, flexible cantilever 644. The piezoelectric tube 642 may be associated with four quadrant electrodes (not shown). The electrodes may be plated, for example, on the outside, outer surface or outer periphery, or diameter of the tube 642. A core electrode (not shown) may also be located in the core, center, inner periphery, or inner diameter of the tube 642.
[0050] For example, drive electronics 650, electrically coupled via wires 660, drive opposing pairs of electrodes to bend piezoelectric tube 642 independently in two axes. The protruding distal tip of optical fiber 644 has a mechanical resonant mode. The frequency of the resonance may depend on the diameter, length, and material properties of optical fiber 644. By oscillating piezoelectric tube 642 near the first mechanical resonant mode of fiber cantilever 644, fiber cantilever 644 may be caused to oscillate and sweep through a large deflection.
[0051] By stimulating resonant vibrations in two axes, the tip of fiber cantilever 644 is scanned biaxially within an area filling a two-dimensional (2-D) scan. By modulating the intensity of light source 610 synchronously with the scanning of fiber cantilever 644, light emitted from fiber cantilever 644 can form an image. A description of such a setup is provided in U.S. Patent Publication No. 2014 / 0003762, which is incorporated herein by reference in its entirety.
[0052] Components of the optical coupler subsystem can collimate light emitted from the scanning fiber cantilever 644. The collimated light can be reflected by a mirrored surface 648 into a narrow dispersive planar waveguide 622b containing at least one diffractive optical element (DOE) 622a. The collimated light can propagate perpendicularly (with respect to the view of FIG. 6) along the dispersive planar waveguide 622b via TIR, and in doing so, repeatedly intersect with the DOE 622a. The DOE 622a preferably has a low diffraction efficiency. This causes a portion of the light (e.g., 10%) to diffract toward the edge of the larger primary planar waveguide 632b at each point of intersection with the DOE 622a, allowing a portion of the light to continue on its original trajectory down the length of the dispersive planar waveguide 622b via TIR.
[0053] At each point of intersection with DOE 622a, additional light can be diffracted toward the entrance of primary waveguide 632b. By splitting the incident light into multiple outcoupled sets, the exit pupil of the light can be vertically expanded by DOE 4 within dispersive planar waveguide 622b. This vertically expanded light outcoupled from dispersive planar waveguide 622b can enter the edge of primary planar waveguide 632b.
[0054] Light entering the primary waveguide 632b can propagate horizontally (with respect to the illustration of FIG. 6) along the primary waveguide 632b via TIR. The light propagates horizontally along at least a portion of the length of the primary waveguide 632b via TIR as it intersects the DOE 632a at multiple points. The DOE 632a advantageously has a phase profile that is the sum of a linear diffraction pattern and a radially symmetric diffraction pattern, and may be designed or configured to produce both deflection and focusing of the light. The DOE 632a advantageously may have a low diffraction efficiency (e.g., 10%) so that only a portion of the light in the beam is deflected toward the viewer's eye at each intersection of the DOE 632a, while the remainder of the light continues to propagate through the primary waveguide 632b via TIR.
[0055] At each point of intersection between the propagating light and the DOE 632a, a portion of the light is diffracted toward the adjacent face of the primary waveguide 632b, allowing the light to escape the TIR and emerge from the face of the primary waveguide 632b. In some embodiments, the radially symmetric diffraction pattern of the DOE 632a additionally imparts a focal level to the diffracted light, both shaping (e.g., imparting curvature) the optical wavefronts of the individual beams and steering the beams to angles that match the designed focal level.
[0056] Thus, these different paths can couple light out of the primary planar waveguide 632b by providing different fill patterns at the DOE 632a's multiplicity, focal level, and / or exit pupil at different angles. Different fill patterns at the exit pupil can be advantageously used to generate light field displays with multiple depth planes. Each layer in the waveguide assembly or set of layers (e.g., three layers) in the stack may be employed to generate distinct colors (e.g., red, blue, and green). Thus, for example, a first set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a first focal depth. A second set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a second focal depth. Multiple sets may be employed to generate full 3D or 4D color image light fields with various focal depths. (Other components of the wearable system)
[0057] In many implementations, the wearable system may include other components in addition to or as an alternative to the components of the wearable system described above. The wearable system may include, for example, one or more tactile devices or components. The tactile device or component may be operable to provide a tactile sensation to the user. For example, the tactile device or component may provide a tactile sensation of pressure and / or texture upon touching virtual content (e.g., a virtual object, virtual tool, other virtual structure). The tactile sensation may replicate the sensation of a physical object represented by the virtual object, or may replicate the sensation of an imaginary object or character (e.g., a dragon) represented by the virtual content. In some implementations, the tactile device or component may be worn by the user (e.g., a user-wearable glove). In some implementations, the tactile device or component may be held by the user.
[0058] A wearable system may include, for example, one or more physical objects that can be manipulated by a user to enable input to or interaction with the wearable system. These physical objects may be referred to herein as totems. Some totems may take the form of inanimate objects, such as, for example, a piece of metal or plastic, a wall, the surface of a table, etc. In some implementations, a totem may not actually have any physical input structures (e.g., keys, triggers, joysticks, trackballs, rocker switches). Instead, the totem may simply provide a physical surface, and the wearable system may render a user interface to appear to the user on one or more surfaces of the totem. For example, the wearable system may render an image of a computer keyboard and trackpad to appear to reside on one or more surfaces of the totem. For example, the wearable system may render a virtual computer keyboard and virtual trackpad to appear on the surface of a thin rectangular plate of aluminum that serves as the totem. The rectangular plate itself does not have any physical keys or trackpads or sensors. However, the wearable system may detect user manipulation or interaction or touch with the rectangular plate as a selection or input made via a virtual keyboard or virtual trackpad. User input device 466 (shown in FIG. 4) may be an embodiment of a totem, which may include a trackpad, touchpad, trigger, joystick, trackball, rocker or virtual switch, mouse, keyboard, multi-degree-of-freedom controller, or another physical input device. A user may use the totem alone or in combination with posture to interact with the wearable system or other users.
[0059] Examples of tactile devices and totems usable with the wearable devices, HMDs, and display systems of the present disclosure are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. Exemplary Wearable Systems, Environments, and Interfaces
[0060] The wearable system may employ various mapping-related techniques to achieve a high depth of field within the rendered light field. When mapping a virtual world, it is advantageous to capture all features and points in the real world and accurately depict virtual objects in relation to the real world. To achieve this goal, FOV images captured from a user of the wearable system can be added to the world model by including new photos that convey information about various points and features in the real world. For example, the wearable system can collect a set of map points (such as 2D or 3D points), find new map points, and render a more accurate version of the world model. The world model of a first user can be communicated to a second user (e.g., via a network such as a cloud network) so that the second user can experience the world surrounding the first user.
[0061] 7 is a block diagram of an example MR environment 700. The MR environment 700 may be configured to receive inputs (e.g., visual input 702 from a user's wearable system, stationary input 704 such as a room camera, sensory input 706 from various sensors, gestures, totems, eye tracking, user input, etc. from user input device 466) from one or more user-wearable systems (e.g., wearable system 200 or display system 220) or stationary room systems (e.g., room cameras, etc.). The wearable systems can determine the location and various other attributes of the user's environment using various sensors (e.g., accelerometers, gyroscopes, temperature sensors, movement sensors, depth sensors, GPS sensors, inward-facing imaging systems, outward-facing imaging systems, etc.). This information may be further supplemented with information from stationary cameras in the room, which may provide images from different perspectives or various cues. Image data acquired by the cameras (e.g., the room cameras and / or the outward-facing imaging system cameras) may be reduced to a set of mapping points.
[0062] One or more object recognizers 708 can crawl through the received data (e.g., a collection of points), recognize or map the points, tag images, and associate semantic information with objects with the assistance of a map database 710. The map database 710 may comprise various points and their corresponding objects collected over time. The various devices and the map database may be interconnected through a network (e.g., a LAN, a WAN, etc.) and accessible to the cloud.
[0063] Based on this information and the collection of points in the map database, the object recognizers 708a-708n may recognize objects in the environment. For example, the object recognizers may recognize faces, people, windows, walls, user input devices, televisions, other objects in the user's environment, etc. One or more object recognizers may be specialized for objects with certain characteristics. For example, object recognizer 708a may be used to recognize faces, while another object recognizer may be used to recognize totems.
[0064] Object recognition may be performed using various computer vision techniques. For example, the wearable system may analyze images acquired by the outward-facing imaging system 464 (shown in FIG. 4) and perform scene reconstruction, event detection, video tracking, object recognition, object pose estimation, learning, indexing, motion estimation, or image restoration, etc. One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale Invariant Feature Transform (SIFT), Speed-Up Robust Features (SURF), Orientation FAST and Rotation BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retinal Keypoints (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous localization and mapping (vSLAM) techniques, sequential Bayes estimators (e.g., Kalman filter, extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), iterative nearest neighbor (ICP), semi-global matching (SGM), semi-global block matching (SGBM), feature point histograms, various machine learning algorithms (e.g., support vector machines, k-nearest neighbor algorithms, naive Bayes, neural networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), etc.
[0065] Object recognition can additionally or alternatively be performed by various machine learning algorithms. Once trained, the machine learning algorithms can be stored by the HMD. Some examples of machine learning algorithms can include supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., a priori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural networks, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., stacked generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models can be customized for individual datasets. For example, the wearable device can generate or store a base model. The base model may be used as a starting point to generate additional models specific to a data type (e.g., a particular user in a telepresence session), a data set (e.g., a set of additional images acquired of a user in a telepresence session), a conditional situation, or other variations. In some embodiments, the wearable HMD can be configured to generate models for analysis of aggregated data using multiple techniques. Other techniques may include using predefined thresholds or data values.
[0066] Based on this information and the set of points in the map database, the object recognizers 708a-708n may recognize objects, complement them with semantic information, and bring them to life. For example, if the object recognizer recognizes that a set of points is a door, the system may associate some semantic information (e.g., a door has a hinge and 90-degree movement around the hinge). If the object recognizer recognizes that a set of points is a mirror, the system may associate the semantic information that a mirror has a reflective surface that can reflect images of objects in the room. Over time, the map database grows as the system (which may reside locally or be accessible through a wireless network) accumulates more data from the world. Once an object is recognized, the information may be transmitted to one or more wearable systems. For example, the MR environment 700 may contain information about a scene taking place in California. The environment 700 may be transmitted to one or more users in New York. Based on the data received from the FOV camera and other inputs, the object recognizer and other software components can map points collected from various images, recognize objects, etc. so that the scene can be accurately "passed" to a second user who may be in a different part of the world. The environment 700 may also use a topology map for localization purposes.
[0067] 8 is a process flow diagram of an example method 800 for rendering virtual content in relation to recognized objects. Method 800 describes how a virtual scene can be presented to a user of a wearable system. The user may be geographically remote from the scene. For example, a user may be in New York but may want to view a scene currently occurring in California, or may want to go for a walk with a friend who is in California.
[0068] In block 810, the wearable system may receive input from the user and other users regarding the user's environment. This may be accomplished through various input devices and knowledge already held in a map database. The user's FOV camera, sensors, GPS, eye tracking, etc., communicate information to the system in block 810. The system may determine coarse points based on this information in block 820. The coarse points may be used to determine pose data (e.g., head pose, eye pose, body pose, or hand gestures) that may be used in displaying and understanding the orientation and position of various objects in the user's surroundings. The object recognizers 708a-708n may crawl through these collected points and recognize one or more objects using the map database in block 830. This information may then be communicated to the user's respective wearable system in block 840, and the desired virtual scene may be displayed to the user appropriately in block 850. For example, a desired virtual scene (eg, a user in CA) may be displayed in the proper orientation, position, etc., relative to various objects and other surroundings of the user in New York.
[0069] FIG. 9 is a block diagram of another example of a wearable system. In this example, the wearable system 900 includes a map, which may include map data about the world. The map may reside partially locally on the wearable system and partially in a networked storage location (e.g., in a cloud system) accessible by a wired or wireless network. An attitude process 910 may run on the wearable computing architecture (e.g., processing module 260 or controller 460) and utilize data from the map to determine the position and orientation of the wearable computing hardware or the user. The attitude data may be calculated from data collected on the fly as the user experiences the system and moves within its world. The data may include images, data from sensors (such as inertial measurement units, which generally include accelerometer and gyroscope components), and surface information about objects in the real or virtual environment.
[0070] The coarse point representation may be the output of a simultaneous localization and mapping (SLAM or V-SLAM, which refers to configurations where the input is image / vision only) process. The system can be configured to find what the world is made of, not just the locations of various components within the world. Poses may be building blocks that accomplish many goals, including populating a map and using data from the map.
[0071] In one embodiment, the rough point locations may not be entirely adequate by themselves, and additional information may be required to generate a multifocal AR, VR, or MR experience. A dense representation, generally referring to depth map information, may be utilized to fill in this gap, at least in part. Such information may be calculated from a process referred to as stereoscopic vision 940, with depth information determined using techniques such as triangulation or time-of-flight sensing. Image information and active patterns (such as infrared patterns generated using an active projector) may serve as inputs to the stereoscopic vision process 940. A significant amount of depth map information may be fused together, and some of this may be summarized using a surface representation. For example, a mathematically definable surface may be an efficient (e.g., compared to a large-scale point cloud) and digestible input to other processing devices, such as a game engine. Thus, the outputs of the stereoscopic vision process (e.g., depth map) 940 may be combined in a fusion process 930. The pose may also be input to this fusion process 930, the output of which is the input to populate the map process 920. Sub-surfaces may connect to each other to form larger surfaces, such as in topographic mapping, and the map becomes a large-scale hybrid of points and surfaces.
[0072] Various inputs may be utilized to resolve various aspects of the mixed reality process 960. For example, in the embodiment depicted in Figure 9, game parameters may be inputs for determining whether a user of the system is playing a monster battle game with one or more monsters in various locations, whether monsters are dead or fleeing under various conditions (such as when a user shoots a monster), walls or other objects in various locations, and the like. A world map may contain information about where such objects are located relative to one another, which is another useful input for mixed reality. Attitude relative to the world is likewise an input and plays an important role for nearly any interactive system.
[0073] Controls or inputs from the user are another input to the wearable system 900. As described herein, user inputs can include visual inputs, gestures, totems, audio inputs, sensory inputs (e.g., physiological data obtained by sensors 232 in FIG. 2 ), etc. To move around or play a game, for example, the user may need to command the wearable system 900 as to what they want to do. There are various forms of user control that can be utilized beyond just moving around in space. In one embodiment, a totem (e.g., a user input device) or an object such as a toy gun may be held by the user and tracked by the system. The system would preferably be configured to know that the user is holding an item and understand the type of interaction the user is having with the item (e.g., if the totem or object is a gun, the system may be configured to understand not only the location and orientation, but also whether the user is clicking a trigger or other sensitive button or element, which may be equipped with sensors such as an IMU, which can help determine what is happening even when such activity is not within the field of view of any of the cameras).
[0074] Hand gesture tracking or recognition may also provide input information. The wearable system 900 may be configured to track and interpret hand gestures to gesture for button presses, left or right, stop, grasp, hold, etc. For example, in one configuration, a user may wish to flip through email or calendar in a non-gaming environment or perform a “fist bump” with another person or performer. The wearable system 900 may be configured to utilize a minimal amount of hand gestures, which may or may not be dynamic. For example, gestures may be simple static gestures, such as extending the hand to indicate stop, thumbs up to indicate OK, thumbs down to indicate not OK, or flipping the hand left and right or up and down to indicate a directional command.
[0075] Eye tracking is another input (e.g., to track where the user is looking and control display technology to render at a specific depth or range). In one embodiment, eye convergence may be determined using triangulation, and then accommodation may be determined using a convergence / accommodation model developed for that particular person.
[0076] The totem can also be used by the user to provide input to the wearable system, which can track the totem's movement, position, or orientation, as well as the user's actuation of the totem (such as pressing keys, buttons, or touch surfaces on the totem) to determine user interface interactions in the mixed reality process 960.
[0077] In some implementations, the wearable system can also use the user's physiological data in the mixed reality process 960. The physiological data may be requested by sensors 232 (which may include physiological sensors 468). The wearable system can determine content to present based on an analysis of the physiological data. For example, if the wearable system determines that the user is getting angry while playing a game (e.g., due to an increased heart rate, changes in blood pressure, etc.), the wearable system can automatically reduce the difficulty level of the game to keep the user engaged in the game.
[0078] With respect to the camera system, the exemplary wearable system 900 shown in FIG. 9 may include three pairs of cameras: a pair of relatively wide-FOV or passive SLAM cameras arranged on either side of the user's face, and a different pair of cameras oriented in front of the user to handle the stereoscopic imaging process 940 and capture hand gestures and totem / object trajectories in front of the user's face. The FOV cameras and pair of cameras for the stereo process 940 may be part of the outward-facing imaging system 464 (shown in FIG. 4). The wearable system 900 may include an eye-tracking camera (which may be part of the inward-facing imaging system 462 shown in FIG. 4) oriented toward the user's eyes to triangulate eye vectors and other information. The wearable system 900 may also include one or more textured light projectors (such as infrared (IR) projectors) to inject texture into the scene.
[0079] 10 is a process flow diagram of an example embodiment of a method 1000 for determining user input to a wearable system. In this example, a user may interact with a totem. A user may have multiple totems. For example, a user may have one totem designated for social media applications, another totem for playing games, etc. In block 1010, the wearable system may detect movement of the totem. The movement of the totem may be recognized through an outward-facing system or may be detected through sensors (e.g., tactile gloves, image sensors, hand tracking devices, eye tracking cameras, head pose sensors, etc.).
[0080] Based at least in part on the detected gestures, eye poses, head poses, or inputs through the totem, the wearable system detects the position, orientation, and / or movement of the totem (or the user's eyes or head or gestures) relative to a frame of reference in block 1020. The frame of reference may be a set of map points based on which the wearable system translates the totem's (or the user's) movements into actions or commands. In block 1030, the user's interactions with the totem are mapped. Based on the mapping of the user interactions to the frame of reference 1020, the system determines the user input in block 1040.
[0081] For example, a user may move a totem or physical object back and forth, turn a virtual page, move to the next page, or move from one user interface (UI) display screen to another. As another example, a user may move their head or eyes to view different real or virtual objects within the user's FOR. If the user's gaze at a particular real or virtual object is longer than a threshold time, that real or virtual object may be selected as user input. In some implementations, the user's eye vergence-divergence can be tracked, and an accommodation / vergence-divergence model can be used to determine the user's eye accommodation state, which provides information about the depth plane the user is focusing on. In some implementations, the wearable system can use ray-casting techniques to determine real or virtual objects that are aligned with the user's head or eye pose. In various implementations, ray casting techniques can include casting a thin bundle of rays with substantially little lateral width, or casting rays with substantial lateral width (e.g., a virtual cone or truncated cone).
[0082] The user interface may be projected by a display system (such as display 220 in FIG. 2) as described herein. It may also be displayed using a variety of other techniques, such as one or more projectors. A projector may project an image onto a physical object, such as a canvas or a sphere. Interactions with the user interface may be tracked using one or more cameras outside or part of the system (e.g., using inward-facing imaging system 462 or outward-facing imaging system 464).
[0083] 11 is a process flow diagram of an example method 1100 for interacting with a virtual user interface. Method 1100 may be performed by a wearable system described herein.
[0084] In block 1110, the wearable system may identify a particular UI. The type of UI may be provided by the user. The wearable system may identify that a particular UI needs to be populated based on user input (e.g., gestures, visual data, audio data, sensory data, direct commands, etc.). In block 1120, the wearable system may generate data for a virtual UI. For example, data associated with the UI's boundaries, general structure, shape, etc. may be generated. Additionally, the wearable system may determine map coordinates of the user's physical location so that the wearable system may display the UI in relation to the user's physical location. For example, if the UI is body-centered, the wearable system may determine coordinates of the user's physical stance, head pose, or eye pose so that a ring UI may be displayed around the user or a planar UI may be displayed on a wall or in front of the user. If the UI is hand-centered, map coordinates of the user's hand may be determined. These map points may be derived through an FOV camera, data received through sensory input, or any other type of collected data.
[0085] In block 1130, the wearable system may send data from the cloud to the display, or the data may be sent from a local database to the display component. In block 1140, a UI is displayed to the user based on the sent data. For example, a light field display can project the virtual UI into one or both of the user's eyes. Once the virtual UI is generated, the wearable system may simply wait for commands from the user and generate more virtual content on the virtual UI in block 1150. For example, the UI may be a body-centered ring around the user's body. The wearable system may then wait for a command (such as a gesture, head or eye movement, input from a user input device, etc.) and, if recognized (block 1160), the virtual content associated with the command may be displayed to the user (block 1170). As an example, the wearable system may wait for a user's hand gesture before mixing multiple stem tracks.
[0086] Additional examples of wearable systems, UIs, and user experiences (UX) are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. (Example objects in the environment)
[0087] 12 illustrates an example of user interaction with a virtual user interface in an office environment. In FIG. 12, a user 210 wearing a wearable device 1270 is standing in an office 1200. The wearable device may be part of a wearable system 200, 400 as described herein. The office 1200 may include multiple physical objects, such as a chair 1244, a mirror 1242, a wall 1248, a table 1246, a swivel chair 1240, etc., and a virtual screen 1250 presented to the user by the wearable device 1270. Example objects in the oculomotor field
[0088] A user 210 wearing a wearable device 1270 can have a field of view (FOV) and a field of view (FOR). As discussed with reference to FIG. 4, the FOR includes the portion of the user's surrounding environment that is perceptible by the user via the wearable device 1270. For an HMD, the FOR can include substantially all of the 4π steradian solid angle surrounding the wearer, since the wearer can move their body, head, or eyes to perceive virtually any direction in space. In other situations, the user's movement may be more constricted, and therefore the user's FOR may cover a smaller solid angle.
[0089] The FOR can contain a group of objects that can be perceived by a user through the ARD. The objects may be virtual and / or physical objects. Virtual objects may include operating system objects, such as a trash can for deleted files, a terminal for entering commands, a file manager for accessing files or directories, icons, menus, applications for audio or video streaming, notifications from the operating system, etc. Virtual objects may also include objects within an application, such as avatars, widgets (e.g., a virtual representation of a wall clock), virtual objects, graphics, or images within a game, etc. Some virtual objects can be both operating system objects and objects within an application.
[0090] In some embodiments, virtual objects may be associated with physical objects. For example, as shown in Figure 12, a virtual screen 1250 may be placed on a table 1246. The virtual screen may include a virtual menu 1210 with selectable options such as office productivity tools 1212, applications for conducting telepresence 1214, and email tools 1216.
[0091] A virtual object may be a three-dimensional (3D), two-dimensional (2D), or one-dimensional (1D) object. A virtual object may be a 3D coffee mug (which may represent virtual controls for a physical coffee maker). A virtual object may also be a 2D menu 1210 (shown in FIG. 12). In some implementations, one or more virtual objects may be displayed within (or associated with) another virtual object. For example, referring to FIG. 12, a virtual menu 5110 is shown inside a virtual screen 1250. In another example, a virtual application for telepresence 1214 may include another menu 1220 with contact information.
[0092] In some implementations, some objects in a user's environment may be interactable. For example, referring to FIG. 1 , a user can interact with some of the virtual objects by, for example, pointing out a finger to land an avatar 140 or pulling up a menu that provides information about the statue 130. A user can interact with the interactable objects by performing user interface actions, such as selecting or moving the interactable object, activating a menu associated with the interactable object, or selecting an action to be performed with the interactable object. A user may perform these user interface actions using head pose, eye pose, body pose, voice, and hand gestures on a user input device, alone or in combination. For example, a user may move a virtual object from one location to another using a change in body pose (e.g., a change in hand gesture, such as waving their hand toward the virtual object). In another example, as shown in FIG. 12 , a user can use a hand gesture to activate a user input device and open a virtual screen 1250 when the user is standing near a table 1246. A user may actuate the user input device 466, such as by clicking a mouse, tapping a touchpad, swiping a touchscreen, hovering over or touching a capacitive button, pressing a key on a keyboard or game controller (e.g., a five-way d-pad), pointing a joystick, wand, or totem toward an object, pressing a button on a remote control, or other interaction with the user input device. In one implementation, the wearable device 1270 can automatically present the virtual menu 1210 in response to detecting the table 1246 (e.g., using one or more object recognizers 708). After the menu is opened, the user can browse the menu 1210 by moving their finger along a trajectory on the user input device.When the user decides to close the virtual screen 1250, the user may utter a word (e.g., "exit") and / or actuate a user input device to indicate an intention to close the virtual screen 1250. After receiving the indication, the ARD may stop projecting the screen 1250 onto the table 1246. (Example object in field of view)
[0093] Within a FOR, the portion of the world that a user perceives at a given time is referred to as the FOV (e.g., the FOV may encompass the portion of the FOR that the user is currently looking at). The FOV may depend on the size or optical properties of the display within the ARD. For example, an AR display may include an optical device that only provides AR functionality when the user is looking through a particular portion of the display. The FOV may correspond to the solid angle perceivable by the user when looking through an AR display such as, for example, stacked waveguide assembly 480 (FIG. 4) or planar waveguide 632b (FIG. 6).
[0094] As the user's posture changes, the FOV changes correspondingly, and the objects within the FOV may also change. Referring to FIG. 12 , when user 210 is standing in front of table 1246, user 210 can perceive virtual screen 1250. However, when user 210 walks toward mirror 1242, virtual screen 1250 may move outside of its FOV. Thus, user 210 would not be able to perceive virtual screen 1250 when standing in front of mirror 1242. In some embodiments, virtual screen 1250 may follow user 210 as they move around office 1200. For example, virtual screen 1250 may move from table 1246 to wall 1248 as user 210 moves and stands in front of mirror 1242. The content of the virtual screen, such as options on virtual menu 1210, may change as virtual screen 1250 changes its location. 12, when virtual screen 1250 is on table 1250, user 210 may perceive virtual menu 1210 containing various office productivity items. However, when the user walks toward mirror 1242, the user may be able to interact with a virtual wardrobe application that allows the user to simulate different outfit looks using wearable device 1270. In one implementation, once wearable device 1270 detects mirror 1242, the wearable system can automatically initiate communication (e.g., a telepresence session) with another user (e.g., the user's 210 personal assistant). Example of rendering virtual objects within the FOV based on contextual factors
[0095] As described herein, there are often multiple virtual objects or user interaction options associated with an object (e.g., physical or virtual) or the user's environment. For example, referring to Figure 12, a virtual menu 1210 includes multiple interaction options, such as office productivity tools 1212 (such as a word processor, file folders, a calendar, etc.), a telepresence application 1214 (via the wearable system) that allows the user to communicate with other users as if the other users were present in the user's 210 environment (e.g., the wearable system can project images of the other users to the user of the wearable system), and a mail tool that allows the user 210 to send and receive electronic mail (email) or text messages. In another example, the living room 1300 shown in FIGS. 13 and 14 may include virtual objects such as a digital frame 1312 a, a telepresence tool 1314 a, a race car driving game 1316 a, a television (TV) application 1312 b, a home management tool 1314 b (which can control the temperature for the room 1300, project wallpaper, etc.), and a music application 1316 b.
[0096] However, as described herein, a virtual user interface may not be able to display all available virtual objects or user interaction options to a user and simultaneously provide a satisfying user experience. For example, as shown in Figures 13 and 14, there are six virtual objects (1312a, 1314a, 1316a, 1312b, 1314b, and 1316b) associated with user interface 1310, which may be virtual menus. However, the virtual menu on wall 1350 may only fit three options readably (see, e.g., the virtual menus in Figures 13 and 14). As a result, the wearable system may need to filter the number of available options and display only a subset of the available options.
[0097] Advantageously, in some embodiments, the wearable system can filter or select user interaction options or virtual objects to be presented on the user interface 1310 based on contextual information. The filtered or selected user interface interaction options or virtual objects may be presented in various layouts. For example, the wearable device can present the options and virtual objects in a list format (such as the virtual menus shown in FIGS. 12-14). In some embodiments, rather than displaying virtual objects on the virtual user interface as a vertical list, the virtual menu can employ a circular representation of the virtual objects (e.g., see the virtual menu 1530 shown in FIG. 15). The virtual objects can be rotated around the center of the circular representation to aid in identifying and selecting the desired virtual object. Contextual information can include information associated with the user's environment, the user, objects in the user's environment, etc. Exemplary contextual information can include the affordances of the physical object with which the option is associated, the user's environment (such as whether the environment is a home or office environment), characteristics of the user, the user's current interaction with objects in the environment, the user's physiological state, the user's psychological state, combinations thereof, or the like. A more detailed description of the various types of context information is provided below. (user environment)
[0098] The wearable system may filter or select virtual objects in an environment based on the user's environment and present only a subset of the virtual objects for user interaction. This is because different environments may have different functionality. For example, a user's contact list may include contact information for family, friends, and professional contacts. An office environment, such as the office 1200 shown in FIG. 12, is typically more suited to work-related activities instead of entertainment activities. As a result, when the user 210 uses the telepresence tool 1214, the wearable device 1270 may present a list of work-related contacts in the menu 1220 for the telepresence session, but the user's contact list also includes contacts for family and friends. In contrast, FIG. 13 depicts a living room 1300 where a user typically relaxes and interacts with people outside of work. As a result, when the user selects the telepresence tool 1314a, the wearable device may present contact information for friends and family in the menu 1320.
[0099] As another example, a user's music collection may include a variety of music, such as country music, jazz, pop, and classical music. When the user is in living room 1300, the wearable device may present jazz and pop music to the user (as shown in virtual menu 1430). However, when the user is in bedroom 1500, a different set of music options may be presented. For example, as shown in virtual menu 1530 in FIG. 15 , the wearable device may present country music and classical music because these types of music may have a relaxing effect and may help the user fall asleep.
[0100] In addition to, or as an alternative to, filtering the virtual objects available in an environment, the wearable device may show only menu options relevant to the functionality of the environment. For example, virtual menu 1210 (in FIG. 12) in office 1200 may include options related to a work environment. On the other hand, virtual user interface 1310 (in FIG. 14) in living room 1300 may include entertainment items such as virtual TV 1312b, music 1316b, and home management tools 1314b.
[0101] 12-15 are illustrative and not intended to limit the types of environments in which a wearable device may be used to contextually interact with physical and virtual content within such environments. Other environments may include other parts of a home or office, a vehicle (e.g., a car, subway, boat, train, or airplane), an entertainment venue (e.g., a movie theater, nightclub, gaming establishment), a retail establishment (e.g., a store or mall), or outdoors (e.g., a park or garden), etc. (Object affordances)
[0102] The wearable system can identify objects in the environment that the user may be interested in interacting with or is currently interacting with. The wearable system can identify objects based on the user's posture, such as eye gaze, body posture, or head pose. For example, the wearable device may track the user's eye posture using an inward-facing imaging system 462 (shown in FIG. 4). When the wearable system determines that the user is looking in a certain direction for an extended period of time, the wearable system may use ray casting or cone casting techniques to identify objects that intersect the user's gaze direction. For example, the wearable system may cast a virtual cone / ray and identify objects that intersect with a portion of the virtual cone / ray. The wearable device may also track the user's head pose using an IMU (e.g., as described with reference to FIGS. 2, 4, and 9). When the wearable device detects a change in the user's head pose, the wearable device may identify an object in the vicinity of the user's head as an object that the user is interested in interacting with. As an example, when a user of a wearable device looks at a refrigerator at home for an extended period of time, the wearable device can recognize that the refrigerator may be an object of interest to the user. In another example, the user may stand in front of the refrigerator. The wearable may detect nodding by the user and identify the refrigerator in front of the user as an object of interest to the user. In yet another example, the object recognizer 708 may track the movement of the user's hand (e.g., based on data from the outward-facing imaging system 464). The object recognizer may recognize a hand gesture (e.g., a finger pointing at the refrigerator) that provides an indication of an object for user interaction.
[0103] The wearable system can recognize the affordances of the identified object. The affordances of an object include relationships between the object and its environment that provide opportunities for actions or uses associated with the object. Affordances may be determined based on, for example, the function, orientation, type, location, shape, and / or size of the object. Affordances may also be based on the environment in which the physical object is located. The wearable device can narrow down available virtual objects in the environment and present the virtual objects according to the affordances of the objects. As an example, the affordance of a horizontal table is that an object can be set on the table, and the affordance of a vertical wall is that an object can be hung from or projected onto the wall.
[0104] For example, the wearable device may identify the function of an object and present a menu with only objects related to the object's function. As an example, when a user of the wearable device interacts with a refrigerator at home, the wearable device may identify that one of the refrigerator's functionalities is storing food. The ability to store food is an affordance of the refrigerator. When the user decides to view options associated with the refrigerator, for example, by activating a user input device, the wearable device may present the user with food-specific options, such as a list of foods currently available in the refrigerator, a cooking application with various recipes, a grocery list of food items, a reminder to change the water filter in the refrigerator, etc. Additional examples of affordances of a refrigerator include that it is heavy and therefore difficult to move, that it has a vertical front surface to which objects can be attached, that the front surface is often metallic and magnetic so that magnetic objects can adhere to the front surface, etc.
[0105] In some situations, the functionality of the same object may vary based on the environment. The wearable system can generate a virtual menu by considering the functionality of the object in light of the environment. For example, the affordances of a table may include that it can be used for writing and eating. When the table is in an office 1200 (shown in FIG. 12 ), the affordances of the table may suggest that the table should be used for writing because an office environment is typically associated with word processing. Thus, the display 220 of the wearable device may present a word processing application under office tools 1212 or an email application 1216 in the virtual menu 1210. However, when the same table is located in a kitchen, the affordances of the table may suggest that the table can be used for eating because people typically do not write documents in the kitchen. As a result, the wearable device may display food-related virtual objects to the user instead of office tools.
[0106] The wearable system can use the object's orientation to determine which options to present, as some activities (such as drawing and writing) may be more appropriate on a horizontal surface (such as a floor or table), while other activities (such as watching TV or playing a driving game) may have a better user experience on a vertical surface (such as a wall). The wearable system can detect the orientation of the object's surface (e.g., horizontal vs. vertical) and display a group of options appropriate for that orientation.
[0107] 12 , office 1200 may include virtual objects such as office tools 1212 for word processing and a virtual TV application (not shown in FIG. 12 ). Because virtual screen 1250 is on table 1246, which has a horizontal surface, wearable device 1270 may present menu 1210 with office tools 1212 because word processing is more appropriately performed on a horizontal surface. On the other hand, wearable device 12700 may be configured not to present the virtual TV application because virtual TV may be more appropriate for a vertical surface and the user is currently interacting with an object having a horizontal surface. However, if the user is standing in front of wall 1248, wearable device 1270 may include virtual TV in menu 1210 while excluding office tools 1212.
[0108] 13 and 14, a virtual user interface 1310 is on a wall 1350 having a vertical surface. As a result, the user interface 1310 may include a driving game 1316a as shown in Figure 13 and a virtual TV application 1316b as shown in Figure 14 because a user may have a better experience when performing these activities on a vertical surface.
[0109] Additionally or alternatively, the function, orientation, location, and affordances of an object may also be determined based on the type of object. For example, a sofa may be associated with entertainment activities such as watching TV, while a desk chair may be associated with work-related activities such as preparing financial documents. The wearable system may also determine affordances based on the size of the object. For example, a small table may be used to hold decorative items such as vases, while a large table may be used for family meals. As another example, affordances may also be based on the shape of the object. A table with a circular top may be associated with certain group games such as poker, while a table with a rectangular top may be associated with single-player games such as Tetris. (User characteristics)
[0110] The wearable system may also present options based on user characteristics, such as age, gender, education level, occupation, preferences, etc. The wearable system may identify these characteristics based on profile information provided by the user. In some embodiments, the wearable system may infer these characteristics based on the user's interactions with the wearable system (e.g., frequently viewed content, etc.). Based on the user's characteristics, the wearable system can present content that matches the user's characteristics. For example, as shown in FIG. 15 , if the user of the wearable device is a young child, the wearable device may provide options for children's music and lullabies in menu 1530 in bedroom 1500.
[0111] In some implementations, an environment may be shared by multiple people. The wearable system may analyze the characteristics of the people sharing the space and present only content appropriate for the people sharing the space. For example, the living room 1300 may be shared by an entire family. If the family has young children, the wearable device may present only movies with a rating suitable for unaccompanied children (e.g., “G” rated movies). In some embodiments, the wearable system may identify people in the same environment as the wearable system images the environment and present options based on the people present in the environment. For example, the wearable system may obtain images of the environment using an outward-facing imaging system, analyze the images using facial recognition techniques, and identify one or more people present in the images. If the wearable system determines that a child is wearing an HMD and is seated with their parents in the same living room, the wearable system may present movies designated as appropriate for children accompanied by an adult (e.g., “PG” rated movies) in addition to, or as an alternative to, G-rated movies.
[0112] The wearable system may present a virtual menu based on the user's preferences. The wearable system may infer the user's preferences based on previous usage patterns. The previous usage patterns may include information about the location and / or time when the virtual object was used. For example, every morning, when user 210 launches virtual screen 1250 in their office 1200 (shown in FIG. 12 ), user 210 typically checks their email first. Based on this usage pattern, the wearable system may display email application 1216 in menu 1210. In another example, every morning, when user 210 enters living room 1300, the user typically watches the news on virtual TV screen 1312 b (shown in FIG. 14 ). The wearable system may therefore display TV application 1312 b on virtual user interface 1310 based on the user's frequent use. However, the wearable system would not show email application 1210 in virtual user interface 1310 because user 210 does not typically check their email in their living room 1300. On the other hand, if user 210 frequently checks their email regardless of their location, the wearable system may also show email application 1216 in virtual user interface 1310.
[0113] The options in the menu may vary according to the time of day. For example, if the user 210 typically listens to jazz or pop music in the morning and plays driving games in the evening, the wearable system may present options for jazz and pop music in the menu 1430 in the morning, while presenting driving games 1316a in the evening.
[0114] The AR system may also allow users to input their preferences. For example, a user 210 may add the contact information of their boss to their contact list 1220 (shown in FIG. 12 ) even though they may not speak with the boss frequently. (Interaction between the user and objects in the environment)
[0115] The wearable device may present a subset of virtual objects in the user's environment based on the current user interaction. The wearable device may present virtual objects based on the person the user is interacting with. For example, when a user has a telepresence session with one of their family members in their living room 1300, the wearable device may automatically erect a photo album on the wall because the user may want to talk about the shared experience captured by the photo album. On the other hand, when a user has a telepresence session with one of their coworkers in their office 1200, the wearable device may automatically present a document on which the coworker is collaborating.
[0116] The wearable device may also present virtual objects based on the virtual object the user is interacting with. For example, if the user is currently preparing a financial document, the wearable device may present the user with a data analysis tool such as a calculator. However, if the user is currently writing a novel, the wearable device may present the user with a word processing tool. (the user's physiological or psychological state)
[0117] The contextual information may include the user's physiological state, psychological state, or autonomic nervous system activity, a combination thereof, or the like. As described with reference to FIG. 2, the wearable system can use various sensors 232 to measure the user's response to virtual content or the environment. For example, one or more sensors 232 may acquire data of the user's eye region and use such data to determine the user's mood. The wearable system can acquire images of the eye using an inward-facing imaging system 462 (shown in FIG. 4). The ARD can use the images to determine eye movement, mydriasis, and heart rate variability. In some implementations, when the inward-facing imaging system 462 has a sufficiently large field of view, the wearable system can use images acquired by the inward-facing imaging system 462 to determine the user's facial expression. The wearable system can also use an outward-facing imaging system 464 (shown in FIG. 2) to determine the user's facial expression. For example, the outward-facing imaging system 464 can obtain a reflected image of the user's face when the user is standing near a reflective surface (such as a mirror). The wearable system can analyze the reflected image and determine the user's facial expression. Additionally or alternatively, the wearable system can include a sensor that measures skin potential (such as a galvanic skin response). The sensor may be part of the user's wearable glove and / or user input device 466 (described with reference to FIG. 4). The wearable system can use the skin potential data to determine the user's emotion.
[0118] The wearable system may also include sensors for electromyography (EMG), electroencephalography (EEG), functional near-infrared (fNIR), etc. The wearable system can use data obtained from these sensors to determine the user's psychological and physiological state. This data may be used alone or in combination with data obtained from other sensors, such as inward-facing imaging systems, outward-facing imaging systems, and sensors for measuring electrodermal activity. The ARD can use information about the user's psychological and physiological state to present virtual content (such as a virtual menu) to the user.
[0119] As an example, the wearable system can suggest a set of entertainment content (e.g., games, movies, music, scenes to be displayed) based on the user's mood. The entertainment content may be suggested to improve the user's mood. For example, the wearable system may determine that the user is currently under stress based on the user's physiological data (such as sweating). The wearable system can further determine that the user is at work based on information obtained from a location sensor (such as a GPS) or images obtained from the outward-facing imaging system 464. The wearable system can combine these two sets of information and determine that the user is experiencing stress at work. Thus, the wearable system can suggest that the user play a slow-paced, exploratory, and non-restrictive game, display relaxing scenes on a virtual display, play calming music in the user's environment, etc. to calm the user during their lunch break.
[0120] As another example, multiple users may be present together or interact with each other in a physical or virtual space. The wearable system can determine one or more shared moods among the users (such as whether the group is happy or angry). The shared mood may be a combination or fusion of the moods of multiple users or may target a common mood or theme among the users. The wearable system can present a virtual activity (such as a game) based on the shared moods among the users. (reference marker)
[0121] In some implementations, the contextual information may be encoded within a fiducial marker (also referred to herein as a label). The fiducial marker may be associated with a physical object. The fiducial marker may be an optical marker, such as a quick response (QR) code, a barcode, an ArUco marker (which can be reliably detected under occlusion), etc. The fiducial marker may also comprise an electromagnetic marker (e.g., a radio frequency identification tag) that may emit or receive an electromagnetic signal detectable by the wearable device. Such a fiducial marker may be physically affixed on or near the physical object. The wearable device can detect the fiducial marker using an outward-facing imaging system 464 (shown in FIG. 4) or one or more sensors that receive or transmit signals from or to the fiducial marker (e.g., transmit a signal to the marker, which may then relay a signal back).
[0122] When the wearable device detects the fiducial marker, the wearable device can decode the fiducial marker and present a group of virtual objects based on the decoded fiducial marker. In some embodiments, the fiducial marker may contain a reference to a database that includes an association between the virtual object to be displayed and an associated contextual factor. For example, the fiducial marker may include an identifier for a physical object (e.g., a table). The wearable device can use the identifier to access contextual characteristics of the physical object. In this example, the contextual characteristics of the table include the size of the table and the horizontal surface. The accessed characteristics can be used to determine user interface actions or virtual objects supported by the physical object. Because the table has a horizontal surface, the wearable device can present office processing tools on the surface of the table rather than paintings, because paintings are typically associated with vertical, not horizontal, surfaces.
[0123] For example, in FIG. 15 , an ArUco marker 1512 is attached to a window 1510 in a bedroom 1500. The wearable system can identify the ArUco marker 1512 (e.g., based on an image acquired by the outward-facing imaging system 464) and extract information encoded within the ArUco marker 1512. In some embodiments, the extracted information may be sufficient for the wearable system to render a virtual menu associated with the reference marker without analyzing other contextual information of the window or the user's environment. For example, the extracted information may include the orientation of the window 1510 (e.g., the window 1510 has a vertical surface). Based on the orientation, the wearable system can present a virtual object associated with the vertical surface. In some implementations, the wearable system may need to communicate with another data source (e.g., remote data repository 280) to obtain contextual information such as the type of object (e.g., a mirror), the environment associated with the object, the user's preferences, etc., in order for the wearable system to render and identify the relevant virtual object. Exemplary Method for Rendering Virtual Objects Based on Contextual Factors
[0124] FIG. 16 is a flowchart of an exemplary method for generating a virtual menu based on context information. Process 1600 can be implemented by a wearable system described herein. The wearable system can include a user input device (e.g., see user input device 466 in FIG. 4 ) configured to receive indications of various user interactions, a display that can display virtual objects in proximity to physical objects, and a posture sensor. The posture sensor can detect and track a user's posture, which can include the user's body orientation relative to the environment or the position or movement of a part of the user's body, such as a gesture made by the user's hand or the user's gaze direction. The posture sensor can include an inertial measurement unit (IMU), an outward-facing imaging system, and / or an eye-tracking camera (e.g., camera 464 shown in FIG. 4 ), as described with reference to FIGS. 2, 4, and 9 . The IMU can include an accelerometer, a gyroscope, and other sensors.
[0125] In block 1610, the wearable system may determine the user's posture using one or more posture sensors. As described herein, the posture may include eye posture, head posture, body posture, a combination thereof, or the like. Based on the user's posture, in block 1620, the wearable system may identify interactable objects in the user's environment. For example, the wearable system may use a cone projection technique to identify objects that intersect with the user's line of sight.
[0126] In block 1630, a user may actuate a user input device to provide an indication to open a virtual menu associated with the interactable object. The virtual menu may include multiple virtual objects as menu options. The multiple virtual objects may be a subset of the virtual objects in the user's environment or a subset of the virtual objects associated with the interactable object. The virtual menu may have many graphical representations. Some examples of virtual menus are shown as object 1220 and object 1210 in FIG. 12, virtual user interface 1310 in FIGS. 13 and 14, object 1320 in FIG. 13, object 1430 in FIG. 14, and object 1530 in FIG. 15. In some implementations, the indication to open the virtual menu need not be received from a user input device. The indication may be associated with direct user input, such as the user's head posture, eye gaze, body posture, gesture, voice command, etc.
[0127] In block 1640, the wearable system can determine contextual information associated with the interactable object. The contextual information can include the affordances of the interactable object, features of the environment (e.g., work environment or living environment), characteristics of the user (such as the user's age or preferences), or the user's current interaction with objects in the environment, combinations thereof, or the like. For example, the wearable system can determine the affordances of the interactable object by analyzing characteristics of the interactable object, such as its function, orientation (horizontal vs. vertical), location, shape, size, etc. The wearable system can also determine the affordances of the interactable object by analyzing its relationship to the environment. For example, an end table in a living room environment may be used for entertainment purposes, while an end table in a bedroom environment may be used to hold items before a person goes to bed.
[0128] Contextual information associated with an interactable object may also be determined from characteristics of the user or the user's interactions with objects in the environment. For example, the wearable system may identify the user's age and present only information appropriate to the user's age. As another example, the wearable system may analyze the user's previous usage patterns with a virtual menu (such as the types of virtual objects the user frequently uses) and adjust the content of the virtual menu according to the previous usage patterns.
[0129] In block 1650, the wearable system can identify a list of virtual objects to be included in the virtual menu based on the context information. For example, the wearable system can identify a list of applications with features related to the interactable objects and / or features related to the user's environment. As another example, the wearable system can identify virtual objects to be included in the virtual menu based on user characteristics such as age, gender, and previous usage patterns.
[0130] In block 1660, the wearable system may generate a virtual menu based on the identified list of virtual objects. The virtual menu may include all virtual objects on the identified list. In some embodiments, the menu may be limited in space. The wearable system may prioritize different types of contextual information so that only a subset of the list is shown to the user. For example, the wearable system may determine that previous usage patterns are the most important contextual information and therefore display only the top five virtual objects based on previous usage patterns.
[0131] A user can perform various actions with the menu, such as browsing through the menu, showing available virtual objects not previously selected based on an analysis of some of the context information, exiting the menu, or selecting one or more objects on the menu to interact with. Exemplary Method for Rendering Virtual Objects Based on Physiological Data of a User
[0132] 17 is a flowchart of an example method for selecting virtual content based, at least in part, on physiological data of a user. Process 1700 can be implemented by a wearable system described herein. The wearable system may include various sensors 232, such as physiological sensors configured to measure physiological parameters of the user, an inward-facing imaging system configured to track the user's eye region, etc.
[0133] In block 1710, the wearable system can acquire physiological data of the user. As described with reference to Figures 2 and 4, the wearable system can use physiological sensors, alone or in combination with an inward-facing imaging system, to measure the user's physiological data. For example, the wearable system can use one or more physiological sensors to determine the user's galvanic skin response. The wearable system can also use an inward-facing imaging system to determine the user's eye movement.
[0134] As shown in block 1720, the wearable system can use physiological data to determine the user's physiological or psychological state. For example, the wearable system can use the user's galvanic skin response and / or the user's eye movements to determine whether the user is excited by certain content.
[0135] In block 1730, the wearable system may determine virtual content to be presented to the user. The wearable system may make such a determination based on physiological data. For example, using the physiological data, the wearable system may determine that the user is under stress. The wearable system may then present a virtual object (such as music or a video game) associated with relieving the user's stress.
[0136] The wearable system can also present virtual content based on an analysis of physiological data in combination with other contextual information. For example, based on a user's location, the wearable system may determine that the user is experiencing stress at work. Because users typically do not play games or listen to music during work hours, the wearable system may suggest video games and music only during the user's breaks to help the user reduce their stress levels.
[0137] In block 1740, the wearable system can generate a 3D user interface including the virtual content. For example, the wearable system may display icons for music and video games when it detects that the user is experiencing stress. The wearable system can present a virtual menu while the user interacts with physical objects. For example, the wearable system can display icons for music and video games when the user activates a user input device in front of their desk during their work break.
[0138] Techniques in various embodiments described herein can provide a user with a subset of available virtual objects or user interface interaction options. This subset of virtual objects or user interface interaction options can be provided in various forms. While embodiments are described primarily with reference to presenting menus, other types of user interface presentations are also available. For example, the wearable system can render icons of virtual objects within the subset of virtual objects. In some implementations, the wearable system can automatically perform actions based on context information. For example, the wearable system may automatically initiate a telepresence session with the user's most frequent contacts when the user is near a mirror. As another example, the wearable system can automatically invoke a virtual object if the wearable system determines that the user is likely to be most interested in that virtual object. (Other embodiments)
[0139] In a first aspect, a method for generating a virtual menu in a user's environment in three-dimensional (3D) space, the method comprising: an augmented reality (AR) system comprising computer hardware; the AR system configured to enable user interaction with objects in the user's environment; the AR system comprising a user input device, an AR display, and an inertial measurement unit (IMU) configured to detect a user's posture; determining a user's posture using the IMU under control of the AR system; identifying a physical object in the user's environment in the 3D space based at least in part on the user's posture; receiving an indication via the user input device to open a virtual menu associated with the physical object; determining context information associated with the physical object; determining virtual objects to be included in the virtual menu based at least in part on the determined context information; determining a spatial location for displaying the virtual menu based at least in part on the determined context information; generating the virtual menu including at least the determined virtual objects; and displaying the generated menu to a user at the spatial location via the AR display.
[0140] In a second aspect, the pose comprises one or more of a head pose or a body pose.
[0141] In a third aspect, the method of aspect 1 or aspect 2, wherein the ARD further comprises an eye-tracking camera configured to track the eye posture of the user.
[0142] In a fourth aspect, the method of aspect 3, wherein the posture includes an eye posture.
[0143] In a fifth aspect, the method of any one of aspects 1-4, wherein the contextual information includes one or more of affordances of physical objects, features of the environment, characteristics of the user, or current or past interactions of the user with the AR system.
[0144] In a sixth aspect, the method of aspect 5, wherein the affordance of a physical object includes a relationship between the physical object and the physical object's environment that provides an opportunity for an action or use associated with the physical object.
[0145] In a seventh aspect, the method of aspect 5 or aspect 6, wherein the affordance of the physical object is based, at least in part, on one or more of the function, orientation, type, location, shape, size of the physical object, or the environment in which the physical object is located.
[0146] In an eighth aspect, the method of aspect 7, wherein the orientation of the physical object comprises horizontal or vertical.
[0147] In a ninth aspect, the method of any one of aspects 5-8, wherein the environment is a living environment or a working environment.
[0148] In a tenth aspect, the method of any one of aspects 5-9, wherein the environment is a private environment or a public environment.
[0149] In an eleventh aspect, the method of aspect 5, wherein the user characteristics include one or more of age, gender, education level, occupation, or preferences.
[0150] In a twelfth aspect, the preference is based, at least in part, on the user's previous usage patterns, the previous usage patterns including information about where or when the virtual object was used.
[0151] In a thirteenth aspect, the method of aspect 5, wherein the current interaction includes a telepresence session between the user of the AR system and another user.
[0152] In a fourteenth aspect, the method of any one of aspects 1-13, wherein the context information is encoded within a reference marker, and the reference marker is associated with a physical object.
[0153] In a fifteenth aspect, the method of aspect 14, wherein the fiducial marker comprises an optical marker or an electromagnetic marker.
[0154] In a sixteenth aspect, the method of any one of aspects 1-15, wherein the object includes at least one of a physical object or a virtual object.
[0155] In a seventeenth aspect, a method for rendering a plurality of virtual objects in a user's environment in three-dimensional (3D) space, the method comprising: an augmented reality (AR) system comprising computer hardware; the AR system configured to enable user interaction with objects in the user's environment; the AR system comprising a user input device, an AR display, and an attitude sensor configured to detect the user's attitude; determining a user's attitude using the attitude sensor under control of the AR system; identifying an interactable object in the user's environment in the 3D space based at least in part on the user's attitude; receiving, via the user input device, an indication to present a plurality of virtual objects associated with the interactable object; determining contextual information associated with the interactable object; determining a plurality of virtual objects to be displayed to the user based at least in part on the determined contextual information; and displaying the determined plurality of virtual objects to the user via the AR display.
[0156] In an eighteenth aspect, the method of aspect 17, wherein the attitude sensor includes one or more of an inertial measurement unit, an eye-tracking camera, or an outward-facing imaging system.
[0157] In a nineteenth aspect, the method of aspect 17 or aspect 18, wherein the posture includes one or more of head posture, eye posture, or body posture.
[0158] In a twentieth aspect, the method of any one of aspects 17-19, wherein the context information includes one or more of affordances of interactable objects, features of the environment, characteristics of the user, or current or past interactions of the user with the AR system.
[0159] In a twenty-first aspect, the method of aspect 20, wherein the affordance of an interactable object includes a relationship between the interactable object and the environment of the interactable object that provides an opportunity for an action or use associated with the interactable object.
[0160] In a twenty-second aspect, the method of aspect 20 or aspect 21, wherein the affordance of the interactable object is based, at least in part, on one or more of the function, orientation, type, location, shape, size of the interactable object, or the environment in which the physical object is located.
[0161] In a twenty-third aspect, the method of aspect 22, wherein the orientation of the interactable object comprises horizontal or vertical.
[0162] In a 24th aspect, the method of any one of aspects 20-23, wherein the environment is a living environment or a working environment.
[0163] In a 25th aspect, the method of any one of aspects 20-24, wherein the environment is a private environment or a public environment.
[0164] In a twenty-sixth aspect, the method of aspect 20, wherein the user characteristics include one or more of age, gender, education level, occupation, or preferences.
[0165] In a twenty-seventh aspect, the method of aspect 26, wherein the preference is based, at least in part, on the user's previous usage patterns, including information about where or when the virtual object was used.
[0166] In a twenty-eighth aspect, the method of aspect 20, wherein the current interaction includes a telepresence session.
[0167] In a 29th aspect, the method of any one of aspects 17-28, wherein the context information is encoded within a reference marker, and the reference marker is associated with the interactable object.
[0168] In a 30th aspect, the method of aspect 29, wherein the reference marker includes an optical marker or an electromagnetic marker.
[0169] In a thirty-first aspect, the method of any one of aspects 17-30, wherein the object includes at least one of a physical object or a virtual object.
[0170] In a thirty-second aspect, the method of any one of aspects 17-31, wherein the interactable object includes at least one of a physical object or a virtual object.
[0171] In a thirty-third aspect, an augmented reality (AR) system includes computer hardware, a user input device, an AR display, and an orientation sensor, and is configured to perform any one of the methods described in aspects 1-32.
[0172] In a thirty-fourth aspect, a method for selectively presenting virtual content to a user in three-dimensional space (3D), the method including: under control of a wearable device comprising a computer processor, a display, and a physiological sensor configured to measure a physiological parameter of the user, using the physiological sensor to acquire data associated with the physiological parameter; determining a physiological state of the user based at least in part on the data; determining virtual content to be presented to the user based at least in part on the physiological state; determining a spatial location for displaying the virtual content within the 3D space; generating a virtual user interface including at least the determined virtual content; and displaying the virtual content to the user at the determined spatial location via a display of the wearable device.
[0173] In a thirty-fifth aspect, the method of aspect 34, wherein the physiological parameters include at least one of heart rate, mydriasis, galvanic skin response, blood pressure, electroencephalogram, respiratory rate, or eye movement.
[0174] In a 36th aspect, a method described in any one of aspects 34-35, further comprising acquiring data associated with the user's physiological parameters using an inward-facing imaging system configured to image one or both of the user's eyes.
[0175] In a 37th aspect, the method of any one of aspects 34-36 further comprises determining a psychological state based on the data.
[0176] In a thirty-eighth aspect, the method of any one of aspects 34-37, wherein the virtual content includes a virtual menu.
[0177] In a thirty-ninth aspect, the method of any one of aspects 34-38, wherein the virtual content is further determined based on at least one of affordances of a physical object associated with the virtual content, features of the user's environment, characteristics of the user, individuals present in the environment, information encoded in fiducial markers associated with the physical object, or the user's current or past interactions with the wearable system.
[0178] In a fortieth aspect, the method of aspect 39, wherein the affordance of a physical object includes a relationship between the physical object and the physical object's environment that provides an opportunity for an action or use associated with the physical object.
[0179] In a forty-first aspect, the method of any one of aspects 39 or 40, wherein the affordance of the physical object is based, at least in part, on one or more of the function, orientation, type, location, shape, size, or environment in which the physical object is located.
[0180] In a forty-second aspect, the orientation of the physical object comprises horizontal or vertical, as described in aspect 41.
[0181] In a 43rd aspect, the method of any one of aspects 39-42, wherein the environment is a living environment or a working environment.
[0182] In a forty-fourth aspect, the method of any one of aspects 39-43, wherein the environment is a private environment or a public environment.
[0183] In a forty-fifth aspect, the method of aspect 39, wherein the user characteristics include one or more of age, gender, education level, occupation, or preferences.
[0184] In a forty-sixth aspect, the preference is based, at least in part, on the user's previous usage patterns, the previous usage patterns including information about where or when the virtual object was used, according to aspect 45.
[0185] In a forty-seventh aspect, the method of aspect 39, wherein the current interaction includes a telepresence session between a user of the AR system and another user.
[0186] In a forty-eighth aspect, the method of any one of aspects 34-47, wherein the wearable device includes an augmented reality system.
[0187] In a forty-ninth aspect, a wearable device is provided comprising a computer processor, a display, and a physiological sensor configured to measure physiological parameters of a user, the wearable device being configured to perform any one of the methods described in aspects 34-48.
[0188] In a fiftieth aspect, a wearable system for generating virtual content within a user's three-dimensional (3D) environment, the wearable system comprising: an augmented reality display configured to present the virtual content to the user in a 3D view; an attitude sensor configured to obtain user position or orientation data, analyze the position or orientation data, and identify the user's posture; and a hardware processor in communication with the attitude sensor and the display, the hardware processor being programmed to: identify a physical object in the user's environment within the 3D environment based at least in part on the user's posture; receive an indication to initiate an interaction with the physical object; identify a set of virtual objects in the user's environment associated with the physical object; determine contextual information associated with the physical object; filter the set of virtual objects and identify a subset of virtual objects from the set of virtual objects based on the contextual information; generate a virtual menu including the subset of virtual objects; determine a spatial location within the 3D environment for presenting the virtual menu based at least in part on the determined contextual information; and present the virtual menu at the spatial location via the augmented reality display.
[0189] In a fifty-first aspect, the wearable system of aspect 50, wherein the contextual information includes an affordance of the physical object, including a relationship between the physical object and the physical object's environment that provides an opportunity for an action or use associated with the physical object, and the affordance of the physical object is based, at least in part, on one or more of the physical object's function, orientation, type, location, shape, size, or the environment in which the physical object is located.
[0190] In a 52nd aspect, the wearable system described in aspect 51, wherein the context information includes an orientation of a surface of the physical object, and to filter the set of virtual objects, the hardware processor is programmed to identify a subset of virtual objects that support user interface interaction on a surface having the orientation.
[0191] In aspect 53, a wearable system described in any one of aspects 50-52, wherein the posture sensor includes an inertial measurement unit configured to measure the user's head posture and identify physical objects, and the hardware processor is programmed to project a virtual cone and select a physical object that intersects with a portion of the virtual cone based, at least in part, on the user's head posture.
[0192] In a 54th aspect, a wearable system described in any one of aspects 50-53 further comprises a physiological sensor configured to measure physiological parameters of the user, and the hardware processor is programmed to determine a psychological state of the user, use the psychological state as part of the context information, and identify a subset of virtual objects for inclusion in the virtual menu.
[0193] In a fifty-fifth aspect, the wearable system of aspect 54, wherein the physiological parameter is related to at least one of heart rate, mydriasis, galvanic skin response, blood pressure, brain wave state, respiratory rate, or eye movement.
[0194] In a 56th aspect, the 3D environment includes a plurality of users, and the hardware processor is programmed to determine common characteristics of the plurality of users and filter the set of virtual objects based on the common characteristics of the plurality of users, a wearable system described in any one of aspects 50-55.
[0195] In a 57th aspect, the wearable system of any one of aspects 50-56 is configured such that the context information includes the user's past interactions with the set of virtual objects, and the hardware processor is programmed to identify one or more virtual objects with which the user frequently interacts, and include the one or more virtual objects within the subset of virtual objects for the virtual menu.
[0196] In aspect 58, a wearable system described in any one of aspects 50-57 is programmed to: identify a reference marker associated with the physical object, the reference marker encoding an identifier of the physical object, decode the reference marker to extract the identifier, access a database storing contextual information associated with the physical object using the identifier, and analyze the contextual information stored in the database to filter the set of virtual objects, in order to determine contextual information associated with the physical object and filter the set of virtual objects.
[0197] In a fifty-ninth aspect, the wearable system of aspect 58, wherein the reference marker includes an ArUco marker.
[0198] In a 60th aspect, a wearable system described in any one of aspects 50-59, wherein the spatial location for rendering the virtual menu includes a position or orientation of the virtual menu relative to a physical object.
[0199] In aspect 61, a wearable system as described in aspect 60, wherein to determine a spatial location for rendering the virtual menu, the hardware processor is programmed to identify a space on the surface of the physical object using an object recognition device associated with the physical object.
[0200] In a sixty-second aspect, a method for generating virtual content within a user's three-dimensional (3D) environment includes analyzing data obtained from a posture sensor to identify a posture of the user; identifying an interactable object within the user's 3D environment based, at least in part, on the posture; receiving an indication to initiate an interaction with the interactable object; determining context information associated with the interactable object; selecting a subset of user interface actions from a set of user interface actions available on the interactable object based on the context information; and generating instructions for presenting the subset of user interface actions to the user in a 3D view.
[0201] In a 63rd aspect, the method of aspect 62, wherein the posture includes at least one of eye gaze, head posture, or gesture.
[0202] In aspect 64, the method of any one of aspects 62-63, wherein the step of identifying the interactable object includes the steps of performing a cone projection based on the user's head pose, and selecting an object in the user's environment as an interactable object that intersects at least a portion of a virtual cone used in the cone projection.
[0203] In aspect 65, the context information includes an orientation of a surface of the interactable object, and the step of selecting a subset of user interface actions includes a step of identifying user interface actions that can be performed on a surface having the orientation.
[0204] In a 66th aspect, the method of any one of aspects 62-65, wherein the indication to initiate an interaction with the interactable object includes actuation of a user input device or a change in the user's posture.
[0205] In aspect 67, a method described in any one of aspects 62-66 further includes steps of receiving physiological parameters of a user and determining a psychological state of the user, the psychological state being part of the context information for selecting a subset of user interactions.
[0206] In a 68th aspect, the method of aspect 67, wherein the physiological parameter is related to at least one of heart rate, mydriasis, galvanic skin response, blood pressure, electroencephalogram (EEG) status, respiratory rate, or eye movement.
[0207] In a 69th aspect, the method of any one of aspects 62-68, wherein the step of generating instructions for presenting a subset of user interactions to a user in a 3D view includes the steps of generating a virtual menu including the subset of user interface actions, determining a spatial location of the virtual menu based on characteristics of the interactable objects, and generating display instructions for presentation of the virtual menu at the spatial location within the user's 3D environment. (Other considerations)
[0208] Each of the processes, methods, and algorithms described herein and / or depicted in the accompanying figures may be embodied in code modules executed by one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to execute specific and particular computer instructions, and thereby may be fully or partially automated. For example, a computing system may include a general-purpose computer (e.g., a server) or a special-purpose computer, special-purpose circuitry, etc., programmed with specific computer instructions. Code modules may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language. In some implementations, particular operations and methods may be performed by circuitry specific to a given function.
[0209] Furthermore, certain implementations of the functionality of the present disclosure may be sufficiently mathematically, computationally, or technically complex that special-purpose hardware (utilizing appropriate specialized executable instructions) or one or more physical computing devices may be required to implement the functionality, for example, due to the amount or complexity of the calculations involved, or to provide results in substantially real time.
[0210] Code modules or any type of data may be stored on any type of non-transitory computer-readable medium, such as physical computer storage devices, including hard drives, solid-state memory, random-access memory (RAM), read-only memory (ROM), optical disks, volatile or non-volatile storage devices, combinations of the same, and / or the like. The methods and modules (or data) may also be transmitted as data signals (e.g., as part of a carrier wave or other analog or digital propagated signal) generated over various computer-readable transmission media, including wireless-based and wired / cable-based media, and may take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed processes or process steps may be stored, persistently or otherwise, in any type of non-transitory tangible computer storage device or communicated via a computer-readable transmission medium.
[0211] Any process, block, state, step, or functionality in the flow diagrams described herein and / or depicted in the accompanying figures should be understood as potentially representing a code module, segment, or portion of code, comprising one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in the process. Various processes, blocks, states, steps, or functionality can be combined, rearranged, added, deleted, modified, or otherwise changed from the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may perform some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the blocks, steps, or states associated therewith can be performed in other suitable sequences, e.g., serially, in parallel, or in some other manner. Tasks or events may be added to or removed from the disclosed exemplary embodiments. Furthermore, the separation of various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems may generally be integrated together in a single computer product or packaged in multiple computer products. Many implementation variations are possible.
[0212] The processes, methods, and systems can be implemented in a network (or distributed) computing environment. Network environments include enterprise-wide computer networks, intranets, local area networks (LANs), wide area networks (WANs), personal area networks (PANs), cloud computing networks, crowdsourced computing networks, the Internet, and the World Wide Web. The network can be a wired or wireless network or any other type of communication network.
[0213] The systems and methods of the present disclosure each have several innovative aspects, none of which is solely responsible for or required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure. Various modifications of the implementations described in the present disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations without departing from the spirit or scope of the present disclosure. Therefore, the claims are not intended to be limited to the implementations shown herein, but are to be accorded the widest scope consistent with the present disclosure, the principles, and novel features disclosed herein.
[0214] Certain features described herein in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented separately in multiple implementations or in any suitable subcombination. Furthermore, while features may be described above as operative in a combination and may even be initially claimed as such, one or more features from the claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination. No single feature or group of features is required or essential to every embodiment.
[0215] Conditional statements used herein, such as "can," "could," "might," "may," "e.g.," and the like, among others, are intended to generally convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context as used. Thus, such conditional statements are not generally intended to imply that features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps should be included or performed in any particular embodiment, with or without authorial input or prompting. The terms "comprise," "include," "have," and the like are synonymous and used inclusively in a non-limiting manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not its exclusive sense), so, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Additionally, the articles "a," "an," and "the," as used in this application and the appended claims, should be interpreted to mean "one or more" or "at least one," unless otherwise specified.
[0216] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single elements. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Transitional phrases such as "at least one of X, Y, and Z," unless specifically stated otherwise, are generally understood differently in the context in which they are used to convey that an item, term, etc. may be at least one of X, Y, or Z. Thus, such transitional phrases generally are not intended to suggest that an embodiment requires that at least one of X, at least one of Y, and at least one of Z, respectively, be present.
[0217] Similarly, while operations may be depicted in the figures in a particular order, it should be recognized that such operations need not be performed in the particular order shown, or in sequential order, or that all of the depicted operations need not be performed to achieve desirable results. Furthermore, the figures may diagrammatically depict one or more example processes in the form of a flowchart. However, other operations not depicted may be incorporated within the diagrammatically depicted example methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any of the depicted operations. Additionally, operations may be rearranged or reordered in other implementations. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. (Item 1) 1. A wearable system for generating virtual content within a user's three-dimensional (3D) environment, the wearable system comprising: an augmented reality display configured to present the virtual content to the user in a 3D view; a posture sensor configured to obtain position or orientation data of a user, analyze said position or orientation data, and identify a posture of said user; a hardware processor in communication with the attitude sensor and the display, the hardware processor comprising: identifying physical objects in the user's environment within the 3D environment based at least in part on the user's pose; receiving an indication to initiate an interaction with the physical object; identifying a set of virtual objects in the user's environment that are associated with the physical object; determining context information associated with the physical object; filtering the set of virtual objects to identify a subset of virtual objects from the set of virtual objects based on the context information; generating a virtual menu including a subset of the virtual objects; determining a spatial location within the 3D environment for presenting the virtual menu based at least in part on the determined context information; and presenting the virtual menu at the spatial location by the augmented reality display; and a hardware processor programmed to perform the A wearable system comprising: (Item 2) Item 1. The wearable system of item 1, wherein the context information includes an affordance of the physical object, including a relationship between the physical object and the environment of the physical object that provides an opportunity for an action or use associated with the physical object, and the affordance of the physical object is based, at least in part, on one or more of the function, orientation, type, location, shape, size, or environment in which the physical object is located. (Item 3) Item 3. The wearable system of item 2, wherein the context information includes an orientation of a surface of the physical object, and to filter the set of virtual objects, the hardware processor is programmed to identify a subset of the virtual objects that support user interface interactions on the surface having the orientation. (Item 4) Item 10. The wearable system of item 1, wherein the posture sensor comprises an inertial measurement unit configured to measure the user's head posture and identify the physical object, and the hardware processor is programmed to project a virtual cone and select the physical object that intersects with a portion of the virtual cone based, at least in part, on the user's head posture. (Item 5) Item 10. The wearable system of item 1, further comprising a physiological sensor configured to measure physiological parameters of the user, wherein the hardware processor is programmed to determine a psychological state of the user, use the psychological state as part of the context information, and identify a subset of the virtual objects for inclusion in the virtual menu. (Item 6) Item 6. The wearable system of item 5, wherein the physiological parameter relates to at least one of heart rate, mydriasis, galvanic skin response, blood pressure, brain wave state, respiratory rate, or eye movement. (Item 7) Item 10. The wearable system of item 1, wherein the 3D environment includes a plurality of users, and the hardware processor is programmed to determine common characteristics of the plurality of users and filter the set of virtual objects based on the common characteristics of the plurality of users. (Item 8) Item 10. The wearable system of item 1, wherein the context information includes the user's past interactions with the set of virtual objects, and the hardware processor is programmed to identify one or more virtual objects with which the user frequently interacts and include the one or more virtual objects within the subset of virtual objects for the virtual menu. (Item 9) To determine context information associated with the physical objects and filter the set of virtual objects, the hardware processor: identifying a fiducial marker associated with the physical object, the fiducial marker encoding an identifier of the physical object; decoding the fiducial marker to extract the identifier; accessing a database using the identifier that stores context information associated with the physical object; analyzing context information stored in said database to filter said set of virtual objects; Item 1. The wearable system of item 1, programmed to: (Item 10) Item 10. The wearable system of item 9, wherein the reference marker includes an ArUco marker. (Item 11) Item 1. The wearable system of item 1, wherein the spatial location for rendering the virtual menu includes a position or orientation of the virtual menu relative to the physical object. (Item 12) Item 12. The wearable system of item 11, wherein the hardware processor is programmed to identify a space on a surface of the physical object using an object recognizer associated with the physical object to determine the spatial location for rendering the virtual menu. (Item 13) 1. A method for generating virtual content within a user's three-dimensional (3D) environment, the method comprising: analyzing data obtained from the posture sensor to identify a posture of the user; Identifying an interactable object within the user's 3D environment based at least in part on the pose; and receiving an indication to initiate an interaction with the interactable object; determining context information associated with the interactable object; selecting a subset of user interface actions from a set of user interface actions available on the interactable object based on the context information; generating instructions for presenting the subset of user interface actions to the user in a 3D view; A method comprising: (Item 14) Item 14. The method of item 13, wherein the posture includes at least one of eye gaze, head posture, or gesture. (Item 15) Item 14. The method of item 13, wherein identifying the interactable object includes performing a cone projection based on the user's head pose, and selecting an object in the user's environment as the interactable object that intersects at least a portion of a virtual cone used in the cone projection. (Item 16) Item 14. The method of item 13, wherein the context information includes an orientation of a surface of the interactable object, and selecting the subset of user interface actions includes identifying user interface actions that can be performed on a surface having the orientation. (Item 17) Item 14. The method of item 13, wherein the indication to initiate an interaction with the interactable object includes an actuation of a user input device or a change in the user's posture. (Item 18) 14. The method of claim 13, further comprising receiving physiological parameters of the user and determining a psychological state of the user, the psychological state being part of the context information for selecting the subset of user interactions. (Item 19) 20. The method of claim 18, wherein the physiological parameter relates to at least one of heart rate, mydriasis, galvanic skin response, blood pressure, electroencephalogram, respiratory rate, or eye movement. (Item 20) Item 14. The method of item 13, wherein generating instructions for presenting a subset of the user interactions to the user in a 3D view includes generating a virtual menu including the subset of user interface actions, determining a spatial location of the virtual menu based on characteristics of the interactable objects, and generating display instructions for presentation of the virtual menu at a spatial location within the user's 3D environment.
Claims
1. 1. A method for generating virtual content within a user's three-dimensional (3D) physical environment, the method comprising: Under the control of a hardware processor, analyzing data obtained from the posture sensor to identify a posture of the user; identifying a physical surface within the user's 3D physical environment based at least in part on the pose; receiving an indication to initiate an interaction with the physical surface; determining an orientation of the physical surface; accessing information indicative of a first set of virtual objects associated with a first surface orientation and a second set of virtual objects associated with a second surface orientation; selecting the first set of virtual objects in response to determining that the orientation of the physical surface is the first surface orientation; selecting the second set of virtual objects in response to determining that the orientation of the physical surface is the second surface orientation; and determining a set of available user interface options corresponding to the selected first or second set of virtual objects; determining a first number of user interface options from the set of available user interface options that can readably fit into a virtual user interface associated with an area available on the physical surface, the first number being less than a number of user interface options in the set of available user interface options; selecting a subset of user interface options from the set of available user interface options equal to the first number; generating the virtual user interface including the subset of user interface options; generating instructions for presenting to the user in a 3D view the virtual user interface containing the subset of user interface options that can be readably fitted within the virtual user interface; generating instructions for displaying a corresponding first virtual object on the physical surface in response to a selection of a first user interface option from the subset of user interface options; A method comprising:
2. The method of claim 1 , wherein the pose includes at least one of an eye gaze, a head pose, or a gesture.
3. The method of claim 1 , wherein identifying the physical surface comprises performing a cone projection based on the pose.
4. The method of claim 1 , wherein the indication to initiate the interaction with the physical surface comprises an actuation of a user input device or a change in the user's posture.
5. 10. The method of claim 1, further comprising receiving physiological parameters of the user and determining a psychological state of the user, wherein determining the set of available user interface options is based at least in part on the psychological state of the user.
6. The method of claim 5 , wherein the physiological parameter relates to at least one of heart rate, mydriasis, galvanic skin response, blood pressure, electroencephalogram, respiratory rate, or eye movement.
7. Generating the instructions for presenting the subset of user interface options to the user in the 3D view includes: generating a virtual menu including said subset of user interface options; determining a spatial location of the virtual menu based on characteristics of the physical surface; generating display instructions for presentation of the virtual menu at the spatial location within the 3D physical environment of the user; The method of claim 1 , comprising:
8. 1. A system for generating virtual content within a user's physical environment, the system comprising: a display system of a wearable device configured to present virtual content; a hardware processor in communication with the display system; Equipped with The hardware processor includes: determining a posture of the user based on data obtained from the posture sensor and identifying the posture of the user; identifying a physical surface within the user's 3D physical environment based at least in part on the pose; receiving an indication to initiate an interaction with the physical surface; determining an orientation of the physical surface; accessing information indicative of a first set of virtual objects associated with a first surface orientation and a second set of virtual objects associated with a second surface orientation; selecting the first set of virtual objects in response to determining that the orientation of the physical surface is the first surface orientation; selecting the second set of virtual objects in response to determining that the orientation of the physical surface is the second surface orientation; and determining a set of available user interface options corresponding to the selected first or second set of virtual objects; determining a first number of user interface options from the set of available user interface options that can readably fit into a virtual user interface associated with an area available on the physical surface, the first number being less than a number of user interface options in the set of available user interface options; selecting a subset of user interface options from the set of available user interface options equal to the first number; generating the virtual user interface including the subset of user interface options; generating instructions for presenting to the user in a 3D view the virtual user interface containing the subset of user interface options that can be readably fitted within the virtual user interface; generating instructions for displaying a corresponding first virtual object on the physical surface in response to a selection of a first user interface option from the subset of user interface options; A system that is programmed to:
9. The system of claim 8 , wherein the pose includes at least one of an eye gaze, a head pose, or a gesture.
10. The system of claim 8 , wherein the hardware processor is programmed to perform a cone projection based on the pose to identify the physical surface.
11. The system of claim 8 , wherein the indication to initiate the interaction with the physical surface comprises an actuation of a user input device or a change in the posture of the user.
12. 10. The system of claim 8, wherein the hardware processor is further programmed to receive physiological parameters of the user, and wherein determining the set of available user interface options is based at least in part on a psychological state of the user.
13. The system of claim 12 , wherein the user's physiological parameters relate to at least one of heart rate, mydriasis, galvanic skin response, blood pressure, brain wave state, respiratory rate, or eye movement.
14. The hardware processor includes:
10. The system of claim 8, further programmed to: generate a virtual menu including the subset of user interface options; determine a spatial location of the virtual menu based on characteristics of the physical surface; and generate display instructions for presentation of the virtual menu at the spatial location within the physical environment of the user, thereby generating the instructions for presenting the subset of user interface options to the user.
15. The hardware processor further comprises:
10. The system of claim 8, programmed to determine the set of available user interface options based on user information including one or more of: user age, user mood, user preferences, information associated with the user, physiological parameters of the user, psychological state of the user, emotional state of the user, mood of a user group, or user characteristics.
Citation Information
Patent Citations
Image processing method and image processing device
JP2006126936A
Image projection device and its control method
JP2009295031A
Projector and display device
JP2013152711A
Electronic device, method of controlling display of display screen in electronic device, and program
JP2016031650A
Planar waveguide apparatus with diffraction element(s) and system employing same
US20150016777A1