Context-aware user interface menus
The wearable system addresses the challenges of presenting VR, AR, and MR experiences by using a posture sensor and hardware processor to generate a virtual menu of relevant objects based on user interaction and context, enhancing the user experience through reduced clutter and improved interaction.
Patent Information
- Application Number
- JP2021119483
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-08-29
- Filing Date
- 2021-07-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2037-05-18
AI Technical Summary
Modern computing and display technologies face challenges in creating virtual reality (VR), augmented reality (AR), and mixed reality (MR) experiences that are comfortable, natural-feeling, and rich, due to the complexity of human visual perception and the need to present relevant virtual objects in a cluttered environment.
A wearable system that includes an augmented reality display, a posture sensor to analyze user position and orientation, and a hardware processor to identify physical objects, determine context information, and generate a virtual menu of relevant virtual objects based on user interaction and environmental context.
The wearable system effectively presents a relevant and user-friendly subset of virtual objects, enhancing the VR, AR, and MR experience by reducing clutter and improving interaction with the virtual environment.
Smart Images

Figure 0007682048000001 
Figure 0007682048000002 
Figure 0007682048000003
Abstract
Description
[Technical field]
[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 62 / 339,572, filed May 20, 2016, and entitled “CONTEXTUAL AWARENESS OF USER INTERFACE MENUS,” and U.S. Provisional Application No. 62 / 380,869, filed August 29, 2016, and entitled “AUGMENTED COGNITION USER INTERFACE,” the disclosures of which are incorporated by reference in their entireties herein.
[0002] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly to presenting and selecting virtual objects based on contextual information. [Background technology]
[0003] Modern computing and display technologies have facilitated the development of systems for so-called "virtual reality", "augmented reality", or "mixed reality" experiences, in which digitally reproduced images or portions thereof are presented to a user in a manner that appears or may be perceived as real. Virtual reality or "VR" scenarios typically involve the presentation of digital or virtual image information without transparency to other actual real-world visual inputs. Augmented reality or "AR" scenarios typically involve the presentation of digital or virtual image information as an extension to the visualization of the real world around the user. Mixed reality or "MR" relates to the merging of real and virtual worlds to generate new environments in which physical and virtual objects coexist and interact in real time. In conclusion, the human visual perception system is highly complex, making it difficult to create VR, AR, or MR technologies that facilitate comfortable, natural-feeling, and rich presentation of virtual image elements among other virtual or real-world image elements. The systems and methods disclosed herein address various challenges associated with VR, AR, and MR technologies. Summary of the Invention [Means for solving the problem]
[0004] In one embodiment, a wearable system for generating virtual content within a user's three-dimensional (3D) environment is disclosed. The wearable system may include an augmented reality display configured to present the virtual content to a user in a 3D view, a posture sensor configured to obtain user position or orientation data, analyze the position or orientation data, and identify a posture of the user, and a hardware processor in communication with the posture sensor and the display. The hardware processor may be programmed to: identify a physical object in the user's environment within the 3D environment based at least in part on the user's posture, receive an indication to initiate an interaction with the physical object, identify a set of virtual objects in the user's environment associated with the physical object, determine context information associated with the physical object, filter the set of virtual objects, identify a subset of the virtual objects from the set of virtual objects based on the context information, generate a virtual menu including the subset of the virtual objects, determine a spatial location within the 3D environment for presenting the virtual menu based at least in part on the determined context information, and present the virtual menu at the spatial location via the augmented reality display.
[0005] In another embodiment, a method for generating virtual content within a user's three-dimensional (3D) environment is disclosed. The method can include analyzing data obtained from a posture sensor to identify a posture of the user, identifying an interactable object within the user's 3D environment based at least in part on the posture, receiving an indication to initiate an interaction with the interactable object, determining context information associated with the interactable object, selecting a subset of user interface actions from a set of user interface actions available on the interactable object based on the context information, and generating instructions for presenting the subset of user interface actions to the user in a 3D view.
[0006] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages will be apparent from the description, drawings, and claims. Neither this summary nor the following detailed description purport to define or limit the scope of the inventive subject matter. The present specification also provides, for example, the following items: (Item 1) 1. A method for generating virtual content within a user's three-dimensional (3D) environment, the method comprising: analyzing data obtained from the posture sensor to identify a posture of the user; Identifying walls within the user's 3D environment based at least in part on the pose; and receiving an indication to initiate an interaction with the wall; determining context information associated with the wall; selecting a subset of user interface actions from a set of user interface actions available on the wall based on the context information; generating instructions for presenting the subset of user interface actions to the user in a 3D view; A method comprising: (Item 2) 2. The method of claim 1, wherein the pose includes at least one of eye gaze, head pose, or gesture. (Item 3) 2. The method of claim 1, wherein identifying the walls includes performing a cone projection based on a head pose of the user. (Item 4) 2. The method of claim 1, wherein the contextual information includes an orientation of a surface of the wall, and selecting the subset of user interface actions includes identifying user interface actions that may be performed on a surface having the orientation. (Item 5) 2. The method of claim 1, wherein the indication to initiate an interaction with the wall includes an actuation of a user input device or a change in the user's posture. (Item 6) 2. The method of claim 1, further comprising receiving physiological parameters of the user and determining a psychological state of the user, the psychological state being part of the context information for selecting the subset of user interactions. (Item 7) 7. The method of claim 6, wherein the physiological parameter relates to at least one of heart rate, pupil dilation, galvanic skin response, blood pressure, electroencephalogram status, respiratory rate, or eye movement. (Item 8) 2. The method of claim 1, wherein generating instructions for presenting a subset of the user interactions to the user in a 3D view includes generating a virtual menu including the subset of user interface actions, determining a spatial location of the virtual menu based on characteristics of the wall, and generating display instructions for presentation of the virtual menu at the spatial location within the user's 3D environment. [Brief description of the drawings]
[0007] [Figure 1] FIG. 1 depicts an illustration of a mixed reality scenario involving a virtual reality object and a physical object viewed by a person. [Diagram 2] FIG. 2 illustrates diagrammatically an embodiment of a wearable system. [Diagram 3] FIG. 3 diagrammatically illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes. [Figure 4] FIG. 4 diagrammatically illustrates an embodiment of a waveguide stack for outputting image information to a user. [Diagram 5] FIG. 5 shows an exemplary output beam that may be output by a waveguide. [Figure 6]FIG. 6 is a schematic diagram showing an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem for use in generating a multifocal volumetric display, image, or light field. [Figure 7] FIG. 7 is a block diagram of an embodiment of a wearable system. [Figure 8] FIG. 8 is a process flow diagram of an embodiment of a method for rendering virtual content in relation to recognized objects. [Figure 9] FIG. 9 is a block diagram of another embodiment of a wearable system. [Figure 10] FIG. 10 is a process flow diagram of an example method for determining user input to a wearable system. [Figure 11] FIG. 11 is a process flow diagram of an embodiment of a method for interacting with a virtual user interface. [Figure 12] FIG. 12 illustrates an example of user interaction with a virtual user interface in an office environment. [Figure 13] 13 and 14 illustrate an example of user interaction with a virtual user interface within a living room environment. [Figure 14] 13 and 14 illustrate an example of user interaction with a virtual user interface within a living room environment. [Figure 15] FIG. 15 illustrates an example of user interaction with a virtual user interface within a bedroom environment. [Figure 16] FIG. 16 is a flow chart of an example method for generating a virtual menu based on context information. [Figure 17] FIG. 17 is a flowchart of an example method for selecting virtual content based, at least in part, on a physiological and / or psychological state of a user. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure. Additionally, the figures in this disclosure are for illustrative purposes and are not to scale. (overview)
[0009] Modern computer interfaces support a wide range of functionality. However, users can be overwhelmed by the number of options and cannot quickly identify objects of interest. In an AR / VR / MR environment, the user's field of view (FOV) as perceived through the AR / VR / MR display of a wearable device may be smaller than the user's natural FOV, making it more difficult to provide a relevant set of virtual objects than in a typical computing environment.
[0010] The wearable system described herein can alleviate this problem by analyzing the user's environment and providing a smaller and more relevant subset of functionality on the user interface. The wearable system may provide this subset of functionality based on contextual information, such as the user's environment or objects in the user's environment. The wearable system can recognize physical objects (such as tables and walls) and their relationship to the environment. For example, the wearable system can recognize that a cup should be placed on a table (instead of a wall) and a painting should be placed on a vertical wall (instead of a table). Based on this relationship, the wearable system can project a virtual cup onto a table in the user's room and a virtual painting onto the vertical wall.
[0011] In addition to the relationship between the object and its environment, other factors such as the orientation of the object (e.g., horizontal or vertical), the nature of the user's environment, and previous usage patterns (time, location, etc.) can also be used to determine the virtual object to be shown by the wearable system. The nature of the user's environment can include whether it is a private environment (e.g., in the user's home or office) where the user may interact with the wearable device in a relatively secure and private manner, or a public environment where there may be others in the vicinity (where the user may not want the user's interaction with the device to be seen or heard). The distinction between private or public environments is not exclusive. For example, a park is a public environment if a large number of people are in the vicinity of the user, but may be a private environment if the user is alone or no others are in the vicinity. The distinction may be made based at least in part on the number of people in the vicinity, their proximity to the user, their relationship to the user (e.g., whether they are friends, family, or strangers), etc. Additionally or alternatively, the wearable system may identify a label (such as a fiducial mark) associated with the object. The label may contain information about a virtual object (such as an item on a virtual menu) that should be displayed for the object associated with the reference marker.
[0012] In some embodiments, the wearable system may also include various physiological sensors. These sensors may measure or estimate the user's physiological parameters, such as heart rate, respiration rate, galvanic skin response, blood pressure, brainwave state, etc. These sensors may be used in conjunction with an inwardly facing imaging system to determine the user's eye movement and pupil dilation, which may also be reflective of the user's physiological or psychological state. Data obtained by the physiological sensors or the inwardly facing imaging system may be analyzed by the wearable system to determine the user's psychological state, such as mood and interests. The wearable system may use the user's psychological state as part of the contextual information to present a set of virtual objects based, at least in part, on the user's psychological state. (Example of a 3D display for a wearable system)
[0013] A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present 2D or 3D virtual images to a user. The images may be still images, frames of video, or videos, in combination or the like. A wearable system can include a wearable device that may present a VR, AR, or MR environment, alone or in combination, for user interaction. The wearable device can be a head-mounted device (HMD), which is used synonymously with an AR device (ARD). The wearable device may be in the form of a helmet, glasses, a headset, or any other wearable configuration.
[0014] Figure 1 depicts an illustration of a mixed reality scenario with certain virtual reality objects and certain physical objects viewed by a person. In Figure 1, an MR scene 100 is depicted in which a user of the MR technology sees a real-world park-like setting 110 featuring people, trees, buildings in the background, and a concrete platform 120. In addition to these items, the user of the MR technology also perceives that they "see" a robotic figure 130 standing on the real-world platform 120, and a flying cartoon-like avatar character 140 that appears to be an anthropomorphic bumblebee, although these elements do not exist in the real world.
[0015] In order for a 3D display to produce a true depth sensation, and more specifically a simulated sensation of surface depth, it may be desirable to generate, for each point in the display's field of view, an accommodation response that corresponds to that point's virtual depth. If the accommodation response to a display point does not correspond to that point's virtual depth as determined by the binocular depth cues of convergence and stereopsis, the human eye may experience accommodation conflicts, resulting in unstable imaging, deleterious eye strain, headaches, and, in the absence of accommodative information, a near-complete lack of surface depth.
[0016] VR, AR, and MR experiences can be provided by a display system having a display in which images corresponding to multiple depth planes are provided to the viewer. The images may be different for each depth plane (e.g., providing slightly different presentations of a scene or object) and may be focused separately by the viewer's eyes, thereby serving to provide depth cues to the user based on the ocular accommodation required to focus on different image features of a scene located on different depth planes, or based on observing different image features on different depth planes that are out of focus. As discussed elsewhere herein, such depth cues provide a believable perception of depth.
[0017] FIG. 2 illustrates an example of a wearable system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems to support the functionality of the display 220. The display 220 may be coupled to a frame 230 that is wearable by a user, wearer, or viewer 210. The display 220 can be positioned in front of the eyes of the user 210. The display 220 can present AR / VR / MR content to the user. The display 220 can include a head mounted display (HMD) that is worn on the user's head. In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the user's ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable sound control).
[0018] The wearable system 200 may include an outwardly facing imaging system 464 (shown in FIG. 4) that observes the world in the user's surrounding environment. The wearable system 200 may also include an inwardly facing imaging system 462 (shown in FIG. 4) that may track the user's eye movements. The inwardly facing imaging system may track either one eye's movements or both eyes' movements. The inwardly facing imaging system 462 may be mounted to the frame 230 and may be in electrical communication with a processing module 260 or 270 that may process image information acquired by the inwardly facing imaging system and determine, for example, pupil diameter or orientation of the user's 210 eyes, eye movements, or eye pose.
[0019] As an example, the wearable system 200 can capture an image of the user's posture using the outward facing imaging system 464 or the inward facing imaging system 462. The image may be a still image, a frame of video, or a video, a combination thereof, or the like.
[0020] In some embodiments, the wearable system 200 can include one or more physiological sensors 232. Examples of such sensors include sensors configured for ophthalmic testing, such as confocal microscopy sensors, electronystagmography (ENG) sensors, electrooculography (EOG) sensors, electroretinogram (ERG) sensors, laser Doppler flowmeter (LDF) sensors, photoacoustic imaging and pressure reading sensors, two-photon excitation microscopy sensors, and / or ultrasound sensors. Other examples of sensors include sensors configured for other electrodiagnostic techniques, such as electrocardiogram (ECG) sensors, electroencephalogram (EEG) sensors, electromyography (EMG) sensors, electrophysiology testing (EP) sensors, event-related potential (ERP) sensors, near-infrared neuroimaging (fNIR) sensors, low-resolution electroencephalography (LORETA) sensors, and / or optical coherence tomography (OCT) sensors. Further examples of the sensor 232 include physiological sensors such as a blood glucose sensor, a blood pressure sensor, a skin potential sensor, a photoplethysmography device, a sensing device for computer-aided auscultation, a galvanic skin response sensor, and / or a temperature sensor. The sensor 232 may also include a CO 2 Monitoring sensor, respiratory rate sensor, end-tidal CO 2 It may include a sensor and / or a breathalyzer.
[0021] An example of a sensor 232 is shown diagrammatically as being connected to the frame 230. This connection may take the form of a physical attachment to the frame 230 and may be anywhere on the frame 230. As an example, the sensor 232 may be mounted on the frame 230 so as to be positioned adjacent the user's temple or at the contact point between the frame 230 and the user's nose. As another example, the sensor 232 may be positioned on a portion of the frame 230 that extends over the user's ear. In some other embodiments, the sensor 232 may extend from the frame 230 and contact the user 210. For example, the sensor 232 may touch a part of the user's body (such as the user's arm) and connect to the frame 230 via a wired connection. In other embodiments, the sensor 232 may not be physically attached to the frame 230. Rather, the sensor 232 may communicate with the wearable system 200 via a wireless connection. In some embodiments, the wearable system 200 may have the form of a helmet, and the sensor 232 may be positioned towards the top or sides of the user's head.
[0022] In some implementations, the sensor 232 makes a direct measurement of a physiological parameter that is used as context information by the wearable system. For example, a heart rate sensor may directly measure the user's heart rate. In other implementations, the sensor 232 (or a group of sensors) may make a measurement that is used to estimate another physiological parameter. For example, stress may be estimated as a combination of heart rate measurements and galvanic skin response measurements. Statistical techniques can be applied to the sensor data to estimate a physiological (or psychological) state. As an example, the sensor data can be combined using machine learning techniques (e.g., decision trees, neural networks, support vector machines, Bayesian techniques) to estimate the user's state. The state estimate may provide a binary state (e.g., stress or baseline), multiple states (e.g., stress, baseline, or relaxed), or a probabilistic measure (e.g., the probability that the user is under stress). The physiological or psychological state may reflect any emotional state of the user, such as, for example, anxiety, stress, anger, love, boredom, despair or disappointment, happiness, sadness, loneliness, shock, or surprise.
[0023] The display 220 is operably coupled (250) to a local data processing module 260, which may be mounted in a variety of configurations, such as fixedly attached to the frame 230, such as by wired or wireless connection, fixedly attached to a helmet or hat worn by the user, integrated into headphones, or otherwise removably attached to the user 210 (e.g., in a backpack configuration, in a belt-coupled configuration).
[0024] The local processing and data module 260 may comprise a hardware processor and digital memory such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storage of data. The data may include a) data captured from sensors (e.g., which may be operatively coupled to the frame 230 or otherwise attached to the user 210), such as an image capture device (e.g., a camera in an inwardly facing imaging system or an outwardly facing imaging system), a microphone, an inertial measurement unit (IMU), an accelerometer, a compass, a global positioning system (GPS) unit, a wireless device, or a gyroscope, or b) data acquired or processed using the remote processing module 270 or a remote data repository 280, possibly for transmission to the display 220 after processing or retrieval. The local processing and data module 260 may be operatively coupled to a remote processing module 270 or a remote data repository 280 by a communication link 262 or 264, such as via a wired or wireless communication link, such that these remote modules are available as resources to the local processing and data module 260. In addition, the remote processing module 280 and the remote data repository 280 may be operatively coupled to each other.
[0025] In some embodiments, remote processing module 270 may comprise one or more hardware processors configured to analyze and process data and / or image information. In some embodiments, remote data repository 280 may comprise a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, allowing for fully autonomous use from the remote module.
[0026] The human visual system is complex and difficult to provide a realistic perception of depth. Without being limited by theory, it is believed that a viewer of an object may perceive the object as three-dimensional due to a combination of vergence and accommodation. The vergence of the two eyes relative to one another (i.e., the rolling of the pupils toward or away from one another to converge the gaze of the eyes and fixate on an object) is closely coupled to the focusing (or "accommodation") of the eye's lenses. Under normal conditions, changing the focus of the eye's lenses, or accommodating the eyes to change focus from one object to another at a different distance, will automatically produce a matching change in vergence to the same distance, in a relationship known as the "accommodation-vergence reflex." Similarly, a change in vergence will trigger a matching change in accommodation under normal conditions. Display systems that provide better matching between accommodation and convergence-divergence movements may produce more realistic and comfortable simulations of three-dimensional images.
[0027] FIG. 3 illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes. With reference to FIG. 3, objects at various distances from the eyes 302 and 304 on the z-axis are accommodated by the eyes 302 and 304 such that the objects are in focus. The eyes 302 and 304 assume particular accommodated states that focus objects at different distances along the z-axis. As a result, a particular accommodated state may be said to be associated with a particular one of the depth planes 306 with an associated focal length such that an object or part of an object at a particular depth plane is in focus when the eye is in an accommodated state relative to that depth plane. In some embodiments, a three-dimensional image may be simulated by providing a different presentation of an image to each eye 302 and 304, and by providing a different presentation of an image corresponding to each of the depth planes. Although shown as separate for clarity of illustration, it should be understood that the fields of view of the eyes 302 and 304 may overlap, for example, as the distance along the z-axis increases. In addition, while shown to be flat for ease of illustration, it should be understood that the contours of the depth plane may be curved in physical space such that all features within the depth plane are in focus with the eye in a particular accommodation state. Without being limited by theory, it is believed that the human eye may typically interpret a finite number of depth planes to provide depth perception. As a result, a highly realistic simulation of perceived depth may be achieved by providing the eye with a different representation of an image corresponding to each of these limited number of depth planes. (Waveguide Stack Assembly)
[0028] FIG. 4 illustrates an example of a waveguide stack for outputting image information to a user. The wearable system 400 includes a stack of waveguides or a stacked waveguide assembly 480 that can be utilized to provide a three-dimensional perception to the eye / brain using multiple waveguides 432b, 434b, 436b, 438b, 4400b. In some embodiments, the wearable system 400 may correspond to the wearable system 200 of FIG. 2, and FIG. 4 diagrammatically illustrates some portions of the wearable system 200 in more detail. For example, in some embodiments, the waveguide assembly 480 may be integrated into the display 220 of FIG. 2.
[0029] 4, the waveguide assembly 480 may also include a number of features 458, 456, 454, 452 between the waveguides. In some embodiments, the features 458, 456, 454, 452 may be lenses. In other embodiments, the features 458, 456, 454, 452 may not be lenses. Rather, they may simply be spacers (e.g., cladding layers or structures to form air gaps).
[0030] The waveguides 432b, 434b, 436b, 438b, 440b or multiple lenses 458, 456, 454, 452 may be configured to transmit image information to the eye with various levels of wavefront curvature or light beam divergence. Each waveguide level may be associated with a particular depth plane and configured to output image information corresponding to that depth plane. Image injection devices 420, 422, 424, 426, 428 may be utilized to inject image information into the waveguides 440b, 438b, 436b, 434b, 432b, respectively, which may be configured to disperse incident light across each individual waveguide for output toward the eye 410. Light exits output surfaces of image launch devices 420, 422, 424, 426, 428 and is launched into corresponding input edges of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, a single beam of light (e.g., a collimated beam) may be launched into each waveguide to output an entire field of cloned collimated beams that are directed towards the eye 410 at a particular angle (and divergence) that corresponds to the depth plane associated with the particular waveguide.
[0031] In some embodiments, the image input devices 420, 422, 424, 426, 428 are each discrete displays that generate image information for input into a respective waveguide 440b, 438b, 436b, 434b, 432b, etc. In some other embodiments, the image input devices 420, 422, 424, 426, 428 are the output of a single multiplexed display that may, for example, send image information to each of the image input devices 420, 422, 424, 426, 428 via one or more optical conduits (such as fiber optic cables).
[0032] A controller 460 controls the operation of the stacked waveguide assembly 480 and the image input devices 420, 422, 424, 426, 428. The controller 460 includes programming (e.g., instructions in a non-transitory computer-readable medium) that coordinates the timing and provision of image information to the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the controller 460 may be a single integrated device or a distributed system connected by a wired or wireless communication channel. The controller 460 may be part of the processing module 260 or 270 (illustrated in FIG. 2) in some embodiments.
[0033] The waveguides 440b, 438b, 436b, 434b, 432b may be configured to propagate light within each respective waveguide by total internal reflection (TIR). The waveguides 440b, 438b, 436b, 434b, 432b may each be planar or have another shape (e.g., curved) with major top and bottom surfaces and edges extending between the major top and bottom surfaces. In the illustrated configuration, the waveguides 440b, 438b, 436b, 434b, 432b may each include a light extraction optical element 440a, 438a, 436a, 434a, 432a configured to extract light from the waveguide by redirecting the light to propagate within each respective waveguide and outputting image information from the waveguide to the eye 410. The extracted light may also be referred to as out-coupled light, and the light extraction optical element may also be referred to as out-coupling optical element. The extracted light beam is output by the waveguide to where the light propagating in the waveguide strikes the light redirecting element. The light extraction optical element (440a, 438a, 436a, 434a, 432a) may be, for example, a reflective or diffractive optical feature. Although shown disposed on the bottom major surface of the waveguides 440b, 438b, 436b, 434b, 432b for ease of explanation and clarity of drawing, in some embodiments the light extraction optical element 440a, 438a, 436a, 434a, 432a may be disposed on the top or bottom major surface, or may be disposed directly within the volume of the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed in a layer of material that is attached to a transparent substrate and forms the waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, the waveguides 440b, 438b, 436b, 434b, 432b may be a monolithic piece of material and the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed on and / or within that piece of material.
[0034] Continuing with reference to FIG. 4, as discussed herein, each waveguide 440b, 438b, 436b, 434b, 432b is configured to output light and form an image corresponding to a particular depth plane. For example, the waveguide 432b closest to the eye may be configured to deliver collimated light to the eye 410 as it is launched into such waveguide 432b. The collimated light may represent an optical infinity focal plane. The next upper waveguide 434b may be configured to send collimated light that passes through a first lens 452 (e.g., a negative lens) before it can reach the eye 410. The first lens 452 may be configured to generate a slight convex wavefront curvature so that the eye / brain interprets the light originating from the next upper waveguide 434b as originating from a first focal plane closer inward from optical infinity toward the eye 410. Similarly, the third upper waveguide 436b passes its output light through both a first lens 452 and a second lens 454 before reaching the eye 410. The combined refractive powers of the first and second lenses 452 and 454 may be configured to produce another incremental amount of wavefront curvature such that the eye / brain interprets the light emerging from the third upper waveguide 436b as originating from a second focal plane that is closer inward from optical infinity towards the person than was the light from the next upper waveguide 434b.
[0035] Other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are similarly configured, with the highest waveguide 440b in the stack sending its output through all of the lenses between it and the eye for an aggregate focal force that represents the focal plane closest to the person. To compensate for the stack of lenses 458, 456, 454, 452 when viewing / interpreting light originating from the world 470 on the other side of the stacked waveguide assembly 480, a compensating lens layer 430 may be placed on top of the stack to compensate for the aggregate force of the lens stacks 458, 456, 454, 452 below. Such a configuration provides as many perceived focal planes as there are waveguide / lens pairs available. Both the light extraction optical elements of the waveguides and the focusing sides of the lenses may be static (e.g., not dynamic or electroactive). In some alternative embodiments, either or both may be dynamic using electroactive features.
[0036] Continuing with reference to FIG. 4, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be configured to both redirect light from its respective waveguide and output the light with an appropriate amount of divergence or collimation for a particular depth plane associated with the waveguide. As a result, waveguides with different associated depth planes may have different configurations of light extraction optical elements that output light with different amounts of divergence depending on the associated depth plane. In some embodiments, as discussed herein, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume or surface features that may be configured to output light at specific angles. For example, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume holograms, surface holograms, and / or diffraction gratings. Light extraction optical elements such as diffraction gratings are described in U.S. Patent Publication No. 2015 / 0178939, published June 25, 2015, which is incorporated by reference in its entirety.
[0037] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a are diffractive features that form a diffraction pattern, i.e., a "diffractive optical element" (also referred to herein as a "DOE"). Preferably, the DOE has a relatively low diffraction efficiency such that only a portion of the light of the beam is deflected towards the eye 410 at each intersection of the DOE, while the remainder continues to travel through the waveguide via total internal reflection. The light carrying the image information is thus split into a plurality of associated output beams that exit the waveguide at multiple locations, and the result can be a very uniform pattern of output emission towards the eye 304 with respect to this particular collimated beam that bounces within the waveguide.
[0038] In some embodiments, one or more DOEs may be switchable between an "on" state in which they actively diffract and an "off" state in which they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer dispersed liquid crystal in which the microdroplets have a diffraction pattern in a host medium, and the refractive index of the microdroplets may be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract the incident light), or the microdroplets may be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts the incident light).
[0039] In some embodiments, the number and distribution of depth planes or depths of field may be dynamically varied based on the size or orientation of the pupil of the viewer's eye. The depth of field may vary inversely with the size of the viewer's pupil. As a result, as the size of the viewer's eye pupil decreases, the depth of field increases so that one plane that is indistinguishable because its location exceeds the focal depth of the eye may become distinguishable and appear more focused with a reduction in pupil size and a corresponding increase in depth of field. Similarly, the number of spaced depth planes used to present different images to the viewer may be reduced with a reduced pupil size. For example, the viewer may not be able to clearly perceive the details of both the first and second depth planes at one pupil size without adjusting the accommodation of the eye from one depth plane to the other. However, these two depth planes may be sufficient to simultaneously focus the user at another pupil size without changing accommodation.
[0040] In some embodiments, the display system may vary the number of waveguides that receive image information based on a determination of pupil size and / or orientation, or in response to receiving an electrical signal indicative of a particular pupil size and / or orientation. For example, if the user's eye is unable to distinguish between two depth planes associated with two waveguides, the controller 460 may be configured or programmed to stop providing image information to one of those waveguides. Advantageously, this may reduce the processing burden on the system, thereby increasing the responsiveness of the system. In embodiments in which the DOE for a waveguide is switchable between on and off states, the DOE may be switched to the off state when the waveguide receives image information.
[0041] In some embodiments, it may be desirable to have the exit beam meet the condition of having a diameter less than the diameter of the viewer's eye. However, meeting this condition may be difficult in light of the variability in the size of the viewer's pupil. In some embodiments, this condition is met over a wide range of pupil sizes by varying the size of the exit beam in response to a determination of the size of the viewer's pupil. For example, as the pupil size decreases, the size of the exit beam may also decrease. In some embodiments, the exit beam size may be varied using a variable aperture.
[0042] The wearable system 400 may include an outwardly facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 may be referred to as the field of view (FOV) of the world camera, and the imaging system 464 is sometimes also referred to as the FOV camera. The entire area available for viewing or imaging by a viewer may be referred to as the ocular field of view (FOR). The FOR may include a solid angle of 4π steradians surrounding the wearable system 400 as the wearer moves their body, head, or eyes to perceive virtually any direction in space. In other situations, the wearer's movements may be more constrained, and accordingly, the wearer's FOR may subtend a smaller solid angle. Images obtained from the outwardly facing imaging system 464 may be used to track gestures (e.g., hand or finger gestures) made by the user, detect objects in the world 470 in front of the user, etc.
[0043] The wearable system 400 may also include an inwardly facing imaging system 466 (e.g., a digital camera) that observes user movements, such as eye and face movements. The inwardly facing imaging system 466 may be used to capture images of the eye 410 and determine the size and / or orientation of the pupil of the eye 304. The inwardly facing imaging system 466 may be used to obtain images for use in determining the direction the user is looking (e.g., eye pose) or for biometric identification of the user (e.g., via iris identification). In some embodiments, at least one camera may be utilized for each eye independently to separately determine the pupil size or eye pose of each eye, thereby allowing the presentation of image information to each eye to be dynamically adjusted for that eye. In some other embodiments, the pupil diameter or orientation of only a single eye 410 (e.g., using only a single camera per pair of eyes) is determined and assumed to be similar for both eyes of the user. Images obtained by the inward-facing imaging system 466 may be analyzed to determine the user's eye posture or mood, which may be used by the wearable system 400 to determine audio or visual content to be presented to the user. The wearable system 400 may also determine head pose (e.g., head position or head orientation) using sensors such as an IMU, accelerometer, gyroscope, etc.
[0044] The wearable system 400 may include a user input device 466 through which a user may input commands into the controller 460 and interact with the wearable system 400. For example, the user input device 466 may include a trackpad, a touch screen, a joystick, a multi-degree-of-freedom (DOF) controller, a capacitive sensing device, a game controller, a keyboard, a mouse, a directional pad (D-pad), a wand, a tactile device, a totem (e.g., serving as a virtual user input device), and the like. A multi-DOF controller may sense user input in possible translation (e.g., left / right, forward / backward, or up / down) or rotation (e.g., yaw, pitch, or roll) of some or all of the controller. A multi-DOF controller that supports translation may be referred to as 3DOF, while a multi-DOF controller that supports translation and rotation may be referred to as 6DOF. In some cases, a user may use a finger (e.g., thumb) to press or swipe on a touch-sensitive input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input device 466 may be held by the user's hand during use of the wearable system 400. The user input device 466 may communicate with the wearable system 400 in wired or wireless communication.
[0045] The wearable system 400 may also include a physiological sensor 468 (which may be an exemplary embodiment of the sensor 232 in FIG. 2 ) configured to measure physiological parameters of the user, such as heart rate, galvanic skin response, respiration rate, etc. The physiological sensor may communicate acquired data to the controller 460. The controller 460 may use the data acquired by the physiological sensor, alone or in combination with data acquired by other sensors, to determine the physiological and / or psychological state of the user. For example, the controller 460 may combine heart rate data acquired by the physiological sensor 468 with pupillary information acquired by the inwardly facing imaging system 462 to determine whether the user is amused or angry. As described further below, the wearable system may selectively present virtual content to the user based on the physiological and / or psychological state of the user.
[0046] 5 shows an example of an output beam output by a waveguide. Although one waveguide is shown, it should be understood that other waveguides in the waveguide assembly 480 may function similarly and that the waveguide assembly 480 includes multiple waveguides. Light 520 is launched into the waveguide 432b at the input edge 432c of the waveguide 432b and propagates within the waveguide 432b by TIR. At the point where the light 520 strikes the DOE 432a, a portion of the light exits the waveguide as an output beam 510. Although the output beams 510 are shown as approximately parallel, they may also be redirected to propagate to the eye 410 at an angle (e.g., divergent output beam formation) depending on the depth plane associated with the waveguide 432b. It should be understood that a nearly collimated exit beam may refer to a waveguide with light extraction optics that outcouples light to form an image that appears to be set at a depth plane at a large distance (e.g., optical infinity) from the eye 410. Other waveguides or other sets of light extraction optics may output a more divergent exit beam pattern, which would require the eye 410 to accommodate to a closer distance and focus on the retina, and would be interpreted by the brain as light from a distance closer to the eye 410 than optical infinity.
[0047] FIG. 6 is a schematic diagram showing an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal volumetric display, image, or light field. The optical system can include a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem. The optical system can be used to generate a multifocal volumetric display, image, or light field. The optical system can include one or more primary planar waveguides 632a (only one is shown in FIG. 6) and one or more DOEs 632b associated with each of at least some of the primary waveguides 632a. The planar waveguides 632b can be similar to the waveguides 432b, 434b, 436b, 438b, 440b discussed with reference to FIG. 4. The optical system may employ a dispersive waveguide device to relay light along a first axis (vertical or Y-axis in the illustration of FIG. 6) and expand the effective exit pupil of the light along the first axis (e.g., Y-axis). The dispersive waveguide device may include, for example, a dispersive planar waveguide 622b and at least one DOE 622a (illustrated by a double dashed line) associated with the dispersive planar waveguide 622b. The dispersive planar waveguide 622b may be similar or the same in at least some respects as the primary planar waveguide 632b, which has a different orientation therefrom. Similarly, the at least one DOE 622a may be similar or the same in at least some respects as the DOE 632a. For example, the dispersive planar waveguide 622b or the DOE 622a may be made of the same material as the primary planar waveguide 632b or the DOE 632a, respectively. The embodiment of the optical display system 600 shown in FIG. 6 can be integrated into the wearable system 200 shown in FIG.
[0048] The relayed and exit pupil expanded light may be optically coupled from the dispersive waveguide device into one or more primary planar waveguides 632b. The primary planar waveguides 632b may relay the light along a second axis (e.g., the horizontal or X-axis in the diagram of FIG. 6) that is preferably orthogonal to the first axis. It should be noted that the second axis may be a non-orthogonal axis to the first axis. The primary planar waveguides 632b may expand the effective exit pupil of the light along its second axis (e.g., the X-axis). For example, the dispersive planar waveguide 622b may relay and expand the light along the vertical or Y-axis and pass the light to the primary planar waveguide 632b, which may relay and expand the light along the horizontal or X-axis.
[0049] The optical system may include one or more colored light sources (e.g., red, green, and blue laser light) 610, which may be optically coupled into a proximal end of a single mode optical fiber 640. The distal end of the optical fiber 640 may be threaded or received through a hollow tube 642 of piezoelectric material. The distal end protrudes from the tube 642 as an unfixed flexible cantilever 644. The piezoelectric tube 642 may be associated with four quadrant electrodes (not shown). The electrodes may be, for example, plated on the outside, outer surface or outer periphery, or diameter of the tube 642. A core electrode (not shown) may also be located in the core, center, inner periphery, or inner diameter of the tube 642.
[0050] For example, drive electronics 650, electrically coupled via wires 660, drive opposing pairs of electrodes to bend the piezoelectric tube 642 independently in two axes. The protruding distal tip of the optical fiber 644 has a mechanical resonant mode. The frequency of the resonance may depend on the diameter, length, and material properties of the optical fiber 644. By oscillating the piezoelectric tube 642 near the first mechanical resonant mode of the fiber cantilever 644, the fiber cantilever 644 may be caused to oscillate and sweep through a large deflection.
[0051] By stimulating resonant vibrations in two axes, the tip of the fiber cantilever 644 is scanned in two axial directions within an area filling a two-dimensional (2-D) scan. By modulating the intensity of the light source 610 synchronously with the scanning of the fiber cantilever 644, the light emitted from the fiber cantilever 644 can form an image. A description of such a setup is provided in U.S. Patent Publication No. 2014 / 0003762, which is incorporated herein by reference in its entirety.
[0052] The optical coupler subsystem components can collimate the light emitted from the scanning fiber cantilever 644. The collimated light can be reflected by a mirrored surface 648 into a narrow dispersive planar waveguide 622b, which contains at least one diffractive optical element (DOE) 622a. The collimated light can propagate perpendicularly (with respect to the view of FIG. 6) along the dispersive planar waveguide 622b by TIR, and in so doing repeatedly intersect with the DOE 622a. The DOE 622a preferably has a low diffraction efficiency. This causes a portion of the light (e.g., 10%) to diffract toward the edge of the larger primary planar waveguide 632b at each point of intersection with the DOE 622a, and allows a portion of the light to continue on its original trajectory down the length of the dispersive planar waveguide 622b via TIR.
[0053] At each point of intersection with DOE 622a, additional light can be diffracted towards the entrance of primary planar waveguide 632b. By splitting the incoming light into multiple outcoupled sets, the exit pupil of the light can be vertically expanded by DOE 4 in dispersive planar waveguide 622b. This vertically expanded light outcoupled from dispersive planar waveguide 622b can enter the edge of primary planar waveguide 632b.
[0054] Light entering the first waveguide 632b can propagate horizontally along the first waveguide 632b (with respect to the figure in FIG. 6) via TIR. As the light intersects the DOE 632a at multiple points, it propagates horizontally along at least a portion of the length of the first waveguide 632b via TIR. The DOE 632a preferably has a phase profile that is advantageously the sum of a linear diffraction pattern and a radially symmetric diffraction pattern and can be designed or configured to generate both deflection and focusing of the light. The DOE 632a preferably has a low diffraction efficiency (e.g., 10%) such that only a portion of the light of the beam is deflected towards the viewer's eye at each intersection of the DOE 632a, while the rest of the light continues to propagate through the first waveguide 632b via TIR.
[0055] At each point of intersection between the propagating light and the DOE 632a, a portion of the light is diffracted towards the adjacent surface of the first waveguide 632b, allowing the light to escape from TIR and be emitted from the surface of the first waveguide 632b. In some embodiments, the radially symmetric diffraction pattern of the DOE 632a additionally imparts a certain focal level to the diffracted light and both shapes (e.g., imparts curvature to) the wavefronts of the individual beams and steers the beams to an angle that matches the designed focal level.
[0056] Thus, these different paths can couple light out of the primary planar waveguide 632b by resulting in different fill patterns at the exit pupil, focus levels, and / or multiplicities of the DOE 632a at different angles. Different fill patterns at the exit pupil can be advantageously used to generate light field displays with multiple depth planes. Each layer or set of layers in a stack (e.g., three layers) in a waveguide assembly may be employed to generate individual colors (e.g., red, blue, green). Thus, for example, a first set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a first focal depth. A second set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a second focal depth. Multiple sets may be employed to generate full 3D or 4D color image light fields with various focal depths. (Other components of the wearable system)
[0057] In many implementations, the wearable system may include other components in addition to or as an alternative to the components of the wearable system described above. The wearable system may include, for example, one or more tactile devices or components. The tactile devices or components may be operable to provide a haptic sensation to the user. For example, the tactile device or component may provide a haptic sensation of pressure and / or texture upon touching the virtual content (e.g., a virtual object, virtual tool, other virtual structure). The haptic sensation may replicate the sensation of a physical object that the virtual object represents, or may replicate the sensation of an imaginary object or character (e.g., a dragon) that the virtual content represents. In some implementations, the tactile device or component may be worn by the user (e.g., a user-wearable glove). In some implementations, the tactile device or component may be held by the user.
[0058] A wearable system may include one or more physical objects that can be manipulated by, for example, a user to allow input to or interaction with the wearable system. These physical objects may be referred to herein as totems. Some totems may take the form of inanimate objects, such as, for example, pieces of metal or plastic, walls, the surface of a table, etc. In some implementations, a totem may not actually have any physical input structures (e.g., keys, triggers, joysticks, trackballs, rocker switches). Instead, a totem may simply provide a physical surface, and the wearable system may render a user interface to appear to the user as being on one or more surfaces of the totem. For example, the wearable system may render an image of a computer keyboard and trackpad to appear to reside on one or more surfaces of the totem. For example, the wearable system may render a virtual computer keyboard and virtual trackpad to appear on the surface of a thin rectangular plate of aluminum that serves as the totem. The rectangular plate itself does not have any physical keys or trackpads or sensors. However, the wearable system may detect user manipulation or interaction or touch with the rectangular plate as a selection or input made via a virtual keyboard or virtual trackpad. User input device 466 (shown in FIG. 4) may be an embodiment of the totem, which may include a trackpad, touchpad, trigger, joystick, trackball, rocker or virtual switch, mouse, keyboard, multi-degree-of-freedom controller, or another physical input device. A user may use the totem alone or in combination with posture to interact with the wearable system or other users.
[0059] Examples of tactile devices and totems usable with the wearable devices, HMDs, and display systems of the present disclosure are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated by reference in its entirety. Exemplary Wearable Systems, Environments, and Interfaces
[0060] The wearable system may employ various mapping-related techniques to achieve a high depth of field in the rendered light field. When mapping a virtual world, it is advantageous to capture all features and points in the real world and accurately depict the virtual objects in relation to the real world. To this end, FOV images captured from a user of the wearable system can be added to the world model by including new photos that convey information about various points and features of the real world. For example, the wearable system can collect a set of map points (such as 2D or 3D points), find new map points, and render a more accurate version of the world model. The world model of the first user can be communicated (e.g., via a network such as a cloud network) to a second user so that the second user can experience the world surrounding the first user.
[0061] 7 is a block diagram of an example of a MR environment 700. The MR environment 700 may be configured to receive inputs (e.g., visual inputs 702 from a user's wearable system, stationary inputs 704 such as room cameras, sensory inputs 706 from various sensors, gestures, totems, eye tracking, user inputs, etc. from user input devices 466) from one or more user wearable systems (e.g., wearable system 200 or display system 220) or stationary room systems (e.g., room cameras, etc.). The wearable systems can use various sensors (e.g., accelerometers, gyroscopes, temperature sensors, movement sensors, depth sensors, GPS sensors, inward facing imaging systems, outward facing imaging systems, etc.) to determine the location and various other attributes of the user's environment. This information may be further supplemented with information from stationary cameras in the room, which may provide images from different perspectives or various cues. Image data acquired by the cameras (e.g., room cameras and / or cameras of the outward facing imaging systems) may be reduced to a set of mapping points.
[0062] One or more object recognizers 708 can crawl through the received data (e.g., a collection of points), recognize or map the points, tag the images, and associate semantic information with the objects with the aid of a map database 710. The map database 710 may comprise various points and their corresponding objects collected over time. The various devices and the map database may be interconnected through a network (e.g., LAN, WAN, etc.) and accessible to the cloud.
[0063] Based on this information and the collection of points in the map database, the object recognizers 708a-708n may recognize objects in the environment. For example, the object recognizers may recognize faces, people, windows, walls, user input devices, televisions, other objects in the user's environment, etc. One or more object recognizers may be specialized for objects with certain characteristics. For example, object recognizer 708a may be used to recognize faces, while another object recognizer may be used to recognize totems.
[0064] Object recognition may be performed using a variety of computer vision techniques. For example, the wearable system may analyze images acquired by the outward-facing imaging system 464 (shown in FIG. 4) and perform scene reconstruction, event detection, video tracking, object recognition, object pose estimation, learning, indexing, motion estimation, image restoration, etc. One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale Invariant Feature Transform (SIFT), Speed Up Robust Features (SURF), Orientation FAST and Rotation BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retinal Keypoints (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous localization and mapping (vSLAM) techniques, sequential Bayes estimators (e.g., Kalman filter, extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), iterative nearest neighbor (ICP), semi-global matching (SGM), semi-global block matching (SGBM), feature point histograms, various machine learning algorithms (e.g., support vector machines, k-nearest neighbor algorithms, naive Bayes, neural networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), and the like.
[0065] Object recognition can additionally or alternatively be performed by various machine learning algorithms. Once trained, the machine learning algorithms can be stored by the HMD. Some examples of machine learning algorithms can include supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., a priori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural networks, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., stacked generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models can be customized for individual datasets. For example, the wearable device can generate or store a base model. The base model may be used as a starting point to generate additional models specific to a data type (e.g., a particular user in a telepresence session), a data set (e.g., a set of additional images acquired of a user in a telepresence session), a conditional situation, or other variations. In some embodiments, the wearable HMD can be configured to generate models for analysis of the aggregated data using multiple techniques. Other techniques may include using predefined thresholds or data values.
[0066] Based on this information and the set of points in the map database, the object recognizer 708a-708n may recognize objects, complement the objects with semantic information, and bring them to life. For example, if the object recognizer recognizes that a set of points is a door, the system may attach some semantic information (e.g., a door has a hinge and has 90 degree movement around the hinge). If the object recognizer recognizes that a set of points is a mirror, the system may attach the semantic information that a mirror has a reflective surface that can reflect an image of an object in a room. Over time, the map database grows as the system (which may reside locally or be accessible through a wireless network) accumulates more data from the world. Once an object is recognized, the information may be transmitted to one or more wearable systems. For example, the MR environment 700 may contain information about a scene taking place in California. The environment 700 may be transmitted to one or more users in New York. Based on the data received from the FOV camera and other inputs, an object recognizer and other software components can map points collected from various images, recognize objects, etc., so that the scene can be accurately "passed" to a second user who may be in a different part of the world. The environment 700 may also use a topology map for localization purposes.
[0067] 8 is a process flow diagram of an example embodiment of a method 800 for rendering virtual content in relation to recognized objects. Method 800 illustrates how a virtual scene may be presented to a user of a wearable system. The user may be geographically remote from the scene. For example, a user may be in New York but may wish to view a scene currently occurring in California, or may wish to go for a walk with a friend who is in California.
[0068] In block 810, the wearable system may receive input from the user and other users regarding the user's environment. This may be accomplished through various input devices and knowledge already held in a map database. The user's FOV camera, sensors, GPS, eye tracking, etc. communicate information to the system in block 810. The system may determine rough points based on this information in block 820. The rough points may be used in determining pose data (e.g., head pose, eye pose, body pose, or hand gestures) that may be used in displaying and understanding the orientation and position of various objects in the user's surroundings. The object recognizers 708a-708n may crawl through these collected points and recognize one or more objects using the map database in block 830. This information may then be communicated to the user's respective wearable system in block 840, and the desired virtual scene may be displayed to the user in block 850 as appropriate. For example, a desired virtual scene (eg, a user in CA) may be displayed in the proper orientation, position, etc., relative to various objects and other surroundings of the user in New York.
[0069] FIG. 9 is a block diagram of another example of a wearable system. In this example, the wearable system 900 comprises a map, which may include map data about the world. The map may reside in part locally on the wearable system and in part in a networked storage location (e.g., in a cloud system) accessible by a wired or wireless network. An attitude process 910 may run on the wearable computing architecture (e.g., processing module 260 or controller 460) and utilize data from the map to determine the position and orientation of the wearable computing hardware or the user. The attitude data may be calculated from data collected on the fly as the user experiences the system and operates within its world. The data may comprise images, data from sensors (such as inertial measurement units, which generally include accelerometer and gyroscope components) and surface information about objects in the real or virtual environment.
[0070] The rough point representation may be the output of a simultaneous localization and mapping (SLAM or V-SLAM, which refers to configurations where the input is image / vision only) process. The system can be configured to find what the world is made of, not just the locations of various components within the world. Poses may be building blocks that accomplish many goals, including populating maps and using data from maps.
[0071] In one embodiment, the rough point locations may not be entirely adequate by themselves, and further information may be required to generate a multifocal AR, VR, or MR experience. A dense representation, generally referring to depth map information, may be utilized to fill this gap, at least in part. Such information may be calculated from a process referred to as stereoscopic vision 940, where depth information is determined using techniques such as triangulation or time-of-flight sensing. Image information and active patterns (such as infrared patterns generated using an active projector) may serve as inputs to the stereoscopic vision process 940. A significant amount of depth map information may be fused together, some of which may be summarized using a surface representation. For example, a mathematically definable surface may be an efficient (e.g., compared to a large-scale point cloud) and applicable input to other processing devices such as a game engine. Thus, the outputs of the stereoscopic vision process (e.g., depth map) 940 may be combined in a fusion process 930. The poses may also be input to this fusion process 930, the output of which becomes the input to populate the map process 920. Sub-surfaces may connect to each other to form larger surfaces, such as in topographical mapping, and the map becomes a large-scale hybrid of points and surfaces.
[0072] Various inputs may be utilized to resolve various aspects in the mixed reality process 960. For example, in the embodiment depicted in FIG. 9, game parameters may be inputs to determine that a user of the system is playing a monster battle game with one or more monsters in various locations, whether a monster is dead or fleeing under various conditions (such as when the user shoots the monster), walls or other objects in various locations, and the like. A world map may contain information about where such objects are relative to each other, which is another useful input for mixed reality. Orientation relative to the world is an input as well, and plays an important role for nearly any interactive system.
[0073] Controls or inputs from the user are another input to the wearable system 900. As described herein, user inputs can include visual inputs, gestures, totems, audio inputs, sensory inputs (e.g., physiological data obtained by sensors 232 in FIG. 2, etc.), and the like. To move around or play a game, for example, the user may need to command the wearable system 900 as to what they want to do. There are various forms of user control that can be utilized beyond just moving around in space themselves. In one embodiment, a totem (e.g., a user input device), or an object such as a toy gun, may be held by the user and tracked by the system. The system would preferably be configured to know that the user is holding an item and understand the type of interaction the user is having with the item (e.g., if the totem or object is a gun, the system may be configured to understand not only the location and orientation, but also whether the user is clicking a trigger or other sensitive button or element, which may be equipped with sensors such as an IMU that can help determine what is happening even when such activity is not within the field of view of any of the cameras).
[0074] Hand gesture tracking or recognition may also provide input information. The wearable system 900 may be configured to track and interpret hand gestures to gesture left or right, stop, grab, hold, etc. for a button press. For example, in one configuration, a user may want to flip through email or calendar in a non-gaming environment, or to "fist bump" with another person or performer. The wearable system 900 may be configured to utilize a minimal amount of hand gestures, which may or may not be dynamic. For example, gestures may be simple static gestures, such as spreading the hand to indicate stop, thumbs up to indicate OK, thumbs down to indicate not OK, or flipping the hand left and right or up and down to indicate a directional command.
[0075] Eye tracking is another input (e.g., to track where the user is looking, control display technology, render at a specific depth or range). In one embodiment, eye vergence may be determined using triangulation, and then accommodation may be determined using a vergence / accommodation model developed for that particular person.
[0076] The totem can also be used by the user to provide input to the wearable system, which can track the movement, position, or orientation of the totem, as well as user actuations of the totem (such as pressing keys, buttons, or touch surfaces on the totem) to determine user interface interactions in the mixed reality process 960.
[0077] In some implementations, the wearable system can also use the user's physiological data in the mixed reality process 960. The physiological data may be requested by the sensors 232 (which may include physiological sensors 468). The wearable system can determine the content to present based on an analysis of the physiological data. For example, if the wearable system determines that the user is angry while playing a game (e.g., due to an increased heart rate, blood pressure changes, etc.), the wearable system can automatically reduce the difficulty level of the game to keep the user engaged in the game.
[0078] With regard to the camera system, the exemplary wearable system 900 shown in FIG. 9 may include three paired cameras, namely, a relative wide FOV or passive SLAM paired camera arranged on either side of the user's face, and a different paired camera oriented in front of the user to handle the stereoscopic imaging process 940 and also capture hand gestures and totem / object trajectories in front of the user's face. The FOV cameras and paired cameras for the stereo process 940 may be part of the outward facing imaging system 464 (shown in FIG. 4). The wearable system 900 may include an eye tracking camera (which may be part of the inward facing imaging system 462 shown in FIG. 4) oriented towards the user's eyes to triangulate eye vectors and other information. The wearable system 900 may also include one or more textured light projectors (such as infrared (IR) projectors) to inject texture into the scene.
[0079] FIG. 10 is a process flow diagram of an example of a method 1000 for determining user input to a wearable system. In this example, a user may interact with a totem. A user may have multiple totems. For example, a user may have one totem designated for social media applications, another totem for playing games, etc. In block 1010, the wearable system may detect the movement of the totem. The movement of the totem may be recognized through an outward facing system or may be detected through sensors (e.g., tactile gloves, image sensors, hand tracking devices, eye tracking cameras, head pose sensors, etc.).
[0080] Based at least in part on the detected gestures, eye poses, head poses, or inputs through the totem, the wearable system detects the position, orientation, and / or movement of the totem (or the user's eyes or head or gestures) relative to a reference frame at block 1020. The reference frame may be a set of map points based on which the wearable system translates the totem's (or user's) movements into actions or commands. At block 1030, the user's interactions with the totem are mapped. Based on the mapping of the user interactions to the reference frame 1020, the system determines the user input at block 1040.
[0081] For example, a user may move a totem or physical object back and forth, turn a virtual page, indicate moving to the next page, or move from one user interface (UI) display screen to another UI screen. As another example, a user may move their head or eyes to look at different real or virtual objects in the user's FOR. If the user's gaze at a particular real or virtual object is longer than a threshold time, that real or virtual object may be selected as the user input. In some implementations, the vergence-divergence movement of the user's eyes can be tracked and an accommodation / vergence-divergence movement model can be used to determine the accommodation state of the user's eyes, which provides information about the depth plane the user is focusing on. In some implementations, the wearable system can use ray casting techniques to determine real or virtual objects that are in line with the direction of the user's head pose or eye pose. In various implementations, ray casting techniques can include casting a thin bundle of rays with substantially no lateral width, or casting a ray with substantial lateral width (e.g., a virtual cone or truncated cone).
[0082] The user interface may be projected by a display system as described herein (such as display 220 in FIG. 2). It may also be displayed using a variety of other techniques, such as one or more projectors. A projector may project an image onto a physical object, such as a canvas or a sphere. Interactions with the user interface may be tracked using one or more cameras outside or part of the system (e.g., using inward-facing imaging system 462 or outward-facing imaging system 464).
[0083] 11 is a process flow diagram of an example method 1100 for interacting with a virtual user interface. Method 1100 may be performed by a wearable system described herein.
[0084] In block 1110, the wearable system may identify a particular UI. The type of UI may be provided by the user. The wearable system may identify that a particular UI needs to be populated based on user input (e.g., gestures, visual data, audio data, sensory data, direct commands, etc.). In block 1120, the wearable system may generate data for a virtual UI. For example, data associated with the UI's boundaries, general structure, shape, etc. may be generated. In addition, the wearable system may determine map coordinates of the user's physical location so that the wearable system may display the UI in relation to the user's physical location. For example, if the UI is body-centered, the wearable system may determine coordinates of the user's physical stance, head pose, or eye pose so that a ring UI may be displayed around the user or a planar UI may be displayed on a wall or in front of the user. If the UI is hand-centered, map coordinates of the user's hand may be determined. These map points may be derived through data received through the FOV camera, sensory input, or any other type of collected data.
[0085] In block 1130, the wearable system may send data from the cloud to the display, or data may be sent from a local database to the display component. In block 1140, a UI is displayed to the user based on the sent data. For example, a light field display may project the virtual UI into one or both of the user's eyes. Once the virtual UI is generated, the wearable system may simply wait for commands from the user and generate more virtual content on the virtual UI in block 1150. For example, the UI may be a body-centered ring around the user's body. The wearable system may then wait for a command (gesture, head or eye movement, input from a user input device, etc.) and if recognized (block 1160), the virtual content associated with the command may be displayed to the user (block 1170). As an example, the wearable system may wait for a user's hand gesture before mixing multiple stem tracks.
[0086] Additional examples of wearable systems, UI, and user experience (UX) are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated by reference in its entirety. (Example objects in the environment)
[0087] Fig. 12 illustrates an example of user interaction with a virtual user interface in an office environment. In Fig. 12, a user 210 wearing a wearable device 1270 is standing in an office 1200. The wearable device may be part of a wearable system 200, 400 as described herein. The office 1200 may include multiple physical objects such as a chair 1244, a mirror 1242, a wall 1248, a table 1246, a swivel chair 1240, etc., and a virtual screen 1250 presented to the user by the wearable device 1270. Example objects in the ocular field
[0088] A user 210 wearing a wearable device 1270 may have a field of view (FOV) and a field of view (FOR). As discussed with reference to FIG. 4, the FOR includes the portion of the user's surrounding environment that is perceptible by the user via the wearable device 1270. For an HMD, the FOR may include substantially all of the 4π steradian solid angle surrounding the wearer, since the wearer can move their body, head, or eyes to perceive substantially any direction in space. In other situations, the user's movement may be more constricted, and thus the user's FOR may be for a smaller solid angle.
[0089] The FOR may contain a group of objects, which may be perceived by the user through the ARD. The objects may be virtual and / or physical objects. Virtual objects may include operating system objects, such as, for example, a trash can for deleted files, a terminal for entering commands, a file manager for accessing files or directories, icons, menus, applications for audio or video streaming, notifications from the operating system, etc. Virtual objects may also include objects within an application, such as, for example, avatars, widgets (e.g., a virtual representation of a wall clock), virtual objects, graphics or images within a game, etc. Some virtual objects may be both operating system objects and objects within an application.
[0090] In some embodiments, virtual objects may be associated with physical objects. For example, as shown in Figure 12, a virtual screen 1250 may be placed on a table 1246. The virtual screen may include a virtual menu 1210 with selectable options such as office productivity tools 1212, applications for performing telepresence 1214, and email tools 1216.
[0091] The virtual object may be a three-dimensional (3D), two-dimensional (2D), or one-dimensional (1D) object. The virtual object may be a 3D coffee mug (which may represent a virtual control for a physical coffee maker). The virtual object may also be a 2D menu 1210 (shown in FIG. 12). In some implementations, one or more virtual objects may be displayed within (or associated with) another virtual object. For example, referring to FIG. 12, a virtual menu 5110 is shown inside a virtual screen 1250. In another example, a virtual application for telepresence 1214 may include another menu 1220 with contact information.
[0092] In some implementations, some objects in the user's environment may be interactable. For example, referring to FIG. 1, the user may interact with some of the virtual objects, such as by pointing out a finger for the avatar 140 to land on or by pulling up a menu that provides information about the statue 130. The user may interact with the interactable objects by performing user interface actions, such as selecting or moving the interactable object, activating a menu associated with the interactable object, selecting an action to be performed with the interactable object, etc. The user may perform these user interface actions using head pose, eye pose, body pose, voice, hand gestures on a user input device, alone or in combination. For example, the user may move a virtual object from one location to another using a change in body pose (e.g., a change in hand gesture, such as waving its hand toward the virtual object). In another example, as shown in FIG. 12, the user may use a hand gesture to activate the user input device to open a virtual screen 1250 when the user is standing near a table 1246. A user may actuate the user input device 466, such as by clicking a mouse, tapping a touchpad, swiping a touchscreen, hovering over or touching a capacitive button, pressing a key on a keyboard or game controller (e.g., a five-way d-pad), pointing a joystick, wand, or totem toward an object, pressing a button on a remote control, or other interaction with the user input device. In some implementations, the wearable device 1270 can automatically present a virtual menu 1210 in response to detecting the table 1246 (e.g., using one or more object recognizers 708). After the menu is opened, the user can browse the menu 1210 by moving their finger along a trajectory on the user input device.When the user decides to close the virtual screen 1250, the user may utter a word (e.g., "Exit") and / or actuate a user input device to indicate an intent to close the virtual screen 1250. After receiving the indication, the ARD may stop projecting the screen 1250 onto the table 1246. (Illustrative objects in view)
[0093] Within a FOR, the portion of the world that a user perceives at a given time is referred to as the FOV (e.g., the FOV may encompass the portion of the FOR that the user is currently looking at). The FOV may depend on the size or optical properties of the display within the ARD. For example, an AR display may include an optical device that only provides AR functionality when the user is looking through a particular portion of the display. The FOV may correspond to the solid angle perceivable by a user when looking through an AR display, such as, for example, stacked waveguide assembly 480 (FIG. 4) or planar waveguide 632b (FIG. 6).
[0094] As the user's posture changes, the FOV changes accordingly, and the objects within the FOV may also change. Referring to FIG. 12, the user 210 may perceive a virtual screen 1250 when standing in front of a table 1246. However, as the user 210 walks towards a mirror 1242, the virtual screen 1250 may move out of its FOV. Thus, the user 210 would not be able to perceive the virtual screen 1250 when standing in front of the mirror 1242. In some embodiments, the virtual screen 1250 may follow the user 210 as he moves around in the office 1200. For example, the virtual screen 1250 may move from the table 1246 to the wall 1248 as the user 210 moves and stands in front of the mirror 1242. The content of the virtual screen, such as the options on the virtual menu 1210, may change as the virtual screen 1250 changes location. 12, when the virtual screen 1250 is on the table 1250, the user 210 may perceive a virtual menu 1210 that includes various office productivity items. However, when the user walks towards the mirror 1242, the user may be able to interact with a virtual wardrobe application that allows the user to simulate different outfit looks using the wearable device 1270. In one implementation, once the wearable device 1270 detects the mirror 1242, the wearable system can automatically initiate a communication (e.g., a telepresence session) with another user (e.g., the user's 210 personal assistant). (Example of rendering virtual objects within the FOV based on contextual factors)
[0095] As described herein, there are often multiple virtual objects or user interaction options associated with an object (e.g., physical or virtual) or environment of the user. For example, referring to Fig. 12, a virtual menu 1210 includes multiple interaction options such as office productivity tools 1212 (word processor, file folders, calendar, etc.), a telepresence application 1214 (via the wearable system) that allows the user to communicate with another user as if the other user were present in the user's 210 environment (e.g., the wearable system can project images of the other user to the user of the wearable system), and a mail tool that allows the user 210 to send and receive electronic mail (email) or text messages. In another example, the living room 1300 shown in FIGS. 13 and 14 may include virtual objects such as a digital frame 1312a, a telepresence tool 1314a, a race car driving game 1316a, a television (TV) application 1312b, a home management tool 1314b (which can control the temperature for the room 1300, project wallpaper, etc.), and a music application 1316b.
[0096] However, as described herein, a virtual user interface may not be able to display all available virtual objects or user interaction options to the user and at the same time provide a satisfying user experience. For example, as shown in Figures 13 and 14, there are six virtual objects (1312a, 1314a, 1316a, 1312b, 1314b, and 1316b) associated with the user interface 1310, which may be virtual menus. However, the virtual menu on the wall 1350 may only fit three options readably (see, for example, the virtual menus in Figures 13 and 14). As a result, the wearable system may need to filter the number of available options and display only a subset of the available options.
[0097] Advantageously, in some embodiments, the wearable system can filter or select user interaction options or virtual objects to be presented on the user interface 1310 based on the context information. The filtered or selected user interface interaction options or virtual objects may be presented in various layouts. For example, the wearable device can present the options and virtual objects in a list form (such as the virtual menus shown in FIGS. 12-14). In an embodiment, rather than displaying on the virtual user interface as a vertical list of virtual objects, the virtual menu can employ a circular representation of the virtual objects (see, for example, the virtual menu 1530 shown in FIG. 15). The virtual objects can be rotated around the center of the circular representation to aid in identifying and selecting the desired virtual object. The context information can include information associated with the user's environment, the user, objects in the user's environment, and the like. Exemplary context information can include the affordances of a physical object with which the option is associated, the user's environment (such as whether the environment is a home or office environment), characteristics of the user, the user's current interaction with objects in the environment, the user's physiological state, the user's psychological state, combinations thereof, or the like. A more detailed description of the various types of context information is provided below. (User environment)
[0098] The wearable system may filter or select virtual objects in an environment based on the user's environment and present only a subset of virtual objects for user interaction. This is because different environments may have different functionality. For example, a user's contact list may include contact information for family, friends, and professional contacts. In an office environment, such as the office 1200 shown in FIG. 12, is typically more suitable for work-related activities instead of entertainment activities. As a result, when the user 210 uses the telepresence tool 1214, the wearable device 1270 may present a list of work-related contacts in the menu 1220 for the telepresence session, but the user's contact list also includes contacts of family and friends. In contrast, FIG. 13 depicts a living room 1300 where a user typically relaxes and interacts with people outside of work. As a result, when the user selects the telepresence tool 1314a, the wearable device may present contact information for friends and family in the menu 1320.
[0099] As another example, a user's music collection may include a variety of music, such as country music, jazz, pop, and classical music. When the user is in the living room 1300, the wearable device may present the user with jazz and pop music (as shown in virtual menu 1430). However, when the user is in the bedroom 1500, a different set of music options may be presented. For example, as shown in virtual menu 1530 in FIG. 15, the wearable device may present country music and classical music because these types of music may have a relaxing effect and may help the user fall asleep.
[0100] In addition to or as an alternative to filtering the virtual objects available in an environment, the wearable device may show only menu options related to the functionality of the environment. For example, the virtual menu 1210 in the office 1200 (in FIG. 12) may include options related to a work environment. On the other hand, the virtual user interface 1310 in the living room 1300 (in FIG. 14) may include entertainment items such as a virtual TV 1312b, music 1316b, as well as home management tools 1314b.
[0101] 12-15 are illustrative and not intended to limit the types of environments in which a wearable device may be used to contextually interact with physical and virtual content within such environments. Other environments may include other parts of a home or office, a vehicle (e.g., a car, subway, boat, train, or airplane), an entertainment venue (e.g., a movie theater, a nightclub, a gaming venue), a retail establishment (e.g., a store or mall), or outdoors (e.g., a park or garden), etc. (Object Affordance)
[0102] The wearable system can identify objects in the environment that the user may be interested in interacting with or is currently interacting with. The wearable system can identify objects based on the user's pose, such as, for example, eye gaze, body posture, or head pose. For example, the wearable device may track the user's eye pose using an inward-facing imaging system 462 (shown in FIG. 4). Once the wearable system determines that the user is looking in a direction for an extended period of time, the wearable system may use ray casting or cone casting techniques to identify objects that intersect the user's gaze direction. For example, the wearable system may cast a virtual cone / ray and identify objects that intersect with a portion of the virtual cone / ray. The wearable device may also track the user's head pose using an IMU (e.g., as described with reference to FIGS. 2, 4, and 9). Once the wearable device detects a change in the user's head pose, the wearable device may identify objects in the vicinity of the user's head as objects that the user is interested in interacting with. As an example, when a user of a wearable device looks at a refrigerator at home for an extended period of time, the wearable device can recognize that the refrigerator may be an object of interest to the user. In another example, the user may stand in front of the refrigerator. The wearable may detect a nod by the user and identify the refrigerator in front of the user as an object of interest to the user. In yet another example, the object recognizer 708 can track the movement of the user's hand (e.g., based on data from the outward-facing imaging system 464). The object recognizer can recognize hand gestures (e.g., a finger pointing at the refrigerator) that provide an indication of an object for user interaction.
[0103] The wearable system can recognize affordances of the identified object. The affordances of an object include relationships between the object and the object's environment that provide opportunities for actions or uses associated with the object. Affordances may be determined based on, for example, the function, orientation, type, location, shape, and / or size of the object. Affordances may also be based on the environment in which the physical object is located. The wearable device can filter available virtual objects in the environment and present the virtual objects according to the affordances of the object. As an example, the affordance of a horizontal table is that an object can be set on the table, and the affordance of a vertical wall is that an object can be hung from or projected onto the wall.
[0104] For example, the wearable device may identify the function of the object and present a menu with only objects related to the function of the object. As an example, when a user of the wearable device interacts with a refrigerator at home, the wearable device may identify that one of the refrigerator's functionalities is to store food. The ability to store food is an affordance of the refrigerator. When the user decides to view options associated with the refrigerator, for example by activating a user input device, the wearable device may present the user with food specific options, such as a list of foods currently available in the refrigerator, a cooking application with various recipes, a grocery list of food items, a reminder to change the water filter in the refrigerator, etc. Additional examples of affordances of the refrigerator include that it is heavy and therefore difficult to move, that it has a vertical front surface to which objects can be attached, that the front surface is often metallic and magnetic such that magnetic objects can stick to the front surface, etc.
[0105] In some situations, the functionality of the same object may vary based on the environment. The wearable system may generate a virtual menu by considering the functionality of the object in light of the environment. For example, the affordances of a table include that it may be used for writing and eating. When the table is in an office 1200 (shown in FIG. 12 ), the affordances of the table may suggest that the table should be used for writing, since an office environment is typically associated with word processing. Thus, the display 220 of the wearable device may present a word processing application under office tools 1212 or an email application 1216 in the virtual menu 1210. However, when the same table is located in a kitchen, the affordances of the table may suggest that the table may be used for eating, since people do not typically write documents in the kitchen. As a result, the wearable device may display to the user food-related virtual objects instead of office tools.
[0106] The wearable system can use the orientation of the object to determine which options should be presented because some activities (such as drawing pictures and writing documents) may be more appropriate on a horizontal surface (such as a floor or table), while other activities (such as watching TV or playing a driving game) may have a better user experience on a vertical surface (such as a wall). The wearable system can detect the orientation of the object's surface (e.g., horizontal vs. vertical) and display a group of options appropriate for that orientation.
[0107] 12, the office 1200 may include virtual objects such as office tools 1212 for word processing and a virtual TV application (not shown in FIG. 12). Because the virtual screen 1250 is on a table 1246, which has a horizontal surface, the wearable device 1270 may present the menu 1210 with the office tools 1212 because word processing is more appropriately performed on a horizontal surface. On the other hand, the wearable device 12700 may be configured not to present the virtual TV application because a virtual TV may be more appropriate for a vertical surface and the user is currently interacting with an object having a horizontal surface. However, if the user is standing in front of the wall 1248, the wearable device 1270 may include the virtual TV in the menu 1210 while excluding the office tools 1212.
[0108] As another example, in Figures 13 and 14, a virtual user interface 1310 is on a wall 1350, which has a vertical surface. As a result, the user interface 1310 may include a driving game 1316a, as shown in Figure 13, and a virtual TV application 1316b, as shown in Figure 14. This is because a user may have a better experience when performing these activities on a vertical surface.
[0109] Additionally or alternatively, the function, orientation, location, and affordances of an object may also be determined based on the type of object. For example, a sofa may be associated with entertainment activities such as watching TV, while a desk chair may be associated with work-related activities such as creating financial documents. The wearable system may also determine affordances based on the size of the object. For example, a small table may be used to hold decorative items such as vases, while a large table may be used for family meals. As another example, affordances may also be based on the shape of the object. A table with a circular top may be associated with certain group games such as poker, while a table with a rectangular top may be associated with single player games such as Tetris. (User characteristics)
[0110] The wearable system may also present options based on characteristics of the user, such as age, gender, education level, occupation, preferences, etc. The wearable system may identify these characteristics based on profile information provided by the user. In some embodiments, the wearable system may infer these characteristics based on the user's interactions with the wearable system (e.g., frequently viewed content, etc.). Based on the user's characteristics, the wearable system can present content that matches the user's characteristics. For example, as shown in FIG. 15, if the user of the wearable device is a young child, the wearable device may provide options for children's music and lullabies in a menu 1530 in a bedroom 1500.
[0111] In some implementations, an environment may be shared by multiple people. The wearable system may analyze characteristics of the people sharing the space and present only content suitable for the people sharing the space. For example, the living room 1300 may be shared by all family members. If the family has young children, the wearable device may present only movies that have a rating such that the movie is suitable for children without adult accompaniment (e.g., "G" rated movies). In some embodiments, the wearable system may identify people in the same environment as the wearable system images the environment and present options based on who is present in the environment. For example, the wearable system may obtain images of the environment using an outward-facing imaging system, and the wearable system may analyze the images using facial recognition techniques to identify one or more people present in the images. If the wearable system determines that a child is wearing an HMD and is seated with its parents in the same living room, the wearable system may present movies that are designated suitable for children with adult accompaniment (e.g., "PG" rated movies) in addition to or as an alternative to G rated movies.
[0112] The wearable system may present a virtual menu based on the user's preferences. The wearable system may infer the user's preferences based on previous usage patterns. The previous usage patterns may include information about the location and / or time when the virtual object was used. For example, every morning when the user 210 launches the virtual screen 1250 in his office 1200 (shown in FIG. 12), the user 210 typically checks his email first. Based on this usage pattern, the wearable system may display the email application 1216 in the menu 1210. In another example, every morning when the user 210 walks into the living room 1300, the user typically watches the news on the virtual TV screen 1312b (shown in FIG. 14). The wearable system may therefore display the TV application 1312b on the virtual user interface 1310 based on the user's frequent usage. However, the wearable system would not show the email application 1210 in the virtual user interface 1310 because the user 210 does not typically check his / her email in his / her living room 1300. On the other hand, if the user 210 frequently checks his / her email regardless of his / her location, the wearable system may show the email application 1216 in the virtual user interface 1310 as well.
[0113] Options in a menu may vary according to the time of day. For example, if a user 210 typically listens to jazz or pop music in the morning and plays driving games in the evening, the wearable system may present options in the menu 1430 for jazz music and pop music in the morning, while presenting the driving game 1316a in the evening.
[0114] The AR system may also allow a user to input their preferences. For example, a user 210 may add contact information for their boss to their contact list 1220 (shown in FIG. 12 ) even though they may not talk to the boss often. (Interaction between the user and objects in the environment)
[0115] The wearable device may present a subset of virtual objects in the user's environment based on the current user interaction. The wearable device may present virtual objects based on the person the user is interacting with. For example, when the user has a telepresence session with one of his family members in his living room 1300, the wearable device may automatically put up a photo album on the wall because the user may want to talk about the shared experience captured by the photo album. On the other hand, when the user has a telepresence session with one of his colleagues in his office 1200, the wearable device may automatically present a document that the colleague is collaborating on.
[0116] The wearable device may also present virtual objects based on the virtual object the user is interacting with. For example, if the user is currently preparing a financial document, the wearable device may present the user with a data analysis tool, such as a calculator, but if the user is currently writing a novel, the wearable device may present the user with a word processing tool. (The user's physiological or psychological state)
[0117] The contextual information may include the user's physiological state, psychological state, or autonomic nervous system activity, combinations thereof, or the like. As described with reference to FIG. 2, the wearable system may use various sensors 232 to measure the user's reaction to the virtual content or environment. For example, one or more sensors 232 may acquire data of the user's eye region and use such data to determine the user's mood. The wearable system may acquire images of the eye using an inwardly facing imaging system 462 (shown in FIG. 4). The ARD may use the images to determine eye movement, pupil dilation, and heart rate variability. In some implementations, when the inwardly facing imaging system 462 has a sufficiently large field of view, the wearable system may use images acquired by the inwardly facing imaging system 462 to determine the user's facial expression. The wearable system may also use an outwardly facing imaging system 464 (shown in FIG. 2) to determine the user's facial expression. For example, the outwardly facing imaging system 464 can obtain a reflected image of the user's face when the user is standing near a reflective surface (such as a mirror). The wearable system can analyze the reflected image and determine the user's facial expression. Additionally or alternatively, the wearable system can include a sensor that measures skin potential (such as a galvanic skin response). The sensor may be part of the user's wearable glove and / or the user input device 466 (described with reference to FIG. 4). The wearable system can use the skin potential data to determine the user's emotion.
[0118] The wearable system may also include sensors for electromyography (EMG), electroencephalography (EEG), functional near-infrared (fNIR), and the like. The wearable system may use data obtained from these sensors to determine the user's psychological and physiological state. These data may be used alone or in combination with data obtained from other sensors, such as inwardly facing imaging systems, outwardly facing imaging systems, and sensors for measuring skin potential. The ARD may use information of the user's psychological and physiological state to present virtual content (such as a virtual menu) to the user.
[0119] As an example, the wearable system can suggest a set of entertainment content (e.g., games, movies, music, scenes to be displayed) based on the user's mood. The entertainment content may be suggested to improve the user's mood. For example, the wearable system may determine that the user is currently under stress based on the user's physiological data (sweating, etc.). The wearable system can further determine that the user is at work based on information obtained from a location sensor (e.g., GPS) or images obtained from the outward-facing imaging system 464. The wearable system can combine these two sets of information and determine that the user is experiencing stress at work. Thus, the wearable system can suggest that the user play a slow-paced, exploratory and non-restrictive game, display relaxing scenes on a virtual display, play calming music in the user's environment, etc. to calm down during their lunch break.
[0120] As another example, multiple users may be present together or interact with each other in a physical or virtual space. The wearable system can determine one or more shared moods between the users (such as whether the group is happy or angry). The shared moods may be a combination or fusion of the moods of multiple users, or may target a common mood or theme between the users. The wearable system can present a virtual activity (such as a game) based on the shared moods between the users. (Reference marker)
[0121] In some implementations, the contextual information may be encoded in a fiducial marker (also referred to herein as a label). The fiducial marker may be associated with a physical object. The fiducial marker may be an optical marker, such as a Quick Response (QR) code, a barcode, an ArUco marker (which can be reliably detected under occlusion), etc. The fiducial marker may also comprise an electromagnetic marker (e.g., a radio frequency identification tag) that may emit or receive an electromagnetic signal detectable by the wearable device. Such a fiducial marker may be physically affixed on or near the physical object. The wearable device may detect the fiducial marker using an outwardly facing imaging system 464 (shown in FIG. 4) or one or more sensors that receive or transmit signals from or to the fiducial marker (e.g., transmit a signal to the marker, which may then relay the signal back).
[0122] When the wearable device detects the fiducial marker, the wearable device can decode the fiducial marker and present a group of virtual objects based on the decoded fiducial marker. In some embodiments, the fiducial marker may contain a reference to a database that includes an association between the virtual object to be displayed and the associated contextual factors. For example, the fiducial marker can include an identifier of a physical object (e.g., a table, etc.). The wearable device can use the identifier to access contextual properties of the physical object. In this example, the contextual properties of the table can include the size of the table and the horizontal surface. The accessed properties can be used to determine the user interface actions or virtual objects supported by the physical object. Because the table has a horizontal surface, the wearable device can present an office processing tool on the surface of the table, but not a painting, because a painting is typically associated with a vertical surface, rather than a horizontal surface.
[0123] For example, in FIG. 15, an ArUco marker 1512 is attached to a window 1510 in a bedroom 1500. The wearable system can identify the ArUco marker 1512 (e.g., based on an image acquired by the outward-facing imaging system 464) and extract information encoded within the ArUco marker 1512. In some embodiments, the extracted information may be sufficient for the wearable system to render a virtual menu associated with the reference marker without analyzing other contextual information of the window or the user's environment. For example, the extracted information may include the orientation of the window 1510 (e.g., the window 1510 has a vertical surface). Based on the orientation, the wearable system can present a virtual object associated with the vertical surface. In some implementations, the wearable system may need to communicate with another data source (e.g., remote data repository 280) to obtain contextual information, such as the type of object (e.g., a mirror), the environment associated with the object, user preferences, etc., in order for the wearable system to render and identify the associated virtual object. Exemplary Method for Rendering Virtual Objects Based on Contextual Factors
[0124] FIG. 16 is a flowchart of an exemplary method for generating a virtual menu based on context information. Process 1600 can be implemented by a wearable system described herein. The wearable system can include a user input device (see, for example, user input device 466 in FIG. 4) configured to receive indications of various user interactions, a display that can display virtual objects in proximity to physical objects, and a posture sensor. The posture sensor can detect and track a user's posture, which can include the orientation of the user's body relative to the environment or the position or movement of a part of the user's body, such as a gesture made by the user's hand or the user's gaze direction. The posture sensor can include an inertial measurement unit (IMU), an outward-facing imaging system, and / or an eye-tracking camera (e.g., camera 464 shown in FIG. 4), described with reference to FIGS. 2, 4, and 9. The IMU can include an accelerometer, a gyroscope, and other sensors.
[0125] In block 1610, the wearable system can determine a user's posture using one or more posture sensors. As described herein, the posture may include eye posture, head posture, body posture, combinations thereof, or the like. Based on the user's posture, in block 1620, the wearable system can identify interactable objects in the user's environment. For example, the wearable system may use cone projection techniques to identify objects that intersect with the user's gaze direction.
[0126] In block 1630, the user can operate the user input device and provide an indication for opening a virtual menu associated with the interactive object. The virtual menu may include a plurality of virtual objects as menu options. The plurality of virtual objects may be a subset of the virtual objects in the user's environment or a subset of the virtual objects associated with the interactive object. The virtual menu may have many graphic representations. Some examples of the virtual menu are shown as object 1220 and object 1210 in FIG. 12, virtual user interface 1310 in FIGS. 13 and 14, object 1320 in FIG. 13, object 1430 in FIG. 14, and object 1530 in FIG. 15. In one implementation, the indication for opening the virtual menu need not be received from the user input device. The indication may be associated directly with user input, such as, for example, the user's head pose, eye gaze, body pose, gesture, voice command, etc.
[0127] In block 1640, the wearable system can determine context information associated with the interactive object. The context information can include combinations of, or equivalents to, the affordances of the interactive object, the functions of the environment (e.g., work environment or living environment), the characteristics of the user (such as the user's age or preferences), or the current user interaction with objects in the environment. For example, the wearable system can determine the affordances of the interactive object by analyzing the characteristics of the interactive object, such as its function, orientation (horizontal vs. vertical), location, shape, size, etc. The wearable system can also determine the affordances of the interactive object by analyzing its relationship with the environment. For example, an end table in a living room environment may be used for entertainment purposes, while an end table in a bedroom environment may be used to hold items before a person goes to bed.
[0128] Contextual information associated with an interactable object may also be determined from a user's characteristics or the user's interactions with objects in the environment. For example, the wearable system may identify the user's age and present only information that corresponds to the user's age. As another example, the wearable system may analyze the user's previous usage patterns with a virtual menu (such as the types of virtual objects the user frequently uses) and tailor the content of the virtual menu according to the previous usage patterns.
[0129] In block 1650, the wearable system can identify a list of virtual objects to be included in the virtual menu based on the context information. For example, the wearable system can identify a list of applications with features related to the interactable objects and / or features related to the user's environment. As another example, the wearable system may identify virtual objects to be included in the virtual menu based on characteristics of the user, such as age, gender, and previous usage patterns.
[0130] In block 1660, the wearable system may generate a virtual menu based on the identified list of virtual objects. The virtual menu may include all virtual objects on the identified list. In some embodiments, the menu may be confined in space. The wearable system may prioritize different types of contextual information such that only a subset of the list is shown to the user. For example, the wearable system may determine that previous usage patterns are the most important contextual information and therefore display only the top five virtual objects based on previous usage patterns.
[0131] A user can perform various actions with the menu, such as, for example, browsing through the menu, showing available virtual objects that were not previously selected based on an analysis of some of the context information, exiting the menu, or selecting one or more objects on the menu for interaction. Exemplary Method for Rendering Virtual Objects Based on Physiological Data of a User
[0132] 17 is a flowchart of an example method for selecting virtual content based, at least in part, on physiological data of a user. Process 1700 can be implemented by a wearable system described herein. The wearable system may include various sensors 232, such as, for example, a physiological sensor configured to measure physiological parameters of the user, an inward-facing imaging system configured to track an eye region of the user, etc.
[0133] In block 1710, the wearable system can acquire physiological data of the user. As described with reference to Figures 2 and 4, the wearable system can use physiological sensors, alone or in combination with an inwardly facing imaging system, to measure the physiological data of the user. For example, the wearable system can use one or more physiological sensors to determine the galvanic skin response of the user. The wearable system can further use an inwardly facing imaging system to determine the eye movement of the user.
[0134] As shown in block 1720, the wearable system can use the physiological data to determine the physiological or psychological state of the user. For example, the wearable system can use the user's galvanic skin response and / or the user's eye movements to determine whether the user is excited by certain content.
[0135] In block 1730, the wearable system can determine virtual content to be presented to the user. The wearable system can make such a determination based on the physiological data. For example, using the physiological data, the wearable system may determine that the user is under stress. The wearable system can therefore present a virtual object (such as music or a video game) associated with relieving the user's stress.
[0136] The wearable system can also present virtual content based on an analysis of the physiological data in combination with other contextual information. For example, based on the user's location, the wearable system may determine that the user is experiencing stress while at work. Because users typically do not play games or listen to music during work hours, the wearable system may suggest video games and music only during the user's breaks to help the user relieve their stress levels.
[0137] In block 1740, the wearable system can generate a 3D user interface that includes the virtual content. For example, the wearable system may show icons for music and video games when it detects that the user is experiencing stress. The wearable system can present a virtual menu while the user interacts with a physical object. For example, the wearable system can show icons for music and video games when the user activates a user input device in front of a desk during their work break.
[0138] Techniques in various embodiments described herein can provide a user with a subset of available virtual objects or user interface interaction options. This subset of virtual objects or user interface interaction options can be provided in various forms. Although embodiments are described primarily with reference to presenting menus, other types of user interface presentations are also available. For example, the wearable system can render icons of virtual objects within the subset of virtual objects. In some implementations, the wearable system can automatically perform actions based on the context information. For example, the wearable system may automatically initiate a telepresence session with the user's most frequent contacts when the user is in the vicinity of a mirror. As another example, the wearable system can automatically invoke a virtual object when the wearable system determines that the user is most likely to be interested in that virtual object. (Other embodiments)
[0139] In a first aspect, a method for generating a virtual menu in a user's environment in three-dimensional (3D) space, the method comprising: an augmented reality (AR) system comprising computer hardware, the AR system configured to enable user interaction with objects in the user's environment, the AR system comprising a user input device, an AR display, and an inertial measurement unit (IMU) configured to detect a user's posture; determining a user's posture using the IMU under control of the AR system; identifying a physical object in the user's environment in the 3D space based at least in part on the user's posture; receiving an indication via the user input device to open a virtual menu associated with the physical object; determining context information associated with the physical object; determining virtual objects to be included in the virtual menu based at least in part on the determined context information; determining a spatial location for displaying the virtual menu based at least in part on the determined context information; generating the virtual menu including at least the determined virtual objects; and displaying the generated menu to a user at the spatial location via the AR display.
[0140] In a second aspect, the method of aspect 1, wherein the pose includes one or more of a head pose or a body pose.
[0141] In a third aspect, the method of aspect 1 or aspect 2, wherein the ARD further comprises an eye tracking camera configured to track the user's eye posture.
[0142] In a fourth aspect, the method of aspect 3, wherein the posture includes an eye posture.
[0143] In a fifth aspect, the method of any one of aspects 1-4, wherein the contextual information includes one or more of affordances of physical objects, features of the environment, characteristics of the user, or current or past interactions of the user with the AR system.
[0144] In a sixth aspect, the method of aspect 5, wherein the affordance of a physical object includes a relationship between the physical object and the physical object's environment that provides an opportunity for an action or use associated with the physical object.
[0145] In a seventh aspect, the method of aspect 5 or aspect 6, wherein the affordance of the physical object is based, at least in part, on one or more of a function, an orientation, a type, a location, a shape, a size of the physical object, or an environment in which the physical object is located.
[0146] In an eighth aspect, the method of aspect 7, wherein the orientation of the physical object includes horizontal or vertical.
[0147] In a ninth aspect, the method of any one of aspects 5-8, wherein the environment is a living environment or a working environment.
[0148] In a tenth aspect, the method of any one of aspects 5-9, wherein the environment is a private environment or a public environment.
[0149] In an eleventh aspect, the method of aspect 5, wherein the user characteristics include one or more of age, gender, education level, occupation, or preferences.
[0150] In a twelfth aspect, the method of aspect 11, wherein the preference is based, at least in part, on the user's previous usage patterns, the previous usage patterns including information about where or when the virtual object was used.
[0151] In a thirteenth aspect, the method of aspect 5, wherein the current interaction includes a telepresence session between the user of the AR system and another user.
[0152] In a fourteenth aspect, the method of any one of aspects 1-13, wherein the context information is encoded within a reference marker, the reference marker being associated with the physical object.
[0153] In a fifteenth aspect, the method of aspect 14, wherein the reference marker includes an optical marker or an electromagnetic marker.
[0154] In a sixteenth aspect, the method of any one of aspects 1-15, wherein the object includes at least one of a physical object or a virtual object.
[0155] In a seventeenth aspect, a method for rendering a plurality of virtual objects in a user's environment in three-dimensional (3D) space, the method comprising: an augmented reality (AR) system comprising computer hardware, the AR system configured to enable user interaction with objects in the user's environment, the AR system comprising a user input device, an AR display, and a posture sensor configured to detect a posture of the user; determining a posture of a user using the posture sensor under control of the AR system; identifying an interactable object in the user's environment in the 3D space based at least in part on the user's posture; receiving an indication via the user input device to present a plurality of virtual objects associated with the interactable object; determining contextual information associated with the interactable object; determining a plurality of virtual objects to be displayed to the user based at least in part on the determined contextual information; and displaying the determined plurality of virtual objects to the user via the AR display.
[0156] In an eighteenth aspect, the method of aspect 17, wherein the attitude sensor includes one or more of an inertial measurement unit, an eye tracking camera, or an outward-facing imaging system.
[0157] In a nineteenth aspect, the method of aspect 17 or aspect 18, wherein the posture includes one or more of a head posture, an eye posture, or a body posture.
[0158] In a twentieth aspect, the method of any one of aspects 17-19, wherein the contextual information includes one or more of affordances of an interactable object, features of the environment, characteristics of the user, or a user's current or past interactions with the AR system.
[0159] In a twenty-first aspect, the method of aspect 20, wherein the affordance of an interactable object includes a relationship between the interactable object and the environment of the interactable object that provides an opportunity for an action or use associated with the interactable object.
[0160] In a twenty-second aspect, the method of aspect 20 or aspect 21, wherein the affordance of the interactable object is based, at least in part, on one or more of the function, orientation, type, location, shape, size, or environment in which the physical object is located.
[0161] In a twenty-third aspect, the method of aspect 22, wherein the orientation of the interactable object includes horizontal or vertical.
[0162] In a 24th aspect, the method of any one of aspects 20-23, wherein the environment is a living environment or a working environment.
[0163] In a 25th aspect, the method of any one of aspects 20-24, wherein the environment is a private environment or a public environment.
[0164] In a twenty-sixth aspect, the method of aspect 20, wherein the user characteristics include one or more of age, gender, education level, occupation, or preferences.
[0165] In a twenty-seventh aspect, the method of aspect 26, wherein the preferences are based, at least in part, on the user's previous usage patterns, including information about where or when the virtual object was used.
[0166] In a twenty-eighth aspect, the method of aspect 20, wherein the current interaction includes a telepresence session.
[0167] In a twenty-ninth aspect, the method of any one of aspects 17-28, wherein the context information is encoded within a reference marker, the reference marker being associated with the interactable object.
[0168] In a 30th aspect, the method of aspect 29, wherein the reference marker includes an optical marker or an electromagnetic marker.
[0169] In a thirty-first aspect, the method of any one of aspects 17-30, wherein the object includes at least one of a physical object or a virtual object.
[0170] In a thirty-second aspect, the method of any one of aspects 17-31, wherein the interactable object includes at least one of a physical object or a virtual object.
[0171] In a thirty-third aspect, an augmented reality (AR) system includes computer hardware, a user input device, an AR display, and an orientation sensor, and is configured to perform any one of the methods described in aspects 1-32.
[0172] In a thirty-fourth aspect, a method for selectively presenting virtual content to a user in a three-dimensional space (3D), the method including: acquiring data associated with a physiological parameter of the user using the physiological sensor under control of a wearable device comprising a computer processor, a display, and a physiological sensor configured to measure a physiological parameter of the user; determining a physiological state of the user based at least in part on the data; determining virtual content to be presented to the user based at least in part on the physiological state; determining a spatial location for displaying the virtual content within the 3D space; generating a virtual user interface including at least the determined virtual content; and displaying the virtual content to the user at the determined spatial location via a display of the wearable device.
[0173] In a thirty-fifth aspect, the method of aspect 34, wherein the physiological parameters include at least one of heart rate, pupil dilation, galvanic skin response, blood pressure, brainwave status, respiratory rate, or eye movement.
[0174] In a 36th aspect, a method described in any one of aspects 34-35, further comprising a step of acquiring data associated with the user's physiological parameters using an inward-facing imaging system configured to image one or both eyes of the user.
[0175] In a 37th aspect, the method of any one of aspects 34-36, further comprising determining a psychological state based on the data.
[0176] In a thirty-eighth aspect, the method of any one of aspects 34-37, wherein the virtual content includes a virtual menu.
[0177] In the 39th aspect, the virtual content is further determined based on at least one of the affordances of the physical object associated with the virtual content, the functions of the user's environment, the user's characteristics, the individuals present in the environment, the information encoded in the reference markers associated with the physical object, or the user's current or past interactions with the wearable system, according to the method of any one of aspects 34 - 38.
[0178] In the 40th aspect, the affordance of the physical object includes the relationship between the physical object and the physical object's environment that provides an opportunity for an action or use associated with the physical object, according to the method of aspect 39.
[0179] In the 41st aspect, the affordance of the physical object is at least partially based on one or more of the function, orientation, type, location, shape, size of the physical object, or the environment in which the physical object is located, according to the method of aspect 39 or aspect 40.
[0180] In the 42nd aspect, the orientation of the physical object includes horizontal or vertical, according to the method of aspect 41.
[0181] In the 43rd aspect, the environment is a living environment or a working environment, according to the method of any one of aspects 39 - 42.
[0182] In the 44th aspect, the environment is a private environment or a public environment, according to the method of any one of aspects 39 - 43.
[0183] In the 45th aspect, the user's characteristics include one or more of age, gender, education level, occupation, or preferences, according to the method of aspect 39.
[0184] In the 46th aspect, the preference is at least partially based on the user's previous usage pattern, and the previous usage pattern includes information about the location or time where the virtual object was used, according to the method of aspect 45.
[0185] In a forty-seventh aspect, the method of aspect 39, wherein the current interaction includes a telepresence session between a user of the AR system and another user.
[0186] In a 48th aspect, the method of any one of aspects 34-47, wherein the wearable device includes an augmented reality system.
[0187] In a forty-ninth aspect, a wearable device comprising a computer processor, a display, and a physiological sensor configured to measure physiological parameters of a user, the wearable device being configured to perform any one of the methods described in aspects 34-48.
[0188] In a fiftieth aspect, a wearable system for generating virtual content within a user's three-dimensional (3D) environment, the wearable system comprising: an augmented reality display configured to present the virtual content to a user in a 3D view; a posture sensor configured to obtain user position or orientation data, analyze the position or orientation data, and identify the user's posture; and a hardware processor in communication with the posture sensor and the display, the hardware processor being programmed to: identify a physical object in the user's environment within the 3D environment based at least in part on the user's posture; receive an indication to initiate an interaction with the physical object; identify a set of virtual objects in the user's environment associated with the physical object; determine context information associated with the physical object; filter the set of virtual objects and identify a subset of the virtual objects from the set of virtual objects based on the context information; generate a virtual menu including the subset of the virtual objects; determine a spatial location within the 3D environment for presenting the virtual menu based at least in part on the determined context information; and present the virtual menu at the spatial location via the augmented reality display.
[0189] In a fifty-first aspect, the wearable system of aspect fifty, wherein the contextual information includes an affordance of the physical object including a relationship between the physical object and the physical object's environment that provides an opportunity for an action or use associated with the physical object, and the affordance of the physical object is based, at least in part, on one or more of the physical object's function, orientation, type, location, shape, size, or the environment in which the physical object is located.
[0190] In a fifty-second aspect, the wearable system of aspect fifty-one, wherein the context information includes an orientation of a surface of the physical object, and to filter the set of virtual objects, the hardware processor is programmed to identify a subset of the virtual objects that support user interface interaction on the surface having the orientation.
[0191] In a 53rd aspect, a wearable system described in any one of aspects 50-52, wherein the posture sensor includes an inertial measurement unit configured to measure the user's head posture and identify the physical object, and the hardware processor is programmed to project a virtual cone and select a physical object that intersects with a portion of the virtual cone based, at least in part, on the user's head posture.
[0192] In a 54th aspect, a wearable system as described in any one of aspects 50-53, further comprising a physiological sensor configured to measure physiological parameters of the user, and the hardware processor is programmed to determine a psychological state of the user, use the psychological state as part of the context information, and identify a subset of virtual objects for inclusion in the virtual menu.
[0193] In a fifty-fifth aspect, the wearable system of aspect 54, wherein the physiological parameter is related to at least one of heart rate, pupil dilation, galvanic skin response, blood pressure, brainwave state, respiratory rate, or eye movement.
[0194] In a 56th aspect, the 3D environment includes a plurality of users, and the hardware processor is programmed to determine common characteristics of the plurality of users and filter the set of virtual objects based on the common characteristics of the plurality of users, a wearable system described in any one of aspects 50-55.
[0195] In a fifty-seventh aspect, the wearable system of any one of aspects 50-56, wherein the context information includes the user's past interactions with the set of virtual objects, and the hardware processor is programmed to identify one or more virtual objects with which the user frequently interacts, and to include the one or more virtual objects within the subset of virtual objects for the virtual menu.
[0196] In a 58th aspect, a wearable system as described in any one of aspects 50-57, wherein to determine contextual information associated with the physical object and filter the set of virtual objects, the hardware processor is programmed to identify a reference marker associated with the physical object, where the reference marker encodes an identifier of the physical object, decode the reference marker and extract the identifier, access a database using the identifier that stores contextual information associated with the physical object, and analyze the contextual information stored in the database and filter the set of virtual objects.
[0197] In a fifty-ninth aspect, the wearable system of aspect 58, wherein the reference marker includes an ArUco marker.
[0198] In a 60th aspect, a wearable system described in any one of aspects 50-59, wherein the spatial location for rendering the virtual menu includes a position or orientation of the virtual menu relative to a physical object.
[0199] In a 61st aspect, the wearable system of aspect 60, wherein to determine a spatial location for rendering the virtual menu, the hardware processor is programmed to identify a space on a surface of the physical object using an object recognition device associated with the physical object.
[0200] In a sixty-second aspect, a method for generating virtual content within a user's three-dimensional (3D) environment includes analyzing data obtained from a posture sensor to identify a posture of the user; identifying an interactable object within the user's 3D environment based, at least in part, on the posture; receiving an indication to initiate an interaction with the interactable object; determining context information associated with the interactable object; selecting a subset of user interface actions from a set of user interface actions available on the interactable object based on the context information; and generating instructions for presenting the subset of user interface actions to the user in a 3D view.
[0201] In a 63rd aspect, the method of aspect 62, wherein the posture includes at least one of eye gaze, head posture, or gesture.
[0202] In a 64th aspect, the method of any one of aspects 62-63, wherein the step of identifying the interactable object includes the steps of performing a cone projection based on the user's head pose, and selecting an object in the user's environment as an interactable object that intersects with at least a portion of a virtual cone used in the cone projection.
[0203] In a 65th aspect, the method of any one of aspects 62-64, wherein the context information includes an orientation of a surface of the interactable object, and the step of selecting a subset of user interface actions includes a step of identifying user interface actions that can be performed on a surface having the orientation.
[0204] In a 66th aspect, the method of any one of aspects 62-65, wherein the indication to initiate an interaction with the interactable object includes actuation of a user input device or a change in the user's posture.
[0205] In a 67th aspect, a method described in any one of aspects 62-66, further comprising the steps of receiving physiological parameters of a user and determining a psychological state of the user, the psychological state being part of the context information for selecting a subset of user interactions.
[0206] In a 68th aspect, the method of aspect 67, wherein the physiological parameter is related to at least one of heart rate, pupil dilation, galvanic skin response, blood pressure, brainwave status, respiratory rate, or eye movement.
[0207] In a sixty-ninth aspect, the method of any one of aspects 62-68, wherein the step of generating instructions for presenting a subset of user interactions to a user in a 3D view includes the steps of generating a virtual menu including the subset of user interface actions, determining a spatial location of the virtual menu based on characteristics of the interactable objects, and generating display instructions for presentation of the virtual menu at the spatial location within the user's 3D environment. (Other considerations)
[0208] Each of the processes, methods, and algorithms described herein and / or depicted in the accompanying figures may be embodied in code modules executed by one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to execute specific and particular computer instructions, and thus may be fully or partially automated. For example, the computing system may include a general-purpose computer (e.g., a server) or a special-purpose computer programmed with specific computer instructions, special-purpose circuits, etc. The code modules may be written in a programming language that may be compiled and linked into an executable program, installed in a dynamic link library, or interpreted. In some implementations, certain operations and methods may be performed by circuitry specific to a given function.
[0209] Furthermore, certain implementations of the functionality of the present disclosure may be sufficiently mathematically, computationally, or technically complex that special purpose hardware (utilizing appropriate specialized executable instructions) or one or more physical computing devices may be required to implement the functionality, e.g., due to the amount or complexity of the calculations involved, or to provide results in substantially real time.
[0210] The code modules or any type of data may be stored on any type of non-transitory computer readable medium, such as physical computer storage devices, including hard drives, solid state memory, random access memory (RAM), read only memory (ROM), optical disks, volatile or non-volatile storage devices, combinations of the same, and / or the like. The methods and modules (or data) may also be transmitted as data signals (e.g., as part of a carrier wave or other analog or digital propagating signal) generated over a variety of computer readable transmission media, including wireless-based and wired / cable-based media, and may take a variety of forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed processes or process steps may be stored persistently or otherwise in any type of non-transitory tangible computer storage device or communicated via a computer readable transmission medium.
[0211] Any process, block, state, step, or functionality in the flow diagrams described herein and / or depicted in the accompanying figures should be understood as potentially representing a code module, segment, or portion of code that includes one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in the process. Various processes, blocks, states, steps, or functionality can be combined, rearranged, added, deleted, modified, or otherwise altered from the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may perform some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the blocks, steps, or states associated therewith can be performed in other sequences as appropriate, e.g., serially, in parallel, or in some other manner. Tasks or events may be added to or removed from the disclosed exemplary embodiments. Furthermore, the separation of various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems may generally be integrated together in a single computer product or packaged into multiple computer products. Many implementation variations are possible.
[0212] The process, method, and system may be implemented in a network (or distributed) computing environment. Network environments include enterprise-wide computer networks, intranets, local area networks (LANs), wide area networks (WANs), personal area networks (PANs), cloud computing networks, crowdsourced computing networks, the Internet, and the World Wide Web. The network may be a wired or wireless network or any other type of communication network.
[0213] The systems and methods of the present disclosure each have several innovative aspects, none of which is solely responsible for or required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure. Various modifications of the implementations described in the present disclosure may be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations without departing from the spirit or scope of the present disclosure. Thus, the claims are not intended to be limited to the implementations shown herein, but should be accorded the widest scope consistent with the present disclosure, the principles, and novel features disclosed herein.
[0214] Certain features described herein in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented separately in multiple implementations or in any suitable subcombination. Furthermore, although features may be described above as acting in a combination and may even be initially claimed as such, one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination. No single feature or group of features is required or essential to every embodiment.
[0215] In particular, conditional statements used herein, such as "can," "could," "might," "may," "eg," and the like, are generally intended to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context as used. Thus, such conditional statements are generally not intended to suggest that features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps should be included or performed in any particular embodiment, with or without authorial input or prompting. The terms "comprise," "include," "have," and the like, are synonymous and used inclusively in a non-limiting manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not its exclusive sense), so that, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. In addition, the articles "a," "an," and "the," as used in this application and the appended claims, unless otherwise specified, should be interpreted to mean "one or more" or "at least one."
[0216] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single elements. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Transitive phrases such as "at least one of X, Y, and Z" are generally understood differently in the context in which they are used to convey that an item, term, etc. may be at least one of X, Y, or Z, unless specifically stated otherwise. Thus, such transitive phrases are generally not intended to suggest that an embodiment requires that at least one of X, at least one of Y, and at least one of Z, respectively, be present.
[0217] Similarly, although operations may be depicted in the figures in a particular order, it should be appreciated that such operations need not be performed in the particular order depicted, or in sequential order, or that all of the depicted operations need not be performed to achieve desirable results. Additionally, the figures may diagrammatically depict one or more exemplary processes in the form of a flow chart. However, other operations not depicted may be incorporated within the diagrammatically depicted exemplary methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or during any of the depicted operations. Additionally, operations may be rearranged or reordered in other implementations. In some circumstances, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve desirable results. (Item 1) 1. A wearable system for generating virtual content within a user's three-dimensional (3D) environment, the wearable system comprising: an augmented reality display configured to present the virtual content to a user in a 3D view; a posture sensor configured to obtain position or orientation data of a user, analyze said position or orientation data, and identify a posture of said user; a hardware processor in communication with the orientation sensor and the display, the hardware processor comprising: identifying physical objects in the user's environment within the 3D environment based at least in part on a pose of the user; receiving an indication to initiate an interaction with the physical object; identifying a set of virtual objects in the user's environment that are associated with the physical object; determining context information associated with the physical object; filtering the set of virtual objects to identify a subset of virtual objects from the set of virtual objects based on the context information; generating a virtual menu including a subset of the virtual objects; determining a spatial location within the 3D environment for presenting the virtual menu based, at least in part, on the determined context information; and presenting said virtual menu at said spatial location by said augmented reality display; a hardware processor programmed to A wearable system comprising: (Item 2) 2. The wearable system of claim 1, wherein the contextual information includes an affordance of the physical object including a relationship between the physical object and an environment of the physical object that provides an opportunity for an action or use associated with the physical object, the affordance of the physical object being based, at least in part, on one or more of the function, orientation, type, location, shape, size, or environment in which the physical object is located. (Item 3) 3. The wearable system of claim 2, wherein the context information includes an orientation of a surface of the physical object, and to filter the set of virtual objects, the hardware processor is programmed to identify a subset of the virtual objects that support user interface interactions on the surface having the orientation. (Item 4) 2. The wearable system of claim 1, wherein the posture sensor comprises an inertial measurement unit configured to measure the user's head posture and identify the physical object, and the hardware processor is programmed to project a virtual cone and select the physical object that intersects with a portion of the virtual cone based, at least in part, on the user's head posture. (Item 5) 2. The wearable system of claim 1, further comprising a physiological sensor configured to measure physiological parameters of the user, wherein a hardware processor is programmed to determine a psychological state of the user, use the psychological state as part of the context information, and identify a subset of the virtual objects for inclusion in the virtual menu. (Item 6) 6. The wearable system of item 5, wherein the physiological parameter relates to at least one of heart rate, pupil dilation, galvanic skin response, blood pressure, brainwave state, respiratory rate, or eye movement. (Item 7) The 3D environment includes a plurality of users, and the hardware processor is programmed to determine common characteristics of the plurality of users and filter the set of virtual objects based on the common characteristics of the plurality of users, the wearable system according to item 1. (Item 8) The context information includes past interactions of the user with the set of virtual objects, and the hardware processor is programmed to identify one or more virtual objects with which the user has interacted frequently and include the one or more virtual objects within the subset of virtual objects for the virtual menu, the wearable system according to item 1. (Item 9) To determine context information associated with the physical object and filter the set of virtual objects, the hardware processor identifies a reference marker associated with the physical object, the reference marker encoding an identifier of the physical object, decodes the reference marker and extracts the identifier, accesses, using the identifier, a database storing context information associated with the physical object, analyzes the context information stored in the database and filters the set of virtual objects and is programmed to perform, the wearable system according to item 1. (Item 10) The reference marker includes an ArUco marker, the wearable system according to item 9. (Item 11) The spatial location for rendering the virtual menu includes the position or orientation of the virtual menu with respect to the physical object, the wearable system according to item 1. (Item 12) Item 12. The wearable system of item 11, wherein to determine the spatial location for rendering the virtual menu, the hardware processor is programmed to identify a space on a surface of the physical object using an object recognizer associated with the physical object. (Item 13) 1. A method for generating virtual content within a user's three-dimensional (3D) environment, the method comprising: analyzing data obtained from the posture sensor to identify a posture of the user; Identifying an interactable object within the user's 3D environment based at least in part on the pose; and receiving an indication to initiate an interaction with the interactable object; determining context information associated with the interactable object; selecting a subset of user interface actions from a set of user interface actions available on the interactable object based on the context information; generating instructions for presenting the subset of user interface actions to the user in a 3D view; A method comprising: (Item 14) Item 14. The method of item 13, wherein the pose includes at least one of eye gaze, head pose, or gesture. (Item 15) Item 14. The method of item 13, wherein identifying the interactable object includes performing a cone projection based on the user's head pose, and selecting an object in the user's environment as the interactable object, the object intersecting at least a portion of a virtual cone used in the cone projection. (Item 16) 14. The method of claim 13, wherein the contextual information includes an orientation of a surface of the interactable object, and selecting the subset of user interface actions includes identifying user interface actions that can be performed on a surface having the orientation. (Item 17) Item 14. The method of item 13, wherein the indication to initiate an interaction with the interactable object includes an actuation of a user input device or a change in the user's posture. (Item 18) 14. The method of claim 13, further comprising receiving physiological parameters of the user and determining a psychological state of the user, the psychological state being part of the context information for selecting the subset of user interactions. (Item 19) 20. The method of claim 18, wherein the physiological parameter is related to at least one of heart rate, pupil dilation, galvanic skin response, blood pressure, electroencephalogram status, respiratory rate, or eye movement. (Item 20) Item 14. The method of item 13, wherein generating instructions for presenting a subset of the user interactions to the user in a 3D view includes generating a virtual menu including the subset of the user interface actions, determining a spatial location of the virtual menu based on characteristics of the interactable objects, and generating display instructions for presentation of the virtual menu at a spatial location within the user's 3D environment.
Claims
1. 1. A method for generating a virtual menu in a user's environment in three-dimensional (3D) space, the method comprising: under control of an augmented reality (AR) system comprising computer hardware, the AR system configured to enable user interaction with objects in the user's environment, the AR system comprising a user input device, an AR display, and an inertial measurement unit (IMU) configured to detect a user's pose; determining the pose of the user using the IMU; and identifying physical objects in the environment of the user in the 3D space based at least in part on the pose of the user; receiving, via the user input device, an indication to open a virtual menu associated with the physical object; determining contextual information associated with the physical object and associated with a psychological or physiological state of the user, the physiological state including at least one of the user's heart rate, pupil dilation, galvanic skin response, blood pressure, electroencephalographic state, or respiratory rate; determining a priority of the determined context information associated with a psychological or physiological state of the user; determining a virtual object to be included within the virtual menu based, at least in part, on the determined context information having a highest priority; determining a spatial location for displaying the virtual menu based at least in part on the determined context information; generating the virtual menu including at least the determined virtual object; displaying the generated menu to the user via the AR display at the spatial location; The method of claim 1,
2. The method of claim 1 , wherein the poses include one or more of a head pose or a body pose.
3. The method of claim 1 , wherein the AR system further comprises an eye-tracking camera configured to track an eye posture of the user.
4. The method of claim 3 , wherein the pose comprises an eye pose.
5. The context information further comprises: the affordances of said physical object; the functionality of said environment, or The user's current or past interactions with the AR system The method of claim 1 , comprising one or more of:
6. The method of claim 5 , wherein the affordance of the physics object comprises a relationship between the physics object and the environment of the physics object that provides an opportunity for an action or use associated with the physics object.
7. The method of claim 5 , wherein the affordances of the physical object are based, at least in part, on one or more of the function, orientation, type, location, shape, size of the physical object, or an environment in which the physical object is located.
8. The method of claim 7 , wherein the orientation of the physical object comprises horizontal or vertical.
9. The method of claim 5 , wherein the environment is a living environment or a working environment.
10. The method of claim 5 , wherein the environment is a private environment or a public environment.
11. The method of claim 5 , wherein the current interaction includes a telepresence session between the user of the AR system and another user.
12. The method of claim 1 , wherein the contextual information is encoded in a fiducial marker, the fiducial marker being associated with the physical object.
13. The method of claim 12 , wherein the fiducial markers include optical or electromagnetic markers.
14. The method of claim 1 , wherein the object comprises at least one of a physical object or a virtual object.
15. 1. A method for rendering a plurality of virtual objects into a user's environment in three-dimensional (3D) space, the method comprising:
1. An augmented reality (AR) system comprising computer hardware configured to enable user interaction with objects in the user's environment, the AR system comprising a user input device, an AR display, and a posture sensor configured to detect a posture of the user, under control of the AR system: determining the posture of the user using the posture sensor; identifying interactable objects in the environment of the user in the 3D space based at least in part on the pose of the user; receiving an indication via the user input device to present a plurality of virtual objects associated with the interactable object; determining context information associated with the interactable object and associated with a psychological or physiological state of the user, the physiological state including at least one of the user's heart rate, pupil dilation, galvanic skin response, blood pressure, electroencephalographic state, or respiratory rate; determining a priority of the determined context information associated with a psychological or physiological state of the user; determining a plurality of virtual objects to be displayed to the user based, at least in part, on the determined context information having a highest priority; and displaying the determined plurality of virtual objects to the user via the AR display; The method of claim 1,
16. The method of claim 15 , wherein the attitude sensor includes one or more of an inertial measurement unit, an eye-tracking camera, or an outward-facing imaging system.
17. The method of claim 15 , wherein the pose includes one or more of a head pose, an eye pose, or a body pose.
18. The context information further comprises: the affordances of said interactable object; the functionality of said environment, or The user's current or past interactions with the AR system The method of claim 15 , comprising one or more of the following:
Citation Information
Patent Citations
Information processing apparatus, information processing method, and program
JP2014134922A
Information processing apparatus, information processing method, and program
JP2014174747A
User interface for augmented reality enabled devices
JP2016508257A
Manipulation of virtual object in augmented reality via intent
WO2014197387A1