Systems and Methods for Virtual and Augmented Reality
By using head-wearable devices to capture environmental and emotional inputs, the system generates virtual companions that provide context-based stimuli, effectively humanizing AI and enhancing the personalization and immersion of mixed reality experiences.
Patent Information
- Application Number
- JP2024030016
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-09
- Filing Date
- 2024-02-29
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2039-10-08
AI Technical Summary
Current mobile computing systems, particularly those designed for virtual and augmented reality, lack the ability to effectively humanize artificial intelligence by capturing the duality of human thought and emotion, leading to a lack of personal and context-based experiences.
The development of systems and methods that utilize head-wearable devices equipped with sensors to receive inputs from the user's environment and emotional reactions, allowing for the generation and display of virtual companions that provide context-based stimuli based on determined associations between emotional reactions and events.
This approach enables the creation of more immersive and personalized mixed reality experiences by allowing AI to respond to user emotions and environmental contexts, thereby enhancing user engagement and memory formation.
Smart Images

Figure 0007697085000001 
Figure 0007697085000002 
Figure 0007697085000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority of U.S. Provisional Application No. 62 / 743,492, filed on October 9, 2018, the content of which is incorporated herein by reference in its entirety.
[0002] The present invention relates to mobile computing systems, methods, and configurations, and more particularly to mobile computing systems, methods, and configurations characterized by at least one wearable component that can be utilized for virtual and / or augmented reality operations.
Background Art
[0003] One goal of artificial intelligence, i.e., "AI", is to perform tasks defined by humans in a way that meets or exceeds the capabilities of the humans performing them. Self - driving cars, music recommendation systems, and other sophisticated computing systems can be examples where AI has greatly contributed to successful examples that many humans like and enjoy. Such artificial intelligence systems can mimic the functions of the human brain and can often be configured to outperform humans in certain tasks, such as in aspects of face recognition or information extraction, to name a few. Artificial intelligence can be a computational model aimed at achieving results that humans can define as rewards (other examples include winning at the Jeopardy game or the Alpha Go game). Such systems cannot be "conscious" or "aware", and they can be described as pattern - matching machines.
[0004] A human-centered artificial intelligence system or configuration can include both a brain and a mind and can include a computational model that captures both. The mind can be the duality of the brain and the self-awareness involved. The mind can be synonymous with human thought, emotion, memory, and / or experience and can be the source of human behavior. By capturing this duality, the embodiments described herein can humanize AI using the subject system and its configuration. To borrow the words of famous chef Anthony Bourdain, a perfect meal occurs within context, leaves a memory, and this is often little to do with the food itself. The brain processes the food, and the mind is involved in the rest. The experiences that stay with the mind can be more desirable to the user and can be what remains in memory. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0005] Embodiments of the present disclosure describe systems and methods for generating and displaying a virtual companion. In an exemplary method, a first input from a user's environment is received at a first time via a first sensor on a head-wearable device. The occurrence of an event within the environment is determined based on the first input. A second input from the user is received via a second sensor on the head-wearable device, and the user's emotional reaction is identified based on the second input. An association is determined between the emotional reaction and the event. A view of the environment is presented at a second time later than the first time via a see-through display of the head-wearable device. A stimulus is presented at the second time via a virtual companion displayed via the see-through display, and the stimulus is determined based on the determined association between the emotional reaction and the event is determined. The present invention provides, for example, the following items. (Item 1) A method comprising: receiving, at a first time, a first input from a user's environment via a first sensor on a head-wearable device; and Determining the occurrence of an event in the environment based on the first input; Receiving a second input from the user via a second sensor on the head wearable device; Identifying an emotional reaction of the user based on the second input; Determining an association between the emotional reaction and the event; Presenting a view of the environment via a see-through display of the head wearable device at a second time after the first time; Presenting a stimulus via a virtual companion displayed via the see-through display at the second time, the stimulus being determined based on the determined association between the emotional reaction and the event; A method comprising. (Item 2) The method according to item 1, wherein the first input comprises an image of a physical object. (Item 3) The method according to item 1, wherein the first input comprises an audio signal. (Item 4) The method according to item 1, wherein the second input comprises speech from the user, and identifying the emotional reaction comprises determining content related to at least a portion of the speech. (Item 5) The method according to item 1, wherein the second input comprises eye movements of the user, and identifying the emotional reaction comprises determining a direction of gaze with respect to the user. (Item 6) The method according to item 1, wherein the second input comprises a field of view of the user, and identifying the emotional reaction comprises identifying at least one object within the field of view. (Item 7) The method according to item 1, further comprising determining an intensity of the emotional reaction, wherein the stimulus is further determined based on the intensity. (Item 8) The method according to item 1, wherein the association between the emotional reaction and the event is a temporal association. (Item 9) The method according to item 1, wherein the association between the emotional reaction and the event is a spatial association. (Item 10) The event is a first event, and the method further includes storing the association between the emotional reaction and the first event in a memory graph, the memory graph comprising an association between the first event and a second event. The method according to item 1. (Item 11) A system, A first sensor on a head-wearable device, A second sensor on the head-wearable device, A see-through display of the head-wearable device, One or more processors, the one or more processors being Receiving a first input from a user's environment via the first sensor on the head-wearable device at a first time; Determining the occurrence of an event in the environment based on the first input; Receiving a second input from the user via the second sensor on the head-wearable device; Identifying the user's emotional reaction based on the second input; Determining an association between the emotional reaction and the event; Presenting a view of the environment via the see-through display of the head-wearable device at a second time after the first time; Presenting a stimulus via a virtual companion displayed via the see-through display at the second time, the stimulus being determined based on the determined association between the emotional reaction and the event; One or more processors configured to execute a method including A system comprising (Item 12) The system according to item 11, wherein the first input comprises an image of a physical object. (Item 13) The system according to item 11, wherein the first input comprises an audio signal. (Item 14) The system according to item 11, wherein the second input comprises speech from the user, and determining the emotional reaction includes determining content related to at least a part of the speech. (Item 15) The system according to item 11, wherein the second input comprises eye movements of the user, and determining the emotional reaction includes determining a gaze direction regarding the user. (Item 16) The system according to item 11, wherein the second input comprises the user's visual field, and determining the emotional reaction includes identifying at least one object within the visual field. (Item 17) The system according to item 11, wherein the method further includes determining an intensity of the emotional reaction, and the stimulus is further determined based on the intensity. (Item 18) The system according to item 11, wherein the association between the emotional reaction and the event is a temporal association. (Item 19) The system according to item 11, wherein the association between the emotional reaction and the event is a spatial association. (Item 20) The event is a first event, and the method further includes storing the association between the emotional reaction and the first event in a memory graph, the memory graph comprising an association between the first event and a second event. (Item 21) A non-transitory computer-readable medium that stores instructions which, when executed by one or more processors, cause the one or more processors to receive a first input from a user's environment via a first sensor on a head-wearable device at a first time; determine the occurrence of an event in the environment based on the first input; receive a second input from the user via a second sensor on the head-wearable device; identify an emotional reaction of the user based on the second input; determine an association between the emotional reaction and the event; present a view of the environment via a see-through display of the head-wearable device at a second time that is after the first time; present a stimulus via a virtual companion displayed via the see-through display at the second time, the stimulus being determined based on the determined association between the emotional reaction and the event; A non-transitory computer-readable medium that causes execution of a method comprising the above. (Item 22) The non-transitory computer-readable medium according to item 21, wherein the first input comprises an image of a physical object. (Item 23) The non-transitory computer-readable medium according to item 21, wherein the first input comprises an audio signal. (Item 24) The non-transitory computer-readable medium according to item 21, wherein the second input comprises speech from the user, and identifying the emotional reaction comprises determining content related to at least a portion of the speech. (Item 25) The non-transitory computer-readable medium according to item 21, wherein the second input comprises eye movements of the user, and identifying the emotional reaction comprises determining a direction of gaze with respect to the user. (Item 26) The non-transitory computer-readable medium of item 21, wherein the second input comprises the user's field of view, and identifying the emotional reaction comprises identifying at least one object within the field of view. (Item 27) The non-transitory computer-readable medium of item 21, wherein the method further comprises determining the intensity of the emotional reaction, and the stimulus is further determined based on the intensity. (Item 28) The non-transitory computer-readable medium of item 21, wherein the association between the emotional reaction and the event is a temporal association. (Item 29) The non-transitory computer-readable medium of item 21, wherein the association between the emotional reaction and the event is a spatial association. (Item 30) The event is a first event, and the method further comprises storing the association between the emotional reaction and the first event in a memory graph, the memory graph comprising an association between the first event and a second event. The non-transitory computer-readable medium of item 21.
Brief Description of the Drawings
[0006]
Figure 1
[0007]
Figure 2A
Figure 2B
Figure 2C
[0008]
Figure 3A
Figure 3B
Figure 3C
Figure 3D
[0009]
Figure 4A
[0010]
Figure 4B
[0011]
Figure 5
[0012]
Figure 6
[0013]
Figure 7
[0014]
Figure 8
[0015]
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 9E
Figure 9F
Figure 9G
Figure 9H
Figure 9I
Figure 9J
[0016]
Figure 10A
Figure 10B
[0017] When addressing the challenges of human-centered AI, many problems and variations that need to be addressed can exist. For example, what is the best experience for a particular person or group of people? Regarding this, some immediate answers exist based on typically available systems (such as those described in the aforementioned incorporated applications, or other available computing systems configured for human interaction, etc.) and the human use of such systems, i.e., use related to screens such as watching TV on a 2D monitor or conventional computing, playing games, web browsing, etc. These experiences are expected on any modern computing platform, including systems such as those illustrated in Figure 1. In systems such as those illustrated in Figure 1, there are systems that collect a lot of information about the surrounding world, but also, perhaps more importantly, such systems may be configured to collect a lot of information about the user. The user can be at the center of many mixed reality experiences, and the world can serve as the backdrop for these experiences. Some of the most appealing mixed reality experiences can be those where the content presented to the user can be "intelligent" and context-based. In other words, in such a configuration, there can be cause and effect where the user does something and the experience responds to that input. The "intelligence" in the experience can arise not only from the experience itself but also from the platform. For example, having a certain degree of information / knowledge at the system level regarding where a person is present in the environment, what surrounds the person, or about the person can be very useful. This system may also be configured to infer or recall information regarding the user's emotions and related associations. This system may be configured to collect information related to the person speaking and the content being spoken. These can be fundamental problems of the challenges of human-centered AI. One important question to answer when developing an experience can be, "What needs are we trying to fulfill?" Various answers can include entertainment, communication, understanding of information or knowledge. These needs can each be caused by perception, emotion, and thought.There are many examples where AI behaves very "mechanically". Many can ingest large amounts of data and, based on that data, create resulting models. In many cases, humans do not properly understand how this functions. Conversely, computers generally do not "understand" how humans function. Generally, AI systems can be configured to efficiently reach accurate answers based on the data they are trained on. One challenge can be to perform machine learning in conjunction with the rich outputs of computing systems and use them to meet human needs in a mixed reality experience. It may be desirable to do so in a way that the AI becomes invisible or integrated into the computing engagement. Thus, the goal is to design a system that can be easily understood by people or, better yet, become transparent to people (e.g., fully integrated into the user's experience so that the user is not explicitly aware of the system's presence) and generally focus on providing a better experience.
[0018] It is desirable for a mixed reality or augmented reality head-mounted display to be lightweight, low-cost, have a small form factor, have a wide virtual image field of view, and be as transparent as possible. Additionally, in some embodiments, it is desirable to have a configuration that presents virtual image information at multiple focal planes (e.g., two or more) to be practical for a wide variety of use cases without exceeding an acceptable range regarding convergence / divergence motion / accommodation mismatch. Referring to FIG. 1, an augmented reality system is illustrated that features a head-mounted viewing component (2), a hand-held controller component (4), and an interconnected auxiliary computing or controller component (6) that can be configured to be worn as a belt pack or equivalent on the user. These components are each operatively coupled to one another (10, 12, 14, 16, 17, 18) and the IEEE It may be coupled to other connected resources (8) such as cloud computing or cloud storage resources via a wired or wireless communication configuration as defined by 802.11, Bluetooth® (RTM), and other connectivity standards and configurations. For example, as described in U.S. Patent Application Nos. 14 / 555,585, 14 / 690,401, 14 / 331,218, 15 / 481,255, and 62 / 518,539 (each of which is incorporated herein by reference in its entirety), through which a user can see the surrounding world along with visual components that may be generated by associated system components for an augmented reality experience, various embodiments of the two depicted optical elements (20), etc., various aspects of such components are described. There is a need for very high-performance systems and assemblies optimized for use in wearable computing systems. In certain embodiments, such systems and subsystems may be configured and utilized for certain "artificial intelligence" related tasks.
[0019] The various components can be used in connection with providing a user with an augmented or mixed reality experience. For example, as shown in FIG. 1, a see-through wearable display system may be able to present a user with a combination of a view of the physical world around the user along with combined virtual content within the user's field of view in a perceptually meaningful way. Using the same system (i.e., such as that depicted in FIG. 1), a spatial computing platform can be used to receive or "sense" information regarding various physical aspects related to the environment and viewer simultaneously. By combining a wearable computing display with a machine learning-powered spatial computing platform, a feedback loop can be established between the user and the experience.
[0020] Mixed reality environment
[0021] Like all people, users of a mixed reality system exist within the physical environment, i.e., the three-dimensional portion of the "real world" and all of its contents that are perceivable by the user. For example, the user perceives the physical environment using their normal human senses, i.e., vision, hearing, touch, taste, and smell, and interacts with the physical environment by moving their own body within the physical environment. Locations within the physical environment can be described as coordinates within a coordinate space, for example, the coordinates can include latitude, longitude, and altitude relative to sea level, distances in three orthogonal dimensions from a reference point, or other suitable values. Similarly, a vector can describe a quantity having a direction and magnitude within a coordinate space.
[0022] A computing device can maintain a representation of a virtual environment, for example, within a memory associated with the device. As used herein, a virtual environment is a computer representation of a three-dimensional space. The virtual environment can include a representation of any object, action, signal, parameter, coordinate, vector, or other characteristic associated with that space. In some embodiments, the circuitry (e.g., a processor) of the computing device can maintain and update the state of the virtual environment, i.e., the processor can determine the state of the virtual environment at a second time t1 based on data associated with the virtual environment and / or inputs provided by a user at a first time t0. For example, if an object within the virtual environment is located at a first coordinate at time t0, has certain programmed physical parameters (e.g., mass, coefficient of friction), and an input received from the user indicates that a force should be applied to the object in a certain direction vector, the processor can apply the laws of kinematics and use basic mechanics to determine the location of the object at time t1. The processor can use any suitable information known about the virtual environment and / or any suitable inputs to determine the state of the virtual environment at time t1. When maintaining and updating the state of the virtual environment, the processor can execute any suitable software, including software related to the creation and deletion of virtual objects within the virtual environment, software (e.g., a script) for defining the behavior of virtual objects or characters within the virtual environment, software for defining the behavior of signals (e.g., audio signals) within the virtual environment, software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling inputs and outputs, software for implementing network operations, software for applying asset data (e.g., animation data for moving virtual objects over time), or many other possibilities.
[0023] Output devices such as displays or speakers can present any or all aspects of the virtual environment to the user. For example, the virtual environment may include virtual objects (which may include representations of inanimate objects, people, animals, light, etc.) that can be presented to the user. The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" with origin coordinates, viewing axis, and frustum), and render on the display a visible scene of the virtual environment corresponding to that view. Any suitable rendering technique may be used for this purpose. In some embodiments, the visible scene may include only some of the virtual objects within the virtual environment and may exclude other virtual objects. Similarly, the virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects within the virtual environment may generate sounds arising from the location coordinates of the objects (e.g., a virtual character may speak or cause sound effects), or the virtual environment may be associated with musical cues or ambient sounds that may or may not be associated with a particular location. The processor can determine an audio signal corresponding to the "listener" coordinates, e.g., an audio signal that is mixed and processed to simulate an audio signal corresponding to a composite of the sounds within the virtual environment and that would be audible to a listener at the listener coordinates, and present the audio signal to the user via one or more speakers.
[0024] Since a virtual environment exists only as a computer construct, a user cannot directly perceive the virtual environment using their normal senses. Instead, a user can only indirectly perceive the virtual environment as presented to the user, for example, by way of a display, speakers, a tactile output device, and the like. Similarly, a user cannot directly touch, manipulate, or otherwise interact with the virtual environment, but can provide input data to a processor that can use the device or sensor data to update the virtual environment via an input device or sensor. For example, a camera sensor can provide optical data indicating that a user is attempting to move an object within the virtual environment, and the processor can use that data to cause the object to respond accordingly within the virtual environment.
[0025] A mixed reality system can present a user with a mixed reality environment ("MRE") that combines aspects of the real environment and the virtual environment, for example, using a see-through display and / or one or more speakers (e.g., that can be incorporated into a wearable head device). In some embodiments, one or more speakers may be external to the head-mounted wearable unit. As used herein, an MRE is a simultaneous representation of the real environment and the corresponding virtual environment. In some examples, the corresponding real and virtual environments share a single coordinate space, and in some examples, the real coordinate space and the corresponding virtual coordinate space are related to each other by a transformation matrix (or other suitable representation). Thus, a single coordinate can (in some examples, together with a transformation matrix) define a first location within the real environment and also a second corresponding location within the virtual environment, and vice versa.
[0026] In MRE, a virtual object (e.g., within a virtual environment associated with MRE) can correspond to a real object (e.g., within a real environment associated with MRE). For example, if the real environment of MRE includes a real street lamp post (real object) at a certain location coordinate, the virtual environment of MRE may include a virtual street lamp post (e.g., virtual object) at the corresponding location coordinate. As used herein, the real object in combination with its corresponding virtual object together constitute a "composite reality object". It is not necessary for the virtual object to exactly match or align with the corresponding real object. In some embodiments, the virtual object can be a simplified version of the corresponding real object. For example, if the real environment includes a real street lamp post, the corresponding virtual object may include a cylinder with approximately the same height and radius as the real street lamp post (reflecting that the street lamp post can be approximately cylindrical in shape). Simplifying the virtual object in this way can enable calculation efficiency improvement and simplify the calculations to be performed on such virtual objects. Further, in some embodiments of MRE, not all real objects in the real environment may be associated with corresponding virtual objects. Similarly, in some embodiments of MRE, not all virtual objects in the virtual environment may be associated with corresponding real objects. That is, some virtual objects may be alone within the virtual environment of MRE without any real-world counterpart.
[0027] In some embodiments, a virtual object may have characteristics that are sometimes dramatically different from those of the corresponding real object. For example, while the real environment within the MRE may comprise a cactus with two green branches, i.e., an inanimate object with needles, the corresponding virtual object within the MRE may have the characteristics of a virtual character with two green branches with human facial features and a grumpy expression. In this embodiment, the virtual object is similar to its corresponding real object in some characteristics (color, number of branches), but different from the real object in other characteristics (facial features, personality). Thus, virtual objects have the potential to represent real objects in a creative, abstract, exaggerated, or fantastical manner, or to endow inanimate real objects with behavior (e.g., human personality) that they would not otherwise have. In some embodiments, a virtual object may be a purely imaginary creation without any real-world counterpart (e.g., perhaps a virtual monster within the virtual environment at a location corresponding to an empty space within the real environment).
[0028] Compared to a VR system that presents a virtual environment while obscuring the real environment, a mixed reality system that presents MRE offers the advantage that the real environment remains perceivable while the virtual environment is presented. Thus, a user of a mixed reality system can use visual and audio cues associated with the real environment to experience and interact with the corresponding virtual environment. As an example, a user of a VR system may struggle to perceive or interact with virtual objects displayed within the virtual environment as, as described above, the user cannot directly perceive or interact with the virtual environment, whereas a user of an MR system may find it intuitive and natural to interact with virtual objects by looking at, listening to, and touching the corresponding real objects within their own real environment. This level of bidirectionality can enhance the user's sense of immersion, connection, and engagement with the virtual environment. Similarly, by presenting the real and virtual environments simultaneously, a mixed reality system can reduce negative psychological sensations (e.g., cognitive dissonance) and negative physical sensations (e.g., motion sickness) associated with VR systems. The mixed reality system further offers many possibilities for applications that can extend or modify real-world experiences.
[0029] Figure 2A illustrates an exemplary real-world environment 200 in which a user 210 uses a mixed reality system 212. The mixed reality system 212 may include, for example, a display (e.g., a transparent display), one or more speakers, and one or more sensors (e.g., cameras), as described below. The illustrated real-world environment 200 includes a rectangular room 204A in which the user 210 stands, and real objects 222A (lamp), 224A (table), 226A (sofa), and 228A (painting). The room 204A further includes location coordinates 206 that may be considered the origin of the real-world environment 200. As shown in Figure 2A, an environment / world coordinate system 208 (including an x-axis 208X, a y-axis 208Y, and a z-axis 208Z) with its origin at point 206 (world coordinates) can define a coordinate space for the real-world environment 200. In some embodiments, the origin 206 of the environment / world coordinate system 208 may correspond to the location where the mixed reality system 212 is powered on. In some embodiments, the origin 206 of the environment / world coordinate system 208 may be reset during operation. In some examples, the user 210 may be considered a real object within the real-world environment 200, and similarly, body parts of the user 210 (e.g., hands, feet) may also be considered real objects within the real-world environment 200. In some examples, a user / listener / head coordinate system 214 (including an x-axis 214X, a y-axis 214Y, and a z-axis 214Z) with its origin at point 215 (e.g., user / listener / head coordinates) can define a coordinate space for the user / listener / head where the mixed reality system 212 is located. The origin 215 of the user / listener / head coordinate system 214 may be defined relative to one or more components of the mixed reality system 212. For example, the origin 215 of the user / listener / head coordinate system 214 may be defined relative to the display of the mixed reality system 212, such as during initial calibration of the mixed reality system 212. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the user / listener / head coordinate system 214 space and the environment / world coordinate system 208 space. In some embodiments, left ear coordinates 216 and right ear coordinates 217 may be defined relative to the origin 215 of the user / listener / head coordinate system 214.A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the left ear coordinates 216 and the right ear coordinates 217 and the user / listener / head coordinate system 214 space. The user / listener / head coordinate system 214 can simplify the representation of the location of the user's head or the head-mounted device with respect to, for example, the environment / world coordinate system 208. Using simultaneous localization and mapping (SLAM), visual odometry, or other techniques, the transformation between the user coordinate system 214 and the environment coordinate system 208 can be determined and updated in real time.
[0030] FIG. 2B illustrates an exemplary virtual environment 230 corresponding to the real environment 200. The illustrated virtual environment 230 includes a virtual rectangular room 204B corresponding to the real rectangular room 204A, a virtual object 222B corresponding to the real object 222A, a virtual object 224B corresponding to the real object 224A, and a virtual object 226B corresponding to the real object 226A. The metadata associated with the virtual objects 222B, 224B, 226B can include information derived from the corresponding real objects 222A, 224A, 226A. The virtual environment 230 further includes a virtual monster 232 that does not correspond to any real object within the real environment 200. The real object 228A within the real environment 200 does not correspond to any virtual object within the virtual environment 230. A persistent coordinate system 233 (comprising an x-axis 233X, a y-axis 233Y, and a z-axis 233Z) with its origin at point 234 (persistent coordinates) can define a coordinate space for the virtual content. The origin 234 of the persistent coordinate system 233 may be defined with respect to one or more real objects such as the real object 226A. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the persistent coordinate system 233 space and the environment / world coordinate system 208 space. In some embodiments, the virtual objects 222B, 224B, 226B, and 232 may each have their own persistent coordinate points with respect to the origin 234 of the persistent coordinate system 233. In some embodiments, multiple persistent coordinate systems may exist, and the virtual objects 222B, 224B, 226B, and 232 may each have their own persistent coordinate points with respect to one or more of the persistent coordinate systems.
[0031] With respect to FIGS. 2A and 2B, the environment / world coordinate system 208 defines a shared coordinate space for both the real environment 200 and the virtual environment 230. In the illustrated embodiment, the coordinate space has its origin at point 206. Further, the coordinate space is defined by the same three orthogonal axes (208X, 208Y, 208Z). Thus, a first location within the real environment 200 and a second corresponding location within the virtual environment 230 can be described with respect to the same coordinate space. This simplifies identifying and displaying corresponding locations within the real and virtual environments since the same coordinates can be used to identify both locations. However, in some embodiments, the corresponding real and virtual environments need not use a shared coordinate space. For example, in some embodiments (not shown), a matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the real environment coordinate space and the virtual environment coordinate space.
[0032] FIG. 2C illustrates an exemplary MRE 250 that simultaneously presents aspects of the real environment 200 and the virtual environment 230 to the user 210 via the composite reality system 212. In the illustrated embodiment, the MRE 250 simultaneously presents real objects 222A, 224A, 226A, and 228A from the real environment 200 to the user 210 (e.g., via the transmissive portion of the display of the composite reality system 212) and virtual objects 222B, 224B, 226B, and 232 from the virtual environment 230 (e.g., via the active display portion of the display of the composite reality system 212). As described above, the origin 206 acts as the origin with respect to the coordinate space corresponding to the MRE 250, and the coordinate system 208 defines the x-axis, y-axis, and z-axis with respect to the coordinate space.
[0033] In the illustrated embodiments, the composite reality object comprises corresponding pairs of real and virtual objects (i.e., 222A / 222B, 224A / 224B, 226A / 226B) that occupy corresponding locations within the coordinate space 208. In some embodiments, both the real object and the virtual object may be simultaneously visible to the user 210. This may be desirable, for example, in cases where the virtual object presents information designed to augment the view of the corresponding real object (such as in museum applications where the virtual object presents missing portions of an ancient damaged sculpture). In some embodiments, the virtual objects (222B, 224B, and / or 226B) may be displayed so as to occlude the corresponding real objects (222A, 224A, and / or 226A) (e.g., via active pixelization occlusion using a pixelization occlusion shutter). This may be desirable, for example, in cases where the virtual object acts as a visual replacement for the corresponding real object (such as in two-way storytelling applications where an inanimate real object becomes a "living" character).
[0034] In some embodiments, the real objects (e.g., 222A, 224A, 226A) may be associated with virtual content or helper data that does not necessarily constitute the virtual object. The virtual content or helper data can facilitate the processing or handling of the virtual object within the composite reality environment. For example, such virtual content may include a two-dimensional representation of the corresponding real object, a custom asset type associated with the corresponding real object, or statistical data associated with the corresponding real object. This information can enable or facilitate calculations involving the real object without introducing unnecessary computational overhead.
[0035] In some embodiments, the presentation described above may also incorporate an audio aspect. For example, in the MRE250, the virtual monster 232 may be associated with one or more audio signals such as the sound effects of footsteps generated when the monster walks around the MRE250. As further described below, the processor of the mixed reality system 212 may calculate an audio signal corresponding to the mixture of all such sounds in the MRE250 and the processed composite, and present the audio signal to the user 210 via one or more speakers included within the mixed reality system 212 and / or one or more external speakers.
[0036] Exemplary Mixed Reality System
[0037] The exemplary mixed reality system 212 can include a display (which can include left and right transmissive displays, which can be near-eye displays, and associated components for coupling light from the display to the user's eyes), left and right speakers (e.g., positioned adjacent to the user's left and right ears, respectively), an inertial measurement unit (IMU) (e.g., mounted on the stem of the head device), an orthogonal coil electromagnetic receiver (e.g., mounted on the left stem portion), left and right cameras (e.g., depth (time-of-flight) cameras) oriented away from the user, and left and right eye cameras (e.g., for detecting the user's eye movements) oriented towards the user, in a wearable head device (e.g., a wearable augmented reality or mixed reality head device). However, the mixed reality system 212 can incorporate any suitable display technology and any suitable sensors (e.g., optical, infrared, acoustic, LIDAR, EOG, GPS, magnetic). Additionally, the mixed reality system 212 can incorporate networking features (e.g., Wi-Fi capabilities) to communicate with other devices and systems, including other mixed reality systems. The mixed reality system 212 can further include a battery (which can be mounted within an auxiliary unit such as a belt pack designed to be worn around the user's waist), a processor, and a memory. The wearable head device of the mixed reality system 212 can include a tracking component, such as an IMU or other suitable sensor, configured to output a set of coordinates of the wearable head device relative to the user's environment. In some embodiments, the tracking component can provide an input to the processor and perform simultaneous localization and mapping (SLAM) and / or visual odometry algorithms. In some embodiments, the mixed reality system 212 can also include a handheld controller 400 and / or an auxiliary unit 420, which can be a wearable belt pack, as further described below.
[0038] Figures 3A-3D illustrate components of an exemplary mixed reality system 300 (which may correspond to the composite reality system 212) that may be used to present to a user an MRE (which may correspond to MRE250) or other virtual environment. FIG. 3A illustrates a perspective view of a wearable head device 2102 included within the exemplary mixed reality system 300. FIG. 3B illustrates a top view of the wearable head device 2102 worn on a user's head 2202. FIG. 3C illustrates a front view of the wearable head device 2102. FIG. 3D illustrates an end view of an exemplary eyepiece 2110 of the wearable head device 2102. As shown in FIGS. 3A-3C, the exemplary wearable head device 2102 includes an exemplary left eyepiece (e.g., a left transparent waveguide set eyepiece) 2108 and an exemplary right eyepiece (e.g., a right transparent waveguide set eyepiece) 2110. Each eyepiece 2108 and 2110 can include a transmissive element through which a real environment can be visible, and a display element for presenting an overlaying display to the real environment (e.g., via light modulated for each image). In some embodiments, such a display element can include a surface diffractive optical element for controlling the flow of light modulated for each image. For example, the left eyepiece 2108 can include a left internal coupling grating set 2112, a left orthogonal pupil expansion (OPE) grating set 2120, and a left exit (output) pupil expansion (EPE) grating set 2122. Similarly, the right eyepiece 2110 can include a right internal coupling grating set 2118, a right OPE grating set 2114, and a right EPE grating set 2116. Light modulated for each image can be transmitted to the user's eyes through the internal coupling gratings 2112 and 2118, the OPEs 2114 and 2120, and the EPEs 2116 and 2122. Each internal coupling grating set 2112, 2118 can be configured to deflect light toward its corresponding OPE grating set 2120, 2114. Each OPE grating set 2120, 2114 is designed to incrementally deflect light downward toward its associated EPE 2122, 2116, thereby enabling the formed exit pupil to extend horizontally.Each of the EPEs 2122, 2116 is configured to incrementally redirect at least a portion of the light received from its corresponding OPE grating set 2120, 2114 outwardly to a user eye box position (not shown) defined behind the eyepieces 2108, 2110, and can vertically extend the exit pupil formed at the eye box. Alternatively, instead of the internal coupling grating sets 2112 and 2118, the OPE grating sets 2114 and 2120, and the EPE grating sets 2116 and 2122, the eyepieces 2108 and 2110 can include other arrangements of gratings and / or refractive and reflective features for controlling the coupling of light modulated for each image to the user's eyes.
[0039] In some embodiments, the wearable head device 2102 can include a left arm 2130 and a right arm 2132, the left arm 2130 including a left speaker 2134 and the right arm 2132 including a right speaker 2136. The orthogonal coil electromagnetic receiver 2138 can be located within the left arm portion or at another suitable location within the wearable head unit 2102. The inertial measurement unit (IMU) 2140 can be located within the right arm 2132 or at another suitable location within the wearable head device 2102. The wearable head device 2102 can also include a left depth (e.g., time-of-flight) camera 2142 and a right depth camera 2144. The depth cameras 2142, 2144 can be suitably oriented in different directions so as to together cover a wider field of view.
[0040] In the embodiments shown in FIGS. 3A - 3D, the left source 2124 of light modulated for each image can be optically coupled to the left eyepiece lens 2108 through the left internal coupling grating set 2112, and the right source 2126 of light modulated for each image can be optically coupled to the right eyepiece lens 2110 through the right internal coupling grating set 2118. The sources 2124, 2126 of light modulated for each image can include, for example, a projector including an electro - optical modulator such as an optical fiber scanning device, a digital light processing (DLP) chip, or a liquid crystal on silicon (LCoS) modulator, or a light - emitting display such as a micro - light - emitting diode (μLED) or a micro - organic light - emitting diode (μOLED) panel coupled to the internal coupling grating sets 2112, 2118 using one or more lenses per side. The input coupling grating sets 2112, 2118 can deflect the light from the sources 2124, 2126 of light modulated for each image at an angle exceeding the critical angle for total internal reflection (TIR) with respect to the eyepiece lenses 2108, 2110. The OPE grating sets 2114, 2120 incrementally deflect downward the light propagating by TIR towards the EPE grating sets 2116, 2122. The EPE grating sets 2116, 2122 incrementally couple the light towards the user's face including the pupil of the user's eye.
[0041] In some embodiments, as shown in FIG. 3D, the left eyepiece lens 2108 and the right eyepiece lens 2110 each include a plurality of waveguides 2402. For example, each eyepiece lens 2108, 2110 can include a plurality of individual waveguides dedicated to individual color channels (e.g., red, blue, and green), respectively. In some embodiments, each eyepiece lens 2108, 2110 can include a plurality of sets of such waveguides, and each set is configured to impart a different wavefront curvature to the emitted light. The wavefront curvature can be convex with respect to the user's eye, for example, to present a virtual object positioned at a certain distance in front of the user (e.g., by a distance corresponding to the reciprocal of the wavefront curvature). In some embodiments, the EPE grating sets 2116, 2122 can include curved grating grooves to provide a convex wavefront curvature by modifying the pointing vector of the emitted light across each EPE.
[0042] In some embodiments, stereoscopically adjusted left and right eye images can be presented to the user through the image-by-image light modulators 2124, 2126 and the eyepiece lenses 2108, 2110 to create the perception that the displayed content is three-dimensional. The perceived realism of the presentation of the three-dimensional virtual object can be improved by selecting the waveguides (and thus the corresponding wavefront curvature) such that the virtual object is displayed at a distance approximating the distance indicated by the stereoscopic left and right images. The technique can also reduce the motion sickness experienced by some users, which can be caused by the difference between the depth perception cues provided by the stereoscopic left and right eye images and the natural accommodation of the human eye (e.g., object distance-dependent focusing).
[0043] FIG. 3D illustrates an end view from above the right eyepiece lens 2110 of the exemplary wearable head device 2102. As shown in FIG. 3D, the plurality of waveguides 2402 can include a first subset 2404 of three waveguides and a second subset 2406 of three waveguides. The two subsets 2404, 2406 of waveguides can be distinguished by different EPE gratings characterized by different grating line curvatures to impart different wavefront curvatures to the outgoing light. Within each of the subsets 2404, 2406 of waveguides, each waveguide can be used to couple a different spectral channel (e.g., one of the red, green, and blue spectral channels) to the user's right eye 2206. (Although not shown in FIG. 3D, the structure of the left eyepiece lens 2108 is similar to the structure of the right eyepiece lens 2110.)
[0044] FIG. 4A illustrates an exemplary hand-held controller component 400 of the mixed reality system 300. In some embodiments, the hand-held controller 400 includes a grip portion 446 and one or more buttons 450 disposed along an upper surface 448. In some embodiments, the buttons 450 may be configured for use as optical tracking targets, for example, in conjunction with a camera or other optical sensor (which may be mounted within a head unit of the mixed reality system 300, such as the wearable head device 2102), to track the six degrees of freedom (6DOF) movement of the hand-held controller 400. In some embodiments, the hand-held controller 400 includes a tracking component (such as an IMU or other suitable sensor) for detecting a position or orientation, such as a position or orientation relative to the wearable head device 2102. In some embodiments, such a tracking component may be positioned within the handle of the hand-held controller 400 and / or mechanically coupled to the hand-held controller. The hand-held controller 400 can be configured to provide one or more output signals corresponding to one or more of the button depressed state, or the position, orientation, and / or movement of the hand-held controller 400 (e.g., via an IMU). Such output signals may be used as an input to a processor of the mixed reality system 300. Such an input may correspond to the position, orientation, and / or movement of the hand-held controller (and in turn, the position, orientation, and / or movement of the hand of the user holding the controller). Such an input may also correspond to the user depressing the button 450.
[0045] FIG. 4B illustrates an exemplary auxiliary unit 420 of the mixed reality system 300. The auxiliary unit 420 can include a battery to provide energy for operating the system 300, and can include a processor for executing a program for operating the system 300. As shown, the exemplary auxiliary unit 420 includes a clip 2128, such as for attaching the auxiliary unit 420 to a user's belt. Other form factors, including form factors that do not involve mounting the unit on a user's belt, will also be suitable and will be apparent for the auxiliary unit 420. In some embodiments, the auxiliary unit 420 is coupled to the wearable head device 2102 through a multi-conduit cable that can include, for example, electrical wires and optical fibers. A wireless connection between the auxiliary unit 420 and the wearable head device 2102 can also be used.
[0046] In some embodiments, the mixed reality system 300 can include one or more microphones to detect sound and provide a corresponding signal to the mixed reality system. In some embodiments, the microphone may be attached to or integrated with the wearable head device 2102 and may be configured to detect the user's voice. In some embodiments, the microphone may be attached to or integrated with the handheld controller 400 and / or the auxiliary unit 420. Such a microphone may be configured to detect ambient sound, ambient noise, the voice of the user or a third party, or other sounds.
[0047] FIG. 5 shows an exemplary functional block diagram of an exemplary mixed reality system 300 or the like (which may correspond to the mixed reality system 212 related to FIG. 2A) that may correspond to the mixed reality systems described above. As shown in FIG. 5, an exemplary handheld controller 500B (which may correspond to the handheld controller 400 (“totem”)) includes a totem / wearable head device 6 degrees of freedom (6DOF) totem subsystem 504A, and an exemplary wearable head device 500A (which may correspond to the wearable head device 2102) includes a totem / wearable head device 6DOF subsystem 504B. In an embodiment, the 6DOF totem subsystem 504A and the 6DOF subsystem 504B cooperate to determine six coordinates of the handheld controller 500B relative to the wearable head device 500A (e.g., offsets in three translational directions and rotations along three axes). The six degrees of freedom may be represented relative to the coordinate system of the wearable head device 500A. The three translational offsets may be represented as X, Y, and Z offsets, a translation matrix, or some other representation within such a coordinate system. The rotational degrees of freedom may be represented as a sequence of yaw, pitch, and roll rotations, a rotation matrix, a quaternion, or some other representation. In some embodiments, the wearable head device 500A, one or more depth cameras 544 (and / or one or more non-depth cameras) included within the wearable head device 500A, and / or one or more optical targets (e.g., buttons 450 of the handheld controller 500B as described above or dedicated optical targets included within the handheld controller 500B) may be used for 6DOF tracking. In some embodiments, the handheld controller 500B may include a camera as described above, and the wearable head device 500A may include an optical target for optical tracking in conjunction with the camera. In some embodiments, the wearable head device 500A and the handheld controller 500B each include a set of three orthogonally oriented solenoids, which are used to wirelessly transmit and receive three distinguishable signals.The 6DOF of the wearable head device 500A relative to the handheld controller 500B may be determined by measuring the relative magnitudes of three distinguishable signals received within each of the coils for reception. Additionally, the 6DOF totem subsystem 504A can include an inertial measurement unit (IMU) that is useful for providing improved accuracy and / or more timely information regarding rapid movement of the handheld controller 500B.
[0048] In some embodiments, for example, in order to compensate for the movement of the wearable head device 500A relative to the coordinate system 208, it may be necessary to convert coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the wearable head device 500A) to an inertial coordinate space (e.g., a coordinate space fixed relative to the real environment). For example, such a conversion is necessary to ensure that the display of the wearable head device 500A presents virtual objects at the expected positions and orientations relative to the real environment (e.g., a virtual person sitting on a real chair facing forward, regardless of the position and orientation of the wearable head device), rather than at fixed positions and orientations on the display (e.g., at the same position in the lower right corner of the display), and that the virtual objects appear to exist within the real environment (and, for example, do not appear unnaturally positioned within the real environment as the wearable head device 500A shifts and rotates). In some embodiments, the compensation transformation between coordinate spaces can be determined by processing images from the depth camera 544 using SLAM and / or visual odometry procedures to determine the transformation of the wearable head device 500A relative to the coordinate system 208. In the embodiment shown in FIG. 5, the depth camera 544 is coupled to the SLAM / visual odometry block 506 and can provide the images to the block 506. The SLAM / visual odometry block 506 implementation can include a processor configured to process the images and then determine the position and orientation of the user's head, which can be used to identify the transformation between the head coordinate space and another coordinate space (e.g., the inertial coordinate space). Similarly, in some embodiments, an additional source of information regarding the user's head pose and location is obtained from the IMU 509. The information from the IMU 509 is integrated with the information from the SLAM / visual odometry block 506 to provide improved accuracy and / or more timely information regarding the high-speed adjustment of the user's head pose and position.
[0049] In some embodiments, the depth camera 544 can supply a 3D image to a hand gesture tracker 511 that can be implemented within the processor of the wearable head device 500A. The hand gesture tracker 511 can identify a user's hand gesture, for example, by matching the 3D image received from the depth camera 544 to a stored pattern representing the hand gesture. Other suitable techniques for identifying a user's hand gesture will also be apparent.
[0050] In some embodiments, one or more processors 516 may be configured to receive data from the 6DOF headgear subsystem 504B of the wearable head device, the IMU 509, the SLAM / visual odometry block 506, the depth camera 544, and / or the hand gesture tracker 511. The processor 516 may also be able to send and receive control signals to and from the 6DOF totem system 504A. The processor 516 may be wirelessly coupled to the 6DOF totem system 504A, such as in embodiments where the handheld controller 500B is not tethered. The processor 516 may further communicate with additional components such as the audiovisual content memory 518, the graphics processing unit (GPU) 520, and / or the digital signal processor (DSP) audio spatializer 522. The DSP audio spatializer 522 may be coupled to the head-related transfer function (HRTF) memory 525. The GPU 520 may include a left channel output coupled to the left source 524 of light modulated per image and a right channel output coupled to the right source 526 of light modulated per image. The GPU 520 may output stereoscopic image data to the sources 524, 526 of light modulated per image, as described above with respect to FIGS. 3A-3D, for example. The DSP audio spatializer 522 may output audio to the left speaker 512 and / or the right speaker 514. The DSP audio spatializer 522 may receive an input from the processor 519 indicating a direction vector from the user to a virtual sound source (e.g., that can be moved by the user via the handheld controller 420). Based on the direction vector, the DSP audio spatializer 522 may be able to determine the corresponding HRTF (e.g., by accessing the HRTF or by interpolating multiple HRTFs). The DSP audio spatializer 522 may then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the credibility and realism of virtual sounds by incorporating the user's relative position and orientation with respect to the virtual sounds within the composite reality environment, i.e., by presenting virtual sounds that match the user's expectations of what the virtual sounds would sound like if they were real sounds within the real environment.
[0051] In some embodiments, such as those shown in FIG. 5, one or more of the processor 516, GPU 520, DSP audio spatializer 522, HRTF memory 525, and audio / visual content memory 518 may be included within an auxiliary unit 500C (which may correspond to the auxiliary unit 420 described above). The auxiliary unit 500C may include a battery 527 to power its components and / or supply power to the wearable head device 500A or the handheld controller 500B. Including such components within an auxiliary unit that can be mounted on the user's waist can limit the size and weight of the wearable head device 500A, which can, in turn, reduce fatigue in the user's head and neck.
[0052] FIG. 5 presents elements corresponding to various components of an exemplary composite reality system, but various other suitable arrangements of these components will be apparent to those skilled in the art. For example, the elements presented in FIG. 5 as being associated with the auxiliary unit 500C may instead be associated with the wearable head device 500A or the handheld controller 500B. Further, some composite reality systems may eliminate the handheld controller 500B or the auxiliary unit 500C entirely. Such changes and modifications are understood to be within the scope of the disclosed embodiments.
[0053] Human-centered AI
[0054] To extend and interact with the real world at a deeper and more personal level, the user can disclose to the platform (which may correspond to the MR systems 212, 300) data about the environment and themselves. In various embodiments, the user owns their data, but at least one important use may be to improve the user's experience using the system, and thus the system may be configured to allow the user to control who and when has access to this information and to allow the user to share both their virtual and physical data.
[0055] Referring to FIG. 6, in various embodiments, a human-centered AI configuration for wearable computing may be built on three fundamental pillars, namely, the user 602, the AI companion 604 for the user, and the environment or space 606 around the user (which may be an MRE and may include both the user's physical environment and the user's virtual environment). Since the AI system of the subject matter and its configuration can be human-centered, the user can be the main focus for such a configuration. In various embodiments, the user can be characterized by their behavior, emotions, preferences, social graph, temperament, and physical attributes. In one embodiment, the virtual AI companion may be characterized by a similar set of attributes in order to make it more "human-like". These may include personality, memory, knowledge, state, actions, and the ability to interact with humans and machines (which may be called "Oz" and may be associated with the concept of a "possible world" or a part thereof as described in the incorporated references mentioned above). Further, similar to the state of "According to the theory of general relativity, a space without ether is unthinkable", the interaction between the user and the AI may be something that is unthinkable without an environment. The environment may be used to determine context and provide the boundaries of the experience. The environment around the user may be parameterized by 3D reconstruction and scene understanding, and understanding of humans and their interactions. The interaction between these three aforementioned pillars facilitates human-centered AI as a platform.
[0056] Memory graph
[0057] FIG. 7 depicts an exemplary system 700 for creating an exemplary memory graph 701. The memory graph 701 can comprise one or more nodes 716 that can have one or more associations with other nodes. In some embodiments, the memory graph 701 can represent all information about a user captured by an MR system (e.g., MR systems 212, 300). In some embodiments, the memory graph 701 can receive inputs from at least three sources, namely, an environmental observation module 702, a user observation module 708, and an external resource 714.
[0058] The exemplary environmental observation module 702 can receive one or more sensor inputs 704a - 704n. The sensor inputs 704a - 704n can comprise inputs related to SLAM. SLAM can be used by an MR system (e.g., MR systems 212, 300) to identify physical features in a physical environment and to locate those physical features relative to the physical environment and to each other. At the same time, the MR system (e.g., MR systems 212, 300) can locate itself within the physical environment and relative to the physical features. SLAM can build an understanding of the user's physical environment, which can enable the MR system (e.g., MR systems 212, 300) to create a virtual environment that respects and interacts with the user's physical environment. For example, it may be desirable for the MR system (e.g., MR systems 212, 300) to identify the physical floor of the user's physical environment and to display a virtual human avatar as if standing on the physical floor in order to display a virtual AI companion near the user. In some embodiments, as the user walks around the room, it may be desirable for the virtual human avatar to be able to move with the user (like a physical companion) and for the virtual human avatar to recognize physical obstacles (e.g., a table) so that the virtual human avatar does not appear to pass through the table. In some embodiments, it may be desirable for the virtual human avatar to appear seated when the user is seated. Thus, it may be beneficial for SLAM to recognize a physical object as a chair and to recognize the dimensions of the chair so that the MR system (e.g., MR systems 212, 300) can display a virtual human avatar as if seated on the chair. Integrating the virtual environment presented to the user with the user's physical environment can create a seamless experience that feels natural to the user as if the user were interacting with physical entities.
[0059] SLAM can rely on, for example, visual input from one or more cameras that use visual odometry. The cameras can capture images of the user's environment, and cameras mounted on an MR system (e.g., MR systems 212, 300) can capture images in the direction the user is looking. Images captured by the SLAM cameras can be fed to a computer vision module, which can identify features captured by the SLAM cameras. The identified features can be tracked across multiple images to determine the location of the features in the physical environment and the location and orientation of the user relative to the features and / or the physical environment. It may be useful to utilize at least two SLAM cameras positioned apart from each other such that each SLAM camera can capture images from different viewpoints. Such stereoscopic imaging can provide additional depth information regarding the location and orientation of features in the physical environment.
[0060] Other sensor inputs can similarly assist SLAM. For example, sensor data from an IMU can be used for SLAM using visual inertial odometry. The IMU can provide information such as the acceleration and rotational velocity of an MR system (e.g., MR systems 212, 300) and correspondingly, a user wearing the MR system. The IMU information can be combined with visual information to determine the position and / or orientation of identified features within the physical environment. For example, the IMU information and visual information can be used to determine a vector regarding gravity that can fix a constructed map of the physical environment. The IMU information can also be used to determine the distance moved and / or rotated between visual frames captured by a user and provide additional information for locating and positioning features within the physical environment. Additional sensor inputs that can assist SLAM can include, for example, depth information from a depth sensor, a LIDAR sensor, and / or a time-of-flight sensor. These sensors can provide additional information for locating and orienting features within the physical environment. The depth information can be particularly useful when a visual sensor captures relatively few features (e.g., an image of a wall without an opening) to track across multiple images.
[0061] In some embodiments, sensor inputs 704a - 704n can include other ways to create a map of the user's environment. For example, sensor inputs 704a - 704n can comprise inputs from a GPS sensor and / or a WiFi chip that can geolocate the MR system (e.g., MR systems 212, 300). The geolocated MR system can then download existing information about the location and environment from a server based on that location information. For example, the MR system can download a 3D map from an online mapping service based on its location. The existing information can be modified or overwritten based on observations from sensor inputs 704a - 704n. Specific examples are used, but any sensor input captured by the MR system (e.g., MR systems 212, 300) and used to determine the user's environment is envisioned to be within the scope of this disclosure. Sensor inputs 704a - 704n can be used to create a map of the user's physical environment at block 706, and this information can be fed into memory graph 701.
[0062] The exemplary user observation module 708 can receive one or more sensor inputs 710a - 710n (which may correspond to sensor inputs 704a - 704n). The sensor inputs 710a - 710n can capture information about the user and the user's responses to various stimuli within the MRE. In some embodiments, the sensor inputs 710a - 710n can capture explicit responses of the user to various stimuli within the MRE. For example, the sensor inputs 710a - 710n can comprise audio signals captured by one or more microphones on the MR system (e.g., MR systems 212, 300). In some embodiments, the user can say out loud "I like it", which can be recorded by one or more microphones on the MR system (e.g., MR systems 212, 300). The one or more microphones can process the audio signal and transcribe the user's utterance, and this transcription can be fed to a natural language processing unit, for example, to determine the meaning behind the spoken words. In some embodiments, the MR system (e.g., MR systems 212, 300) can determine that the audio signal originated from the user wearing the MR system. For example, the audio signal can be processed and compared to one or more previous known recordings of the user's voice to determine whether the user is the speaker. In other embodiments, two microphones positioned on the MR system (e.g., MR systems 212, 300) can be equidistant from the user's mouth, and the audio signals captured by the two microphones can thus contain substantially the same speech signal at substantially the same amplitude, and this information can be used to determine that the user is the speaker.
[0063] In some embodiments, the sensor inputs 710a - 710n can capture other ways the user can use to explicitly indicate a response to one or more stimuli. For example, the user can perform a "thumbs up" gesture, and the MR system (e.g., MR systems 212, 300) can capture this via one or more cameras. The captured image can be processed using computer vision methods to determine that the user has performed a thumbs up gesture, and the MR system (e.g., MR systems 212, 300) can determine that the user is indicating approval through the gesture. The gesture can be either prompted or unprompted by the system. A prompted gesture can include the system indicating that the user can perform a particular gesture if the user prefers the stimulus. In another example, the user can press a button on a controller, which can be part of the MR system (e.g., MR systems 212, 300). In another example, the user can nod. The MR system (e.g., MR systems 212, 300) can capture this information using, for example, cameras and / or an IMU and can determine that the user is indicating approval. Specific examples are used, but any explicit response that can be captured by the MR system (e.g., MR systems 212, 300) is envisioned to be within the scope of this disclosure.
[0064] The sensor inputs 710a - 710n can also capture the user's implicit responses to various stimuli within the MRE. For example, the sensor inputs 710a - 710n can capture information about the user's gaze and determine the degree to which the user is focused (e.g., an eye - tracking sensor can determine the direction of the user's gaze, determine the object the user is looking at, and / or determine the duration of the user's gaze). The sensor inputs 710a - 710n can comprise inputs from one or more outward - facing cameras mounted on an MR system (e.g., MR systems 212, 300) that can capture information about physical objects within the user's field of view. The sensor inputs 710a - 710n can further comprise inputs from one or more inward - facing cameras mounted on an MR system (e.g., MR systems 212, 300) that can capture information about the user's eye movements. These inputs can be combined to determine the user's gaze and what the user is looking at within the MRE (e.g., physical and / or virtual objects the user is viewing). In some embodiments, the MR system (e.g., MR systems 212, 300) can determine the time the user is looking at a physical or virtual object and determine the level of focus. For example, if the user gazes at a physical or virtual object for an extended period of time, the MR system (e.g., MR systems 212, 300) can determine a high level of focus. In another example, one or more inward - facing cameras mounted on an MR system (e.g., MR systems 212, 300) can capture information about the user's mouth movements. If the user is smiling, the MR system (e.g., MR systems 212, 300) can determine an affinity level based on the user's mouth movements. In another example, one or more inward - facing cameras mounted on an MR system (e.g., MR systems 212, 300) can capture information about the user's facial color.When the user's color turns red, the MR system (e.g., MR systems 212, 300) can determine the emotional intensity level, and the appropriate emotion can be determined using other sensor inputs 710a - 710n (e.g., whether the user is smiling, what the user is saying, and the loudness of the user's voice, whether the user is speaking, and / or whether the user is laughing). Implicit responses can also include sounds emitted by the user such as laughter, gasps, moans, etc., which are captured as sensor inputs and can be interpreted to determine the user's emotional state. Specific examples are used, but any implicit response that can be captured by the MR system (e.g., MR systems 212, 300) is assumed to be within the scope of the present disclosure. The sensor inputs 710a - 710n can be used to determine the user response at block 712, and this information can be fed into the memory graph 701.
[0065] External resource 714 can provide additional information to memory graph 701. For example, external resource 714 can comprise an existing social graph. A social graph can represent relationships between entities. For example, a social graph can link various literary works to a common author, a social graph can link various sounds to a common artist, a social graph can link people together (e.g., as colleagues, friends, or family), and a social graph can link images together (e.g., as all images of the Washington Monument or all images of dogs). A social graph can be cited from a social media site, a web crawl algorithm, or any available source. A social graph can also be created and / or modified by an MR system (e.g., MR systems 212, 300) using sensor inputs (e.g., sensor inputs 704a - 704n and / or sensor inputs 710a - 710n). The external resource can also comprise other forms of information. For example, external resource 714 can comprise a connected email account that can provide access to the user's email content. External resource 714 can be fed into memory graph 701.
[0066] The environmental observation module 702, the user observation module 708, and the external resource 714 can be used to create an exemplary memory graph 701, as shown in FIG. 7. The exemplary memory graph 701 can include one or more nodes 716. The nodes 716 can represent physical objects, virtual objects, events, actions, sounds, user responses, and / or generally any experience that a user may encounter. The nodes 716 can be connected to one or more nodes, and these connections can represent any way in which the nodes can be linked to each other. The connections can represent spatial connections (e.g., a table and a chair are located in close proximity to each other), temporal connections (e.g., the rain stopped and the user immediately went out running), semantic connections (e.g., the identified person is the user's colleague), or any other connection.
[0067] The memory graph can represent all known and / or learned information about a user, and how that information relates to itself and other information. For example, node 716c can represent the user's previous vacation in London. Node 716c can be generated using sensor inputs 704a - 704n (e.g., a GPS sensor identifies that the user wearing the MR system is in London, and / or a camera identifies landmarks in London such as Buckingham Palace) and / or external resources 714 (e.g., a connected email account contains travel itineraries to and from London including flights and / or hotel itineraries in London). Node 716e can represent the hotel where the user stayed during the user's trip to London, and node 716e can be connected to node 716c via spatial (e.g., the hotel is located in London), temporal (e.g., while the user was visiting London, the user was at the hotel), semantic (e.g., the hotel had the word "London" in its name), and / or other connections. Node 716d can represent a soccer match the user attended during the user's trip to London, and node 716f can represent the soccer team that played during the soccer match. Node 716d can be connected to node 716c via spatial (e.g., the stadium is located in London), temporal (e.g., while the user was visiting London, the user was at the stadium), semantic, and / or other connections. Node 716d can be connected to node 716f via spatial (e.g., the team was in London), temporal (e.g., the team was in London during the match), semantic (e.g., the team is based in London), and / or other connections. Similarly, node 716f can be connected to node 716c via spatial (e.g., the user was in the same city as the team), temporal (e.g., while the user was visiting London, the user watched the team play), semantic, and / or other connections.
[0068] Each node can have an associated user reaction that can be determined from sensor inputs 710a - 710n, and the association can be generated from the environmental observation module 702 and / or external resource 714. For example, an MR system (e.g., MR systems 212, 300) can observe (e.g., using an inward-facing camera) that the user smiled, and the user observation module can determine a user reaction (e.g., that the user is happy). The environmental observation module 702 and / or external resource 714 can observe (e.g., using GPS and / or WiFi information to determine that the user is at a hotel and using a visual camera to determine that the user entered a room where the number on the door matches the room number provided in the user's email) that the user entered their hotel room. The information provided by the user observation module 708 can be associated with the information provided by the environmental observation module 702 and / or external resource 714, for example, based on their temporal relationship. When the user enters their room and is determined to have been smiling, the user can be determined to be satisfied with the hotel. The association between the user reaction and the node can be a temporal association (e.g., the reaction occurred temporally adjacent to the event represented by the node), a spatial association (e.g., the reaction occurred when the user was physically adjacent to the location represented by the node or when the user was in the physical vicinity of the object represented by the node), or any other association or combination of associations.
[0069] The connections between nodes can be weighted based on the degree to which the nodes are connected and / or based on the user's reaction to the associated nodes. For example, if a user determines that they particularly enjoyed an event represented by a node, the connected nodes may have their connection weighted higher. In some embodiments, a negative reaction by the user can result in one or more associated connections being weighted higher (e.g., for a virtual companion to recommend that the user avoid an object / event / experience) or lower (e.g., to avoid recommending that the user try an object / event / experience).
[0070] Presentation of a virtual companion within the MRE
[0071] FIG. 8 illustrates an exemplary system 800 for presenting a virtual companion to a user within the MRE. Presenting a virtual companion to a user can utilize information from a database 802, an environment observation module 808 (which may correspond to environment observation module 702), and / or a user observation module 814 (which may correspond to user observation module 708). It may be desirable to utilize information from an MR system (e.g., MR systems 212, 300), respect the user's physical environment, present a virtual companion that exists in, and interacts with, and reproduces interactions with physical content, creating a seamless interaction with virtual content that appears to be alive. Seamlessness can be a result of the large number and variety of sensors that may be present on an MR system (e.g., MR systems 212, 300), combined with the "always-on" nature of the MR system (e.g., the user does not need to intentionally interact with the MR system for the MR system to receive input about the user and the user's environment).
[0072] The database 802 can be used to present a virtual companion within the MRE, and the database 802 can comprise various information. For example, the database 802 can comprise a memory graph 804a (which may correspond to the memory graph 701), and the memory graph 804a can represent all (or at least some) of the known and / or learned information about the user. The database 802 can also comprise scripted information 804b. The scripted information 804b can include scripted animations and / or poses that can be used by an MR system (e.g., MR systems 212, 300) to render the virtual companion as a human avatar. For example, the scripted information 804b can comprise recordings of a human actor walking, sitting, and running, which may be animated (e.g., for mesh animation). The scripted information 804b can also comprise voice recordings of a human actor that are decomposed into linguistic basic elements and can be used to synthesize human voice for the virtual companion. The database 802 can also comprise learned information 804c. In some embodiments, the learned information 804c can supplement and / or overwrite the scripted information 804b. For example, the learned information 804c can comprise information about the user speaking in a particular natural language and / or with a particular accent. The MR system (e.g., MR systems 212, 300) can learn the language and / or accent through an audio recording of the user's speech (e.g., via machine learning), and can modify the scripted voice recordings and / or generate new voice recordings to synthesize human speech with the appropriate language and / or accent. The database 802 can further comprise information from user prompts 804d. The user prompts 804d can comprise information directly obtained from the user. For example, the virtual companion may ask the user questions as part of an initialization process (e.g., the virtual companion can "introduce" itself to the user and ask questions that may be typical of an introduction).In some embodiments, some or all of the information contained in 804b - 804d may also be represented in the memory graph 804a.
[0073] The information stored within the database 802 can be used to present a user with a large amount of detailed and personalized information. For example, a user can ask a virtual companion "Where did I stay when I went to London last year?" The database 802 and / or the memory graph 804a can be queried, and the virtual companion can tell the user the hotel where the user stayed based on the information collected about the user.
[0074] The environment module 808 can also be used to present the virtual companion within the MRE in a seamless manner such that the virtual companion appears as a real companion within the real environment. For example, the environment module 808 can determine the presence of an empty chair in the vicinity of the user. When the user sits down, the MR system (e.g., MR systems 212, 300) can display a human avatar that exists in the same space as the user and sits on the empty chair as well. Similarly, when the user walks around, the human avatar can be displayed to move with the user, and the human avatar can be displayed to avoid physical obstacles such as chairs and generally respect the physical environment (e.g., cross over a staircase rather than pass through it).
[0075] The user observation module 814 can also be used to present the virtual companion within the MRE in a seamless manner such that the presented emotional state of the virtual companion mirrors (or at least approximates) that of the user as determined as described above based on explicit and / or implicit cues from the user. For example, the user observation module 814 can determine the general mood of the user (e.g., based on an inward-facing camera that captures information about the user's smile, the user is determined to be happy), and the virtual companion can mirror the user's behavior (e.g., the virtual companion can also be displayed as smiling).
[0076] In some embodiments, the database 802, the environmental observation module 808, and the user observation module 814 can provide information that can be combined to present a seamless virtual companion experience within the MRE where the user is present. In some embodiments, sensors on the MR system (e.g., MR systems 212, 300) enable the virtual companion to present information within the user's MRE in some cases without requiring any prompts from the user. For example, the MR system (e.g., MR systems 212, 300) can determine that the user is discussing a London accommodation with another person (e.g., the microphone on the MR system detects an audio signal that is transcribed and sent to a natural language processor, and the camera on the MR system detects and identifies the person within the user's field of view), and that the user is trying to recall information (e.g., the inward-facing camera on the MR system detects that the user's eyes are looking upward). The database 802 can then be accessed, and the context information can be used from the environmental observation module 808 and the user observation module 814 to determine the hotel where the user stayed during their previous trip to London. This information can then be presented to the user in a non-intrusive and accessible manner in real time (e.g., via a virtual text bubble presented to the user, or via an information card held up by the virtual companion). In other embodiments, the virtual companion can present information (explicitly and / or implicitly learned) to the user within its MRE through an explicit prompt by the user (e.g., the user may ask the virtual companion where the user stayed in London).
[0077] In some embodiments, the virtual companion can interact with the user and the user's MRE. For example, the virtual companion can present itself as a virtual avatar of a dog, and the user can play fetch with the virtual companion. The user can throw a virtual or physical stick, and the virtual companion can move within the physical environment where the user is present and can be presented to respect obstacles within the physical environment (e.g., by moving around the obstacles). In another example, the MR system (e.g., MR systems 212, 300) can connect to other devices (e.g., smart light bulbs), and the user can request that the virtual companion turn on the light. A virtual companion that can access data provided by the MR system (e.g., MR systems 212, 300) has many benefits. For example, information can be continuously recorded by the MR system without user intervention (regardless of whether the virtual companion is currently being displayed). Similarly, information can be presented to the user without user intervention based on the continuously recorded information.
[0078] Examples of Virtual Companions
[0079] Referring to FIG. 9A, a human user (“Alex”) is shown seated on a couch in his actual living room. Alex is wearing a wearable computing system (e.g., MR systems 212, 300), which creates a mesh of the room and objects around him as shown in FIG. 9B. Also referring to FIG. 9B, a virtual companion (which may be named “Aya”) appears and looks like a hologram in the depicted illustration. Referring to FIGS. 9C - 9E, in this embodiment, Aya notices that the room is unusually dark (e.g., via a camera on the MR systems 212, 300), and utilizes observations of Alex's preferences (e.g., observations stored and associated in the memory graphs 701, 804a) to brighten the actual / physical lights in Alex's room (e.g., via a wireless connection to smart bulbs). Aya proceeds to scan the environment (e.g., using SLAM and sensors on the MR systems 212, 300) and understand its context. The scene is segmented, objects are detected, and stored in Aya's memory as an association of information nodes on the right in FIGS. 9C - 9I, which may be called a “Lifestream”. The Lifestream can correspond to the memory graphs 701, 804a. In one embodiment, the Lifestream may be defined as a theoretically complete dataset that captures the entire experience flow of a person (e.g., from birth to death), including both physical and virtual observations and experiences.
[0080] Referring to FIG. 9F, Alex looks at Aya and asks, "Aya, what was the song I liked that was being played at the Pink Floyd concert last summer?" Referring to FIG. 9G, Aya queries Lifestream, reads out the concert memory, and Aya replies, "It's Another brick in the wall." Referring to FIG. 9H, Alex states, "Ah, that's amazing! I wouldn't have been able to remember without your help. Could you play that on the TV?" Aya can stream the music video on the actual TV in the room or, alternatively, present the video via Alex's augmented reality TV. The audio may be presented to Alex, for example, through Alex's headset or other speakers.
[0081] Referring to FIG. 9I, after that conversation, another actual person ("Erica") enters the room and greets Alex. Aya scans Erica's face and recognizes Erica. Aya perceives Alex's reaction to Erica through a camera positioned adjacent to Alex's eye on a wearable computer system component (e.g., MR systems 212, 300) and "sees" that Alex is happy to see Erica. Aya creates another memory snapshot and stores this in Lifestream.
[0082] Referring to FIG. 9J, after Alex greets Erica, Alex tells Erica that Aya just reminded Alex of the song Alex liked at the Pink Floyd concert. Erica replies that she also wants to hear it, so Alex asks Aya to play the song through the physical speakers in the room so that Erica can also hear it. Aya plays the song so that everyone can hear it, tells John that they'll talk later, and disappears.
[0083] Referring to FIGS. 10A and 10B, a virtual, digital, and / or composite or augmented reality assistant or companion, such as the one emphasized in this specification and called "Mica", can preferably be configured to have certain capabilities and characteristics such as approachability, empathy, understanding, memory, and expressions. Various factors such as lighting and realistic blushing, realistic walking movement models, user-based reaction models, and consideration models may provide input for the presentation of such an AI assistant or companion. Computer graphics, animation, capture, and operating systems may be important for creating a virtual companion that seems alive, and thorough details may be required to achieve a persuasive experience. Experts from various fields may need to cooperate closely. If done wrong, the character can become aloof, but if done correctly, its existence and function can be achieved. Among any other type of character, the digital human is perhaps the most difficult, but this is also what the user can most identify with, and thus can be the most rewarding means for developing an approachable AI. Compared to movies, the hurdles in mixed reality are perhaps higher. The interaction with the character is not scripted, and by definition, the user should influence the way the character responds. For example, after developing an accurate synthetic eye expression system, the character and the AI system can be set to track the gaze with the user. The user can have strong opinions about the character and can state them in a way, for example, that the user would describe a human. This can be important for developing a human-centered interface for AI. In these developments, it may be required that AI-related systems with particular focus on important attributes are designed and evolved. As described above, it may be desirable for the system to present the user with a persona that is approachable, empathetic, persistent (i.e., has memory and utilizes the concept of Lifestream), knowledgeable, and useful. These developments can serve as an entry point for making AI less aloof and more natural for the user.The task of representing a human or character is numerous, but the embodiment of a character also brings out the subtle nuances of knowledge and understanding that all people possess.
[0084] Context, details, and nuances can be important, and intelligence does not exist in isolation. Just like human intelligence, AI can emerge not only from one system but also from the interaction of multiple components and agents. It may be desirable to develop the subject system and its configuration as an important benchmark for a human-centered AI interface for composite reality, and it may also be desirable to develop a software system to help creators and developers create a human-centered experience. It may be desirable to help developers create and build experiences induced by humanized AI, that is, experiences that induce realistic emotions and moods and promote the very efficient use of information and computing systems.
[0085] Various exemplary embodiments of the present invention are described herein. These examples are referred to in a non-limiting sense. They are provided to illustrate the more broadly applicable aspects of the present invention. Various changes may be made to the present invention described, and equivalents may be substituted without departing from the true spirit and scope of the present invention. In addition, many modifications may be made to adapt a particular situation, material, composition, process, process action, or step to the purpose, spirit, or scope of the present invention. Furthermore, as will be understood by those skilled in the art, each of the individual variations described and illustrated herein can be readily separated from or combined with features of any of several other embodiments without departing from the scope or spirit of the present invention, having discrete components and features. All such modifications are intended to be within the scope of the claims associated with this disclosure.
[0086] The present invention includes methods that can be implemented using the subject device. The method may include the act of providing such a suitable device. Such provision may be performed by an end user. In other words, the act of "providing" simply requires that the end user acts to obtain, access, approach, locate, configure, activate, power on, or otherwise provide the device required in the subject method. The methods recited herein may be performed in any order of the recited logical possible events and in the recited order of events.
[0087] Exemplary aspects of the present invention are described above, along with details regarding material selection and manufacture. Regarding other details of the present invention, these are understood in relation to the patents and publications referenced above and are generally known or understandable by those skilled in the art. The same can apply to the method-based aspects of the present invention from the perspective of additional acts that are commonly or logically employed.
[0088] In addition, the present invention is optionally described with reference to several embodiments incorporating various features, but the present invention is not limited to what is described or illustrated as would be considered for each variation of the present invention. Various changes may be made to the present invention as described, and equivalents (whether listed herein or not for some brevity purposes) may be substituted without departing from the true spirit and scope of the present invention. Additionally, when a range of values is provided, it is to be understood that all intervening values between the upper and lower limits of that range and any other stated value or intervening value within the stated range are included within the present invention.
[0089] Also, any optional feature of the described variations of the present invention may be considered to be described and claimed independently or in combination with any one or more of the features described herein. References to singular items include the possibility that multiple identical items exist. More specifically, as used in this specification and the claims associated herewith, the singular forms "a", "an", "said", and "the" include plural references unless specifically stated otherwise. In other words, the use of the article enables "at least one" of the subject items in the claims associated with the above description and this disclosure. Further, note that such claims may be drafted to exclude any optional element. Accordingly, the text is intended to serve as a precursor for the use of exclusive terms such as "merely", "only", and equivalents in relation to the recitation of claim elements, or for the use of "negative" limitations.
[0090] Without using such exclusive terms, the term "comprising" in the claims associated with this disclosure shall be taken to enable the inclusion of any additional elements, regardless of whether a given number of elements are recited in such claims, or the addition of a feature may be considered to transform the nature of the elements recited in such claims. Unless specifically defined herein, all technical and scientific terms used herein should be given the broadest generally understood meaning possible while maintaining the validity of the claims.
[0091] The scope of the present invention is not limited to the provided examples and / or this specification, but rather is limited only by the scope of the terms of the claims associated with this disclosure.
Claims
1. 1. A method executed by one or more processors of a head wearable device, the method comprising: receiving, by the one or more processors, a first input from a user's environment via one or more sensors on the head wearable device at a first time; The one or more processors receiving information via an external resource in communication with the head wearable device; determining, by the one or more processors, an occurrence of an event in the environment based on the first input and the information; receiving, by the one or more processors, a second input from the user via the one or more sensors on the head wearable device; identifying a response of the user based on the second input, the one or more processors; and the one or more processors determining a strength of the response; and the one or more processors determining associations between the reactions and the events, the associations being stored in a database and weighted based on the strength of the reactions; receiving, by the one or more processors, a third input from the user via the one or more sensors at a second time after the first time; the one or more processors constructing a query based on the third input, the query being associated with the event; and presenting, at the second time, a response to the query via a see-through display of the head wearable device, the response being determined based on the determined association between the reaction and the event; and A method comprising:
2. The method of claim 1 , wherein the first input comprises at least one of an image of a physical object and an audio signal.
3. The method of claim 1 , wherein the second input comprises an utterance from the user, and identifying the response includes determining content relating to at least a portion of the utterance.
4. The method of claim 1 , wherein the second input comprises eye movements of the user, and identifying the response includes determining a gaze direction for the user based on the eye movements.
5. The method of claim 1 , wherein the second input comprises a field of view of the user, and wherein identifying the response includes identifying at least one object within the field of view.
6. The method of claim 1, wherein the response to the query is determined further based on the strength.
7. The method of claim 1 , wherein the association between the reaction and the event is at least one of a temporal association and a spatial association.
8. The method of claim 1 , wherein the event comprises a first event and the database comprises an association between the first event and a second event.
9. 1. A system comprising: one or more sensors on a head wearable device; A see-through display of the head-wearable device; and one or more processors, the one or more processors comprising: receiving a first input from a user's environment via the one or more sensors on a head wearable device at a first time; receiving information via an external resource in communication with the head wearable device; determining an occurrence of an event in the environment based on the first input and the information; receiving a second input from the user via the one or more sensors on the head wearable device; identifying a response of the user based on the second input; and determining the intensity of the reaction; and determining an association between the reaction and the event, the association being stored in a database and weighted based on the strength of the reaction; receiving a third input from the user via the one or more sensors at a second time after the first time; constructing a query based on the third input, the query being associated with the event; and presenting, at the second time, a response to the query via the see-through display, the response being determined based on the determined association between the reaction and the event; and one or more processors configured to execute a method comprising: A system comprising:
10. The system of claim 9 , wherein the first input comprises at least one of an image of a physical object and an audio signal.
11. The system of claim 9 , wherein the second input comprises an utterance of the user, and identifying the response includes determining content relating to at least a portion of the utterance.
12. The system of claim 9 , wherein the second input comprises eye movements of the user, and identifying the response includes determining a gaze direction for the user based on the eye movements.
13. The system of claim 9 , wherein the second input comprises a field of view of the user, and wherein identifying the response includes identifying at least one object within the field of view.
14. The system of claim 9, wherein the response to the query is further determined based on the strength.
15. A non-transitory computer readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: receiving a first input from a user's environment via one or more sensors on a head wearable device at a first time; receiving information via an external resource in communication with the head wearable device; determining an occurrence of an event in the environment based on the first input and the information; receiving a second input from the user via the one or more sensors on the head wearable device; identifying a response of the user based on the second input; and determining the intensity of the reaction; and determining an association between the reaction and the event, the association being stored in a database and weighted based on the strength of the reaction; receiving a third input from the user via the one or more sensors at a second time after the first time; constructing a query based on the third input, the query being associated with the event; and presenting a response to the query via a see-through display at the second time, the response being determined based on the determined association between the reaction and the event; and A non-transitory computer readable medium for carrying out a method comprising:
16. 16. The non-transitory computer-readable medium of claim 15, wherein the first input comprises at least one of an image of a physical object and an audio signal.
17. 16. The non-transitory computer-readable medium of claim 15, wherein the second input comprises an utterance of the user, and wherein identifying the response includes determining content relating to at least a portion of the utterance.
18. 16. The non-transitory computer-readable medium of claim 15, wherein the second input comprises eye movements of the user, and identifying the response includes determining a gaze direction for the user based on the eye movements.
19. 16. The non-transitory computer-readable medium of claim 15, wherein the second input comprises a field of view of the user, and wherein identifying the response includes identifying at least one object within the field of view.
20. The non-transitory computer-readable medium of claim 15, wherein the response to the query is determined further based on the strength.
Citation Information
Patent Citations
Information processing method, information processing device and program causing computer to execute information processing method
JP2018124666A
Systems and methods for augmented and virtual reality
US20150302652A1
Facial expression identification system, facial expression identification method, and facial expression identification program
WO2017006872A1
Information processing device, information processing method and program
WO2018154933A1
Information processing system, information processing method, and recording medium
WO2019207896A1