Systems and methods for augmented and virtual reality

The system facilitates seamless interaction and dynamic rendering in virtual and augmented reality environments by using a computer network with user devices and wearable sensors, addressing limitations in existing systems for multiple user interaction and physical environment integration.

JP7850320B2Active Publication Date: 2026-04-22MAGIC LEAP INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MAGIC LEAP INC
Filing Date
2025-05-14
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing virtual and augmented reality systems lack the capability to enable seamless interaction between multiple users in different physical locations, with limited flexibility in rendering and interaction modes, and inadequate integration of physical environment data.

Method used

A system comprising a computer network with computing devices and user devices that process and transmit virtual world data, allowing users to interact in augmented or virtual reality modes, with features like wearable devices for input and sensor integration, and dynamic rendering based on physical objects and user interactions.

Benefits of technology

Enables bidirectional virtual and augmented reality experiences for multiple users, allowing real-time interaction and dynamic rendering of virtual objects relative to physical environments, supporting high-definition and low-latency data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850320000001
    Figure 0007850320000001
  • Figure 0007850320000002
    Figure 0007850320000002
  • Figure 0007850320000003
    Figure 0007850320000003
Patent Text Reader

Abstract

To provide systems and methods configured to facilitate interactive virtual or augmented reality environments for one or more users.SOLUTION: One embodiment is directed to a system for enabling two or more users to interact within a virtual world comprising virtual world data, the system comprising a computer network with one or more computing devices comprising memory, processing circuitry, and software that is stored at least in part in the memory and executable by the processing circuitry to process at least a portion of the virtual world data. At least a first portion of the virtual world data originates from a first user virtual world local to a first user, and the computer network is operable to transmit the first portion to a user device for presentation to a second user.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Related Application Data) This application claims priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 61 / 552,941, filed October 28, 2011. Accordingly, the above application is hereby incorporated by reference in its entirety.

[0002] (Background of the Invention) The present invention generally relates to systems and methods configured to facilitate a bidirectional virtual reality or augmented reality environment for one or more users.

Background Art

[0003] (Background) Virtual reality and augmented reality environments are generated by a computer using, in part, data representing the environment. This data may represent, for example, various objects that a user can perceive and interact with. Examples of such objects include objects that are rendered and displayed for the user to view, audio that is played for the user to hear, and haptic (or tactile) feedback that the user can feel. A user may perceive and interact with virtual reality and augmented reality environments through various visual, auditory, and tactile means.

Summary of the Invention

Means for Solving the Problems

[0004] One embodiment relates to a system for enabling two or more users to interact with a virtual world containing virtual world data, the system comprising a computer network comprising one or more computing devices, one or more computing devices comprising memory, processing circuitry, and software at least partially stored in memory and executable by the processing circuitry to process at least a portion of the virtual world data, at least a first portion of the virtual world data originating from a first user virtual world local to a first user, the computer network being operable to transmit the first portion to a user device for presentation to a second user, thereby allowing the second user to experience the first portion from the second user's location, and aspects of the first user virtual world being effectively passed to the second user. The first and second users may be in different physical locations or substantially the same physical locations. At least a portion of the virtual world may be configured to change in response to changes in the virtual world data. At least a portion of the virtual world may be configured to change in response to physical objects sensed by the user device. Changes in the virtual world data may represent virtual objects having a predetermined relationship with physical objects. Changes in virtual world data may be presented to a second user device for presentation to a second user according to a predetermined relationship. The virtual world may be operable to be rendered by at least one of a computer server or a user device. The virtual world may be presented in a two-dimensional format. The virtual world may be presented in a three-dimensional format. The user device may be operable to provide an interface to enable interaction between the user and the virtual world in augmented reality mode. The user device may be operable to provide an interface to enable interaction between the user and the virtual world in virtual reality mode. The user device may be operable to provide an interface to enable interaction between the user and the virtual world in a combination of augmented reality mode and virtual reality mode.Virtual world data may be transmitted over a data network. The computer network may be capable of receiving at least a portion of the virtual world data from a user device. At least a portion of the virtual world data transmitted to the user device may include instructions for generating at least a portion of the virtual world. At least a portion of the virtual world data may be transmitted to a gateway for at least one of processing or distribution. At least one of one or more computer servers may be capable of processing the virtual world data distributed by the gateway.

[0005] Another embodiment relates to a system for virtual and / or augmented user experiences in which a remote avatar is animated, at least partially, based on data on a wearable device, using voluntary input from voice intonation and facial recognition software.

[0006] Another embodiment relates to a system for a virtual user experience and / or an augmented user experience, in which the camera pose or viewpoint position and vector may be located at any location within the world sector.

[0007] Another embodiment relates to a system for virtual and / or augmented user experiences, in which the world or a part thereof may be rendered on a diverse and selectable scale for the observing user.

[0008] Another embodiment relates to a system for virtual and / or augmented user experiences, where, in addition to pose-tagged images, features such as points or parametric lines may be used as foundational data for a world model, from which a software robot or object recognition device may use to create a parametric representation of real-world objects, tagging source features for mutual inclusion in segmented objects and the world model. The present invention provides, for example, the following: (Item 1) A system for enabling two or more users to interact within a virtual world containing virtual world data, wherein the system is A computer network comprising one or more computing devices, wherein each of the one or more computing devices comprises memory, processing circuitry, and software at least partially stored in the memory and executable by the processing circuitry to process at least a portion of the virtual world data. A system comprising, wherein at least a first portion of the virtual world data originates from a first user virtual world that is local to a first user, and the computer network is operable to transmit the first portion to a user device for presentation to a second user, thereby allowing the second user to experience the first portion from the second user's location, and aspects of the first user virtual world are effectively passed to the second user. (Item 2) The first user and the second user are in different physical locations, as described in item 1 of the system. (Item 3) The system described in item 1, wherein the first user and the second user are in substantially the same physical location. (Item 4) The system described in item 1, wherein at least a portion of the virtual world changes in response to changes in the virtual world data. (Item 5) The system according to item 1, wherein at least a portion of the virtual world changes in response to physical objects sensed by the user device. (Item 6) The system described in item 5, wherein the changes in the virtual world data represent virtual objects having a predetermined relationship with the physical objects. (Item 7) The system according to item 6, wherein changes in the virtual world data are presented to a second user device for presentation to the second user according to the predetermined relationship. (Item 8) The system described in item 1, wherein the virtual world is operable to be rendered by at least one of the computer server or user device. (Item 9) The aforementioned virtual world is the system described in item 1, presented in a two-dimensional format. (Item 10) The aforementioned virtual world is the system described in item 1, presented in a three-dimensional format. (Item 11) The system according to item 1, wherein the user device is operable to provide an interface for enabling interaction between the user and the virtual world in augmented reality mode. (Item 12) The system according to item 1, wherein the user device is operable to provide an interface for enabling interaction between the user and the virtual world in virtual reality mode. (Item 13) The system according to item 11, wherein the user device is operable to provide an interface for enabling interaction between the user and the virtual world in a combination of augmented reality mode and virtual reality mode. (Item 14) The aforementioned virtual world data is transmitted over a data network by the system described in item 1. (Item 15) The system according to item 1, wherein the computer network is capable of receiving at least a portion of the virtual world data from user devices. (Item 16) The system according to item 1, wherein at least a portion of the virtual world data transmitted to the user device comprises instructions for generating at least a portion of the virtual world. (Item 17) The system according to item 1, wherein at least a part of the virtual world data is transmitted to a gateway for at least one of processing or distribution. (Item 18) The system according to item 17, wherein at least one of the one or more computer servers is operable to process virtual world data distributed by the gateway.

Brief Description of the Drawings

[0009] [Figure 1] FIG. 1 illustrates an exemplary embodiment of the disclosed system for facilitating a two-way virtual reality or augmented reality environment for multiple users. [Figure 2] FIG. 2 illustrates an example of a user device interacting with the system illustrated in FIG. 1. [Figure 3] FIG. 3 illustrates an exemplary embodiment of a mobile wearable user device. [Figure 4] FIG. 4 illustrates an example of an object visible to a user when the mobile wearable user device of FIG. 3 is operating in an extended mode. [Figure 5] FIG. 5 illustrates an example of an object visible to a user when the mobile wearable user device of FIG. 3 is operating in a virtual mode. [Figure 6] FIG. 6 illustrates an example of an object visible to a user when the mobile wearable user device of FIG. 3 is operating in a mixed virtual interface mode. [Figure 7] FIG. 7 illustrates an embodiment in which two users located at different geographical locations interact with each other and a common virtual world through their respective user devices. [Figure 8] FIG. 8 illustrates an embodiment in which the embodiment of FIG. 7 is extended to include the use of a tactile device. [Figure 9A]Figure 9A illustrates an example of a hybrid mode interface connection where a first user is interface-connected to a digital world in a mixed virtual interface mode and a second user is interface-connected to the same digital world in a virtual reality mode. [Figure 9B] Figure 9B illustrates another example of a hybrid mode interface connection where a first user is interface-connected to a digital world in a mixed virtual interface mode and a second user is interface-connected to the same digital world in an augmented reality mode. [Figure 10] Figure 10 illustrates an exemplary illustration of a user's field of view when interface-connected to the system in an augmented reality mode. [Figure 11] Figure 11 illustrates an exemplary illustration of a user's field of view showing virtual objects induced by physical objects when the user is interface-connected to the system in an augmented reality mode. [Figure 12] Figure 12 illustrates an embodiment of an integrated configuration of augmented reality and virtual reality where a user in an augmented reality experience visualizes the presence of another user in a virtual reality experience. [Figure 13] Figure 13 illustrates an embodiment of a time- and / or event-based augmented reality experience configuration. [Figure 14] Figure 14 illustrates an embodiment of a user display configuration suitable for virtual reality and / or augmented reality experiences. [Figure 15] Figure 15 illustrates an embodiment of local and cloud-based computational collaboration. [Figure 16] Figure 16 illustrates various aspects of an alignment configuration.

Mode for Carrying Out the Invention

[0010] (Detailed Description) Referring to Figure 1, System 100 is typical hardware for implementing the processes described below. This typical system comprises a computing network 105 consisting of one or more computer servers 110 connected through one or more high-bandwidth interfaces 115. The servers in the computing network do not need to be located in the same place. Each of the one or more servers 110 comprises one or more processors for executing program instructions. The servers also include memory for storing program instructions and data used and / or generated by processes executed by the servers under the direction of the program instructions.

[0011] The computing network 105 communicates data between servers 110 and between servers and one or more user devices 120 over one or more data network connections 130. Examples of such data networks include, but are not limited to, all types of public and private data networks, both mobile and wired, including many interconnections of such networks, such as the internet. No particular media, topology, or protocol is intended to be implied by the diagram.

[0012] User devices are configured to communicate directly with either the computing network 105 or the server 110. Alternatively, user device 120 communicates with the remote server 110 and, if necessary, with other user devices locally through a specially programmed local gateway 140 for processing data and / or for communicating data between the network 105 and one or more local user devices 120.

[0013] As illustrated, the gateway 140 is implemented as a separate hardware component including a processor for executing software instructions and memory for storing software instructions and data. The gateway has its own wired and / or wireless connectivity to a data network for communicating with servers 110 that constitute the computing network 105. Alternatively, the gateway 140 can be integrated with a user device 120 that is worn or carried by the user. For example, the gateway 140 may be implemented as a downloadable software application that is installed and runs on a processor contained in the user device 120. In one embodiment, the gateway 140 provides access to the computing network 105 to one or more users via the data network 130.

[0014] Each server 110 includes, for example, working memory and storage devices for storing data and software programs, a microprocessor for executing program instructions, and a graphics processor and other specialized processors for rendering and generating graphics, images, video, audio, and multimedia files. The computing network 105 may also include devices for storing data accessed, used, or created by the servers 110.

[0015] The server, optionally, the user device 120, and the software programs running on the gateway 140 are used to generate a digital world (also referred to herein as a virtual world) in which the user interacts with the user device 120. The digital world is represented by data and processes that represent and / or define virtual, non-existent entities, environments, and conditions that can be presented to the user through the user device 120 for the user to experience and interact with. For example, any type of object, entity, or item that appears to be physically present when instantiated in a scene viewed or experienced by the user may include a description of its appearance, its behavior, how the user is permitted to interact with it, and other characteristics. The data used to create the environment of the virtual world (including virtual objects) may include, for example, atmospheric data, topographic data, weather data, temperature data, location data, and other data used to define and / or represent the virtual environment. In addition, the data defining the various conditions governing the behavior of the virtual world may include, for example, physical laws, time, spatial relationships, and other data that can be used to define and / or create the various conditions governing the behavior of the virtual world (including virtual objects).

[0016] Entities, objects, conditions, characteristics, behaviors, or other features of the digital world are generally referred to as objects (e.g., digital objects, virtual objects, rendered physical objects, etc.) as herein, unless the context specifically indicates otherwise. Objects may be any type of living or inanimate object, including, but not limited to, buildings, plants, vehicles, people, animals, living things, machines, data, videos, text, photographs, and other users. Objects may also be defined in the digital world to store information about items, behaviors, or conditions that actually exist in the physical world. Data that represents or defines an entity, object, or item, or data that stores its current state, is generally referred to as object data as herein. Depending on the implementation, this data is processed by the server 110, or by the gateway 140 or user device 120, to instantiate instances of the object and render the object in a manner appropriate for the user to experience through the user device.

[0017] A programmer who develops and / or creates a digital world creates or defines objects and the conditions under which those objects are instantiated. However, the digital world can allow others to create or modify objects. Once an object is instantiated, its state may be allowed to be changed, controlled, or manipulated by one or more users experiencing the digital world.

[0018] For example, in one embodiment, the development, production, and management of the digital world are generally provided by one or more system administrator programmers. In some embodiments, this may include the development, design, and / or execution of storylines, themes, and events in the digital world, as well as the distribution of the narrative through various forms of events and media, such as movies, digital, network, mobile, augmented reality, and live entertainment. The system administrator programmer may also handle the technical management, moderation, and curation of the digital world and its associated user community, as well as other tasks typically performed by network administrators.

[0019] The user interacts with one or more digital worlds using some type of local computing device, generally designated as user device 120. Embodiments of such user devices include, but are not limited to, smartphones, tablet devices, head-up displays (HUDs), game consoles, or any other devices capable of communicating data and providing an interface or display to the user, or combinations of such devices. In some embodiments, user device 120 may include, or communicate with, local peripheral components or input / output components such as, for example, keyboards, mice, joysticks, game controllers, haptic interface devices, motion capture controllers, optical tracking devices (e.g., those available from Leap Motion, Inc. or those available from Microsoft under the trade name Kinect (RTM)), audio equipment, voice equipment, projector systems, 3D displays, and holographic 3D contact lenses.

[0020] An embodiment of a user device 120 for interacting with system 100 is illustrated in Figure 2. In the exemplary embodiment shown in Figure 2, a user 210 may interface with one or more digital worlds through a smartphone 220. The gateway is implemented by a software application 230 stored and operating on the smartphone 220. In this particular embodiment, the data network 130 includes a wireless mobile network that connects the user device (i.e., the smartphone 220) to the computer network 105.

[0021] In one preferred embodiment, the system 100 is capable of supporting a large number of simultaneous users (e.g., millions of users), each of which interfaces with the same digital world or multiple digital worlds using some type of user device 120.

[0022] The user device provides the user with an interface to enable visual, audible, and / or physical interaction between the user and the digital world generated by the server 110 (including other users and objects (real or virtual) presented to the user). The interface provides the user with a rendered scene that can be seen, heard, or otherwise perceived, and the ability to interact with that scene in real time. The manner in which the user interacts with the rendered scene may be determined by the capabilities of the user device. For example, if the user device is a smartphone, the user interaction may be implemented by the user touching the touchscreen. In another embodiment, if the user device is a computer or game console, the user interaction may be implemented using a keyboard or game controller. The user device may include additional components that enable user interaction, such as sensors, and objects and information (including gestures) detected by the sensors may be provided as inputs representing user interaction with the virtual world using the user device.

[0023] The rendered scene can be presented in various forms, such as two-dimensional or three-dimensional visual displays (including projections), sound, and tactile or haptic feedback. The rendered scene may be interfaced by the user in one or more modes, including, for example, augmented reality, virtual reality, and combinations thereof. The form of the rendered scene, as well as the interface mode, may be determined by one or more of the following: user device, data processing capacity, user device connectivity, network capacity, and system workload. The simultaneous interaction of multiple users with the digital world and the real-time nature of data exchange are made possible by the computing network 105, server 110, gateway component 140 (if necessary), and user device 120.

[0024] In one embodiment, the computing network 105 comprises a large-scale computing system having single and / or multicore servers (i.e., server 110) connected through high-speed connections (e.g., high-bandwidth interface 115). The computing network 105 may form a cloud or grid network. Each server contains memory or is coupled with computer-readable memory for storing software for implementing data to create, design, modify, or process objects in the digital world. These objects and their instantiations may be dynamic, appearing, disappearing, changing over time, and changing in response to other conditions. Embodiments of the dynamic capabilities of objects are generally discussed herein with respect to various embodiments. In some embodiments, each user interfaced with system 100 may also be represented as an object and / or a collection of objects in one or more digital worlds.

[0025] Server 110 within the computing network 105 also stores computed state data for each of the digital worlds. Computed state data (also referred to herein as state data) may be components of object data and generally define the state of an instance of an object in a given instance over time. Thus, computed state data may change over time and may be affected by the actions of one or more users and / or programmers who maintain the system 100. When a user affects computed state data (or other data that constitutes a digital world), the user either directly modifies or otherwise manipulates the digital world. If the digital world is shared with or interfaced with other users, the user's actions may affect what is experienced by other users interacting with the digital world. Thus, in some embodiments, changes made to the digital world by a user are experienced by other users interfaced with the system 100.

[0026] In one embodiment, data stored in one or more servers 110 within the computing network 105 is transmitted or unpacked to one or more user devices 120 and / or gateway components 140 at high speed and with low latency. In one embodiment, object data shared by the servers may be complete or compressed and include instructions for recreating the complete object data on the user side, which may be rendered and visualized by the user's local computer device (e.g., gateway 140 and / or user device 120). In some embodiments, software running on the servers 110 of the computing network 105 may adapt the data that the computing network 105 generates and transmits to a particular user's device 120 for objects in the digital world (or any other data exchanged by the computing network 105) depending on the user's specific device and bandwidth. For example, when a user interacts with the digital world through a user device 120, the server 110 may recognize the specific type of device used by the user, the device connectivity, and / or the available bandwidth between the user device and the server, and appropriately determine and balance the size of the data being delivered to the device to optimize user interaction. Embodiments of this may include reducing the size of the transmitted data to a low-resolution quality so that the data can be displayed on a particular user device having a low-resolution display. In a preferred embodiment, the computing network 105 and / or gateway component 140 deliver the data to the user device 120 at a speed fast enough to present an interface operating at 15 frames per second or faster and at high-definition quality or higher resolution.

[0027] The gateway 140 provides local connectivity to the computing network 105 for one or more users. In some embodiments, it may be implemented by a downloadable software application running on a user device 120 or another local device such as the one shown in Figure 2. In other embodiments, it may be implemented by a hardware component (a component having a processor with appropriate software / firmware stored on the component) that communicates with the user device 120 but is either built into the user device 120, not attached to it, or built into the user device 120. The gateway 140 communicates with the computing network 105 via the data network 130 and provides data exchange between the computing network 105 and one or more local user devices 120. As will be discussed in more detail below, the gateway component 140 may include software, firmware, memory, and processing circuitry that may be capable of processing the data communicated between the network 105 and one or more local user devices 120.

[0028] In some embodiments, the gateway component 140 monitors and adjusts the rate of data exchanged between the user device 120 and the computer network 105 to enable optimal data processing capabilities for a particular user device 120. For example, in some embodiments, the gateway 140 buffers and downloads both static and dynamic aspects of the digital world, even beyond the field of view presented to the user through the interface connected to the user device. In such embodiments, instances of static objects (structured data, software implementations, or both) may be stored in memory (local to the gateway component 140, the user device 120, or both) and referenced relative to the local user's current location, as indicated by the data provided by the computing network 105 and / or the user's device 120. Instances of dynamic objects, which may include intelligent software agents and objects controlled by other users and / or the local user, are stored in a high-speed memory buffer. Dynamic objects, representing two-dimensional or three-dimensional objects in the view presented to the user, can be classified into component shapes, such as static shapes that are moving but not changing, and dynamic shapes that are changing. Some of the changing dynamic objects can be updated by a real-time, thread-high-priority data stream from the server 110 through the computing network 105, managed by the gateway component 140. As one embodiment of the priority thread data stream, data within 60 degrees of the user's field of view may be given higher priority than data further out. Another embodiment includes prioritizing dynamic characters and / or objects within the user's field of view over static background objects.

[0029] In addition to managing the data connection between the computing network 105 and the user device 120, the gateway component 140 may store and / or process data that may be presented to the user device 120. For example, in some embodiments, the gateway component 140 may receive compressed data from the computing network 105, for example, representing graphical objects to be rendered for user viewing, and perform advanced rendering techniques to reduce the data load transmitted from the computing network 105 to the user device 120. In another embodiment, where the gateway 140 is a separate device, the gateway 140 may store and / or process data of local instances of objects, rather than transmitting data to the computing network 105 for processing.

[0030] Referring here to Figure 3, the digital world may be experienced by one or more users in various forms that may depend on the capabilities of the user's device. In some embodiments, the user device 120 may include, for example, a smartphone, a tablet device, a head-up display (HUD), a game console, or a wearable device. Generally, the user device includes a processor for executing program code stored in memory on the device, coupled with a display, and a communication interface. An exemplary embodiment of a user device is shown in Figure 3, which comprises a mobile wearable device, i.e., a head-mounted display system 300. According to embodiments of this disclosure, the head-mounted display system 300 includes a user interface 302, a user sensing system 304, an environment sensing system 306, and a processor 308. The processor 308 is shown in Figure 3 as an isolated component separate from the head-mounted system 300, but in alternative embodiments, the processor 308 may be integrated with one or more components of the head-mounted system 300, or incorporated into a component of another system 100, such as a gateway 140.

[0031] The user device presents the user with an interface 302 for interacting with and experiencing the digital world. Such interactions may include the user and the digital world, one or more other users interfaced with the system 100, and objects within the digital world. Interface 302 generally provides the user with image and / or audio sensory input (and in some embodiments, physical sensory input). Thus, interface 302 may include a speaker (not shown) and, in some embodiments, a display component 303 that can enable stereoscopic 3D viewing and / or 3D viewing that embodies more natural characteristics of the human visual system. In some embodiments, the display component 303 may have a transparent interface (such as a transparent OLED) that, when in the "off" setting, allows for an optically correct view of the physical environment around the user with little or no optical distortion or computing overlay. As will be discussed in more detail below, interface 302 may include additional settings that enable various visual / interface performance and functionality.

[0032] In some embodiments, the user sensing system 304 may include one or more sensors 310 that are operable to detect specific features, characteristics, or information relating to an individual user wearing the system 300. For example, in some embodiments, the sensor 310 may include a camera or optical detection / scanning circuit capable of detecting real-time optical characteristics / measurements of the user, such as pupil constriction / dilation, angle measurement / position of each pupil, sphericity, eye shape (as changes in eye shape over time), and one or more other anatomical data. This data may provide information (e.g., the user's visual focus) that can be used by the head-mounted system 300 and / or interface system 100 to optimize the user's visual experience, or may be used to calculate such information. For example, in one embodiment, each sensor 310 may measure the pupil constriction rate of each of the user's eyes. This data may be transmitted to the processor 308 (or to the gateway component 140, or to the server 110), and the data may be used, for example, to determine the user's response to the brightness setting of the interface display 303. Interface 302 may be adjusted according to user responses, for example, by dimming the display 303 if the user's response indicates that the brightness level of the display 303 is too high. User sensing system 304 may include other components other than those discussed above or illustrated in Figure 3. For example, in some embodiments, user sensing system 304 may include a microphone for receiving voice input from the user. User sensing system may also include one or more infrared camera sensors, one or more visible spectrum camera sensors, structural light emitters and / or sensors, infrared light emitters, coherent light emitters and / or sensors, gyroscopes, accelerometers, magnetometers, proximity sensors, GPS sensors, ultrasonic emitters and detectors, and a haptic interface.

[0033] The environmental sensing system 306 includes one or more sensors 312 for acquiring data from the physical environment surrounding the user. Objects or information detected by the sensors may be provided to the user device as input. In some embodiments, this input may represent user interaction with a virtual world. For example, a user viewing a virtual keyboard on a desk may make finger gestures as if typing on the virtual keyboard. The movement of the fingers may be captured by the sensors 312 and provided to the user device or system as input, which may be used to change the virtual world or to create new virtual objects. For example, the finger movements may be recognized as typing (using a software program), and the recognized typing gestures may be combined with known positions of virtual keys on the virtual keyboard. The system may then render a virtual monitor that is displayed to the user (or other users interfaced with the system), which displays the text being typed by the user.

[0034] The sensor 312 may include, for example, a substantially outward-facing camera or scanner to interpret scene information through continuously and / or intermittently projected infrared structured light. The environment sensing system 306 may be used to map one or more elements of the physical environment around the user by detecting and registering static objects, dynamic objects, people, gestures, and the local environment, including various lighting, atmospheric, and acoustic conditions. Thus, in some embodiments, the environment sensing system 306 may include image-based 3D reconstruction software, which is integrated into a local computing system (e.g., gateway component 140 or processor 308) and is operable to digitally reconstruct one or more objects or information detected by the sensor 312. In one exemplary embodiment, the environment sensing system 306 provides motion capture data (including gesture recognition), depth sensing, face recognition, object recognition, unique object feature recognition, voice / audio recognition and processing, sound source localization, noise reduction, infrared laser or similar laser projection, and one or more of monochrome and / or color CMOS sensors (or other similar sensors), field of view sensors, and various other light-enhanced sensors. It should be understood that the environment sensing system 306 may include other components other than those discussed above or illustrated in Figure 3. For example, in some embodiments, the environment sensing system 306 may include a microphone for receiving sound from the local environment. The user sensing system may also include one or more infrared camera sensors, one or more visible spectrum camera sensors, structural light emitters and / or sensors, infrared light emitters, coherent light emitters and / or sensors, gyroscopes, infrared light emitters, accelerometers, magnetometers, proximity sensors, GPS sensors, ultrasonic emitters and detectors, and a tactile interface.

[0035] As described above, in some embodiments, the processor 308 may be integrated with other components of the head-mounted system 300, integrated with other components of the interface system 100, or may be an isolated device (wearable or separate from the user) as shown in Figure 3. The processor 308 may be connected to various components of the head-mounted system 300 and / or components of the interface system 100 via physical wired connections or via wireless connections such as mobile network connections (including cellular and data networks), Wi-Fi, or Bluetooth®. The processor 308 may include a memory module, an integrated and / or additional graphics processing unit, wireless and / or wired internet connectivity, and a codec and / or firmware capable of converting data from a source (e.g., computing network 105, user sensing system 304, environment sensing system 306, or gateway component 140) into image and audio data, which may be presented to the user via interface 302.

[0036] The processor 308 handles data processing for various components of the head-mounted system 300, as well as data exchange between the head-mounted system 300 and the gateway component 140 (in some embodiments, the computing network 105). For example, the processor 308 may be used to buffer and process data streaming between the user and the computing network 105, thereby enabling a smooth, continuous, and high-fidelity user experience. In some embodiments, the processor 308 may process data at a speed sufficient to achieve any or more between 8 frames / second at 320x240 resolution and 24 frames / second at high-definition resolution (1280x720) (e.g., 60-120 frames / second and 4k resolution or higher (10k+ resolution and 50,000 frames / second)). In addition, the processor 308 may store and / or process data that may be presented to the user rather than being streamed in real time from the computing network 105. For example, in some embodiments, the processor 308 may receive compressed data from the computing network 105 and reduce the data load transmitted from the computing network 105 to the user device 120 by performing advanced rendering techniques (such as brightness or shading). In another embodiment, the processor 308 may store and / or process local object data instead of transmitting data to the gateway component 140 or the computing network 105.

[0037] In some embodiments, the head-mounted system 300 may include various settings or modes that enable various visual / interface performance and functionality. The modes may be selected manually by the user or automatically by the components of the head-mounted system 300 or the gateway component 140. As described above, one embodiment of the head-mounted system 300 includes an "off" mode in which the interface 302 provides substantially no digital or virtual content. In off mode, the display component 303 may be transparent, thereby enabling an optically correct view of the user's surrounding physical environment with little or no optical distortion or computing overlay.

[0038] In one exemplary embodiment, the head-mounted system 300 includes an "augmented" mode in which interface 302 provides an augmented reality interface. In augmented mode, the interface display 303 may be substantially transparent, thereby allowing the user to view the local physical environment. Simultaneously, virtual object data provided by the computing network 105, processor 308, and / or gateway component 140 is presented on the display 303 in combination with the local physical environment.

[0039] Figure 4 illustrates an exemplary embodiment of objects visible to the user when interface 302 is operating in extended mode. As shown in Figure 4, interface 302 presents physical objects 402 and virtual objects 404. In the embodiment illustrated in Figure 4, physical objects 402 are actual physical objects present in the user's local environment, while virtual objects 404 are objects created by system 100 and displayed via user interface 302. In some embodiments, virtual objects 404 may be displayed at a fixed position or location within the physical environment (e.g., a virtual monkey standing next to a specific road sign located within the physical environment), or they may be displayed to the user as objects located relative to the user interface / display 303 (e.g., a virtual clock or thermometer visible in the upper left corner of display 303).

[0040] In some embodiments, a virtual object may be signaled to or triggered by a physically existing object within or outside the user's field of view. For example, the virtual object 404 is signaled to or triggered by a physical object 402. For instance, the physical object 402 may actually be a stool, and the virtual object 404 may be presented to the user (or, in some embodiments, to another user interfaced with the system 100) as a virtual animal standing on the stool. In such embodiments, the environment sensing system 306 may use, for example, software and / or firmware stored in the processor 308 to recognize various features and / or shape patterns (captured by the sensor 312) in order to identify the physical object 402 as a stool. These recognized shape patterns, for example, the top of the stool, may be used to trigger the placement of the virtual object 404. Other embodiments may use any visible object, including walls, tables, furniture, cars, buildings, people, floors, plants, and animals, to trigger an augmented reality experience in which there is some relationship with one or more objects.

[0041] In some embodiments, the specific virtual object 404 to be triggered may be selected by the user or automatically selected by other components of the head-mounted system 300 or the interface system 100. In addition, in embodiments in which the virtual object 404 is automatically triggered, the specific virtual object 404 may be selected based on a specific physical object 402 (or its characteristics) to which the virtual object 404 is signaled or triggered. For example, if the physical object is identified as a diving board extending over a pool, the triggered virtual object may be a creature wearing a snorkel, swimsuit, flotation device, or other related item.

[0042] In another exemplary embodiment, the head-mounted system 300 may include a “virtual” mode in which interface 302 provides a virtual reality interface. In virtual mode, the physical environment is omitted from display 303, and virtual object data provided by the computing network 105, processor 308, and / or gateway component 140 is presented on display 303. The omission of the physical environment may be achieved by physically blocking the visual display 303 (e.g., by a cover) or through a feature of interface 302 in which display 303 transitions to an opaque setting. In virtual mode, live visual and auditory sensations and / or stored visual and auditory sensations may be presented to the user through interface 302, and the user experiences and interacts with the digital world (digital objects, other users, etc.) through the virtual mode of interface 302. Thus, the interface provided to the user in virtual mode consists of virtual object data, including the virtual digital world.

[0043] Figure 5 illustrates an exemplary embodiment of the user interface when the head-mounted interface 302 is operating in virtual mode. As shown in Figure 5, the user interface presents a virtual world 500 composed of digital objects 510, which may include atmosphere, weather, terrain, buildings, and people. Although not shown in Figure 5, the digital objects may also include, for example, plants, vehicles, animals, organisms, machines, artificial intelligence, location information, and any other objects or information that define the virtual world 500.

[0044] In another exemplary embodiment, the head-mounted system 300 may include a “mixed” mode, and various features of the head-mounted system 300 (as well as features of the virtual mode and extended mode) may be combined to create one or more custom interface modes. In one example of a custom interface mode, the physical environment is omitted from the display 303, and virtual object data is presented on the display 303 in a manner similar to that of the virtual mode. However, in this example of a custom interface mode, the virtual objects may be entirely virtual (i.e., they do not exist in the local physical environment), or they may be actual local physical objects that are rendered as virtual objects in the interface 302 instead of physical objects. Thus, in a particular custom mode (referred to herein as a mixed virtual interface mode), live visual and auditory sensations and / or stored visual and auditory sensations may be presented to the user through the interface 302, and the user experiences and interacts with a digital world that includes fully virtual objects and rendered physical objects.

[0045] Figure 6 illustrates an exemplary embodiment of a user interface operating according to a mixed virtual interface mode. As shown in Figure 6, the user interface presents a virtual world 600 consisting of fully virtual objects 610 and rendered physical objects 620 (renderings of objects that otherwise exist physically in the scene). According to the embodiment illustrated in Figure 6, the rendered physical objects 620 include a building 620A, ground 620B, and platform 620C, and are indicated by thick outlines 630 to show the user that the objects are rendered. In addition, the fully virtual objects 610 include an additional user 610A, clouds 610B, sun 610C, and flames 610D on platform 620C. It should be understood that the fully virtual objects 610 may include, for example, the atmosphere, weather, terrain, buildings, people, plants, vehicles, animals, organisms, machines, artificial intelligence, location information, and any other objects or information that are not rendered from objects that exist in the local physical environment of the virtual world 600. Conversely, the rendered physical object 620 is an actual local physical object that is rendered as a virtual object within interface 302. The thick contour 630 represents one embodiment for showing the rendered physical object to the user. Thus, the rendered physical object may be shown using methods other than those disclosed herein.

[0046] In some embodiments, the rendered physical objects 620 may be detected using the sensor 312 of the environment sensing system 306 (or using other devices such as a motion or image capture system) and converted into digital object data by software and / or firmware stored in the processing circuit 308, for example. Thus, when a user interfaces with the system 100 in a mixed virtual interface mode, various physical objects may be displayed to the user as rendered physical objects. This can be particularly useful in allowing the user to interface with the system 100 while still being able to safely navigate the local physical environment. In some embodiments, the user may be able to selectively remove or add rendered physical objects to the interface display 303.

[0047] In another example of a custom interface mode, the interface display 303 may be substantially transparent, thereby allowing the user to perceive the local physical environment while various local physical objects are displayed to the user as rendered physical objects. This example of a custom interface mode is similar to the extended mode, except that one or more of the virtual objects may be rendered physical objects, as discussed above with respect to the previous embodiment.

[0048] The aforementioned examples of custom interface modes represent some exemplary embodiments of the various custom interface modes that can be provided by the mixed modes of the head-mounted system 300. Accordingly, various other custom interface modes may be created from various combinations of the components of the head-mounted system 300 and the features and functionalities provided by the various modes discussed above, without departing from the scope of this disclosure.

[0049] The embodiments discussed herein merely describe some examples of providing an interface that operates in off-mode, extended-mode, virtual-mode, or mixed-mode, and are not intended to limit the scope or content of each interface mode or the functionality of the components of the head-mounted system 300. For example, in some embodiments, virtual objects may include data displayed to the user (time, temperature, altitude, etc.), objects created and / or selected by system 100, objects created and / or selected by the user, or even objects representing other users interfaced with system 100. In addition, virtual objects may include extensions of physical objects (e.g., virtual statues growing from a physical platform) and may be visually connected to or separated from physical objects.

[0050] Virtual objects are also dynamic and change over time, according to various relationships between the user, other users, physical objects, and other virtual objects (e.g., position, distance, etc.), and / or according to other variables defined in the software and / or firmware of the head-mounted system 300, gateway component 140, or server 110. For example, in certain embodiments, virtual objects may respond to user devices or their components (e.g., a virtual ball moves when a tactile device is placed next to it), physical or verbal user interactions (e.g., a virtual creature runs away when a user approaches it, or speaks when a user speaks to it), a chair being thrown at the virtual creature causing it to avoid the chair, other virtual objects (e.g., a first virtual creature reacts when it sees a second virtual creature), physical variables such as position, distance, temperature, time, or other physical objects in the user's environment (e.g., a virtual creature shown standing on a physical road flattens when a physical car passes by).

[0051] The various modes discussed herein may also be applied to user devices other than the head-mounted system 300. For example, an augmented reality interface may be provided via a mobile phone or tablet device. In such embodiments, the phone or tablet may use a camera to capture the physical environment around the user, and virtual objects may be overlaid on the phone / tablet display screen. In addition, the virtual mode may be provided by displaying a digital world on the phone / tablet display screen. Thus, these modes may be mixed to create various custom interface modes, as described above, using the phone / tablet components discussed herein, as well as other components connected to or used in combination with the user device. For example, a mixed virtual interface mode may be provided by a computer monitor, television screen, or other device without a camera, operating in combination with a motion or image capture system. In this exemplary embodiment, the virtual world may be visible from the monitor / screen, and object detection and rendering may be performed by the motion or image capture system.

[0052] Figure 7 illustrates an exemplary embodiment of the present invention in which two users located in different geographical locations interact with the other user and a common virtual world through their respective user devices. In this embodiment, the two users 701 and 702 throw a virtual ball 703 (a type of virtual object) back and forth, and each user can observe the other user's influence on the virtual world (for example, each user observes the virtual ball changing direction, being caught by the other user, etc.). Since the movement and position of the virtual object (i.e., the virtual ball 703) are tracked by a server 110 in a computing network 105, the system 100 may, in some embodiments, communicate to users 701 and 702 the exact location and timing of the ball 703's arrival for each user. For example, if the first user 701 is located in London, user 701 may throw the ball 703 to the second user 702, located in Los Angeles, at a speed calculated by the system 100. Therefore, system 100 may communicate the exact time and location of the ball's arrival to a second user 702 (e.g., via email, text message, instant message, etc.). In this case, the second user 702 may use their device to see the ball 703 arrive at the specified time and location. One or more users may also use geolocation mapping software (or similar) to track one or more virtual objects when virtually traveling around the Earth. An example of this might be a user wearing a 3D head-mounted display, looking up at the sky and seeing a virtual airplane flying overhead, superimposed on the real world. The virtual airplane may be flown by the user, an intelligent software agent (software running on the user device or gateway), other users who may be locally and / or remotely, and / or any combination thereof.

[0053] As described above, the user device may include a haptic interface device that provides feedback (e.g., resistance, vibration, light, sound, etc.) to the user when the system 100 determines that the haptic device is located in a physical spatial position relative to a virtual object. For example, the embodiment described above with respect to Figure 7 may be extended to include the use of a haptic device 802, as shown in Figure 8.

[0054] In this exemplary embodiment, the haptic device 802 may be represented in the virtual world as a baseball bat. When a ball 703 arrives, the user 702 may swing the haptic device 802 towards the virtual ball 703. If the system 100 determines that the virtual bat provided by the haptic device 802 has "contacted" the ball 703, the haptic device 802 may vibrate or provide other feedback to the user 702, and the virtual ball 703 may bounce off the virtual bat in a direction calculated by the system 100 according to the detected velocity, direction, and timing of the contact between the ball and the bat.

[0055] In some embodiments, the disclosed system 100 may facilitate mixed-mode interface connections, allowing multiple users to interface with a common virtual world (and the virtual objects contained therein) using different interface modes (e.g., augmented, virtual, mixed, etc.). For example, a first user interface-connecting to a particular virtual world in virtual interface mode may interact with a second user interface-connecting to the same virtual world in augmented reality mode.

[0056] Figure 9A illustrates an embodiment in which a first user 901 (interfacing with the digital world of system 100 in mixed virtual interface mode) and a first object 902 appear as virtual objects to a second user 922 (interfacing with the same digital world of system 100 in full virtual reality mode). As described above, when interfacing with the digital world via mixed virtual interface mode, local physical objects (e.g., the first user 901 and the first object 902) may be scanned and rendered as virtual objects in the virtual world. The first user 901 may be scanned, for example, by a motion capture system or similar device and rendered in the virtual world as a first rendered physical object 931 (by software / firmware stored in the motion capture system, gateway component 140, user device 120, system server 110, or other devices). Similarly, the first object 902 may be scanned, for example, by the environment sensing system 306 of the head-mounted interface 300 and rendered in the virtual world as a second rendered physical object 932 (by software / firmware stored in the processor 308, gateway component 140, system server 110, or other devices). The first user 901 and the first object 902 are shown as physical objects in the physical world in the first part 910 of Figure 9A. In the second part 920 of Figure 9A, the first user 901 and the first object 902 are shown as a first rendered physical object 931 and a second rendered physical object 932, appearing to a second user 922 interfaced with the same digital world of the system 100 in full virtual reality mode.

[0057] Figure 9B illustrates another exemplary embodiment of a mixed-mode interface connection, in which a first user 901 interfaces with the digital world in a mixed-virtual interface mode, as discussed above, and a second user 922 interfaces with the same digital world (and the second user's physical local environment 925) in an augmented reality mode. In the embodiment of Figure 9B, the first user 901 and the first object 902 are located at a first physical location 915, and the second user 922 is located at a different second physical location 925, separated by a certain distance from the first location 915. In this embodiment, virtual objects 931 and 932 may be transposed in real time (or near real time) to their locations in the virtual world corresponding to the second location 925. Therefore, the second user 922 may observe and interact with the rendered physical object 931 representing the first user 901 and the rendered physical object 932 representing the first object 902 in the second user's local physical environment 925.

[0058] Figure 10 illustrates an exemplary example of a user's field of view when interfaced with system 100 in augmented reality mode. As shown in Figure 10, the user sees a local physical environment (i.e., a city with multiple buildings) as well as a virtual character 1010 (i.e., a virtual object). The position of the virtual character 1010 may be triggered by 2D visual targets (e.g., signs, postcards, or magazines) and / or one or more 3D reference coordinate systems (e.g., buildings, cars, people, animals, airplanes, parts of buildings, and / or 3D physical objects, virtual objects, and / or combinations thereof). In the embodiment illustrated in Figure 10, known locations of buildings in the city may provide alignment references and / or information and important features for rendering the virtual character 1010. In addition, the user's geospatial position relative to the buildings (e.g., provided by GPS, attitude / position sensors, etc.) or mobile position may include data used by the computing network 105 to trigger the transmission of data used to display the virtual character 1010(one or more). In some embodiments, the data used to display the virtual character 1010 may include instructions (executed by the gateway component 140 and / or the user device 120) for rendering the rendered character 1010 and / or a portion of the virtual character 1010. In some embodiments, if the user's geospatial location is unavailable or unknown, the server 110, gateway component 140, and / or the user device 120 may still display the virtual object 1010 using an estimation algorithm that estimates where a particular virtual object and / or physical object could be located, using the user's last known location as a function of time and / or other parameters. This may also be used to determine the location of any virtual object if the user's sensors are interfered with and / or experience other malfunctions.

[0059] In some embodiments, a virtual character or virtual object may have a virtual image, and the rendering of the virtual image is triggered by a physical object. For example, referring here to Figure 11, the virtual image 1110 may be triggered by an actual physical platform 1120. The triggering of the image 1110 may respond to a visual object or feature (e.g., reference point, design feature, geometric shape, pattern, physical location, altitude, etc.) detected by a user device or other components of the system 100. When a user views the platform 1120 without using a user device, the user sees the platform 1120 without the image 1110. However, when a user views the platform 1120 through a user device, the user sees the image 1110 on the platform 1120, as shown in Figure 11. The image 1110 is a virtual object and therefore may be stationary, active, change over time, change relative to the user's viewing position, or change depending on which particular user is viewing the image 1110. For example, if the user is a young child, the image may be a dog, and if the viewer is an adult male, the image may be a large robot as shown in Figure 11. These are examples of user-dependent and / or state-dependent experiences. This allows one or more users to perceive one or more virtual objects, individually and / or in combination with physical objects, and to experience customized and personalized versions of the virtual objects. The image 1110 (or a part thereof) may be rendered by various components of the system, including, for example, software / firmware installed on the user device. Using data indicating the position and orientation of the user device, combined with the alignment features of the virtual object (i.e., the image 1110), the virtual object (i.e., the image 1110) forms a relationship with a physical object (i.e., the platform 1120).For example, the relationship between one or more virtual objects and one or more physical objects may be a function of distance, positioning, time, geolocation information, proximity to one or more other virtual objects, and / or any other functional relationship including any kind of virtual data and / or physical data. In some embodiments, image recognition software in a user device may further enhance the digital-physical object relationship.

[0060] The interactive interfaces provided by the disclosed systems and methods may be implemented to facilitate various activities, such as interacting with one or more virtual environments and objects, interacting with other users, and experiencing various forms of media content, including advertisements, music concerts, and movies. Thus, the disclosed systems facilitate user interaction so that users not only watch or listen to media content, but rather actively participate in and experience it. In some embodiments, user participation may include modifying existing content or creating new content to be rendered in one or more virtual worlds. In some embodiments, the media content, and / or the user creating the content, may be themed around the creation of one or more virtual worlds.

[0061] In one embodiment, a musician (or other user) may create musical content that is rendered for a user interacting with a specific virtual world. The musical content may include, for example, various singles, EPs, albums, videos, short films, and concert performances. In one embodiment, multiple users may interface with system 100 to simultaneously experience a virtual concert performed by the musician.

[0062] In some embodiments, the media produced may include a unique identifier code associated with a specific entity (e.g., a band, artist, user, etc.). The code may be a set of alphanumeric characters, a UPC code, a QR code (registered trademark), a 2D image trigger, a 3D physical object feature trigger, or other form of digital mark, as well as sound, images, and / or both. In some embodiments, the code may also be embedded in digital media that can be interfaced using system 100. A user may obtain a code (e.g., by payment of a fee) and, in exchange for the code, access media content produced by the entity associated with the identifier code. The media content may be added to or removed from the user's interface.

[0063] In one embodiment, to avoid computational and bandwidth limitations in transferring real-time or near-real-time video data from one computing system to another (e.g., from a cloud computing system to a user-coupled local processor) with low latency, parameter information relating to various shapes and geometric forms may be transferred and used to define surfaces, while textures may be transferred and added to these surfaces, resulting in static or dynamic details such as bitmap-based video details of an individual's face mapped onto a parameterically reproduced facial geometry. In another embodiment, if the system is configured to recognize an individual's face and understands that the individual's avatar is located in an augmented world, the system may be configured to transfer relevant world information and individual avatar information in one relatively large setup transfer, and the remaining transfer to a local computing system, such as the 308 depicted in Figure 1, for local rendering may be limited to parameter and texture updates such as motion parameters of the individual's skeletal structure and motion bitmaps of the individual's face, with much less bandwidth compared to the initial setup transfer or the transfer of real-time video. Therefore, cloud-based and local computing assets may be used in an integrated manner in which the cloud handles computations that do not require relatively low latency, and local processing assets handle tasks where low latency is critical, in which case the form of data transferred to the local system is preferably delivered with relatively low bandwidth in a certain amount of such data (i.e., parameter information, textures, etc., compared to all real-time video).

[0064] Referring first to Figure 15, the schematic diagram illustrates the coordination between a cloud computing asset (46) and local processing assets (308, 120). In one embodiment, the cloud (46) asset is directly (40, 42) operably coupled to one or both of local computing assets (120, 308), such as a processor and memory configuration, which may be housed in a structure configured to be coupled to the user's head (120) or belt (308), for example, via wired or wireless networking (wireless is preferred for mobility, and wired for certain high bandwidth or large data capacity transfers). These computing assets, which are local to the user, may also be operably coupled to each other via wired and / or wireless connectivity configurations (44). In one embodiment, in order to maintain a low-inertia and compact head-mounted subsystem (120), the primary transfer between the user and the cloud (46) may also be via a link between a belt-based subsystem (308) and the cloud, and the head-mounted subsystem (120) is primarily data-tethered to the belt-based subsystem (308) using a wireless connection such as an ultra-wideband ("UWB") connection, as currently employed in personal computing peripheral connectivity applications.

[0065] With efficient local and remote processing coordination, and using a suitable display device for the user (e.g., the user interface 302 or user “display device” featured in Figure 3, the display display 14 described below with reference to Figure 14, or variations thereof), aspects of one world related to the user’s current actual or virtual location may be transferred to or “passed” to the user and efficiently updated. In fact, in one embodiment, if one individual uses a virtual reality system ("VRS") in augmented reality mode and another individual uses the VRS in full virtual mode to explore the same world local to the first individual, the two users may experience that world with each other in various ways. For example, with reference to Figure 12, a scenario similar to that described with reference to Figure 11 is depicted, with the addition of visualization of a second user’s avatar 2 flying through an augmented reality world depicted from a full virtual reality scenario. In other words, the scene depicted in Figure 12 may be experienced and displayed in augmented reality for the first individual, and in addition to the actual physical elements surrounding the local world within the scene, such as the ground, background buildings, and statue platform 1120, two augmented reality elements (the statue 1110 and the flying bumblebee avatar 2 of the second individual) are displayed. If avatar 2 flies through the world which is local to the first individual, dynamic updates may be used to allow the first individual to visualize the progress of avatar 2 of the second individual.

[0066] Again, using the configuration described above, where there is a single world model residing on cloud computing resources and from which it can be delivered, such a world could be "delivered" to one or more users in a relatively low-bandwidth format, which is preferable for distributing real-time video data or similar. The augmented experience of an individual standing near the statue (i.e., as shown in Figure 12) may be informed by the cloud-based world model, a subset of which may be delivered to that individual and their local display device to complete their view. An individual sitting in front of a remote display device, which may be as simple as a personal computer located on a desk, can efficiently download identical sections of information from the cloud and have them rendered on the display. In fact, one individual who is actually present in a park near the statue may bring a friend who is located remotely to take a walk in that park, with the friend participating through virtual and augmented reality. The system needs to know where the paths are, where the trees are, where the statue is, but with that information in the cloud, the participating friend can download the scenario aspects from the cloud and then begin walking together as augmented reality that is local to the individual who is actually in the park.

[0067] Referring to Figure 13, an embodiment based on time and / or chance parameters is depicted in which an individual interacting with a virtual reality and / or augmented reality interface such as the user interface 302 or user display device featured in Figure 3, the display device 14 described below with reference to Figure 14, or variations thereof is using the System (4) and enters a coffee shop to order a cup of coffee (6). The VRS may be configured to utilize locally and / or remote sensing and data collection capabilities to provide enhanced augmented reality and / or virtual reality display performance for the individual, such as a highlighted location on the coffee shop door or a bubble window of the relevant coffee menu (8). When the individual receives the cup of coffee they ordered, or when the System detects any other relevant parameters, the System may be configured to display one or more time-based augmented reality or virtual reality images, videos, and / or sounds (e.g., a view of the Madagascar jungle from the walls and ceiling, either static or dynamic, with or without jungle sounds and other effects) in the local environment using the display devices (10). Such presentation to the user may be interrupted based on timing parameters (i.e., 5 minutes after a full coffee cup is recognized and handed to the user, 10 minutes after the system recognizes a user walking through the front door of the store) or other parameters (e.g., the system's recognition that the user has finished drinking the coffee by noticing the inverted orientation of the coffee cup when the user takes the last sip of coffee from the cup, or the system's recognition that the user has left the front door of the store) (12).

[0068] Referring to Figure 14, one embodiment of a preferred user display device (14) is shown, comprising a display lens (82) that can be mounted on the user's head or eyes by a housing or frame (84). The display lens (82) comprises one or more transparent mirrors positioned by the housing (84) in front of the user's eye (20), the one or more transparent mirrors may be configured to reflect projected light (38) into the eye (20) to facilitate beam shaping, while also allowing some light to pass through from the local environment in an augmented reality configuration (in a virtual reality configuration, it may be desirable that the display system 14 be able to block substantially all light from the local environment by a darkened visor, a blackout curtain, a fully black LCD panel mode, or the like). In the embodiment depicted, two wide-field machine vision cameras (16) are coupled to the housing (84) to image the environment around the user. In one embodiment, these cameras (16) are dual-capture visible light / infrared cameras. The described embodiment also includes a pair of scanning laser wavefront shape (i.e., for depth) projector modules, along with a display mirror and optics configured to project light (38) into the eye (20), as shown. The described embodiment also includes two miniature infrared cameras (24) paired with infrared light sources (26, light-emitting diodes "LEDs" etc.) configured to track the user's eye (20) to assist rendering and user input. The system (14) further features a sensor assembly (39), which comprises X-axis, Y-axis, and Z-axis accelerometer capabilities, a magnetic compass, and X-axis, Y-axis, and Z-axis gyroscope capabilities, and is preferably capable of providing data at relatively high frequencies such as 200 Hz. The described system (14) also includes a head pose processor (36), such as an ASIC (Application-Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), and / or an ARM processor (Highly Reduced Instruction Set Machine), which may be configured to calculate real-time or near-real-time user head pose from wide-field image information output from a capture device (16).Also shown is another processor (32) configured to perform digital and / or analog processing to derive attitude from gyro, compass, and / or accelerometer data from the sensor assembly (39). The described embodiment also features a GPS (37, Global Positioning Satellite System) subsystem to assist in attitude and positioning. Finally, the described embodiment may feature hardware that activates a software program configured to provide rendering information local to the user to facilitate the operation of the scanner and imaging into the user's eye for the user's field of view of the world. The rendering engine (34) is operably coupled to the sensor attitude processor (32), image attitude processor (36), target tracking camera (24), and projection subsystem (18) so that the light of the rendered augmented reality and / or virtual reality objects is projected using a scanning laser array (18) as well as a retinal scanning display (81, 70, 76 / 78, 80, i.e., via a wired or wireless connection). The wavefront of the projected light ray (38) may be bent or focused to match the desired focal length of the augmented reality and / or virtual reality object. A mini infrared camera (24) may be used to track the eye to assist with rendering and user input (i.e., where the user is looking, at what depth the user is focusing, (the edge of the eye may be used to estimate the depth of focus, as discussed below)). A GPS (37), gyroscope, compass, and accelerometer (39) may be used to provide trajectory estimation and / or fast attitude estimation. Images and attitudes from the camera (16), along with data from associated cloud computing resources, may be used to map the local world and share the user's field of view with the virtual or augmented reality community.Most of the hardware within the display system (14) featured in Figure 14 is depicted as being directly coupled to a housing (84) adjacent to the display (82) and the user's eyes (20). However, the depicted hardware components may be mounted on other components, such as belt-mounted components, as shown in Figure 3, or housed within other components. In one embodiment, all components of the system (14) featured in Figure 14, except for the image pose processor (36), sensor pose processor (32), and rendering engine (34), are directly coupled to the display housing (84). Communication between the image pose processor (36), sensor pose processor (32), and rendering engine (34) and the remaining components of the system (14) may be via wireless communication such as ultra-wideband or wired communication. The depicted housing (84) is preferably head-mounted and wearable by the user. It may also feature a speaker, which may be inserted into the user's ear and used to provide the user with sounds that may be related to an augmented reality or virtual reality experience, such as jungle sounds as referenced in Figure 13, and a microphone, which may be used to capture sounds that are local to the user.

[0069] Regarding the projection of light (38) into the user's eye (20), in one embodiment, a mini-camera (24) may be used to measure the location where the center of the user's eye (20) geometrically touches, which generally coincides with the focal point of the eye (20), i.e., the “depth of focus.” The three-dimensional surface of all points touched by the eye is called a “holopter.” The focal length may exhibit a finite number of depths or may vary infinitely. Light projected from the convergence distance appears to be focused onto the target eye (20), while light in front of or behind the convergence distance is blurred. Furthermore, it has been found that spatially coherent light with a beam diameter of less than approximately 0.7 millimeters is correctly resolved by the human eye, regardless of where the eye focuses. With this understanding in mind, in order to create illusions with appropriate depth of field, eye convergence may be tracked using a mini-camera (24), and the rendering engine (34) and projection subsystem (18) may be used to render all objects on or near the focused holopter, and all other objects that are out of focus to varying degrees (i.e., using intentionally created blur). A translucent light guiding optical element configured to project coherent light into the eye may be provided by a supplier such as Lumus, Inc. Preferably, the system (14) renders to the user at a frame rate of about 60 frames per second or greater. As described above, preferably, the mini-camera (24) may be used for target tracking, and the software may be configured to take up not only the convergence geometry but also a cue of the focal position that serves as user input. Preferably, such a system is configured with brightness and contrast suitable for daytime or nighttime use. In one embodiment, such a system preferably has a latency of less than about 20 milliseconds for visual object alignment, angular alignment of less than about 0.1 degrees, and a resolution of about 1 minute, which is roughly the limit of the human eye.The display system (14) may be integrated with a localization system which may include GPS elements, optical tracking, a compass, an accelerometer, and / or other data sources to assist in position and orientation determination, and the localization information may be used to facilitate accurate rendering within the user's field of view of the relevant world (i.e., such information helps the glasses to understand where they are in relation to the real world).

[0070] Other suitable display devices include desktop and mobile computers, smartphones, smartphones which may be extended with additional software and hardware features to facilitate or simulate 3D far vision (for example, in one embodiment, a frame may be detachably coupled to the smartphone, and the frame features a subset of 200Hz gyro and accelerometer sensors, two small machine vision cameras with wide-field lenses, and an ARM processor to simulate some of the functionality of the configuration featured in Figure 14), tablet computers, and tablet computers which may be extended with smartphones as described above. This includes, but is not limited to, tablet computers augmented by additional processing and sensing hardware, head-mounted systems using smartphones and / or tablets to display augmented and virtual viewpoints (visual adaptation via magnifying optics, mirrors, contact lenses, or light-structured elements), non-transparent displays of light-emitting elements (LCDs, OLEDs, vertical-cavity surface-emitting lasers, LED laser beams, etc.), transparent displays that enable humans to see the natural world and artificially generated images simultaneously (e.g., light-guided optical elements, transparent and polarized OLEDs that shine into close-focus contact lenses, LED laser beams, etc.), contact lenses with light-emitting elements (such as those available from Innega, Inc (Bellevue, WA) under the trade name Ioptik RTM, which may be combined with special complementary spectacle components), implantable devices with light-emitting elements, and implantable devices that simulate photoreceptors in the human brain.

[0071] Using a system such as those depicted in Figures 3 and 14, 3D points can be captured from the environment, and the orientation of the camera capturing these images or points (i.e., vector and / or origin position information relative to the world) can be determined, so that these points or images may be "tagged" or associated with this orientation information. Points captured by a second camera may then be used to determine the orientation of the second camera. In other words, the second camera can be oriented and / or localized based on a comparison with the tagged image from the first camera. This knowledge may then be used to extract textures, create maps, and create virtual copies of the real world (because there are two cameras to be aligned around it). Thus, at a basic level, in one embodiment, there is a wearable system that can be used to capture both 3D points and the 2D images that generated the points, and these points and images may be sent to cloud storage and processing resources. These may also be cached locally along with embedded pose information (i.e., cache tagged images), so the cloud may have, along with the 3D points, tagged (i.e., tagged with 3D poses) 2D images that are ready (i.e., in the available cache). If the user is observing something dynamic, the user may send additional information to the cloud related to the motion (for example, if looking at another person's face, the user may take a texture map of the face and push it up to an optimized frequency, even if the surrounding world is otherwise essentially static).

[0072] The cloud system may be configured to store several points as posture-only references, thereby reducing overall posture tracking calculations. Generally, it may be desirable to have several contour features so that it is possible to track key items in the user's environment, such as walls and tables, as the user moves around a room, and the user may wish to "share" the world, allowing other users to enter the room and show them these points. Such useful and key points can be referred to as "references" because they are extremely useful as anchoring points. They relate to features that can be recognized by machine vision and can be consistently and repeatedly extracted from the world on different parts of the user's hardware. Therefore, these references may preferably be stored in the cloud for further use.

[0073] In one embodiment, since the reference is an item of a type that the camera can easily use to recognize its position, it is preferable to have a relatively even distribution of the reference throughout the relevant world.

[0074] In one embodiment, the relevant cloud computing configuration may be configured to periodically update the database of 3D points and any relevant metadata to use the best data from various users for both the refinement of the criteria and the creation of the world. In other words, the system may be configured to obtain the best dataset by using input from various users who look at the relevant world and work within it. In one embodiment, the database is fractal in nature, and as a user approaches an object, the cloud passes higher-resolution information to such user. As a user maps an object more precisely, that data is sent to the cloud, and the cloud can add new 3D points and image-based texture maps to the database if the new 3D points and image-based texture maps are better than those previously stored in the database. All of this may be configured to happen simultaneously from many users.

[0075] As described above, augmented reality or virtual reality experiences may be based on recognizing certain types of objects. For example, understanding that a particular object has depth may be important for recognizing and understanding that object. Recognition device software objects ("Recognition Devices") may be deployed on cloud or local resources to specifically assist in the recognition of various objects on either one or both platforms as a user navigates data within the world. For example, if a system has data of a world model comprising a 3D point cloud and pose-tagged images, and there is a desk with a large number of points on it, as well as an image of the desk, a person perceiving the desk may not make the determination that what is being observed is actually a desk. In other words, a few 3D points in space, and an image from somewhere in space showing most of the desk may not be sufficient for an instantaneous recognition that the desk is being observed. To assist in this identification, a specific object recognition device may be created that goes into the raw 3D point cloud, segments a set of points, and extracts, for example, the plane of the top surface of the desk. Similarly, a recognition device may be created to segment a wall from 3D points so that a user can change the wallpaper in virtual or augmented reality, or remove a portion of a wall, and have an entrance to another room that is not actually there in the real world. Such a recognition device operates within the data of a world model, crawls the world model, and such a recognition device may be thought of as a software “robot,” which implants semantic information, or ontology, of what is thought to exist between points in space, into its world model. Such recognition devices or software robots may be configured such that their entire existence relates to crawling the data of the relevant world and finding what is thought to be a wall, or a chair, or other item. They may be configured to tag sets of points with functional equivalence such as “this set of points belongs to a wall,” and may have a combination of point-based algorithms and pose-tagged image analysis to manually inform the system about what is in the points.

[0076] Object recognition devices may be created for a variety of purposes and uses, depending on the perspective. For example, in one embodiment, a coffee shop such as Starbucks may invest in creating an accurate recognition device for Starbucks coffee cups within a relevant world of data. Such a recognition device is configured to crawl large and small worlds of data and search for Starbucks coffee cups, thereby segmenting them and allowing them to be identified to a user as they operate within a relevant neighborhood space (i.e., perhaps, when a user sees a Starbucks coffee cup over a certain period of time, it can be quickly recognized when a user moves it to their desk. Such a recognition device may be configured to operate or function on local resources and data, or both cloud and local, as well as on cloud computing resources and data, depending on the available computing resources. In one embodiment, there is a global copy of the world model on the cloud to which millions of users contribute, but for smaller worlds or subworlds, such as the office of a particular individual in a particular town, the majority of the global world does not care what that office looks like, so the system may be configured to organize the data and move it to local cache information that is considered most locally relevant to a given user. In one embodiment, for example, when a user walks up to a desk, an object that is often identified as moving (e.g., a cup on the desk) does not need to burden the cloud model and impose a transmission burden between the cloud and local resources, so the relevant information (such as a segment of a particular cup on the desk) may be configured to reside only on local computing resources and not on the cloud.Therefore, cloud computing resources may be configured to segment 3D points and images, and thus decompose permanent (i.e., generally, non-moving) objects from movable ones, which affects where the relevant data remains, where it is processed, and removes the processing burden from wearable / local systems for certain data related to more permanent objects, then enables one-time processing of locations that can be shared with an infinite number of other users, and allows multiple data sources to simultaneously build databases of fixed and movable objects at a given physical location, and may segment objects from the background to create object-specific criteria and texture maps.

[0077] In one embodiment, the system may be configured to query the user for input regarding the identification of a particular object, so that the user can create the system and enable the system to associate semantic information with objects in the real world (for example, the system may ask the user questions such as, "Is that a Starbucks coffee cup?"). The ontology may provide guidance on what segmented objects from the world can do, how they behave, etc. In one embodiment, the system may feature a wirelessly connected keypad, connectivity to a smartphone keypad, or a virtual or physical keypad, such as the same, to facilitate specific user input to the system.

[0078] The system may be configured to share basic elements (such as the geometric shapes of walls, windows, and desks) with any user entering the room in virtual or augmented reality. In one embodiment, the individual's system is configured to capture images from a specific viewpoint and upload them to the cloud. The cloud then combines old and new sets of data, activates optimization routines, and establishes standards for each individual object.

[0079] GPS and other localization information may be used as input to such processing. Furthermore, other computing systems and data, such as personal online calendar or Facebook account information, may be used as input (for example, in one embodiment, the cloud and / or local system may be configured to analyze the contents of the user's calendar for flights, dates, and destinations so that information can be moved from the cloud to the user's local system over time to prepare for the user's arrival time at a given destination).

[0080] In one embodiment, tags such as QR codes (registered trademark) and similar tags may be inserted into the world for use in conjunction with non-statistical pose calculation, security / access control, transmission of special information, spatial messaging, non-statistical object recognition, etc.

[0081] In one embodiment, cloud resources may be configured to pass digital models of the real and virtual worlds between users, as described above with reference to a “passable world,” and the models are rendered by individual users based on parameters and textures. This reduces bandwidth compared to the transmission of real-time video, enables the rendering of virtual viewpoints of the scene, and allows millions of users to participate in a single virtual gathering without each of them having to send data (e.g., video) that they need to see, because their viewpoint is rendered by local computing resources.

[0082] A virtual reality system ("VRS") may be configured to register user position and field of view (both known as "attitude") through real-time metric computer vision using a camera, co-localization and mapping techniques, maps, and data from one or more of the following: sensors (e.g., gyroscope, accelerometer, compass, barometer, GPS), radio signal strength triangulation, signal time-of-flight analysis, LiDAR ranging, RADAR ranging, odometer, and sonar ranging. The wearable device system may be configured to map and orient simultaneously. For example, in an unknown environment, the VRS may be configured to collect information about the environment and examine images to provide a suitable reference point for user attitude calculation, other points for world modeling, and a texture map of the world. The reference point may be used to optically calculate the attitude. As the world is mapped in more detail, more objects may be segmented and given their own texture maps, but the world can still preferably be represented with a low spatial resolution as a simple polygon with a low-resolution texture map. Other sensors, such as those discussed above, may be used to assist this modeling effort. The world can be fractal in that moving (through viewpoints, "surveillance" modes, zooming, etc.) or otherwise seeking a better view requires high-resolution information from cloud resources. By getting closer to an object, higher-resolution data can be captured, which can be sent to the cloud, where the cloud can compute the new data and / or insert the new data into gaps in the world model.

[0083] Referring to Figure 16, the wearable system may be configured to capture image information and extract reference and recognized points (52). The wearable local system may calculate the pose using one of the pose calculation techniques described below. The cloud (54) may be configured to use images and references to segment 3D objects from a more static 3D background, where the images provide texture maps of the objects and the world (the textures may be real-time video). The cloud resource (56) may be configured to store and make available static references and textures for world alignment. The cloud resource may be configured to trim the point cloud for an optimal point density for alignment. The cloud resource (60) may be configured to store and make available object references and textures for object alignment and manipulation, where the cloud may trim the point cloud for an optimal density for alignment. The cloud resource may be configured to use all valid points and textures to generate a fractal solid model of the object (62), where the cloud may trim the point cloud information for an optimal reference density. Cloud resources (64) may be configured to query the user for tailoring the identification of segmented objects and worlds, and an ontology database may use the answers to infuse the objects and worlds with implementable properties.

[0084] The following specific alignment and mapping modes feature “O attitude,” representing the attitude determined by the optical or camera system; “s attitude,” representing the attitude determined by sensors (i.e., a combination of data such as GPS, gyroscope, compass, accelerometer, etc., as discussed above); and “MLC,” representing cloud computing and data management resources.

[0085] 1. Orientation: Create a base map for the new environment. Objective: To establish an attitude (or similar) when the environment is not mapped or not connected to an MLC. • Extract points from the image, track them from frame to frame, and triangulate the reference point using S-pose. Since there is no standard, the S posture is used. Based on sustainability, substandard criteria will be eliminated. This is the most basic mode. It always operates on low-precision attitudes. With minimal time and some relative motion, this establishes the minimum reference set for O-attitude and / or mapping. As soon as the posture is secure, it exits this mode.

[0086] 2. Mapping and O-pose: Mapping the environment Objective: To establish high-precision attitudes, map the environment, and provide the map (with images) to the MLC. • Calculate O-pose from mature global standards. Use S-pose as a check for O-pose resolution and to accelerate calculation (O-pose is a nonlinear gradient search). • Mature standards may originate from MLC or be locally determined. • Extract points from images, track them from frame to frame, and triangulate the standards using O-pose. Based on sustainability, substandard criteria will be eliminated. • Provide reference and posture tagged images to MLC. The last three steps do not need to happen in real time.

[0087] 3. Posture: Determining your posture Objective: To establish high-precision attitudes within a pre-mapped environment using minimal processing power. To estimate the pose at n, past S-poses and O-poses (n-1, n-2, n-3, etc.) are used. • Use the pose at n to project the reference onto the image captured at n, and then create an image mask from the projection. • Extract points from a masked area (by searching / extracting points only from a masked subset of the image, the processing load is greatly reduced). • Calculate the O-attitude from the extracted points and mature global standards. To estimate the posture at n+1, we use the S-posture and O-posture at n. • Optional: Provide posture-tagged images / videos to the MLC cloud.

[0088] 4. Super-resolution: Determining super-resolution images and standards. Objective: To create super-resolution images and references. • Synthesize pose-tagged images to create a super-resolution image. • Use super-resolution images to enhance reference position estimation. • Iterate through super-resolution criteria and O-pose estimation from images. • Optional: Loop the above steps (in real time) on a wearable device or (for a better world) on an MLC.

[0089] In one embodiment, the VLS system may be configured to have certain basic functionalities and functionalities facilitated by “apps” or applications that can be delivered through the VLS to provide certain special functionalities. For example, the following apps may be installed on the target VLS to provide special functionalities.

[0090] A painting-style rendering app. Artists create image transformations that represent the world as they see it. Users enable these transformations and thus view the world "through the artist's eyes."

[0091] A desktop modeling app. Users "construct" objects from physical objects placed on a table.

[0092] A virtual existence application. Users pass a virtual model of a space to other users, who then move around the space using virtual avatars.

[0093] An avatar emotion app. Measurements such as subtle vocal intonation, slight head movements, body temperature, and heart rate are used to create tangible effects on the virtual avatar. Digitizing human state information and transmitting it to a remote avatar uses less bandwidth than video. Furthermore, such data can be mapped to emotional, non-human avatars. For example, a dog avatar could show excitement by wagging its tail based on excited vocal intonation.

[0094] An efficient mesh network may be desirable for moving data, as opposed to sending everything back to the server. However, many mesh networks have suboptimal performance because their location information and topology are not well characterized. In one embodiment, the system may be used to determine the location of all users with relatively high accuracy, and therefore a mesh network configuration may be used for high performance.

[0095] In one embodiment, the system may be used for searching. Using augmented reality, for example, a user generates and leaves behind content relating to many aspects of the physical world. Much of this content is not text and is therefore not easily searchable by typical methods. The system may be configured to provide equipment for continuously tracking personal and social network content for search and reference purposes.

[0096] In one embodiment, if the display device tracks 2D points through a series of frames and then fits a vector-valued function to the time evolution of these points, it is possible to sample the vector-valued function at any point in time (e.g., between frames) or at some point in the near future (by projecting the vector-valued function in advance). This enables the creation of high-resolution post-processing and prediction of future poses before the next image is actually captured (e.g., it is possible to double the alignment speed without doubling the camera frame rate).

[0097] For body-fixed rendering (in contrast to head-fixed or world-fixed rendering), accurate representation of the body is desirable. In one embodiment, instead of measuring the body, its location can be derived through the average position of the user's head. If the user's face is facing forward most of the time, a multi-day average of the head position reveals its orientation. Combined with the gravity vector, this provides a reasonably stable coordinate system for body-fixed rendering. By using the current scale of the head position against this long-term coordinate system, consistent rendering of objects on / around the user's body is possible without the use of extra equipment. For the implementation of this embodiment, a single-registration average of the head direction vector may be started, and the cumulative data divided by delta t gives the current average head position. By keeping approximately five registrations, starting at n-5 days and continuing at n-4 days, n-3 days, n-2 days, and n-1 days, it is possible to use a rolling average of only the past "n" days.

[0098] In one embodiment, the scene may be scaled down and presented to the user in a smaller space than it actually is. For example, in situations where a scene must be rendered in a huge space (i.e., a soccer stadium), an equivalently huge space may not exist, or such a large space may be inconvenient for the user. In one embodiment, the system may be configured to scale down the scene so that the user can observe a scaled-down scene. For example, an individual could play a video game from a god's-eye view, or a match of the World Championship soccer, in an unscaled stadium, or in a stadium that is scaled down and presented on the floor of a living room. The system may be configured to simply shift the viewpoint, scale, and associated adaptive distance.

[0099] The system may also be configured to draw the user's attention to specific items within a presented scene by manipulating the focus of virtual reality or augmented reality objects, highlighting them, and changing their contrast, brightness, size, etc.

[0100] Preferably, the system may be configured to achieve the following modes:

[0101] Open space rendering: • Capture key points from the structured environment, and then fill in the spaces between them using ML rendering. • Potential venues: stages, output spaces, large indoor spaces (stadiums).

[0102] Object wrapping: • Recognize 3D objects in the real world and then extend them. • Here, "recognition" means identifying 3D blobs with a sufficiently high degree of accuracy to connect them to the image. There are two types of recognition: 1) classifying the type of object (e.g., "face"), and 2) classifying a specific instance of an object (e.g., Joe, an individual). • Create recognition device software objects for various things such as walls, ceilings, floors, faces, roads, the sky, skyscrapers, lunch houses, tables, chairs, cars, road signs, billboards, doors, windows, bookshelves, etc. Some recognition devices are Type I and have general functionality, such as "Put my video on that wall" or "That's a dog." Other recognition devices are Type II and have specific functionalities, such as "My TV is on the living room wall 3.2 feet from the ceiling" or "That's Fido" (these are more capable versions of general recognition devices). By constructing the recognition device as a software object, quantitative release of functionality and finer control over the user experience become possible.

[0103] Body-centered rendering • Renders virtual objects fixed to the user's body. Several devices, such as a digital tool belt, should float around the user's body. This requires knowing the location of the body as well as the head. A reasonably accurate desired position may be obtained by taking a long-term average of the user's head position (the head is usually facing forward parallel to the ground). A simple example is an object floating around a head.

[0104] Transparency / Breakdown Diagram • For Type II recognition objects, a fracture diagram is shown. • Link Type II recognized objects to an online database of 3D models. • You should start with objects that have commonly available 3D models, such as cars and public facilities.

[0105] Virtual existence • Project avatars of people in remote locations into an open space. ○ A subset of "Open Space Rendering" (above). ○ The user creates a rough geometric shape in their local environment and repeatedly sends both the geometric shape and texture map to others. ○ Users must grant permission for others to enter their environment. Subtle voice cues, hand tracking, and head movements are transmitted to a remote avatar. The avatar then visualizes these fuzzy inputs. ○ The above minimizes bandwidth. • Create an "entrance" to another room in the wall. ○As with other methods, you pass the geometric shape and texture map. Instead of displaying avatars within a local room, recognized objects (such as walls) are designated as entry points to other people's environments. In this way, multiple people can sit in their own rooms and view each other's environments "through" the walls.

[0106] Virtual viewpoint When a group of cameras (people) view a scene from different perspectives, a high-density digital model of the area is created. This rich digital model can be rendered from any favorable point that at least one camera can see. Example: People at a wedding. The scene is modeled collectively by all attendees. The recognition device distinguishes stationary objects from moving objects and creates texture maps for them (for example, walls have a stable texture map, while people have a higher frequency moving texture map). • With a wealth of digital models updated in real time, the scene can be rendered from any viewpoint. Attendees at the back can fly through the air to the front row for a better view. • Attendees can either show their moving avatar or hide their viewpoint. • Off-site attendees can find their "seat" using their avatar or, if permitted by the organizers, invisiblely. • This likely requires extremely high bandwidth. Conceptually, high-frequency data is streamed to the crowd over high-speed local radio. Low-frequency data originates from MLC. Since all attendees have high-precision location information, it is obvious that an optimal routing path for local networking can be created.

[0107] Messaging Simple, silent messaging may be preferable. • For this and other applications, it may be desirable to have a finger coding keyboard. • Tactile grab solutions may offer enhanced performance.

[0108] Fully virtual reality (VR): • When the vision system gets dark, it displays a view that does not overlap with the real world. • An alignment system is still necessary to track head position. • "Couch Mode" allows users to fly. • To prevent users from colliding with the real world, "walking mode" re-renders real-world objects as virtual ones. Rendering body parts is essential to believing in fiction. This suggests having a method for tracking and rendering body parts within the field of view. Non-transparent visors are a form of VR that offers many image quality improvements not possible with direct overlays. • A wide field of vision, perhaps even the ability to see behind you. • Various forms of "super vision": telescopes, X-ray vision, infrared, God's-eye view, etc.

[0109] In one embodiment, a system for a virtual user experience and / or an enhanced user experience is configured such that a remote avatar associated with a user can be animated, at least in part, based on data on a wearable device having input from sources such as voice intonation analysis and facial recognition analysis, which are performed by the relevant software module. For example, referring again to Figure 12, the bee avatar (2) may be animated to smile in a friendly manner, based on facial recognition of the user's smile, or based on a friendly tone of voice or intonation, which is determined by software configured to analyze voice input to a microphone that can locally capture voice samples from the user. Furthermore, the avatar character may be animated in a manner in which the avatar would express a particular emotion. For example, in an embodiment where the avatar is a dog, a happy smile or tone of voice detected by a system local to the human user may be represented by the dog avatar wagging its tail.

[0110] Various exemplary embodiments of the present invention are described herein. These embodiments are referenced in a non-limiting sense. They are provided to illustrate broader and more available aspects of the present invention. Various modifications may be made to the described invention, and similar ones may be substituted without departing from the true spirit and scope of the invention. In addition, many modifications may be made to adapt a particular situation, material, material composition, process, process act, or step to the object, spirit, or scope of the invention. Furthermore, as will be understood by those skilled in the art, each of the individual modifications described and illustrated herein has distinct components and features that can be readily separated from or combined with features of any of several other embodiments without departing from the scope or spirit of the invention. All such modifications are intended to be within the scope of the claims associated with this disclosure.

[0111] The present invention includes methods that may be performed using the device of interest. The methods may include the act of providing such a suitable device. Such provision may be performed by an end user. In other words, the act of “providing” simply requires the end user to acquire, access, approach, position, configure, activate, power on, or otherwise operate the essential device in the method of interest. The methods described herein may be performed not only in the order in which the events are described, but in any logically possible order of the described events.

[0112] Exemplary aspects of the present invention, along with details relating to material selection and manufacturing, are described above. Further details of the present invention are understood in connection with the patents and publications referenced above and are generally known or understandable to those skilled in the art. The same may apply to the method-based aspects of the present invention in terms of additional actions that may be generally or logically adopted.

[0113] In addition, while the present invention has been described with reference to several embodiments that incorporate various features as needed, the present invention is not limited to those described and indicated, as assumed with respect to each modification of the present invention. Various modifications may be made to the described invention, and equivalents (whether described herein or not for some simplicity) may be substituted without departing from the true spirit and scope of the invention. Furthermore, where a range of values ​​is provided, it should be understood that all intermediate values ​​between the upper and lower limits of that range, and any other specified or intermediate values ​​within that specified range, are encompassed within the present invention.

[0114] Furthermore, it is assumed that any features of the modifications of the invention described herein, as needed, may be described and claimed independently or in combination with one or more of the features described herein. References to singular items include the possibility that multiple identical items exist. More specifically, as used herein and in the claims associated herein, the singular forms “a,” “an,” “said,” and “the” include plural referents unless otherwise specified. In other words, the use of articles allows for “at least one” of the items of interest in the above description and in the claims associated with this disclosure. Furthermore, it should be noted that such claims may be drafted to exclude any optional elements. Thus, this statement is intended to function as an antecedent for the use of exclusive terms such as “simply,” “only,” and similar, or for the use of “negative” restrictions, in connection with the description of elements of the claims.

[0115] Without using such exclusive terminology, the term “comprising” in the claims associated with this disclosure shall allow for the inclusion of any additional elements, whether a given number of elements are enumerated in such claims or whether the addition of features can be considered a transformation of the properties of the elements described in such claims. Unless specifically defined herein, all technical and chemical terms used herein are given the broadest possible generally understood meaning while maintaining the validity of the claims.

[0116] The scope of the present invention is not limited to the provided examples and / or the subject specification, but rather is limited only by the scope of the claims language associated with this disclosure.

Claims

1. A system, wherein the system is A first user device configured to communicate with a computer network comprising one or more computing devices, wherein the one or more computing devices are One or more processors, A memory for storing instructions, wherein, when the instructions are executed by one or more processors, the memory causes one or more processors to process first virtual world data. A first user device comprising The computer network is equipped with, Receiving a first input from a first user via a user sensing system, The system receives a second input from the local environment of the first user device via the environmental sensing system. The computing devices are configured to perform the following, and one or more of the one or more computing devices are configured to generate second virtual world data based on the first virtual world data and further based on at least one of the first input and the second input. The first user device is further configured to present virtual content to the first user based on the second virtual world data. Presenting the virtual content to the first user includes presenting a visual rendering of the virtual content in 3D format. Presenting the visual rendering of the virtual content in the 3D format includes presenting the visual rendering on a display based on the position and orientation of the first user device, The second virtual world data described above includes virtual objects, The one or more computing devices are To predict the time and location of events associated with the virtual object for a second user, Communicating a written message to a second user device associated with the second user, wherein the written message includes the predicted time and place of the event for the second user. After communicating the written message to the second user device, the event is presented to the second user at the predicted time and place. A system including a server configured to perform this task.

2. The system according to claim 1, wherein the written message includes an email.

3. The system according to claim 1, wherein the written message includes a text message.

4. The system according to claim 1, wherein the written message includes an instant message.

5. The system according to claim 1, wherein the virtual object includes a vehicle, and the location of the event includes the location of the vehicle determined using geolocation mapping software.

6. The system according to claim 5, wherein the vehicle is configured to be operated by the first user.

7. The system according to claim 5, wherein the vehicle is configured to be operated by the second user.

8. The system according to claim 5, wherein the vehicle is configured to be operated by an intelligent software agent configured to run via the computer network in conjunction with one or more of the first user and the second user.

9. The system according to claim 5, wherein the vehicle is configured to be operated by a plurality of users remotely from the second user.

10. The server is To predict the second time and place of the event for the first user, Communicating a written message to the first user device, wherein the written message includes the predicted second time and place of the event for the first user, Presenting the event to the first user at the predicted second time and place. The system according to claim 1, configured to perform the following:

11. The system according to claim 10, wherein the first user and the second user are located in different geographical areas.

12. The system according to claim 10, wherein the first user and the second user are located in different time zones.

13. The system according to claim 10, wherein the first user and the second user are located in different countries.

14. The system according to claim 1, wherein the server is further configured to track the virtual object.

Citation Information

Patent Citations

  • Method and system for three-dimensional virtual reality space, medium and method for recording information, medium and method for transmitting information, information processing method, client terminal, and common-use server terminal

    JP1997081781A

  • Object controller, object control method and recording medium for the same

    JP2002063125A

  • Electronic shopping system for multiple persons

    JP2005182231A

  • Information processing method, information processor, and remote mixed reality sharing device

    JP2006293604A

  • Method and Apparatus to Facilitate a Differently Configured Virtual Reality Experience for Some Participants in a Communication Session

    US20080231626A1