System and method for augmented and virtual reality

The system facilitates interaction and dynamic updates in virtual worlds by processing user-generated data across a computer network, addressing the limitations of existing systems in enabling multi-user interaction and location-independent experiences.

JP2025122040AActive Publication Date: 2025-08-20MAGIC LEAP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025081379
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2011-10-28
Filing Date
2025-05-14
Publication Date
2025-08-20
Estimated Expiration
2032-10-29

AI Technical Summary

Technical Problem

Existing systems and methods for virtual and augmented reality environments do not effectively enable interaction between multiple users in different or same physical locations, and do not allow for dynamic changes in the virtual world based on physical objects or user interactions.

Method used

A system comprising a computer network with computing devices and software that processes virtual world data, allowing transmission of user-generated data to other users, enabling interaction and dynamic changes in the virtual world based on physical objects, and supporting augmented and virtual reality modes.

Benefits of technology

Enables seamless interaction and dynamic updates in virtual worlds across multiple users, regardless of location, enhancing user experience through real-time data processing and rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025122040000001
    Figure 2025122040000001
  • Figure 2025122040000002
    Figure 2025122040000002
  • Figure 2025122040000003
    Figure 2025122040000003
Patent Text Reader

Abstract

To provide systems and methods configured to facilitate interactive virtual or augmented reality environments for one or more users.SOLUTION: One embodiment is directed to a system for enabling two or more users to interact within a virtual world comprising virtual world data, the system comprising a computer network with one or more computing devices comprising memory, processing circuitry, and software that is stored at least in part in the memory and executable by the processing circuitry to process at least a portion of the virtual world data. At least a first portion of the virtual world data originates from a first user virtual world local to a first user, and the computer network is operable to transmit the first portion to a user device for presentation to a second user.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Related application data) This application claims priority under U.S.C. §119 to U.S. Provisional Patent Application No. 61 / 552,941, filed October 28, 2011, which is hereby incorporated by reference in its entirety into this application.

[0002] BACKGROUND OF THE INVENTION The present invention relates generally to systems and methods configured to facilitate an interactive virtual or augmented reality environment for one or more users. [Background technology]

[0003] (background) Virtual and augmented reality environments are generated by a computer using, in part, data that represents the environment. This data may represent, for example, various objects that a user may sense and interact with. Examples of these objects include objects that are rendered and displayed for the user to see, audio that is played for the user to hear, and tactile (or haptic) feedback that the user feels. A user may sense and interact with virtual and augmented reality environments through various visual, auditory, and tactile means. Summary of the Invention [Means for solving the problem]

[0004] One embodiment relates to a system for enabling two or more users to interact with a virtual world including virtual world data, the system comprising a computer network with one or more computing devices, the one or more computing devices comprising a memory, a processing circuit, and software at least partially stored in the memory and executable by the processing circuit to process at least a portion of the virtual world data, wherein at least a first portion of the virtual world data originates in a first user virtual world local to the first user, and the computer network is operable to transmit the first portion to the user device for presentation to a second user, whereby the second user may experience the first portion from the second user's location, effectively passing aspects of the first user virtual world to the second user. The first user and the second user may be in different physical locations or substantially the same physical location. At least a portion of the virtual world may be configured to change in response to changes in the virtual world data. At least a portion of the virtual world may be configured to change in response to physical objects sensed by the user device. The changes in the virtual world data may represent virtual objects having a predetermined relationship to the physical objects. The changes in the virtual world data may be presented to a second user device for presentation to a second user according to the predetermined relationship. The virtual world may be operable to be rendered by at least one of a computer server or the user device. The virtual world may be presented in a two-dimensional format. The virtual world may be presented in a three-dimensional format. The user device may be operable to provide an interface for enabling interaction between a user and the virtual world in an augmented reality mode. The user device may be operable to provide an interface for enabling interaction between a user and the virtual world in a virtual reality mode. The user device may be operable to provide an interface for enabling interaction between a user and the virtual world in a combination of the augmented reality mode and the virtual reality mode.The virtual world data may be transmitted over a data network. The computer network may be operable to receive at least a portion of the virtual world data from the user device. The at least a portion of the virtual world data transmitted to the user device may comprise instructions for generating at least a portion of the virtual world. The at least a portion of the virtual world data may be transmitted to a gateway for at least one of processing or distribution. At least one of the one or more computer servers may be operable to process the virtual world data distributed by the gateway.

[0005] Another embodiment is directed to a system for virtual and / or augmented user experiences, in which remote avatars are animated based at least in part on data on a wearable device, with optional input from voice intonation and facial recognition software.

[0006] Another embodiment is directed to a system for virtual and / or augmented user experiences, in which a camera pose or viewpoint position and vector may be located anywhere within a world sector.

[0007] Another embodiment is directed to a system for virtual and / or augmented user experiences, in which a world or portions thereof may be rendered at various and selectable scales for an observing user.

[0008] Another embodiment is directed to a system for virtual and / or augmented user experiences, in which features such as points or parametric lines, in addition to pose-tagged images, may be utilized as the basis for a world model, from which a software robot or object recognizer may create parametric representations of real-world objects that tag source features for inclusion in segmented objects and the world model. The present invention provides, for example, the following. (Item 1) 1. A system for enabling two or more users to interact in a virtual world including virtual world data, the system comprising: a computer network comprising one or more computing devices, the one or more computing devices comprising a memory, a processing circuit, and software at least partially stored in the memory and executable by the processing circuit to process at least a portion of the virtual world data; wherein at least a first portion of the virtual world data originates from a first user virtual world that is local to a first user, and wherein the computer network is operable to transmit the first portion to a user device for presentation to a second user, whereby the second user may also experience the first portion from the second user's location, effectively passing aspects of the first user virtual world to the second user. (Item 2) Item 10. The system of item 1, wherein the first user and the second user are in different physical locations. (Item 3) Item 10. The system of item 1, wherein the first user and the second user are in substantially the same physical location. (Item 4) Item 10. The system of item 1, wherein at least a portion of the virtual world changes in response to changes in the virtual world data. (Item 5) Item 10. The system of item 1, wherein at least a portion of the virtual world changes in response to physical objects sensed by the user device. (Item 6) Item 6. The system of item 5, wherein the changes in the virtual world data represent virtual objects having a predetermined relationship with the physical objects. (Item 7) 7. The system of claim 6, wherein the changes in the virtual world data are presented to a second user device for presentation to the second user in accordance with the predetermined relationship. (Item 8) Item 10. The system of item 1, wherein the virtual world is operable to be rendered by at least one of the computer server or a user device. (Item 9) Item 10. The system of item 1, wherein the virtual world is presented in two-dimensional format. (Item 10) Item 10. The system of item 1, wherein the virtual world is presented in three-dimensional format. (Item 11) Item 10. The system of claim 1, wherein the user device is operable to provide an interface to enable interaction between a user and the virtual world in an augmented reality mode. (Item 12) Item 10. The system of item 1, wherein the user device is operable to provide an interface in a virtual reality mode to enable interaction between a user and the virtual world. (Item 13) Item 12. The system of item 11, wherein the user device is operable to provide an interface to enable interaction between a user and the virtual world in a combination of augmented reality and virtual reality modes. (Item 14) Item 10. The system of item 1, wherein the virtual world data is transmitted over a data network. (Item 15) 2. The system of claim 1, wherein the computer network is operable to receive at least a portion of the virtual world data from a user device. (Item 16) Item 10. The system of claim 1, wherein at least a portion of the virtual world data transmitted to the user device comprises instructions for generating at least a portion of the virtual world. (Item 17) Item 10. The system of claim 1, wherein at least a portion of the virtual world data is transmitted to a gateway for at least one of processing or distribution. (Item 18) Item 18. The system of item 17, wherein at least one of the one or more computer servers is operable to process virtual world data delivered by the gateway. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 illustrates an exemplary embodiment of the disclosed system for facilitating an interactive virtual reality or augmented reality environment for multiple users. [Figure 2] FIG. 2 illustrates an example of a user device that interacts with the system illustrated in FIG. [Figure 3] FIG. 3 illustrates an exemplary embodiment of a mobile wearable user device. [Figure 4] FIG. 4 illustrates an example of objects viewed by a user when the mobile wearable user device of FIG. 3 is operating in an enhanced mode. [Figure 5] FIG. 5 illustrates an example of objects viewed by a user when the mobile wearable user device of FIG. 3 is operating in virtual mode. [Figure 6] FIG. 6 illustrates an example of objects viewed by a user when the mobile wearable user device of FIG. 3 is operating in a mixed virtual interface mode. [Figure 7] FIG. 7 illustrates an embodiment in which two users located in different geographic locations each interact with the other user and a common virtual world through their respective user devices. [Figure 8] FIG. 8 illustrates an embodiment in which the embodiment of FIG. 7 is extended to include the use of a haptic device. [Figure 9A]FIG. 9A illustrates an example of a mixed mode interface connection in which a first user interfaces with a digital world in a mixed virtual interface mode and a second user interfaces with the same digital world in a virtual reality mode. [Figure 9B] FIG. 9B illustrates another example of a mixed mode interface in which a first user interfaces with a digital world in a mixed virtual interface mode and a second user interfaces with the same digital world in an augmented reality mode. [Figure 10] FIG. 10 illustrates an exemplary illustration of a user's field of view when interfacing with the system in augmented reality mode. [Figure 11] FIG. 11 illustrates an example illustration of a user's view showing virtual objects triggered by physical objects when the user is interfacing with the system in augmented reality mode. [Figure 12] FIG. 12 illustrates one embodiment of an integrated augmented reality and virtual reality configuration in which one user in an augmented reality experience visualizes the presence of another user in a virtual reality experience. [Figure 13] FIG. 13 illustrates one embodiment of a time and / or contingency-based augmented reality experience configuration. [Figure 14] FIG. 14 illustrates one embodiment of a user display configuration suitable for virtual reality and / or augmented reality experiences. [Figure 15] FIG. 15 illustrates one embodiment of local and cloud-based computational collaboration. [Figure 16] FIG. 16 illustrates various aspects of the alignment configuration. DETAILED DESCRIPTION OF THE INVENTION

[0010] (Detailed explanation) Referring to Figure 1, system 100 is exemplary hardware for implementing the processes described below. This exemplary system includes a computing network 105 consisting of one or more computer servers 110 connected through one or more high-bandwidth interfaces 115. The servers in the computing network need not be co-located. Each of the one or more servers 110 includes one or more processors for executing program instructions. The servers also include memory for storing program instructions and data used and / or generated by processes being executed by the servers under the direction of the program instructions.

[0011] The computing network 105 communicates data among servers 110 and between the servers and one or more user devices 120 over one or more data network connections 130. Examples of such data networks include, but are not limited to, all types of public and private data networks, both mobile and wired, including, for example, the many interconnections of such networks commonly referred to as the Internet. No particular media, topology, or protocol is intended to be implied by the diagram.

[0012] The user devices are configured to communicate directly with either the computing network 105 or the server 110. Alternatively, the user devices 120 communicate with the remote server 110, and, if necessary, with other user devices locally, through a specially programmed local gateway 140 for processing data and / or communicating data between the network 105 and one or more local user devices 120.

[0013] As illustrated, gateway 140 is implemented as a separate hardware component including a processor for executing software instructions and memory for storing software instructions and data. The gateway has its own wired and / or wireless connection to a data network for communicating with servers 110 that comprise computing network 105. Alternatively, gateway 140 can be integrated with user device 120 worn or carried by a user. For example, gateway 140 may be implemented as a downloadable software application that is installed and operates on a processor included in user device 120. Gateway 140, in one embodiment, provides access to computing network 105 to one or more users via data network 130.

[0014] The servers 110 each include, for example, working memory and storage devices for storing data and software programs, a microprocessor for executing program instructions, and graphics processors and other specialized processors for rendering and generating graphics, images, video, audio, and multimedia files. The computing network 105 may also include devices for storing data that is accessed, used, or created by the servers 110.

[0015] Software programs running on the server, and optionally the user device 120 and gateway 140, are used to generate a digital world (also referred to herein as a virtual world) in which a user interacts with the user device 120. The digital world is represented by data and processes that represent and / or define virtual, non-existent entities, environments, and conditions that may be presented to a user through the user device 120 for the user to experience and interact with. For example, any type of object, entity, or item that appears to be physically present when instantiated in a scene viewed or experienced by a user may include a description of its appearance, its behavior, how the user is allowed to interact with it, and other characteristics. Data used to create the environment of a virtual world (including virtual objects) may include, for example, atmospheric data, terrain data, weather data, temperature data, location data, and other data used to define and / or represent the virtual environment. Additionally, the data defining the various conditions governing the operation of the virtual world may include, for example, laws of physics, time, spatial relationships, and other data that may be used to define and / or create the various conditions governing the operation of the virtual world (including virtual objects).

[0016] Entities, objects, conditions, properties, behaviors, or other features of the digital world are generally referred to herein as objects (e.g., digital objects, virtual objects, rendered physical objects, etc.) unless the context dictates otherwise. Objects may be any type of animate or inanimate object, including, but not limited to, structures, plants, vehicles, people, animals, living things, machines, data, video, text, photographs, and other users. Objects may also be defined in the digital world to store information about items, behaviors, or conditions that actually exist in the physical world. Data that represents or defines an entity, object, or item, or that stores its current state, is generally referred to herein as object data. This data is processed by the server 110, or by the gateway 140 or user device 120, depending on the implementation, to instantiate an instance of the object and render the object in an appropriate manner for a user to experience through the user device.

[0017] Programmers who develop and / or create a digital world create or define objects and the conditions under which the objects are instantiated. However, a digital world may allow others to create or modify objects. Once an object is instantiated, the state of the object may be permitted to be changed, controlled, or manipulated by one or more users experiencing the digital world.

[0018] For example, in one embodiment, the development, production, and management of the digital world is generally provided by one or more system administration programmers. In some embodiments, this may include the development, design, and / or execution of storylines, themes, and events in the digital world, as well as the distribution of the story through various forms of events and media, such as, for example, film, digital, network, mobile, augmented reality, and live entertainment. The system administration programmer may also handle the technical management, moderation, and curation of the digital world and its associated user community, as well as other tasks typically performed by a network administrator.

[0019] Users interact with one or more digital worlds using some type of local computing device, generally designated as user device 120. Examples of such user devices include, but are not limited to, smartphones, tablet devices, heads-up displays (HUDs), game consoles, or any other device capable of communicating data and providing an interface or display to a user, or a combination of such devices. In some embodiments, user device 120 may include or communicate with local peripheral or input / output components, such as, for example, a keyboard, a mouse, a joystick, a game controller, a haptic interface device, a motion capture controller, an optical tracking device (e.g., those available from Leap Motion, Inc. or from Microsoft under the trade name Kinect (RTM)), audio equipment, voice equipment, a projector system, a 3D display, and holographic 3D contact lenses.

[0020] An example of a user device 120 for interacting with system 100 is illustrated in Figure 2. In the exemplary embodiment shown in Figure 2, a user 210 may interface with one or more digital worlds through a smartphone 220. The gateway is implemented by a software application 230 stored and running on smartphone 220. In this particular example, data network 130 includes a wireless mobile network that connects the user device (i.e., smartphone 220) to computer network 105.

[0021] In one implementation of a preferred embodiment, system 100 is capable of supporting a large number of concurrent users (e.g., millions of users) who each interface with the same digital world or multiple digital worlds using some type of user device 120.

[0022] The user device provides the user with an interface to enable visual, audible, and / or physical interaction between the user and the digital world generated by the server 110, including other users and objects (real or virtual) presented to the user. The interface provides the user with a rendered view that can be seen, heard, or otherwise sensed, and the ability to interact with that view in real time. The manner in which the user interacts with the rendered view may be dictated by the capabilities of the user device. For example, if the user device is a smartphone, user interaction may be implemented by the user touching a touchscreen. In another example, if the user device is a computer or game console, user interaction may be implemented using a keyboard or game controller. The user device may include additional components that enable user interaction, such as sensors, and objects and information (including gestures) detected by the sensors may be provided as input representing the user's interaction with the virtual world using the user device.

[0023] The rendered view can be presented in a variety of formats, such as, for example, two-dimensional or three-dimensional visual displays (including projections), sound, and haptic or tactile feedback. The rendered view may be interfaced with by the user in one or more modes, including, for example, augmented reality, virtual reality, and combinations thereof. The format of the rendered view, as well as the interface mode, may be dictated by one or more of the user device, data processing power, user device connectivity, network capacity, and system workload. Having multiple users simultaneously interacting with the digital world and the real-time nature of the data exchange is made possible by the computing network 105, the server 110, the gateway component 140 (if necessary), and the user device 120.

[0024] In one example, computing network 105 is comprised of a large-scale computing system having single and / or multi-core servers (i.e., servers 110) connected through high-speed connections (e.g., high-bandwidth interfaces 115). Computing network 105 may form a cloud or grid network. Each of the servers includes memory or is coupled with computer-readable memory for storing software for implementing data to create, design, modify, or process objects in the digital world. These objects and their instantiations may be dynamic, appearing, disappearing, changing over time, and changing in response to other conditions. Examples of object dynamic capabilities are generally discussed herein with respect to various embodiments. In some embodiments, each user that interfaces with system 100 may also be represented as an object and / or collection of objects within one or more digital worlds.

[0025] The servers 110 in the computing network 105 also store computational state data for each of the digital worlds. Computational state data (also referred to herein as state data) may be a component of object data and generally defines the state of an instance of an object at a given instance in time. Thus, computational state data may change over time and may be affected by the actions of one or more users and / or programmers maintaining system 100. When a user affects computational state data (or other data that make up a digital world), the user directly modifies or otherwise manipulates the digital world. If the digital world is shared with or interfaces with other users, the user's actions may affect what is experienced by other users who interact with the digital world. Thus, in some embodiments, changes to the digital world made by a user are experienced by other users who interface with system 100.

[0026] Data stored on one or more servers 110 in the computing network 105 is transmitted or deployed to one or more user devices 120 and / or gateway components 140, in one embodiment, at high speed and with low latency. In one embodiment, object data shared by a server may be complete or compressed and include instructions for recreating the complete object data at the user's end, which may be rendered and visualized by the user's local computing device (e.g., gateway 140 and / or user device 120). Software running on the servers 110 of the computing network 105 may, in some embodiments, adapt the data that the computing network 105 generates and sends to a particular user's device 120 about objects in the digital world (or any other data exchanged by the computing network 105) depending on the user's particular device and bandwidth. For example, as a user interacts with the digital world through a user device 120, the server 110 may recognize the particular type of device being used by the user, the device's connectivity, and / or the available bandwidth between the user device and the server, and appropriately size and balance the data being delivered to the device to optimize the user interaction. An example of this may include reducing the size of the transmitted data to a lower-resolution quality so that the data can be displayed on a particular user device with a lower-resolution display. In a preferred embodiment, the computing network 105 and / or gateway component 140 delivers data to the user device 120 at 15 frames per second or faster and at a rate sufficient to present an interface operating at high-definition quality or higher resolution.

[0027] The gateway 140 provides local connectivity to the computing network 105 for one or more users. In some embodiments, it may be implemented by a downloadable software application running on the user device 120 or another local device such as that shown in FIG. 2 . In other embodiments, it may be implemented by a hardware component (a component having a processor with appropriate software / firmware stored thereon) that is in communication with the user device 120 but is either integrated into, not attached to, or integrated into the user device 120. The gateway 140 communicates with the computing network 105 via the data network 130 and provides data exchange between the computing network 105 and one or more local user devices 120. As discussed in more detail below, the gateway component 140 may include software, firmware, memory, and processing circuitry and may be capable of processing data communicated between the network 105 and one or more local user devices 120.

[0028] In some embodiments, the gateway component 140 monitors and adjusts the rate at which data is exchanged between the user device 120 and the computer network 105 to enable optimal data throughput for a particular user device 120. For example, in some embodiments, the gateway 140 buffers and downloads both static and dynamic aspects of the digital world, even beyond the field of view presented to the user through an interface connected to the user device. In such embodiments, instances of static objects (structured data, software-implemented methods, or both) may be stored in memory (local to the gateway component 140, the user device 120, or both) and are referenced to the local user's current location as indicated by data provided by the computing network 105 and / or the user's device 120. Instances of dynamic objects, which may include, for example, intelligent software agents and objects controlled by other users and / or the local user, are stored in a high-speed memory buffer. Dynamic objects, which represent two-dimensional or three-dimensional objects within the view presented to the user, can be categorized into component shapes, such as, for example, static shapes that move but do not change, and dynamic shapes that change. Portions of the changing dynamic objects can be updated by a real-time threaded high-priority data stream from server 110 over computing network 105, managed by gateway component 140. As an example of a priority threaded data stream, data within 60 degrees of the user's eye field of view may be given higher priority than data further out. Another example includes prioritizing dynamic characters and / or objects within the user's field of view over static objects in the background.

[0029] In addition to managing the data connection between the computing network 105 and the user device 120, the gateway component 140 may store and / or process data that may be presented to the user device 120. For example, the gateway component 140, in some embodiments, may receive compressed data from the computing network 105, e.g., representing graphical objects to be rendered for viewing by a user, and may perform advanced rendering techniques to reduce the data load transmitted from the computing network 105 to the user device 120. In another example where the gateway 140 is a separate device, the gateway 140 may store and / or process data for local instances of objects rather than communicating the data to the computing network 105 for processing.

[0030] Referring now also to FIG. 3 , the digital world may be experienced by one or more users in various forms, which may depend on the capabilities of the user's device. In some embodiments, user device 120 may include, for example, a smartphone, a tablet device, a head-up display (HUD), a gaming console, or a wearable device. Generally, a user device includes a processor for executing program code stored in memory on the device, coupled to a display, and a communication interface. An exemplary embodiment of a user device is illustrated in FIG. 3 , which includes a mobile wearable device, i.e., a head-mounted display system 300. According to an embodiment of the present disclosure, head-mounted display system 300 includes a user interface 302, a user sensing system 304, an environmental sensing system 306, and a processor 308. While processor 308 is shown in FIG. 3 as a standalone component separate from head-mounted system 300, in alternative embodiments, processor 308 may be integrated with one or more components of head-mounted system 300 or incorporated into other system 100 components, such as gateway 140, for example.

[0031] The user device presents the user with an interface 302 for interacting with and experiencing the digital world. Such interactions may include the user and the digital world, one or more other users interfacing with system 100, and objects within the digital world. Interface 302 generally provides visual and / or audio sensory input (and in some embodiments, physical sensory input) to the user. Accordingly, interface 302 may include speakers (not shown) and, in some embodiments, a display component 303 that can enable stereoscopic 3D viewing and / or 3D viewing that embodies the more natural characteristics of the human visual system. In some embodiments, display component 303 may comprise a transparent interface (such as a transparent OLED) that, when in an “off” setting, enables an optically correct view of the user's surrounding physical environment with little optical distortion or computing overlay. As discussed in more detail below, interface 302 may include additional settings that enable various visual / interface capabilities and functionality.

[0032] The user sensing system 304, in some embodiments, may include one or more sensors 310 operable to detect specific characteristics, properties, or information related to an individual user wearing the system 300. For example, in some embodiments, the sensors 310 may include a camera or optical detection / scanning circuitry capable of detecting real-time optical properties / measurements of the user, such as one or more of pupil constriction / dilation, angular measurement / position of each pupil, sphericity, eye shape (as eye shape changes over time), and other anatomical data. This data may provide or be used to calculate information (e.g., the user's visual focus) that can be used by the head-worn system 300 and / or interface system 100 to optimize the user's viewing experience. For example, in one embodiment, the sensors 310 may each measure the pupil constriction rate of each of the user's eyes. This data may be transmitted to the processor 308 (or to the gateway component 140 or to the server 110), where it is used to determine the user's response to, for example, the brightness setting of the interface display 303. The interface 302 may be adjusted according to the user's response, for example, by dimming the display 303 if the user's response indicates that the brightness level of the display 303 is too high. The user sensing system 304 may include other components other than those discussed above or illustrated in FIG. 3 . For example, in some embodiments, the user sensing system 304 may include a microphone for receiving audio input from the user. The user sensing system may also include one or more infrared camera sensors, one or more visible spectrum camera sensors, structured light emitters and / or sensors, infrared light emitters, coherent light emitters and / or sensors, gyros, accelerometers, magnetometers, proximity sensors, GPS sensors, ultrasonic emitters and detectors, and a tactile interface.

[0033] The environmental sensing system 306 includes one or more sensors 312 for acquiring data from the physical environment around the user. Objects or information detected by the sensors may be provided as input to the user device. In some embodiments, this input may represent user interaction with the virtual world. For example, a user viewing a virtual keyboard on a desk may gesture with their fingers as if they were typing on the virtual keyboard. The moving finger movements may be captured by the sensors 312 and provided as input to the user device or system, which may be used to change the virtual world or create new virtual objects. For example, the finger movements may be recognized (using a software program) as typing, and the recognized typing gestures may be combined with known locations of virtual keys on the virtual keyboard. The system may then render a virtual monitor that is displayed to the user (or other users interfacing with the system), displaying the text being typed by the user.

[0034] The sensor 312 may include, for example, a generally outward-facing camera or scanner to interpret sight information, for example, through continuously and / or intermittently projected infrared structured light. The environmental sensing system 306 may be used to map one or more elements of the user's surrounding physical environment by detecting and registering static objects, dynamic objects, people, gestures, and the local environment, including various lighting, atmospheric, and acoustic conditions. Thus, in some embodiments, the environmental sensing system 306 may include image-based 3D reconstruction software that is incorporated into a local computing system (e.g., the gateway component 140 or the processor 308) and operable to digitally reconstruct one or more objects or information detected by the sensor 312. In one exemplary embodiment, the environmental sensing system 306 provides one or more of motion capture data (including gesture recognition), depth sensing, facial recognition, object recognition, unique object feature recognition, voice / audio recognition and processing, sound source localization, noise reduction, infrared or similar laser projection, as well as monochrome and / or color CMOS sensors (or other similar sensors), field of view sensors, and various other light-enhancing sensors. It should be understood that the environmental sensing system 306 may include other components besides those discussed above or illustrated in FIG. 3 . For example, in some embodiments, the environmental sensing system 306 may include a microphone for receiving sound from the local environment. The user sensing system may also include one or more infrared camera sensors, one or more visible spectrum camera sensors, structured light emitters and / or sensors, infrared light emitters, coherent light emitters and / or sensors, gyros, infrared light emitters, accelerometers, magnetometers, proximity sensors, GPS sensors, ultrasonic emitters and detectors, and a haptic interface.

[0035] As mentioned above, processor 308, in some embodiments, may be integrated with other components of head-mounted system 300, integrated with other components of interface system 100, or may be a stand-alone device (wearable or separate from the user) as shown in FIG. 3. Processor 308 may be connected to various components of head-mounted system 300 and / or components of interface system 100 through a physical wired connection or through a wireless connection, such as, for example, a mobile network connection (including cellular and data networks), Wi-Fi, or Bluetooth. Processor 308 may include memory modules, integrated and / or additional graphics processing units, wireless and / or wired Internet connectivity, and codecs and / or firmware capable of converting data from sources (e.g., computing network 105, user sensing system 304, environmental sensing system 306, or gateway component 140) into image and audio data, which may be presented to the user via interface 302.

[0036] The processor 308 handles data processing for the various components of the head-mounted system 300, as well as data exchange between the head-mounted system 300 and the gateway component 140 (and in some embodiments, the computing network 105). For example, the processor 308 may be used to buffer and process data streaming between a user and the computing network 105, thereby enabling a smooth, continuous, and high-fidelity user experience. In some embodiments, the processor 308 may process data at a rate sufficient to achieve anywhere between 8 frames per second at 320x240 resolution to 24 frames per second at high-definition resolution (1280x720), or more (e.g., 60-120 frames per second and 4k resolution or higher (10k+ resolution and 50,000 frames per second)). Additionally, the processor 308 may store and / or process data that may be presented to a user rather than being streamed in real time from the computing network 105. For example, the processor 308, in some embodiments, may receive compressed data from the computing network 105 and perform advanced rendering techniques (such as lighting or shading) to reduce the data load transmitted from the computing network 105 to the user device 120. In another example, the processor 308 may store and / or process local object data rather than transmitting the data to the gateway component 140 or to the computing network 105.

[0037] The head-mounted system 300, in some embodiments, may include various settings or modes that enable different visual / interface capabilities and functionality. The modes may be selected manually by the user or automatically by components of the head-mounted system 300 or the gateway component 140. As mentioned above, one example of the head-mounted system 300 includes an “off” mode in which the interface 302 does not provide substantially any digital or virtual content. In the off mode, the display component 303 may be transparent, thereby enabling an optically correct view of the user's surrounding physical environment with little or no optical distortion or computing overlay.

[0038] In one exemplary embodiment, head-mounted system 300 includes an "augmented" mode in which interface 302 provides an augmented reality interface. In the augmented mode, interface display 303 may be substantially transparent, thereby allowing the user to view the local physical environment. Simultaneously, virtual object data provided by computing network 105, processor 308, and / or gateway component 140 is presented on display 303 in combination with the local physical environment.

[0039] 4 illustrates an example embodiment of objects viewed by a user when interface 302 is operating in augmented mode. As shown in FIG. 4, interface 302 presents physical object 402 and virtual object 404. In the embodiment illustrated in FIG. 4, physical object 402 is an actual physical object that exists in the user's local environment, while virtual object 404 is an object created by system 100 and displayed via user interface 302. In some embodiments, virtual object 404 may be displayed at a fixed position or location within the physical environment (e.g., a virtual monkey standing next to a particular road sign located in the physical environment) or may be displayed to the user as an object located at a position relative to user interface / display 303 (e.g., a virtual clock or thermometer visible in the upper left corner of display 303).

[0040] In some embodiments, the virtual object may be cue from or triggered by a physically present object within or outside the user's field of view. The virtual object 404 may be cue from or triggered by the physical object 402. For example, the physical object 402 may actually be a stool, and the virtual object 404 may appear to the user (and, in some embodiments, to other users interfacing with the system 100) as a virtual animal standing on the stool. In such an embodiment, the environmental sensing system 306 may use software and / or firmware stored, for example, in the processor 308, to recognize various features and / or shape patterns (captured by the sensor 312) to identify the physical object 402 as a stool. These recognized shape patterns, such as the top of a stool, may be used to trigger the placement of the virtual object 404. Other examples include walls, tables, furniture, cars, buildings, people, floors, plants, animals, and any visible object may be used to trigger an augmented reality experience having some relationship with one or more objects.

[0041] In some embodiments, the particular virtual object 404 to be evoked may be selected by the user or may be automatically selected by other components of the head-mounted system 300 or the interface system 100. Additionally, in embodiments in which a virtual object 404 is automatically evoked, the particular virtual object 404 may be selected based on the particular physical object 402 (or characteristics thereof) from which the virtual object 404 is cued or evoked. For example, if the physical object is identified as a diving board extending over a pool, the evoked virtual object may be a living creature wearing a snorkel, swimsuit, flotation device, or other related item.

[0042] In another exemplary embodiment, the head-mounted system 300 may include a “virtual” mode in which the interface 302 provides a virtual reality interface. In the virtual mode, the physical environment is omitted from the display 303, and virtual object data provided by the computing network 105, the processor 308, and / or the gateway component 140 is presented on the display 303. The omission of the physical environment may be achieved by physically blocking the visual display 303 (e.g., by a cover) or through a feature of the interface 302 that causes the display 303 to transition to an opaque setting. In the virtual mode, live and / or stored visual and audio sensations may be presented to the user through the interface 302, and the user experiences and interacts with the digital world (digital objects, other users, etc.) through the virtual mode of the interface 302. Thus, the interface presented to the user in the virtual mode is composed of virtual object data that comprises the virtual digital world.

[0043] 5 illustrates an example embodiment of a user interface when head-mounted interface 302 is operating in virtual mode. As shown in FIG. 5, the user interface presents a virtual world 500 made up of digital objects 510, which may include atmosphere, weather, terrain, buildings, and people. Although not shown in FIG. 5, the digital objects may also include, for example, plants, vehicles, animals, living things, machines, artificial intelligence, location information, and any other objects or information that define virtual world 500.

[0044] In another exemplary embodiment, the head-mounted system 300 may include a “mixed” mode, in which various features of the head-mounted system 300 (as well as features of the virtual and augmented modes) may be combined to create one or more custom interface modes. In one example custom interface mode, the physical environment is omitted from the display 303, and virtual object data is presented on the display 303 in a manner similar to the virtual mode. However, in this example custom interface mode, the virtual objects may be completely virtual (i.e., they do not exist in the local physical environment), or they may be actual local physical objects that are rendered as virtual objects in the interface 302 in place of physical objects. Thus, in a particular custom mode (referred to herein as a mixed virtual interface mode), live and / or stored visual and audio sensations may be presented to the user through the interface 302, and the user experiences and interacts with a digital world that includes fully virtual objects and rendered physical objects.

[0045] FIG. 6 illustrates an exemplary embodiment of a user interface operating according to a mixed virtual interface mode. As shown in FIG. 6, the user interface presents a virtual world 600 composed of complete virtual objects 610 and rendered physical objects 620 (renderings of objects that otherwise physically exist in the scene). According to the example illustrated in FIG. 6, the rendered physical objects 620 include a building 620A, a ground 620B, and a platform 620C, and are indicated with a thick outline 630 to indicate to the user that the objects are rendered. In addition, the complete virtual objects 610 include an additional user 610A, clouds 610B, a sun 610C, and a flame 610D above the platform 620C. It should be understood that the complete virtual objects 610 may include, for example, atmosphere, weather, terrain, buildings, people, plants, vehicles, animals, creatures, machines, artificial intelligence, location information, and any other objects or information that define the virtual world 600 and that are not rendered from objects present in the local physical environment. Conversely, rendered physical object 620 is an actual local physical object that is rendered as a virtual object in interface 302. Thick outline 630 represents one example for showing the rendered physical object to the user. Thus, the rendered physical object may be shown using methods other than those disclosed herein.

[0046] In some embodiments, the rendered physical objects 620 may be detected using the sensors 312 of the environment sensing system 306 (or using other devices, such as a motion or image capture system) and converted into digital object data, for example, by software and / or firmware stored in the processing circuit 308. Thus, when a user interfaces with the system 100 in a mixed virtual interface mode, various physical objects may be displayed to the user as rendered physical objects. This may be particularly useful for allowing the user to interface with the system 100 while still being able to safely navigate the local physical environment. In some embodiments, the user may be able to selectively remove or add rendered physical objects to the interface display 303.

[0047] In another example custom interface mode, interface display 303 may be substantially transparent, thereby allowing the user to view the local physical environment while various local physical objects are displayed to the user as rendered physical objects. This example custom interface mode is similar to the augmented mode, except that one or more of the virtual objects may be rendered physical objects, as discussed above with respect to the previous example.

[0048] The foregoing example custom interface modes represent some illustrative embodiments of the various custom interface modes that can be provided by the mixed mode of head-mounted system 300. Accordingly, various other custom interface modes may be created from various combinations of the components of head-mounted system 300 and the features and functionality provided by the various modes discussed above without departing from the scope of this disclosure.

[0049] The embodiments discussed herein merely illustrate some examples for providing an interface that operates in off mode, augmented mode, virtual mode, or mixed mode and are not intended to limit the scope or content of each interface mode or the functionality of the components of head-mounted system 300. For example, in some embodiments, virtual objects may include data displayed to a user (e.g., time, temperature, altitude, etc.), objects created and / or selected by system 100, objects created and / or selected by a user, or even objects representing other users interfacing with system 100. Additionally, virtual objects may include extensions of physical objects (e.g., virtual statues growing from a physical platform) and may be visually connected to or detached from physical objects.

[0050] Virtual objects may also be dynamic, changing over time, and according to various relationships (e.g., position, distance, etc.) between the user and other users, physical objects, and other virtual objects, and / or according to other variables defined in the software and / or firmware of head-mounted system 300, gateway component 140, or server 110. For example, in particular embodiments, virtual objects may respond to a user device or its components (e.g., a virtual ball moves when a haptic device is placed next to it), physical or verbal user interactions (e.g., a virtual creature runs away when a user approaches it or speaks when a user speaks to it), a chair being thrown at a virtual creature and the creature dodging the chair, other virtual objects (e.g., a first virtual creature reacts when it sees a second virtual creature), physical variables such as position, distance, temperature, time, or other physical objects in the user's environment (e.g., a virtual creature shown standing in a physical road flattens when a physical car passes by).

[0051] The various modes discussed herein may be applied to user devices other than head-mounted system 300. For example, an augmented reality interface may be provided via a mobile phone or tablet device. In such an embodiment, the phone or tablet may use a camera to capture the user's surrounding physical environment, and virtual objects may be overlaid on the phone / tablet display screen. Additionally, virtual modes may be provided by displaying a digital world on the phone / tablet display screen. Thus, these modes may be mixed to create various custom interface modes as described above using the phone / tablet components discussed herein, as well as other components connected to or used in combination with the user device. For example, mixed virtual interface modes may be provided by a computer monitor, television screen, or other camera-less device operating in combination with a motion or image capture system. In this exemplary embodiment, the virtual world may be viewed from the monitor / screen, and object detection and rendering may be performed by the motion or image capture system.

[0052] 7 illustrates an exemplary embodiment of the present invention in which two users located in different geographic locations each interact with the other user and a common virtual world through their respective user devices. In this embodiment, two users 701 and 702 are tossing a virtual ball 703 (a type of virtual object) back and forth, and each user is able to observe the other user's effect on the virtual world (e.g., each user observes the virtual ball changing direction, being caught by the other user, etc.). Because the movement and position of the virtual object (i.e., virtual ball 703) is tracked by server 110 in computing network 105, system 100, in some embodiments, may communicate to users 701 and 702 the precise location and timing of ball 703's arrival at each user. For example, if a first user 701 is located in London, user 701 may throw ball 703 at a second user 702 located in Los Angeles at a velocity calculated by system 100. Thus, system 100 may communicate the exact time and location of the ball's arrival to second user 702 (e.g., via email, text message, instant message, etc.). In this case, second user 702 may use their device to watch ball 703 arrive at the specified time and location. One or more users may also use geolocation mapping software (or the like) to track one or more virtual objects as they virtually travel around the globe. An example of this might be a user wearing a 3D head-mounted display looking up into the sky and seeing a virtual airplane flying overhead, superimposed on the real world. The virtual airplane may be flown by the user, by an intelligent software agent (software running on the user device or gateway), by other users, which may be locally and / or remotely present, and / or any combination thereof.

[0053] As mentioned above, the user device may include a haptic interface device that provides feedback (e.g., resistance, vibration, light, sound, etc.) to the user when the haptic device is determined by system 100 to be located at a physical spatial position relative to the virtual object. For example, the embodiment described above with respect to FIG. 7 may be extended to include the use of a haptic device 802, as shown in FIG. 8.

[0054] In this exemplary embodiment, haptic device 802 may be displayed in the virtual world as a baseball bat. When ball 703 arrives, user 702 may swing haptic device 802 toward virtual ball 703. If system 100 determines that the virtual bat provided by haptic device 802 has "made contact" with ball 703, haptic device 802 may vibrate or provide other feedback to user 702, and virtual ball 703 may bounce off the virtual bat in a direction calculated by system 100 according to the detected speed, direction, and timing of the contact between the ball and the bat.

[0055] The disclosed system 100, in some embodiments, may facilitate mixed-mode interfacing, where multiple users may interface with a common virtual world (and the virtual objects contained therein) using different interface modes (e.g., augmented, virtual, mixed, etc.). For example, a first user interfacing with a particular virtual world in a virtual interface mode may interact with a second user interfacing with the same virtual world in an augmented reality mode.

[0056] 9A illustrates an example in which a first user 901 (interfacing with the digital world of system 100 in a mixed virtual interface mode) and a first object 902 appear as virtual objects to a second user 922 interfacing with the same digital world of system 100 in a full virtual reality mode. As described above, when interfacing with the digital world via the mixed virtual interface mode, local physical objects (e.g., first user 901 and first object 902) may be scanned and rendered as virtual objects in the virtual world. First user 901 may be scanned, for example, by a motion capture system or similar device and rendered in the virtual world (by software / firmware stored in the motion capture system, gateway component 140, user device 120, system server 110, or other device) as a first rendered physical object 931. Similarly, the first object 902 may be scanned, for example, by the environmental sensing system 306 of the head-mounted interface 300 and rendered in the virtual world (by software / firmware stored on the processor 308, the gateway component 140, the system server 110, or another device) as a second rendered physical object 932. The first user 901 and the first object 902 are shown in a first portion 910 of FIG. 9A as physical objects in the physical world. In a second portion 920 of FIG. 9A , the first user 901 and the first object 902 are shown as first rendered physical object 931 and second rendered physical object 932, as they would appear to a second user 922 interfacing with the same digital world of the system 100 in full virtual reality mode.

[0057] 9B illustrates another example embodiment of a mixed-mode interface in which a first user 901 interfaces with a digital world in a mixed virtual interface mode, as discussed above, and a second user 922 interfaces with the same digital world (and the second user's local physical environment 925) in an augmented reality mode. In the embodiment of FIG. 9B , the first user 901 and a first object 902 are located at a first physical location 915, and the second user 922 is located at a different second physical location 925 separated by some distance from the first location 915. In this embodiment, the virtual objects 931 and 932 may be transposed in real time (or near real time) to a location in the virtual world corresponding to the second location 925. Thus, the second user 922 may observe and interact with a rendered physical object 931 representing the first user 901 and a rendered physical object 932 representing the first object 902 in the second user's local physical environment 925.

[0058] FIG. 10 illustrates an exemplary illustration of a user's field of view when interfacing with system 100 in augmented reality mode. As shown in FIG. 10, the user views the local physical environment (i.e., a city with multiple buildings) and a virtual character 1010 (i.e., a virtual object). The position of the virtual character 1010 may be triggered by 2D visual targets (e.g., signs, postcards, or magazines) and / or one or more 3D reference frames (e.g., buildings, cars, people, animals, airplanes, portions of buildings, and / or 3D physical objects, virtual objects, and / or combinations thereof). In the example illustrated in FIG. 10, known locations of buildings within the city may provide alignment references and / or information and key features for rendering the virtual character 1010. Additionally, the user's geospatial location (e.g., provided by GPS, attitude / position sensors, etc.) or mobile location relative to the buildings may include data used by computing network 105 to trigger the transmission of data used to display the virtual character(s) 1010. In some embodiments, the data used to display the virtual character 1010 may include a rendered character 1010 and / or instructions (executed by the gateway component 140 and / or the user device 120) for rendering the virtual character 1010 or portions thereof. In some embodiments, if the user's geospatial location is unavailable or unknown, the server 110, the gateway component 140, and / or the user device 120 may still display the virtual object 1010 using an estimation algorithm that uses the user's last known location as a function of time and / or other parameters to estimate where certain virtual and / or physical objects may be located. This may also be used to determine the location of any virtual objects if the user's sensors are obstructed and / or experience other malfunctions.

[0059] In some embodiments, the virtual character or virtual object may comprise a virtual figurine, and the rendering of the virtual figurine is triggered by a physical object. For example, referring now to FIG. 11 , a virtual figurine 1110 may be triggered by an actual physical platform 1120. The triggering of the figurine 1110 may be in response to a visual object or feature (e.g., a fiducial, a design feature, a geometric shape, a pattern, a physical location, an elevation, etc.) detected by a user device or other component of the system 100. When a user views the platform 1120 without a user device, the user sees the platform 1120 without the figurine 1110. However, when a user views the platform 1120 through a user device, the user sees the figurine 1110 on the platform 1120, as shown in FIG. 11 . The figurine 1110 is a virtual object and, as such, may be static, animated, change over time or relative to the user's viewing position, or even change depending on which particular user is viewing the figurine 1110. For example, if the user is a small child, the statue may be a dog, whereas if the viewer is an adult male, the statue may be a large robot as shown in FIG. 11. These are examples of user-dependent and / or state-dependent experiences, which allow one or more users to perceive one or more virtual objects, alone and / or in combination with physical objects, and to experience customized and personalized versions of the virtual objects. The statue 1110 (or portions thereof) may be rendered by various components of the system, including, for example, software / firmware installed on the user device. Using data indicating the position and pose of the user device, in combination with alignment features of the virtual object (i.e., statue 1110), the virtual object (i.e., statue 1110) forms a relationship with the physical object (i.e., platform 1120).For example, the relationship between one or more virtual objects and one or more physical objects may be a function of distance, positioning, time, geolocation, proximity to one or more other virtual objects, and / or any other functional relationship involving any type of virtual and / or physical data. In some embodiments, image recognition software in the user device may further enhance the digital-physical object relationship.

[0060] The interactive interface provided by the disclosed systems and methods may be implemented to facilitate a variety of activities, such as, for example, interacting with one or more virtual environments and objects, interacting with other users, and experiencing various forms of media content, including advertisements, musical concerts, and movies. Thus, the disclosed systems facilitate user interaction so that users not only watch or listen to media content, but rather actively participate in and experience the media content. In some embodiments, user participation may include modifying existing content or creating new content to be rendered in one or more virtual worlds. In some embodiments, media content, and / or users who create content, may theme the creation of one or more virtual worlds.

[0061] In one embodiment, a musician (or other user) may create musical content that is rendered for users interacting with a particular virtual world. The musical content may include, for example, various singles, EPs, albums, videos, short films, and concert performances. In one embodiment, multiple users may interface with system 100 to simultaneously experience a virtual concert performed by a musician.

[0062] In some embodiments, produced media may include a unique identifier code associated with a particular entity (e.g., a band, artist, user, etc.). The code may be a set of alphanumeric characters, a UPC code, a QR code, a 2D image trigger, a 3D physical object feature trigger, or other form of digital mark, as well as sound, image, and / or both. In some embodiments, the code may also be embedded in digital media that may be interfaced using system 100. A user may obtain the code (e.g., by paying a fee) and redeem the code to access media content produced by the entity associated with the identifier code. Media content may be added to or removed from the user's interface.

[0063] In one embodiment, to avoid the computational and bandwidth limitations of passing real-time or near-real-time video data from one computing system to another (e.g., from a cloud computing system to a local processor coupled to a user) with low latency, parametric information about various shapes and geometries is transferred and may be utilized to define surfaces, while textures are transferred and added to these surfaces to provide static or dynamic details, such as bitmap-based video details of an individual's face mapped onto the parametrically reconstructed facial geometry. As another example, if a system is configured to recognize an individual's face and knows that the individual's avatar is located in the augmented world, the system may be configured to pass relevant world information and individual's avatar information in one relatively large setup transfer, after which the remaining transfers to a local computing system, such as 308 depicted in FIG. 1 , for local rendering may be limited to parameter and texture updates, such as motion parameters of the individual's skeletal structure and movement bitmaps of the individual's face, at a much smaller bandwidth than the initial setup transfer or real-time video passing. Thus, cloud-based and local computing assets may be used in an integrated fashion, with the cloud handling computations that do not require relatively low latency and the local processing assets handling low latency critical tasks, and in such cases, the type of data transferred to the local system is preferably passed at a relatively low bandwidth, with some amount of such data type (i.e., parameter information, textures, etc. as opposed to all real-time video).

[0064] Referring first to Figure 15, a schematic diagram illustrates the coordination between cloud computing assets (46) and local processing assets (308, 120). In one embodiment, the cloud (46) assets are operatively coupled directly (40, 42) to one or both of the local computing assets (120, 308), such as processor and memory configurations that may be housed in structures configured to be coupled to a user's head (120) or belt (308), via wired or wireless networking (wireless being preferred for mobility, wired being preferred for certain high-bandwidth or large data capacity transfers that may be desired). These computing assets that are local to the user may likewise be operatively coupled to each other via wired and / or wireless connection configurations (44). In one embodiment, to maintain low inertia and a small head-mounted subsystem (120), the primary transfer between the user and the cloud (46) may be via a link between the belt-based subsystem (308) and the cloud, with the head-mounted subsystem (120) being primarily data-tethered to the belt-based subsystem (308) using a wireless connection, such as, for example, an ultra-wideband ("UWB") connection as currently employed in personal computing peripheral connectivity applications.

[0065] Using efficient local and remote processing coordination and an appropriate display device for the user (e.g., the user interface 302 or user “display device” featured in FIG. 3 , the display 14 described below with reference to FIG. 14 , or variations thereof), aspects of a world relevant to the user's current real or virtual location may be transferred or “handed” to the user and efficiently updated. Indeed, in one embodiment, when one individual utilizes a virtual reality system (“VRS”) in an augmented reality mode and another individual utilizes the VRS in a fully virtual mode to explore the same world local to the first individual, the two users may mutually experience the world in various ways. For example, with reference to FIG. 12 , a scenario similar to that described with reference to FIG. 11 is depicted, with the addition of a visualization of the second user's avatar 2 flying through the augmented reality world depicted from the fully virtual reality scenario. 12 may be experienced and displayed in augmented reality to a first individual, with two augmented reality elements (statue 1110 and the flying bumblebee avatar 2 of the second individual) displayed in addition to the actual physical elements of the local world surroundings within the scene, such as the ground, background buildings, statue platform 1120, etc. Dynamic updates may be utilized to allow the first individual to visualize the progress of the second individual's avatar 2 as it flies through the world local to the first individual.

[0066] Again, with a configuration such as that described above, where there is one world model that resides on and can be distributed from cloud computing resources, such a world can be “handed off” to one or more users in a relatively low-bandwidth format that is preferable for distributing real-time video data or the like. An individual standing near a statue (i.e., as shown in FIG. 12 ) may have their augmented experience informed by the cloud-based world model, a subset of which may be handed off to that individual and their local display device to complete the view. An individual sitting at a remote display device, which may be as simple as a personal computer located on a desk, can efficiently download the same section of information from the cloud and have it rendered on their display. In fact, an individual physically present in the park near the statue may bring a remotely located friend for a walk in the park, with the friend participating through virtual and augmented reality. The system needs to know where the paths are, where the trees are, and where the statues are, but with that information on the cloud, participating friends can download aspects of the scenario from the cloud and then begin walking together as augmented reality that is local to the individual physically in the park.

[0067] Referring to Figure 13, an embodiment based on time and / or contingency parameters is depicted in which an individual interacting with a virtual reality and / or augmented reality interface, such as the user interface 302 or user display device featured in Figure 3, the display device 14 described below with reference to Figure 14, or variations thereof, is using the system (4) and enters a coffee shop to order a cup of coffee (6). The VRS may be configured to utilize sensing and data collection capabilities locally and / or remotely to provide augmented reality and / or virtual reality display enhancements for the individual, such as a highlighted location of the coffee shop door or a bubble window of a related coffee menu (8). When the individual receives the ordered cup of coffee, or upon detection of some other relevant parameter by the system, the system may be configured to display, using the display device (10), one or more time-based augmented reality or virtual reality images, videos, and / or sounds in the local environment (e.g., views of the Madagascar jungle from the walls and ceiling, either static or dynamic, with or without jungle sounds and other effects). Such presentation to the user may be interrupted based on timing parameters (i.e., 5 minutes after a full coffee cup is recognized and handed to the user, 10 minutes after the system recognizes the user walking through the store's front door, etc.) or other parameters (e.g., the system's recognition that the user has finished drinking their coffee by noting the upside-down orientation of the coffee cup as the user takes the last sip of coffee from the cup, or the system's recognition that the user has walked out the store's front door) (12).

[0068] Referring to FIG. 14 , one embodiment of a suitable user display device 14 is shown, comprising a display lens 82 that can be attached to a user's head or eye by a housing or frame 84. The display lens 82 comprises one or more transparent mirrors positioned by the housing 84 in front of the user's eye 20, which may be configured to reflect projected light 38 into the eye 20 to facilitate beam shaping, while also allowing transmission of at least some light from the local environment in an augmented reality configuration. (In a virtual reality configuration, it may be desirable for the display system 14 to be capable of blocking substantially all light from the local environment, such as by a darkened visor, blocking curtains, a fully black LCD panel mode, or the like.) In the depicted embodiment, two wide-field machine vision cameras 16 are coupled to the housing 84 to image the user's surrounding environment. In one embodiment, the cameras 16 are dual-capture visible / infrared cameras. The depicted embodiment also includes a pair of scanning laser wavefront shaping (i.e., for depth) projector modules, along with viewing mirrors and optics configured to project light (38) into the eye (20), as shown. The depicted embodiment also includes two miniature infrared cameras (24) paired with infrared light sources (26, light emitting diodes "LEDs," etc.) configured to track the user's eye (20) to assist in rendering and user input. The system (14) further features a sensor assembly (39) with X-, Y-, and Z-axis accelerometer capabilities, a magnetic compass, and X-, Y-, and Z-axis gyro capabilities, and may preferably provide data at a relatively high frequency, such as 200 Hz. The depicted system (14) also includes a head pose processor (36), such as an ASIC (application specific integrated circuit), FPGA (field programmable gate array), and / or ARM processor (advanced reduced instruction set machine), which may be configured to calculate real-time or near-real-time user head pose from the wide field of view image information output from the capture device (16).Also shown is another processor (32) configured to perform digital and / or analog processing to derive attitude from gyro, compass, and / or accelerometer data from the sensor assembly (39). The depicted embodiment also features a GPS (37, Global Positioning Satellite) subsystem to assist with attitude and positioning. Finally, the depicted embodiment includes a rendering engine (34), which may feature hardware that runs a software program configured to provide rendering information local to the user to facilitate operation of the scanner and imaging into the user's eye for the user's view of the world. The rendering engine (34) is operably coupled (81, 70, 76 / 78, 80, i.e., via wired or wireless connections) to the sensor attitude processor (32), image attitude processor (36), eye-tracking camera (24), and projection subsystem (18) so that light of the rendered augmented reality and / or virtual reality object is projected using the scanning laser array (18) similar to a retinal scanning display. The wavefront of the projected light beam (38) may be bent or focused to match the desired focal length of the augmented reality and / or virtual reality object. A mini infrared camera (24) may be used to track the eyes to assist with rendering and user input (i.e., where the user is looking, at what depth the user is focusing (as discussed below, the edge of the eye may be used to estimate focal depth)). A GPS (37), gyro, compass, and accelerometer (39) may be used to provide heading estimation and / or fast pose estimation. The camera's (16) images and pose, along with data from associated cloud computing resources, may be used to map the local world and share the user's view with the virtual reality or augmented reality community.While most of the hardware in the display system 14 featured in FIG. 14 is depicted as directly coupled to the display 82 and the housing 84 adjacent to the user's eyes 20, the depicted hardware components may be mounted on or housed within other components, such as, for example, belt-mounted components as shown in FIG. 3. In one embodiment, all of the components of the system 14 featured in FIG. 14 are directly coupled to the display housing 84, except for the image pose processor 36, the sensor pose processor 32, and the rendering engine 34, and communication between the image pose processor 36, the sensor pose processor 32, and the rendering engine 34 and the remaining components of the system 14 may be via wireless communication, such as ultra-wideband, or wired communication. The depicted housing 84 is preferably head-mounted and wearable by a user. It may also feature speakers, such as those inserted into the user's ears, that may be utilized to provide the user with sounds that may be associated with the augmented reality or virtual reality experience, such as the jungle sounds referenced in reference to FIG. 13, and a microphone that may be utilized to capture sounds that are local to the user.

[0069] With regard to projecting light 38 into the user's eye 20, in one embodiment, a miniature camera 24 may be utilized to measure the location where the center of the user's eye 20 is geometrically tangent, which generally coincides with the location of the eye's 20 focal point, or "depth of focus." The three-dimensional surface of all points of contact with the eye is called the "horopter." The focal distance may exhibit a finite number of depths or may vary infinitely. Light projected from the convergence distance appears to be focused on the subject's eye 20, while light in front of or behind the convergence distance is blurred. Furthermore, it has been discovered that spatially coherent light with a beam diameter of less than approximately 0.7 millimeters is properly resolved by the human eye, regardless of where the eye focuses. With this understanding in mind, to create the illusion of proper depth of focus, eye convergence may be tracked using a mini-camera (24), and the rendering engine (34) and projection subsystem (18) may be used to render all objects on or near the horopter in focus, and all other objects to various degrees of defocus (i.e., using intentionally created blur). Perspective light-guiding optics configured to project coherent light into the eye may be provided by suppliers such as Lumus, Inc. Preferably, the system (14) renders to the user at a frame rate of about 60 frames per second or greater. As described above, the mini-camera (24) may preferably be utilized for eye tracking, and software may be configured to capture not only convergence geometry but also focus position cues that serve as user input. Preferably, such a system is configured with brightness and contrast suitable for daytime or nighttime use. In one embodiment, such a system preferably has a latency of less than about 20 milliseconds for visual object alignment, an angular alignment of less than about 0.1 degrees, and a resolution of about 1 arc minute, which is approximately the limit of the human eye.The display system (14) may be integrated with a localization system, which may include a GPS element, optical tracking, a compass, an accelerometer, and / or other data sources to assist in position and attitude determination, and the localization information may be utilized to facilitate accurate rendering within the user's field of view of the relevant world (i.e., such information facilitates the glasses' understanding of where they are relative to the real world).

[0070] Other suitable display devices include desktop and mobile computers, smartphones that can be augmented with additional software and hardware features that facilitate or simulate 3D perspective viewing (e.g., in one embodiment, a frame may be removably coupled to the smartphone, the frame featuring a 200 Hz gyro and accelerometer sensor subset, two small machine vision cameras with wide field of view lenses, and an ARM processor to simulate some of the functionality of the configuration featured in FIG. 14 ), tablet computers, tablet computers that can be augmented as described above for smartphones. These include, but are not limited to, computers, tablet computers augmented with additional processing and sensing hardware, head-mounted systems that use smartphones and / or tablets to display augmented and virtual viewpoints (visual adaptation via magnifying optics, mirrors, contact lenses, or light structuring elements), non-see-through displays of light-emitting elements (LCD, OLED, vertical cavity surface-emitting laser, guided laser beams, etc.), see-through displays that allow a person to simultaneously see the natural world and artificially generated images (e.g., light-directing optics, clear and polarized OLEDs shining into near-focus contact lenses, guided laser beams, etc.), contact lenses with light-emitting elements (such as those available from Innovega, Inc. (Bellevue, WA) under the trade name Ioptik R™, which may be combined with specialized complementary eyeglass components), implantable devices with light-emitting elements, and implantable devices that simulate the photoreceptors of the human brain.

[0071] Using systems such as those depicted in FIGS. 3 and 14, 3D points can be captured from the environment, and the pose (i.e., vector and / or origin position information relative to the world) of the camera capturing these images or points can be determined, so that these points or images can be "tagged" or associated with this pose information. Points captured by a second camera can then be used to determine the pose of the second camera. In other words, the second camera can be oriented and / or localized based on a comparison with the tagged image from the first camera. This knowledge can then be used to extract textures (since there are two aligned cameras around), create maps, and create virtual copies of the real world. Thus, at a basic level, one embodiment has a personally worn system that can be used to capture both the 3D points and the 2D images that generated the points, which can then be sent to cloud storage and processing resources. They may also be cached locally (i.e., caching tagged images) with embedded pose information, so the cloud may have tagged (i.e., tagged with 3D pose) 2D images ready (i.e., available in cache) along with the 3D points. If the user is observing something dynamic, the user may send additional information up to the cloud related to the movement (e.g., when looking at another individual's face, the user can take the texture map of the face and boost it to an optimized frequency, even if the surrounding world is otherwise essentially static).

[0072] The cloud system may be configured to store some points as pose-only fiducials to reduce overall pose tracking computations. Generally, it may be desirable to have some contour features so that key items in the user's environment, such as walls, tables, etc., can be tracked as the user moves around a room; the user may want to "share" the world and let other users into the room and see these points. Such useful and key points may be referred to as "fiducials" because they are extremely useful as anchoring points. They pertain to features that can be recognized by machine vision and that can be consistently and repeatedly extracted from the world on different pieces of user hardware. Therefore, these fiducials may preferably be stored in the cloud for further use.

[0073] In one embodiment, it is preferable to have a relatively even distribution of fiducials throughout the relevant universe, since fiducials are the kinds of items that a camera can easily use to recognize locations.

[0074] In one embodiment, the associated cloud computing arrangement may be configured to periodically refine the database of 3D points and any associated metadata to use the best data from various users for both criteria refinement and world creation. In other words, the system may be configured to obtain the best data set by using input from various users who view and work within the associated world. In one embodiment, the database is fractal in nature, and as users get closer to an object, the cloud passes on higher resolution information to those users. As users map an object more closely, that data is sent to the cloud, which can add new 3D points and image-based texture maps to the database if they are better than those previously stored in the database. All of this may be configured to happen simultaneously from many users.

[0075] As described above, an augmented reality or virtual reality experience may be based on recognizing certain types of objects. For example, to recognize and understand certain objects, it may be important to understand that such objects have depth. A recognizer software object (“recognizer”) may be deployed on a cloud or local resource to specifically assist in recognizing various objects on either or both platforms as a user navigates the data within the world. For example, if a system has world model data comprising a 3D point cloud and pose-tagged images of a desk with many points on it, as well as an image of the desk, there may be no determination that what is being observed is actually a desk when a human is grasping the desk. In other words, a few 3D points in space and an image from somewhere in the space showing most of the desk may not be sufficient to instantly recognize that a desk is being observed. To assist in this identification, a specific object recognizer may be created that goes into the raw 3D point cloud, segments a set of points, and extracts, for example, the plane of the desk's top surface. Similarly, recognizers may be created to segment walls from 3D points so that a user can change wallpaper in virtual or augmented reality or remove part of a wall and have an entrance to another room that is not actually there in the real world. Such recognizers operate within the data of a world model, crawling the world model; such recognizers may be thought of as software “robots” that instill semantic information, or ontologies, into the world model regarding what is believed to exist between points in space. Such recognizers or software robots may be configured such that their entire existence is about crawling through the relevant world data and finding what is believed to be a wall, or a chair, or other item. They may be configured to tag a set of points with the functional equivalence “this set of points belongs to a wall,” or may comprise a combination of point-based algorithms and pose-tagging image analysis to manually inform the system regarding what is among the points.

[0076] Object recognizers may be created for many purposes of varying utility, depending on the perspective. For example, in one embodiment, a specialty coffee shop such as Starbucks may invest in creating an accurate recognizer of Starbucks coffee cups within a relevant universe of data. Such a recognizer would be configured to crawl universes of data, large and small, searching for Starbucks coffee cups, so that Starbucks coffee cups may be segmented and identified to a user when operating within the relevant neighborhood space (i.e., perhaps providing the user with coffee at a nearby Starbucks outlet when the user sees the Starbucks coffee cup for a particular period of time). Once the cup is segmented, it may be quickly recognized when the user moves it onto their desk. Such a recognizer may be configured to operate or operate on cloud computing resources and data, as well as local resources and data, or both cloud and local, depending on available computational resources. In one embodiment, there is a global copy of the world model on the cloud with millions of users contributing to the global model, but for smaller worlds or sub-worlds, such as a particular person's office in a particular town, the system may be configured to arrange the data and move it to a local cache of information that is deemed most locally relevant to a given user, since most of the global world does not care what that office looks like. In one embodiment, for example, an object identified as moving often (e.g., a cup on a desk) as a user walks up to a desk may be configured so that relevant information (such as a segment of a particular cup on a desk) resides only on local computing resources, not on the cloud, so that it does not need to burden the cloud model and incur transmission burdens between the cloud and local resources.Thus, cloud computing resources may be configured to segment 3D points and images, thus decomposing permanent (i.e., generally non-moving) objects from movable ones, which affects where relevant data remains and where it is processed, removing the processing burden from wearable / local systems for specific data related to more permanent objects, allowing one-time processing of locations that can then be shared with an unlimited number of other users, allowing multiple data sources to simultaneously build a database of fixed and movable objects at a particular physical location, and segmenting objects from their background to create object-specific fiducials and texture maps.

[0077] In one embodiment, the system may be configured to query the user for input regarding the identity of a particular object (e.g., the system may pose a question to the user such as, "Is that a Starbucks coffee cup?"), so that the user can tailor the system and enable the system to associate semantic information with objects in the real world. The ontology may provide guidance regarding what objects segmented from the world can do, how they behave, etc. In one embodiment, the system may feature a virtual or actual keypad, such as a wirelessly connected keypad, connectivity to a smartphone keypad, or the like, to facilitate specific user input into the system.

[0078] The system may be configured to share primitives (walls, windows, desk geometry, etc.) with any user who enters the room in virtual or augmented reality; in one embodiment, that individual's system is configured to take images from specific viewpoints and upload them to the cloud. The cloud can then combine the old and new sets of data to run optimization routines and establish criteria that exist on individual objects.

[0079] GPS and other localization information may be used as inputs to such processing. Additionally, other computing systems and data, such as a person's online calendar or Facebook account information, may be used as inputs (e.g., in one embodiment, the cloud and / or local system may be configured to analyze the contents of a user's calendar for flights, dates, and destinations so that information can be moved from the cloud to the user's local system over time to prepare for the user's arrival time at a given destination).

[0080] In one embodiment, tags such as QR codes and the like may be inserted into the world for use with non-statistical pose calculations, security / access control, conveying special information, spatial messaging, non-statistical object recognition, etc.

[0081] In one embodiment, cloud resources may be configured to pass digital models of the real and virtual worlds between users, as described above with reference to "passable worlds," with the models rendered by individual users based on parameters and textures. This reduces bandwidth compared to passing real-time video, enables the rendering of virtual perspectives of the scene, and allows millions or more users to participate in a single virtual gathering without transmitting the data (e.g., video) that each user needs to see, because their view is rendered by local computing resources.

[0082] A virtual reality system (“VRS”) may be configured to register a user's position and field of view (together known as “pose”) through one or more of real-time metric computer vision using cameras, simultaneous localization and mapping techniques, maps, and data from sensors (e.g., gyros, accelerometers, compasses, barometers, GPS), radio signal strength triangulation, signal time-of-flight analysis, LIDAR ranging, RADAR ranging, odometry, and sonar ranging. A wearable device system may be configured to simultaneously map and orient. For example, in an unknown environment, a VRS may be configured to gather information about the environment and identify images to provide suitable reference points for user pose calculation, other points for world modeling, and a texture map of the world. The reference points may be used to optically calculate pose. As the world is mapped in greater detail, more objects may be segmented and given their own texture maps, although the world can still preferably be represented at a low spatial resolution with simple polygons using low-resolution texture maps. Other sensors, such as those discussed above, may be utilized to assist in this modeling effort. The world may be fractal in nature, in that moving (through viewpoint, "surveillance" mode, zooming, etc.) or otherwise seeking a better view requires higher resolution information from cloud resources. Getting closer to an object captures higher resolution data, which can be transmitted to the cloud, which can compute new data and / or insert new data into gaps in the world model.

[0083] Referring to FIG. 16 , the wearable system may be configured to capture image information and extract fiducials and recognized points (52). The wearable local system may calculate the pose using one of the pose calculation techniques described below. The cloud (54) may be configured to use the images and fiducials to segment the 3D object from a more static 3D background, with the images providing a texture map of the object and the world (the texture may be real-time video). The cloud resource (56) may be configured to store and make available the static fiducials and texture for world registration. The cloud resource may be configured to trim the point cloud for optimal point density for registration. The cloud resource (60) may be configured to store and make available object fiducials and texture for object registration and manipulation, with the cloud trimming the point cloud for optimal density for registration. The cloud resource may be configured to use all valid points and textures to generate a fractal solid model of the object (62), with the cloud trimming the point cloud information for optimal fiducial density. The cloud resources (64) may be configured to query users for tailoring regarding the identification of segmented objects and worlds, and the ontology database may use the answers to instill actionable properties into the objects and worlds.

[0084] The following specific alignment and mapping modes feature "O attitude," which represents attitude determined from an optical or camera system; "s attitude," which represents attitude determined from sensors (i.e., a combination of data from GPS, gyro, compass, accelerometer, etc., as discussed above); and "MLC," which represents cloud computing and data management resources.

[0085] 1. Orientation: Creating a base map of the new environment Purpose: To establish attitude (or similar) when the environment is not mapped or connected to an MLC. Extract points from images, track them from frame to frame, and triangulate fiducials using the S-pose. · Use S position as there is no standard. · Weed out poor criteria based on persistence. This is the most basic mode. It always operates on a low-precision attitude. Over a small amount of time and some relative motion, this establishes an attitude and / or a minimal reference set for mapping. Exit this mode as soon as you are comfortable with the position.

[0086] 2. Map and O-pose: Map the environment Goal: Establish high-precision attitude, map the environment, and provide the map (with images) to the MLC. ·Calculate O pose from the maturity world reference. Use S pose as a check on the O pose solution and to accelerate the calculation (O pose is a nonlinear gradient search). ·The maturity reference may be derived from MLC or locally determined. ·Extract points from the image, track them from frame to frame, and triangulate the reference using the O pose. · Weed out poor criteria based on persistence. · Provides reference and attitude tagged images to MLC. The last three steps do not need to happen in real time.

[0087] 3. Posture: Determine your posture Goal: Establish high-accuracy pose within an already mapped environment using minimal processing power. Use the previous S and O poses (n-1, n-2, n-3, etc.) to estimate the pose at n. Use the pose at n to project the fiducial onto the image captured at n, then create an image mask from the projection. Extract points from masked regions (by searching / extracting points only from a masked subset of the image, the processing load is greatly reduced). Calculate O-pose from extracted points and mature global reference. · Use the S and O poses at n to estimate the pose at n+1. Optional: Provide posture-tagged images / videos to the MLC Cloud.

[0088] 4. Super-resolution: Determine the super-resolution image and the standard Objective: To create super-resolution images and standards. · Synthesize pose-tagged images to create super-resolution images. Use super-resolution images to enhance reference position estimation. · Iterative pose estimation from super-resolution reference and images. Optional: Loop the above steps on a wearable device (in real time) or on an MLC (for a better world).

[0089] In one embodiment, a VLS system may be configured with certain base functionality and functionality facilitated by "apps" or applications that may be delivered through the VLS to provide certain specialized functionality. For example, the following apps may be installed on a target VLS to provide specialized functionality:

[0090] A painterly rendering app. Artists create image transformations that represent the world as they see it. Users activate these transformations, thus seeing the world "through" the artist's eyes.

[0091] Tabletop modeling apps, where users "build" objects from physical objects placed on a table.

[0092] A virtual presence app where users hand over a virtual model of a space to other users who then move around the space using a virtual avatar.

[0093] Avatar emotion app. Subtle voice inflections, slight head movements, body temperature, heart rate, and other measurements animate subtle effects on virtual presence avatars. Digitizing human state information and passing it to a remote avatar uses less bandwidth than video. In addition, such data can be mapped to a non-human avatar with emotions. For example, a dog avatar can show excitement by wagging its tail based on excited voice inflections.

[0094] An efficient mesh network may be desirable for moving data as opposed to sending everything back to a server. However, many mesh networks have suboptimal performance because location information and topology are not well characterized. In one embodiment, the system may be utilized to determine the location of all users with relatively high accuracy, and thus a mesh network configuration may be utilized for high performance.

[0095] In one embodiment, the system may be utilized for search. With augmented reality, for example, users generate and leave content related to many aspects of the physical world. Much of this content is not text and therefore not easily searchable by typical methods. The system may be configured to provide a facility for keeping track of personal and social network content for search and reference purposes.

[0096] In one embodiment, if a display device tracks 2D points through successive frames and then fits a vector-valued function to the time evolution of these points, it is possible to sample the vector-valued function at any point in time (e.g., between frames) or at some point in the near future (by projecting the vector-valued function forward). This allows for the creation of high-resolution post-processing and the prediction of future poses before the next image is actually captured (e.g., it is possible to double the alignment speed without doubling the camera frame rate).

[0097] For body-fixed rendering (as opposed to head-fixed or world-fixed rendering), an accurate representation of the body is desirable. In one embodiment, rather than measuring the body, its location can be derived through the average position of the user's head. If the user's face faces forward most of the time, a multi-day average of head position will reveal that orientation. In conjunction with the gravity vector, this provides a reasonably stable coordinate system for body-fixed rendering. Using a current measure of head position relative to this long-term coordinate system allows for consistent rendering of objects on / around the user's body without extra instrumentation. For implementation of this embodiment, a single registration average of the head direction vector may be started, and the cumulative sum of the data divided by delta t gives the current average head position. Keeping approximately five registrations, starting on day n-5, on days n-4, n-3, n-2, and n-1, allows for the use of a rolling average of only the past "n" days.

[0098] In one embodiment, a scene may be scaled down and presented to a user in a smaller space than it actually is. For example, in situations where a scene must be rendered in a huge space (i.e., a soccer stadium, etc.), an equivalent huge space may not exist, or such a large space may be inconvenient for the user. In one embodiment, the system may be configured to reduce the scale of the scene so that a user may observe a scaled down view. For example, an individual may play a video game from a God's eye view, or a world championship soccer match, on an unscaled stadium, or on a scaled stadium presented on a living room floor. The system may be configured to simply shift the viewpoint, scale, and associated adaptive distance.

[0099] The system may also be configured to draw the user's attention to particular items within the presented scene by manipulating the focus of virtual or augmented reality objects, highlighting them, changing their contrast, brightness, scale, etc.

[0100] Preferably, the system may be configured to achieve the following modes:

[0101] Open Space Rendering: Capture key points from a structured environment and then use ML rendering to fill in the spaces in between. · Potential venues: stages, output spaces, large indoor spaces (stadiums).

[0102] Object Wrapping: Recognize 3D objects in the real world and then augment them. "Recognition" here means identifying 3D blobs with sufficient accuracy to connect the images. There are two kinds of recognition: 1) classifying a type of object (e.g., "face"), and 2) classifying a specific instance of an object (e.g., Joe, an individual). Build recognizer software objects for a variety of objects, including walls, ceilings, floors, faces, roads, skies, skyscrapers, ranch houses, tables, chairs, cars, road signs, billboards, doors, windows, bookshelves, etc. Some recognizers are I-type and have general functionality, such as "put my video on that wall" or "that's a dog." Other recognizers are Type II and have specific functionality, such as "My TV is on the living room wall 3.2 feet from the ceiling" or "That's Fido" (which are more capable versions of general recognizers). Building perceptors as software objects allows for quantified release of functionality and finer-grained control of the experience.

[0103] Body-centric rendering Rendering virtual objects fixed to the user's body. Some things, such as a digital tool belt, should float around the user's body. This requires knowing where the body is, not just the head. By taking a long-term average of the user's head position (which is usually facing forward and parallel to the ground), you can get a reasonably accurate desired position. A simple example is objects floating around the head.

[0104] Transparency / Broken Views For type II recognition objects, a cutaway view is shown. Linking type II recognized objects to an online database of 3D models. You should start with objects that have commonly available 3D models, such as cars and utilities.

[0105] Virtual Being Draw avatars of remote people in open space. ○ A subset of "Free Space Rendering" (above). o A user creates the rough geometry of a local environment and iteratively sends both the geometry and texture maps to others. Users must grant permission for others to enter their environment. Subtle vocal cues, hand tracking, and head movements are transmitted to a remote avatar, which is animated from these fuzzy inputs. ○The above minimizes bandwidth. Create a "doorway" in the wall to another room Pass in geometry and texture maps as you would with any other method. Instead of showing avatars in a local room, designate recognized objects (e.g., walls) as portals to other people's environments. This way, multiple people can sit in their own rooms and see other people's environments "through" the walls.

[0106] Virtual Viewpoint As a group of cameras (people) view the scene from different perspectives, a dense digital model of the area is created. This rich digital model can be rendered from any vantage point that at least one camera can see. Example: People at a wedding. The scene is modeled jointly by all attendees. The recognizer distinguishes static objects from moving objects and creates texture maps (e.g., walls have a stable texture map, people have a higher frequency moving texture map). With a rich digital model updated in real time, the view can be rendered from any vantage point. Attendees in the back can even fly up to the front rows for a better view. Attendees can show their moving avatar or hide their viewpoint. Off-site attendees can find their "seat" using their avatar, or invisibly if the organizer allows. Likely to require extremely high bandwidth. Conceptually, high frequency data is streamed to the crowd over high speed local radio. Low frequency data comes from MLC. · Since all attendees have high-precision location information, creating optimal routing paths for local networking is trivial.

[0107] Messaging Simple silent messaging may be desirable. For this and other uses, it may be desirable to have a finger-chording keyboard. Tactile grab solutions may provide improved performance.

[0108] Full Virtual Reality (VR): When the Vision System goes dark, it shows a non-overlapping view of the real world. A registration system is still required to track head position. "Couch Mode" allows users to fly. "Walk Mode" re-renders real-world objects as virtual ones to prevent users from colliding with the real world. Rendering body parts is essential to believing in fiction. This suggests having a method for tracking and rendering body parts in the field of view. Non-see-through visors are a form of VR that offer many image quality enhancement benefits not possible with direct overlay. · Wide field of vision, perhaps even the ability to see behind you. Various forms of "super vision": telescopic, clairvoyant, infrared, God's eye view, etc.

[0109] In one embodiment, the system for virtual and / or augmented user experiences is configured so that a remote avatar associated with a user may be animated based at least in part on data on the wearable device with input from sources such as voice intonation analysis and facial recognition analysis as performed by an associated software module. For example, referring again to FIG. 12 , a bee avatar (2) may be animated to smile in a friendly manner based on facial recognition of a smile on the user's face or based on a friendly tone of voice or tone as determined by software configured to analyze audio input into a microphone that may capture audio samples locally from the user. Additionally, the avatar character may be animated in a manner that would cause the avatar to express a particular emotion. For example, in an embodiment in which the avatar is a dog, a happy smile or tone detected by a system local to the human user may be represented by the avatar as the dog avatar wagging its tail.

[0110] Various exemplary embodiments of the present invention are described herein. Reference is made to these examples in a non-limiting sense. They are provided to illustrate the broader and more applicable aspects of the present invention. Various changes may be made to the invention described, and like substitutes may be made without departing from the true spirit and scope of the invention. In addition, many modifications may be made to adapt a particular situation, material, composition of matter, process, process act, or step to the objective, spirit, or scope of the present invention. Moreover, as will be understood by those skilled in the art, each of the individual variations described and illustrated herein has distinct elements and characteristics that may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. All such modifications are intended to be within the scope of the claims associated with this disclosure.

[0111] The present invention includes methods that may be performed using the subject devices. The methods may include the act of providing such suitable devices. Such providing may be performed by an end user. In other words, the act of "providing" merely requires the end user to obtain, access, access, position, configure, activate, power on, or otherwise act upon the requisite devices in the subject methods. Methods described herein may be carried out in any order of the recited events that is logically possible, not just the recited order of events.

[0112] Exemplary aspects of the invention, along with details regarding material selection and manufacturing, have been described above. As for other details of the invention, these may be understood in connection with the above-referenced patents and publications and are generally known or may be understood by those skilled in the art. The same may be true with respect to method-based aspects of the invention in terms of additional acts as commonly or logically adopted.

[0113] Additionally, while the present invention has been described with reference to several embodiments incorporating various features as appropriate, the present invention is not limited to that described and indicated, as envisioned with respect to each variation of the present invention. Various modifications may be made to the invention as described, and equivalents (whether described herein or not included for purposes of simplicity) may be substituted without departing from the true spirit and scope of the invention. Additionally, when a range of values is provided, it is to be understood that all intervening values between the upper and lower limits of that range, and any other stated or intervening value within that stated range, are encompassed within the invention.

[0114] It is also contemplated that any optional features of the described inventive variations may be described and claimed independently or in combination with any one or more of the features described herein. Reference to a singular item includes the possibility that there are plurals of the same items. More specifically, as used herein and in the claims associated herewith, the singular forms "a," "an," "said," and "the" include plural referents unless otherwise specified. In other words, the use of articles in the above description and in the claims associated with this disclosure allows for "at least one" of the subject item. Furthermore, it should be noted that such claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a predicate for the use of exclusive terminology such as "solely," "only," and the like, or the use of a "negative" limitation, in connection with the recitation of claim elements.

[0115] Without using such exclusive language, the term "comprising" in the claims associated with this disclosure shall be construed as allowing for the inclusion of any additional elements, regardless of whether a given number of elements are recited in such claims or whether the addition of features can be considered as a transformation of the nature of the elements recited in such claims. Except as specifically defined herein, all technical and scientific terms used herein shall be given the broadest possible commonly understood meaning while maintaining the validity of the claims.

[0116] The breadth of the present invention is not limited to the examples and / or subject specification provided, but rather is limited only by the scope of the claims language associated with this disclosure.

Claims

1. A system, comprising: A first user device configured to communicate with a computer network comprising one or more computing devices, the one or more computing devices comprising: one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to process first virtual world data; a first user device comprising: The computer network comprises: receiving a first input from a first user via a user sensing system; receiving a second input from a local environment of the first user device via an environmental sensing system; wherein one or more of the one or more computing devices are configured to generate second virtual world data based on the first virtual world data and further based on at least one of the first input and the second input; the first user device is further configured to present virtual content to the first user based on the second virtual world data; presenting the virtual content to the first user includes presenting a visual rendering of the virtual content in a 3D format; presenting the visual rendering of the virtual content in the 3D form includes presenting the visual rendering on a display based on a position and orientation of the first user device; the second virtual world data includes a virtual object; the one or more computing devices: predicting a time and location of an event associated with the virtual object for a second user; communicating a written message to a second user device associated with the second user, the written message including the predicted time and location of the event for the second user; presenting the event to the second user at the predicted time and location after communicating the written message to the second user device; 1. A system comprising: a server configured to:

2. The system of claim 1, wherein the written message comprises an email.

3. The system of claim 1, wherein the written message includes a text message.

4. The system of claim 1, wherein the written message includes an instant message.

5. The system of claim 1, wherein the virtual object includes a vehicle and the location of the event includes the location of the vehicle determined using geolocation mapping software.

6. The system described in claim 5, wherein the vehicle is configured to be operated by the first user.

7. The system described in claim 5, wherein the vehicle is configured to be operated by the second user.

8. The system described in claim 5, wherein the vehicle is configured to be operated by an intelligent software agent configured to execute via the computer network in conjunction with one or more of the first user and the second user.

9. The system described in claim 5, wherein the vehicle is configured to be operated by multiple users remote to the second user.

10. The server predicting a second time and location of the event for the first user; communicating a written message to the first user device, the written message including the predicted second time and location of the event for the first user; presenting the event to the first user at the predicted second time and location; and The system of claim 1 configured to:

11. The system of claim 10, wherein the first user and the second user are located in different geographical regions.

12. The system of claim 10, wherein the first user and the second user are located in different time zones.

13. The system of claim 10, wherein the first user and the second user are located in different countries.

14. The system of claim 1, wherein the server is further configured to track the virtual object.

Citation Information

Patent Citations

  • Method and system for three-dimensional virtual reality space, medium and method for recording information, medium and method for transmitting information, information processing method, client terminal, and common-use server terminal

    JP1997081781A

  • Object controller, object control method and recording medium for the same

    JP2002063125A

  • Electronic shopping system for multiple persons

    JP2005182231A

  • Information processing method, information processor, and remote mixed reality sharing device

    JP2006293604A

  • Method and Apparatus to Facilitate a Differently Configured Virtual Reality Experience for Some Participants in a Communication Session

    US20080231626A1