Systems and methods for augmented and virtual reality - Patents.com
The computer network system enables seamless interaction between multiple users in virtual and augmented reality environments by processing and transmitting virtual world data, addressing the limitations of existing systems and enhancing immersive and collaborative experiences.
Patent Information
- Application Number
- JP2023145375
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2011-10-28
- Filing Date
- 2023-09-07
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2032-10-29
AI Technical Summary
Existing virtual and augmented reality systems lack the capability to enable seamless interaction between multiple users in different physical locations, limiting the immersive and collaborative aspects of these environments.
A computer network system comprising one or more computing devices with memory, processing circuits, and software that processes virtual world data, allowing for the transmission and presentation of virtual world data from one user to another, regardless of their physical location, and enabling interactive experiences in augmented and virtual reality modes.
The system facilitates immersive and collaborative interactions between multiple users in virtual and augmented reality environments, allowing for real-time data exchange and dynamic changes in the virtual world based on user interactions and physical objects, thereby enhancing the overall user experience.
Smart Images

Figure 0007682963000001 
Figure 0007682963000002 
Figure 0007682963000003
Abstract
Description
[Technical field]
[0001] (Related Application Data) This application claims priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 61 / 552,941, filed October 28, 2011, the entirety of which is hereby incorporated by reference into this application.
[0002] BACKGROUND OF THEINVENTION The present invention relates generally to systems and methods configured to facilitate an interactive virtual or augmented reality environment for one or more users. [Background technology]
[0003] (background) Virtual and augmented reality environments are generated by a computer using, in part, data that represents the environment. This data may represent, for example, various objects that a user may sense and interact with. Examples of these objects include objects that are rendered and displayed for the user to see, audio that is played for the user to hear, and tactile (or haptic) feedback that the user feels. A user may sense and interact with virtual and augmented reality environments through a variety of visual, auditory, and tactile means. Summary of the Invention [Means for solving the problem]
[0004] One embodiment relates to a system for enabling two or more users to interact with a virtual world including virtual world data, the system comprising a computer network comprising one or more computing devices, the one or more computing devices comprising a memory, a processing circuit, and software at least partially stored in the memory and executable by the processing circuit to process at least a portion of the virtual world data, where at least a first portion of the virtual world data originates from a first user virtual world that is local to the first user, and the computer network is operable to transmit the first portion to the user device for presentation to a second user, whereby the second user may experience the first portion from a location of the second user, and aspects of the first user virtual world are effectively passed on to the second user. The first user and the second user may be in different physical locations, or substantially the same physical location. At least a portion of the virtual world may be configured to change in response to a change in the virtual world data. At least a portion of the virtual world may be configured to change in response to a physical object sensed by the user device. The change in the virtual world data may represent a virtual object having a predetermined relationship to the physical object. The changes in the virtual world data may be presented to a second user device for presentation to a second user according to the predetermined relationship. The virtual world may be operable to be rendered by at least one of the computer server or the user device. The virtual world may be presented in a two-dimensional format. The virtual world may be presented in a three-dimensional format. The user device may be operable to provide an interface for enabling interaction between a user and the virtual world in an augmented reality mode. The user device may be operable to provide an interface for enabling interaction between a user and the virtual world in a virtual reality mode. The user device may be operable to provide an interface for enabling interaction between a user and the virtual world in a combination of the augmented reality mode and the virtual reality mode.The virtual world data may be transmitted over a data network. The computer network may be operable to receive at least a portion of the virtual world data from a user device. At least a portion of the virtual world data transmitted to the user device may comprise instructions for generating at least a portion of the virtual world. At least a portion of the virtual world data may be transmitted to a gateway for at least one of processing or distribution. At least one of the one or more computer servers may be operable to process the virtual world data distributed by the gateway.
[0005] Another embodiment is directed to a system for virtual and / or augmented user experiences, in which remote avatars are animated based at least in part on data on a wearable device, with optional input from voice intonation and facial recognition software.
[0006] Another embodiment is directed to a system for virtual and / or augmented user experiences, where the camera pose or viewpoint position and vector may be located anywhere within the world sector.
[0007] Another embodiment is directed to a system for virtual and / or augmented user experiences, in which a world, or portions thereof, may be rendered at various and selectable scales for an observing user.
[0008] Another embodiment is directed to a system for virtual and / or augmented user experiences, where features such as points or parametric lines, in addition to pose-tagged images, may be utilized as basis data for a world model from which a software robot or object recognizer may be utilized to create parametric representations of real-world objects that tag source features for mutual inclusion in segmented objects and the world model. The present invention provides, for example, the following: (Item 1) 1. A system for enabling two or more users to interact within a virtual world including virtual world data, the system comprising: A computer network comprising one or more computing devices, the one or more computing devices comprising a memory, a processing circuit, and software at least partially stored in the memory and executable by the processing circuit to process at least a portion of the virtual world data. wherein at least a first portion of the virtual world data originates from a first user virtual world that is local to a first user, and the computer network is operable to transmit the first portion to a user device for presentation to a second user, whereby the second user may experience the first portion from the location of the second user, effectively passing aspects of the first user virtual world to the second user. (Item 2) 2. The system of claim 1, wherein the first user and the second user are in different physical locations. (Item 3) 2. The system of claim 1, wherein the first user and the second user are in substantially the same physical location. (Item 4) 2. The system of claim 1, wherein at least a portion of the virtual world changes in response to changes in the virtual world data. (Item 5) 2. The system of claim 1, wherein at least a portion of the virtual world changes in response to a physical object sensed by the user device. (Item 6) 6. The system of claim 5, wherein the changes in the virtual world data represent virtual objects having a predetermined relationship with the physical objects. (Item 7) 7. The system of claim 6, wherein changes in the virtual world data are presented to a second user device for presentation to the second user in accordance with the predetermined relationship. (Item 8) 2. The system of claim 1, wherein the virtual world is operable to be rendered by at least one of the computer server or a user device. (Item 9) 2. The system of claim 1, wherein the virtual world is presented in two-dimensional format. (Item 10) 2. The system of claim 1, wherein the virtual world is presented in three-dimensional format. (Item 11) 2. The system of claim 1, wherein the user device is operable to provide an interface to enable interaction between a user and the virtual world in an augmented reality mode. (Item 12) 2. The system of claim 1, wherein the user device is operable to provide an interface in a virtual reality mode to enable interaction between a user and the virtual world. (Item 13) Item 12. The system of item 11, wherein the user device is operable to provide an interface to enable interaction between a user and the virtual world in a combination of augmented reality and virtual reality modes. (Item 14) 2. The system of claim 1, wherein the virtual world data is transmitted over a data network. (Item 15) 2. The system of claim 1, wherein the computer network is operable to receive at least a portion of the virtual world data from a user device. (Item 16) 2. The system of claim 1, wherein at least a portion of the virtual world data transmitted to the user device comprises instructions for generating at least a portion of the virtual world. (Item 17) 2. The system of claim 1, wherein at least a portion of the virtual world data is transmitted to a gateway for at least one of processing or distribution. (Item 18) 20. The system of claim 17, wherein at least one of the one or more computer servers is operable to process virtual world data delivered by the gateway. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 illustrates an exemplary embodiment of the disclosed system for facilitating an interactive virtual reality or augmented reality environment for multiple users. [Diagram 2] FIG. 2 illustrates an example of a user device that interacts with the system illustrated in FIG. [Diagram 3] FIG. 3 illustrates an exemplary embodiment of a mobile wearable user device. [Figure 4] FIG. 4 illustrates an example of objects viewed by a user when the mobile wearable user device of FIG. 3 is operating in an enhanced mode. [Diagram 5] FIG. 5 illustrates an example of objects viewed by a user when the mobile wearable user device of FIG. 3 is operating in virtual mode. [Figure 6] FIG. 6 illustrates an example of objects viewed by a user when the mobile wearable user device of FIG. 3 is operating in a mixed virtual interface mode. [Figure 7] FIG. 7 illustrates an embodiment in which two users located in different geographic locations each interact with the other user and a common virtual world through their respective user devices. [Figure 8] FIG. 8 illustrates an embodiment in which the embodiment of FIG. 7 is extended to include the use of a haptic device. [Figure 9A]FIG. 9A illustrates an example of a mixed mode interfacing where a first user is interfacing with a digital world in a mixed virtual interface mode and a second user is interfacing with the same digital world in a virtual reality mode. [Figure 9B] FIG. 9B illustrates another example of a mixed mode interface in which a first user interfaces with a digital world in a mixed virtual interface mode and a second user interfaces with the same digital world in an augmented reality mode. [Figure 10] FIG. 10 illustrates an example illustration of a user's field of view when interfacing with the system in augmented reality mode. [Figure 11] FIG. 11 illustrates an example illustration of a user's view showing a virtual object triggered by a physical object when the user is interfacing with the system in an augmented reality mode. [Figure 12] FIG. 12 illustrates one embodiment of an integrated augmented reality and virtual reality configuration in which one user in an augmented reality experience visualizes the presence of another user in a virtual reality experience. [Figure 13] FIG. 13 illustrates one embodiment of a time and / or contingency based augmented reality experience configuration. [Figure 14] FIG. 14 illustrates one embodiment of a user display configuration suitable for virtual reality and / or augmented reality experiences. [Figure 15] FIG. 15 illustrates one embodiment of local and cloud-based computational collaboration. [Figure 16] FIG. 16 illustrates various aspects of the alignment configuration. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] (Detailed Description) Referring to Figure 1, system 100 is representative hardware for implementing the processes described below. The representative system comprises a computing network 105 consisting of one or more computer servers 110 connected through one or more high bandwidth interfaces 115. The servers in a computing network need not be co-located. The one or more servers 110 each comprise one or more processors for executing program instructions. The servers also include memory for storing program instructions and data used and / or generated by processes being executed by the servers under the direction of the program instructions.
[0011] The computing network 105 communicates data among the servers 110, and between the servers and one or more user devices 120 over one or more data network connections 130. Examples of such data networks include, without limitation, all manner of public and private data networks, both mobile and wired, including the many interconnections of such networks commonly referred to as, for example, the Internet. No particular media, topology, or protocol is intended to be implied by the diagram.
[0012] The user devices are configured to communicate directly with either the computing network 105, or the server 110. Alternatively, the user devices 120 communicate with the remote server 110, and, if necessary, locally with other user devices, through a specially programmed local gateway 140 for processing data and / or communicating data between the network 105 and one or more local user devices 120.
[0013] As illustrated, the gateway 140 is implemented as a separate hardware component including a processor for executing software instructions and a memory for storing software instructions and data. The gateway has its own wired and / or wireless connection to a data network for communicating with the servers 110 that make up the computing network 105. Alternatively, the gateway 140 can be integrated with a user device 120 worn or carried by a user. For example, the gateway 140 may be implemented as a downloadable software application that is installed and operates on a processor included in the user device 120. The gateway 140, in one embodiment, provides access to the computing network 105 to one or more users via the data network 130.
[0014] The servers 110 each include, for example, working memory and storage devices for storing data and software programs, microprocessors for executing program instructions, and graphics processors and other specialized processors for rendering and generating graphics, image, video, audio, and multimedia files. The computing network 105 may also include devices for storing data that is accessed, used, or created by the servers 110.
[0015] The server, and optionally software programs running on the user device 120 and gateway 140, are used to generate a digital world (also referred to herein as a virtual world) in which the user interacts with the user device 120. The digital world is represented by data and processes that represent and / or define virtual, non-existent entities, environments, and conditions that may be presented to the user through the user device 120 for the user to experience and interact with. For example, any type of object, entity, or item that appears to be physically present when instantiated in a scene viewed or experienced by the user may include a description of its appearance, its behavior, how the user is permitted to interact with it, and other characteristics. The data used to create the environment of the virtual world (including virtual objects) may include, for example, atmospheric data, terrain data, weather data, temperature data, location data, and other data used to define and / or represent the virtual environment. Additionally, the data defining various conditions governing operation of the virtual world may include, for example, physical laws, time, spatial relationships, and other data that may be used to define and / or create various conditions governing operation of the virtual world (including virtual objects).
[0016] Entities, objects, conditions, properties, behaviors, or other features of the digital world are generally referred to herein as objects (e.g., digital objects, virtual objects, rendered physical objects, etc.) unless the context dictates otherwise. Objects may be any type of animate or inanimate object, including, but not limited to, structures, plants, vehicles, people, animals, creatures, machines, data, videos, text, photos, and other users. Objects may also be defined in the digital world to store information about items, behaviors, or conditions that actually exist in the physical world. Data that represents or defines an entity, object, or item, or that stores its current state, is generally referred to herein as object data. This data is processed by the server 110, or by the gateway 140 or user device 120, depending on the implementation, to instantiate an instance of the object and render the object in an appropriate manner for a user to experience through the user device.
[0017] Programmers who develop and / or create the digital world create or define the objects and the conditions under which the objects are instantiated. However, the digital world may allow others to create or modify the objects. Once an object is instantiated, the state of the object may be permitted to be changed, controlled, or manipulated by one or more users experiencing the digital world.
[0018] For example, in one embodiment, the development, production, and management of the digital world is generally provided by one or more system administration programmers. In some embodiments, this may include the development, design, and / or execution of storylines, themes, and events in the digital world, as well as the delivery of the stories through various forms of events and media, such as, for example, film, digital, networked, mobile, augmented reality, and live entertainment. System administration programmers may also handle the technical management, moderation, and curation of the digital world and its associated user community, as well as other tasks typically performed by network administrators.
[0019] A user interacts with one or more digital worlds using some type of local computing device, generally designated as user device 120. Examples of such user devices include, but are not limited to, a smartphone, a tablet device, a heads-up display (HUD), a gaming console, or any other device capable of communicating data and providing an interface or display to a user, or a combination of such devices. In some embodiments, user device 120 may include or communicate with local peripheral or input / output components, such as, for example, a keyboard, a mouse, a joystick, a game controller, a haptic interface device, a motion capture controller, an optical tracking device (e.g., those available from Leap Motion, Inc. or from Microsoft under the trade name Kinect (RTM)), audio equipment, voice equipment, projector systems, 3D displays, and holographic 3D contact lenses.
[0020] An example of a user device 120 for interacting with the system 100 is illustrated in Fig. 2. In the exemplary embodiment shown in Fig. 2, a user 210 may interface with one or more digital worlds through a smartphone 220. The gateway is implemented by a software application 230 stored and running on the smartphone 220. In this particular example, the data network 130 includes a wireless mobile network that connects the user device (i.e., smartphone 220) to the computer network 105.
[0021] In one implementation of a preferred embodiment, the system 100 is capable of supporting a large number of concurrent users (e.g., millions of users) each interfacing with the same digital world or with multiple digital worlds using some type of user device 120.
[0022] The user device provides the user with an interface to enable visual, audible, and / or physical interaction between the user and the digital world generated by the server 110, including other users and objects (real or virtual) presented to the user. The interface provides the user with a rendered view that can be seen, heard, or otherwise sensed, and the ability to interact with that view in real time. The manner in which the user interacts with the rendered view may be dictated by the capabilities of the user device. For example, if the user device is a smartphone, user interaction may be implemented by the user touching a touch screen. In another example, if the user device is a computer or gaming console, user interaction may be implemented using a keyboard or game controller. The user device may include additional components that enable user interaction, such as sensors, and objects and information (including gestures) detected by the sensors may be provided as input representing the user interaction with the virtual world using the user device.
[0023] The rendered views can be presented in a variety of formats, such as, for example, two-dimensional or three-dimensional visual displays (including projections), sound, and haptic or tactile feedback. The rendered views may be interfaced with by the user in one or more modes, including, for example, augmented reality, virtual reality, and combinations thereof. The format of the rendered views, as well as the interface modes, may be dictated by one or more of the user device, data processing capabilities, user device connectivity, network capacity, and system workload. Having multiple users simultaneously interacting with the digital world, and the real-time nature of the data exchange, is made possible by the computing network 105, the server 110, the gateway component 140 (as needed), and the user devices 120.
[0024] In one embodiment, the computing network 105 is comprised of a large-scale computing system having single and / or multi-core servers (i.e., servers 110) connected through high-speed connections (e.g., high-bandwidth interfaces 115). The computing network 105 may form a cloud or grid network. Each of the servers includes memory or is coupled with computer-readable memory for storing software for implementing data to create, design, modify, or process objects in the digital world. These objects and their instantiations may be dynamic, appearing, disappearing, changing over time, and changing in response to other conditions. Examples of dynamic capabilities of objects are generally discussed herein with respect to various embodiments. In some embodiments, each user that interfaces with the system 100 may also be represented as an object and / or collection of objects within one or more digital worlds.
[0025] The servers 110 in the computing network 105 also store computational state data for each of the digital worlds. Computational state data (also referred to herein as state data) may be a component of object data and generally defines the state of an instance of an object at a given instance in time. Thus, the computational state data may change over time and may be affected by the actions of one or more users and / or programmers maintaining the system 100. When a user affects the computational state data (or other data that makes up the digital world), the user directly modifies or otherwise manipulates the digital world. If the digital world is shared or interfaces with other users, the user's actions may affect what is experienced by other users interacting with the digital world. Thus, in some embodiments, changes made to the digital world by a user are experienced by other users that interface with the system 100.
[0026] Data stored on one or more servers 110 in the computing network 105 is transmitted or deployed, in one embodiment, at high speed and low latency to one or more user devices 120 and / or gateway components 140. In one embodiment, the object data shared by the server may be complete or may be compressed and include instructions to recreate the complete object data at the user's end, and may be rendered and visualized by the user's local computing device (e.g., gateway 140 and / or user device 120). Software running on the servers 110 of the computing network 105 may, in some embodiments, adapt the data that the computing network 105 generates and sends to a particular user's device 120 for objects in the digital world (or any other data exchanged by the computing network 105) depending on the user's particular device and bandwidth. For example, as a user interacts with the digital world through a user device 120, the server 110 may recognize the particular type of device being used by the user, the device's connectivity and / or the available bandwidth between the user device and the server, and appropriately determine and balance the size of the data being delivered to the device to optimize the user interaction. An example of this may include reducing the size of the transmitted data to a lower resolution quality so that the data can be displayed on a particular user device having a lower resolution display. In a preferred embodiment, the computing network 105 and / or gateway component 140 delivers data to the user device 120 at 15 frames per second or faster and at a rate sufficient to present an interface operating at high definition quality or higher resolution.
[0027] The gateway 140 provides a local connection to the computing network 105 for one or more users. In some embodiments, it may be implemented by a downloadable software application running on the user device 120 or another local device such as that shown in FIG. 2. In other embodiments, it may be implemented by a hardware component (a component having a processor with appropriate software / firmware stored on the component) that is in communication with the user device 120 but is either not integrated into or attached to the user device 120 or is integrated into the user device 120. The gateway 140 communicates with the computing network 105 via the data network 130 and provides data exchange between the computing network 105 and one or more local user devices 120. As discussed in more detail below, the gateway component 140 may include software, firmware, memory, and processing circuitry and may be capable of processing data communicated between the network 105 and one or more local user devices 120.
[0028] In some embodiments, the gateway component 140 monitors and adjusts the rate at which data is exchanged between the user device 120 and the computer network 105 to enable optimal data throughput for a particular user device 120. For example, in some embodiments, the gateway 140 buffers and downloads both static and dynamic aspects of the digital world, even beyond the field of view presented to the user through an interface connected to the user device. In such embodiments, instances of static objects (structured data, software-implemented methods, or both) may be stored in memory (local to the gateway component 140, the user device 120, or both) and are referenced to the local user's current location as indicated by data provided by the computing network 105 and / or the user's device 120. Instances of dynamic objects, which may include, for example, intelligent software agents and objects controlled by other users and / or the local user, are stored in a high-speed memory buffer. Dynamic objects, which represent two- or three-dimensional objects within the view presented to the user, may be categorized into component shapes, such as, for example, static shapes that move but do not change, and dynamic shapes that change. Some of the dynamic objects that are changing can be updated by a real-time threaded high priority data stream from the server 110 through the computing network 105 managed by the gateway component 140. As an example of a priority threaded data stream, data that is within 60 degrees of the field of view of the user's eye may be given higher priority than data that is further out. Another example includes prioritizing dynamic characters and / or objects within the user's field of view over static objects in the background.
[0029] In addition to managing the data connection between the computing network 105 and the user device 120, the gateway component 140 may store and / or process data that may be presented to the user device 120. For example, the gateway component 140 may in some embodiments receive compressed data from the computing network 105, e.g., representing graphical objects to be rendered for viewing by a user, and perform advanced rendering techniques to reduce the data load transmitted from the computing network 105 to the user device 120. In another example where the gateway 140 is a separate device, the gateway 140 may store and / or process data for local instances of objects rather than communicating the data to the computing network 105 for processing.
[0030] Referring now also to FIG. 3, the digital world may be experienced by one or more users in various forms that may depend on the capabilities of the user's device. In some embodiments, the user device 120 may include, for example, a smartphone, a tablet device, a head-up display (HUD), a gaming console, or a wearable device. Generally, the user device includes a processor for executing program code stored in a memory on the device, coupled with a display, and a communication interface. An exemplary embodiment of a user device is illustrated in FIG. 3, where the user device comprises a mobile wearable device, i.e., a head-mounted display system 300. According to an embodiment of the present disclosure, the head-mounted display system 300 includes a user interface 302, a user sensing system 304, an environmental sensing system 306, and a processor 308. Although the processor 308 is illustrated in FIG. 3 as a standalone component separate from the head-mounted system 300, in alternative embodiments, the processor 308 may be integrated with one or more components of the head-mounted system 300 or may be incorporated into other system 100 components, such as, for example, the gateway 140.
[0031] The user device presents the user with an interface 302 for interacting with and experiencing the digital world. Such interactions may include the user and the digital world, one or more other users interfacing with the system 100, and objects within the digital world. The interface 302 generally provides visual and / or audio sensory input (and in some embodiments physical sensory input) to the user. Thus, the interface 302 may include a speaker (not shown) and a display component 303 that, in some embodiments, may enable stereoscopic 3D viewing and / or 3D viewing that embodies more natural properties of the human visual system. In some embodiments, the display component 303 may comprise a transparent interface (such as a transparent OLED) that, when in an "off" setting, enables an optically correct view of the user's surrounding physical environment with little optical distortion or computing overlay. As discussed in more detail below, the interface 302 may include additional settings that enable various visual / interface capabilities and functionality.
[0032] The user sensing system 304 may, in some embodiments, include one or more sensors 310 operable to detect certain characteristics, properties, or information relating to an individual user wearing the system 300. For example, in some embodiments, the sensor 310 may include a camera or optical detection / scanning circuitry capable of detecting real-time optical properties / measurements of the user, such as, for example, one or more of pupil contraction / dilation, angular measurements / position of each pupil, sphericity, eye shape (as eye shape changes over time), and other anatomical data. This data may provide or be used to calculate information (e.g., the user's visual focus) that may be used by the head-worn system 300 and / or the interface system 100 to optimize the user's viewing experience. For example, in one embodiment, the sensors 310 may each measure the pupil contraction rate of each of the user's eyes. This data may be transmitted to the processor 308 (or to the gateway component 140 or to the server 110), and the data may be used to determine the user's reaction to, for example, the brightness settings of the interface display 303. The interface 302 may be adjusted according to the user's response, for example, by dimming the display 303 if the user's response indicates that the brightness level of the display 303 is too high. The user sensing system 304 may include other components than those discussed above or illustrated in FIG. 3. For example, in some embodiments, the user sensing system 304 may include a microphone for receiving audio input from the user. The user sensing system may also include one or more infrared camera sensors, one or more visible spectrum camera sensors, structured light emitters and / or sensors, infrared light emitters, coherent light emitters and / or sensors, gyros, accelerometers, magnetometers, proximity sensors, GPS sensors, ultrasonic emitters and detectors, and a haptic interface.
[0033] The environmental sensing system 306 includes one or more sensors 312 for acquiring data from the physical environment around the user. Objects or information detected by the sensors may be provided as input to the user device. In some embodiments, this input may represent user interaction with the virtual world. For example, a user viewing a virtual keyboard on a desk may gesture with their fingers as if they were typing on the virtual keyboard. The motion of the moving fingers may be captured by the sensors 312 and provided as input to the user device or system, where the input may be used to change the virtual world or to create new virtual objects. For example, the finger motions may be recognized (using a software program) as typing, and the recognized typing gestures may be combined with known positions of virtual keys on the virtual keyboard. The system may then render a virtual monitor that is displayed to the user (or other users interfacing with the system), where the virtual monitor displays the text being typed by the user.
[0034] The sensor 312 may include, for example, a generally outward-facing camera, or a scanner, to interpret sight information, for example, through continuously and / or intermittently projected infrared structured light. The environmental sensing system 306 may be used to map one or more elements of the user's surrounding physical environment by detecting and registering static objects, dynamic objects, people, gestures, and the local environment, including various lighting, atmospheric, and acoustic conditions. Thus, in some embodiments, the environmental sensing system 306 may include image-based 3D reconstruction software that is incorporated into a local computing system (e.g., the gateway component 140 or the processor 308) and is operable to digitally reconstruct one or more objects or information detected by the sensor 312. In one exemplary embodiment, the environmental sensing system 306 provides one or more of motion capture data (including gesture recognition), depth sensing, face recognition, object recognition, unique object feature recognition, voice / audio recognition and processing, sound source localization, noise reduction, infrared or similar laser projection, as well as monochrome and / or color CMOS sensors (or other similar sensors), field of view sensors, and various other light-enhancing sensors. It should be understood that the environmental sensing system 306 may include other components than those discussed above or illustrated in FIG. 3. For example, in some embodiments, the environmental sensing system 306 may include a microphone for receiving audio from the local environment. The user sensing system may also include one or more infrared camera sensors, one or more visible spectrum camera sensors, structured light emitters and / or sensors, infrared light emitters, coherent light emitters and / or sensors, gyros, infrared light emitters, accelerometers, magnetometers, proximity sensors, GPS sensors, ultrasonic emitters and detectors, and haptic interfaces.
[0035] As mentioned above, the processor 308 may in some embodiments be integrated with other components of the head-mounted system 300, integrated with other components of the interface system 100, or may be a stand-alone device (wearable or separate from the user) as shown in FIG. 3. The processor 308 may be connected to various components of the head-mounted system 300 and / or components of the interface system 100 through a physical wired connection or through a wireless connection such as, for example, a mobile network connection (including cellular and data networks), Wi-Fi, or Bluetooth. The processor 308 may include memory modules, integrated and / or additional graphics processing units, wireless and / or wired Internet connectivity, and codecs and / or firmware capable of converting data from a source (e.g., the computing network 105, the user sensing system 304, the environmental sensing system 306, or the gateway component 140) into image and audio data, and the images / video and audio may be presented to the user via the interface 302.
[0036] The processor 308 handles data processing for the various components of the head-mounted system 300, as well as data exchange between the head-mounted system 300 and the gateway component 140 (the computing network 105 in some embodiments). For example, the processor 308 may be used to buffer and process data streaming between the user and the computing network 105, thereby enabling a smooth, continuous, and high-fidelity user experience. In some embodiments, the processor 308 may process data at a rate sufficient to achieve anywhere between 8 frames per second at 320×240 resolution to 24 frames per second at high definition resolution (1280×720), or more (e.g., 60-120 frames per second and 4k resolution or higher (10k+ resolution and 50,000 frames per second)). Additionally, the processor 308 may store and / or process data that may be presented to the user rather than being streamed in real time from the computing network 105. For example, the processor 308, in some embodiments, may receive compressed data from the computing network 105 and perform advanced rendering techniques (such as lighting or shading) to reduce the data load transmitted from the computing network 105 to the user device 120. In another example, the processor 308 may store and / or process local object data rather than transmitting the data to the gateway component 140 or to the computing network 105.
[0037] The head-mounted system 300 may, in some embodiments, include various settings or modes that enable various visual / interface capabilities and functionality. The modes may be selected manually by the user or automatically by components of the head-mounted system 300 or the gateway component 140. As previously mentioned, one embodiment of the head-mounted system 300 includes an "off" mode in which the interface 302 does not provide substantially any digital or virtual content. In the off mode, the display component 303 may be transparent, thereby enabling an optically correct view of the user's surrounding physical environment with little optical distortion or computing overlay.
[0038] In one exemplary embodiment, the head-mounted system 300 includes an "augmented" mode in which the interface 302 provides an augmented reality interface. In the augmented mode, the interface display 303 may be substantially transparent, thereby allowing the user to view the local physical environment. At the same time, virtual object data provided by the computing network 105, the processor 308, and / or the gateway component 140 is presented on the display 303 in combination with the local physical environment.
[0039] 4 illustrates an example embodiment of objects viewed by a user when the interface 302 is operating in an augmented mode. As shown in FIG. 4, the interface 302 presents a physical object 402 and a virtual object 404. In the embodiment illustrated in FIG. 4, the physical object 402 is an actual physical object that exists in the user's local environment, while the virtual object 404 is an object created by the system 100 and displayed via the user interface 302. In some embodiments, the virtual object 404 may be displayed at a fixed position or location within the physical environment (e.g., a virtual monkey standing next to a particular road sign located in the physical environment) or may be displayed to the user as an object located at a position relative to the user interface / display 303 (e.g., a virtual clock or thermometer visible in the top left corner of the display 303).
[0040] In some embodiments, the virtual object may be cued from or triggered by an object that is physically present within or outside the user's field of view. The virtual object 404 is cued from or triggered by the physical object 402. For example, the physical object 402 may actually be a stool, and the virtual object 404 may be displayed to the user (and in some embodiments, to other users interfacing with the system 100) as a virtual animal standing on the stool. In such an embodiment, the environmental sensing system 306 may use software and / or firmware stored, for example, in the processor 308, to recognize various features and / or shape patterns (captured by the sensor 312) to identify the physical object 402 as a stool. These recognized shape patterns, such as the top of the stool, may be used to trigger the placement of the virtual object 404. Other examples include walls, tables, furniture, cars, buildings, people, floors, plants, animals, and any object that can be seen may be used to trigger an augmented reality experience that has some relationship to one or more objects.
[0041] In some embodiments, the particular virtual object 404 to be evoked may be selected by a user or may be automatically selected by other components of the head-mounted system 300 or the interface system 100. Additionally, in embodiments in which a virtual object 404 is automatically evoked, the particular virtual object 404 may be selected based on the particular physical object 402 (or characteristics thereof) from which the virtual object 404 is cued or evoked. For example, if the physical object is identified as a diving board extending over a pool, the evoked virtual object may be a living being wearing a snorkel, swimsuit, flotation device, or other related item.
[0042] In another exemplary embodiment, the head-mounted system 300 may include a "virtual" mode in which the interface 302 provides a virtual reality interface. In the virtual mode, the physical environment is omitted from the display 303, and virtual object data provided by the computing network 105, the processor 308, and / or the gateway component 140 is presented on the display 303. Omission of the physical environment may be achieved by physically blocking the visual display 303 (e.g., by a cover) or through a feature of the interface 302 in which the display 303 transitions to an opaque setting. In the virtual mode, live and / or stored visual and audio sensations may be presented to the user through the interface 302, and the user experiences and interacts with the digital world (digital objects, other users, etc.) through the virtual mode of the interface 302. Thus, the interface provided to the user in the virtual mode is composed of virtual object data that includes the virtual digital world.
[0043] Figure 5 illustrates an example embodiment of a user interface when the head-mounted interface 302 is operating in a virtual mode. As shown in Figure 5, the user interface presents a virtual world 500 made up of digital objects 510, which may include atmosphere, weather, terrain, structures, and people. Although not shown in Figure 5, the digital objects may also include, for example, plants, vehicles, animals, creatures, machines, artificial intelligence, location information, and any other objects or information that define the virtual world 500.
[0044] In another exemplary embodiment, the head-mounted system 300 may include a "mixed" mode, where various features of the head-mounted system 300 (as well as features of the virtual and augmented modes) may be combined to create one or more custom interface modes. In one example custom interface mode, the physical environment is omitted from the display 303, and virtual object data is presented on the display 303 in a manner similar to the virtual mode. However, in this example custom interface mode, the virtual objects may be fully virtual (i.e., they do not exist in the local physical environment), or they may be actual local physical objects that are rendered as virtual objects in the interface 302 instead of physical objects. Thus, in a particular custom mode (referred to herein as a mixed virtual interface mode), live and / or stored visual and audio sensations may be presented to the user through the interface 302, and the user experiences and interacts with a digital world that includes fully virtual objects and rendered physical objects.
[0045] FIG. 6 illustrates an exemplary embodiment of a user interface operating according to a mixed virtual interface mode. As shown in FIG. 6, the user interface presents a virtual world 600 composed of full virtual objects 610 and rendered physical objects 620 (renderings of objects that are otherwise physically present in the scene). According to the example illustrated in FIG. 6, the rendered physical objects 620 include a building 620A, a ground 620B, and a platform 620C, and are shown with a thick outline 630 to indicate to the user that the objects are rendered. In addition, the full virtual objects 610 include an additional user 610A, a cloud 610B, a sun 610C, and a flame 610D above the platform 620C. It should be understood that the full virtual objects 610 may include, for example, atmosphere, weather, terrain, buildings, people, plants, vehicles, animals, creatures, machines, artificial intelligence, location information, and any other objects or information that define the virtual world 600 and that are not rendered from objects present in the local physical environment. Conversely, rendered physical object 620 is an actual local physical object that is rendered as a virtual object in interface 302. Thick outline 630 represents one embodiment for showing the rendered physical object to the user. Thus, the rendered physical object may be shown using methods other than those disclosed herein.
[0046] In some embodiments, the rendered physical objects 620 may be detected using the sensors 312 of the environment sensing system 306 (or using other devices such as a motion or image capture system) and converted into digital object data, for example, by software and / or firmware stored in the processing circuitry 308. Thus, when a user interfaces with the system 100 in a mixed virtual interface mode, various physical objects may be displayed to the user as rendered physical objects. This may be particularly useful to allow the user to interface with the system 100 while still being able to safely navigate the local physical environment. In some embodiments, the user may be able to selectively remove or add rendered physical objects to the interface display 303.
[0047] In another example custom interface mode, the interface display 303 may be substantially transparent, thereby allowing the user to view the local physical environment while various local physical objects are displayed to the user as rendered physical objects. This example custom interface mode is similar to the augmented mode, except that one or more of the virtual objects may be rendered physical objects, as discussed above with respect to the previous example.
[0048] The foregoing custom interface mode examples represent some example embodiments of the various custom interface modes that can be provided by the mixed mode of head-mounted system 300. Thus, various other custom interface modes may be created from various combinations of the components of head-mounted system 300 and the features and functionality provided by the various modes discussed above without departing from the scope of this disclosure.
[0049] The embodiments discussed herein merely illustrate some examples for providing an interface that operates in off mode, augmented mode, virtual mode, or mixed mode, and are not intended to limit the scope or content of each interface mode or the functionality of the components of head-mounted system 300. For example, in some embodiments, virtual objects may include data displayed to a user (time, temperature, altitude, etc.), objects created and / or selected by system 100, objects created and / or selected by a user, or even objects representing other users interfacing with system 100. In addition, virtual objects may include extensions of physical objects (e.g., virtual statues growing out of a physical platform) and may be visually connected to or separate from physical objects.
[0050] Virtual objects may also be dynamic, changing over time, changing according to various relationships (e.g., position, distance, etc.) between the user or other users, physical objects, and other virtual objects, and / or changing according to other variables defined in the software and / or firmware of the head-mounted system 300, the gateway component 140, or the server 110. For example, in certain embodiments, virtual objects may respond to the user device or its components (e.g., a virtual ball moves when a haptic device is placed next to it), physical or verbal user interactions (e.g., a virtual creature runs away when a user approaches it or speaks when a user speaks to it), a chair is thrown at a virtual creature and the creature avoids the chair, other virtual objects (e.g., when a first virtual creature sees a second virtual creature, it reacts), physical variables such as position, distance, temperature, time, or other physical objects in the user's environment (e.g., a virtual creature shown standing on a physical road becomes flattened when a physical car passes by).
[0051] The various modes discussed herein may be applied to user devices other than the head-mounted system 300. For example, an augmented reality interface may be provided via a mobile phone or tablet device. In such an embodiment, the phone or tablet may use a camera to capture the physical environment around the user and virtual objects may be overlaid on the phone / tablet display screen. In addition, virtual modes may be provided by displaying a digital world on the phone / tablet display screen. Thus, these modes may be mixed to create various custom interface modes as described above using the phone / tablet components discussed herein, as well as other components connected to or used in combination with the user device. For example, mixed virtual interface modes may be provided by a computer monitor, television screen, or other device without a camera operating in combination with a motion or image capture system. In this exemplary embodiment, the virtual world may be viewed from the monitor / screen and object detection and rendering may be performed by the motion or image capture system.
[0052] 7 illustrates an exemplary embodiment of the present invention in which two users located at different geographic locations each interact with the other user and a common virtual world through their respective user devices. In this embodiment, two users 701 and 702 are throwing a virtual ball 703 (a type of virtual object) back and forth, and each user is able to observe the other user's effect on the virtual world (e.g., each user observes the virtual ball change direction, be caught by the other user, etc.). Because the movement and location of the virtual object (i.e., virtual ball 703) is tracked by server 110 in computing network 105, system 100 may, in some embodiments, communicate to users 701 and 702 the exact location and timing of the arrival of ball 703 to each user. For example, if a first user 701 is located in London, user 701 may throw ball 703 to a second user 702 located in Los Angeles at a speed calculated by system 100. Thus, the system 100 may communicate (e.g., via email, text message, instant message, etc.) the exact time and location of the ball's arrival to the second user 702. In this case, the second user 702 may use his device to watch the ball 703 arrive at the specified time and location. One or more users may also use geolocation mapping software (or the like) to track one or more virtual objects as they virtually travel around the globe. An example of this may be a user wearing a 3D head-mounted display looking up into the sky and seeing a virtual airplane flying overhead, superimposed on the real world. The virtual airplane may be flown by the user, by an intelligent software agent (software running on the user device or gateway), by other users, which may be locally and / or remotely, and / or any combination of these.
[0053] As mentioned above, the user device may include a haptic interface device that provides feedback (e.g., resistance, vibration, light, sound, etc.) to the user when the haptic device is determined by system 100 to be located at a physical spatial location relative to the virtual object. For example, the embodiment described above with respect to FIG. 7 may be extended to include the use of a haptic device 802, as shown in FIG.
[0054] In this exemplary embodiment, haptic device 802 may be displayed in the virtual world as a baseball bat. When ball 703 arrives, user 702 may swing haptic device 802 toward virtual ball 703. If system 100 determines that the virtual bat provided by haptic device 802 has "made contact" with ball 703, haptic device 802 may vibrate or provide other feedback to user 702, and virtual ball 703 may bounce off the virtual bat in a direction calculated by system 100 according to the detected speed, direction, and timing of contact between the ball and the bat.
[0055] The disclosed system 100, in some embodiments, may facilitate mixed-mode interfacing, where multiple users may interface with a common virtual world (and virtual objects contained therein) using different interface modes (e.g., augmented, virtual, mixed, etc.). For example, a first user interfacing with a particular virtual world in a virtual interface mode may interact with a second user interfacing with the same virtual world in an augmented reality mode.
[0056] 9A illustrates an example in which a first user 901 (interfacing with the digital world of the system 100 in a mixed virtual interface mode) and a first object 902 appear as virtual objects to a second user 922 interfacing with the same digital world of the system 100 in a full virtual reality mode. As described above, when interfacing with the digital world via the mixed virtual interface mode, local physical objects (e.g., the first user 901 and the first object 902) may be scanned and rendered as virtual objects in the virtual world. The first user 901 may be scanned, for example, by a motion capture system or similar device and rendered in the virtual world (by software / firmware stored in the motion capture system, the gateway component 140, the user device 120, the system server 110, or other devices) as a first rendered physical object 931. Similarly, the first object 902 may be scanned, for example, by the environmental sensing system 306 of the head-mounted interface 300 and rendered in the virtual world (by software / firmware stored in the processor 308, the gateway component 140, the system server 110, or other device) as a second rendered physical object 932. The first user 901 and the first object 902 are shown in the first portion 910 of FIG. 9A as physical objects in the physical world. In the second portion 920 of FIG. 9A, the first user 901 and the first object 902 are shown as first rendered physical object 931 and second rendered physical object 932 as they would appear to a second user 922 interfacing with the same digital world of the system 100 in a full virtual reality mode.
[0057] 9B illustrates another example embodiment of a mixed mode interface connection in which a first user 901 interfaces with a digital world in a mixed virtual interface mode, as discussed above, and a second user 922 interfaces with the same digital world (and the second user's local physical environment 925) in an augmented reality mode. In the embodiment of FIG. 9B, the first user 901 and the first object 902 are located at a first physical location 915, and the second user 922 is located at a different second physical location 925 separated by some distance from the first location 915. In this embodiment, the virtual objects 931 and 932 may be transposed in real time (or near real time) to a location in the virtual world that corresponds to the second location 925. Thus, the second user 922 may observe and interact with a rendered physical object 931 representing the first user 901 and a rendered physical object 932 representing the first object 902 in the second user's local physical environment 925.
[0058] FIG. 10 illustrates an exemplary illustration of a user's field of view when interfacing with the system 100 in an augmented reality mode. As shown in FIG. 10, the user sees a local physical environment (i.e., a city with multiple buildings) as well as a virtual character 1010 (i.e., a virtual object). The location of the virtual character 1010 may be triggered by 2D visual targets (e.g., a billboard, a postcard, or a magazine) and / or one or more 3D reference frames (e.g., buildings, cars, people, animals, airplanes, parts of buildings, and / or 3D physical objects, virtual objects, and / or combinations thereof). In the example illustrated in FIG. 10, known locations of buildings in the city may provide alignment references and / or information and key features for rendering the virtual character 1010. Additionally, the user's geospatial location (e.g., provided by GPS, attitude / position sensors, etc.) or mobile location relative to the buildings may include data used by the computing network 105 to trigger the transmission of data used to display the virtual character(s) 1010. In some embodiments, the data used to display the virtual character 1010 may include a rendered character 1010 and / or instructions (executed by the gateway component 140 and / or the user device 120) for rendering the virtual character 1010 or portions thereof. In some embodiments, if the user's geospatial location is unavailable or unknown, the server 110, the gateway component 140, and / or the user device 120 may still display the virtual object 1010 using an estimation algorithm that uses the user's last known location as a function of time and / or other parameters to estimate where certain virtual and / or physical objects may be located. This may also be used to determine the location of any virtual objects in the event that the user's sensors are obstructed and / or experience other malfunctions.
[0059] In some embodiments, the virtual character or virtual object may comprise a virtual figurine, and the rendering of the virtual figurine is triggered by a physical object. For example, referring now to FIG. 11, a virtual figurine 1110 may be triggered by an actual physical platform 1120. The triggering of the figurine 1110 may be in response to a visual object or feature (e.g., a fiducial, a design feature, a geometric shape, a pattern, a physical location, an elevation, etc.) detected by a user device or other component of the system 100. When a user views the platform 1120 without a user device, the user sees the platform 1120 without the figurine 1110. However, when a user views the platform 1120 through a user device, the user sees the figurine 1110 on the platform 1120, as shown in FIG. 11. The figurine 1110 is a virtual object and therefore may be stationary, active, change over time or relative to the user's viewing position, or even change depending on which particular user is viewing the figurine 1110. For example, if the user is a small child, the statue may be a dog, whereas if the viewer is an adult male, the statue may be a large robot as shown in FIG. 11. These are examples of user-dependent and / or state-dependent experiences, which allow one or more users to perceive one or more virtual objects, alone and / or in combination with physical objects, and experience customized and personalized versions of the virtual objects. The statue 1110 (or portions thereof) may be rendered by various components of the system, including, for example, software / firmware installed on the user device. Using data indicative of the position and pose of the user device, in combination with the alignment features of the virtual object (i.e., statue 1110), the virtual object (i.e., statue 1110) forms a relationship with the physical object (i.e., platform 1120).For example, the relationship between one or more virtual objects and one or more physical objects may be a function of distance, positioning, time, geolocation, proximity to one or more other virtual objects, and / or any other functional relationship involving any type of virtual and / or physical data. In some embodiments, image recognition software in the user device may further enhance the digital-physical object relationship.
[0060] The interactive interfaces provided by the disclosed systems and methods may be implemented to facilitate a variety of activities, such as, for example, interacting with one or more virtual environments and objects, interacting with other users, and experiencing various forms of media content, including advertisements, musical concerts, and movies. Thus, the disclosed systems facilitate user interaction such that users do not only watch or listen to media content, but rather actively participate in and experience the media content. In some embodiments, user participation may include modifying existing content or creating new content that is rendered in one or more virtual worlds. In some embodiments, the media content, and / or users creating content, may theme the creation of one or more virtual worlds.
[0061] In one embodiment, a musician (or other user) may create musical content that is rendered for users interacting with a particular virtual world. Musical content may include, for example, various singles, EPs, albums, videos, short films, and concert performances. In one embodiment, multiple users may interface with system 100 to simultaneously experience a virtual concert performed by a musician.
[0062] In some embodiments, produced media may include a unique identifier code associated with a particular entity (e.g., a band, artist, user, etc.). The code may be a set of alphanumeric characters, a UPC code, a QR code, a 2D image trigger, a 3D physical object feature trigger, or other form of digital mark, as well as sound, image, and / or both. In some embodiments, the code may also be embedded in digital media that may be interfaced using system 100. A user may obtain the code (e.g., by paying a fee) and redeem the code to access media content produced by the entity associated with the identifier code. Media content may be added to or removed from the user's interface.
[0063] In one embodiment, to avoid the computational and bandwidth limitations of passing real-time or near real-time video data from one computing system to another (e.g., from a cloud computing system to a local processor coupled to a user) with low latency, parametric information for various shapes and geometries may be transferred and utilized to define surfaces while textures are transferred and added to these surfaces resulting in static or dynamic details such as bitmap-based video details of the individual's face mapped onto the parametrically recreated facial geometry. As another example, if the system is configured to recognize the individual's face and knows that the individual's avatar is located in the augmented world, the system may be configured to pass the relevant world information and the individual's avatar information in one relatively large setup transfer, after which the remaining transfers to the local computing system, such as 308 depicted in FIG. 1, for local rendering may be limited to parameter and texture updates such as motion parameters of the individual's skeletal structure and movement bitmaps of the individual's face, at a much smaller bandwidth compared to the initial setup transfer or passing real-time video. Thus, cloud-based and local computing assets may be used in an integrated fashion where the cloud handles computations that do not require relatively low latency and the local processing assets handle low latency critical tasks, and in such cases the type of data transferred to the local system is preferably passed at a relatively low bandwidth due to the amount of such data (i.e., parametric information, textures, etc. as opposed to all real-time video).
[0064] Referring back to Figure 15, a schematic diagram illustrates the coordination between cloud computing assets (46) and local processing assets (308, 120). In one embodiment, the cloud (46) assets are operatively coupled directly (40, 42) to one or both of the local computing assets (120, 308), such as processor and memory configurations that may be housed in a structure configured to be coupled to a user's head (120) or belt (308), via wired or wireless networking (wireless being preferred for mobility and wired being preferred for certain high bandwidth or large data capacity transfers that may be desired). These computing assets that are local to the user may likewise be operatively coupled to each other via wired and / or wireless connection configurations (44). In one embodiment, to maintain low inertia and a small size head-mounted subsystem (120), the primary transfer between the user and the cloud (46) may be via a link between the belt-based subsystem (308) and the cloud, with the head-mounted subsystem (120) being primarily data tethered to the belt-based subsystem (308) using a wireless connection, such as, for example, an ultra-wideband ("UWB") connection as currently employed in personal computing peripheral connectivity applications.
[0065] With efficient local and remote processing coordination and an appropriate display device for the user (e.g., the user interface 302 or user "display device" featured in FIG. 3, the display 14 described below with reference to FIG. 14, or a variation thereof), aspects of one world relevant to the user's current real or virtual location may be transferred or "handed" to the user and efficiently updated. Indeed, in one embodiment, when one individual utilizes a virtual reality system ("VRS") in an augmented reality mode and another individual utilizes the VRS in a fully virtual mode to explore the same world local to the first individual, the two users may mutually experience the world in a variety of ways. For example, with reference to FIG. 12, a scenario similar to that described with reference to FIG. 11 is depicted with the addition of a visualization of the second user's avatar 2 flying through the augmented real world depicted from the fully virtual reality scenario. 12 may be experienced and displayed in augmented reality to a first individual, with two augmented reality elements (statue 1110 and the flying bumblebee avatar 2 of the second individual) displayed in addition to the actual physical elements of the local world surroundings within the view, such as the ground, background buildings, statue platform 1120, etc. Dynamic updates may be utilized to allow the first individual to visualize the progress of the second individual's avatar 2 as the avatar 2 flies through the world local to the first individual.
[0066] Again, with an arrangement as described above, where there is one world model that resides on and can be distributed from cloud computing resources, such a world may be "hand-over" to one or more users in a relatively low bandwidth format that is preferable for trying to distribute real-time video data or the like. An individual standing near the statue (i.e., as shown in FIG. 12) may have their augmented experience informed by the cloud-based world model, a subset of which may be hand-over to that individual and their local display device to complete the view. An individual sitting at a remote display device, which may be as simple as a personal computer located on a desk, can effectively download the same section of information from the cloud and have it rendered on the display. In fact, an individual physically present in the park near the statue may take a remotely located friend for a walk in the park, with the friend participating through virtual and augmented reality. The system needs to figure out where the paths are, where the trees are, where the statues are, but with that information on the cloud, a participating friend can download aspects of the scenario from the cloud and then start walking together as an augmented reality that is local to the individual actually in the park.
[0067] With reference to Figure 13, an embodiment based on time and / or contingency parameters is depicted in which an individual interacting with a virtual and / or augmented reality interface, such as the user interface 302 or user display device featured in Figure 3, the display device 14 described below with reference to Figure 14, or variations thereof, is utilizing the system (4) and enters a coffee shop to order a cup of coffee (6). The VRS may be configured to utilize sensing and data collection capabilities locally and / or remotely to provide an augmented and / or virtual reality display enhancement for the individual, such as a highlighted location of the coffee shop door, or a bubble window of a related coffee menu (8). Upon receipt of the individual's ordered cup of coffee, or upon detection of some other relevant parameter by the system, the system may be configured to display with the display device one or more time-based augmented or virtual reality images, videos, and / or sounds in the local environment (e.g., views of the Madagascar jungle from the walls and ceiling, either static or dynamic, with or without jungle sounds and other effects) (10). Such presentation to the user may be interrupted based on timing parameters (i.e., 5 minutes after a full coffee cup has been recognized and handed to the user, 10 minutes after the system recognizes the user walking through the front door of the store, etc.) or other parameters (e.g., the system's recognition that the user has finished drinking their coffee by noting the upside-down orientation of the coffee cup as the user takes a final sip of coffee from the cup, or the system's recognition that the user has walked out the front door of the store) (12).
[0068] Referring to FIG. 14, one embodiment of a suitable user display device (14) is shown that includes a display lens (82) that may be attached to a user's head or eye by a housing or frame (84). The display lens (82) includes one or more transparent mirrors positioned by the housing (84) in front of the user's eye (20), which may be configured to reflect the projected light (38) into the eye (20) to facilitate beam shaping, while also allowing at least some light transmission from the local environment in an augmented reality configuration (in a virtual reality configuration, it may be desirable for the display system 14 to be capable of blocking substantially all light from the local environment, such as by a darkened visor, blocking curtains, a full black LCD panel mode, or the like). In the depicted embodiment, two wide field machine vision cameras (16) are coupled to the housing (84) to image the environment around the user. In one embodiment, the cameras (16) are dual capture visible / infrared cameras. The depicted embodiment also includes a pair of scanning laser wavefront shaping (i.e., for depth) projector modules along with viewing mirrors and optics configured to project light (38) into the eye (20) as shown. The depicted embodiment also includes two miniature infrared cameras (24) paired with infrared light sources (26, light emitting diodes "LEDs," etc.) configured to track the user's eye (20) to aid in rendering and user input. The system (14) further features a sensor assembly (39) that includes X-axis, Y-axis, and Z-axis accelerometer capabilities, a magnetic compass, and X-axis, Y-axis, and Z-axis gyro capabilities, and may preferably provide data at a relatively high frequency, such as 200 Hz. The depicted system (14) also includes a head pose processor (36), such as an ASIC (application specific integrated circuit), FPGA (field programmable gate array), and / or ARM processor (advanced reduced instruction set machine), which may be configured to calculate real-time or near real-time user head pose from the wide field of view image information output from the capture device (16).Also shown is another processor (32) configured to perform digital and / or analog processing to derive pose from gyro, compass, and / or accelerometer data from the sensor assembly (39). The depicted embodiment also features a GPS (37, Global Positioning Satellite) subsystem to assist with pose and positioning. Finally, the depicted embodiment includes a rendering engine (34), which may feature hardware to operate software programs configured to provide rendering information that is local to the user to facilitate operation of the scanner and imaging into the user's eye for the user's view of the world. The rendering engine (34) is operatively coupled (81, 70, 76 / 78, 80, i.e., via wired or wireless connections) to the sensor pose processor (32), image pose processor (36), eye-tracking camera (24), and projection subsystem (18) such that light of the rendered augmented reality and / or virtual reality objects is projected using the scanning laser array (18) in a manner similar to a retinal scanning display. The wavefront of the projected light beam (38) may be bent or focused to match the desired focal length of the augmented and / or virtual reality object. A mini infrared camera (24) may be utilized to track the eye to assist with rendering and user input (i.e., where the user is looking, at what depth the user is focusing (the edge of the eye may be utilized to estimate focal depth, as discussed below)). A GPS (37), gyro, compass, and accelerometer (39) may be utilized to provide heading estimation and / or fast pose estimation. The camera's (16) images and pose, along with data from associated cloud computing resources, may be utilized to map the local world and share the user's view with the virtual or augmented reality community.While most of the hardware in the display system (14) featured in FIG. 14 is depicted as directly coupled to the display (82) and the housing (84) adjacent the user's eyes (20), the depicted hardware components may be mounted on or housed within other components, such as belt-mounted components, as shown in FIG. 3, for example. In one embodiment, all of the components of the system (14) featured in FIG. 14 are directly coupled to the display housing (84), with the exception of the image pose processor (36), the sensor pose processor (32), and the rendering engine (34), and communication between the image pose processor (36), the sensor pose processor (32), and the rendering engine (34) and the remaining components of the system (14) may be by wireless communication, such as ultra-wideband, or by wired communication. The depicted housing (84) is preferably head-mounted and wearable by a user. It may also feature speakers, such as those that can be inserted into the user's ears and utilized to provide the user with sounds that may be associated with an augmented reality or virtual reality experience, such as the jungle sounds referenced in reference to FIG. 13, and a microphone that can be utilized to capture sounds that are local to the user.
[0069] With regard to projecting light 38 into the user's eye 20, in one embodiment, the mini-camera 24 may be utilized to measure where the center of the user's eye 20 is geometrically tangent, which generally coincides with the location of the eye's 20 focal point, or "depth of focus." The three-dimensional surface of all points that the eye tangents to is called the "horopter." The focal distance may exhibit a finite number of depths, or may vary infinitely. Light projected from the vergence distance appears to be focused on the subject's eye 20, while light in front of or behind the vergence distance is blurred. Furthermore, it has been discovered that spatially coherent light with a beam diameter of less than about 0.7 millimeters is properly resolved by the human eye, regardless of where the eye focuses. With this understanding in mind, to create the illusion of proper depth of focus, eye convergence may be tracked using a mini camera (24), and the rendering engine (34) and projection subsystem (18) may be used to render all objects on or near the horopter in focus, and all other objects out of focus to various degrees (i.e., using intentionally created blur). Perspective light directing optical elements configured to project coherent light into the eye may be provided by a supplier such as Lumus, Inc. Preferably, the system (14) renders to the user at a frame rate of about 60 frames per second or greater. As explained above, preferably, the mini camera (24) may be utilized for eye tracking, and software may be configured to pick up not only convergence geometry, but also focal position cues that serve as user input. Preferably, such a system is configured with brightness and contrast suitable for daytime or nighttime use. In one embodiment, such a system preferably has a latency of less than about 20 milliseconds for visual object alignment, an angular alignment of less than about 0.1 degrees, and a resolution of about 1 arcminute, which is approximately the limit of the human eye.The display system (14) may be integrated with a localization system, which may include a GPS element, optical tracking, a compass, an accelerometer, and / or other data sources to assist in position and attitude determination, and the localization information may be utilized to facilitate accurate rendering within the user's field of view of the relevant world (i.e., such information facilitates the glasses knowing where they are relative to the real world).
[0070] Other suitable display devices include desktop and mobile computers, smartphones that may be augmented with additional software and hardware features that facilitate or simulate 3D perspective viewing (e.g., in one embodiment, a frame may be removably coupled to the smartphone, the frame featuring a 200 Hz gyro and accelerometer sensor subset, two small machine vision cameras with wide field of view lenses, and an ARM processor to simulate some of the functionality of the configuration featured in FIG. 14), tablet computers, tablet computers that may be augmented as described above for smartphones. These include, but are not limited to, computer vision devices, tablet computers augmented with additional processing and sensing hardware, head-mounted systems using smartphones and / or tablets to display augmented and virtual viewpoints (visual adaptation via magnifying optics, mirrors, contact lenses, or light structuring elements), non-see-through displays of light-emitting elements (LCD, OLED, vertical cavity surface emitting laser, directed laser beams, etc.), see-through displays that allow a human to simultaneously see the natural world and artificially generated images (e.g., light directing optics, clear and polarized OLEDs shining into close focus contact lenses, directed laser beams, etc.), contact lenses with light-emitting elements (such as those available from Innovega, Inc. (Bellevue, WA) under the trade name Ioptik RTM, which may be combined with specialized complementary eyeglass components), implantable devices with light-emitting elements, and implantable devices that simulate the photoreceptors of the human brain.
[0071] With a system such as those depicted in FIG. 3 and FIG. 14, 3D points may be captured from the environment and the pose (i.e., vector and / or origin position information relative to the world) of the camera capturing these images or points may be determined, so that these points or images may be "tagged" or associated with this pose information. The points captured by the second camera may then be utilized to determine the pose of the second camera. In other words, the second camera may be oriented and / or localized based on a comparison with the tagged image from the first camera. This knowledge may then be utilized to extract textures, create maps, and create a virtual copy of the real world (since there are two aligned cameras around). Thus, at a basic level, in one embodiment, we have a personally worn system that may be utilized to capture both the 3D points and the 2D images that generated the points, which may then be sent out to cloud storage and processing resources. They may also be cached locally (i.e., caching tagged images) with the pose information embedded, so the cloud may have tagged (i.e., tagged with 3D pose) 2D images ready (i.e., in available cache) along with the 3D points. If the user is observing something dynamic, the user may send additional information up to the cloud related to the movement (e.g., when looking at another individual's face, the user may take the texture map of the face and boost it to an optimized frequency, even if the surrounding world is otherwise essentially static).
[0072] The cloud system may be configured to store some points as pose-only fiducials to reduce the overall pose tracking calculations. In general, it may be desirable to have some contour features so that key items in the user's environment, such as walls, tables, etc., can be tracked as the user moves around a room and may want to "share" the world and let other users into the room and see these points. Such useful and key points may be referred to as "fiducials" since they are extremely useful as anchoring points. They relate to features that can be recognized with machine vision and can be consistently and repeatedly extracted from the world on different pieces of user hardware. Thus, these fiducials may preferably be stored in the cloud for further use.
[0073] In one embodiment, it is preferable to have a relatively even distribution of fiducials throughout the relevant universe, since fiducials are the kinds of items that a camera can easily use to recognize locations.
[0074] In one embodiment, the associated cloud computing arrangement may be configured to periodically trim the database of 3D points and any associated metadata to use the best data from various users for both refining the criteria and creating the world. In other words, the system may be configured to obtain the best data set by using input from various users who view and work within the associated world. In one embodiment, the database is fractal in nature, and as users get closer to an object, the cloud passes on higher resolution information to such users. As users map objects more closely, that data is sent to the cloud, which can add new 3D points and image-based texture maps to the database if they are better than those previously stored in the database. All of this may be configured to happen simultaneously from many users.
[0075] As described above, an augmented reality or virtual reality experience may be based on recognizing certain kinds of objects. For example, to recognize and understand a particular object, it may be important to understand that such an object has depth. A recognizer software object ("recognizer") may be deployed on a cloud or local resource to specifically assist in the recognition of various objects on either or both platforms as the user navigates the data in the world. For example, if a system has world model data with 3D point clouds and pose tagged images, and there is a desk with a number of points on it, as well as an image of the desk, there may be no determination that what is being observed is actually a desk when a human is grasping the desk. In other words, a few 3D points in space, and an image from somewhere in the space showing most of the desk, may not be enough to instantly recognize that a desk is being observed. To assist with this identification, a specific object recognizer may be created that goes into the raw 3D point cloud, segments a set of points, and extracts, for example, the plane of the top surface of the desk. Similarly, perceivers may be created to segment walls from 3D points so that a user can change wallpaper in virtual or augmented reality, or remove a portion of a wall, and have an entrance to another room that is not actually there in the real world. Such perceivers operate within the data of the world model, crawling the world model, and such perceivers may be thought of as software "robots" that seed the world model with semantic information, or an ontology of what is believed to exist between points in space. Such perceivers or software robots may be configured such that their entire existence is about crawling through the relevant world data and finding what is believed to be a wall, or a chair, or other item. They may be configured to tag a set of points with the functional equivalence "this set of points belongs to a wall" and may comprise a combination of point-based algorithms and pose tagging image analysis to manually inform the system as to what is in the points.
[0076] Object recognizers may be created for many purposes of various utility depending on the viewpoint. For example, in one embodiment, a specialty coffee shop such as Starbucks may invest in creating an accurate recognizer of Starbucks coffee cups within the relevant world of data. Such a recognizer may be configured to crawl the world of data, large and small, searching for Starbucks coffee cups, whereby Starbucks coffee cups may be segmented and identified to a user when operating within the relevant neighborhood space (i.e., perhaps to provide the user with coffee at a nearby Starbucks outlet when the user sees a Starbucks coffee cup for a certain period of time). Once the cup is segmented, it may be quickly recognized when the user moves it onto his or her desk. Such a recognizer may be configured to operate or operate on cloud computing resources and data as well as on local resources and data, or both on the cloud and locally, depending on available computational resources. In one embodiment, there is a global copy of the world model on the cloud with millions of users contributing to the global model, but for smaller worlds or sub-worlds, such as a particular person's office in a particular town, the system may be configured to trim and move data to locally cache information that is deemed most locally relevant to a given user, since most of the global world does not care what that office looks like. In one embodiment, relevant information (such as segments of a particular cup on a desk) may be configured to reside only on local computing resources, not on the cloud, since an object identified as moving often (e.g., a cup on a desk), for example, as a user walks up to the desk, does not need to burden the cloud model and incur transmission burdens between the cloud and local resources.Thus, cloud computing resources may be configured to segment 3D points and images, thus decomposing permanent (i.e., generally non-moving) objects from movable ones, which affects where the relevant data remains and where it is processed, removing the processing burden from the wearable / local system for certain data related to more permanent objects, allowing one-time processing of locations that can then be shared with an unlimited number of other users, allowing multiple data sources to simultaneously build a database of fixed and movable objects at a particular physical location, and may segment objects from the background to create object-specific fiducials and texture maps.
[0077] In one embodiment, the system may be configured to query the user for input regarding the identity of a particular object (e.g., the system may present the user with questions such as "Is that a Starbucks coffee cup?") so that the user can tailor the system and allow the system to associate semantic information with objects in the real world. The ontology may provide guidance as to what objects segmented from the world can do, how they behave, etc. In one embodiment, the system may feature a virtual or actual keypad, such as a wirelessly connected keypad, connectivity to a smartphone keypad, or the like, to facilitate specific user input into the system.
[0078] The system may be configured to share primitives (walls, windows, desk geometry, etc.) with any user who enters the room in virtual or augmented reality, and in one embodiment, that individual's system is configured to take images from specific viewpoints and upload them to the cloud, which can then be combined with the old and new sets of data to run optimization routines and establish criteria that exist on individual objects.
[0079] GPS and other localization information may be used as inputs to such processing. Additionally, other computing systems and data, such as a person's online calendar or Facebook account information, may be used as inputs (e.g., in one embodiment, the cloud and / or local system may be configured to analyze the contents of a user's calendar for flights, dates, and destinations, so that over time, information may be moved from the cloud to the user's local system so that it is ready for the user's arrival time at a given destination).
[0080] In one embodiment, tags such as QR codes and the like may be inserted into the world for use with non-statistical pose calculations, security / access control, communication of special information, spatial messaging, non-statistical object recognition, etc.
[0081] In one embodiment, cloud resources may be configured to pass digital models of real and virtual worlds between users, as described above with reference to "passable worlds," with the models rendered by individual users based on parameters and textures. This reduces bandwidth compared to passing real-time video, enables rendering of virtual viewpoints of the view, and allows millions or more users to participate in one virtual gathering without transmitting to each of the millions or more users the data (e.g., video) they need to see, because their view is rendered by local computing resources.
[0082] A virtual reality system ("VRS") may be configured to register user position and field of view (together known as "pose") through one or more of real-time metric computer vision using cameras, simultaneous localization and mapping techniques, maps, and data from sensors (e.g., gyros, accelerometers, compasses, barometers, GPS), radio signal strength triangulation, signal time-of-flight analysis, LIDAR ranging, RADAR ranging, odometry, and sonar ranging. A wearable device system may be configured to simultaneously map and orient. For example, in an unknown environment, a VRS may be configured to gather information about the environment and identify suitable reference points for user pose calculation, other points for world modeling, and images to provide a texture map of the world. The reference points may be used to optically calculate the pose. As the world is mapped in greater detail, more objects may be segmented and given their own texture maps, but the world can still preferably be represented at a low spatial resolution with simple polygons with low-resolution texture maps. Other sensors such as those discussed above may be utilized to assist in this modeling effort. The world may be fractal in nature, in that moving (through viewpoint, "surveillance" mode, zooming, etc.) or otherwise seeking a better view requires higher resolution information from cloud resources. Moving closer to an object captures higher resolution data, which may be transmitted to the cloud, which may compute new data and / or insert the new data into gaps in the world model.
[0083] With reference to FIG. 16, the wearable system may be configured to capture image information and extract fiducials and recognized points (52). The wearable local system may calculate the pose using one of the pose calculation techniques described below. The cloud (54) may be configured to use the images and fiducials to segment the 3D object from a more static 3D background, the images providing a texture map of the object and the world (the texture may be real-time video). The cloud resources (56) may be configured to store and make available the static fiducials and textures for world registration. The cloud resources may be configured to trim the point cloud for optimal point density for registration. The cloud resources (60) may be configured to store and make available the object fiducials and textures for object registration and manipulation, the cloud may trim the point cloud for optimal density for registration. The cloud resources may be configured to use all valid points and textures to generate a fractal solid model of the object (62), the cloud may trim the point cloud information for optimal fiducial density. The cloud resources (64) may be configured to query the user for tailoring regarding the identities of the segmented objects and worlds, and the ontology database may use the answers to populate the objects and worlds with actionable properties.
[0084] The following specific alignment and mapping modes feature "O attitude", which represents attitude determined from an optical or camera system, "s attitude", which represents attitude determined from sensors (i.e., a combination of GPS, gyro, compass, accelerometer, etc. data, as discussed above), and "MLC", which represents cloud computing and data management resources.
[0085] 1. Orientation: Creating a base map of the new environment Purpose: To establish attitude (similar) when the environment is not mapped or not connected to an MLC. Extract points from images, track them from frame to frame, and triangulate fiducials using the S-pose. · Since there is no standard, use the S position. · Eliminate poor criteria based on persistence. This is the most basic mode. It always operates on a low precision pose. With a small amount of time and some relative motion, it establishes the O pose and / or a minimum reference set for mapping. -As soon as you feel comfortable with the position, exit this mode.
[0086] 2. Map and O-pose: Map the environment Goal: Establish high accuracy attitude, map the environment, and provide the map (with images) to the MLC. ·Compute O pose from maturity world reference. Use S pose as a check on O pose solution and to accelerate computation (O pose is a non-linear gradient search). ·Maturity reference may be derived from MLC or locally determined. ·Extract points from image, track from frame to frame and triangulate reference using O pose. · Eliminate poor criteria based on persistence. · Provides reference and attitude tagged imagery to the MLC. The last three steps do not have to happen in real time.
[0087] 3. Posture: Determine your posture Goal: Establish high accuracy pose within an already mapped environment, using minimal processing power. Use the previous S and O poses (n-1, n-2, n-3, etc.) to estimate the pose at n. Use the pose at n to project the fiducials onto the image captured at n, then create an image mask from the projection. Extract points from masked regions (by searching / extracting points only from a masked subset of the image, the processing load is greatly reduced). Calculate O-pose from extracted points and mature global reference. · To estimate the pose at n+1, use the S and O poses at n. Optional: Provide posture tagged images / videos to the MLC Cloud.
[0088] 4. Super-resolution: Determine the super-resolution image and criteria Objective: To create super-resolution images and standards. · Synthesize pose-tagged images to create a super-resolution image. Use super-resolution images to enhance reference position estimation. · Iterative pose estimation from super-resolution reference and images. Optional: Loop the above steps on a wearable device (in real time) or on the MLC (for a better world).
[0089] In one embodiment, a VLS system may be configured to have certain base functionality and functionality facilitated by "apps" or applications that may be delivered through the VLS to provide certain specialized functionality. For example, the following apps may be installed on a target VLS to provide specialized functionality:
[0090] A painterly rendering app. Artists create image transforms that represent the world as they see it. Users activate these transforms and thus see the world "through" the artist's eyes.
[0091] Tabletop modeling apps, where the user "builds" an object from physical objects placed on a table.
[0092] A virtual presence app, where users hand off a virtual model of a space to other users who then move around the space using a virtual avatar.
[0093] Avatar emotion apps. Subtle voice inflections, slight head movements, body temperature, heart rate, and other measurements animate subtle effects on virtual presence avatars. Digitizing human state information and passing it to a remote avatar uses less bandwidth than video. In addition, such data can be mapped to sentient, non-human avatars. For example, a dog avatar can indicate excitement by wagging its tail based on an excited voice inflection.
[0094] An efficient mesh network may be desirable for moving data as opposed to sending everything back to a server. However, many mesh networks have suboptimal performance because the location information and topology are not well characterized. In one embodiment, the system may be utilized to determine the location of all users with a relatively high degree of accuracy, and thus a mesh network configuration may be utilized for high performance.
[0095] In one embodiment, the system may be utilized for search. With augmented reality, for example, users generate and leave content related to many aspects of the physical world. Most of this content is not text and therefore not easily searchable by typical methods. The system may be configured to provide a facility for keeping track of personal and social network content for search and reference purposes.
[0096] In one embodiment, if a display device tracks 2D points through successive frames and then fits a vector-valued function to the time evolution of these points, it is possible to sample the vector-valued function at any point in time (e.g., between frames) or at some point in the near future (by projecting the vector-valued function forward). This allows for the creation of high-resolution post-processing, and prediction of future poses before the next image is actually captured (e.g., it is possible to double the alignment speed without doubling the camera frame rate).
[0097] For body-fixed rendering (as opposed to head-fixed or world-fixed rendering), an accurate representation of the body is desirable. In one embodiment, rather than measuring the body, it is possible to derive its location through the average position of the user's head. If the user's face faces forward most of the time, a multi-day average of head position will reveal that orientation. Together with the gravity vector, this provides a reasonably stable coordinate system for body-fixed rendering. Using a current measure of head position relative to this long-term coordinate system allows for consistent rendering of objects on / around the user's body without extra instrumentation. For implementation of this embodiment, a single registration average of the head orientation vector may be started, and the cumulative sum of the data divided by delta t gives the current average head position. Keeping approximately five registrations starting on day n-5, days n-4, days n-3, days n-2, days n-1 allows for the use of a rolling average of only the past "n" days.
[0098] In one embodiment, a view may be scaled down and presented to a user in a space smaller than it actually is. For example, in a situation where a view must be rendered in a huge space (i.e., a soccer stadium, etc.), an equivalent huge space may not exist or such a large space may be inconvenient for the user. In one embodiment, the system may be configured to scale down the view so that a user may observe a scaled down view. For example, an individual may play a video game from a God's eye view, or a world championship soccer match, with the playing field unscaled or scaled and presented on a living room floor. The system may be configured to simply shift the viewpoint, the scale, and the associated adaptive distance.
[0099] The system may also be configured to draw the user's attention to particular items within the presented view by manipulating the focus of virtual reality or augmented reality objects, highlighting them, changing their contrast, brightness, scale, etc.
[0100] Preferably, the system may be configured to achieve the following modes:
[0101] Open Space Rendering: Capture key points from a structured environment and then use ML rendering to fill in the spaces in between. · Potential Venues: stages, output spaces, large indoor spaces (stadiums).
[0102] Object Wrapping: Recognize 3D objects in the real world and then augment them. "Recognition" here means identifying the 3D blob with enough accuracy to connect the image. · There are two kinds of recognition: 1) classifying a type of object (e.g., "face"), and 2) classifying a specific instance of an object (e.g., Joe, an individual). Build recognizer software objects for a variety of things, such as walls, ceilings, floors, faces, roads, skies, skyscrapers, ranch houses, tables, chairs, cars, road signs, billboards, doors, windows, bookshelves, etc. Some recognizers are I-type and have general functionality, e.g. "put my video on that wall" or "that's a dog". Other recognizers are Type II and have specific functionality, e.g., "My TV is on the living room wall 3.2 feet from the ceiling," "That's Fido," etc. (This is a more capable version of the general recognizer). Building perceptors as software objects allows for quantified release of functionality and finer-grained control of the experience.
[0103] Body-centered rendering Rendering virtual objects fixed to the user's body. Some things, like a digital tool belt, should float around the user's body. This requires knowing where the body is, not just the head. By taking a long-term average of the user's head position (which is usually facing forward, parallel to the ground), we may get a reasonably accurate position. A simple example would be objects floating around your head.
[0104] Transparency / Breakout View For type II recognition objects, cutaway views are shown. Linking type II recognized objects to an online database of 3D models. You should start with objects that have commonly available 3D models, such as vehicles and utilities.
[0105] Virtual Being -Draw avatars of remote people into open space. ○ A subset of "Free Space Rendering" (above). ○ A user creates the rough geometry of a local environment and iteratively sends both the geometry and texture maps to others. ○Users must give permission for others to enter their environment. Subtle vocal cues, hand tracking, and head movements are transmitted to a remote avatar, which is animated from these fuzzy inputs. ○The above minimizes bandwidth. Create a "doorway" in the wall to another room ○ Pass in geometry and texture maps as you would with the other methods. Instead of showing avatars in a local room, designate recognized objects (e.g. walls) as portals to other people's environments. This way multiple people can sit in their own rooms and see other people's environments "through" the walls.
[0106] Virtual Viewpoint As a group of cameras (people) view the scene from different perspectives, a dense digital model of the area is created. This rich digital model can be rendered from any vantage point that at least one of the cameras can see. Example: People at a wedding. The scene is modeled by all attendees jointly. The perceiver distinguishes stationary objects from moving objects and creates a texture map (e.g. walls have a stable texture map, people have a higher frequency moving texture map). With a rich digital model updated in real time, the view can be rendered from any vantage point. Rear attendees can even fly up to the front rows for a better view. · Attendees can show their moving avatar or hide their viewpoint. · Off-premise attendees can find their "seat" using their avatar or invisibly if the host allows. Likely to require extremely high bandwidth. Conceptually, high frequency data is streamed to the crowd over high speed local radio. Low frequency data comes from MLC. Since all attendees have highly accurate location information, creating optimal routing paths for local networking is trivial.
[0107] Messaging Simple silent messaging may be desirable. For this and other applications, it may be desirable to have a finger coding keyboard. Tactile grab solutions may provide improved performance.
[0108] Full Virtual Reality (VR): When the Vision System goes dark, it will show a view that doesn't overlap with the real world. A registration system is still needed to track head position. "Couch Mode" allows the user to fly. "Walk Mode" re-renders real-world objects as virtual ones to prevent users from colliding with them in the real world. Rendering body parts is essential to believability in fiction. This suggests having a method for tracking and rendering body parts in the field of view. Non-see-through visors are a form of VR that have many image quality enhancement benefits not possible with direct overlay. · Wide field of vision, perhaps even the ability to see behind you. -Various forms of "super vision": telescopic, clairvoyant, infrared, god's eye view, etc.
[0109] In one embodiment, the system for virtual and / or augmented user experiences is configured such that a remote avatar associated with a user may be animated based at least in part on data on the wearable device with input from sources such as voice intonation analysis and facial recognition analysis as performed by associated software modules. For example, referring again to FIG. 12, the bee avatar (2) may be animated to smile in a friendly manner based on facial recognition of a smile on the user's face, or based on a friendly tone of voice or tone as determined by software configured to analyze audio input into a microphone that may capture audio samples locally from the user. Additionally, the avatar character may be animated in a manner that would cause the avatar to express a particular emotion. For example, in an embodiment in which the avatar is a dog, a happy smile or tone detected by a system local to the human user may be expressed by the avatar as the dog avatar wagging its tail.
[0110] Various exemplary embodiments of the present invention are described herein. Reference is made to these examples in a non-limiting sense. They are provided to illustrate the broader and more applicable aspects of the present invention. Various changes may be made to the described invention, and similar may be substituted without departing from the true spirit and scope of the present invention. In addition, many modifications may be made to adapt a particular situation, material, composition of matter, process, process act, or step to the objective, spirit, or scope of the present invention. Moreover, as will be understood by those skilled in the art, each of the individual variations described and illustrated herein has distinct components and features that can be easily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. All such modifications are intended to be within the scope of the claims associated with this disclosure.
[0111] The present invention includes methods that may be performed using the subject devices. The methods may include the act of providing such a suitable device. Such providing may be performed by an end user. In other words, the act of "providing" merely requires the end user to obtain, access, access, position, configure, activate, power on, or otherwise act to provide the requisite device in the subject method. The methods described herein may be carried out in any order of the described events that is logically possible, not just the recited order of events.
[0112] Exemplary aspects of the invention have been described above, along with details regarding material selection and manufacture. As for other details of the invention, these will be appreciated in conjunction with the above-referenced patents and publications, and will generally be known or understood by those skilled in the art. The same may be true with respect to method-based aspects of the invention in terms of additional acts as typically or logically adopted.
[0113] In addition, although the present invention has been described with reference to several embodiments incorporating various features as appropriate, the present invention is not limited to what has been described and indicated, as envisioned with respect to each variation of the present invention. Various modifications may be made to the invention described, and equivalents may be substituted (whether described herein or not included for some simplification) without departing from the true spirit and scope of the invention. In addition, when a range of values is provided, it is to be understood that all intervening values between the upper and lower limits of that range, and any other stated or intervening values within that stated range, are encompassed within the present invention.
[0114] It is also contemplated that any optional features of the described inventive variations may be described and claimed independently or in combination with any one or more of the features described herein. Reference to a singular item includes the possibility that there are multiple identical items. More specifically, as used herein and in the claims associated herewith, the singular forms "a," "an," "said," and "the" include plural referents unless otherwise specified. In other words, the use of articles allows for "at least one" of the subject items in the above description as well as in the claims associated herewith. Furthermore, it is noted that such claims may be drafted to exclude any optional element. Thus, this statement is intended to serve as a predicate for the use of exclusive terms such as "solely," "only," and the like, or the use of "negative" limitations in connection with the recitation of claim elements.
[0115] Without using such exclusive terms, the term "comprising" in the claims associated with this disclosure shall be construed as allowing for the inclusion of any additional elements, regardless of whether a given number of elements are recited in such claims or whether the addition of features can be considered as a transformation of the nature of the elements recited in such claims. Except as specifically defined herein, all technical and scientific terms used herein are to be given the broadest possible commonly understood meaning while maintaining the validity of the claims.
[0116] The breadth of the present invention is not to be limited by the examples and / or subject specification provided, but rather only by the scope of the claims language associated with this disclosure.
Claims
1. 1. A system for interacting with a virtual world including virtual world data, the system comprising: A first user device operably coupled to a computer network comprising one or more computing devices, the one or more computing devices comprising: one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to process a first portion of the virtual world data; A first user device comprising: The computer network comprises: receiving a first input from a first user via an eye sensor of the first user device, the first input being indicative of a visual focal distance of the first user; determining the visual focal length of the first user based on the first input; receiving, via the first user device, a second input from a local environment of the first user device; wherein one or more of the one or more computing devices are configured to generate modified virtual world data by modifying the virtual world data based on at least one of the first input and the second input; the first user device is further configured to present virtual content to the first user based on the modified virtual world data; presenting the virtual content to the first user includes presenting a visual rendering of the virtual world in a 3D format via projected light; the projected light includes a first light projected as if at a first distance corresponding to the visual focal length of the first user, the first light corresponding to a focused element of the virtual content; The system further includes a second light projected as if at a second distance different from the visual focal length of the first user, the second light corresponding to an unfocused element of the virtual content, and the first light including spatially coherent light of a light directing optical element, the spatially coherent light being presented to the eye of the first user such that the spatially coherent light is resolved by the eye of the first user regardless of the focal position of the first user.
2. 2. The system of claim 1, wherein at least one of the first input and the second input includes audio data, and modifying the virtual world data includes modifying the virtual world data based on the audio data.
3. 2. The system of claim 1 , wherein the first user device comprises a see-through display, and presenting the virtual content to the first user further comprises presenting the virtual content to the first user via the see-through display.
4. The system of claim 1 , wherein the virtual content includes one or more of visual content, audio content, and haptic content.
5. The system of claim 1 , wherein the first input comprises optical property data.
6. The system of claim 1 , wherein the second input comprises one or more of audio data from the local environment, visual data from physical objects in the local environment.
7. The system of claim 1 , further comprising a second user device configured to present virtual content to a second user based on the modified virtual world data.
8. The system of claim 1 , wherein the first user device includes a wearable device having a head-mounted see-through display.
9. The system of claim 8 , wherein the wearable device is configured to communicate with the computer network via a mobile phone.
10. The system of claim 8 , wherein the wearable device is configured to communicate with the computer network via Wi-Fi.
11. The system of claim 8 , wherein the wearable device is configured to communicate with the computer network via Bluetooth.
12. The system of claim 1 , wherein the first user device comprises a mobile phone.
13. The system of claim 12 , wherein the mobile phone comprises an input device, and the first input is received from the first user via the input device.
14. The system of claim 12 , wherein the mobile phone includes a camera, and the second input is received via the camera.
15. The system of claim 12 , wherein the mobile phone comprises a display, and presenting the virtual content to the first user comprises overlaying the virtual content on the display.
16. The system of claim 1 , wherein the first user device is configured to communicate with the computer network via a mobile phone.
17. The system of claim 1 , wherein the first distance corresponds to a distance to a horopter associated with the first user.
18. The system of claim 1 , wherein the first light comprises spatially coherent light having a beam diameter of less than 0.7 millimeters.
19. the first user device includes a head-wearable device; the eye sensor includes a camera of the head wearable device; the head-wearable device comprises a display configured to present the visual rendering of the virtual world in 3D format; 2. The system of claim 1, wherein the computer network is configured to determine the visual focal length based on a convergence distance of the first user's eyes, the convergence distance being determined based on an output of the camera.
20. 20. The system of claim 19, wherein the computer network is configured to track the convergence distance over time.
21. 2. The system of claim 1, wherein the first light comprises a first wavefront that is bent according to the first distance and the second light comprises a second wavefront that is bent according to the second distance.
22. The system of claim 1 , wherein the focused element comprises an augmented reality object.
23. The system of claim 1 , wherein the focused element comprises a virtual reality object.
24. The system of claim 1 , wherein the projected light comprises light projected via a scanned laser.
Citation Information
Patent Citations
Integrated wearable computer provided with image input interface to utilize body
JP2001312356A
Frequency tunable resonant scanner
US6331909B1
Holographic image display systems
WO2009156752A1
Systems and methods for interaction with a virtual environment
WO2011041466A1
Local advertising content on an interactive head-mounted eyepiece
WO2011106798A1