Cross-reality system
By utilizing persistent coordinate frames to manage spatial information, the system addresses computational inefficiencies in XR systems, enabling efficient and immersive cross-reality experiences for multiple users.
Patent Information
- Application Number
- JP2024087120
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-07
- Filing Date
- 2024-05-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-10-04
AI Technical Summary
Existing XR systems face challenges in efficiently creating and maintaining persistent spatial information for multiple users, leading to computational inefficiencies and suboptimal user experiences due to varying orientations and session-based spatial information collection.
The system employs a technique for generating and maintaining persistent spatial information using persistent coordinate frames (PCFs) that are shared among devices, allowing for efficient map creation, update, and rendering of virtual content across multiple user sessions with reduced computational overhead.
This approach enables a more immersive and computationally efficient XR experience by allowing devices to quickly restore and reset head poses, share spatial information, and render virtual content accurately across different user sessions and orientations.
Smart Images

Figure 0007711263000001 
Figure 0007711263000002 
Figure 0007711263000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This patent application claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 742,237, filed on October 5, 2018, entitled "COORDINATE FRAME PROCESSING AUGMENTED REALITY", which is incorporated herein by reference in its entirety. This patent application also claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 812,935, filed on March 1, 2019, entitled "MERGING A PLURALITY OF INDIVIDUALLY MAPPED ENVIRONMENTS", which is incorporated herein by reference in its entirety. This patent application also claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 815,955, filed on March 8, 2019, entitled "VIEWING DEVICE OR VIEWING DEVICES HAVING ONE OR MORE COORDINATE FRAME TRANSFORMERS", which is incorporated herein by reference in its entirety. This patent application also claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 868,786, filed on June 28, 2019, entitled "RANKING AND MERGING A PLURALITY OF ENVIRONMENT MAPS", which is incorporated herein by reference in its entirety. This patent application also claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 870,954, filed on July 5, 2019, entitled "RANKING AND MERGING A PLURALITY OF ENVIRONMENT MAPS", which is incorporated herein by reference in its entirety. This patent application also claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 884,109, filed on August 7, 2019, entitled "A VIEWING SYSTEM", which is incorporated herein by reference in its entirety.
[0002] This application generally relates to cross-reality systems.
Background Art
[0003] A computer can control a human user interface and create an X Reality (XR or cross-reality) environment where, as perceived by the user, part or all of the XR environment is generated by the computer. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments where part or all of the XR environment can be generated by the computer using data that describes the environment in part. This data can describe virtual objects, for example, that can be rendered so that the user can perceive or sense them as part of the physical world and interact with the virtual objects. The user can experience these virtual objects as a result of data being rendered and presented through a user interface device such as a head-mounted display device. The data can control audio that can be displayed to be visible to the user or reproduced to be audible to the user, or control a haptic (or tactile) interface to enable the user to experience a touch sensation that the user perceives or senses as the user feels the virtual object.
[0004] XR systems can be useful for many applications spanning the fields of scientific visualization, medical training, engineering design, and prototyping, teleoperation and telepresence, and personal entertainment. AR and MR, in contrast to VR, involve one or more objects in relation to real objects in the physical world. The experience of virtual objects interacting with real objects generally enhances the user's enjoyment when using XR systems and also expands the possibilities for various applications by presenting realistic and easily understandable information about how the physical world can be modified.
[0005] To realistically render virtual content, an XR system may construct a representation of the physical world around a user of the system. This representation may be constructed, for example, by processed images obtained using sensors on wearable devices that form part of the XR system. In such a system, a user may perform an initialization routine by looking around the room or other physical environment in which the user intends to use the XR system until the system has obtained sufficient information to construct a representation of that environment. As the system operates and the user moves around the environment or to other environments, sensors on the wearable device may obtain additional information and extend or update the representation of the physical world.
SUMMARY OF THE INVENTION
MEANS FOR SOLVING THE PROBLEM
[0006] Aspects of the present application relate to methods and apparatuses for providing an X Reality (Cross Reality or XR) scenario. The techniques described herein may be used together, separately, or in any suitable combination.
[0007] Some embodiments relate to an electronic system including one or more sensors configured to capture information about a three-dimensional (3D) environment. The captured information includes a plurality of images. The electronic system includes at least one processor configured to execute computer-executable instructions to generate a map of at least a portion of the 3D environment based on the plurality of images. The computer-executable instructions further include instructions for identifying a plurality of features in the plurality of images, selecting a plurality of keyframes from the plurality of images based at least in part on the plurality of features of the selected keyframes, generating one or more coordinate frames based at least in part on the identified features of the selected keyframes, and storing the one or more coordinate frames as one or more persistent coordinate frames in association with the map of the 3D environment.
[0008] In some embodiments, one or more sensors comprise a plurality of pixel circuits arranged in a two-dimensional array such that each image of the plurality of images comprises a plurality of pixels. Each feature corresponds to a plurality of pixels.
[0009] In some embodiments, the step of identifying a plurality of features in a plurality of images includes selecting, as the identified features, a number of groups of pixels that is less than a predetermined maximum value based on a measure of similarity to a group of pixels depicting a portion of a persistent object.
[0010] In some embodiments, the step of storing one or more coordinate frames includes, for each of the one or more coordinate frames, storing a descriptor representing at least a subset of the features in a selected key frame from which the coordinate frame was generated.
[0011] In some embodiments, the step of storing one or more coordinate frames includes, for each of the one or more coordinate frames, storing at least a subset of the features in a selected key frame from which the coordinate frame was generated.
[0012] In some embodiments, the step of storing one or more coordinate frames includes, for each of the one or more coordinate frames, storing a transformation between a coordinate frame of a map of the 3D environment and a persistent coordinate frame, and geographic information indicating a location within the 3D environment of the selected key frame from which the coordinate frame was generated.
[0013] In some embodiments, the geographic information comprises a WiFi fingerprint of the location.
[0014] In some embodiments, the computer-executable instructions comprise instructions for calculating feature descriptors for individual features using an artificial neural network.
[0015] In some embodiments, the first artificial neural network is the first artificial neural network. The computer-executable instructions comprise instructions for implementing a second artificial neural network configured to calculate a frame descriptor for representing a keyframe, at least in part, based on a calculated feature descriptor for identified features within the keyframe.
[0016] In some embodiments, the computer-executable instructions further comprise an application programming interface configured to provide information characterizing the persistent coordinate frame of one or more persistent coordinate frames to an application running on a portable electronic system, instructions for refining a map of a 3D environment based on a second plurality of images, instructions for adjusting one or more of the persistent coordinate frames at least in part based on the second plurality of images, and instructions for providing a notification of a persistent coordinate frame adjusted through the application programming interface.
[0017] In some embodiments, the step of adjusting one or more persistent coordinate frames includes the step of adjusting the translation and rotation of one or more persistent coordinate frames relative to the origin of the map of the 3D environment.
[0018] In some embodiments, the electronic system comprises a wearable device, and one or more sensors are mounted on the wearable device. The map is a tracking map calculated on the wearable device. The origin of the map is determined based on the location where the device was powered on.
[0019] In some embodiments, the electronic system comprises a wearable device, and one or more sensors are mounted on the wearable device. The computer-executable instructions further include instructions for tracking the movement of the portable device and for controlling the timing of execution of instructions for generating one or more coordinate frames and / or storing one or more persistent coordinate frames based on the tracked movement indicating movement of the wearable device beyond a threshold distance, the threshold distance being between 2 and 20 meters.
[0020] Some embodiments relate to a method of operating an electronic system to render virtual content in a 3D environment comprising a portable device. The method includes using one or more processors to maintain a coordinate frame local to the portable device on the portable device based on the output of one or more sensors on the portable device, obtaining a stored coordinate frame from stored spatial information about the 3D environment, calculating a transformation between the coordinate frame local to the portable device and the obtained stored coordinate frame, receiving a specification of a virtual object having a coordinate frame local to the virtual object and a location of the virtual object relative to a selected stored coordinate frame, and rendering the virtual object on a display of the portable device at a determined location based at least in part on the calculated transformation and the received location of the virtual object.
[0021] In some embodiments, obtaining the stored coordinate frame includes obtaining the coordinate frame through an application programming interface (API).
[0022] In some embodiments, the portable device comprises a first portable device comprising a first processor of one or more processors. The system further comprises a second portable device comprising a second processor of one or more processors. The processors on each of the first and second devices obtain the same stored coordinate frame, calculate the transformation between the coordinate frame local to the individual device and the same stored coordinate frame obtained, receive the specification of the virtual object, and render the virtual object on the individual display.
[0023] In some embodiments, each of the first and second devices comprises a camera configured to output a plurality of camera images, a keyframe generator configured to convert the plurality of camera images into a plurality of keyframes, a persistent pose computer configured to generate a persistent pose by averaging the plurality of keyframes, a tracking map and persistent pose converter configured to convert the tracking map into the persistent pose and determine the persistent pose relative to the origin of the tracking map, a persistent pose and persistent coordinate frame (PCF) converter configured to convert the persistent pose into a PCF, and a map publisher configured to transmit spatial information including the PCF to a server.
[0024] In some embodiments, the method further comprises executing an application and generating the location of the virtual object relative to the specification of the virtual object and the selected stored coordinate frame.
[0025] In some embodiments, on a portable device, the step of maintaining a coordinate frame local to the portable device includes, for each of the first and second portable devices, capturing a plurality of images of the 3D environment from one or more sensors of the portable device, calculating one or more persistent poses at least in part based on the plurality of images, and generating spatial information about the 3D environment at least in part based on the calculated one or more persistent poses. The method further includes, for each of the first and second portable devices, transmitting the generated spatial information to a remote server, and the step of obtaining the stored coordinate frame includes receiving the stored coordinate frame from the remote server.
[0026] In some embodiments, the step of calculating one or more persistent poses at least in part based on the plurality of images includes extracting one or more features from each of the plurality of images, generating a descriptor for each of the one or more features, generating a keyframe for each of the plurality of images at least in part based on the descriptor, and generating one or more persistent poses at least in part based on the one or more keyframes.
[0027] In some embodiments, the step of generating one or more persistent poses includes selectively generating a persistent pose based on a portable device that travels a predetermined distance from the location of another persistent pose.
[0028] In some embodiments, each of the first and second devices comprises a download system configured to download a stored coordinate frame from a server.
[0029] Some embodiments relate to an electronic system for maintaining persistent spatial information about a 3D environment in order to render virtual content on each of a plurality of portable devices. The electronic system includes networked computing devices. The networked computing devices include at least one processor, at least one storage device connected to the processor, and a map storage routine executable using the at least one processor to receive a plurality of maps from portable devices of the plurality of portable devices and store map information on the at least one storage device, wherein each of the plurality of received maps comprises at least one coordinate frame, and a map transmitter executable using the at least one processor to receive location information from portable devices of the plurality of portable devices, select one or more maps from the stored maps, transmit information from the selected one or more maps to portable devices of the plurality of portable devices, and wherein the transmitted information comprises the coordinate frame of the map of the selected one or more maps.
[0030] In some embodiments, the coordinate frame comprises a computer data structure. The computer data structure comprises a coordinate frame comprising information characterizing a plurality of features of an object within the 3D environment.
[0031] In some embodiments, the information characterizing the plurality of features comprises descriptors characterizing regions of the 3D environment.
[0032] In some embodiments, each coordinate frame of the at least one coordinate frame comprises a persistent point characterized by a feature detected in sensor data representing the 3D environment.
[0033] In some embodiments, each coordinate frame of the at least one coordinate frame comprises a persistent pose.
[0034] In some embodiments, each coordinate frame of at least one coordinate frame comprises a persistent coordinate frame.
[0035] The foregoing description is provided by way of illustration and not by way of limitation. The present invention provides, for example, the following. (Item 1) An electronic system, One or more sensors configured to capture information about a three-dimensional (3D) environment, wherein the captured information comprises a plurality of images, and the sensors, At least one processor configured to execute computer-executable instructions and generate a map of at least a portion of the 3D environment based on the plurality of images, the computer-executable instructions further comprising Identifying a plurality of features within the plurality of images; Selecting a plurality of keyframes from the plurality of images based at least in part on the plurality of features of the selected keyframes; Generating one or more coordinate frames based at least in part on the identified features of the selected keyframes; Storing the one or more coordinate frames as one or more persistent coordinate frames in association with the map of the 3D environment And instructions for performing the operations, at least one processor An electronic system comprising. (Item 2) The one or more sensors comprise a plurality of pixel circuits arranged in a two-dimensional array such that each image of the plurality of images comprises a plurality of pixels, Each feature corresponds to a plurality of pixels, The electronic system according to item 1. (Item 3) Identifying a plurality of features in the plurality of images includes selecting, as the identified features, a number less than a predetermined maximum value of a group of pixels that depict a portion of a persistent object, based on a measure of similarity to the group of pixels, for the electronic system of claim 1. (Item 4) Storing the one or more coordinate frames includes, for each of the one or more coordinate frames, a descriptor representing at least a subset of the features in a selected key frame from which the coordinate frame was generated for the electronic system of claim 1. (Item 5) Storing the one or more coordinate frames includes, for each of the one or more coordinate frames, at least a subset of the features in a selected key frame from which the coordinate frame was generated for the electronic system of claim 1. (Item 6) Storing the one or more coordinate frames includes, for each of the one or more coordinate frames, a transformation between the coordinate frame of the map of the 3D environment and the persistent coordinate frame, and geographic information indicating a location within the 3D environment of the selected key frame from which the coordinate frame was generated for the electronic system of claim 1. (Item 7) The geographic information comprises a WiFi fingerprint of the location, for the electronic system of claim 6. (Item 8) The computer-executable instructions comprise instructions for calculating feature descriptors for individual features using an artificial neural network, for the electronic system of claim 1. (Item 9) The first artificial neural network is a first artificial neural network, The computer-executable instructions comprise instructions for implementing a second artificial neural network configured to calculate a frame descriptor for representing a keyframe, at least in part, based on the calculated feature descriptors for the identified features within the keyframe. The electronic system according to item 8. (Item 10) The computer-executable instructions further An application programming interface configured to provide information characterizing the persistent coordinate frames of the one or more persistent coordinate frames to an application running on a portable electronic system. Instructions for refining a map of the 3D environment based on a second plurality of images. At least in part, adjusting one or more of the persistent coordinate frames based on the second plurality of images. Instructions for providing a notification of the adjusted persistent coordinate frames through the application programming interface. The electronic system according to item 1, comprising (Item 11) Adjusting the one or more persistent coordinate frames includes adjusting a translation and a rotation of the one or more persistent coordinate frames relative to an origin of a map of the 3D environment. The electronic system according to item 10. (Item 12) The electronic system comprises a wearable device, and the one or more sensors are mounted on the wearable device. The map is a tracking map calculated on the wearable device. An origin of the map is determined based on a location where the device is powered on. The electronic system according to item 11. (Item 13) The electronic system includes a wearable device, and the one or more sensors are mounted on the wearable device. The computer-executable instructions further track the movement of the portable device, control the timing of execution of instructions for generating one or more coordinate frames and / or storing one or more persistent coordinate frames based on the tracked movement indicating movement of the wearable device beyond a threshold distance, where the threshold distance is between 2 and 20 meters, The electronic system according to item 1, comprising instructions for performing the above. (Item 14) A method of operating an electronic system to render virtual content in a 3D environment comprising a portable device, the method using one or more processors, maintain a coordinate frame local to the portable device on the portable device based on the output of one or more sensors on the portable device, obtain a stored coordinate frame from stored spatial information about the 3D environment, calculate a transformation between the coordinate frame local to the portable device and the obtained stored coordinate frame, receive a specification of a virtual object having a coordinate frame local to the virtual object and a location of the virtual object relative to the selected stored coordinate frame, render the virtual object on a display of the portable device at a determined location based at least in part on the calculated transformation and the received location of the virtual object, The method includes the above steps. (Item 15) The method according to item 14, wherein obtaining the stored coordinate frame includes obtaining the coordinate frame through an application programming interface (API). (Item 16) The portable device includes a first portable device including a first processor of the one or more processors, the system further includes a second portable device including a second processor of the one or more processors, the processors on each of the first and second devices obtain the same stored coordinate frame, calculate a conversion between the local coordinate frame of an individual device and the obtained same stored coordinate frame, receive the specification of the virtual object, render the virtual object on an individual display, and perform the method according to item 14. (Item 17) Each of the first and second devices includes a camera configured to output a plurality of camera images, a keyframe generator configured to convert the plurality of camera images into a plurality of keyframes, a persistent pose calculator configured to generate a persistent pose by averaging the plurality of keyframes, a tracking map and persistent pose converter configured to convert a tracking map into the persistent pose and determine the persistent pose with respect to the origin of the tracking map, a persistent pose and persistent coordinate frame (PCF) converter configured to convert the persistent pose into a PCF, a map publisher configured to transmit spatial information including the PCF to a server, and perform the method according to item 16. (Item 18) The method according to item 16, further including executing an application and generating a location of the virtual object with respect to the specification of the virtual object and the selected stored coordinate frame. (Item 19) Maintaining a coordinate frame local to the portable device on the portable device is, for each of the first and second portable devices, capturing a plurality of images of the 3D environment from one or more sensors of the portable device; calculating one or more persistent poses, at least in part, based on the plurality of images; generating spatial information about the 3D environment, at least in part, based on the calculated one or more persistent poses and includes The method further includes transmitting the generated spatial information to a remote server for each of the first and second portable devices, Obtaining the stored coordinate frame includes receiving the stored coordinate frame from the remote server. The method according to item 16. (Item 20) Calculating the one or more persistent poses, at least in part, based on the plurality of images includes extracting one or more features from each of the plurality of images; generating a descriptor for each of the one or more features; generating a key frame for each of the plurality of images, at least in part, based on the descriptors; generating the one or more persistent poses, at least in part, based on the one or more key frames and includes the method according to item 19. (Item 21) Generating the one or more persistent poses includes selectively generating a persistent pose based on the portable device that travels a predetermined distance from the location of another persistent pose and includes the method according to item 20. (Item 22) Each of the first and second devices a download system configured to download the stored coordinate frame from a server The method according to item 16, comprising (Item 23) An electronic system for maintaining persistent spatial information about a 3D environment in order to render virtual content on each of a plurality of portable devices, the electronic system comprising: A networked computing device, At least one processor, At least one storage device connected to the processor, A map storage routine that is executable using the at least one processor to receive a plurality of maps from a portable device of the plurality of portable devices and store map information on the at least one storage device, each of the plurality of received maps comprising at least one coordinate frame; A map transmitter that: Receives location information from a portable device of the plurality of portable devices; Selects one or more maps from the stored maps; Transmits information from the selected one or more maps to a portable device of the plurality of portable devices, the transmitted information comprising the coordinate frame of the selected one or more maps; And is executable using the at least one processor to perform the above operations; A networked computing device comprising An electronic system comprising (Item 24) The coordinate frame Is a coordinate frame comprising information characterizing a plurality of features of an object within the 3D environment The electronic system according to item 23, comprising a computer data structure comprising (Item 25) The information characterizing the plurality of features is the electronic system according to item 23, comprising descriptors characterizing regions of the 3D environment. (Item 26) Each coordinate frame of the at least one coordinate frame is the electronic system according to item 23, comprising persistent points characterized by features detected in sensor data representing the 3D environment. (Item 27) Each coordinate frame of the at least one coordinate frame is the electronic system according to item 26, comprising a persistent pose. (Item 28) Each coordinate frame of the at least one coordinate frame is the electronic system according to item 26, comprising a persistent coordinate frame.
Brief Description of the Drawings
[0036] The accompanying drawings are not intended to be drawn to scale. In the drawings, each same or substantially same component illustrated in various figures is represented by a like number. For purposes of clarity, not all components are labeled in all the drawings.
[0037]
Figure 1
[0038]
Figure 2
[0039]
Figure 3
[0040]
Figure 4
[0041]
Figure 5A
[0042]
Figure 5B
[0043]
Figure 6A
[0044]
Figure 6B
[0045]
Figure 7
[0046]
Figure 8
[0047]
Figure 9
[0048]
Figure 10
[0049]
Figure 11
[0050]
Figure 12
[0051]
Figure 13
[0052]
Figure 14
[0053]
Figure 15
[0054]
Figure 16
[0055]
Figure 17
[0056]
Figure 18
[0057]
Figure 19
[0058]
Figure 20
[0059]
Figure 21
[0060]
Figure 22
[0061]
Figure 23
[0062]
Figure 24
[0063]
Figure 25
[0064]
Figure 26
[0065]
Figure 27
[0066]
Figure 28
[0067]
Figure 29
[0068]
Figure 30
[0069]
Figure 31A
[0070]
Figure 31B
[0071]
Figure 32
[0072]
Figure 33
[0073]
Figure 34
[0074]
Figure 35
Figure 36
[0075]
Figure 37
[0076]
Figure 38A
Figure 38B
[0077]
Figure 39A
Figure 39B
Figure 39C
Figure 39D
Figure 39E
Figure 39F
[0078]
Figure 40
[0079]
Figure 41
[0080]
Figure 42
[0081]
Figure 43A
[0082]
Figure 43B
[0083]
Figure 43C
[0084]
Figure 44
[0085]
Figure 45
[0086]
Figure 46A
Figure 46B
[0087]
Figure 47
[0088]
Figure 48
[0089]
Figure 49
[0090]
Figure 50
[0091]
Figure 51
[0092]
Figure 52
[0093]
Figure 53
[0094]
Figure 54
[0095]
Figure 55
Figure 56
[0096]
Figure 57
Figure 58
[0097]
Figure 59
[0098]
Figure 60
DETAILED DESCRIPTION OF THE INVENTION
[0099] What is described in this specification is a method and apparatus for providing an X Reality (XR or Cross Reality) scenario. To provide a realistic XR experience to multiple users, an XR system must understand the user's physical surroundings in order to correctly correlate the location of virtual objects with real objects. The XR system may construct an environmental map of the scenario, which may be created from images and / or depth information collected using sensors that are part of the XR devices worn by the users of the XR system.
[0100] The inventors have realized that it may be beneficial to have an XR system in which each XR device develops a local map of its physical environment by integrating information from one or more images collected during a scan at a given point in time, and have recognized its true value. In some embodiments, the coordinate system of that map is tied to the orientation of the device when the scan is initiated. That orientation can change from moment to moment as the user interacts with the XR system, regardless of whether different moments are associated with different users, their respective wearable devices with sensors scanning the environment, or the same user using the same device at different times. The inventors have realized a technique for operating an XR system based on persistent spatial information that overcomes the limitations of an XR system that relies only on spatial information collected for different orientations that vary for different user instances (e.g., time-based snapshots) or sessions of the system (e.g., the time between on and off). This technique can provide an XR scenario for a single or multiple users that is more computationally efficient and immersive, for example, by enabling persistent spatial information to be created, stored, and retrieved by any of a plurality of users of the XR system.
[0101] Persistent spatial information may be represented by a persistent map that can enable one or more functions to enhance the XR experience. The persistent map may be stored in a remote storage medium (e.g., the cloud). For example, a wearable device worn by a user may, after being turned on, read an appropriate stored map that was previously created and stored from a persistent storage device such as a cloud storage device. The previously stored map may be based on data about the environment collected using sensors on the user's wearable device during a previous session. Reading the stored map may enable the use of the wearable device without involving a scan of the physical world using the sensors on the wearable device. Alternatively, or in addition, the system / device may similarly read an appropriate stored map in response to entering a new area of the physical world.
[0102] The stored map may be represented in a canonical form in which each XR device can be related to its local reference frame. In a multi-device XR system, a stored map accessed by one device may be created and stored by another device and / or may be constructed by aggregating data about the physical world collected by sensors on multiple wearable devices that were previously present within at least a portion of the physical world represented by the stored map.
[0103] Furthermore, sharing data about the physical world among multiple devices can enable a shared user experience of virtual content. Two XR devices having access to the same stored map may both be, for example, localized with respect to the stored map. Once localized, the user device may render virtual content having a location defined by a reference by translating that location parallel to a frame or reference maintained by the user device onto the stored map. The user device may use this local reference frame to control the display of the user device and render the virtual content within the defined location.
[0104] To support these and other functions, the XR system may include components that develop, maintain, and use persistent spatial information, including one or more stored maps, based on data about the physical world collected using sensors on the user device. These components may be distributed across the XR system, with some operating, for example, on the head-mounted portion of the user device. Other components may operate on a computer associated with the user that is coupled to the head-mounted portion via a local or personal area network. Still others may operate at a remote location, such as one or more servers accessible via a wide area network.
[0105] These components may include components that can identify information about the physical world collected by, for example, one or more user devices as a persistent map or as information of sufficient quality to be stored within a persistent map. An example of such a component, described in more detail below, is a map merge component. Such a component may, for example, receive input from a user device and determine the suitability of portions of the input for use in updating a persistent map. A map merge component may, for example, split a local map created by a user device into portions, determine the mergability of one or more of the portions with the persistent map, and merge portions that meet the identified mergability criteria into the persistent map. A map merge component may also, for example, promote portions that are not merged with the persistent map to be separate persistent maps.
[0106] As another example, these components may include components that can assist in determining an appropriate persistent map that can be read and used by a user device. An example of such a component, described in more detail below, is a map ranking component. Such a component may, for example, receive input from a user device and identify one or more persistent maps that are likely to represent the area of the physical world in which the device is operating. A map ranking component may, for example, assist in selecting the persistent map to be used by the local device when rendering virtual content, collecting data about the environment, or performing other actions. As an alternative or in addition, a map ranking component may assist in identifying the persistent map to be updated as additional information about the physical world is collected by one or more user devices.
[0107] Further, other components may determine a transformation that converts information captured or described relative to one reference frame to another reference frame. For example, a sensor may be attached to a head-mounted display such that data read from the sensor indicates the location of an object in the physical world relative to the wearer's head pose. One or more transformations may be applied to relate that location information to a coordinate frame associated with a persistent environmental map. Similarly, data indicating where a virtual object should be rendered when represented within the coordinate frame of the persistent environmental map may undergo one or more transformations to be within the reference frame of a display above the user's head. As described in more detail below, there may be multiple such transformations. These transformations may be partitioned across components of the XR system so that they can be efficiently updated and / or applied within a distributed system.
[0108] In some embodiments, the persistent map may be constructed from information collected by multiple user devices. The XR device may capture local spatial information and construct a separate tracking map using information collected by respective sensors of the XR device at various locations and times. Each tracking map may include points that may be associated with features of real objects, each of which may include multiple features. Potentially, in addition to supplying input for creating and maintaining the persistent map, the tracking map may be used to track the movement of a user within a scene, enabling the XR system to estimate the head pose of an individual user based on the tracking map.
[0109] This co-dependence between map creation and head pose estimation constitutes a significant challenge. Substantial processing may be required to simultaneously create a map and estimate the head pose. To make the waiting time not very realistic for the user in the XR experience, the processing must be performed quickly as objects move within the scene (e.g., moving a cup on a table) and as the user moves within the scene. On the other hand, since the weight of the XR device should be lightweight for the user to wear comfortably, the XR device may provide limited computing resources. The lack of computing resources unfortunately cannot be compensated for by using more sensors, as adding sensors would also add weight. Furthermore, either more sensors or more computing resources leads to heat, which can cause deformation of the XR device.
[0110] The inventors have realized techniques for operating an XR system and providing an XR scene for a more immersive user experience, recognizing their true value, such as low usage of computing resources associated with an XR device that can be configured with four video graphic array (VGA) cameras operating at 30 Hz for head pose estimation at a frequency of 1 kHz, one inertial measurement unit (IMU) operating at 1 kHz, the computing power of a single advanced RISC machine (ARM) core, less than 1 GB of memory, and less than 100 Mbp of network bandwidth. These techniques relate to steps of generating and maintaining a map, reducing the processing required to estimate the head pose, and providing and consuming data with low computational overhead.
[0111] These techniques may include hybrid tracking such that the XR system can utilize both (1) patch-based tracking of distinguishable points between successive images of the environment (e.g., frame / frame tracking), and (2) matching of a descriptor-based map of known real-world locations of points of interest corresponding to points of interest in the current image (e.g., map / frame tracking). In frame / frame tracking, the XR system may track specific points of interest (e.g., salient points) such as corners between captured images of the real-world environment. For example, the display system may identify the location of a visual point of interest in the current image that was included in (e.g., located within) the previous image. This identification may be accomplished, for example, using a photometric error minimization process. In map / frame tracking, the XR system may access map information indicative of the real-world location of a point of interest and match the point of interest included in the current image to the point of interest indicated by the map information. Information regarding the point of interest may be stored in the map database as a descriptor. The XR system may calculate its pose based on the matched visual features. U.S. Patent Application No. 16 / 221,065 describes hybrid tracking and is hereby incorporated by reference in its entirety.
[0112] These techniques may include steps to reduce the amount of data processed when constructing a map, such as constructing a sparse map using a set of mapped points and keyframes and / or dividing the map into blocks and enabling block-by-block updates. The mapped points may be associated with points of interest in the environment. The keyframes may include information selected from camera capture data. U.S. Patent Application No. 16 / 520,582 describes steps for determining and / or evaluating a localization map and is hereby incorporated by reference in its entirety.
[0113] In some embodiments, persistent spatial information may be represented in a way that can be easily shared among users and distributed components including applications. Information about the physical world may be represented, for example, as a Persistent Coordinate Frame (PCF). The PCF may be defined based on one or more points that represent features recognized within the physical world. The features may be selected such that they are likely to be the same for each user session of the XR system. The PCF may be sparsely present so that they can be efficiently processed and transferred, and may provide less than all of the available information about the physical world. Techniques for processing persistent spatial information may include creating a dynamic map based on one or more coordinate systems within the real space across one or more sessions, and generating a Persistent Coordinate Frame (PCF), which can be exposed to an XR application, for example, via an Application Programming Interface (API), across the sparse map. These capabilities may be supported by techniques for ranking and merging multiple maps created by one or more XR devices. Persistent spatial information may also make it possible to quickly restore and reset the head pose on each of one or more XR devices in a computationally efficient manner.
[0114] Furthermore, the techniques may enable efficient comparison of spatial information. In some embodiments, an image frame may be represented by a numerical descriptor. The descriptor may be calculated via a transformation that maps a set of features identified within the image to the descriptor. The transformation may be performed within a trained neural network. In some embodiments, the set of features supplied as input to the neural network may be a filtered set of features extracted from the image using techniques that preferentially select features that are likely to be persistent, for example.
[0115] The representation of an image frame as a descriptor, for example, enables efficient matching of new image information with stored image information. The XR system may store descriptors of one or more frames in the underlying layer of the persistent map, together with the persistent map. Similarly, local image frames obtained by the user device may also be converted into such descriptors. By selecting a stored map with descriptors similar to those of the local image frame, one or more persistent maps that are likely to represent the same physical space as the user device can be selected with relatively little processing. In some embodiments, the descriptors are calculated with respect to key frames in the local map and the persistent map, and may further reduce processing when comparing maps. Such efficient comparison may be used, for example, to simplify finding a persistent map for loading or updating in the local device based on image information obtained using the local device.
[0116] The techniques described herein may be used together or separately with many types of devices, including wearable or portable devices with limited computing resources, for many types of scenarios that provide extended or mixed reality scenes. In some embodiments, the techniques may be implemented by one or more services that form part of the XR system.
[0117] AR System Overview
[0118] Figures 1 and 2 illustrate a scene with virtual content that is displayed together with a portion of the physical world. For illustrative purposes, an AR system is used as an example of an XR system. Figures 3-6B illustrate an exemplary AR system that may operate in accordance with the techniques described herein and includes one or more processors, a memory, sensors, and a user interface.
[0119] Referring to FIG. 1, an outdoor AR scene 354 is depicted, and to a user of AR technology, a physical-world park-like setting 356 is visible, featuring people, trees, buildings in the background, and a concrete platform 358. In addition to these items, the user of AR technology also "sees" and perceives a robot image 357 standing on the physical-world concrete platform 358 and a flying comic-like avatar character 352 that appears to be an anthropomorphization of a honeybee, although these elements (e.g., avatar character 352 and robot image 357) do not exist within the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is difficult to produce AR technology that promotes a comfortable, natural, and rich presentation of virtual image elements among other virtual or physical-world image elements.
[0120] Such an AR scene can be achieved using a system that constructs a map of the physical world based on tracking information, which enables a user to place AR content within the physical world, determines the location within the map of the physical world where the AR content is placed, saves the AR scene so that the placed AR content can be reloaded for display within the physical world, for example, during different AR experience sessions, and enables multiple users to share the AR experience. This system can construct and update a digital representation of the physical-world surface around the user. This representation may be used, in whole or in part, to render virtual content such that it appears to be occluded by physical objects between the user and the rendered location of the virtual content for purposes of placing virtual objects, in physics-based interactions, and for virtual character path planning and navigation, or for other operations where information about the physical world is used.
[0121] Figure 2 depicts another example of an indoor AR scene 400 according to some embodiments and shows an exemplary use case of an XR system. The exemplary scene 400 is a living room having a wall, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of AR technology may also perceive virtual objects such as an image on the wall behind the sofa, a bird flying in through the door, a deer peeking out from the bookshelf, and a decoration in the form of a windmill placed on the coffee table.
[0122] Regarding the image on the wall, AR technology requires information about objects and surfaces within the room, such as the shape of a lamp, not only for the surface of the wall but also for occluding the image to correctly render the virtual object. Regarding the flying bird, AR technology requires information about all objects and surfaces around the room to render the bird using realistic physics so that the bird avoids objects and surfaces or bounces back if it collides. Regarding the deer, AR technology requires information about surfaces such as the floor or the coffee table to calculate where to place the deer. Regarding the windmill, the system may be able to identify that it is a separate object from the table and determine that it is movable, while the corner of the shelf or the wall may be determined to be stationary. Such specificities may be used in determining portions of the scene that are used or updated in each of the various operations.
[0123] Virtual objects may be placed within a previous AR experience session. When a new AR experience session starts in a living room, the AR technology requires that the virtual objects be accurately displayed at the previously placed locations and be realistically visible from different viewpoints. For example, the windmill should be displayed as standing on a book rather than floating above a table at different locations without the book. Such floating may occur when the location of the user within the new AR experience session is not accurately located within the living room. As another example, when the user views the windmill from a viewpoint different from the viewpoint when the windmill was placed, the AR technology requires the corresponding side of the displayed windmill.
[0124] The scene may be presented to the user via a system that includes a plurality of components including a user interface that can stimulate one or more user perceptions such as vision, hearing, and / or touch. Additionally, the system may include one or more sensors that can measure parameters of the physical part of the scene including the position and / or movement of the user within the physical part of the scene. Further, the system may include one or more computing devices with associated computer hardware such as memory. These components may be integrated within a single device or may be distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated within a wearable device.
[0125] Figure 3 depicts an AR system 502 configured to provide an experience of AR content that interacts with the physical world 506, according to some embodiments. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by the user as part of a headset such that the user can wear the display across their eyes, such as a pair of goggles or glasses. At least a portion of the display may be transparent such that the user can observe the see-through reality 510. The see-through reality 510 may correspond to the portion of the physical world 506 within the current viewing perspective of the AR system 502, which may correspond to the user's viewing perspective when the user wears a headset incorporating both the display and sensors of the AR system and obtains information about the physical world.
[0126] The AR content may also be presented on the display 508, overlaid on the see-through reality 510. To provide an accurate interaction between the AR content and the see-through reality 510 on the display 508, the AR system 502 may include a sensor 522 configured to capture information about the physical world 506.
[0127] The sensor 522 may include one or more depth sensors that output a depth map 512. Each depth map 512 may have a plurality of pixels that may each represent the distance to a surface within the physical world 506 in a particular direction relative to the depth sensor. Raw depth data may originate from the depth sensor and may create the depth map. Such depth maps may be updated as fast as the depth sensor can form a new image, which may be hundreds or thousands of times per second. However, the data is noisy and incomplete and may have holes that are shown as black pixels on the illustrated depth maps.
[0128] The system may include other sensors such as an image sensor. The image sensor may obtain monocular or stereoscopic information that can be processed to represent the physical world in other ways. For example, the image may be processed within the world reconstruction component 516 to create a mesh representing the connected parts of the objects in the physical world. For example, metadata about such objects, including color and surface texture, may also be obtained using the sensors and stored as part of the world reconstruction.
[0129] The system may also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, the system's head pose tracking component may be used to calculate the head pose in real time. The head pose tracking component may represent the user's head pose within a coordinate frame with six degrees of freedom, including, for example, translations along three perpendicular axes (e.g., forward / backward, up / down, left / right) and rotations about the three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, the sensor 522 may include an inertial measurement unit that can be used to calculate and / or determine the head pose 514. The head pose 514 for the depth map may indicate, for example, the current viewpoint of the sensor capturing the depth map with six degrees of freedom, but the head pose 514 may also be used for other purposes such as associating image information with a particular part of the physical world or associating the position of a display worn on the user's head with the physical world.
[0130] In some embodiments, the head pose information may be derived by methods other than the IMU, such as from the analysis of objects in the image. For example, the head pose tracking component may calculate the relative position and orientation of the AR device with respect to the physical object based on the visual information captured by the camera and the inertial information captured by the IMU. The head pose tracking component may then calculate the head pose of the AR device, for example, by comparing the calculated relative position and orientation of the AR device with respect to the physical object with the characteristics of the physical object. In some embodiments, the comparison is made using one or more of the sensors 522 that are stable over time, such that changes in the position of these features in the images captured over time can be associated with changes in the user's head pose, by identifying features in the images captured using them.
[0131] In some embodiments, the AR device may construct a map from feature points recognized in consecutive images within a series of image frames captured as the user moves through the physical world with the AR device. Each image frame can be obtained from a different pose as the user moves, but the system may adjust the orientation of the features of each consecutive image frame and match the orientation of the initial image frame by matching the features of the consecutive image frames with the previously captured image frames. Translation of the consecutive image frames can be used to align each consecutive image frame and match the orientation of the previously processed image frames so that points representing the same feature will match the corresponding feature points from the previously collected image frames. The frames in the resulting map may have a common orientation established when the first image frame is added to the map. This map may be used to determine the pose of the user in the physical world by matching features from the current image frame with the set of feature points in the common reference frame. In some embodiments, this map may be referred to as a tracking map.
[0132] In addition to enabling the tracking of a user's pose within the environment, this map may enable other components of the system, such as the world reconstruction component 516, to determine the location of physical objects relative to the user. The world reconstruction component 516 may receive the depth map 512 and the head pose 514 and any other data from the sensors and integrate that data into the reconstruction 518. The reconstruction 518 may be more complete and less noisy than the sensor data. The world reconstruction component 516 may update the reconstruction 518 using the spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0133] The reconstruction 518 may include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats may represent alternative representations of the same portion of the physical world or different portions of the physical world. In the illustrated embodiment, on the left side of the reconstruction 518, a portion of the physical world is presented as a global surface, and on the right side of the reconstruction 518, a portion of the physical world is presented as a mesh.
[0134] In some embodiments, the map maintained by the head pose component 514 may be spaced apart from other maps of the physical world that can be maintained. Instead of providing information about locations and other characteristics of surfaces as possibilities, the sparse map may indicate locations of points of interest and / or structures such as corners or edges. In some embodiments, the map may include image frames as captured by the sensor 522. These frames may be reduced to features that may represent points of interest and / or structures. Along with each frame, information about the user's pose from which the frame was obtained may also be stored as part of the map. In some embodiments, all images obtained by the sensor may or may not be stored. In some embodiments, the system may process the images as they are collected by the sensor and select a subset of the image frames for further calculations. The selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may, for example, add new image frames to the map based on overlap with previous image frames already added to the map or based on an image frame that contains a sufficient number of features determined to be likely to represent stationary objects. In some embodiments, the selected image frames or the group of features from the selected image frames may serve as key frames for the map, which are used to provide spatial information.
[0135] The AR system 502 may integrate sensor data from multiple viewpoints of the physical world over time. The pose of the sensor (e.g., position and orientation) may be tracked as the device containing the sensor is moved. As the frame pose of the sensor and how it relates to other poses are understood, these multiple viewpoints of the physical world may each be fused together into a single combined reconstruction of the physical world, which may serve as an abstraction layer for a map and provide spatial information. The reconstruction may be more complete and less noisy than the original sensor data by using spatial and temporal averaging (i.e., averaging of data from multiple viewpoints over time) or any other suitable method.
[0136] In the embodiment illustrated in FIG. 3, the map represents a portion of the physical world in which a user of a single wearable device is present. In that scenario, the head pose associated with a frame within the map may be represented as a local head pose that indicates the orientation relative to the initial orientation of the single device at the start of the session. For example, the head pose may be tracked relative to the initial head pose when the device is turned on or otherwise operated to scan the environment and construct a representation of that environment.
[0137] In combination with the content characterizing that portion of the physical world, the map may include metadata. The metadata, for example, may indicate the capture time of the sensor information used to form the map. The metadata may alternatively or additionally indicate the location of the sensor at the capture time of the information used to form the map. The location may be represented directly, using information from a GPS chip etc., or indirectly, using a Wi-Fi signature etc. that indicates the strength of a signal received from one or more wireless access points while the sensor data was being collected, and / or using the BSSID of the wireless access point to which the user device was connected while the sensor data was being collected.
[0138] The reconstruction 518 may be used for AR functions, such as the production of a surface representation of the physical world for occlusion processing or physics-based processing. This surface representation may change as the user moves or as objects within the physical world change. The side of the reconstruction 518 may be used, for example, by component 520, which produces a changing global surface representation in world coordinates that may be used by other components.
[0139] AR content may be generated, for example, by AR application 504, etc., based on this information. AR application 504 may be, for example, a game program that implements one or more functions based on information about the physical world, such as visual occlusion, physics-based interactions, and environmental inference. This may be done by querying data in a different format from the reconstruction 518 produced by the world reconstruction component 516 to implement these functions. In some embodiments, component 520 may be configured to output an update when the representation within the area of interest of the physical world changes. The area of interest may be set to approximate a part of the physical world within the vicinity of the system's user, such as a part within the user's field of view, or projected (predicted / decided) to enter the user's field of view.
[0140] AR application 504 may use this information to generate and update AR content. The virtual part of the AR content may be combined with see-through reality 510 and presented on display 508 to create a realistic user experience.
[0141] In some embodiments, the AR experience may be part of a system that may include remote processing and / or remote data storage devices, a wearable display device, and / or, in some embodiments, other wearable display devices worn by other users, and may be provided to the user through an XR device. FIG. 4 illustrates, for illustrative convenience, an example of a system 580 (hereinafter referred to as the "system 580") that includes a single wearable device. The system 580 includes a head-mounted display device 562 (hereinafter referred to as the "display device 562") and various mechanical and electronic modules and systems that support the functions of the display device 562. The display device 562 may be coupled to a frame 564, which is wearable by a user or viewer 560 (hereinafter referred to as the "user 560") of the display system and is configured to position the display device 562 in front of the eyes of the user 560. According to various embodiments, the display device 562 may be a sequential display. The display device 562 may be monocular or binocular. In some embodiments, the display device 562 may be an example of the display 508 in FIG. 3.
[0142] In some embodiments, a speaker 566 is coupled to the frame 564 and positioned proximate to the external auditory canal of the user 560. In some embodiments, another speaker (not shown) is positioned adjacent to the other external auditory canal of the user 560 to provide stereo / adjustable sound control. The display device 562 is operably coupled to a local data processing module 570 by means such as a wired conductor or wireless connectivity 568, which may be mounted in various configurations, such as fixed to the frame 564, fixed to a helmet or hat worn by the user 560, built into headphones, or otherwise removably attached to the user 560 (e.g., in a backpack configuration, in a belt attachment configuration).
[0143] The local data processing module 570 may include a processor and a digital memory such as a non-volatile memory (e.g., flash memory), both of which can be used to assist in data processing, caching, and storage. The data includes a) data captured from sensors such as an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope (e.g., operably coupled to frame 564 or otherwise attachable to user 560), and / or b) data that may be obtained and / or processed using remote processing module 572 and / or remote data repository 574 for passage to display device 562 after processing or reading.
[0144] In some embodiments, the wearable device may communicate with remote components. The local data processing module 570 may be operably coupled to the remote processing module 572 and the remote data repository 574 via communication links 576, 578, respectively, such as a wired or wireless communication link, such that these remote modules 572, 574 are operably coupled to each other and available as resources to the local data processing module 570. In some embodiments, the head pose tracking component described above may be implemented at least partially within the local data processing module 570. In some embodiments, the world reconstruction component 516 in FIG. 3 may be implemented at least partially within the local data processing module 570. For example, the local data processing module 570 may be configured to execute computer-executable instructions and generate a map and / or a physical world representation based at least in part on at least a portion of the data.
[0145] In some embodiments, the processing may be distributed across local and remote processors. For example, local processing may be used to construct a map (e.g., a tracking map) on the user's device based on sensor data collected using sensors on the user's device. Such a map may be used by an application on the user's device. Additionally, previously created maps (e.g., reference maps) may be stored in the remote data repository 574. If a suitable stored or persistent map is available, it may be used instead of, or in addition to, the tracking map created locally on the device. In some embodiments, the tracking map may be geolocated with respect to a stored map such that the correspondence can be oriented with respect to the position of the wearable device at the time the user turned the system on, and with respect to one or more persistent features, and with respect to a reference map that can be oriented with respect to the position of the wearable device at the time the user turned the system on. In some embodiments, the persistent map may be loaded onto the user device and enable the rendering of virtual content without the latency associated with the scanning of the location to construct a tracking map of the user's complete environment from sensor data obtained during the scan. In some embodiments, the user device may access a remote persistent map (e.g., stored in the cloud) without having to download the persistent map onto the user device.
[0146] Alternatively, or in addition, the tracking map may be merged with previously stored maps to extend those maps or improve their quality. The process for determining whether a suitable previously created environmental map is available and / or whether to merge the tracking map with one or more stored environmental maps may be performed within the local data processing module 570 or the remote processing module 572.
[0147] In some embodiments, the local data processing module 570 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the computational budget of the local data processing module 570 but enable a smaller device. In some embodiments, the world reconstruction component 516 may use a computational budget less than that of a single advanced RISC machine (ARM) core so that the remaining computational budget of the single ARM core can be accessed for other uses such as mesh extraction, etc., and generate a physical world representation in real time over an unspecified space.
[0148] In some embodiments, the remote data repository 574 may include a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local data processing module 570, enabling fully autonomous use from the remote module. In some embodiments, all data is stored and all or most of the calculations are performed within the remote data repository 574, enabling a smaller device. The world reconstruction may be stored, for example, in whole or in part, within this repository 574.
[0149] In an embodiment where the data is remotely stored and accessible via a network, the data may be shared by multiple users of the augmented reality system. For example, the user device may upload its tracking map and expand it within the environmental map database. In some embodiments, the upload of the tracking map occurs at the end of the user session with the wearable device. In some embodiments, the upload of the tracking map may occur continuously, semi - continuously, intermittently, at a predefined time, after a predefined period from the previous upload, or triggered by an event. The tracking map uploaded by any user device may be used to expand or improve the previously stored map, regardless of whether it is based on data from that user device or any other user device. Similarly, the persistent map downloaded to the user device may be based on data from that user device or any other user device. Thus, a high - quality environmental map may be readily available to the user to improve the experience using the AR system.
[0150] In some embodiments, the local data processing module 570 is operably coupled to the battery 582. In some embodiments, the battery 582 is a removable power source such as a commercially available battery. In other embodiments, the battery 582 is a lithium - ion battery. In some embodiments, the battery 582 includes both an internal lithium - ion battery that can be charged by the user 560 during the non - operating time of the system 580 and a removable battery, such that the user 560 can operate the system 580 for a longer time period without having to connect to a power source, charge the lithium - ion battery, or shut off the system 580 and replace the battery.
[0151] Figure 5A illustrates a user 530 wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as "environment 532"). Information captured by the AR system along the user's movement path may be processed into one or more tracking maps. The user 530 positions the AR display system at location 534, and the AR display system records ambient information about the passable world for location 534 (e.g., a digital representation of real objects in the physical world that can be memorized and updated as real objects in the physical world change). That information may be stored as a pose in combination with an image, features, directional audio input, or other desired data. Location 534 is aggregated, for example, as part of a tracking map, for data input 536 and processed at least by a passable world module 538, which may be implemented, for example, by processing on the remote processing module 572 of FIG. 4. In some embodiments, the passable world module 538 may include a head pose component 514 and a world reconstruction component 516 such that the processed information can be combined with other information about the physical objects used in the rendered virtual content to indicate the location of the objects in the physical world.
[0152] The passable world module 538 determines, at least in part, where and how the AR content 540 can be placed within the physical world as determined from the data input 536. The AR content is "placed" within the physical world by presenting both a representation of the physical world and the AR content via the user interface, and the AR content is rendered as if interacting with objects within the physical world, and the objects within the physical world are presented as if the AR content obscures the user's view of those objects when appropriate. In some embodiments, the AR content may be placed by appropriately selecting a portion of a fixed element 542 (e.g., a table) from the reconstruction (e.g., reconstruction 518) and determining the shape and position of the AR content 540. As an example, the fixed element may be a table and the virtual content may be positioned to appear on that table. In some embodiments, the AR content may be placed within a structure within the field of view 544, which may be the current field of view or an estimated future field of view. In some embodiments, the AR content may be persisted with respect to a model 546 (e.g., a mesh) of the physical world.
[0153] As described, the fixed element 542 serves as a proxy (e.g., a digital copy) for any fixed element in the physical world that can be stored within the passable world module 538 such that the user 530 can perceive content on the fixed element 542 each time it is visible to the user 530 without the system having to map to the fixed element 542. The fixed element 542 may thus be a mesh model that is stored by the passable world module 538 for future reference by multiple users, even though it is determined from a previous modeling session or from a different user. Thus, the passable world module 538 recognizes the environment 532 from a previously mapped environment and can display AR content without the user 530's device first mapping all or part of the environment 532, saving calculation processes and cycles and avoiding latency for any rendered AR content.
[0154] The mesh model 546 of the physical world may be created by the AR display system, interact with the AR content 540, and the appropriate surfaces and metrics for display can be stored by the passable world module 538 for future retrieval by the user 530 or other users without having to recreate the model completely or partially. In some embodiments, the data input 536 provides the passable world module 538 with inputs such as the geographical location, user identification, and current activity indicating which of the one or more fixed elements 542 are available, the AR content 540 last placed on the fixed element 542, and whether that same content should be displayed (such AR content being "persistent" content regardless of whether the user is viewing a particular passable world model).
[0155] Even in embodiments where an object is considered to be fixed (e.g., a kitchen table), the passable world module 538 may update those objects in the model of the physical world as needed to account for possible changes to the physical world. The models of fixed objects may be updated very infrequently. Other objects in the physical world may be considered to be moving or otherwise not fixed (e.g., a kitchen chair). To render the AR scene with a realistic feel, the AR system may update the positions of these non-fixed objects at a much higher frequency than that used to update the fixed objects. To enable accurate tracking of all objects in the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.
[0156] FIG. 5B is a schematic illustration of the viewing optics assembly 548 and associated components. In some embodiments, two eye tracking cameras 550 are directed toward the user's eyes 549 to detect metrics of the user's eyes 549, such as eye shape, eyelid occlusion, pupil direction, and glints on the user's eyes 549.
[0157] In some embodiments, one of the sensors is a depth sensor 551, such as a time-of-flight sensor, that emits signals into the world and detects the reflections of those signals from neighboring objects to determine the distance to a given object. The depth sensor may, for example, quickly determine whether an object has entered the user's field of view as a result of either the movement of those objects or a change in the user's pose. However, information about the positions of objects within the user's field of view may alternatively or additionally be collected using other sensors. Depth information may, for example, be obtained from a stereoscopic image sensor or a plenoptic sensor.
[0158] In some embodiments, the world camera 552 records, maps, and / or otherwise creates a model of the environment 532 and detects inputs that can affect the AR content, capturing a view wider than the periphery. In some embodiments, the world camera 552 and / or the camera 553 may be grayscale and / or color image sensors, which may output grayscale and / or color image frames at a fixed time interval. The camera 553 may further capture a physical world image within the user's field of view at a specific time. The pixels of the frame-based image sensor may be sampled iteratively even if their values are invariant. The world camera 552, the camera 553, and the depth sensor 551 each have individual fields of view 554, 555, and 556, respectively, and collect and record data from a physical world scene such as the physical world environment 532 depicted in FIG. 34A.
[0159] The inertial measurement unit 557 may determine the movement and orientation of the visual optics assembly 548. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 551 is operably coupled to the eye tracking camera 550 as a confirmation of the measured focusing for the actual distance that the user's eye 549 is looking at.
[0160] It should be understood that the visual optics assembly 548 may include some of the components illustrated in FIG. 34B, and may include components instead of or in addition to the illustrated components. In some embodiments, for example, the visual optics assembly 548 may include two world cameras 552 instead of four. Alternatively, or in addition, the cameras 552 and 553 do not need to capture a visible light image of their full field of view. The visual optics assembly 548 may include other types of components. In some embodiments, the visual optics assembly 548 may include one or more dynamic vision sensors (DVSs), the pixels of which may respond asynchronously to a relative change in light intensity exceeding a threshold.
[0161] In some embodiments, the visual optical system assembly 548 may not include the time-of-flight based depth sensor 551. In some embodiments, for example, the visual optical system assembly 548 may include one or more plenoptic cameras, the pixels of which may capture the light intensity and angle of the incident light, from which depth information can be determined. For example, the plenoptic camera may include an image sensor overlaid with a transmissive diffraction mask (TDM). Alternatively, or in addition, the plenoptic camera may include an image sensor containing angle sensing pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors may serve as a depth information source instead of or in addition to the depth sensor 551.
[0162] Also, it should be understood that the component configuration in FIG. 5B is provided as an example. The visual optical system assembly 548 may include components with any suitable configuration, which may be set to provide the user with a practical maximum field of view for a particular set of components. For example, if the visual optical system assembly 548 has one world camera 552, the world camera may be installed within the central region of the visual optical system assembly instead of on the side.
[0163] Information from sensors within the visual optics assembly 548 may be coupled to one or more of the processors within the system. The processor may generate data that can be rendered to make the user perceive that the virtual content interacts with objects in the physical world. The rendering may be implemented in any suitable manner, including the step of generating image data depicting both physical and virtual objects. In other embodiments, the physical and virtual content may be depicted in one scene by modulating the opacity of the display device through which the user sees through the physical world. The opacity may be controlled to create the appearance of the virtual object and block the view of objects in the physical world that are occluded by the virtual object from the user. In some embodiments, the image data may be modified to be perceived by the user such that the virtual content realistically interacts with the physical world when the virtual content is viewed through the user interface (e.g., clipping the content and taking occlusion into account), and may include only the virtual content.
[0164] The location on the visual optics assembly 548 where the content may be displayed to create the impression of an object at a particular location may depend on the physics of the visual optics assembly. Additionally, the pose of the user's head with respect to the physical world and the direction in which the user's eyes are looking will affect the location within the physical world content that will be displayed at a particular location on the visual optics assembly where the content will appear. Sensors such as those described above may collect this information and / or supply information from which this information can be calculated so that a processor receiving the sensor input can calculate the location where the object should be rendered on the visual optics assembly 548 to create the desired appearance for the user.
[0165] Regardless of how the content is presented to the user, a model of the physical world can be used so that characteristics of virtual objects, including the shape, position, movement, and visibility of the virtual objects, that can be affected by physical objects can be correctly calculated. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 518.
[0166] The model may be created from data collected from sensors on the user's wearable device. However, in some embodiments, the model may be created from data collected by multiple users, which may be aggregated within a computing device remote from all users (and may be "in the cloud").
[0167] The model may be created, at least in part, by a world reconstruction system, such as world reconstruction component 516 of FIG. 3 described in more detail in FIG. 6A. World reconstruction component 516 may include a perception module 660 that can generate, update, and store a representation for a portion of the physical world. In some embodiments, perception module 660 may represent a portion of the physical world within the reconstruction range of the sensors as a plurality of voxels. Each voxel corresponds to a 3D cube of a predetermined volume within the physical world, includes surface information, and can indicate whether a surface exists within the volume represented by the voxel. The voxels may be assigned a value indicating whether the corresponding volume has been determined to include the surface of a physical object, is determined to be empty, or has not yet been measured using the sensors and thus its value is unknown. It should be understood that the values indicating voxels determined to be empty or unknown need not be explicitly stored, and the voxel values may be stored in computer memory in any suitable manner, including not storing information regarding voxels determined to be empty or unknown.
[0168] In addition to generating information for a persistent world representation, the perception module 660 may identify and output an indication of a change in the area around the user of the AR system. Such an indication of change may trigger an update to the volumetric data stored as part of the persistent world, or trigger other functions such as generating and updating AR content, triggering component 604 to generate AR content, etc.
[0169] In some embodiments, the perception module 660 may identify changes based on a signed distance function (SDF) model. The perception module 660 may be configured to receive sensor data such as, for example, a depth map 660a and a head pose 660b, and then fuse the sensor data into the SDF model 660c. The depth map 660a may directly provide SDF information, and the image may be processed to arrive at SDF information. The SDF information represents the distance from the sensor used to capture that information. Since those sensors can be part of a wearable unit, the SDF information may represent the physical world from the perspective of the wearable unit, and thus the user's perspective. The head pose 660b may enable the SDF information to be associated with voxels within the physical world.
[0170] In some embodiments, the perception module 660 may generate, update, and store a representation for a portion of the physical world that is within the perception range. The perception range may be determined based at least in part on the reconstruction range of the sensor, which may be determined based at least in part on the limits of the sensor's observation range. As a specific example, an active depth sensor that operates using active IR pulses can reliably operate over a certain range of distances and can create an observation range of the sensor that can be several centimeters or tens of centimeters to several meters.
[0171] The world reconstruction component 516 may include additional modules that can interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data obtained by the perception module 660. The persistent world module 662 may also include representations of the physical world in various formats. For example, volumetric metadata 662b such as voxels may be stored along with the mesh 662c and the plane 662d. In some embodiments, other information such as depth maps may also be stored.
[0172] In some embodiments, a representation of the physical world, such as that illustrated in FIG. 6A, may provide relatively dense information about the physical world as compared to a sparse map such as a tracking map based on feature points, as described above.
[0173] In some embodiments, the perception module 660 may include modules that generate representations of the physical world in various formats, including, for example, the mesh 660d, the plane, and the semantics 660e. Representations of the physical world may be stored across local and remote storage media. Representations of the physical world may be described in different coordinate frames, for example, depending on the location of the storage media. For example, a representation of the physical world stored within a device may be described within a coordinate frame local to the device. Representations of the physical world may have counterparts stored within the cloud. The counterparts within the cloud may be described within a coordinate frame shared by all devices within the XR system.
[0174] In some embodiments, these modules may generate the representation based on data within the perceptual range of one or more sensors at the time the representation is generated, data captured at previous times, and information within the persistent world module 662. In some embodiments, these components may act on depth information captured using a depth sensor. However, the AR system may include a vision sensor and may generate such a representation by analyzing monocular or binocular vision information.
[0175] In some embodiments, these modules may act on regions of the physical world. Those modules may be triggered to update a sub-region of the physical world when the perception module 660 detects a change in the physical world within that sub-region. Such a change may be detected, for example, by detecting a new surface within the SDF model 660c or by other criteria such as a change in the values of a sufficient number of voxels representing the sub-region.
[0176] The world reconstruction component 516 may include a component 664 that may receive a representation of the physical world from the perception module 660. Information about the physical world may be pulled by these components, for example, according to usage requests from an application. In some embodiments, the information may be pushed to the usage components via an indication of a change in a pre-identified region or a change in the physical world representation within the perceptual range. The component 664 may include, for example, game programs and other components that perform processing for visual occlusion, physics-based interactions, and environmental inference.
[0177] In response to a query from component 664, the perception module 660 may transmit a representation for the physical world in one or more formats. For example, when component 664 indicates that the usage is for visual occlusion or physics-based interaction, the perception module 660 may transmit a surface representation. When component 664 indicates that the usage is for environmental inference, the perception module 660 may transmit a mesh, plane, and semantics of the physical world.
[0178] In some embodiments, the perception module 660 may include a component that provides format information to component 664. An example of such a component may be a raycasting component 660f. The usage component (e.g., component 664) may query, for example, information about the physical world from a particular viewpoint. The raycasting component 660f may select from one or more representations of the physical world data within the field of view from that viewpoint.
[0179] As should be understood from the foregoing description, the perception module 660 or another component of the AR system may process data and create a 3D representation of a portion of the physical world. The data to be processed can be reduced, at least in part, by extracting and persisting planar data that thins out a portion of the 3D reconstruction volume, capturing, persisting, and updating 3D reconstruction data in blocks that enable local updates while maintaining neighborhood coherence, deriving occlusion data from a combination of one or more depth data sources, providing such occlusion data to an application that generates such scenes, and / or performing multistage mesh simplification. The reconstruction may contain data of different levels of sophistication, including, for example, raw data such as live depth data, fused volumetric data such as voxels, and computed data such as meshes.
[0180] In some embodiments, the components of the passable world model may be distributed, with some parts being executed locally on the XR device and some parts being executed remotely, such as on a network connected to a server or otherwise in the cloud. The distribution of information processing and storage between the local XR device and the cloud can affect the functionality of the XR system and the user experience. For example, by distributing processing to the cloud, reducing the processing on the local device can enable a longer battery life and reduce the heat generated on the local device. However, distributing much more processing to the cloud can create unacceptable wait times that cause an undesirable user experience.
[0181] FIG. 6B depicts a distributed component architecture 600 configured for spatial computing, according to some embodiments. The distributed component architecture 600 may include a passable world component 602 (e.g., PW538 in FIG. 5A), a Lumin OS 604, an API 606, an SDK 608, and an application 610. The Lumin OS 604 may include a Linux-based kernel with custom drivers compatible with the XR device. The API 606 may include an application programming interface that provides XR applications (e.g., application 610) access to the spatial computing features of the XR device. The SDK 608 may include a software development kit that enables the creation of XR applications.
[0182] One or more components within the architecture 600 may create and maintain a model of the passable worlds. In this example, sensor data is collected on a local device. The processing of that sensor data may be performed locally on the XR device, at least in part, and in the cloud, at least in part. The PW 538 may include an environmental map that is created, at least in part, based on data captured by an AR device worn by a plurality of users. During a session of an AR experience, an individual AR device (such as the wearable device described above in connection with FIG. 4) may create a tracking map, which is one type of map.
[0183] In some embodiments, the device may include a component that constructs both a sparse map and a dense map. The tracking map may serve as a sparse map and may include information about the head pose of the AR device scanning the environment and the objects detected in that environment at each head pose. Those head poses may be maintained locally for each device. For example, the head pose on each device may be relative to the initial head pose when the device was turned on for that session. As a result, each tracking map may be local to the device that created it. The dense map may include surface information, which may be represented by a mesh or depth information. Alternatively, or in addition, the dense map may include a higher level of information derived from surface or depth information such as the location and / or characteristics of planes and / or other objects.
[0184] In some embodiments, the creation of the dense map may be independent of the creation of the sparse map. The creation of the dense map and the sparse map may be performed, for example, within separate processing pipelines in the AR system. Separating the processing may, for example, enable different types of map generation or processing to be performed at different rates. The sparse map may be refreshed, for example, at a faster rate than the dense map. However, in some embodiments, the processing of the dense and sparse maps may be related even when performed within different pipelines. Changes in the physical world exposed in the sparse map may, for example, trigger an update of the dense map, or vice versa. Further, even when created independently, the maps may be used together. For example, the coordinate system derived from the sparse map may be used to define the position and / or orientation of objects within the dense map.
[0185] The sparse map and / or the dense map may persist for reuse by the same device and / or for sharing with other devices. Such persistence may be achieved by storing the information in the cloud. The AR device may send the tracking map to the cloud and merge it, for example, with an environmental map selected from previously stored persistent maps in the cloud. In some embodiments, the selected persistent map may be sent from the cloud to the AR device for merging. In some embodiments, the persistent maps may be oriented with respect to one or more persistent coordinate frames. Such maps may serve as canonical maps since they can be used by any of a plurality of devices. In some embodiments, a model of the traversable world may include or be created with one or more canonical maps. The device may use the canonical map by determining the transformation between its local coordinate frame, in which the device performs some operations, and the canonical map.
[0186] The reference map may result as a tracking map (TM) (e.g., TM1102 in FIG. 31A), which may be promoted to the reference map. The reference map may be persisted such that a device accessing the reference map can use the information within the reference map to determine the location of objects represented within the reference map in the physical world around the device once the transformation between its local coordinate system and the coordinate system of the reference map is determined. In some embodiments, the TM may be a head pose sparse map created by an XR device. In some embodiments, the reference map may be created when an XR device transmits one or more TMs to a cloud server for merging with additional TMs captured by the XR device at different times or by other XR devices.
[0187] The reference map or other maps may provide information about a portion of the physical world represented by data processed to create the individual maps. FIG. 7 depicts an exemplary tracking map 700 according to some embodiments. The tracking map 700 may provide a top view 706 of physical objects in the corresponding physical world represented by points 702. In some embodiments, the map points 702 may represent features of the physical objects, which may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. The features may be derived from a processed image such as may be obtained using sensors of a wearable device within the augmented reality system. The features may be derived, for example, by processing an image frame output by a sensor and identifying the features based on large gradients or other suitable criteria within the image. Further processing may limit the number of features within each frame. For example, the processing may select features that are likely to represent persistent objects. One or more heuristics may be applied for this selection.
[0188] The tracking map 700 may include data regarding points 702 collected by the device. For each image frame with data points included in the tracking map, the pose may be stored. The pose may represent the orientation from which the image frame was captured such that feature points within each image frame can be spatially correlated. The pose may be determined by positioning information such that it can be derived from sensors such as IMU sensors on the wearable device. Alternatively, or in addition, the pose may be determined by matching the image frame with other image frames depicting overlapping portions of the physical world. By finding such a spatial correlation, which can be accomplished by matching a subset of the feature points in the two frames, the relative pose between the two frames can be calculated. The relative pose may be appropriate for the tracking map since the map may be relative to a local coordinate system of the device established based on the initial pose of the device when construction of the tracking map was initiated.
[0189] Since much of the information collected using sensors is likely to be redundant, not all of the feature points and image frames collected by the device may be retained as part of the tracking map. Rather, only certain frames may be added to the map. Those frames may be selected based on one or more criteria such as the degree of overlap with image frames already in the map, the number of new features they contain, or a quality metric regarding the features within the frame. Image frames not added to the tracking map may be discarded or may be used to revise the location of features. As a further alternative, all or most of the image frames, represented as a set of features, may be retained, but a subset of those frames may be designated as key frames, which are used for further processing.
[0190] The key frames may be processed to produce key rigs 704. The key frames may be processed to produce a 3D set of feature points and stored as key rigs 704. Such processing may involve, for example, comparing image frames simultaneously derived from two cameras and stereoscopically determining the 3D positions of the feature points. Metadata such as pose may be associated with these key frames and / or key rigs.
[0191] The environment map may have any of a plurality of formats, depending on, for example, the storage location of the environment map, including, for example, the local storage device and the remote storage device of the AR device. For example, a map within a remote storage device may have a higher resolution than a map within a local storage device on a wearable device if memory is limited. To transmit a higher resolution map from the remote storage device to the local storage device, the map may be downsampled or otherwise converted to an appropriate format, such as by reducing the number of poses per area of the physical world stored within the map and / or the number of feature points stored per pose. In some embodiments, a slice or portion of the high resolution map from the remote storage device may be transmitted to the local storage device and the slice or portion is not downsampled.
[0192] The database of environment maps may be updated as new tracking maps are created. To determine which of potentially a very large number of environment maps within the database should be updated, the updating step may include efficiently selecting one or more environment maps stored within the database that are associated with the new tracking map. The one or more selected environment maps may be ranked by relevance, and one or more of the highest ranked maps may be selected for processing to merge the higher ranked selected environment maps and the new tracking map to create one or more updated environment maps. When the new tracking map represents a portion of the physical world for which there is no existing environment map to update across it, the tracking map may be stored in the database as a new environment map.
[0193] View-independent display
[0194] What is described herein are methods and apparatus for providing virtual content using an XR system independent of the location of the eye viewing the virtual content. Conventionally, virtual content is re-rendered in response to any movement of the display system. For example, if a user wearing a display system views a virtual representation of a three-dimensional (3D) object on the display and walks around the area where the 3D object appears, the 3D object should be re-rendered for each viewpoint to have the perception that the user is walking around the object while occupying real space. However, re-rendering consumes significant computational resources of the system and causes artifacts due to latency.
[0195] The inventors recognized and appreciated the true value that the head pose (e.g., the location and orientation of a user wearing an XR system) can be used to render virtual content independently of eye rotations in the user's head. In some embodiments, a dynamic map of the scene is independent of eye rotations in the user's head and / or independent of sensor deformations caused by heat generated, for example, during high-speed computationally intensive operations, such that virtual content that interacts with the dynamic map can be robustly rendered across one or more sessions and generated based on multiple coordinate frames in the real space. In some embodiments, the configuration of the multiple coordinate frames can enable a first XR device worn by a first user and a second XR device worn by a second user to recognize a common location within the scene. In some embodiments, the configuration of the multiple coordinate frames can enable a user wearing an XR device to visually perceive virtual content within the same location of the scene.
[0196] In some embodiments, a tracking map may be constructed within a world coordinate frame, which may have a world origin. The world origin may be the first pose of the XR device when the XR device is powered on. The world origin may be aligned with gravity so that the developer of the XR application can obtain gravity alignment without extra work. Different tracking maps may be constructed within different world coordinate frames because the tracking map can be captured by the same XR device in different sessions and / or different XR devices worn by different users. In some embodiments, a session of the XR device may start when the device is powered on and continue until it is powered off. In some embodiments, the XR device may have a head coordinate frame, which may have a head origin. The head origin may be the current pose of the XR device when an image is captured. The difference between the head poses of the world coordinate frame and the head coordinate frame may be used to estimate the tracking route.
[0197] In some embodiments, the XR device may have a camera coordinate frame, which may have a camera origin. The camera origin may be the current pose of one or more sensors of the XR device. The inventors have recognized and appreciated the true value of the configuration of the camera coordinate frame in enabling a robust display of virtual content independently of eye rotations in the user's head. This configuration also enables a robust display of virtual content independently of, for example, sensor deformations due to heat generated during operation.
[0198] In some embodiments, the XR device may have a head unit with a head-mountable frame that the user can attach to their head and that may include two waveguides, one in front of each eye of the user. The waveguides may be transparent such that ambient light from real-world objects can pass through the waveguides and the user can see the real-world objects. Each waveguide may transmit light projected from a projector to the user's individual eye. The projected light may form an image on the retina of the eye. The retina of the eye thus receives both ambient light and the projected light. The user may be able to see both real-world objects and one or more virtual objects created by the projected light at the same time. In some embodiments, the XR device may have sensors that detect real-world objects around the user. These sensors may be, for example, cameras that capture images that can be processed to identify the location of real-world objects.
[0199] In some embodiments, rather than associating virtual content within a world coordinate frame, the XR system may assign a coordinate frame to the virtual content. Such a configuration allows the virtual content to be described regardless of the location at which it is rendered for the user, but may be associated with a more persistent frame location, such as a persistent coordinate frame (PCF) described in connection with FIGS. 14-20C, and rendered at a defined location. As the location of an object changes, the XR device may detect a change in the environmental map and determine the movement of the head unit worn by the user relative to the real-world object.
[0200] FIG. 8 illustrates a user experiencing virtual content as rendered by an XR system 10 within a physical environment, according to some embodiments. The XR system may include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is present within a physical environment with real objects in the form of a table 16.
[0201] In the illustrated embodiment, the first XR device 12.1 includes a head unit 22, a belt pack 24, and a cable connection 26. The first user 14.1 secures the head unit 22 to their head and the belt pack 24, remote from the head unit 22, on their waist. The cable connection 26 connects the head unit 22 to the belt pack 24. The head unit 22 includes technology used to display virtual object(s) to the first user 14.1 while allowing the first user 14.1 to see real objects such as the table 16. The belt pack 24 primarily includes the processing and communication capabilities of the first XR device 12.1. In some embodiments, the processing and communication capabilities may reside wholly or partially within the head unit 22 such that the belt pack 24 may be removable or located within another device such as a backpack.
[0202] In the illustrated embodiment, the belt pack 24 is connected to the network 18 via a wireless connection. The server 20 is connected to the network 18 and holds data representing local content. The belt pack 24 downloads data representing local content from the server 20 via the network 18. The belt pack 24 provides data to the head unit 22 via a cable connection 26. The head unit 22 may include a display having a light source, such as a laser light source or a light emitting diode (LED) light source, and a waveguide for guiding light.
[0203] In some embodiments, the first user 14.1 may mount the head unit 22 on his or her head and the belt pack 24 on his or her waist. The belt pack 24 may download image data representing virtual content from the server 20 via the network 18. The first user 14.1 may be able to see the table 16 through the display of the head unit 22. A projector forming part of the head unit 22 may receive the image data from the belt pack 24 and generate light based on the image data. The light may travel through one or more of the waveguides forming part of the display of the head unit 22. The light may then exit the waveguide and propagate onto the retina of the eyes of the first user 14.1. The projector may generate light in a pattern that is replicated on the retina of the eyes of the first user 14.1. The light hitting the retina of the eyes of the first user 14.1 may have a selected depth of field so that the first user 14.1 perceives the image at a preselected depth behind the waveguide. Additionally, both eyes of the first user 14.1 may receive slightly different images so that the brain of the first user 14.1 perceives a three-dimensional image or images at a selected distance from the head unit 22. In the illustrated embodiment, the first user 14.1 perceives virtual content 28 above the table 16. The virtual content 28 and the ratio of its location and distance from the first user 14.1 are determined by the data representing the virtual content 28 and the various coordinate frames used to display the virtual content 28 to the first user 14.1.
[0204] In the illustrated embodiment, the virtual content 28 is invisible from the perspective of the drawing and is visible to the first user 14.1 through the use of the first XR device 12.1. The virtual content 28 may initially reside as a data structure within the visual data and an algorithm within the belt pack 24. The data structure may then be exposed as light when the projector of the head unit 22 generates light based on the data structure. The virtual content 28 does not exist within the three-dimensional space in front of the first user 14.1, but it should be understood that the virtual content 28 is still represented in FIG. 1 within the three-dimensional space for illustrative purposes of what the wearer of the head unit 22 perceives. The visualization of computer data within the three-dimensional space may be used in this description to illustrate how data structures that facilitate the rendering perceived by one or more users are interrelated among the data structures within the belt pack 24.
[0205] FIG. 9 illustrates components of the first XR device 12.1 according to some embodiments. The first XR device 12.1 may include the head unit 22, for example, a rendering engine 30, various coordinate systems 32, various origin and destination coordinate frames 34, and various origin / destination coordinate frame converters 36, and may include various components that form part of the visual data and algorithms. The various coordinate systems may be based on the inherent properties of the XR device or may be determined by referring to other information such as a persistent pose or a persistent coordinate system as described herein.
[0206] The head unit 22 may include a head-mountable frame 40, a display system 42, a real object detection camera 44, a motion tracking camera 46, and an inertial measurement unit 48.
[0207] The head-mountable frame 40 may have a shape that can be fixed to the head of the first user 14.1 in FIG. 8. The display system 42, the real object detection camera 44, the movement tracking camera 46, and the inertial measurement unit 48 are mounted on the head-mountable frame 40 and can thus move together with the head-mountable frame 40.
[0208] The coordinate system 32 may include a local data system 52, a world frame system 54, a head frame system 56, and a camera frame system 58.
[0209] The local data system 52 may include a data channel 62, a local frame determination routine 64, and a local frame storage instruction 66. The data channel 62 may be a hardware component such as an internal software routine, an external cable, or a radio frequency receiver, or a hybrid component such as an open port. The data channel 62 may be configured to receive image data 68 representing virtual content.
[0210] The local frame determination routine 64 may be connected to the data channel 62. The local frame determination routine 64 may be configured to determine a local coordinate frame 70. In some embodiments, the local frame determination routine may determine the local coordinate frame based on a real-world object or a real-world location. In some embodiments, the local coordinate frame may be based on the upper edge relative to the bottom edge of the browser window, the head or feet of a character, a node on the outer surface of a prism or bounding box surrounding the virtual content, or any other suitable location for installing a coordinate frame that defines the facing direction of the virtual content and the location where the virtual content should be installed (e.g., a node such as an installation node or an anchor node).
[0211] The local frame storage instruction 66 may be connected to the local frame determination routine 64. Those skilled in the art will understand that software modules and routines can be "connected" to each other through subroutines, calls, etc. The local frame storage instruction 66 may store the local coordinate frame 70 as the local coordinate frame 72 within the origin and destination coordinate frames 34. In some embodiments, the origin and destination coordinate frames 34 may be one or more coordinate frames that can be manipulated or transformed in order for virtual content to persist across sessions. In some embodiments, a session may be the time period between the boot-up and shutdown of an XR device. Two sessions may be two startup and shutdown cycles for a single XR device, or startup and shutdown for two different XR devices.
[0212] In some embodiments, the origin and destination coordinate frames 34 may be coordinate frames involved in one or more transformations required for the XR devices of a first user and a second user to recognize a common location. In some embodiments, the destination coordinate frame may be the output of a series of calculations and transformations applied to the target coordinate frame in order for the first and second users to view virtual content at the same location.
[0213] The rendering engine 30 may be connected to the data channel 62. The rendering engine 30 may receive the image data 68 from the data channel 62 such that the rendering engine 30 can render virtual content, at least in part, based on the image data 68.
[0214] The display system 42 may be connected to the rendering engine 30. The display system 42 may include components that convert the image data 68 into visible light. The visible light may form one or two patterns per eye. The visible light may be incident on the eye of the first user 14.1 in FIG. 8 and may be detected on the retina of the eye of the first user 14.1.
[0215] The real object detection camera 44 may include one or more cameras that can capture images from different sides of the head-mounted frame 40. The motion tracking camera 46 may include one or more cameras that capture images on the side surface of the head-mounted frame 40. One set of one or more cameras may be used instead of two sets of one or more cameras representing the real object detection camera 44 and the motion tracking camera 46. In some embodiments, the cameras 44, 46 may capture images. As described above, these cameras may collect data that is used to construct a tracking map.
[0216] The inertial measurement unit 48 may include several devices that are used to detect the movement of the head unit 22. The inertial measurement unit 48 may include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of the inertial measurement unit 48, in combination, track the movement of the head unit 22 in at least three orthogonal directions and about at least three orthogonal axes.
[0217] In the illustrated embodiment, the world frame system 54 includes a world surface determination routine 78, a world frame determination routine 80, and a world frame memory command 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 receives an image and / or keyframes based on the image captured by the real object detection camera 44, processes the image, and identifies the surface within the image. A depth sensor (not shown) may determine the distance to the surface. The surface is thus represented by data in three dimensions, including its size, shape, and distance from the real object detection camera.
[0218] In some embodiments, the world coordinate frame 84 may be based on the origin at the initialization of the head pose session. In some embodiments, the world coordinate frame may be located where the device was booted up, or in the event that the head pose is lost during the boot session, it may be a new location. In some embodiments, the world coordinate frame may be the origin at the start of the head pose session.
[0219] In the illustrated embodiment, the world frame determination routine 80 is connected to the world surface determination routine 78 and determines the world coordinate frame 84 based on the location of the surface as determined by the world surface determination routine 78. The world frame memory command 82 is connected to the world frame determination routine 80 and receives the world coordinate frame 84 from the world frame determination routine 80. The world frame memory command 82 stores the world coordinate frame 84 as the world coordinate frame 86 within the origin and destination coordinate frame 34.
[0220] The head frame system 56 may include a head frame determination routine 90 and a head frame storage command 92. The head frame determination routine 90 may be connected to the motion tracking camera 46 and the inertial measurement unit 48. The head frame determination routine 90 may calculate a head coordinate frame 94 using data from the motion tracking camera 46 and the inertial measurement unit 48. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The motion tracking camera 46 may continuously capture images that are used by the head frame determination routine 90 to refine the head coordinate frame 94. The head unit 22 moves as the first user 14.1 in FIG. 8 moves his or her head. The motion tracking camera 46 and the inertial measurement unit 48 may continuously provide data to the head frame determination routine 90 so that the head frame determination routine 90 can update the head coordinate frame 94.
[0221] The head frame storage command 92 may be connected to the head frame determination routine 90 and may receive the head coordinate frame 94 from the head frame determination routine 90. The head frame storage command 92 may store the head coordinate frame 94 as the head coordinate frame 96 within the origin and destination coordinate frame 34. The head frame storage command 92 may repeatedly store the updated head coordinate frame 94 as the head coordinate frame 96 when the head frame determination routine 90 recalculates the head coordinate frame 94. In some embodiments, the head coordinate frame may be the location of the wearable XR device 12.1 relative to the local coordinate frame 72.
[0222] The camera frame system 58 may include camera intrinsics 98. The camera intrinsics 98 may include the dimensions of the head unit 22, which are characteristics of its design and manufacture. The camera intrinsics 98 may be used to calculate a camera coordinate frame 100 that is stored within the origin and destination coordinate frame 34.
[0223] In some embodiments, the camera coordinate frame 100 may include all the pupil positions of the left eye of the first user 14.1 in FIG. 8. When the left eye moves from left to right or up and down, the pupil position of the left eye is located within the camera coordinate frame 100. Additionally, the pupil position of the right eye is located within the camera coordinate frame 100 for the right eye. In some embodiments, the camera coordinate frame 100 may include the location of the camera relative to the local coordinate frame when an image is captured.
[0224] The origin / destination coordinate frame converter 36 may include a local / world coordinate converter 104, a world / head coordinate converter 106, and a head / camera coordinate converter 108. The local / world coordinate converter 104 may receive the local coordinate frame 72 and convert the local coordinate frame 72 to the world coordinate frame 86. The conversion of the local coordinate frame 72 to the world coordinate frame 86 may be represented as the local coordinate frame that is converted to the world coordinate frame 110 within the world coordinate frame 86.
[0225] The world / head coordinate converter 106 may convert from the world coordinate frame 86 to the head coordinate frame 96. The world / head coordinate converter 106 may convert the local coordinate frame that is converted to the world coordinate frame 110 to the head coordinate frame 96. The conversion may be represented as the local coordinate frame that is converted to the head coordinate frame 112 within the head coordinate frame 96.
[0226] The head / camera coordinate converter 108 may convert from the head coordinate frame 96 to the camera coordinate frame 100. The head / camera coordinate converter 108 may convert the local coordinate frame that is converted to the head coordinate frame 112 to the local coordinate frame that is converted to the camera coordinate frame 114 within the camera coordinate frame 100. The local coordinate frame that is converted to the camera coordinate frame 114 may be incorporated into the rendering engine 30. The rendering engine 30 may render the image data 68 representing the local content 28 based on the local coordinate frame that is converted to the camera coordinate frame 114.
[0227] FIG. 10 is a spatial representation of various origin and destination coordinate frames 34. Local coordinate frame 72, world coordinate frame 86, head coordinate frame 96, and camera coordinate frame 100 are shown in the figure. In some embodiments, the local coordinate frame associated with XR content 28 may have a position and rotation relative to the local and / or world coordinate frame and / or PCF when the virtual content is installed in the real world and thus can be viewed by the user (e.g., can provide a node and facing direction). Each camera may have its own camera coordinate frame 100 that encompasses all pupil positions of one eye. Reference numerals 104A and 106A represent the conversions performed by local / world coordinate converter 104, world / head coordinate converter 106, and head / camera coordinate converter 108 in FIG. 9, respectively.
[0228] FIG. 11 depicts a camera rendering protocol for converting from a head coordinate frame to a camera coordinate frame according to some embodiments. In the illustrated example, the pupil for one eye moves from position A to position B. Virtual objects that are intended to appear stationary will be projected onto a depth plane at one of two positions A or B depending on the position of the pupil (assuming the camera is configured to use a pupil-based coordinate frame). As a result, using a pupil coordinate frame that is converted to the head coordinate frame will introduce jitter into the stationary virtual object as the eye moves from position A to position B. This situation is referred to as view-dependent display or projection.
[0229] As depicted in FIG. 12, a camera coordinate frame (e.g., CR) is positioned to encompass all pupil positions, and the object projection will be consistent here regardless of pupil positions A and B. The head coordinate frame is transformed into the CR frame, which is referred to as a view-independent display or projection. Image reprojection may be applied to virtual content taking into account changes in eye position. However, since the rendering remains at the same position, jitter is minimized.
[0230] FIG. 13 illustrates the display system 42 in further detail. The display system 42 includes a stereoscopic analyzer 144 that is connected to the rendering engine 30 and forms part of the visual data and algorithms.
[0231] The display system 42 further includes left and right projectors 166A and 166B, and left and right light guides 170A and 170B. The left and right projectors 166A and 166B are connected to a power source. Each projector 166A and 166B has an individual input for image data to be provided to the individual projector 166A or 166B. When powered, the individual projector 166A or 166B generates light in a two-dimensional pattern and emits the light therefrom. The left and right light guides 170A and 170B are positioned to receive light from the left and right projectors 166A and 166B, respectively. The left and right light guides 170A and 170B are transparent light guides.
[0232] In use, the user mounts the head-mountable frame 40 on their head. The components of the head-mountable frame 40 may include, for example, a strap (not shown) that wraps around the perimeter of the back of the user's head. The left and right light guides 170A and 170B are then positioned in front of the user's left and right eyes 220A and 220B.
[0233] The rendering engine 30 takes in the image data it receives into the stereoscopic analyzer 144. The image data is the three-dimensional image data of the local content 28 in FIG. 8. The image data is projected onto a plurality of virtual planes. The stereoscopic analyzer 144 analyzes the image data and determines left and right image data sets based on the image data for projection onto each depth plane. The left and right image data sets are data sets that represent two-dimensional images that are projected in three dimensions and give the user a perception of depth.
[0234] The stereoscopic analyzer 144 takes in the left and right image data sets into the left and right projectors 166A and 166B. The left and right projectors 166A and 166B then create left and right light patterns. The components of the display system 42 are shown in a plan view, but it should be understood that the left and right patterns are two-dimensional patterns when shown in a front elevation view. Each light pattern includes a plurality of pixels. For illustrative purposes, light rays 224A and 226A from two of the pixels are shown exiting the left projector 166A and entering the left waveguide 170A. The light rays 224A and 226A are reflected from the side of the left waveguide 170A. Although the light rays 224A and 226A are shown propagating through internal reflection from left to right within the left waveguide 170A, it should be understood that the light rays 224A and 226A also propagate in a direction towards the plane of the paper using refractive and reflective systems.
[0235] The light rays 224A and 226A exit the left light waveguide 170A through the pupil 228A and then enter the left eye 220A through the pupil 230A of the left eye 220A. The light rays 224A and 226A then strike the retina 232A of the left eye 220A. Thus, the left light pattern strikes the retina 232A of the left eye 220A. The user is given the perception that the pixels formed on the retina 232A are pixels 234A and 236A that are at a certain distance on the side of the left waveguide 170A that the user faces the left eye 220A. Depth perception is created by manipulating the focal length of the light.
[0236] Similarly, the stereoscopic analyzer 144 captures the right image data set into the right projector 166B. The right projector 166B transmits a right light pattern, which is represented by pixels in the form of light rays 224B and 226B. The light rays 224B and 226B are reflected within the right waveguide 170B and exit through the pupil 228B. The light rays 224B and 226B then enter through the pupil 230B of the right eye 220B and strike the retina 232B of the right eye 220B. The pixels of the light rays 224B and 226B are perceived as pixels 134B and 236B behind the right waveguide 170B.
[0237] The patterns created on the retinas 232A and 232B are perceived individually as left and right images. The left and right images are slightly different from each other due to the function of the stereoscopic analyzer 144. The left and right images are perceived as a three-dimensional rendering within the user's brain.
[0238] As described, the left and right waveguides 170A and 170B are transparent. Light from real objects such as the table 16 on the sides of the left and right waveguides 170A and 170B facing the eyes 220A and 220B can be projected through the left and right waveguides 170A and 170B and strike the retinas 232A and 232B.
[0239] Persistent Coordinate Frame (PCF)
[0240] What is described herein are methods and apparatuses for providing spatial persistence across user instances in a shared space. Without spatial persistence, virtual content placed by a user within the physical world within a session may not exist within the view of the user in a different session or may be mis-placed. Without spatial persistence, virtual content placed by one user within the physical world may not exist within the view of a second user or may be out of place even when the second user intends to share the experience of the same physical space as the first user.
[0241] The inventors recognize and understand that spatial persistence can be provided through a persistent coordinate frame (PCF). The PCF may be defined based on one or more points representing features recognized in the physical world (e.g., corners, edges). The features may be selected such that they are likely to be the same from one user instance of the XR system to another.
[0242] Furthermore, drift during tracking, which can cause the calculated tracking path (e.g., camera trajectory) to deviate from the actual tracking path, can cause the location of virtual content to appear displaced when rendered relative to a local map based only on the tracking map. The tracking map for the space may be refined as the XR device collects additional information about the scene over time to correct for drift. However, if virtual content is placed on a physical object and stored relative to the world coordinate frame of the device derived from the tracking map before map refinement, the virtual content may appear displaced as if the physical object has moved during map refinement. The PCF may be updated according to map refinement since the PCF is defined based on features and updated as the features move during map refinement.
[0243] The PCF may have six degrees of freedom involving translation and rotation with respect to the map coordinate system. The PCF may be stored in a local and / or remote storage medium. The translation and rotation of the PCF may be calculated with respect to the map coordinate system, e.g., depending on the storage location. For example, the PCF used locally by a device may have translation and rotation with respect to the world coordinate frame of the device. The PCF in the cloud may have translation and rotation with respect to the standard coordinate frame of the standard map.
[0244] The PCF may provide a sparse representation of the physical world, providing less than all of the available information about the physical world, so that they can be efficiently processed and transferred. Techniques for processing persistent spatial information may create a dynamic map based on one or more coordinate systems within the physical space, across one or more sessions, and generate a persistent coordinate frame (PCF) across the sparse map that can be exposed to XR applications, for example, via an application programming interface (API).
[0245] FIG. 14 is a block diagram illustrating the creation of a persistent coordinate frame (PCF) and the association of XR content with the PCF, according to some embodiments. Each block may represent digital information stored in a computer memory. In the case of application 1180, the data may represent computer-executable instructions. In the case of virtual content 1170, the digital information may define a virtual object, for example, as defined by application 1180. In the case of the other boxes, the digital information may characterize some aspects of the physical world.
[0246] In the illustrated embodiments, one or more PCFs are created from images captured using sensors on a wearable device. In the embodiment of FIG. 14, the sensors are visual image cameras. These cameras may be the same cameras used to form a tracking map. Thus, some of the processing proposed by FIG. 14 may be implemented as part of the step of updating the tracking map. However, FIG. 14 illustrates that information providing persistence is generated in addition to the tracking map.
[0247] To derive the 3D PCF, two images 1110 from two cameras, mounted on a wearable device in a configuration enabling stereoscopic image analysis, are both processed. FIG. 14 illustrates image 1 and image 2, respectively derived from one of the cameras. A single image from each camera is illustrated for convenience. However, each camera may output a stream of image frames, and the processing illustrated in FIG. 14 may be performed for multiple image frames within the stream.
[0248] Thus, image 1 and image 2 may each be one frame within a sequence of image frames. The processing as depicted in FIG. 14 may be repeated on successive image frames within the sequence until an image frame containing feature points that form suitable images for deriving continuous spatial information from them is processed. Alternatively, or in addition, the processing of FIG. 14 may be repeated as the user moves such that the user is no longer close enough to the previously identified PCF and cannot reliably use that PCF to determine their position relative to the physical world. For example, an XR system may maintain the current PCF for the user. When the distance exceeds a threshold, the system may switch to a new current PCF closer to the user that may be generated according to the process of FIG. 14 using the image frames obtained at the user's current location.
[0249] Even when generating a single PCF, a stream of image frames may be processed to identify image frames depicting content within the physical world that are likely to be stable and easily identifiable by the device in the vicinity of the region of the physical world depicted in the image frame. In the embodiment of FIG. 14, this processing begins with the identification of features 1120 within the image. Features may be identified, for example, by finding locations of gradients within the image or other characteristics that exceed a threshold, which may correspond to corners of objects, for example. In the illustrated embodiment, the features are points, but other recognizable features such as edges may also be used, alternatively or in addition.
[0250] In the illustrated embodiment, a fixed number N of feature points 1120 are selected for further processing. Those feature points may be selected based on one or more criteria such as the magnitude of the gradient or proximity to other feature points. Alternatively, or in addition, the feature points may be selected heuristically, such as based on characteristics that suggest the feature points are persistent. For example, the heuristic may be defined based on characteristics of the feature points that are likely to correspond to corners of windows or doors or large furniture. Such a heuristic may consider the feature points themselves and what surrounds them. As a specific example, the number of feature points per image may be 100 - 500 or 150 - 250, such as 200.
[0251] Regardless of the number of feature points selected, descriptors 1130 may be calculated for the feature points. In this example, the descriptors are calculated for each selected feature point, but the descriptors may be calculated for groups of feature points, or subsets of feature points, or for all features in the image. The descriptors characterize the feature points such that feature points representing the same object in the physical world are assigned similar descriptors. The descriptors may facilitate the alignment of two frames such as may occur when one map is geolocated relative to another. Instead of searching for the relative orientation of the frames that minimizes the distance between feature points of two images, the initial alignment of the two frames may be done by identifying feature points with similar descriptors. The alignment of the image frames may be based on the step of aligning points with similar descriptors, which may involve less processing of calculating the alignment of all feature points in the image.
[0252] The descriptors may be calculated as a mapping of feature points to the descriptors, or in some embodiments, as a mapping of image patches around the feature points. The descriptors may be numerical quantities. U.S. Patent Application No. 16 / 190,948 describes the step of calculating descriptors for feature points and is incorporated herein by reference in its entirety.
[0253] In the embodiment of FIG. 14, descriptor 1130 is calculated for each feature point within each image frame. Based on the descriptor and / or the feature point and / or the image itself, the image frame may be identified as key frame 1140. In the illustrated embodiment, a key frame is an image frame that meets certain criteria and is then selected for further processing. When creating a tracking map, for example, an image frame that adds meaningful information to the map may be selected as a key frame to be integrated into the map. On the other hand, image frames that substantially overlap an area where the image frames have already been integrated into the map may be discarded so that they do not become key frames. Alternatively, or in addition, key frames may be selected based on the number and / or type of feature points within the image frame. In the embodiment of FIG. 14, key frame 1150 selected for inclusion in the tracking map may also be processed as a key frame for determining the PCF, although different or additional criteria for selecting a key frame for PCF generation may be used.
[0254] FIG. 14 shows that key frames are used for further processing, but the information obtained from the images may be processed in other forms. For example, feature points such as within a key rig may be processed alternatively or in addition. Further, although key frames are described as being derived from a single image frame, it is not necessary that there be a one-to-one relationship between the key frame and the image frame obtained. Key frames may be obtained from multiple image frames, for example, by stitching or aggregating the image frames together such that only features that appear in multiple images are retained within the key frame.
[0255] A keyframe may include image information and / or metadata associated with the image information. In some embodiments, an image captured by cameras 44, 46 (FIG. 9) may be calculated among one or more keyframes (e.g., keyframes 1, 2). In some embodiments, a keyframe may include a camera pose. In some embodiments, a keyframe may include one or more camera images captured in a camera pose. In some embodiments, the XR system determines that a portion of a camera image captured in a camera pose is not useful and thus may not include that portion within the keyframe. Thus, using keyframes to align new images with earlier knowledge of the scene reduces the use of the XR system's computational resources. In some embodiments, a keyframe may include an image and / or image data at a location with a certain direction / angle. In some embodiments, a keyframe may include a location and direction from which one or more map points can be observed. In some embodiments, a keyframe may include a coordinate frame with an ID. U.S. Patent Application No. 15 / 877,359 describes keyframes and is incorporated herein by reference in its entirety.
[0256] Some or all of the keyframes 1140 may be selected for further processing such as the generation of a persistent pose 1150 for the keyframe. The selection may be based on the characteristics of all or a subset of the feature points within the image frame. Those characteristics may be determined from processing descriptors, features, and / or the image frame itself. As a specific example, the selection may be based on a cluster of feature points identified as likely to be associated with a persistent object.
[0257] Each keyframe is associated with the pose of the camera at which the keyframe was obtained. For keyframes selected for processing into a persistent pose, the pose information may be stored along with other metadata about the keyframe such as a WiFi fingerprint and / or GPS coordinates at the time and / or location of acquisition.
[0258] A persistent pose is a source of information that a device can use to orient itself with respect to previously obtained information about the physical world. For example, if the keyframe from which the persistent pose was created is incorporated into a map of the physical world, the device can orient itself with respect to that persistent pose using a sufficient number of feature points within the keyframe that are associated with the persistent pose. The device can align the obtained current image of its surroundings with the persistent pose. This alignment may be based on a matching of the current image with the image 1110, features 1120, and / or descriptors 1130 that gave rise to the persistent pose, or any subset of that image or those features or descriptors. In some embodiments, the current image frame matched to the persistent pose may be another keyframe incorporated within the tracking map of the device.
[0259] The information about the persistent pose may be stored in a format that facilitates sharing among multiple applications that may be executed on the same or different devices. In the example of FIG. 14, some or all of the persistent poses may be reflected as a Persistent Coordinate Frame (PCF) 1160. Similar to a persistent pose, the PCF may also be associated with a map and may include a set of features or other information that a device can use to determine its orientation with respect to that PCF. The PCF may include a transformation that defines its transformation with respect to the origin of the map such that the device can determine its position with respect to any object within the physical world reflected in the map by correlating its position to the PCF.
[0260] Since the PCF provides a mechanism for determining the location of physical objects, an application such as application 1180 may define the position of a virtual object relative to one or more PCFs that serve as anchors for virtual content 1170. FIG. 14 illustrates, for example, that application 1 associates its virtual content 2 with PCF1 and 2. Similarly, application 2 associates its virtual content 3 with PCF1 and 2. Application 1 is also shown to associate its virtual content 1 with PCF4 and 5, and application 2 is shown to associate its virtual content 4 with PCF3. In some embodiments, similar to the methods based on image 1 and image 2 for PCF1 and 2, PCF3 may be based on image 3 (not shown), and PCF4 and 5 may be based on image 4 and image 5 (not shown). When rendering this virtual content, the device may apply one or more transformations and calculate information such as the location of the virtual content relative to the device's display and / or the location of the physical object relative to the desired location of the virtual content. Using the PCF as a reference can simplify such calculations.
[0261] In some embodiments, the persistent pose may be a coordinate location and / or orientation having one or more associated keyframes. In some embodiments, the persistent pose may be automatically created after the user has traveled a certain distance, e.g., 3 meters. In some embodiments, the persistent pose may act as a reference point during localization. In some embodiments, the persistent pose may be stored within a traversable world (e.g., traversable world module 538).
[0262] In some embodiments, the new PCF may be determined based on a predefined distance that is allowed between adjacent PCFs. In some embodiments, one or more persistent postures may be calculated into the PCF when the user travels a predetermined distance, such as 5 meters. In some embodiments, the PCF may be associated with one or more world coordinate frames and / or reference coordinate frames, for example, within a traversable world. In some embodiments, the PCF may be stored in a local and / or remote database, for example, depending on security settings.
[0263] FIG. 15 illustrates a method 4700 for establishing and using a persistent coordinate frame according to some embodiments. The method 4700 may begin with the step of (act 4702) using one or more sensors of an XR device to capture an image of a scene (e.g., image 1 and image 2 in FIG. 14). A plurality of cameras may be used, and one camera may generate a plurality of images, for example, in a stream.
[0264] Method 4700 may include step (4704) of extracting a point of interest (e.g., map point 702 in FIG. 7, feature 1120 in FIG. 14) from a captured image, step (act 4706) of generating a descriptor (e.g., descriptor 1130 in FIG. 14) regarding the extracted point of interest, and step (act 4708) of generating a keyframe (e.g., keyframe 1140) based on the descriptor. In some embodiments, the method may compare the points of interest within the keyframe and form pairs of keyframes that share a predetermined amount of points of interest. The method may use individual pairs of keyframes to reconstruct a part of the physical world. The mapped part of the physical world may be stored as 3D features (e.g., key rig 704 in FIG. 7). In some embodiments, selected portions of pairs of keyframes may be used to construct 3D features. In some embodiments, the results of the mapping may be selectively stored. Keyframes not used to construct 3D features may be associated with the 3D features through their poses, e.g., representing the distance between keyframes with a covariance matrix between the poses of the keyframes. In some embodiments, pairs of keyframes may be selected to construct 3D features such that the distance between each of the constructed 3D features is within a predetermined distance that can balance the amount of calculation required and the level of accuracy of the resulting model. Such an approach makes it possible to provide a model of the physical world with an amount of data suitable for efficient and accurate calculations using an XR system. In some embodiments, the covariance matrix of two images may include the covariance between the poses (e.g., 6 degrees of freedom) of the two images.
[0265] Method 4700 may include a step (act 4710) of generating a persistent pose based on keyframes. In some embodiments, the method may include a step of generating a persistent pose based on 3D features reconstructed from a pair of keyframes. In some embodiments, a persistent pose may be associated with 3D features. In some embodiments, a persistent pose may include the poses of the keyframes used to construct the 3D features. In some embodiments, a persistent pose may include the average pose of the keyframes used to construct the 3D features. In some embodiments, a persistent pose may be generated such that the distance between neighboring persistent poses is within a predetermined value, for example, in the range of 1 meter to 5 meters, any value therebetween, or any other suitable value. In some embodiments, the distance between neighboring persistent poses may be represented by the covariance matrix of the neighboring persistent poses.
[0266] Method 4700 may include a step (act 4712) of generating a PCF based on the persistent pose. In some embodiments, the PCF may be associated with 3D features. In some embodiments, the PCF may be associated with one or more persistent poses. In some embodiments, the PCF may include the pose of one of the associated persistent poses. In some embodiments, the PCF may include the average pose of the poses of the associated persistent poses. In some embodiments, the PCF may be generated such that the distance between neighboring PCFs is within a predetermined value, for example, in the range of 3 meters to 10 meters, any value therebetween, or any other suitable value. In some embodiments, the distance between neighboring PCFs may be represented by the covariance matrix of the neighboring PCFs. In some embodiments, the PCF may be exposed to an XR application, for example, via an application programming interface (API), such that the XR application can access the model of the physical world through the PCF without accessing the model itself.
[0267] Method 4700 may include an act (act 4714) of associating at least one of image data of a virtual object and a PCF for display by an XR device. In some embodiments, the method may include an act of calculating a translation and orientation of a virtual object with respect to an associated PCF. It should be understood that it is not necessary to associate a virtual object with a PCF generated by a device on which the virtual object is installed. For example, the device may read a stored PCF in a reference map in the cloud and associate the virtual object with the read PCF. It should be understood that a virtual object may move with an associated PCF as the PCF is adjusted over time.
[0268] FIG. 16 illustrates a first XR device 12.1, visual data and algorithms of a second XR device 12.2, and a server 20, according to some embodiments. The components illustrated in FIG. 16 may be operative to perform some or all of the operations associated with the acts of generating, updating, and / or using spatial information, such as a persistent pose, a persistent coordinate frame, a tracking map, or a reference map, as described herein. Although not shown, the first XR device 12.1 may be configured identically to the second XR device 12.2. The server 20 may include a map storage routine 118, a reference map 120, a map transmitter 122, and a map merge algorithm 124.
[0269] A second XR device 12.2 that may be in the same scene as the first XR device 12.1 may include a Persistent Coordinate Frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that may be used to render virtual objects, and a frame embedding generator 308 (see FIG. 21). In some embodiments, the map download system 126, the PCF identification system 128, the map 2, the location module 130, the canonical map incorporator 132, the canonical map 133, and the map issuer 136 may be grouped within the passable world unit 1304. The PCF integration unit 1300 may be connected to the passable world unit 1304 and other components of the second XR device 12.2 to enable reading, generating, using, uploading, and downloading of the PCF.
[0270] A map with a PCF may enable more persistence within a changing world. In some embodiments, for example, the step of locating a tracking map, including matching features for an image, may include the step of selecting features representing persistent content from a map configured by the PCF, which enables fast matching and / or location. For example, a world where people move in and out of a scene and objects such as doors move relative to the scene requires less memory space and transmission rate and enables the use of individual PCFs and their relationships to each other (e.g., the integrated constellation of PCFs) to map the scene.
[0271] In some embodiments, the PCF integration unit 1300 may include a PCF 1306 previously stored in data storage on the memory unit of the second XR device 12.2, a PCF tracker 1308, a persistent pose acquirer 1310, a PCF verifier 1312, a PCF generation system 1314, a coordinate frame computer 1316, a persistent pose computer 1318, a tracking map and persistent pose converter 1320, a persistent pose and PCF converter 1322, and a PCF and image data converter 1324, including three converters.
[0272] In some embodiments, the PCF tracker 1308 may have on-prompts and off-prompts that are selectable by the application 1302. The application 1302 may be executable by a processor of the second XR device 12.2 and may, for example, display virtual content. The application 1302 may have a call to switch on the PCF tracker 1308 via an on-prompt. The PCF tracker 1308 may generate a PCF when the PCF tracker 1308 is switched on. The application 1302 may have a subsequent call to switch off the PCF tracker 1308 via an off-prompt. The PCF tracker 1308 terminates PCF generation when the PCF tracker 1308 is switched off.
[0273] In some embodiments, the server 20 may include a plurality of persistent postures 1332 and a plurality of PCFs 1330 that are associated with and previously stored with the reference map 120. The map transmitter 122 may transmit the reference map 120 to the second XR device 12.2 together with the persistent postures 1332 and / or the PCFs 1330. The persistent postures 1332 and the PCFs 1330 may be stored on the second XR device 12.2 in association with the reference map 133. When map 2 is located relative to the reference map 133, the persistent postures 1332 and the PCFs 1330 may be stored in association with map 2.
[0274] In some embodiments, the persistent posture acquirer 1310 may acquire a persistent posture for map 2. The PCF verifier 1312 may be connected to the persistent posture acquirer 1310. The PCF verifier 1312 may read a PCF from the PCF 1306 based on the persistent posture read by the persistent posture acquirer 1310. The PCF read by the PCF verifier 1312 may form an initial group of PCFs for use in an image display based on the PCF.
[0275] In some embodiments, application 1302 may require that an additional PCF be generated. For example, when the user moves to an area that has not been previously mapped, application 1302 may switch on the PCF tracker 1308. The PCF generation system 1314 is connected to the PCF tracker 1308 and may start generating the PCF based on map 2 as map 2 begins to expand. The PCF generated by the PCF generation system 1314 may form a second group of PCFs that can be used for PCF-based image display.
[0276] The coordinate frame computer 1316 may be connected to the PCF verifier 1312. After the PCF verifier 1312 reads the PCF, the coordinate frame computer 1316 may call the head coordinate frame 96 and determine the head pose of the second XR device 12.2. The coordinate frame computer 1316 may also call the persistent pose computer 1318. The persistent pose computer 1318 may be directly or indirectly connected to the frame embedding generator 308. In some embodiments, an image / frame may be designated as a keyframe after a threshold distance, e.g., 3 meters, has been traveled from the previous keyframe. The persistent pose computer 1318 may generate a persistent pose based on a plurality of, e.g., three keyframes. In some embodiments, the persistent pose may essentially be the average of the coordinate frames of a plurality of keyframes.
[0277] The tracking map and persistent pose converter 1320 may be connected to map 2 and the persistent pose computer 1318. The tracking map and persistent pose converter 1320 may convert map 2 to a persistent pose and determine the persistent pose at the origin with respect to map 2.
[0278] The persistent pose and PCF converter 1322 may be connected to the tracking map and persistent pose converter 1320 and further to the PCF verifier 1312 and the PCF generation system 1314. The persistent pose and PCF converter 1322 may convert the persistent pose (for which the tracking map has been converted) into a PCF from the PCF verifier 1312 and the PCF generation system 1314 and determine the PCF for the persistent pose.
[0279] The PCF and image data converter 1324 may be connected to the persistent pose and PCF converter 1322 and the data channel 62. The PCF and image data converter 1324 converts the PCF into image data 68. The rendering engine 30 may be connected to the PCF and image data converter 1324 and display the image data 68 for the PCF to the user.
[0280] The PCF integration unit 1300 may store additional PCFs generated using the PCF generation system 1314 within the PCF 1306. The PCF 1306 may be stored for the persistent pose. The map publisher 136 may read the PCF 1306 and the persistent pose associated with the PCF 1306 when the map publisher 136 transmits the map 2 to the server 20 and the map publisher 136 also transmits the PCF and the persistent pose associated with the map 2 to the server 20. When the map storage routine 118 of the server 20 stores the map 2, the map storage routine 118 may also store the persistent pose and the PCF generated by the second visual device 12.2. The map merge algorithm 124 may create the reference map 120 together with the persistent pose and the PCF of the map 2, which are respectively associated with the reference map 120 and stored within the persistent pose 1332 and the PCF 1330.
[0281] The first XR device 12.1 may include a PCF integration unit similar to the PCF integration unit 1300 of the second XR device 12.2. When the map transmitter 122 transmits the reference map 120 to the first XR device 12.1, the map transmitter 122 may transmit the persistent pose 1332 and the PCF 1330, which are associated with the reference map 120 and originate from the second XR device 12.2. The first XR device 12.1 may store the PCF and the persistent pose in a data storage device on the storage device of the first XR device 12.1. The first XR device 12.1 may then utilize the persistent pose and the PCF originating from the second XR device 12.2 for image display with respect to the PCF. Additionally, or alternatively, the first XR device 12.1 may read, generate, utilize, upload, and download the PCF and the persistent pose in a manner similar to the second XR device 12.2 as described above.
[0282] In the illustrated embodiment, the first XR device 12.1 generates a local tracking map (hereinafter referred to as "Map 1"), and the map storage routine 118 receives Map 1 from the first XR device 12.1. The map storage routine 118 then stores Map 1 as the reference map 120 on the storage device of the server 20.
[0283] The second XR device 12.2 includes a map download system 126, an anchor identification system 128, a location determination module 130, a reference map incorporator 132, a local content positioning system 134, and a map publisher 136.
[0284] In use, the map transmitter 122 transmits the reference map 120 to the second XR device 12.2, and the map download system 126 downloads and stores the reference map 120 from the server 20 as the reference map 133.
[0285] The anchor identification system 128 is connected to the world surface determination routine 78. The anchor identification system 128 identifies an anchor based on an object detected by the world surface determination routine 78. The anchor identification system 128 uses the anchor to generate a second map (Map 2). As indicated by cycle 138, the anchor identification system 128 continues to identify the anchor and update Map 2. The location of the anchor is recorded as three-dimensional data based on the data provided by the world surface determination routine 78. The world surface determination routine 78 receives an image from the real object detection camera 44 and depth data from the depth sensor 135, and determines the location of the surface and its relative distance from the depth sensor 135.
[0286] The localization module 130 is connected to the reference map 133 and Map 2. The localization module 130 repeatedly attempts to localize Map 2 relative to the reference map 133. The reference map incorporator 132 is connected to the reference map 133 and Map 2. When the localization module 130 localizes Map 2 relative to the reference map 133, the reference map incorporator 132 incorporates the reference map 133 into the anchors of Map 2. Map 2 is then updated with the missing data contained within the reference map.
[0287] The local content positioning system 134 is connected to Map 2. The local content positioning system 134 may be, for example, a system by which a user can localize local content at a specific location within the world coordinate frame. The local content itself is then associated with one of the anchors of Map 2. The local / world coordinate converter 104 converts the local coordinate frame to the world coordinate frame based on the settings of the local content positioning system 134. The functions of the rendering engine 30, the display system 42, and the data channel 62 are described with reference to FIG. 2.
[0288] The map issuer 136 uploads map 2 to the server 20. The map storage routine 118 of the server 20 then stores map 2 in the storage medium of the server 20.
[0289] The map merge algorithm 124 merges map 2 and the reference map 120. When more than two maps, for example, three or four maps, related to the same or adjacent areas of the physical world are stored, the map merge algorithm 124 merges all the maps into the reference map 120 and renders a new reference map 120. The map transmitter 122 then transmits the new reference map 120 to any devices 12.1 and 12.2 within the area represented by the new reference map 120. When devices 12.1 and 12.2 locate their individual maps relative to the reference map 120, the reference map 120 becomes the promoted map.
[0290] FIG. 17 illustrates an example of generating keyframes for a map of a scene according to some embodiments. In the illustrated example, the first keyframe KF1 is generated for the door on the left wall of the room. The second keyframe KF2 is generated for the area within the corner where the floor, left wall, and right wall of the room meet. The third keyframe KF3 is generated for the area of the window on the right wall of the room. The fourth keyframe KF4 is generated for the area at the edge of the rug on the floor of the wall. The fifth keyframe KF5 is generated for the area of the rug closest to the user.
[0291] FIG. 18 illustrates an example of generating a persistent pose for the map of FIG. 17 according to some embodiments. In some embodiments, a new persistent pose is created when the device measures a threshold distance traveled and / or when the application requests a new persistent pose (PP). In some embodiments, the threshold distance may be 3 meters, 5 meters, 20 meters, or any other suitable distance. Selecting a smaller threshold distance (e.g., 1m) may result in an increase in computational load because a larger number of PPs can be created and managed compared to a larger threshold distance. Selecting a larger threshold distance (e.g., 40m) results in fewer PPs being created and fewer PCFs being created, meaning that the virtual content associated with the PCF is relatively far (e.g., 30m) from the PCF, and the error may increase as the distance from the PCF to the virtual content increases, which may result in an increase in virtual content installation error.
[0292] In some embodiments, the PP may be created at the start of a new session. This initial PP may be considered zero and can be visualized as the center of a circle having a radius equal to the threshold distance. When the device reaches the perimeter of the circle and, in some embodiments, when the application requests a new PP, the new PP may be placed at the current location of the device (threshold distance). In some embodiments, a new PP will not be created if the device is able to find an existing PP within the threshold distance from the new location of the device. In some embodiments, when a new PP (PP1150 in FIG. 14) is created, the device associates one or more of the nearest keyframes with the PP. In some embodiments, the location of the PP with respect to the keyframe may be based on the location of the device at the time the PP was created. In some embodiments, the PP will not be created unless the application requests the PP, even if the device travels the threshold distance.
[0293] In some embodiments, the application may request the PCF from the device when the application has virtual content to display to the user. The PCF request from the application may trigger a PP request, and a new PP will be created after the device has advanced the threshold distance. FIG. 18 illustrates, for example, a first persistent pose PP1 that can associate the closest key frames (e.g., KF1, KF2, and KF3) by calculating the relative pose between a key frame and a persistent pose. FIG. 18 also illustrates a second persistent pose PP2 that can associate the closest key frames (e.g., KF4 and KF5).
[0294] FIG. 19 illustrates an example of generating a PCF for the map of FIG. 17 according to some embodiments. In the illustrated example, PCF1 may include PP1 and PP2. As described above, the PCF may be used to display image data for the PCF. In some embodiments, each PCF may have coordinates in a different coordinate frame (e.g., a world coordinate frame) and, for example, a PCF descriptor that uniquely identifies the PCF. In some embodiments, the PCF descriptor may be calculated based on the feature descriptors of the features in the frame associated with the PCF. In some embodiments, the various constellations of the PCF may be combined in a persistent manner that requires less data and less data transmission and may represent the real world.
[0295] FIGS. 20A-20C are schematic diagrams illustrating examples of establishing and using a persistent coordinate frame. FIG. 20A shows two users 4802A, 4802B with individual local tracking maps 4804A, 4804B that are not located with respect to a reference map. The origins 4806A, 4806B for the individual users are depicted by a coordinate system (e.g., a world coordinate system) within their individual areas. These origins of each tracking map can be local to each user because the origin depends on the orientation of that individual device when tracking is initiated.
[0296] As the sensors of the user device scan the environment, the device may capture images that contain features representing persistent objects such that those images can be classified as key frames from which persistent poses can be created, as described above in connection with FIG. 14. In this example, the tracking map 4802A includes a persistent pose (PP) 4808A, and the tracking map 4802B includes a PP 4808B.
[0297] Also, as described above in connection with FIG. 14, some of the PPs may be classified as PCFs that are used to determine the orientation of virtual content for rendering it to the user. FIG. 20B shows that XR devices worn by individual users 4802A, 4802B may create local PCFs 4810A, 4810B based on the PPs 4808A, 4808B. FIG. 20C shows that persistent content 4812A, 4812B (e.g., virtual content) may be associated with the PCFs 4810A, 4810B by individual XR devices.
[0298] In this example, the virtual content may have a virtual content coordinate frame that can be used by the application that generates the virtual content, regardless of how the virtual content is to be displayed. The virtual content may be defined, for example, as a surface such as a triangle of a mesh at a particular location and angle with respect to the virtual content coordinate frame. To render that virtual content to the user, the location of those surfaces may be determined with respect to the user who is to perceive the virtual content.
[0299] Associating virtual content with a PCF can simplify the calculations involved in determining the location of the virtual content for a user. The location of the virtual content for a user may be determined by applying a series of transformations. Some of those transformations may vary and may be updated frequently. Others of those transformations may be stable and may not be updated very frequently or at all. Nevertheless, the transformations can be applied with a relatively low computational burden such that the location of the virtual content can be updated frequently for the user and provide a rendered virtual content with a realistic appearance.
[0300] In the embodiments of FIGS. 20A - 20C, the device of user 1 has a coordinate system that may be related to a coordinate system that defines the origin of the map by the transformation rig1_T_w1. The device of user 2 has a similar transformation rig2_T_w2. These transformations are represented as 6 - degree transformations and may define translations and rotations to align the device coordinate system with the map coordinate system. In some embodiments, the transformation may be represented as two separate transformations, one defining a translation and the other defining a rotation. It should be understood, therefore, that the transformation can be represented in a form that simplifies the calculations or otherwise provides an advantage.
[0301] The transformation between the origin of the tracking map and the PCF identified by an individual user device is represented as pcf1_T_w1 and pcf2_T_w2. In this example, the PCF and the PP are the same such that the same transformation also characterizes the PP.
[0302] The location of the user device with respect to the PCF can thus be calculated by successive application of these transformations such as rig1_T_pcf1=(rig1_T_w1) * (pcf1_T_w1).
[0303] As shown in FIG. 20C, the virtual content is located with respect to the PCF using the transformation of obj1_T_pcf1. This transformation may be set by an application that generates virtual content, which may receive information from a world reconstruction system that describes physical objects with respect to the PCF. To render the virtual content to the user, a transformation to the coordinate system of the user's device is calculated, which is the transformation obj1_t_w1=(obj1_T_pcf1) * (pcf1_T_w1) can be calculated by relating the virtual content coordinate frame to the origin of the tracking map. That transformation can then be related to the user's device through a further transformation rig1_T_w1.
[0304] The location of the virtual content can change based on the output from the application that generates the virtual content. When it changes, an end-to-end transformation from the source coordinate system to the destination coordinate system can be recalculated. Additionally, the user's location and / or head pose can also change as the user moves. As a result, any end-to-end transformation that depends on the user's location or head pose, such as the transformation rig1_T_w1, will also change in a similar manner.
[0305] The transformation rig1_T_w1 may be updated as the user moves based on tracking the user's position relative to stationary objects in the physical world. Such tracking may be performed by a headset tracking component that processes a sequence of images, or other components of the system, as described above. Such an update may be performed by determining the user's pose with respect to a stationary reference frame such as the PP.
[0306] In some embodiments, the location and orientation of the user device may be determined relative to the nearest persistent pose, or in this example, the PCF as PP is used as the PCF. Such determination may be made by identifying feature points that characterize the PP in the current image captured using sensors on the device. Image processing techniques such as stereoscopic image analysis may be used to determine the location of the device relative to those feature points. From this data, the system may calculate the change in the transformation associated with the user's movement based on the relationship rig1_T_pcf1=(rig1_T_w1) * (pcf1_T_w1).
[0307] The system may determine and apply the transformation in an order that is computationally efficient. For example, the need to calculate rig1_T_w1 from the measurements that yield rig1_T_pcf1 may be avoided by both tracking the user's pose and defining the location of the virtual content relative to the PP or PCF constructed on the persistent pose. Thus, the transformation from the source coordinate system of the virtual content to the destination coordinate system of the user's device may be based on the measured transformation according to the expression (rig1_T_pcf1) * (obj1_t_pcf1), where the first transformation is measured by the system and the latter transformation is supplied by an application that defines the virtual content for rendering. In embodiments where the virtual content is positioned relative to the origin of the map, the end-to-end transformation may relate the virtual object coordinate system to the PCF coordinate system based on an additional transformation between the map coordinates and the PCF coordinates. In embodiments where the virtual content is positioned relative to a PP or PCF that is different from the one relative to which the user position is being tracked, a transformation between the two may be applied. Such a transformation may be fixed, for example, determined from the map where both appear.
[0308] The transformation-based approach may be implemented within a device with components, for example, that process sensor data and build a tracking map. As part of that process, those components may identify feature points that can be used as persistent poses, which may in turn be transformed into PCFs. Those components limit the number of persistent poses generated for the map and provide a suitable spacing between persistent poses, as described above in connection with FIGS. 17-19, while allowing the user to be close enough to a persistent pose location, regardless of the location within the physical environment, to accurately calculate the user's pose. As the persistent pose closest to the user is updated as a result of user movement, refinement to the tracking map, or other causes, any of the transformations used to calculate the location of virtual content for the user, which depends on the location of the PP (or PCF if used), may be updated and stored for use until at least the user moves away from that persistent pose. Note that by calculating and storing the transformations, the computational burden of calculating how often the location of virtual content is updated can be relatively low such that it can be performed with a relatively short latency.
[0309] FIGS. 20A-20C illustrate the positioning relative to a tracking map, where each device has its own tracking map. However, the transformations may be generated for any map coordinate system. The persistence of content across a user session of an XR system can be achieved by using a persistent map. The shared experience of the user may also be facilitated by using a map to which multiple user devices may be oriented.
[0310] In some embodiments, described in more detail below, the location of virtual content may be defined in relation to coordinates in a canonical map that is formatted so that any of a plurality of devices may use the map. Each device may maintain a tracking map and may determine changes in the user's pose relative to the tracking map. In this example, the transformation between the tracking map and the canonical map may be determined through a "localization" process, which may be performed by matching structures (such as one or more persistent poses) in the tracking map with one or more structures (such as one or more PCFs) in the canonical map.
[0311] What is further described below are techniques for creating and using a canonical map in this way.
[0312] Deep keyframe
[0313] The techniques as described herein rely on the comparison of image frames. For example, to establish the position of a device relative to a tracking map, a new image may be captured using sensors worn by the user, and the XR system may search for an image that shares at least a predetermined amount of points of interest with the new image within the set of images used to create the tracking map. As an example of another scenario involving the comparison of image frames, a tracking map may be localized with respect to a canonical map by first finding an image frame associated with a persistent pose in the tracking map that is similar to the image frame associated with a PCF in the canonical map. Alternatively, the transformation between two canonical maps may be calculated by first finding similar image frames in the two maps.
[0314] The deep keyframe provides a way to reduce the amount of processing required to identify similar image frames. For example, in some embodiments, the comparison can be made between image features (e.g., "2D features") in a new 2D image and 3D features in a map. Such a comparison can be made in any suitable way, such as by projecting the 3D image into a 2D plane. Conventional methods, such as the bag of words (BoW), search for the 2D features of a new image in a database that includes all 2D features in the map, which can require significant computational resources, especially when the map represents a large area. The conventional method then locates an image that shares at least one of the 2D features with the new image, which may include images that are not useful for locating significant 3D features in the map. The conventional method then locates 3D features that are not significant for the 2D features in the new image.
[0315] The inventors recognize and understand a technique for reading images in a map that uses fewer memory resources (e.g., one quarter of the memory resources used by BoW), higher efficiency (e.g., 2.5 ms processing time per keyframe, 100 μs for comparison of 500 keyframes), and higher accuracy (e.g., 20% better recall than BoW for a 1,024-dimensional model, 5% better recall than BoW for a 256-dimensional model).
[0316] Descriptors that can be used to compare an image frame with other image frames to reduce calculations may be calculated for the image frame. The descriptors may be stored instead of, or in addition to, the image frame and feature points. In a map where persistent poses and / or PCFs can be generated from an image frame, the descriptors of the image frame or frames from which each persistent pose or PCF was generated may be stored as part of the persistent pose and / or PCF.
[0317] In some embodiments, the descriptor may be calculated as a function of feature points within an image frame. In some embodiments, a neural network is configured to calculate a unique frame descriptor for representing an image. The image may have a resolution higher than 1 megabyte such that sufficient details of a 3D environment within the field of view of a device worn by a user are captured within the image. The frame descriptor can be much smaller, such as a sequence of numbers, for example, within the range of 128 bytes to 512 bytes or any number in between.
[0318] In some embodiments, the neural network is trained such that the calculated frame descriptor indicates similarity between images. Images within a map can be located by identifying the closest images in a database comprising the images used to generate the map that have frame descriptors within a predetermined distance of the frame descriptor for the new image. In some embodiments, the distance between images may be represented by the difference between the frame descriptors of two images.
[0319] FIG. 21 is a block diagram illustrating a system for generating descriptors for individual images, according to some embodiments. In the illustrated example, a frame embedding generator 308 is shown. The frame embedding generator 308 may be used in conjunction with the server 20 in some embodiments, but alternatively, or in addition, may be executed in whole or in part within one of the XR devices 12.1 and 12.2, or any other device that processes images for comparison with other images.
[0320] In some embodiments, the frame embedding generator may be configured to generate a data representation of an image that is reduced from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes) that still represents the content within the image, despite the reduced size. In some embodiments, the frame embedding generator may be used to generate a data representation for an image, which may be a keyframe or frame used in other methods. In some embodiments, the frame embedding generator 308 may be configured to convert an image at a specific location and orientation into a unique sequence of numbers (e.g., 256 bytes). In the illustrated embodiment, the image 320 captured by the XR device may be processed by the feature extractor 324 to detect the point of interest 322 within the image 320. The point of interest may or may not be derived from the identified feature points as described above with respect to the feature 1120 (FIG. 14) or as otherwise described herein. In some embodiments, the point of interest may be represented by a descriptor as described above with respect to the descriptor 1130 (FIG. 14), which may be generated using a deep sparse feature method. In some embodiments, each point of interest 322 may be represented by a sequence of numbers (e.g., 32 bytes). For example, there may be n features (e.g., 100), and each feature may be represented by a 32-byte sequence.
[0321] In some embodiments, the frame embedding generator 308 may include a neural network 326. The neural network 326 may include a multi-layer perceptron unit 312 and a max pooling unit 314. In some embodiments, the multi-layer perceptron (MLP) unit 312 may comprise a multi-layer perceptron, which may be trained. In some embodiments, the point of interest 322 (e.g., the descriptor for the point of interest) may be reduced by the multi-layer perceptron 312 and output as a weighted combination 310 of the descriptors. For example, the MLP may reduce n features to m features, where m is less than n features.
[0322] In some embodiments, the MLP unit 312 may be configured to perform matrix multiplication. The multi-layer perceptron unit 312 receives a plurality of points of interest 322 of the image 320 and converts each point of interest into a column of individual numbers (e.g., 256). For example, there may be 100 features, and each feature may be represented by a column of 256 numbers. The matrix may be created to have 100 horizontal rows and 256 vertical columns in this example. Each row may have a series of 256 numbers that vary in size, some being smaller and some being larger. In some embodiments, the output of the MLP may be an n×256 matrix, where n represents the number of features extracted from the image. In some embodiments, the output of the MLP may be an m×256 matrix, where m is the number of points of interest reduced from n.
[0323] In some embodiments, the MLP 312 may have a training phase and a usage phase during which model parameters for the MLP are determined. In some embodiments, the MLP may be trained as illustrated in FIG. 25. The input training data may comprise data in three sets, the three sets comprising 1) a query image, 2) a positive sample, and 3) a negative sample. The query image may be regarded as a reference image.
[0324] In some embodiments, the positive sample may comprise an image that is similar to the query image. For example, in some embodiments, being similar means having the same object in both the query and positive sample images, but viewable from different angles. In some embodiments, being similar means having the same object in both the query and positive sample images, but may have an object that is offset with respect to other images (e.g., to the left, right, up, down).
[0325] In some embodiments, the negative sample may comprise an image that does not resemble the query image. For example, in some embodiments, the dissimilar image may not contain any object that is prominent within the query image, or may contain only a small portion (e.g., < 10%, 1%) of the prominent objects within the query image. In contrast, the similar image may, for example, have a majority (e.g., > 50%, or > 75%) of the objects within the query image.
[0326] In some embodiments, the point of interest may be extracted from an image within the input training data and may be transformed into a feature descriptor. These descriptors may be calculated for both the training images, as shown in FIG. 25, and for the features extracted during the operation of the frame embedding generator 308 of FIG. 21. In some embodiments, a deep sparse feature (DSF) process may be used to generate descriptors (e.g., DSF descriptors) as described in U.S. Patent Application No. 16 / 190,948. In some embodiments, the DSF descriptor is of n×32 dimensions. The descriptor may then be passed through a model / MLP to create a 256-byte output. In some embodiments, the model / MLP may have the same structure as MLP312 such that once the model parameters are set through training, the resulting trained MLP can be used as MLP312.
[0327] In some embodiments, the feature descriptor (e.g., 256 bytes output from an MLP model) may then be sent to a triplet margin loss module (which is only used during the training phase of the MLP neural network and cannot be used during the usage phase). In some embodiments, the triplet margin loss module is configured to select parameters for the model such that it reduces the difference between the 256 bytes output from the query image and the 256 bytes output from the positive sample, and increases the difference between the 256 bytes output from the query image and the 256 bytes output from the negative sample. In some embodiments, the training phase may include the step of feeding a plurality of triplet input images into the learning process and determining the model parameters. This training process may continue, for example, until the difference regarding the positive image is minimized and the difference regarding the negative image is maximized, or until another suitable termination criterion is reached.
[0328] Referring back to FIG. 21, the frame embedding generator 308 may here include a pooling layer, illustrated as a max pooling unit 314. The max pooling unit 314 may analyze each column and determine the maximum number within an individual column. The max pooling unit 314 may combine the maximum value of each column of the output matrix of the MLP 312 into a global feature column 316 of a number, e.g., 256. It should be understood that the images processed within the XR system may desirably have high-resolution frames with potentially millions of pixels. The global feature column 316 occupies relatively little memory and is relatively small, making it easily searchable compared to an image (e.g., with a resolution higher than 1 megabyte). Thus, it is possible to search for an image without analyzing each original frame from the camera, and it is also less expensive to store 256 bytes instead of the full frame.
[0329] Figure 22 is a flowchart illustrating a method 2200 for calculating image descriptors according to some embodiments. The method 2200 may begin with receiving a plurality of images captured by an XR device worn by a user (act 2202). In some embodiments, the method 2200 may include determining one or more keyframes from the plurality of images (act 2204). In some embodiments, act 2204 may be skipped and / or may occur after step 2210 instead.
[0330] The method 2200 may include identifying one or more points of interest in the plurality of images using an artificial neural network (act 2206), and calculating a feature descriptor for each individual point of interest using the artificial neural network (act 2208). The method may include, for each image, calculating a frame descriptor for representing the image, at least in part, based on the calculated feature descriptors for the identified points of interest in the image using an artificial neural network (act 2210).
[0331] Figure 23 is a flowchart illustrating a method 2300 for location identification using image descriptors according to some embodiments. In this example, a new image frame depicting the current location of the XR device may be compared with image frames stored in relation to points in the map (such as persistent poses or PCFs as described above). Method 2300 may begin with the step of receiving a new image captured by an XR device worn by a user (act 2302). Method 2300 may include the step of identifying one or more closest key frames in a database that comprise key frames used to generate one or more maps (act 2304). In some embodiments, the closest key frames may be identified based on approximate spatial information and / or previously determined spatial information. For example, the approximate spatial information may indicate that the XR device is within a geographical area represented by a 50m x 50m area of the map. Image matching may be performed only with respect to points within that area. As another example, based on tracking, the XR system may know that the XR device was previously close to a first persistent pose in the map and is moving in the direction of a second persistent pose in the map. That second persistent pose may be considered the closest persistent pose, and the key frame stored with it may be considered the closest key frame. Alternatively, or in addition, other metadata such as GPS data or WiFi fingerprints may also be used to select the closest key frame or set of closest key frames.
[0332] Regardless of how the nearest keyframe is selected, the frame descriptor may be used to determine whether it matches any of the frames selected such that the new image is associated with a neighboring persistent pose. The determination may be made by comparing the frame descriptor of the new image with the frame descriptor of the nearest keyframe or subset of keyframes in a database selected in any other suitable manner, and selecting a keyframe with a frame descriptor within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors may be calculated by obtaining the difference between two columns of numbers that may represent the two frame descriptors. In embodiments where the columns are treated as columns of multiple quantities, the difference may be calculated as a vector difference.
[0333] Once a matching image frame is identified, the orientation of the XR device with respect to that image frame may be determined. Method 2300 may include performing feature matching (act 2306) on 3D features in the map corresponding to the identified nearest keyframe, and calculating the pose of the device worn by the user (act 2308) based on the feature matching results. Thus, the matching that is intensive in the calculation of feature points in two images may be performed for only one image that has already been determined to be a likely match for the new image.
[0334] FIG. 24 is a flowchart illustrating a method 2400 for training a neural network according to some embodiments. Method 2400 may begin with generating a dataset (act 2402) comprising a plurality of image sets. Each of the plurality of image sets may include a query image, a positive sample image, and a negative sample image. In some embodiments, the plurality of image sets may include synthetic recording pairs configured to teach a neural network basic information such as shape. In some embodiments, the plurality of image sets may include real recording pairs that may be recorded from the physical world.
[0335] In some embodiments, the inliers may be calculated by fitting a fundamental matrix between two images. In some embodiments, the sparse overlap may be calculated as the intersection over union (IoU) on the union of the keypoints seen in both images. In some embodiments, the positive samples may include at least 20 keypoints that are the same within the query image and serve as inliers. The negative samples may include less than 10 inlier points. The negative samples may have less than half sparse points that overlap with the analysis points of the query image.
[0336] Method 2400 may include an act (act 2404) of calculating a loss by comparing a query image with positive sample images and negative sample images for each image set. Method 2400 may include an act (act 2406) of modifying the artificial neural network based on the calculated loss such that the distance between the frame descriptor generated by the artificial neural network for the query image and the frame descriptor for the positive sample image is less than the distance between the frame descriptor for the query image and the frame descriptor for the negative sample image.
[0337] A method and apparatus configured to generate global descriptors for individual images are described above, but it should be understood that the method and apparatus may be configured to generate descriptors for individual maps. For example, a map may include a plurality of keyframes, each having a frame descriptor as described above. The max pooling unit may analyze the frame descriptors of the keyframes of the map and combine the frame descriptors into a unique map descriptor for the map.
[0338] Furthermore, it should be understood that other architectures may also be used for processing, as described above. For example, separate neural networks are described for generating DSF descriptors and frame descriptors. Such an approach is computationally efficient. However, in some embodiments, the frame descriptor may be generated from selected feature points without first generating the DSF descriptor. Ranking and Merging of Maps
[0339] What is described herein are methods and apparatuses for ranking and merging multiple environment maps within an X Reality (XR) system. Map merging can enable maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking the maps can enable efficient implementation of techniques as described herein that include map merging, which involves selecting maps from a set of maps based on similarity. In some embodiments, for example, a set of reference maps, formatted in a manner accessible by any of several XR devices, may be maintained by the system. These reference maps may be formed by merging selected tracking maps from those devices with other tracking maps or previously stored reference maps. The reference maps may be ranked, for example, to select one or more reference maps, merge with a new tracking map, and / or select one or more reference maps from the set for use when used within the device.
[0340] To provide a realistic XR experience to the user, the XR system must understand the user's physical surroundings in order to correctly correlate the location of virtual objects with real objects. Information about the user's physical surroundings may be obtained from an environment map regarding the user's location.
[0341] The inventors recognized and appreciated the true value that an XR system can provide an enhanced XR experience to multiple users sharing the same world with real and / or virtual content, regardless of whether those users are present in the world at the same or different times, by enabling efficient sharing of environmental maps of the real / physical world collected by multiple users. However, significant challenges exist in providing such a system. Such a system may store multiple maps generated by multiple users and / or the system may store multiple maps generated at different times. For example, with respect to operations that may be performed using previously generated maps such as localization, substantial processing may be required to identify relevant environmental maps of the same world (e.g., the same real-world location) from all of the environmental maps collected within the XR system. In some embodiments, there may be only a few environmental maps that a device may access, for example, for localization. In some embodiments, there may be a large number of environmental maps that a device may access. The inventors recognized and appreciated the true value of a technique for quickly and accurately ranking the relevance of environmental maps from any possible set of environmental maps, such as the population of all reference maps 120 in FIG. 28. High-ranked maps may then be selected for further processing, such as rendering virtual objects to realistically interact with the physical world around the user on a user display or merging the map data collected by that user with the stored maps to create a larger or more accurate map.
[0342] In some embodiments, a stored map associated with a task for a user at a location in the physical world may be identified by filtering the stored map based on multiple criteria. Those criteria may indicate a comparison of a tracking map generated by the user's wearable device at that location with candidate environmental maps stored in a database. The comparison may be performed based on metadata associated with a map, such as a Wi-Fi fingerprint detected by a device that generated the map, and / or a set of BSSIDs to which the device is connected while forming the map. The comparison may also be performed based on compressed or decompressed content of the map. A comparison based on a compressed representation may be performed, for example, by comparing vectors calculated from the map content. A comparison based on a decompressed map may be performed, for example, by locating the tracking map within the stored map or vice versa. Multiple comparisons may be performed in an order based on the calculation time required to reduce the number of candidate maps under consideration, and comparisons involving less calculation are performed in an order prior to other comparisons that require more calculation.
[0343] Figure 26 depicts an AR system 800 configured to rank and merge one or more environmental maps, according to some embodiments. The AR system may include a passable world model 802 of the AR device. Information for capturing the passable world model 802 may originate from sensors on the AR device, which may include computer-executable instructions for performing some or all of the processing for converting sensor data into a map, stored within a processor 804 (e.g., local data processing module 570 in FIG. 4). Such a map may be a tracking map that can be constructed as sensor data is collected as the AR device operates within the area. Along with the tracking map, area attributes may be provided to indicate the area represented by the tracking map. These area attributes may be geographical location identifiers, such as IDs used by the AR system to represent coordinates or locations presented as latitude and longitude. Alternatively, or in addition, the area attributes may be measured characteristics that have a high likelihood of being unique with respect to that area. The area attributes may be derived, for example, from parameters of wireless networks detected within the area. In some embodiments, the area attributes may be associated with unique addresses of access points that the AR system is adjacent to and / or connected to. For example, the area attributes may be associated with MAC addresses or basic service set identifiers (BSSIDs) of 5G base stations / routers, Wi-Fi routers, and the like.
[0344] In the example of FIG. 26, the tracking map may be merged with other maps of the environment. The map ranking portion 806 receives the tracking map from the device PW802, communicates with the map database 808, and selects and ranks environmental maps from the map database 808. The selected maps that are ranked higher are sent to the map merging portion 810.
[0345] The map merge part 810 may perform the merge process on the map transmitted from the map ranking part 806. The merge process may involve the step of merging some or all of the tracking map and the ranked map and transmitting the new merged map to the passable world model 812. The map merge part may merge the maps by identifying the maps depicting the overlapping parts of the physical world. Those overlapping parts may be aligned so that the information in both maps can be aggregated into the final map. The reference map may be merged with other reference maps and / or tracking maps.
[0346] Aggregation may involve the step of extending one map with information from another map. Alternatively, or in addition, aggregation may involve the step of adjusting the representation of the physical world in one map based on information in another map. The latter map may represent, for example, that an object generating the feature points has moved so that the map can be updated based on the latter information. Alternatively, two maps may characterize the same area with different feature points, and aggregation may involve the step of selecting a set of feature points from the two maps to better represent that area. Regardless of the specific processing occurring in the process of merging, in some embodiments, the PCF from all the maps being merged may be retained so that an application positioning the content relative thereto can continue to do so. In some embodiments, the merging of the maps may result in redundant persistent postures, and some of the persistent postures may be deleted. When the PCF is associated with the persistent postures to be deleted, the step of merging the maps may involve the step of modifying the PCF so that it is associated with the persistent postures remaining in the map after the merge.
[0347] In some embodiments, as maps are extended and / or updated, they may be refined. The refinement may involve calculations to reduce internal inconsistencies between feature points that are likely to represent the same object in the physical world. The inconsistencies can arise from inaccuracies in the pose associated with keyframes that supply feature points representing the same object in the physical world. Such inconsistencies can occur, for example, from an XR device calculating a pose relative to a tracking map, which in turn is based on the step of estimating a pose such that errors in the pose estimation accumulate over time and create a "drift" in the pose accuracy. The map may be refined by performing bundle adjustment or other operations to reduce the inconsistencies between feature points from multiple keyframes.
[0348] In response to the refinement, the location of persistent points relative to the origin of the map may change. Thus, the persistent pose or the transformation associated with such persistent points, such as a PCF, may also change. In some embodiments, the XR system may recalculate the transformation associated with any persistent points that have changed in relation to map refinement (whether performed as part of a merge operation or for other reasons). These transformations may be pushed from the component that calculates the transformation to the component that uses the transformation so that any use of the transformation can be based on the updated location of the persistent points.
[0349] The passable world model 812 may be a cloud model, which may be shared by a plurality of AR devices. The passable world model 812 may store or otherwise have access to an environmental map in the map database 808. In some embodiments, when a previously calculated environmental map is updated, the previous version of the map may be deleted so as to remove the old map from the database. In some embodiments, when a previously calculated environmental map is updated, the previous version of the map may be archived to enable reading / browsing of the previous version of the environment. In some embodiments, permissions may be set such that only an AR system having certain read / write access may trigger deletion / archiving of the previous version of the map.
[0350] These environmental maps created from the tracking maps provided by one or more AR devices / systems may be accessed by AR devices within the AR system. The map ranking portion 806 may also be used when supplying the environmental map to the AR device. The AR device may send a message requesting an environmental map regarding its current location, and the map ranking portion 806 may be used to select and rank the environmental map associated with the requesting device.
[0351] In some embodiments, the AR system 800 may include a downsampling portion 814 configured to receive the merged map from the cloud PW812. The merged map received from the cloud PW812 may be in a storage format for the cloud, which may include high-resolution information such as a large number of PCFs per square meter or multiple image frames or a large set of feature points associated with the PCFs. The downsampling portion 814 may be configured to downsample the cloud format map to a format suitable for storage on the AR device. The device format map may have less data, such as fewer PCFs or less data stored per PCF, and may accommodate the limited local computing power and storage space of the AR device.
[0352] FIG. 27 is a simplified block diagram illustrating a plurality of reference maps 120 that may be stored in a remote storage medium, such as in the cloud. Each reference map 120 may include a plurality of reference map identifiers indicating the location of the reference map in physical space, such as any location on Earth that is a planet. These reference map identifiers may include one or more of the following identifiers: an area identifier represented by a range of longitude and latitude, a frame descriptor (e.g., the global feature column 316 in FIG. 21), a Wi-Fi fingerprint, a feature descriptor (e.g., the feature descriptor 310 in FIG. 21), and a device identifier indicating one or more devices that contributed to the map.
[0353] In the illustrated embodiment, since the reference maps 120 may exist on the surface of the Earth, they are geographically arranged in a two-dimensional pattern. The reference maps 120 may be uniquely identifiable by corresponding longitude and latitude because any reference map having overlapping longitude and latitude may be merged into a new reference map.
[0354] FIG. 28 is a schematic diagram illustrating a method of selecting a reference map that can be used to locate a new tracking map relative to one or more reference maps, according to some embodiments. The method may begin, as an example, with an act (act 120) of accessing a parent set of reference maps 120 that may be stored in a database within a passable world (e.g., passable world module 538). The parent set of reference maps may include reference maps from all previously visited locations. The XR system may filter the parent set of all reference maps to a small subset or only a single map. It should be understood that in some embodiments, due to bandwidth limitations, it may not be possible to send all reference maps to the viewing device. Selecting a subset that is selected as a likely candidate for matching to the tracking map for transmission to the device can reduce the bandwidth and latency associated with accessing the remote database of maps.
[0355] This method may include a step (act 300) of filtering a parent set of reference maps based on an area with a predetermined size and shape. In the embodiment illustrated in FIG. 27, each square may represent an area. Each square may cover 50m × 50m. Each square may have six neighboring areas. In some embodiments, act 300 may select at least one matching reference map 120 that covers the longitude and latitude, including the longitude and latitude of the position identifier received from the XR device, as long as at least one map exists at that longitude and latitude. In some embodiments, act 300 may select at least one neighboring reference map that covers the longitude and latitude and is adjacent to the matching reference map. In some embodiments, act 300 may select a plurality of matching reference maps and a plurality of neighboring reference maps. Act 300 may, for example, reduce the number of reference maps to about one-tenth, for example, from thousands to hundreds, to form a first filtered selection. Alternatively, or in addition, criteria other than latitude and longitude may be used to identify neighboring maps. The XR device may have been previously located using the reference maps in the set, for example, as part of the same session. The cloud service may retain information about the XR device, including the previously located maps. In this example, the maps selected in act 300 may include those that cover an area adjacent to the map at which the XR device was located.
[0356] This method may include a step (act 302) of filtering a first filtered selection of reference maps based on a Wi-Fi fingerprint. Act 302 may determine latitude and longitude based on a Wi-Fi fingerprint received as part of a location identifier from an XR device. Act 302 may compare the latitude and longitude from the Wi-Fi fingerprint with the latitude and longitude of reference map 120 and determine one or more reference maps that form a second filtered selection. Act 302 may reduce the number of reference maps to about one tenth, for example, from hundreds to dozens (e.g., 50) of reference maps that form a second selection. For example, the first filtered selection may include 130 reference maps, the second filtered selection may include 50 of the 130 reference maps, and may not include the remaining 80 of the 130 reference maps.
[0357] This method may include a step (act 304) of filtering a second filtered selection of reference maps based on keyframes. Act 304 may compare data representing an image captured by an XR device with data representing reference map 120. In some embodiments, the data representing the image and / or the map may include feature descriptors (e.g., DSF descriptors in FIG. 25) and / or global feature sequences (e.g., 316 in FIG. 21). Act 304 may provide a third filtered selection of reference maps. In some embodiments, the output of act 304 may be, for example, only 5 of the 50 reference maps identified following the second filtered selection. Map transmitter 122 then transmits one or more reference maps to the viewing device based on the third filtered selection. Act 304 may reduce the number of reference maps to about one tenth, forming a third selection, for example, from dozens to single-digit numbers of reference maps (e.g., 5). In some embodiments, the XR device may receive the reference maps within the third filtered selection and attempt to locate itself within the received reference maps.
[0358] For example, the act 304 may filter the reference map 120 based on the global feature sequence 316 of the reference map 120 and the global feature sequence 316 based on an image captured by the vision device (e.g., an image that may be part of a local tracking map for the user). Each of the reference maps 120 in FIG. 27 thus has one or more global feature sequences 316 associated therewith. In some embodiments, the global feature sequence 316 may be obtained when the XR device submits an image or feature details to the cloud, and the cloud processes the image or feature details and generates the global feature sequence 316 for the reference map 120.
[0359] In some embodiments, the cloud may receive feature details of live / new / current images captured by the vision device, and the cloud may generate the global feature sequence 316 for the live image. The cloud may then filter the reference map 120 based on the live global feature sequence 316. In some embodiments, the global feature sequence may be generated on a local vision device. In some embodiments, the global feature sequence may be generated remotely, e.g., on the cloud. In some embodiments, the cloud may transmit the filtered reference map, along with the global feature sequence 316 associated with the filtered reference map, to the XR device. In some embodiments, when the vision device locates its tracking map relative to the reference map, this may be done by matching the global feature sequence 316 of the local tracking map with the global feature sequence of the reference map.
[0360] It should be understood that the XR device's operation may not perform all of the actions (300, 302, 304). For example, if the parent set of reference maps is relatively small (e.g., 500 maps), the XR device attempting to localize may filter the parent set of reference maps based on Wi-Fi fingerprints (e.g., action 302) and keyframes (e.g., action 304), but may omit filtering based on area (e.g., action 300). Further, the maps need not be compared as a whole. In some embodiments, for example, the comparison of two maps may result in the identification of common persistent points such as persistent poses or PCFs that appear in both the new map and the map selected from the parent set of maps. In that case, descriptors may be associated with the persistent points and those descriptors may be compared.
[0361] FIG. 29 is a flowchart illustrating a method 900 for selecting one or more ranked environmental maps, according to some embodiments. In the illustrated embodiment, the ranking step is performed for the user's AR device that creates a tracking map. Thus, the tracking map is available for use in ranking the environmental maps. In embodiments where the tracking map is unavailable, some or all of the steps of selecting and ranking the environmental maps that do not explicitly rely on the tracking map may be used.
[0362] Method 900 may begin with act 902, where a set of maps (which may be formatted as reference maps) from a database of environmental maps in the vicinity of where the tracking map was formed is accessed and then may be filtered for ranking. Additionally, in act 902, at least one area attribute regarding the area in which the user's AR device is operating is determined. In a scenario where the user's AR device is constructing a tracking map, the area attribute may correspond to the area over which the tracking map was created. As a specific example, the area attribute may be calculated based on signals received from access points to a computer network during the time the AR device was calculating the tracking map.
[0363] FIG. 30 depicts an exemplary map ranking portion 806 of an AR system 800 according to some embodiments. The map ranking portion 806 may include portions executed on the AR device and portions executed on a remote computing system such as the cloud and thus may be executed within a cloud computing environment. The map ranking portion 806 may be configured to implement at least a portion of method 900.
[0364] FIG. 31A depicts an example of a tracking map (TM) 1102 and area attributes AA1-AA8 of environmental maps CM1-CM4 in a database according to some embodiments. As shown, an environmental map may be associated with a plurality of area attributes. Area attributes AA1-AA8 may include parameters of a wireless network detected by an AR device that calculates the tracking map 1102, e.g., the basic service set identifier (BSSID) of the network to which the AR device is connected, and / or, e.g., the strength of the signal of an access point received by the wireless network through network tower 1104. The parameters of the wireless network may conform to protocols including Wi-Fi and 5G NR. In the example illustrated in FIG. 32, the area attribute is a fingerprint of the area in which the user AR device collected sensor data and formed a tracking map.
[0365] FIG. 31B depicts an example of a determined geographic location 1106 of a tracking map 1102 according to some embodiments. In the example shown, the determined geographic location 1106 includes a centroid point 1110 and an area 1108 surrounding the centroid point. It should be understood that the determination of the geographic location in the present application is not limited to the illustrated format. The determined geographic location may have any suitable format, e.g., including different area shapes. In this example, the geographic location is determined from the area attributes using a database that associates the area attributes with the geographic location. The database is commercially available and is, e.g., a database that associates Wi-Fi fingerprints, represented as latitude and longitude, with locations and can be used for this operation.
[0366] In the embodiment of FIG. 29, the map database containing the environmental map may also include location data regarding those maps, including the latitude and longitude covered by the map. The processing in act 902 may involve selecting a set of environmental maps from the database that cover the same latitude and longitude determined for the area attributes of the tracking map.
[0367] Action 904 is a first filtering of a set of environmental maps accessed in Action 902. In Action 902, the environmental maps are retained in the set based on their proximity to the geographical location of the tracking map. This filtering step may be performed by comparing the latitude and longitude associated with the tracking map and the environmental maps in the set.
[0368] FIG. 32 depicts an example of Action 904 according to some embodiments. Each area attribute may have a corresponding geographical location 1202. The set of environmental maps may include environmental maps with at least one area attribute having a geographical location that overlaps with the determined geographical location of the tracking map. In the illustrated example, the identified set of environmental maps includes environmental maps CM1, CM2, and CM4, each having at least one area attribute having a geographical location that overlaps with the determined geographical location of the tracking map 1102. The environmental map CM3, associated with the area attribute AA6, is not included in the set because it is outside the determined geographical location of the tracking map.
[0369] Other filtering steps may also be performed on the set of environment maps to reduce / rank the number of environment maps within the set that are ultimately processed (for map merging, or for providing passable world information to the user device, etc.). Method 900 may include a step (act 906) of filtering the set of environment maps based on the similarity of one or more identifiers of network access points associated with the environment maps of the set of tracking maps and environment maps. During map formation, a device that collects sensor data and generates a map may be connected to the network through a network access point through Wi-Fi or a similar wireless communication protocol, etc. The access point may be identified by a BSSID. As the user device moves through an area, collects data, and forms a map, it may connect to multiple different access points. Similarly, when multiple devices supply information for forming a map, the devices may be connected through different access points, and thus, for the same reason, there may be multiple access points used when forming the map. Therefore, there may be multiple access points associated with the map, and the set of access points may be an indication of the location of the map. The strength of the signal from the access point, which may be reflected as an RSSI value, may provide additional geographical information. In some embodiments, a list of BSSIDs and RSSI values may form area attributes for the map.
[0370] In some embodiments, the step of filtering a set of environmental maps based on the similarity of one or more identifiers of network access points may include retaining, within the set of environmental maps, an environmental map with the highest Jaccard similarity with at least one area attribute of a tracking map, based on one or more identifiers of network access points. FIG. 33 depicts an example of act 906 according to some embodiments. In the illustrated example, the network identifier associated with area attribute AA7 can be determined as the identifier for tracking map 1102. The set of environmental maps after act 906 may include environmental map CM2, which has an area attribute within a higher Jaccard similarity with AA7, and environmental map CM4, which also includes area attribute AA7. Environmental map CM1 is not included in the set because it has the lowest Jaccard similarity with AA7.
[0371] The processing in acts 902-906 may be performed without actually accessing the content of the maps stored in the map database, based on the metadata associated with the maps. Other processing may involve accessing the content of the maps. Act 908 shows the step of accessing the environmental maps remaining in the subset after filtering based on the metadata. It should be understood that this act may be performed either earlier or later in the process, if the subsequent operations can be performed using the content being accessed.
[0372] Method 900 may include a step (act 910) of filtering a set of environment maps based on the similarity of a metric representing the content of the environment map of the set of tracking maps and environment maps. The metric representing the content of the tracking maps and environment maps may include a vector of values calculated from the content of the maps. For example, the deep keyframe descriptors as described above, calculated for one or more keyframes used in forming the maps, may provide a metric for comparison of maps or portions of maps. The metric may be calculated from the maps read in at act 908, or may be pre-calculated and stored as metadata associated with those maps. In some embodiments, the step of filtering the set of environment maps based on the similarity of a metric representing the content of the environment map of the set of tracking maps and environment maps may include retaining in the set of environment maps those environment maps with a minimum vector distance between a vector of characteristics of the tracking map and a vector representing the environment map within the set of environment maps.
[0373] Method 900 may further include a step (act 912) of filtering the set of environment maps based on a degree of matching between a portion of the tracking map and a portion of the environment map of the set of environment maps. The degree of matching may be determined as part of a localization process. As a non-limiting example, localization may be performed by identifying significant points within the tracking map and environment map that are similar enough that they may represent the same portion of the physical world. In some embodiments, the significant points may be features, feature descriptors, keyframes, key rigs, persistent poses, and / or PCFs. The set of significant points within the tracking map may then be aligned to produce a best fit with the set of significant points within the environment map. An average squared distance between corresponding significant points may be calculated and used as an indication that the tracking map and environment map represent the same region of the physical world if it is below a threshold for a particular region of the tracking map.
[0374] In some embodiments, the step of filtering the set of environmental maps based on a degree of matching between a portion of the tracking map and a portion of the environmental map of the set of environmental maps may include calculating a volume of the physical world represented by the tracking map that is also represented within the environmental map of the set of environmental maps, and retaining within the set of environmental maps an environmental map having a calculated volume larger than the filtered-out environmental maps of the set. FIG. 34 depicts an example of act 912 according to some embodiments. In the illustrated example, the set of environmental maps after act 912 includes environmental map CM4 having an area 1402 that matches an area of tracking map 1102. Environmental map CM1 is not included in the set because it does not have an area that matches the area of tracking map 1102.
[0375] In some embodiments, the set of environmental maps may be filtered in the order of act 906, act 910, and act 912. In some embodiments, the set of environmental maps may be filtered based on act 906, act 910, and act 912, which may be performed in an order based on the processing required to perform the filtering from lowest to highest. Method 900 may include a step (act 914) of loading a set of environmental maps and data.
[0376] In the illustrated embodiment, the user database stores an area identification indicating the area where the AR device was used. The area identification may be an area attribute, which may include parameters of the wireless network detected by the AR device during use. The map database may store a plurality of environmental maps constructed from the data supplied by the AR device and the associated metadata. The associated metadata may include an area identification derived from the area identification of the AR device that supplied the data from which the environmental map was constructed. The AR device may send a message to the PW module indicating that a new tracking map is being created or is in the process of being created. The PW module may calculate an area identifier for the AR device and update the user database based on the received parameters and / or the calculated area identifier. The PW module may also determine an area identifier associated with the AR device requesting an environmental map, identify a set of environmental maps from the map database based on the area identifier, filter the set of environmental maps, and transmit the filtered set of environmental maps to the AR device. In some embodiments, the PW module may filter the set of environmental maps based on one or more criteria, including, for example, the geographical location of the tracking map, the similarity of one or more identifiers of network access points associated with the environmental maps of the set of tracking maps and environmental maps, the similarity of metrics representing the content of the environmental maps of the set of tracking maps and environmental maps, and the degree of matching between a portion of the tracking map and a portion of the environmental maps of the set of environmental maps.
[0377] Although some aspects of some embodiments have been described so far, it should be understood that various modifications, corrections, and improvements will readily occur to those skilled in the art. As one example, the embodiments are described in relation to an augmented (AR) environment. It should be understood that some or all of the techniques described herein may be applied within an MR environment, and more generally, within other XR environments and VR environments.
[0378] As another example, embodiments are described in relation to devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), discrete applications, and / or any suitable combination of devices, networks, and discrete applications.
[0379] Further, FIG. 29 provides an example of a criterion that can be used to filter a candidate map and result in a set of high-ranked maps. Other criteria may be used instead of, or in addition to, the described criteria. For example, if multiple candidate maps have similar values of a metric used to filter out less desirable maps, the characteristics of the candidate maps may be used to determine which candidate maps are retained as candidate maps or filtered out. For example, larger or denser candidate maps may be prioritized over smaller candidate maps. In some embodiments, FIGS. 27-28 may illustrate all or part of the systems and methods described in FIGS. 29-34.
[0380] FIGS. 35 and 36 are schematic diagrams illustrating an XR system configured to rank and merge a plurality of environmental maps according to some embodiments. In some embodiments, a passable world (PW) may determine when to trigger steps to rank and / or merge maps. In some embodiments, the step of determining which maps are to be used may, according to some embodiments, be at least partially based on the deep keyframes described above in relation to FIGS. 21-25.
[0381] FIG. 37 is a block diagram illustrating a method 3700 for creating an environmental map of the physical world, according to some embodiments. The method 3700 may begin with the step of (act 3702) locating a tracking map captured by an XR device worn by a user with respect to a group of reference maps (e.g., reference maps selected by the method of FIG. 28 and / or method 900 of FIG. 29). Act 3702 may include the step of locating key rigs of the tracking map among the group of reference maps. The location result of each key rig may include the located pose of the key rig and a set of 2D / 3D feature correspondences.
[0382] In some embodiments, the method 3700 may include the step of (act 3704) splitting the tracking map into connected components, which may enable the map to be robustly merged by merging the connected fragments. Each connected component may include key rigs that are within a predetermined distance. The method 3700 may include the step of (act 3706) merging connected components that are larger than a predetermined threshold into one or more reference maps, and the step of removing the merged connected components from the tracking map.
[0383] In some embodiments, the method 3700 may include the step of (act 3708) merging the reference maps of the group that are merged with the same connected component of the tracking map. In some embodiments, the method 3700 may include the step of (act 3710) promoting the remaining connected components of the tracking map that are not merged with any reference map to the reference map. In some embodiments, the method 3700 may include the step of (act 3712) merging the reference maps that are merged with the persistent pose and / or PCF of the tracking map and at least one connected component of the tracking map. In some embodiments, the method 3700 may include the step of (act 3714) completing the reference map, for example, by fusing map points and pruning redundant key rigs.
[0384] Figures 38A and 38B illustrate an environmental map 3800 created by updating a reference map 700 that can be promoted from a tracking map 700 (FIG. 7) with a new tracking map, according to some embodiments. As illustrated and described with respect to FIG. 7, the reference map 700 can provide a floor plan 706 of a reconstructed physical object in the corresponding physical world, represented by points 702. In some embodiments, the map points 702 can represent features of a physical object that can include a plurality of features. The new tracking map can be captured with respect to the physical world, uploaded to the cloud, and merged with the map 700. The new tracking map can include map points 3802 and key rigs 3804, 3806. In the illustrated example, the key rig 3804 represents a key rig that is properly located with respect to the reference map, for example, by establishing a correspondence with the key rig 704 of the map 700 (as illustrated in FIG. 38B). On the other hand, the key rig 3806 represents a key rig that is not located with respect to the map 700. The key rig 3806 can be promoted to a separate reference map in some embodiments.
[0385] FIGS. 39A - 39F are schematic diagrams illustrating examples of a cloud-based persistent coordinate system that provides a shared experience for users within the same physical space. FIG. 39A shows, for example, that a reference map 4814 from the cloud is received by XR devices worn by users 4802A and 4802B of FIGS. 20A - 20C. The reference map 4814 can have a reference coordinate frame 4806C. The reference map 4814 can have a PCF 4810C with a plurality of associated PPs (e.g., 4818A, 4818B in FIG. 39C).
[0386] FIG. 39B shows that the XR device has established a relationship between its individual world coordinate systems 4806A, 4806B and the reference coordinate frame 4806C. This may be done, for example, by localizing the reference map 4814 on the individual device. Localizing the tracking map with respect to the reference map can result in a transformation between the local world coordinate system of each device and the coordinate system of the reference map.
[0387] FIG. 39C shows that as a result of the localization, a transformation can be calculated between the local PCF (e.g., PCF 4810A, 4810B) on the individual device and the individual persistent poses (e.g., PP 4818A, 4818B) on the reference map (e.g., transformations 4816A, 4816B). By using these transformations, each device can use its local PCF, which processes the images detected using sensors on the device, determines the location relative to the local device, and can be locally detected on the device by displaying virtual content associated with PP 4818A, 4818B or other persistent points on the reference map. Such an approach can accurately localize the virtual content for each user and can enable each user to have the same experience of the virtual content within the physical space.
[0388] FIG. 39D shows a persistent pose snapshot from a reference map to a local tracking map. As can be seen from the figure, the local tracking maps are interconnected via the persistent pose. FIG. 39E shows that the PCF4810A on the device worn by user 4802A is accessible through PP4818A within the device worn by user 4802B. FIG. 39F shows that the tracking maps 4804A, 4804B and the reference 4814 can be merged. In some embodiments, some PCFs may be removed as a result of the merge. In the illustrated example, the merged map includes the PCF4810C of the reference map 4814 but does not include the PCFs 4810A, 4810B of the tracking maps 4804A, 4804B. The PPs previously associated with the PCFs 4810A, 4810B may be associated with the PCF4810C after the map merge.
[0389]
Example
[0390] FIGS. 40 and 41 illustrate an example of generating a tracking map by the first XR device 12.1 of FIG. 9. FIG. 40 is a two-dimensional representation of a three-dimensional first local tracking map (Map 1) according to some embodiments, which can be generated by the first XR device of FIG. 9. FIG. 41 is a block diagram illustrating the step of uploading Map 1 from the first XR device to the server of FIG. 9 according to some embodiments.
[0391] FIG. 40 illustrates Map 1 and virtual content (Content 123 and Content 456) on the first XR device 12.1. Map 1 has an origin (Origin 1). Map 1 includes several PCFs (PCFa - PCFd). From the perspective of the first XR device 12.1, as an example, PCFa is located at the origin of Map 1 and has X, Y, and Z coordinates of (0, 0, 0), and PCFb has X, Y, and Z coordinates of (-1, 0, 0). Content 123 is associated with PCFa. In this embodiment, Content 123 has X, Y, and Z relationships with respect to PCFa of (1, 0, 0). Content 456 has a relationship with respect to PCFb. In this embodiment, Content 456 has X, Y, and Z relationships of (1, 0, 0) with respect to PCFb.
[0392] In FIG. 41, the first XR device 12.1 uploads Map 1 to the server 20. The server 20, here, has a reference map based on Map 1. The first XR device 12.1 has a reference map that is empty at this stage. The server 20, for the purpose of discussion, in some embodiments, does not include other maps other than Map 1. The map is not stored on the second XR device 12.2.
[0393] The first XR device 12.1 also transmits its Wi-Fi signature data to the server 20. The server 20 may use the Wi-Fi signature data to determine the approximate location of the first XR device 12.1 based on the GPS locations of such other devices that were recorded together with the intelligence collected from other devices that were connected to the server 20 or other servers in the past. The first XR device 12.1 may here end the first session (see FIG. 8) and disconnect from the server 20.
[0394] FIG. 42 is a schematic diagram illustrating the XR system of FIG. 16 according to some embodiments, showing that after the first user 14.1 ended the first session, the second user 14.2 started a second session using the second XR device of the XR system. FIG. 43A is a block diagram showing the start of the second session by the second user 14.2. The first user 14.1 is shown by a phantom line because the first session by the first user 14.1 has ended. The second XR device 12.2 begins to record an object. Various systems with variable granularity may be used by the server 20 to determine that the second session by the second XR device 12.2 is within the same vicinity as the first session by the first XR device 12.1. For example, Wi-Fi signature data, global positioning system (GPS) positioning data, GPS data based on Wi-Fi signature data, or any other data indicating location may be included within the first and second XR devices 12.1 and 12.2 to record that location. Alternatively, the PCF identified by the second XR device 12.2 may show similarity to the PCF of Map 1.
[0395] As shown in FIG. 43B, the second XR device boots up and begins to collect data such as image 1110 from one or more cameras 44, 46. As shown in FIG. 14, in some embodiments, an XR device (e.g., the second XR device 12.2) may collect one or more images 1110, perform image processing, and extract one or more features / keypoints 1120. Each feature may be converted into a descriptor 1130. In some embodiments, the descriptor 1130 may be used to describe a keyframe 1140 that may have the position and orientation of the associated associated image. One or more keyframes 1140 may correspond to a single persistent pose 1150 that may be automatically generated after a threshold distance from a previous persistent pose 1150, e.g., 3 meters. One or more persistent poses 1150 may correspond to a single PCF 1160 that may be automatically generated after a predetermined distance, e.g., every 5 meters. Over time, as the user continues to move around the user's environment and the XR device continues to collect more data such as image 1110, additional PCFs (e.g., PCF3 and PCF4, 5) may be created. An application, i.e., two 1180, may be launched on the XR device and provide virtual content 1170 to the XR device for presentation to the user. The virtual content may have an associated content coordinate frame, which may be established with respect to one or more PCFs. As shown in FIG. 43B, the second XR device 12.2 creates three PCFs. In some embodiments, the second XR device 12.2 may attempt to localize with respect to one or more reference maps stored on server 20.
[0396] In some embodiments, as shown in FIG. 43C, the second XR device 12.2 may download the reference map 120 from the server 20. The map 1 on the second XR device 12.2 includes PCFs a-d and the origin 1. In some embodiments, the server 20 may have multiple reference maps for various locations, and the second XR device 12.2 determines that it is in the same vicinity as the first XR device 12.1 during the first session, and the server 20 may send the reference map regarding that vicinity to the second XR device 12.2.
[0397] FIG. 44 shows that the second XR device 12.2 starts to identify the PCFs for the purpose of generating map 2. The second XR device 12.2 identifies only a single PCF, i.e., PCFs 1, 2. The X, Y, and Z coordinates of PCFs 1, 2 for the second XR device 12.2 can be (1, 1, 1). Map 2 has its own origin (origin 2), which may be based on the head pose of device 2 at the start of the device for the current head pose session. In some embodiments, the second XR device 12.2 may immediately attempt to localize map 2 relative to the reference map. In some embodiments, map 2 may be impossible to localize relative to the reference map (map 1) (i.e., the localization may fail) because the system does not recognize any or sufficient overlap between the two maps. In some embodiments, the system may perform localization based on a PCF comparison between the local map and the reference map. In some embodiments, the system may perform localization based on a continuous pose comparison between the local map and the reference map. In some embodiments, the system may perform localization based on a keyframe comparison between the local map and the reference map.
[0398] Figure 45 shows Map 2 after the second XR device 12.2 has identified further PCFs (PCF1, 2, PCF3, PCF4, 5) of Map 2. The second XR device 12.2 again attempts to localize Map 2 with respect to the reference map. Since Map 2 has been extended to overlap at least a portion of the reference map, the localization attempt will succeed. In some embodiments, the overlap between the local tracking map, Map 2, and the reference map may be represented by PCFs, persistent poses, keyframes, or any other suitable intermediate or derived constructs.
[0399] Further, the second XR device 12.2 associates content 123 and content 456 with PCF1, 2, and PCF3 of Map 2. Content 123 has X, Y, and Z coordinates for PCF1, 2 of (1, 0, 0). Similarly, the X, Y, and Z coordinates for PCF3 within Map 2 are also (1, 0, 0).
[0400] Figures 46A and 46B illustrate the successful localization of Map 2 with respect to the reference map. The overlapping area / volume / section of Map 1410 represents the common portion with Map 1 and the reference map. Different PCFs were created to represent the same volume within the real space (e.g., different maps) since Map 2 created PCF3 and 4, 5 before localization, and the reference map created PCFa and c before Map 2 was created.
[0401] As shown in FIG. 47, the second XR device 12.2 expands map 2 to include PCFa-d from the reference map. The inclusion of PCFa-d represents the localization of map 2 relative to the reference map. In some embodiments, the XR system may perform an optimization step to remove PCFs within 1410, i.e., duplicate PCFs such as PCF3 and PCF4, 5, etc., from the overlapping areas. After map 2 is localized, the placement of virtual content such as content 456 and content 123 will be relative to the nearest updated PCF within the updated map 2. The virtual content will appear to the user within the same real-world location despite the changed PCF associations for the content and the updated PCF for map 2.
[0402] As shown in FIG. 48, the second XR device 12.2 continues to expand map 2 as additional PCFs (PCFe, f, g, and h) are identified by the second XR device 12.2, e.g., as the user walks around the real world. Note that map 1 is not expanded in FIGS. 47 and 48.
[0403] Referring to FIG. 49, the second XR device 12.2 uploads map 2 to the server 20. The server 20 stores map 2 together with the reference map. In some embodiments, map 2 may be uploaded to the server 20 when the session for the second XR device 12.2 ends.
[0404] The reference map within the server 20 here includes PCFi, which is not included within map 1 on the first XR device 12.1. The reference map on the server 20 can be expanded to include PCFi when a third XR device (not shown) uploads a map to the server 20 and such a map includes PCFi.
[0405] In FIG. 50, the server 20 merges Map 2 with the reference map to form a new reference map. The server 20 determines that PCFs a-d are common to the reference map and Map 2. The server expands the reference map to include PCFs e-h and PCF 1, 2 from Map 2, forming a new reference map. The reference maps on the first and second XR devices 12.1 and 12.2 become obsolete based on Map 1.
[0406] In FIG. 51, the server 20 transmits the new reference map to the first and second XR devices 12.1 and 12.2. In some embodiments, this may occur when the first XR device 12.1 and the second device 12.2 attempt to localize during a different or new or subsequent session. The first and second XR devices 12.1 and 12.2 proceed to localize their respective local maps (Map 1 and Map 2, respectively) with respect to the new reference map as described above.
[0407] As shown in FIG. 52, the head coordinate frame 96 or "head pose" is related to the PCF within Map 2. In some embodiments, the origin of the map, i.e., origin 2, is based on the head pose of the second XR device 12.2 at the start of the session. As PCFs are created during the session, the PCFs are established with respect to the world coordinate frame, i.e., origin 2. The PCFs of Map 2 serve as a persistent coordinate frame with respect to the reference coordinate frame, and the world coordinate frame may be the world coordinate frame of the previous session (e.g., origin 1 of Map 1 in FIG. 40). The transformation from the world coordinate frame to the head coordinate frame 96 has been described above with reference to FIG. 9. The head coordinate frame 96 shown in FIG. 52 is at a specific coordinate position with respect to the PCF of Map 2 and at a specific angle with respect to Map 2, having only two orthogonal axes. However, it should be understood that the head coordinate frame 96 is within a certain three-dimensional location with respect to the PCF of Map 2 and has three orthogonal axes within the three-dimensional space.
[0408] In FIG. 53, the head coordinate frame 96 is moving relative to the PCF of Map 2. The head coordinate frame 96 is moving because the second user 14.2 has moved their head. The user can move their head in six degrees of freedom (6dof). The head coordinate frame 96 can thus move in 6dof, i.e., three-dimensionally and in approximately three orthogonal axes relative to the PCF of Map 2, from its previous location in FIG. 52. The head coordinate frame 96 is adjusted whenever the real object detection camera 44 and the inertial measurement unit 48 in FIG. 9 respectively detect the movement of the real object and the head unit 22. Further information regarding head pose tracking is disclosed in and incorporated herein by reference in its entirety in U.S. Patent Application No. 16 / 221,065, entitled "Enhanced Pose Determination for Display Device".
[0409] FIG. 54 shows that sound may be associated with one or more PCFs. The user may wear, for example, headphones or earphones with stereo sound. The location of the sound through the headphones can be simulated using conventional techniques. The location of the sound may be stationary such that when the user rotates their head to the left, the location of the sound rotates to the right, and thus the user perceives the sound originating from the same location in the real world. In this embodiment, the locations of the sound are represented by sound 123 and sound 456. For purposes of discussion, FIG. 54 is similar to FIG. 48 in its analysis. When the first and second users 14.1 and 14.2 are located in the same room at the same or different times, they perceive the sounds 123 and 456 originating from the same location in the real world.
[0410] Figures 55 and 56 illustrate further implementations of the technology described above. The first user 14.1 initiated a first session as described with reference to FIG. 8. As shown in FIG. 55, the first user 14.1 ended the first session as indicated by the imaginary line. At the end of the first session, the first XR device 12.1 uploaded Map 1 to the server 20. The first user 14.1 then initiated a second session at a time after the first session. Since Map 1 is already stored on the first XR device 12.1, the first XR device 12.1 does not download Map 1 from the server 20. If Map 1 is lost, the first XR device 12.1 downloads Map 1 from the server 20. The first XR device 12.1 then proceeds to build a PCF for Map 2, localize with respect to Map 1, and further expand the reference map as described above. The Map 2 of the first XR device 12.1 is then used to associate local content, head coordinate frame, local sound, etc. as described above.
[0411] Referring to FIGS. 57 and 58, it is also possible to consider that more than one user may interact with the server in the same session. In this embodiment, a third user 14.3 with a third XR device 12.3 is added to the first user 14.1 and the second user 14.2. Each of the XR devices 12.1, 12.2, and 12.3 begins to generate its own map, namely, Map 1, Map 2, and Map 3, respectively. As the XR devices 12.1, 12.2, and 12.3 continue to expand Maps 1, 2, and 3, the maps are gradually uploaded to the server 20. The server 20 merges Maps 1, 2, and 3 to form a reference map. The reference map is then transmitted from the server 20 to each of the XR devices 12.1, 12.2, and 12.3.
[0412] Figure 59 illustrates aspects of a visual recognition method for restoring and / or resetting a head pose according to some embodiments. In the illustrated embodiment, in act 1400, the visual recognition device is powered on. In act 1410, in response to being powered on, a new session is started. In some embodiments, the new session may include steps to establish a head pose. One or more capture devices on a head-mounted frame that is affixed to the user's head first capture an image of the environment and then capture the surface of the environment by determining the surface from the image. In some embodiments, the surface data may also be combined with data from a gravity sensor to establish the head pose. Other suitable methods for establishing the head pose may be used.
[0413] In act 1420, the processor of the visual recognition device enters a routine for tracking the head pose. The capture device continues to capture the surface of the environment and determine the orientation of the head-mounted frame relative to the surface as the user moves their head.
[0414] In act 1430, the processor determines whether the head pose has been lost. The head pose may be lost due to "edge" cases such as too many reflective surfaces, low light levels, empty walls, outdoors, etc., which can result in low feature acquisition, or due to dynamic cases such as moving clusters that form part of the map. The routine in 1430 allows a certain amount of time, for example, 10 seconds, to elapse to allow sufficient time to determine whether the head pose has been lost. If the head pose has not been lost, the processor returns to 1420 and enters the head pose tracking again.
[0415] If the head pose is lost in act 1430, the processor enters a routine for restoring the head pose in 1440. If the head pose is lost due to low light levels, a message such as the following message is displayed to the user through the display of the visual recognition device:
[0416] The system is detecting low light conditions. Please move to an area with more light.
[0417] The system will continue to monitor whether sufficient light is available and whether the head pose can be restored. Alternatively, the system may determine that the low texture of the surface is causing the loss of the head pose, in which case the user will be given the following prompt within the display as a suggestion to improve the capture of the surface:
[0418] The system cannot detect a sufficient surface with fine texture. Please move to an area where the surface texture is not rough and the texture is more refined.
[0419] In action 1450, the processor enters a routine to determine whether the head pose restoration has failed. If the head pose restoration has not failed (i.e., the head pose restoration has been successful), the processor returns to action 1420 by entering the head pose tracking again. If the head pose restoration has failed, the processor returns to action 1410 to establish a new session. As part of the new session, all cached data is invalidated, and thereafter, the head pose is newly established. Any suitable method of head tracking may be used in combination with the process described in FIG. 59. U.S. Patent Application No. 16 / 221,065 describes head tracking and is incorporated herein by reference in its entirety.
[0420] FIG. 60 shows a schematic representation of a machine in an exemplary form of computer system 1900, where a set of instructions for causing the machine to carry out any one or more of the methodologies discussed herein may be executed according to some embodiments. In alternative embodiments, the machine may operate as a stand-alone device or may be connected (e.g., networked) to other machines. Further, although only a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set of instructions (or multiple sets) to perform any one or more of the methodologies discussed herein.
[0421] Exemplary computer system 1900 includes processor 1902 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), main memory 1904 (e.g., read only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), and static memory 1906 (e.g., flash memory, static random access memory (SRAM), etc.), which communicate with each other via bus 1908.
[0422] Computer system 1900 may further include disk drive unit 1916 and network interface device 1920.
[0423] Disk drive unit 1916 includes machine-readable medium 1922 on which is stored a set of instructions 1924 (e.g., software) embodying any one or more of the methodologies or functions described herein. The software may also reside, completely or at least partially, within main memory 1904 and / or within processor 1902 during execution by computer system 1900, main memory 1904, and processor 1902, and in so doing may also constitute a machine-readable medium.
[0424] The software may also be transmitted or received via network 18 through network interface device 1920.
[0425] Computer system 1900 includes driver chip 1950 that drives a projector and is used to generate light. Driver chip 1950 includes its own data storage device 1960 and its own processor 1962.
[0426] Although machine-readable medium 1922 is shown as a single medium in the exemplary embodiment, the term "machine-readable medium" should be construed to include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) that store a set of one or more instructions. The term "machine-readable medium" should also be construed to include any medium that can store, encode, or carry a set of instructions for machine execution and cause a machine to perform any one or more of the methodologies of the present invention. The term "machine-readable medium" should, therefore, be construed to include, without limitation, solid-state memory, optical and magnetic media, and carrier wave signals.
[0427] Although some aspects of some embodiments have been described heretofore, it should be understood that various modifications, corrections, and improvements will readily occur to those skilled in the art.
[0428] As an example, embodiments are described in relation to an augmented (AR) environment. It should be understood that some or all of the techniques described herein may be applied within an MR environment, or more generally, within other XR environments and VR environments.
[0429] As another example, the embodiments are described in relation to devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), discrete applications, and / or any suitable combination of devices, networks, and discrete applications.
[0430] Further, FIG. 29 provides an example of a criterion that can be used to filter a candidate map and result in a set of highly ranked maps. Other criteria may be used instead of or in addition to the criteria described. For example, if multiple candidate maps have similar values of a metric used to filter out less desirable maps, the characteristics of the candidate maps may be used to determine which candidate maps are retained as candidate maps or filtered out. For example, larger or denser candidate maps may be preferred over smaller candidate maps.
[0431] Such modifications, corrections, and improvements are intended to be part of the present disclosure and are intended to be within the spirit and scope of the present disclosure. Further, while the advantages of the present disclosure are shown, it should be understood that not all embodiments of the present disclosure include all of the advantages described. Some embodiments may not implement any of the features described as advantageous herein and in some cases. Therefore, the foregoing description and drawings are merely examples.
[0432] Some embodiments relate to a portable electronic system that includes a sensor configured to capture information about a three-dimensional (3D) environment and output images, where each image comprises a plurality of pixels, and at least one processor configured to execute computer-executable instructions and process the images output by the sensor. The computer-executable instructions comprise instructions for receiving a plurality of images captured by the sensor, identifying, for at least a subset of the plurality of images, for each image of the subset of images, one or more features within the plurality of pixels, where each feature corresponds to one or more pixels, calculating a feature descriptor for each of the one or more features, and calculating a frame descriptor for representing the image, at least in part, based on the calculated feature descriptors within the image, for each image of the subset of images.
[0433] In some embodiments, the sensor comprises at least a million pixel circuit. The frame descriptor for each of the plurality of images comprises 512 or fewer numbers.
[0434] In some embodiments, the computer-executable instructions further comprise instructions for constructing a map of at least a portion of the 3D environment and associating a feature descriptor for an individual frame with a portion of the map generated at least in part from the individual frame.
[0435] In some embodiments, the computer-executable instructions comprise instructions for selecting, as a subset of the plurality of images, one or more keyframes from the plurality of images, at least in part based on the location of the images with respect to the 3D environment and the plurality of pixels of the plurality of images.
[0436] In some embodiments, the computer-executable instructions comprise instructions for identifying one or more frames associated with a map of a 3D environment for one or more keyframes, the one or more frames having a frame descriptor less than a threshold distance from a frame descriptor for a keyframe.
[0437] In some embodiments, the computer-executable instructions for calculating a frame descriptor comprise an artificial neural network.
[0438] In some embodiments, the artificial neural network is trained based on similar and different images, configured to receive, as input, a plurality of values representing features within an image, and provide, as output, a weighted combination of the plurality of values representing the features, and comprises a multi-layer perceptron unit and a max pooling unit configured to select a subset of the output of the multi-layer perceptron unit as the frame descriptor.
[0439] Some embodiments relate to a method of operating a computing system to generate a map of at least a portion of a three-dimensional (3D) environment based on sensor data collected by a device worn by a user. The method includes receiving a plurality of images captured by a device worn by the user, determining one or more keyframes from the plurality of images, identifying, using a first artificial neural network, one or more points of interest within the one or more keyframes, calculating, using the first artificial neural network, a feature descriptor for an individual point of interest, and for each of the one or more keyframes, calculating, using a second artificial neural network, a frame descriptor for representing the keyframe, at least in part, based on the calculated feature descriptors for the identified points of interest within the keyframe.
[0440] In some embodiments, the first and second artificial neural networks are sub-networks of an artificial neural network.
[0441] In some embodiments, the frame descriptor is unique to each key frame.
[0442] In some embodiments, one or more key frames each have a resolution higher than 1 megabyte. The frame descriptor for each one or more key frames is a column less than the number 512.
[0443] In some embodiments, each feature descriptor is a 32-byte column.
[0444] In some embodiments, the frame descriptor is generated by max-pooling the feature descriptors.
[0445] In some embodiments, the method includes receiving a new image captured by a device worn by a user, and identifying one or more nearest key frames in a database that include key frames used to generate a map, wherein the one or more nearest key frames have frame descriptors within a predetermined distance of the frame descriptor for the new image.
[0446] In some embodiments, the method includes performing feature matching on 3D map points of the map corresponding to the identified one or more nearest key frames, and calculating the pose of the device worn by the user based on the feature matching results.
[0447] In some embodiments, determining one or more key frames from a plurality of images includes comparing pixels of a first image with pixels of a second image captured immediately after the first image, and identifying the second image as a key frame when the difference between the pixels of the first image and the pixels of the second image exceeds or falls below a threshold.
[0448] In some embodiments, the method comprises training a second artificial neural network by generating a dataset comprising a plurality of image sets, each of the plurality of image sets including a query image, a positive sample image, and a negative sample image; calculating a loss for each of the plurality of image sets in the dataset by comparing the query image with the positive sample image and the negative sample image; and modifying the second artificial neural network based on the calculated loss such that a distance between a frame descriptor generated by the second artificial neural network for the query image and a frame descriptor for the positive sample image is greater than a distance between the frame descriptor for the query image and a frame descriptor for the negative sample image.
[0449] Some embodiments relate to a computing environment for a cross-reality system. The computing environment includes a database storing a plurality of maps. Each map comprises information representing an area of a 3D environment. The information representing each area comprises a frame descriptor representing an image of the area. The non-transitory computer storage medium stores computer-executable instructions that, when executed by at least one processor, process an image captured by a portable device by identifying a plurality of features in the image, calculate a feature descriptor for each of the plurality of features, calculate a frame descriptor for representing the image, at least in part, based on the calculated feature descriptors for one or more identified points of interest in the image, and select a map in the database based on a comparison between the calculated frame descriptor and a frame descriptor stored in the database of maps.
[0450] In some embodiments, the frame descriptor is unique to a frame stored in the database.
[0451] In some embodiments, the image has a resolution higher than 1 megabyte. The frame descriptor calculated for representing the image is a column less than the number 512.
[0452] In some embodiments, the computer-executable instructions include steps of processing a data set comprising a plurality of image sets, each of the plurality of image sets including a query image, a positive sample image, and a negative sample image; calculating a loss for a set of images of the plurality of image sets within the data set; comparing the query image with the positive sample image and the negative sample image; and modifying an artificial neural network based on the calculated loss such that a distance between a frame descriptor generated by the artificial neural network for the query image and a frame descriptor for the positive sample image is less than a distance between the frame descriptor for the query image and a frame descriptor for the negative sample image. The artificial neural network is trained by the above steps.
[0453] In some embodiments, the step of modifying the artificial neural network includes modifying a copy of the artificial neural network on a portable device within a computing environment.
[0454] In some embodiments, the computing environment includes a cloud platform and a plurality of portable devices communicating with the cloud platform. The cloud platform includes a database and computer-executable instructions for selecting a map. Computer-executable instructions for processing images captured by the portable devices are stored on the portable devices.
[0455] Some embodiments relate to an XR system including a first XR device including a first processor, a first computer-readable medium connected to the first processor, a first origin coordinate frame stored on the first computer-readable medium, a first destination coordinate frame stored on the computer-readable medium, a first data channel that receives data representing local content, a first coordinate frame converter executable by the first processor to transform the positioning of the local content from the first origin coordinate frame to the first destination coordinate frame, and a first display system adapted to display the local content to a first user after transforming the positioning of the local content from the first origin coordinate frame to the first destination coordinate frame.
[0456] Some embodiments relate to a viewing method including storing a first origin coordinate frame, storing a first destination coordinate frame, receiving data representing local content, transforming the positioning of the local content from the first origin coordinate frame to the first destination coordinate frame, and displaying the local content to a first user after transforming the positioning of the local content from the first origin coordinate frame to the first destination coordinate frame.
[0457] A map storage routine that stores a first map, which is a reference map having a plurality of persistent coordinate frames (PCFs), wherein each PCF of the first map has a set of coordinates; an object detection device positioned to detect the location of a real object; a PCF identification system connected to the object detection device and detecting a PCF of a second map based on the location of the real object, wherein each PCF of the second map has a set of coordinates; and a localization module connected to the reference map and the second map and executable to localize the second map relative to the reference map by matching a first PCF of the second map to a first PCF of the reference map and matching a second PCF of the second map to a second PCF of the reference map. Relates to an XR system.
[0458] In some embodiments, the object detection device is an object detection camera.
[0459] In some embodiments, the XR system further comprises a reference map incorporator connected to the reference map and the second map and executable to incorporate a third PCF of the reference map into the second map.
[0460] In some embodiments, the XR system further comprises a head unit comprising a head-mountable frame, wherein the object detection device is mounted on the head-mountable frame; a data channel receiving image data of local content; a local content positioning system connected to the data channel and executable to associate the local content with one PCF of the reference map; and a display system connected to the local content positioning system and displaying the local content.
[0461] In some embodiments, the XR system further comprises a local / world coordinate converter that converts the local coordinate frame of the local content into the world coordinate frame of the second map.
[0462] In some embodiments, the XR system further comprises a first world frame determination routine that calculates a first world coordinate frame based on the PCF of the second map, a first world frame storage instruction that stores the world coordinate frame, a head frame determination routine that calculates a head coordinate frame that changes in response to the movement of the head-mounted frame, a head frame storage instruction that stores the first head coordinate frame, and a world / head coordinate converter that converts the world coordinate frame into the head coordinate frame.
[0463] In some embodiments, the head coordinate frame changes with respect to the world coordinate frame when the head-mounted frame moves.
[0464] In some embodiments, the XR system further comprises at least one sound element associated with at least one PCF of the second map.
[0465] In some embodiments, the first and second maps are created by the XR device.
[0466] In some embodiments, the XR system further comprises first and second XR devices. Each XR device includes a head unit having a head-mounted frame, a data channel that receives image data of local content, a local content positioning system connected to the data channel and executable to associate the local content with one PCF of a reference map, and a display system connected to the local content positioning system and that displays the local content.
[0467] In some embodiments, the first XR device creates a PCF for a first map, the second XR device creates a PCF for a second map, and the positioning module forms part of the second XR device.
[0468] In some embodiments, the first and second maps are created in first and second sessions, respectively.
[0469] In some embodiments, the XR system further includes a server and a map download system that forms part of the XR device and downloads the first map from the server via a network.
[0470] In some embodiments, the positioning module repeatedly attempts to position the second map relative to the reference map.
[0471] In some embodiments, the XR system further includes a map publisher that uploads the second map to the server via a network.
[0472] Some embodiments relate to a visual recognition method, including the steps of storing a first map, which is a reference map having a plurality of PCFs, each PCF of the reference map having a set of coordinates; detecting the location of an actual object; detecting a PCF of a second map based on the location of the actual object, each PCF of the second map having a set of coordinates; and positioning the second map relative to the reference map by matching a first PCF of the second map to a first PCF of the first map and matching a second PCF of the second map to a second PCF of the reference map.
[0473] Some embodiments include a processor, a computer-readable medium connected to the processor, a plurality of reference maps on the computer-readable medium, and individual reference map identifiers on the computer-readable medium associated with each individual reference map, the reference map identifiers being different from each other and uniquely identifying the reference maps, a position detector on the computer-readable medium that is executable by the processor to receive and store a position identifier from an XR device, a first filter on the computer-readable medium that is executable by the processor to compare the position identifier and the reference map identifier and determine one or more reference maps that form a first filtered selection, and a map transmitter on the computer-readable medium that is executable by the processor to transmit one or more of the reference maps to the XR device based on the first filtered selection. An XR system includes a server that may have these components.
[0474] In some embodiments, each reference map identifier includes longitude and latitude, and the position identifier includes longitude and latitude.
[0475] In some embodiments, the first filter is a neighborhood area filter that selects at least one matching reference map that encompasses the longitude and latitude of the position identifier and at least one neighboring map that encompasses the longitude and latitude adjacent to the first matching reference map.
[0476] In some embodiments, the location identifier includes a WiFi fingerprint. The XR system further includes a WiFi fingerprint filter that is on a computer-readable medium and that, by a processor, determines latitude and longitude based on the WiFi fingerprint, compares the latitude and longitude from the WiFi fingerprint filter with the latitude and longitude of a reference map, determines one or more reference maps that form a second filtered selection within a first filtered selection, and a map transmitter that is executable to transmit the one or more reference maps based on the second selection and not transmit reference maps based on a first selection outside the second selection, comprising a second filter.
[0477] In some embodiments, the first filter is a WiFi fingerprint filter that is on a computer-readable medium and that, by a processor, is executable to determine latitude and longitude based on the WiFi fingerprint, compare the latitude and longitude from the WiFi fingerprint filter with the latitude and longitude of a reference map, and determine one or more reference maps that form a first filtered selection.
[0478] In some embodiments, the XR system further includes a multi-layer perceptron unit on a computer-readable medium, executable by a processor, that receives a plurality of features of an image and converts each feature into a separate column of numbers, and a max-pooling unit on a computer-readable medium, executable by a processor, that combines the maximum value of each column of numbers into a global feature column representing the image, wherein each reference map has at least one of the global feature columns, and the position identifier received from the XR device is advanced by the multi-layer perceptron unit and the max-pooling unit to determine the global feature column of the image, including features of an image captured by the XR device, a max-pooling unit, and a keyframe filter that compares the global feature column of the image with the global feature columns of the reference maps to determine one or more reference maps and forms a third filtered selection within a second filtered selection, wherein the map transmitter transmits one or more reference maps based on the third selection and does not transmit reference maps based on a second selection outside the third selection.
[0479] In some embodiments, the XR system further includes a multi-layer perceptron unit on a computer-readable medium, executable by a processor, that receives a plurality of features of an image and converts each feature into a separate column of numbers, and a max-pooling unit on a computer-readable medium, executable by a processor, that combines the maximum value of each column of numbers into a global feature column representing the image, wherein each reference map has at least one of the global feature columns, and the position identifier received from the XR device is advanced by the multi-layer perceptron unit and the max-pooling unit to determine the global feature column of the image, including features of an image captured by the XR device, a max-pooling unit, and a first filter that is a keyframe filter that compares the global feature column of the image with the global feature columns of the reference maps to determine one or more reference maps.
[0480] In some embodiments, the XR system comprises a head unit having a head-mounted frame, wherein an actual object detection device is mounted on the head-mounted frame, a data channel for receiving image data of local content, a local content positioning system connected to the data channel and executable to associate the local content with one of the PCFs of a reference map, and a display system connected to the local content positioning system and for displaying the local content.
[0481] In some embodiments, the XR device comprises a map storage routine for storing a first map, which is a reference map having a plurality of PCFs, wherein each PCF of the first map has a set of coordinates, an actual object detection device positioned to detect the location of an actual object, a PCF identification system connected to the actual object detection device and for detecting a PCF of a second map based on the location of the actual object, wherein each PCF of the second map has a set of coordinates, and a positioning module connected to the reference and second maps and executable to identify the second map relative to the reference map by matching a first PCF of the second map to a first PCF of the reference map and a second PCF of the second map to a second PCF of the reference map.
[0482] In some embodiments, the actual object detection device is an actual object detection camera.
[0483] In some embodiments, the XR system comprises a reference map incorporator connected to the reference map and the second map and executable to incorporate a third PCF of the reference map into the second map.
[0484] Some embodiments include storing a plurality of reference maps on a computer-readable medium, each reference map having an individual reference map identifier associated therewith, the reference map identifiers being different from one another and uniquely identifying the reference maps; receiving and storing a location identifier from an XR device using a processor connected to the computer-readable medium; determining one or more reference maps by comparing the location identifier and the reference map identifiers using the processor to form a first filtered selection; and transmitting the plurality of reference maps to the XR device based on the first filtered selection using the processor. This relates to a visualization method.
[0485] Some embodiments relate to an XR system including a processor, a computer-readable medium connected to the processor, a multi-layer perceptron unit on the computer-readable medium and executable by the processor to receive a plurality of features of an image and convert each feature into an individual number of columns, and a max pooling unit on the computer-readable medium and executable by the processor to combine the maximum value of each number of columns into a global feature column representing the image.
[0486] In some embodiments, the XR system includes a plurality of reference maps on a computer-readable medium, each reference map having at least one of the global feature columns associated therewith; a position detector that receives from the XR device features of an image captured by the XR device and that is processed by the multi-layer perceptron unit and the max pooling unit to determine a global feature column of the image; a keyframe filter that compares the global feature column of the image and the global feature columns of the reference maps to determine one or more reference maps that form part of a filtered selection; and a map transmitter on the computer-readable medium and executable by the processor to transmit one or more of the reference maps to the XR device based on the filtered selection.
[0487] In some embodiments, the XR system comprises a head unit having a head-mountable frame with an actual object detection device mounted on the head-mountable frame, a data channel for receiving image data of local content, a local content positioning system connected to the data channel and executable to associate the local content with one of the PCFs of the reference map, and a display system connected to the local content positioning system for displaying the local content.
[0488] In some embodiments, the XR system comprises a head unit having a head-mountable frame with an actual object detection device mounted on the head-mountable frame, a data channel for receiving image data of local content, a local content positioning system connected to the data channel and executable to associate the local content with one of the PCFs of the reference map, and a display system connected to the local content positioning system for displaying the local content, wherein the step of matching is performed by matching the global feature sequence of the second map to the global feature sequence of the reference map. The XR device includes the display system.
[0489] Some embodiments relate to a visual recognition method including receiving, using a processor, a plurality of features of an image, converting, using the processor, each feature into a column of individual numbers, and combining, using the processor, the maximum value of each column of numbers into a global feature sequence representing the image.
[0490] Some embodiments are methods of operating a computing system to identify one or more environmental maps stored in a database and merge them with a tracking map calculated based on sensor data collected by a device worn by a user, the device receiving a signal of an access point to a computer network while calculating the tracking map and determining at least one area attribute of the tracking map based on characteristics of communication with the access point, determining a geographical location of the tracking map based on the at least one area attribute, identifying a set of environmental maps stored in a database corresponding to the determined geographical location, filtering the set of environmental maps based on similarity of one or more identifiers of a network access point associated with the environmental maps of the tracking map and the set of environmental maps, filtering the set of environmental maps based on similarity of a metric representing the content of the environmental maps of the tracking map and the set of environmental maps, and filtering the set of environmental maps based on a degree of matching between a part of the tracking map and a part of the environmental maps of the set of environmental maps.
[0491] In some embodiments, the step of filtering the set of environmental maps based on similarity of one or more identifiers of a network access point includes retaining, within the set of environmental maps, an environmental map with the highest Jaccard similarity to at least one area attribute of the tracking map based on the one or more identifiers of the network access point.
[0492] In some embodiments, the step of filtering the set of environmental maps based on similarity of a metric representing the content of the environmental maps of the tracking map and the set of environmental maps includes retaining, within the set of environmental maps, an environmental map with a minimum vector distance between a vector of characteristics of the tracking map and a vector representing the environmental map within the set of environmental maps.
[0493] In some embodiments, the metrics representing the content of the tracking map and the environmental map include a vector of values calculated from the content of the map.
[0494] In some embodiments, the step of filtering the set of environmental maps based on the degree of matching between a portion of the tracking map and a portion of the environmental maps of the set of environmental maps includes calculating the volume of the physical world represented by the tracking map, which is also represented within the environmental maps of the set of environmental maps, and retaining, within the set of environmental maps, the environmental maps with a calculated volume larger than the environmental maps filtered out from the set.
[0495] In some embodiments, the set of environmental maps is first filtered based on the similarity of one or more identifiers, subsequently based on the similarity of the metrics representing the content, and subsequently based on the degree of matching between a portion of the tracking map and a portion of the environmental map.
[0496] In some embodiments, the filtering of the set of environmental maps based on the similarity of one or more identifiers, the similarity of the metrics representing the content, and the degree of matching between a portion of the tracking map and a portion of the environmental map is performed in an order based on the processing required to perform the filtering.
[0497] In some embodiments, the environmental map is selected based on the filtering of the set of environmental maps based on the similarity of one or more identifiers, the similarity of the metrics representing the content, and the degree of matching between a portion of the tracking map and a portion of the environmental map, and the information is loaded onto the user device from the selected environmental map.
[0498] In some embodiments, the environmental map is selected based on the filtering of the set of environmental maps based on the si...
Claims
1. A method for rendering virtual content in a 3D environment with a portable device by operating an electronic system, the method comprising: using one or more processors to maintain, on the portable device, a coordinate frame local to the portable device based on outputs of one or more sensors on the portable device, the coordinate frame local to the portable device describing the location of the electronic system in the 3D environment; obtaining a stored coordinate frame from stored spatial information about the 3D environment; calculating a transformation between the coordinate frame local to the portable device and the obtained and stored coordinate frame; receiving a specification of a virtual object having a coordinate frame local to the virtual object and a location of the virtual object relative to the obtained and stored coordinate frame; rendering the virtual object on a display of the portable device at a determined location based at least in part on the calculated transformation and the received location of the virtual object. A method comprising the above.
2. The method according to claim 1, wherein obtaining the stored coordinate frame comprises obtaining the coordinate frame through an application programming interface (API).
3. The portable device comprises a first portable device having a first processor among the one or more processors, the electronic system further comprises a second portable device having a second processor among the one or more processors, the processors on each of the first portable device and the second portable device obtain the same stored coordinate frame; calculate a transformation between a coordinate frame local to an individual device and the obtained same stored coordinate frame; receive the specification of the virtual object; render the virtual object on an individual display. The method according to claim 1, wherein the above steps are performed.
4. Each of the first portable device and the second portable device A camera configured to output a plurality of camera images, A keyframe generator configured to convert the plurality of camera images into a plurality of keyframes, A persistent pose calculator configured to generate a persistent pose by averaging the plurality of keyframes, A tracking map and persistent pose converter configured to determine the persistent pose with respect to the origin of the tracking map by converting the tracking map into the persistent pose, A persistent pose and persistent coordinate frame (PCF) converter configured to convert the persistent pose into a PCF, A map publisher configured to transmit spatial information including the PCF to a server The method according to claim 3, comprising:
5. The method according to claim 3, further comprising generating, by executing an application, a specification of the virtual object and a location of the virtual object with respect to the acquired and stored coordinate frame.
6. Maintaining a coordinate frame local to the portable device on the portable device includes, for each of the first portable device and the second portable device, Capturing a plurality of images of the 3D environment from one or more sensors of the portable device, Calculating one or more persistent poses based at least in part on the plurality of images, Generating spatial information about the 3D environment based at least in part on the calculated one or more persistent poses including, The method further includes transmitting the generated spatial information to a remote server for each of the first portable device and the second portable device, Obtaining the stored coordinate frame includes receiving the stored coordinate frame from the remote server, the method according to claim 3.
7. Calculating the one or more persistent poses based at least in part on the plurality of images includes Extracting one or more features from each of the plurality of images, Generating a descriptor for each of the one or more features, Generating a keyframe for each of the plurality of images based at least in part on the descriptor generating the one or more persistent postures based at least in part on the one or more key frames The method according to claim 6, comprising: **Claim 8** generating the one or more persistent postures comprises: selectively generating a persistent posture based on the portable device traveling a predetermined distance from the location of another persistent posture The method according to claim 7, comprising: **Claim 9** each of the first portable device and the second portable device comprises: a download system configured to download the stored coordinate frame from a server The method according to claim 3, comprising: **Claim 10** A portable device, the portable device comprising: one or more sensors configured to capture information about a 3D (three-dimensional) environment, the captured information comprising a plurality of images; at least one processor configured to generate a map of at least a portion of the 3D environment based on the plurality of images by executing computer-executable instructions comprising: the computer-executable instructions comprising: maintaining a local coordinate frame on the portable device based on an output of the one or more sensors; obtaining a stored coordinate frame from stored spatial information about the 3D environment; calculating a transformation between the local coordinate frame and the obtained and stored coordinate frame; receiving a specification of a virtual object having a virtual object coordinate frame and a location of the virtual object relative to the obtained and stored coordinate frame; and rendering the virtual object on a display of the portable device at a determined location based at least in part on the calculated transformation and the received location of the virtual object The portable device further comprising instructions for performing. **Claim 11** maintaining a local coordinate frame on the portable device comprises: calculating one or more persistent postures based at least in part on the plurality of images; and generating spatial information about the 3D environment based at least in part on the calculated one or more persistent postures The portable device according to claim 10, comprising: **Claim 12** The computer-executable instructions further include instructions for transmitting the generated spatial information to a remote server. The portable device according to claim 11, wherein obtaining the stored coordinate frame includes receiving the stored coordinate frame from the remote server.
13. Calculating the one or more persistent postures based at least in part on the plurality of images includes: extracting one or more features from each of the plurality of images; generating a descriptor for each of the one or more features; generating a keyframe for each of the plurality of images based at least in part on the descriptor; and generating the one or more persistent postures based at least in part on the one or more keyframes. The portable device according to claim 11.
14. Generating the one or more persistent postures includes: selectively generating a persistent posture based on the portable device that progresses a predetermined distance from the location of another persistent posture. The portable device according to claim 13.
15. An electronic system for maintaining persistent spatial information about a 3D environment for rendering virtual content on each of a plurality of portable devices, the electronic system comprising: a networked computing device; The networked computing device: includes at least one processor; at least one storage device connected to the processor; a map storage routine executable by the at least one processor to receive a plurality of maps from a portable device of the plurality of portable devices and store map information on the at least one storage device, each of the received plurality of maps comprising at least one coordinate frame, the at least one coordinate frame describing the location of the 3D environment in the plurality of maps; a map transmitter, the map transmitter: receiving location information from one of the plurality of portable devices; selecting one or more maps from among the stored plurality of maps; Transmitting information from the selected one or more maps to one of the plurality of portable devices, wherein the transmitted information comprises a coordinate frame of one of the selected one or more maps A map transmitter executable by using the at least one processor to perform the above An electronic system comprising the same **Claim 16** The coordinate frame A coordinate frame comprising information characterizing a plurality of features of an object in the 3D environment The electronic system according to claim 15, comprising a computer data structure comprising the same **Claim 17** The information characterizing the plurality of features comprises descriptors characterizing regions of the 3D environment. The electronic system according to claim 16 **Claim 18** Each coordinate frame of the at least one coordinate frame comprises persistent points characterized by features detected in sensor data representing the 3D environment. The electronic system according to claim 15 **Claim 19** Each coordinate frame of the at least one coordinate frame comprises a persistent pose. The electronic system according to claim 18 **Claim 20** Each coordinate frame of the at least one coordinate frame comprises a persistent coordinate frame. The electronic system according to claim 18
Citation Information
Patent Citations
System and method for augmented and virtual reality
JP2017107604A
Methods and systems for creating virtual and augmented reality.
JP2017529635A
Scalable 3D mapping system
US20160179830A1
Systems and methods for utilizing anchor graphs in mixed reality environments
US20180053315A1