Cross reality system with accurate shared maps
Patent Information
- Application Number
- JP2025076183
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-02-13
- Filing Date
- 2025-05-01
- Publication Date
- 2025-09-02
AI Technical Summary
Existing cross-reality systems face challenges in accurately merging environmental maps with tracking maps to ensure alignment with the direction of gravity, leading to potential distortions and reduced immersion in augmented and mixed reality experiences.
A method for merging environmental and tracking maps by determining a transformation that aligns the tracking map with the direction of gravity, ensuring that the transformed tracking map is correctly oriented, thereby preserving the alignment with the environmental map.
This approach enhances the accuracy and immersion of cross-reality experiences by reducing merge errors and maintaining alignment with the gravitational direction, providing a more realistic and seamless integration of virtual content with the physical world.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 975,983, filed on February 13, 2020, entitled "CROSS REALITY SYSTEM WITH ACCURATE SHARED MAPS", which is hereby incorporated by reference in its entirety for all purposes under 35 U.S.C. § 119(e).
[0002] This application generally relates to cross - reality systems.
Background Art
[0003] A computer can create a cross - reality (XR) environment that controls a human - user interface and in which part or all of the XR environment is generated by the computer as it is perceived by the user. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments in which part or all of the XR environment can be generated by the computer using data that describes the environment in part. This data can describe, for example, virtual objects that can be rendered so that a user can perceive or sense them as part of the physical world and interact with the virtual objects. The user can experience these virtual objects as a result of data being rendered and presented through a user - interface device such as a head - mounted display device. The data can be controlled to display so that it is visible to the user, or to play audio so that it is audible to the user, or to control a haptic (or tactile) interface to enable the user to experience a touch sensation as the user senses or perceives the virtual object.
[0004] An XR system can be useful for many applications spanning the fields of scientific visualization, medical training, engineering design and prototyping, remote operation and telepresence, and personal entertainment. In contrast to VR, AR and MR involve one or more objects in relation to real objects in the physical world. The experience of virtual objects interacting with real objects significantly enhances the user's enjoyment when using an XR system and also broadens the possibilities for various applications by presenting realistic and easily understandable information about how the physical world can be altered.
[0005] To realistically render virtual content, an XR system can construct a representation of the physical world around the user of the system. This representation may be constructed, for example, by processed images obtained using sensors on wearable devices that form part of the XR system. In such a system, the user may perform an initialization routine by looking around the room or other physical environment in which the user intends to use the XR system until the system has obtained sufficient information to construct a representation of that environment. As the system operates and the user moves around the environment or to other environments, sensors on the wearable device can obtain additional information and expand or update the representation of the physical world. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0006] Aspects of the present application relate to methods and apparatus for providing cross-reality (XR) scenarios. The techniques described herein may be used together, separately, or in any suitable combination.
[0007] According to one embodiment, a method is provided for merging one or more environmental maps stored in a database and a tracking map calculated based on sensor data collected by a device worn by a user. The method includes receiving the tracking map from the device, where the tracking map is aligned with respect to the direction of gravity; determining a transformation between the tracking map and the environmental map; and determining whether to merge the environmental map and the tracking map, where the step of determining whether to merge includes determining whether applying the transformation to the tracking map produces a transformed tracking map that is aligned with respect to the direction of gravity. The method may further include merging the environmental map and the tracking map based on determining that the transformed tracking map is aligned with the direction of gravity.
[0008] According to one embodiment, the step of determining the transformation may include selecting, as the determined transformation, a transformation that aligns with corresponding features with a metric of error below a threshold with respect to the corresponding features in the tracking map and the environmental map.
[0009] According to one embodiment, the method may further include determining corresponding features based on the similarity of identifiers assigned to the features.
[0010] According to one embodiment, the step of selecting as the determined transformation may further include selecting a transformation that, when applied, produces a transformed tracking map that is aligned with respect to the direction of gravity.
[0011] According to one embodiment, the step of determining the transformation may include applying a plurality of candidate transformations to the tracking map and selecting, as the determined transformation, a candidate transformation among the plurality of candidate transformations.
[0012] According to one embodiment, the step of selecting the determined transformation may further include selecting, as the determined transformation, a candidate transformation that, when applied, produces a transformed tracking map that aligns with the direction of gravity.
[0013] According to one embodiment, the step of determining whether applying a transformation to a tracking map produces a transformed tracking map that aligns with the direction of gravity includes determining whether applying the transformation to the tracking map produces a transformed tracking map that is rotated relative to the direction of gravity by more than a threshold amount, and in response to a determination that applying the transformation to the tracking map does not produce a transformed tracking map that is rotated relative to the direction of gravity by more than the threshold amount, selecting the transformation as the determined transformation, and in response to a determination that applying the transformation to the tracking map produces a transformed tracking map that is rotated relative to the direction of gravity by more than the threshold amount, discarding the transformation.
[0014] According to one embodiment, the method may further include identifying, from a database, a set of environment maps to be merged with the tracking map, determining, for each environment map in the set of environment maps, a transformation between the tracking map and the environment map, determining whether to merge the environment map and the tracking map, and merging the environment map and the tracking map based on a determination that the transformed tracking map aligns with the direction of gravity.
[0015] According to one embodiment, the method may further include, for each environment map in the set of environment maps, preventing the environment map and the tracking map from being merged based on a determination that the direction of gravity of the environment map does not align with the direction of gravity of the transformed tracking map.
[0016] According to one embodiment, the step of identifying a set of environment maps may include determining an area identifier associated with the tracking map and, at least in part, based on the area identifier associated with the tracking map, identifying the set of environment maps from a database.
[0017] According to one embodiment, the step of identifying the set of environment maps from the database may further include filtering the set of environment maps based on the similarity of one or more metrics associated with the tracking map and the environment maps within the set of environment maps.
[0018] According to one embodiment, a computing device is configured for use in a cross-reality system in which a portable device that operates within a three-dimensional (3D) environment renders virtual content. The computing device may include at least one processor, a computer-readable medium connected to the processor, a plurality of environment maps stored within the computer-readable medium, and computer-executable instructions configured to implement a method when executed by the at least one processor. The method implemented by the computer-executable instructions may include receiving, from the portable device, a tracking map, the tracking map being aligned with respect to the direction of gravity, and determining whether to merge the environment map and the tracking map, the step of determining whether to merge including searching for a transformation of the tracking map that aligns the transformed tracking map and the environment map in a manner that preserves the alignment of the transformed tracking map with respect to the direction of gravity, and merging the environment map and the transformed tracking map based on determining that the direction of gravity of the environment map aligns with the direction of gravity of the transformed tracking map.
[0019] According to one embodiment, the step of searching for a transformation may include the step of searching for a transformation that aligns a first set of features associated with a tracking map with a second set of features associated with an environmental map with an error metric below a threshold.
[0020] According to one embodiment, the step of searching for a transformation may include the step of searching for a transformation that does not change the orientation of the transformed tracking map with respect to the direction of gravity.
[0021] According to one embodiment, the step of searching for a transformation includes determining whether applying the transformation to the tracking map produces a transformed tracking map that is rotated by more than a threshold amount with respect to the direction of gravity, and in response to the determination that applying the transformation to the tracking map does not produce a transformed tracking map that is rotated by more than a threshold amount with respect to the direction of gravity, applying the transformation to the tracking map to generate a transformed tracking map, and in response to the determination that applying the transformation to the tracking map produces a transformed tracking map that is rotated by more than a threshold amount with respect to the direction of gravity, discarding the transformation.
[0022] According to one embodiment, the method further includes identifying a set of environmental maps to be merged with the tracking map from a plurality of environmental maps, determining for each environmental map in the set of environmental maps whether to merge the environmental map with the tracking map, and merging the environmental map with the transformed tracking map based on determining that the direction of gravity of the environmental map aligns with the direction of gravity of the transformed tracking map.
[0023] According to one embodiment, the method may further include, for each environmental map in the set of environmental maps, preventing the environmental map from being merged with the transformed tracking map based on determining that the direction of gravity of the environmental map does not align with the direction of gravity of the tracking map.
[0024] According to one embodiment, the step of identifying a set of environment maps may further include determining an area identifier associated with the tracking map, and identifying the set of environment maps based at least in part on the area identifier associated with the tracking map.
[0025] According to one embodiment, the step of identifying a set of environment maps may further include filtering the set of environment maps based on the similarity of one or more metrics associated with the tracking map and the environment maps within the set of environment maps.
[0026] According to one embodiment, a cloud computing environment for an extended reality system is configured for communication with a plurality of user devices equipped with sensors. The cloud computing environment for the extended reality system includes a map database that stores a plurality of environment maps constructed from data supplied by the plurality of user devices, and a non-transitory computer storage medium that stores computer-executable instructions that, when executed by at least one processor within the cloud computing environment, implement a method. The method includes receiving a tracking map from a user device, the tracking map being aligned with respect to the direction of gravity, and updating the map database based on the received tracking map, the step of updating the map database may include determining a transformation between the tracking map and an environment map for each environment map within the set of environment maps, and determining whether to merge the environment map and the tracking map, the step of determining whether to merge may include determining whether applying the transformation to the tracking map produces a transformed tracking map that is aligned with respect to the direction of gravity, and merging the environment map and the tracking map based on determining that the transformed tracking map is aligned with the direction of gravity.
[0027] According to one embodiment, the step of determining the transformation may include selecting, as the determined transformation, the transformation that aligns corresponding features with an error metric below a threshold with respect to corresponding features in the tracking map and the environment map.
[0028] According to one embodiment, the method may further include determining corresponding features based on the similarity of identifiers assigned to the features.
[0029] According to one embodiment, the step of selecting as the determined transformation may further include selecting the transformation that, when applied, produces a transformed tracking map that aligns with the direction of gravity.
[0030] According to one embodiment, the step of determining the transformation may include applying a plurality of candidate transformations to the tracking map and selecting, as the determined transformation, a candidate transformation among the plurality of candidate transformations.
[0031] According to one embodiment, the step of selecting as the determined transformation may further include selecting, as the determined transformation, a candidate transformation that, when applied, produces a transformed tracking map that aligns with the direction of gravity.
[0032] According to one embodiment, the step of determining whether applying the transformation to the tracking map produces a transformed tracking map that aligns with the direction of gravity may include determining whether applying the transformation to the tracking map produces a transformed tracking map that is rotated by more than a threshold amount with respect to the direction of gravity, selecting the transformation as the determined transformation in response to a determination that applying the transformation to the tracking map does not produce a transformed tracking map that is rotated by more than a threshold amount with respect to the direction of gravity, and discarding the transformation in response to a determination that applying the transformation to the tracking map produces a transformed tracking map that is rotated by more than a threshold amount with respect to the direction of gravity. The present invention provides, for example, the following items. (Item 1) A method for merging one or more environmental maps stored in a database and a tracking map calculated based on sensor data collected by a device worn by a user, the method comprising: Receiving the tracking map from the device, the tracking map being aligned with the direction of gravity; Determining a transformation between the tracking map and the environmental map; Determining whether to merge the environmental map and the tracking map, wherein determining whether to merge includes determining whether applying the transformation to the tracking map produces a transformed tracking map that aligns with the direction of gravity; Merging the environmental map and the tracking map based on determining that the transformed tracking map aligns with the direction of gravity comprising the method. (Item 2) The method according to item 1, wherein determining the transformation includes selecting, as the determined transformation, a transformation that aligns corresponding features with a metric of error below a threshold, with respect to corresponding features in the tracking map and the environmental map. (Item 3) The method according to item 2, further comprising determining the corresponding features based on the similarity of identifiers assigned to the features. (Item 4) The method according to item 2, wherein selecting as the determined transformation further includes selecting a transformation that, when applied, produces a transformed tracking map that aligns with the direction of gravity. (Item 5) The method according to item 1, wherein determining the transformation includes applying a plurality of candidate transformations to the tracking map and selecting one of the plurality of candidate transformations as the determined transformation. (Item 6) Selecting the determined transformation further includes selecting, as the determined transformation, a candidate transformation that, when applied, produces a transformed tracking map that aligns with the gravitational direction, as described in item 5 of the method. (Item 7) Determining whether applying the transformation to the tracking map produces a transformed tracking map that aligns with the gravitational direction Determining whether applying the transformation to the tracking map produces a transformed tracking map that is rotated relative to the gravitational direction by more than a threshold amount In response to determining that applying the transformation to the tracking map does not produce a transformed tracking map that is rotated relative to the gravitational direction by more than the threshold amount, selecting the transformation as the determined transformation Discarding the transformation in response to determining that applying the transformation to the tracking map produces a transformed tracking map that is rotated relative to the gravitational direction by more than the threshold amount The method described in item 1, including (Item 8) Identifying a set of environmental maps to be merged with the tracking map from the database For each environmental map in the set of environmental maps Determining a transformation between the tracking map and the environmental map Determining whether to merge the environmental map and the tracking map Based on determining that the transformed tracking map aligns with the gravitational direction, merging the environmental map and the tracking map The method described in item 1, further including (Item 9) For each environmental map in the set of environmental maps Based on determining that the gravitational direction of the environmental map does not align with the gravitational direction of the transformed tracking map, preventing the environmental map and the tracking map from being merged The method described in item 8, further including (Item 10) Identifying the set of environment maps includes determining an area identifier associated with the tracking map, and identifying the set of environment maps from the database, at least in part, based on the area identifier associated with the tracking map The method according to item 8, comprising. (Item 11) Identifying the set of environment maps from the database further includes filtering the set of environment maps based on the similarity of one or more metrics associated within the tracking map and the set of environment maps The method according to item 8, comprising. (Item 12) A portable device operating in a three-dimensional (3D) environment is a computing device configured for use in a cross-reality system that renders virtual content, the computing device comprising at least one processor, and a computer-readable medium connected to the processor, and a plurality of environment maps stored within the computer-readable medium, and computer-executable instructions that, when executed by the at least one processor, receive a tracking map from the portable device, the tracking map being aligned with respect to the direction of gravity, and determine whether to merge the environment map and the tracking map, determining whether to merge includes searching for a transformation of the tracking map that aligns the transformed tracking map and the environment map in a manner that preserves the alignment of the transformed tracking map with respect to the direction of gravity, and Merging the environmental map and the transformed tracking map based on determining that a gravitational direction of the environmental map aligns with a gravitational direction of the transformed tracking map Computer-executable instructions configured to implement a method including A computing device comprising (Item 13) The computing device according to item 12, wherein searching for the transformation includes searching for a transformation that aligns a first set of features associated with the tracking map with a second set of features associated with the environmental map with an error metric below a threshold (Item 14) The computing device according to item 13, wherein searching for the transformation includes searching for a transformation that does not change an orientation of the transformed tracking map with respect to the gravitational direction (Item 15) Searching for the transformation includes Determining whether applying the transformation to the tracking map produces a transformed tracking map that is rotated with respect to the gravitational direction by more than a threshold amount; Applying the transformation to the tracking map and generating the transformed tracking map in response to a determination that applying the transformation to the tracking map does not produce a transformed tracking map that is rotated with respect to the gravitational direction by more than the threshold amount; Discarding the transformation in response to a determination that applying the transformation to the tracking map produces a transformed tracking map that is rotated with respect to the gravitational direction by more than the threshold amount The computing device according to item 12, including (Item 16) The method further includes Identifying, from the plurality of environmental maps, a set of environmental maps to be merged with the tracking map; and For each environmental map in the set of environmental maps Determining whether to merge the environmental map and the tracking map merging the environmental map and the transformed tracking map based on determining that a gravity direction of the environmental map aligns with a gravity direction of the transformed tracking map The computing device according to item 12, comprising the above. (Item 17) The method further includes, for each environmental map in a set of environmental maps, preventing the environmental map from being merged with the transformed tracking map based on determining that a gravity direction of the environmental map does not align with a gravity direction of the tracking map The computing device according to item 16, comprising the above. (Item 18) Identifying the set of environmental maps further includes determining an area identifier associated with the tracking map, and identifying the set of environmental maps at least in part based on the area identifier associated with the tracking map The computing device according to item 16, comprising the above. (Item 19) Identifying the set of environmental maps further includes filtering the set of environmental maps based on a similarity of one or more metrics associated with the tracking map and environmental maps in the set of environmental maps The computing device according to item 16, comprising the above. (Item 20) A cloud computing environment for an augmented reality system configured for communication with a plurality of user devices equipped with sensors, a map database storing a plurality of environmental maps constructed from data supplied by the plurality of user devices, and a non-transitory computer storage medium storing computer-executable instructions that, when executed by at least one processor within the cloud computing environment, Receiving a tracking map from a user device, wherein the tracking map is aligned with the direction of gravity, and Updating the map database based on the received tracking map, and updating the map database is performed for each environmental map within the set of environmental maps, Determining a transformation between the tracking map and the environmental map, and Determining whether to merge the environmental map and the tracking map, and determining whether to merge includes determining whether applying the transformation to the tracking map produces a transformed tracking map that is aligned with the direction of gravity, and Based on determining that the transformed tracking map is aligned with the direction of gravity, merging the environmental map and the tracking map Including, and A non-transitory computer storage medium that implements a method including, and A cloud computing environment comprising. (Item 21) The cloud computing environment according to item 20, wherein determining the transformation includes selecting, as the determined transformation, a transformation that aligns corresponding features with an error metric below a threshold with respect to corresponding features in the tracking map and the environmental map. (Item 22) The cloud computing environment according to item 21, wherein the method further includes determining the corresponding features based on the similarity of identifiers assigned to the features. (Item 23) The cloud computing environment according to item 21, wherein selecting as the determined transformation further includes selecting a transformation that, when applied, produces a transformed tracking map that is aligned with the direction of gravity. (Item 24) Determining the transformation includes applying a plurality of candidate transformations to the tracking map and selecting a candidate transformation among the plurality of candidate transformations as the determined transformation, in the cloud computing environment according to item 20. (Item 25) Selecting as the determined transformation further includes selecting, as the determined transformation, a candidate transformation that, when applied, produces a transformed tracking map that aligns with the direction of gravity, in the cloud computing environment according to item 24. (Item 26) Determining whether applying the transformation to the tracking map produces a transformed tracking map that aligns with the direction of gravity includes determining whether applying the transformation to the tracking map produces a transformed tracking map that is rotated about the direction of gravity by more than a threshold amount and in response to determining that applying the transformation to the tracking map does not produce a transformed tracking map that is rotated about the direction of gravity by more than the threshold amount, selecting the transformation as the determined transformation and in response to determining that applying the transformation to the tracking map produces a transformed tracking map that is rotated about the direction of gravity by more than the threshold amount, discarding the transformation in the cloud computing environment according to item 20.
[0033] The foregoing description is provided by way of illustration and not of limitation.
Brief Description of the Drawings
[0034] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by like numerals. For purposes of clarity, not all components are labeled in all of the drawings.
[0035]
Figure 1
[0036]
Figure 2
[0037]
Figure 3
[0038]
Figure 4
[0039]
Figure 5A
[0040]
Figure 5B
[0041]
Figure 6A
[0042]
Figure 6B
[0043]
Figure 7
[0044]
Figure 8
[0045]
Figure 9
[0046]
Figure 10
[0047]
Figure 11
[0048]
Figure 12
[0049]
Figure 13
[0050]
Figure 14
[0051]
Figure 15
[0052]
Figure 16
[0053]
Figure 17
[0054]
Figure 18
[0055]
Figure 19
[0056]
Figure 20
[0057]
Figure 21
[0058]
Figure 22
[0059]
Figure 23
[0060]
Figure 24
[0061]
Figure 25
[0062]
Figure 26
[0063]
Figure 27
[0064]
Figure 28
[0065]
Figure 29
[0066]
Figure 30
[0067]
Figure 31A
[0068]
Figure 31B
[0069]
Figure 32
[0070]
Figure 33
[0071]
Figure 34
[0072]
Figure 35
Figure 36
[0073]
Figure 37
[0074]
Figure 38A
Figure 38B
[0075]
Figure 39-1
Figure 39-2
[0076]
Figure 40
[0077]
Figure 41
[0078]
Figure 42
[0079]
Figure 43A
[0080]
Figure 43B
[0081]
Figure 43C
[0082]
Figure 44
[0083]
Figure 45
[0084]
Figure 46A
Figure 46B
[0085]
Figure 47
[0086]
Figure 48
[0087]
Figure 49
[0088]
Figure 50
[0089]
Figure 51
[0090]
Figure 52
[0091]
Figure 53
[0092]
Figure 54
[0093]
Figure 55
Figure 56
[0094]
Figure 57
Figure 58
[0095]
Figure 59
[0096]
Figure 60
[0097]
Figure 61
[0098]
Figure 62
[0099]
Figure 63A
Figure 63B
Figure 63C
[0100]
Figure 64
[0101]
Figure 65
[0102]
Figure 66
[0103]
Figure 67
[0104]
Figure 68
[0105]
Figure 69
[0106] Detailed Description What is described herein are methods and apparatuses for providing an XR scene. To provide a realistic XR experience to multiple users, an XR system must determine the location of a user within the physical world in order to correctly correlate the location of virtual objects with real objects. The inventors have recognized and appreciated methods and apparatuses for localizing XR devices within large and very large environments (e.g., neighborhoods, cities, countries, the world) with reduced time and improved accuracy.
[0107] An XR system may construct an environmental map of a scene that can be created from images and / or depth information collected using sensors that are part of an XR device worn by a user of the XR system. Each XR device may develop a local map of its physical environment by integrating information from one or more images collected as the device moves. In some embodiments, the coordinate system of the map is tied to the orientation of the device when the device begins to scan the physical world. That orientation may change from session to session as the user interacts with the XR system, whether different sessions are associated with different users, each with their own wearable device with sensors for scanning the environment, or the same user using the same device at different times.
[0108] The XR system may implement one or more techniques to enable operations based on persistent spatial information. The techniques may, for example, enable persistent spatial information to be created, stored, and read by any of a plurality of users of the XR system, thereby providing an XR scene for a single or multiple users that is more computationally efficient and immersive. Persistent spatial information may also enable the rapid restoration and reset of head poses on each of one or more XR devices in a computationally efficient manner.
[0109] Persistent spatial information may be represented by a persistent map. The persistent map may be stored in a remote storage medium (e.g., the cloud). For example, a wearable device worn by a user may, after being turned on, read an appropriate map previously created and stored from a persistent storage device such as a cloud storage device. The previously stored map may be based on data about the environment collected using sensors on the user's wearable device during a previous session. Reading the stored map may enable the use of the wearable device without completing a scan of the physical world using sensors on the wearable device. Alternatively, or in addition, the system / device may similarly read an appropriate stored map in response to entering a new area of the physical world.
[0110] The stored map may be represented in a canonical form to which a local reference frame on each XR device can be related. In a multi-device XR system, a stored map accessed by one device may be created and stored by another device and / or may be constructed by aggregating data about the physical world collected by sensors on a plurality of wearable devices that pre-exist within at least a portion of the physical world represented by the stored map.
[0111] In some embodiments, the persistent spatial information may be represented in a way that can be easily shared among users and among distributed components, including applications. The reference map may provide information about the physical world, for example, as a Persistent Coordinate Frame (PCF). The PCF may be defined based on a set of features recognized within the physical world. The features may be selected such that they are likely to be the same for each user session of the XR system. The PCF may be sparsely present so that they can be efficiently processed and transferred, and may provide less than all of the available information about the physical world. Techniques for processing the persistent spatial information may include creating a dynamic map based on the local coordinate systems of one or more devices across one or more sessions. These maps may be sparse maps representing the physical world, based on a subset of feature points detected in the images used to form the map. The Persistent Coordinate Frame (PCF) may be generated from the sparse map and may be exposed to the XR application, for example, via an Application Programming Interface (API). These capabilities may be supported by techniques for forming a reference map by merging multiple maps created by one or more XR devices.
[0112] The relationship between the reference map and the local map may be determined for each device through a localization process. The localization process may be performed on each XR device based on a selected set of reference maps that are sent to the device. Alternatively, or in addition, the localization service may be provided on a remote processor such that it can be implemented within the cloud.
[0113] Sharing data about the physical world among multiple devices can enable a shared user experience of virtual content. For example, two XR devices with access to the same stored map can both be localized with respect to the stored map. Once localized, the user device may render virtual content having a location defined by reference to the stored map by converting that location to a reference frame maintained by the user device. The user device may use this local reference frame to control the display of the user device and render the virtual content at the defined location.
[0114] To support these and other functions, the XR system may include components that develop, maintain, and use persistent spatial information, including one or more stored maps, based on data about the physical world collected using sensors on the user device. These components may be distributed across the XR system, and some may operate, for example, on the head-mounted portion of the user device. Other components may operate on a computer associated with the user that is coupled to the head-mounted portion via a local or personal area network. Still others may operate at a remote location, such as one or more servers accessible via a wide area network.
[0115] These components may include components that can identify information of sufficient quality to be stored as, or within, a persistent map, for example, from information about the physical world collected by one or more user devices. An example of such a component, described in more detail below, is a map merge component. Such a component may, for example, receive input from a user device and determine the suitability of portions of the input for use in updating a persistent map, which may be in a canonical form. The map merge component may also, for example, promote a local map from a user device that is not merged with the persistent map to a separate persistent map.
[0116] As another example, these components may include components that can assist in selecting an appropriate set of one or more persistent maps that are likely to represent the same region of the physical world represented by location information provided by a user device. An example of such a component, described in more detail below, is a map ranking and map selection component. Such a component may, for example, receive input from a user device and identify one or more persistent maps that are likely to represent the region of the physical world in which the device is operating. The map ranking component may, for example, assist in selecting a persistent map to be used by the local device as it renders virtual content, collects data about the environment, or performs other actions. The map ranking component may alternatively, or in addition, assist in identifying a persistent map to be updated as additional information about the physical world is collected by one or more user devices.
[0117] Further, other components may determine a transformation that converts information captured or described relative to one reference frame to another reference frame. For example, sensors may be attached to a head-mounted display such that data read from those sensors indicates the location of an object within the physical world relative to the wearer's head pose. One or more transformations may be applied to relate that location information to a coordinate frame associated with a persistent environment map. Similarly, data indicating where a virtual object should be rendered when represented within the coordinate frame of the persistent environment map may also undergo one or more transformations to be within the reference frame of a display above the user's head. As described in more detail below, there may be multiple such transformations. These transformations may be partitioned across components of the XR system so that they can be efficiently updated and / or applied within a distributed system.
[0118] In some embodiments, the persistent map may be constructed from information collected by multiple user devices. Each XR device may capture local spatial information and construct a separate tracking map using information collected by each device's sensors at various locations and times. Each tracking map may include points that may be associated with features of real objects, each of which may include multiple features. Potentially, in addition to supplying input for creating and maintaining the persistent map, the tracking maps may be used to track the movement of users within the scene and enable the XR system to estimate the head pose of individual users relative to the reference frame established by the tracking map on that user's device.
[0119] This co-dependency between map creation and head pose estimation constitutes a significant challenge. Substantial processing may be required to simultaneously create a map and estimate the head pose. The processing must be performed quickly as objects move within the scene (e.g., moving a cup on a table) and as the user moves within the scene, since latency can make the XR experience for the user less realistic. On the other hand, XR devices may provide limited computational resources since the XR device should be lightweight for the user to wear comfortably. The lack of computational resources cannot be compensated for using more sensors since the addition of sensors would also undesirably add weight. Furthermore, either more sensors or more computational resources would lead to heat, which can cause deformation of the XR device.
[0120] An XR system may be configured to create, share, and use persistent spatial information with low usage of computational resources and / or short latency to provide a more immersive user experience. Some such techniques may enable efficient comparison of spatial information. Such a comparison may occur, for example, as part of localization, in which a set of features from a local device is matched to a set of features in a reference map.
[0121] Similarly, in a map merge process, an attempt may be made to match one or more sets of features in a tracking map from a device to corresponding features in a reference map and determine a transformation between the corresponding sets of features that provides a suitably low error between the positions of the transformed features in a first set of features derived from the tracking map and a second set of features derived from the reference map. Subsequent processing to incorporate the tracking map into a set of reference maps may be based on the result of that comparison. For example, determining a transformation with suitably low error may indicate that the region represented by the second set of features derived from the reference map corresponds to the same region represented by the first set of features derived from the tracking map and that the two maps may be merged.
[0122] The inventors recognize that, when aligning sets of features, errors can be introduced into the merge process even when there are low errors. The inventors further recognize and understand that such errors can be detected by a transformation that changes the orientation of the tracking map with respect to the direction of gravity, and that merging such a tracking map with a reference map can result in a distorted merged map.
[0123] By suppressing the merging of a transformed tracking map having an orientation changed with respect to gravity, the reference map can initiate and retain alignment with respect to gravity. The merge error can be reduced by ensuring that the direction of gravity of the transformed (i.e., after applying the determined transformation) tracking map aligns with the direction of gravity of the reference map with which the tracking map is to be merged.
[0124] The techniques described herein may be used together or separately with many types of devices, including wearable or portable devices with limited computational resources that provide augmented or mixed reality scenarios, and for many types of scenarios. In some embodiments, the techniques may be implemented by one or more services that form part of an XR system.
[0125] AR System Overview
[0126] Figures 1 and 2 illustrate scenes with virtual content that are presented in conjunction with a portion of the physical world. For illustrative purposes, an AR system is used as an example of an XR system. Figures 3 - 6B illustrate an exemplary AR system that can operate in accordance with the techniques described herein and includes one or more processors, memory, sensors, and a user interface.
[0127] Referring to FIG. 1, an outdoor AR scene 354 is depicted, and to the user of AR technology, a physical-world park-like setting 356 is visible, featuring people, trees, buildings in the background, and a concrete platform 358. In addition to these items, the user of AR technology also "sees" and perceives a robot image 357 standing on the physical-world concrete platform 358 and a flying comic-like avatar character 352 that appears to be an anthropomorphization of a bumblebee, but these elements (e.g., avatar character 352 and robot image 357) do not exist within the physical world. Due to the extreme complexity of human vision and the nervous system, it is difficult to produce AR technology that promotes a comfortable, natural, and rich presentation of virtual image elements among other virtual or physical-world image elements.
[0128] Such an AR scene can be achieved using a system that enables a user to place AR content within the physical world, determines the location within the map of the physical world where the AR content is placed, saves the AR scene so that the placed AR content can be reloaded for display within the physical world, for example, during different AR experience sessions, and enables multiple users to share the AR experience, based on tracking information to construct a map of the physical world. This system can construct and update a digital representation of the physical-world surface around the user. This representation may be used, in whole or in part, to render virtual content so that it appears to be occluded by physical objects between the user and the rendered location of the virtual content for purposes of placing virtual objects, in physics-based interactions, and for virtual character path planning and navigation, or for other operations where information about the physical world is used.
[0129] Figure 2 depicts another example of an indoor AR scene 400 according to some embodiments and shows an exemplary use case of an XR system. The exemplary scene 400 is a living room having a wall, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of AR technology may also perceive virtual objects such as an image on the wall behind the sofa, a bird flying in through the door, a deer peeking out from the bookshelf, and a decoration in the form of a windmill placed on the coffee table.
[0130] Regarding the image on the wall, AR technology requires information about objects and surfaces within the room, such as the shape of a lamp, not only for the surface of the wall but also to occlude the image in order to correctly render the virtual object. Regarding the flying bird, AR technology requires information about all objects and surfaces around the room to render the bird using realistic physics so that the bird avoids objects and surfaces or bounces back if it collides. Regarding the deer, AR technology requires information about surfaces such as the floor or the coffee table to calculate where to place the deer. Regarding the windmill, the system may be able to identify that it is a separate object from the table and determine that it is movable, while the corner of the shelf or the wall may be determined to be stationary. Such specificities may be used in determining the parts of the scene that are used or updated in each of the various operations.
[0131] The virtual object may be placed within a previous AR experience session. When a new AR experience session starts in the living room, the AR technology requires that the virtual object be accurately displayed at the previously placed location and be realistically visible from different viewpoints. For example, the windmill should be displayed as standing on the book rather than floating above the table even at different locations without the book. Such floating can occur when the location of the user within the new AR experience session is not accurately located within the living room. As another example, when the user views the windmill from a viewpoint different from the viewpoint when the windmill was placed, the AR technology requires the corresponding side of the displayed windmill.
[0132] The scene may be presented to the user via a system that includes a plurality of components, including a user interface that can stimulate one or more user perceptions such as vision, hearing, and / or touch. In addition, the system may include one or more sensors that can measure parameters of the physical part of the scene, including the position and / or movement of the user within the physical part of the scene. Further, the system may include one or more computing devices with associated computer hardware such as memory. These components may be integrated within a single device or distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated within a wearable device.
[0133] Figure 3 depicts an AR system 502 configured to provide an experience of AR content that interacts with the physical world 506 according to some embodiments. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by a user as part of a headset such that the user can wear the display across their eyes, like a pair of goggles or glasses. At least a portion of the display may be transparent such that the user can observe the see-through reality 510. The see-through reality 510 may correspond to the portion of the physical world 506 within the current viewpoint of the AR system 502, which may correspond to the user's viewpoint when the user wears a headset incorporating both the display and sensors of the AR system and obtains information about the physical world.
[0134] The AR content may also be presented on the display 508, overlaid on the see-through reality 510. To provide an accurate interaction between the AR content and the see-through reality 510 on the display 508, the AR system 502 may include a sensor 522 configured to capture information about the physical world 506.
[0135] The sensor 522 may include one or more depth sensors that output a depth map 512. Each depth map 512 may have a plurality of pixels that may each represent the distance to a surface within the physical world 506 in a particular direction relative to the depth sensor. Raw depth data may originate from the depth sensor and may create the depth map. Such depth maps may be updated as fast as the depth sensor can form a new image, which may be hundreds or thousands of times per second. However, the data is noisy and incomplete and may have holes, shown as black pixels on the illustrated depth map.
[0136] The system may include other sensors such as an image sensor. The image sensor may obtain monocular or stereoscopic information that can be processed to represent the physical world in other ways. For example, the image may be processed within a world reconstruction component 516 to create a mesh representing connected portions of objects within the physical world. For example, metadata about such objects, including color and surface texture, may similarly be obtained using the sensor and stored as part of the world reconstruction.
[0137] The system may also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, the system's head pose tracking component may be used to calculate the head pose in real time. The head pose tracking component may represent the user's head pose within a coordinate frame with six degrees of freedom, including, for example, translations along three perpendicular axes (e.g., forward / backward, up / down, left / right) and rotations about the three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, the sensor 522 may include an inertial measurement unit that can be used to calculate and / or determine the head pose 514. The head pose 514 for a depth map may indicate, for example, the current viewpoint of the sensor capturing the depth map with six degrees of freedom, but the head pose 514 may also be used for other purposes such as associating image information with a particular portion of the physical world or associating the position of a display worn on the user's head with the physical world.
[0138] In some embodiments, the head pose information may be derived by methods other than the IMU, such as from the analysis of objects in the image. For example, the head pose tracking component may calculate the relative position and orientation of the AR device with respect to the physical object based on the visual information captured by the camera and the inertial information captured by the IMU. The head pose tracking component may then calculate the head pose of the AR device, for example, by comparing the calculated relative position and orientation of the AR device with respect to the physical object with the features of the physical object. In some embodiments, the comparison is made using one or more of the sensors 522 that are stable over time such that changes in the position of these features in the images captured over time can be associated with changes in the user's head pose, by identifying features in the images captured using the sensors 522.
[0139] The inventors have realized techniques for operating an XR system to provide an XR scenario for a more immersive user experience, such as estimating head pose at a frequency of 1 kHz, with low usage of the calculation resources connected to the XR device, which may be configured with, for example, four video graphics array (VGA) cameras operating at 30 Hz, one inertial measurement unit (IMU) operating at 1 kHz, the computing power of a single advanced RISC machine (ARM) core, less than 1 GB of memory, and a network with a bandwidth of less than 100 Mbps. These techniques relate to reducing the processing required to generate and maintain a map and estimate head pose, and to providing and consuming data with low calculation overhead. The XR system may calculate its pose based on the matched visual features. U.S. Patent Application No. 16 / 221,065 describes hybrid tracking and is incorporated herein by reference in its entirety.
[0140] In some embodiments, the AR device may construct a map from feature points recognized in successive images within a series of image frames captured as the user moves through the physical world with the AR device. Each image frame may be obtained from a different pose as the user moves, but the system may adjust the orientation of the features of each successive image frame and match the orientation of the initial image frame by matching the features of the successive image frames with previously captured image frames. Translation of the successive image frames may be used to align each successive image frame and match the orientation of the previously processed image frame such that points representing the same feature will match corresponding feature points from previously collected image frames. The frames within the resulting map may have a common orientation established when the first image frame is added to the map. The map may be used to determine the pose of the user within the physical world by matching features from the current image frame with the set of feature points within the common reference frame. In some embodiments, the map may be referred to as a tracking map.
[0141] In addition to enabling tracking of the pose of the user within the environment, the map may enable other components of the system, such as the world reconstruction component 516, to determine the location of physical objects relative to the user. The world reconstruction component 516 may receive the depth map 512 and the head pose 514 and any other data from the sensors and integrate that data into the reconstruction 518. The reconstruction 518 may be more complete and less noisy than the sensor data. The world reconstruction component 516 may update the reconstruction 518 using the spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0142] The reconstruction 518 may include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats may represent alternative representations of the same part of the physical world or different parts of the physical world. In the illustrated embodiment, on the left side of the reconstruction 518, a part of the physical world is presented as a global surface, and on the right side of the reconstruction 518, a part of the physical world is presented as a mesh.
[0143] In some embodiments, the map maintained by the head pose component 514 may be sparse relative to other maps of the physical world that can be maintained. Instead of providing information about locations and other characteristics of surfaces as possibilities, the sparse map may indicate locations of points of interest and / or structures such as corners or edges. In some embodiments, the map may include image frames as captured by the sensor 522. These frames may be reduced to features that may represent points of interest and / or structures. Along with each frame, information about the user's pose from which the frame was obtained may also be stored as part of the map. In some embodiments, all images obtained by the sensor may or may not be stored. In some embodiments, the system may process the images as they are collected by the sensor and select a subset of the image frames for further calculations. The selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may, for example, add a new image frame to the map based on its overlap with previous image frames already added to the map or based on an image frame that contains a sufficient number of features determined to be likely to represent stationary objects. In some embodiments, the selected image frame or group of features from the selected image frame may serve as a keyframe for the map, which is used to provide spatial information.
[0144] In some embodiments, the amount of data processed when constructing a map may be reduced by constructing a sparse map with a set of mapped points and keyframes and / or splitting the map into blocks to enable per-block updates. The mapped points may be associated with points of interest in the environment. The keyframes may include information selected from camera capture data. U.S. Patent Application No. 16 / 520,582 describes the steps of determining and / or evaluating a localization map and is hereby incorporated by reference in its entirety.
[0145] The AR system 502 may integrate sensor data from multiple viewpoints of the physical world over time. The pose of the sensor (e.g., position and orientation) may be tracked as the device including the sensor is moved. As the frame pose of the sensor and how it relates to other poses are understood, these multiple viewpoints of the physical world may each be fused together into a single combined reconstruction of the physical world, which may serve as an abstraction layer for the map and provide spatial information. The reconstruction may be more complete and less noisy than the original sensor data by using spatial and temporal averaging (i.e., averaging of data from multiple viewpoints over time) or any other suitable method.
[0146] In the embodiment illustrated in FIG. 3, the map represents a portion of the physical world in which a user of a single wearable device is present. In that scenario, the head pose associated with a frame in the map may be represented as a local head pose indicating the orientation relative to the initial orientation for the single device at the start of the session. For example, the head pose may be tracked relative to the initial head pose when the device is turned on or otherwise operated to scan the environment and construct a representation of that environment.
[0147] In combination with content that characterizes that part of the physical world, the map may include metadata. The metadata may indicate, for example, the capture time of sensor information used to form the map. The metadata may alternatively or additionally indicate the location of the sensor at the capture time of the information used to form the map. The location may be represented directly, using information from a GPS chip etc., or indirectly, using a wireless (e.g., Wi-Fi) signature etc. that indicates the strength of a signal received from one or more wireless access points while the sensor data was being collected, and / or using an identifier such as the BSSID of the wireless access point to which the user device was connected while the sensor data was being collected.
[0148] The reconstruction 518 may be used for AR functions such as the production of a surface representation of the physical world for occlusion processing or physics-based processing. This surface representation may change as the user moves or as objects within the physical world change. The side of the reconstruction 518 may be used, for example, by component 520 to produce a changing global surface representation in world coordinates that can be used by other components.
[0149] AR content may be generated based on this information by an AR application 504 etc. The AR application 504 may be, for example, a game program that implements one or more functions based on information about the physical world such as visual occlusion, physics-based interactions, and environmental inference. This may be done by querying data in a different format from the reconstruction 518 produced by the world reconstruction component 516 to implement these functions. In some embodiments, component 520 may be configured to output an update when the representation within the area of interest of the physical world changes. The area of interest may be set to approximate a part of the physical world in the vicinity of the user of the system, such as a part within the user's field of view, or projected (predicted / decided) to enter the user's field of view.
[0150] The AR application 504 may use this information to generate and update AR content. The virtual part of the AR content may be presented on the display 508 in combination with see-through reality 510 to create a realistic user experience.
[0151] In some embodiments, the AR experience may be part of a system that can include remote processing and / or remote data storage devices, a wearable display device, and / or, in some embodiments, other XR devices worn by other users, and may be provided to the user through an XR device. FIG. 4 illustrates an example of a system 580 (hereinafter referred to as the "system 580") that includes a single wearable device for purposes of illustration. The system 580 includes a head-mounted display device 562 (hereinafter referred to as the "display device 562") and various mechanical and electronic modules and systems that support the functions of the display device 562. The display device 562 may be coupled to a frame 564, which is wearable by a user or viewer 560 (hereinafter referred to as the "user 560") of the display system and is configured to position the display device 562 in front of the eyes of the user 560. According to various embodiments, the display device 562 may be a sequential display. The display device 562 may be monocular or binocular. In some embodiments, the display device 562 may be an example of the display 508 in FIG. 3.
[0152] In some embodiments, speaker 566 is coupled to frame 564 and positioned proximate to the ear canal of user 560. In some embodiments, another speaker, not shown, is positioned adjacent to the other ear canal of user 560 to provide stereo / adjustable sound control. Display device 562 is operably coupled to local data processing module 570 by means such as a wired conductor or wireless connectivity 568, which may be mounted in a variety of configurations, such as fixed to a helmet or hat worn by user 560, fixed to a headset incorporated within frame 564, or otherwise removably attached to user 560 (e.g., in a backpack configuration, in a belt attachment configuration).
[0153] Local data processing module 570 may include digital memory such as a processor and non-volatile memory (e.g., flash memory), both of which may be utilized to assist in the processing, caching, and storage of data. The data includes a) data captured from sensors such as an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope (e.g., operably coupled to frame 564 or otherwise attachable to user 560), and / or b) data that may potentially be obtained and / or processed using remote processing module 572 and / or remote data repository 574 for passage to display device 562 after processing or retrieval.
[0154] In some embodiments, the wearable device may communicate with remote components. The local data processing module 570 may be operably coupled to the remote processing module 572 and the remote data repository 574 via communication links 576, 578, such as a wired or wireless communication link, respectively, such that these remote modules 572, 574 are operably coupled to each other and are available as resources to the local data processing module 570. In further embodiments, in addition to or instead of the remote data repository 574, the wearable device may be able to access cloud-based remote data repositories and / or services. In some embodiments, the head pose tracking component described above may be implemented at least partially within the local data processing module 570. In some embodiments, the world reconstruction component 516 in FIG. 3 may be implemented at least partially within the local data processing module 570. For example, the local data processing module 570 may be configured to execute computer-executable instructions and generate a map and / or a physical world representation based at least in part on at least a portion of the data.
[0155] In some embodiments, the processing may be distributed across local and remote processors. For example, local processing may be used to construct a map (e.g., a tracking map) on the user's device based on sensor data collected using sensors on that user's device. Such a map may be used by an application on that user's device. Additionally, previously created maps (e.g., canonical maps) may be stored within a remote data repository 574. If a suitable stored or persistent map is available, it may be used instead of, or in addition to, a tracking map created locally on the device. In some embodiments, the tracking map may be geolocated with respect to a stored map such that the correspondence can be oriented with respect to the location of the wearable device at the time the user turned the system on, a tracking map, and can be oriented with respect to one or more persistent features, a canonical map. In some embodiments, the persistent map may be loaded onto the user's device and enable the rendering of virtual content without the latency associated with scanning the location where the user's device constructs a tracking map of the user's complete environment from sensor data obtained during the scan. In some embodiments, the user's device may access a remote persistent map (e.g., stored in the cloud) without having to download the persistent map onto the user's device.
[0156] In some embodiments, spatial information may be communicated from the wearable device to a remote service, such as a cloud service, configured to locate the device and store it in a map maintained on the cloud service. According to one embodiment, the location determination process may occur within the cloud, returning a transformation that matches the device location to an existing map, such as a reference map, and links virtual content to the wearable device location. In such embodiments, the system can avoid communicating the map from the remote resource to the wearable device. Other embodiments are configured for both device-based and cloud-based location determination and can enable functionality, for example, when network connectivity is unavailable or the user chooses not to enable cloud-based location determination.
[0157] Alternatively, or in addition, the tracking map may be merged with previously stored maps to extend those maps or improve their quality. The process for determining whether a suitable previously created environmental map is available and / or for merging the tracking map with one or more stored environmental maps may be performed within the local data processing module 570 or the remote processing module 572.
[0158] In some embodiments, the local data processing module 570 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the computational budget of the local data processing module 570 but enable a smaller device. In some embodiments, the world reconstruction component 516 may use a computational budget less than that of a single advanced RISC machine (ARM) core so that the remaining computational budget of the single ARM core can be accessed for other uses such as mesh extraction, etc., and generate a physical world representation in real time in an unspecified space.
[0159] In some embodiments, the remote data repository 574 may include a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, enabling completely autonomous use from the remote module. In some embodiments, all data is stored and all or most of the computations are performed within the remote data repository 574, enabling a smaller device. The world reconstruction may be stored, for example, in whole or in part, within this repository 574.
[0160] In embodiments where the data is remotely stored and accessible via a network, the data may be shared by multiple users of the augmented reality system. For example, the user device may upload its tracking map and it may be extended into a database of environmental maps. In some embodiments, the upload of the tracking map occurs at the end of a user session with the wearable device. In some embodiments, the upload of the tracking map may occur persistently, semi - persistently, intermittently, at a predefined time, after a predefined period from the previous upload, or triggered by an event. The tracking map uploaded by any user device may be used to extend or improve previously stored maps, regardless of whether it is based on data from that user device or any other user device. Similarly, the persistent map downloaded to the user device may be based on data from that user device or any other user device. Thus, a high - quality environmental map may be readily available to the user to improve the experience using the AR system.
[0161] In further embodiments, the download of the persistent map may be limited and / or avoided based on localization performed on remote resources (e.g., in the cloud). In such a configuration, the wearable device or other XR device communicates to the cloud service feature information (e.g., positioning information regarding the device at the time the feature represented in the feature information was sensed) coupled with pose information. One or more components of the cloud service may match the feature information with an individual stored map (e.g., a reference map) and generate a transformation between the coordinate systems of the tracking map maintained by the XR device and the reference map. Each XR device having the tracking map localized with respect to the reference map may accurately render virtual content at locations defined with respect to the reference map based on its own tracking.
[0162] In some embodiments, the local data processing module 570 is operatively coupled to the battery 582. In some embodiments, the battery 582 is a removable power source such as a commercially available battery. In other embodiments, the battery 582 is a lithium-ion battery. In some embodiments, the battery 582 allows the user 560 to operate the system 580 for longer time periods without being connected to a power source, without having to charge a lithium-ion battery, or without having to shut down the system 580 and replace the battery, including both an internal lithium-ion battery rechargeable by the user 560 during non-operating times of the system 580 and a removable battery.
[0163] FIG. 5A illustrates a user 530 wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as "environment 532"). Information captured by the AR system along the user's movement path may be processed into one or more tracking maps. The user 530 positions the AR display system at location 534, and the AR display system records ambient information about the passable world relative to location 534 (e.g., a digital representation of real objects in the physical world that can be memorized and updated as real objects in the physical world change). That information may be stored as a pose in combination with an image, feature, directional audio input, or other desired data. Location 534 is aggregated with respect to data input 536, for example, as part of a tracking map, and processed at least by the passable world module 538, which may be implemented, for example, by processing on the remote processing module 572 of FIG. 4. In some embodiments, the passable world module 538 may include a head pose component 514 and a world reconstruction component 516 such that the processed information can indicate the location of an object in the physical world in combination with other information about the physical object used in the rendered virtual content.
[0164] The passable world module 538 determines, at least in part, the location and manner in which AR content 540 can be placed within the physical world, as determined from the data input 536. The AR content is "placed" within the physical world by presenting both a representation of the physical world and the AR content via a user interface, and the AR content is rendered as if it were interacting with objects within the physical world, and the objects within the physical world are presented as if the AR content were obscuring the user's view of those objects when appropriate. In some embodiments, the AR content may be placed by appropriately selecting a portion of a fixed element 542 (e.g., a table) from the reconstruction (e.g., reconstruction 518) and determining the shape and position of the AR content 540. As an example, the fixed element may be a table, and the virtual content may be positioned to appear on that table. In some embodiments, the AR content may be placed within a structure within the field of view 544, which may be the current field of view or an estimated future field of view. In some embodiments, the AR content may be persisted with respect to a model 546 (e.g., a mesh) of the physical world.
[0165] As described, the fixed element 542 serves as a proxy (e.g., a digital copy) for any fixed element in the physical world that can be stored within the passable world module 538 such that the user 530 can perceive content on the fixed element 542 each time it becomes visible to the user 530 without the system having to map to the fixed element 542. The fixed element 542 may thus be a mesh model that is stored by the passable world module 538 for future reference by multiple users, even though it is determined from a previous modeling session or from a different user. Thus, the passable world module 538 recognizes the environment 532 from a previously mapped environment and can display AR content without the user 530's device having to first map all or part of the environment 532, saving calculation processes and cycles and avoiding latency for any rendered AR content.
[0166] The mesh model 546 of the physical world may be created by the AR display system, interacts with the AR content 540, and the appropriate surfaces and metrics for display can be stored by the passable world module 538 for future retrieval by the user 530 or other users without having to recreate the model completely or partially. In some embodiments, the data input 536 provides the passable world module 538 with input such as the geographical location, user identification, and current activity indicating which of one or more fixed elements 542 are available, the AR content 540 last placed on the fixed element 542, and whether to display that same content (such AR content being "persistent" content regardless of whether the user is viewing a particular passable world model).
[0167] Even in embodiments where an object is considered fixed (e.g., a kitchen table), the passable world module 538 may update those objects in the physical world model as needed to account for possible changes in the physical world. The models of fixed objects may be updated very infrequently. Other objects in the physical world may be considered to be moving or otherwise not fixed (e.g., a kitchen chair). To render the AR scene in a realistic sense, the AR system may update the positions of these non-fixed objects at a much higher frequency than that used to update the fixed objects. To enable accurate tracking of all objects in the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.
[0168] FIG. 5B is a schematic illustration of the viewing optics assembly 548 and associated components. In some embodiments, two eye-tracking cameras 550 are directed towards the user's eyes 549 to detect metrics of the user's eyes 549, such as eye shape, eyelid occlusion, pupil direction, and glints on the user's eyes 549.
[0169] In some embodiments, one of the sensors is a depth sensor 551, such as a time-of-flight sensor, that emits signals into the world and detects the reflections of those signals from neighboring objects to determine the distance to a given object. The depth sensor may, for example, quickly determine whether an object has entered the user's field of view as a result of either the movement of those objects or a change in the user's pose. However, information about the positions of objects within the user's field of view may alternatively or additionally be collected using other sensors. Depth information may be obtained, for example, from a stereoscopic image sensor or a plenoptic sensor.
[0170] In some embodiments, the world camera 552 records, maps, and / or otherwise creates a model of the environment 532 that can detect inputs that can affect AR content and has a wider view than the surroundings. In some embodiments, the world camera 552 and / or the camera 553 may be grayscale and / or color image sensors, which may output grayscale and / or color image frames at a fixed time interval. The camera 553 may further capture a physical world image within the user's field of view at a specific time. The pixels of the frame-based image sensor may be sampled iteratively even if their values are invariant. The world camera 552, the camera 553, and the depth sensor 551 each have individual fields of view 554, 555, and 556, respectively, and collect and record data from a physical world scene such as the physical world environment 532 depicted in FIG. 34A.
[0171] The inertial measurement unit 557 may determine the movement and orientation of the visual optics assembly 548. In some embodiments, the inertial measurement unit 557 may provide an output indicating the direction of gravity. In some embodiments, each component is operatively coupled to at least one other component. For example, the depth sensor 551 is operatively coupled to the eye tracking camera 550 as a confirmation of the measured focus adjustment for the actual distance that the user's eye 549 is looking at.
[0172] The visual optical system assembly 548 may include some of the components illustrated in FIG. 34B, and it should be understood that it may include components instead of, or in addition to, the illustrated components. In some embodiments, for example, the visual optical system assembly 548 may include two world cameras 552 instead of four. Alternatively, or in addition, cameras 552 and 553 need not capture visible light images of their full fields of view. The visual optical system assembly 548 may include other types of components. In some embodiments, the visual optical system assembly 548 may include one or more dynamic vision sensors (DVSs), the pixels of which may respond asynchronously to relative changes in light intensity exceeding a threshold.
[0173] In some embodiments, the visual optical system assembly 548 may not include a depth sensor 551 based on time-of-flight information. In some embodiments, for example, the visual optical system assembly 548 may include one or more plenoptic cameras, the pixels of which may capture the light intensity and angle of incident light, from which depth information may be determined. For example, a plenoptic camera may include an image sensor overlaid with a transmissive diffractive mask (TDM). Alternatively, or in addition, a plenoptic camera may include an image sensor containing angle-sensing pixels and / or phase-detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors may serve as a depth information source instead of, or in addition to, the depth sensor 551.
[0174] Also, it should be understood that the component configuration in FIG. 5B is provided as an example. The visual optics assembly 548 may include components with any suitable configuration, which may be set to provide the user with a practical maximum field of view for a particular set of components. For example, if the visual optics assembly 548 has one world camera 552, the world camera may be installed within the central region of the visual optics assembly instead of on the side.
[0175] Information from sensors within the visual optics assembly 548 may be coupled to one or more than one of the processors within the system. The processor may generate data that can be rendered to make the user perceive the virtual content as interacting with objects in the physical world. That rendering may be implemented in any suitable way, including generating image data that depicts both physical and virtual objects. In other embodiments, the physical and virtual content may be depicted in one scene by modulating the opacity of the display device such that the user can see through the physical world. The opacity may be controlled to create the appearance of the virtual object and block the view of objects in the physical world that are occluded by the virtual object from the user. In some embodiments, the image data may be modified to be perceived by the user as interacting realistically with the physical world when the virtual content is viewed through the user interface (e.g., clipping the content and taking occlusion into account), and may include only the virtual content.
[0176] The location on the visual optics assembly 548 where content can be displayed to create an impression of an object at a particular location may depend on the physics of the visual optics assembly. Additionally, the posture of the user's head and the direction in which the user's eyes are looking with respect to the physical world will affect the location within the physical world content that will be displayed at a particular location on the visual optics assembly where the content will appear. Sensors as described above may supply information such that a processor receiving the sensor input can collect this information and / or calculate information therefrom so that it can calculate the location on the visual optics assembly 548 where an object should be rendered to create the desired appearance for the user.
[0177] Regardless of how content is presented to the user, a model of the physical world can be used such that the characteristics of virtual objects, including the shape, position, motion, and visibility of the virtual objects, can be correctly calculated, which can be affected by physical objects. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 518.
[0178] The model may be created from data collected from sensors on the user's wearable device. However, in some embodiments, the model may be created from data collected by multiple users, which may be aggregated within a computing device remote from all users (and may be "in the cloud").
[0179] The model may be created, at least in part, by a world reconstruction system, such as world reconstruction component 516 of FIG. 3, which is depicted in more detail in FIG. 6A. The world reconstruction component 516 may include a perception module 660 that can generate, update, and store a representation for a portion of the physical world. In some embodiments, the perception module 660 may represent a portion of the physical world within the reconstruction range of the sensors as a plurality of voxels. Each voxel corresponds to a 3D cube of a predetermined volume within the physical world, includes surface information, and may indicate whether a surface exists within the volume represented by the voxel. The voxels may be assigned a value indicating whether the corresponding volume has been determined to contain the surface of a physical object, to be empty, or not yet measured using the sensors and thus its value is unknown. It should be understood that the values indicating voxels determined to be empty or unknown need not be explicitly stored, and the voxel values may be stored in computer memory in any suitable manner, including not storing information regarding voxels determined to be empty or unknown.
[0180] In addition to generating information for the persistent world representation, the perception module 660 may identify and output an indication of a change in the area surrounding the user of the AR system. Such an indication of a change may trigger other functions, such as triggering an update to the volumetric data stored as part of the persistent world, or triggering component 604 to generate and update AR content.
[0181] In some embodiments, the perception module 660 may identify changes based on a signed distance function (SDF) model. The perception module 660 may be configured to receive sensor data such as, for example, a depth map 660a and a head pose 660b, and then fuse the sensor data into an SDF model 660c. The depth map 660a may directly provide SDF information, and the image may be processed to arrive at the SDF information. The SDF information represents the distance from the sensor used to capture that information. Since those sensors can be part of the wearable unit, the SDF information may represent the physical world from the perspective of the wearable unit and thus the user's perspective. The head pose 660b may enable the SDF information to be associated with voxels within the physical world.
[0182] In some embodiments, the perception module 660 may generate, update, and store a representation for a portion of the physical world that is within the perception range. The perception range may be determined based at least in part on the reconstruction range of the sensor, which may be determined based at least in part on the limits of the observation range of the sensor. As a specific example, an active depth sensor that operates using active IR pulses can reliably operate over a certain range of distances, creating an observation range for the sensor that can be several centimeters or from tens of centimeters to several meters.
[0183] The world reconstruction component 516 may include additional modules that can interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data obtained by the perception module 660. The persistent world module 662 may also include representations of the physical world in various formats. For example, volumetric metadata 662b such as voxels may be stored along with a mesh 662c and a plane 662d. In some embodiments, other information such as a depth map may also be stored.
[0184] In some embodiments, a representation of the physical world, such as that illustrated in FIG. 6A, may provide relatively dense information about the physical world as compared to a sparse map, such as a tracking map based on feature points, as described above.
[0185] In some embodiments, the perception module 660 may include modules that generate representations for the physical world in various formats, such as, for example, a mesh 660d, a plane, and a semantics 660e. Representations for the physical world may be stored across local and remote memory media. Representations for the physical world may be described in different coordinate frames, for example, depending on the location of the memory media. For example, a representation for the physical world stored within a device may be described in a coordinate frame local to the device. Representations for the physical world may have counterparts stored within the cloud. The counterparts within the cloud may be described in a coordinate frame shared by all devices within the XR system.
[0186] In some embodiments, these modules may generate a representation based on data within the perception range of one or more sensors at the time the representation is generated, data captured at previous times, and information within the persistent world module 662. In some embodiments, these components may act on depth information captured using a depth sensor. However, the AR system may include visual sensors and may generate such a representation by analyzing monocular or binocular visual information.
[0187] In some embodiments, these modules may act on regions of the physical world. Those modules may be triggered to update a sub-region of the physical world when the perception module 660 detects a change in the physical world within that sub-region. Such changes may be detected, for example, by detecting a new surface within the SDF model 660c or by other criteria such as a change in the values of a sufficient number of voxels representing the sub-region.
[0188] The world reconstruction component 516 may include a component 664 that can receive a representation of the physical world from the perception module 660. Information about the physical world may be pulled by these components, for example, according to usage requests from an application. In some embodiments, the information may be pushed to the using component via, for example, an indication of a change in a pre-identified area or a change in the representation of the physical world within the perception range. The component 664 may include, for example, game programs and other components that perform processing for visual occlusion, physics-based interactions, and environmental inference.
[0189] In response to a query from the component 664, the perception module 660 may send a representation for the physical world in one or more formats. For example, when the component 664 indicates that the usage is for visual occlusion or physics-based interaction, the perception module 660 may send a representation of the surface. When the component 664 indicates that the usage is for environmental inference, the perception module 660 may send a mesh, plane, and semantics of the physical world.
[0190] In some embodiments, the perception module 660 may include a component that provides format information to the component 664. An example of such a component may be the raycasting component 660f. The using component (e.g., the component 664) may query for information about the physical world from a specific perspective, for example. The raycasting component 660f may select from one or more representations of the physical world data within the field of view from that perspective.
[0191] As should be understood from the foregoing description, the perception module 660 or another component of the AR system may process data and create a 3D representation of a portion of the physical world. The data to be processed may, at least in part, be based on the camera frustum and / or depth images to decimate a portion of the 3D reconstruction volume, extract and persist planar data, capture, persist, and update 3D reconstruction data in blocks that allow local updates while maintaining neighborhood coherence, where occlusion data is derived from one or more combinations of depth data sources, provide such occlusion data to an application that generates such a scene, and / or may be reduced by performing multi-level mesh simplification. The reconstruction may contain data of different levels of sophistication, including, for example, raw data such as live depth data, fused volumetric data such as voxels, and computed data such as meshes.
[0192] In some embodiments, components of the passable world model may be distributed, with some parts being executed locally on the XR device and some parts being executed remotely, such as on a network connected to a server or otherwise in the cloud. The distribution of information processing and storage between the local XR device and the cloud can affect the functionality and user experience of the XR system. For example, reducing the processing on the local device by distributing the processing to the cloud can enable a longer battery life and reduce the heat generated on the local device. However, distributing much more processing to the cloud can create unacceptable wait times that cause an unacceptable user experience.
[0193] Figure 6B depicts a distributed component architecture 600 configured for spatial computing, according to some embodiments. The distributed component architecture 600 may include a passable world component 602 (e.g., PW538 in FIG. 5A), Lumin OS 604, an API 606, an SDK 608, and an application 610. Lumin OS 604 may include a Linux-based kernel with a custom driver compatible with XR devices. The API 606 may include an application programming interface that provides XR applications (e.g., application 610) access to the spatial computing features of XR devices. The SDK 608 may include a software development kit that enables the creation of XR applications.
[0194] One or more components within architecture 600 may create and maintain a model of the passable world. In this example, sensor data is collected on a local device. The processing of that sensor data may be performed locally on the XR device, in part, and in the cloud, in part. PW538 may include an environmental map created, at least in part, based on data captured by an AR device worn by a plurality of users. During a session of an AR experience, an individual AR device (such as the wearable device described above in connection with FIG. 4) may create a tracking map, which is one type of map.
[0195] In some embodiments, the device may include components that construct both a sparse map and a dense map. The tracking map may serve as a sparse map and may include information about the head pose of the AR device scanning the environment and the objects detected in that environment at each head pose. Those head poses may be maintained locally for each device. For example, the head pose on each device may be relative to the initial head pose when the device was turned on for that session. As a result, each tracking map may be local to the device that created it and may have its own reference frame defined by its own local coordinate system. However, in some embodiments, the tracking map on each device may be formed such that one coordinate of its local coordinate system is aligned with the direction of gravity as measured by its sensors such as the inertial measurement unit 557.
[0196] In some embodiments, the device may include components that construct both a sparse map and a dense map. The tracking map may serve as a sparse map and may include information about the head pose of the AR device scanning the environment as well as information about the objects detected in that environment at each head pose. Those head poses may be maintained locally for each device. For example, the head pose on each device may be relative to the initial head pose when the device was turned on for that session. As a result, each tracking map may be local to the device that created it. The dense map may include surface information, which may be represented by a mesh or depth information. Alternatively, or in addition, the dense map may include a higher level of information derived from surface or depth information such as the location and / or properties of planes and / or other objects.
[0197] In some embodiments, the creation of the dense map may be independent of the creation of the sparse map. The creation of the dense and sparse maps may be performed, for example, within separate processing pipelines in the AR system. Separating the processing may, for example, enable different types of map generation or processing to be performed at different rates. The sparse map may be refreshed, for example, at a faster rate than the dense map. However, in some embodiments, the processing of the dense and sparse maps may be related even when performed within different pipelines. Changes in the physical world exposed in the sparse map may, for example, trigger an update of the dense map, or vice versa. Further, even when created independently, the maps may be used together. For example, the coordinate system derived from the sparse map may be used to define the position and / or orientation of objects within the dense map.
[0198] The sparse map and / or the dense map may persist for reuse by the same device and / or for sharing with other devices. Such persistence may be achieved by storing the information in the cloud. The AR device may send the tracking map to the cloud and, for example, merge it with an environmental map selected from previously stored persistent maps in the cloud. In some embodiments, the selected persistent map may be sent from the cloud to the AR device for merging. In some embodiments, the persistent maps may be oriented with respect to one or more persistent coordinate frames. Such maps may serve as reference maps since they can be used by any of a plurality of devices. In some embodiments, a model of the traversable world may include or be created with one or more reference maps. The device may use the reference map by determining the transformation between its local coordinate frame, in which the device performs some operations, and the reference map.
[0199] The reference map may arise as a tracking map (TM) (e.g., TM1102 in FIG. 31A), which may be promoted to the reference map. The reference map may be persisted such that a device accessing the reference map can use the information in the reference map to determine the location of objects represented in the reference map within the physical world around the device once the transformation between its local coordinate system and the coordinate system of the reference map is determined. In some embodiments, the TM may be a head pose sparse map created by an XR device. In some embodiments, the reference map may be created when an XR device sends one or more TMs to a cloud server for merging with additional TMs captured at different times by the XR device or by other XR devices.
[0200] In embodiments where the tracking map is formed on a local device with one coordinate of a local coordinate frame aligned with gravity, the orientation with respect to gravity may be saved in response to the creation of the reference map. For example, when the tracking map submitted for merging does not overlap with any previously stored map, that tracking map may be promoted to the reference map. Other tracking maps, which may similarly have an orientation with respect to gravity, may subsequently be merged with that reference map. The merging may be done to ensure that the resulting reference map retains its orientation with respect to gravity. The two maps may not be merged, for example, if the coordinates of each map aligned with gravity do not match each other with a sufficient approximation tolerance, regardless of the correspondence of feature points in those maps.
[0201] A reference map or other map may provide information about a part of the physical world represented by data processed to create an individual map. FIG. 7 depicts an exemplary tracking map 700 according to some embodiments. The tracking map 700 may provide a plan view 706 of a physical object in the corresponding physical world represented by points 702. In some embodiments, the map points 702 may represent features of a physical object that may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. The features may be derived from a processed image such as may be obtained using sensors of a wearable device within an augmented reality system. The features may be derived, for example, by processing an image frame output by a sensor and identifying features based on large gradients or other suitable criteria within the image. Further processing may limit the number of features within each frame. For example, the processing may select features that are likely to represent persistent objects. One or more heuristics may be applied for this selection.
[0202] The tracking map 700 may include data regarding points 702 collected by a device. For each image frame with data points included within the tracking map, a pose may be stored. The pose may represent the orientation from which the image frame was captured such that feature points within each image frame can be spatially correlated. The pose may be determined by positioning information such as may be derived from sensors such as an IMU sensor on the wearable device. Alternatively, or in addition, the pose may be determined by matching an image frame with other image frames depicting an overlapping portion of the physical world. By finding such a spatial correlation, which may be accomplished by matching a subset of the feature points within two frames, a relative pose between the two frames may be calculated. The relative pose may be appropriate for the tracking map as the map may be with respect to a coordinate system local to the device established based on the initial pose of the device when construction of the tracking map was initiated.
[0203] Since much of the information collected using sensors is likely to be redundant, not all of the feature points and image frames collected by the device can be retained as part of the tracking map. Rather, only certain frames may be added to the map. Those frames may be selected based on one or more criteria such as the degree of overlap with image frames already in the map, the number of new features they contain, or a quality metric for the features within the frame. Image frames not added to the tracking map may be discarded or used to revise the location of features. As a further alternative, all or most of the image frames, represented as a set of features, may be retained, but a subset of those frames may be designated as key frames, which are used for further processing.
[0204] The key frames may be processed to produce key rigs 704. The key frames may be processed to produce a three-dimensional set of feature points and stored as key rigs 704. Such processing may involve, for example, comparing image frames derived simultaneously from two cameras and stereoscopically determining the 3D position of feature points. Metadata such as pose may be associated with these key frames and / or key rigs.
[0205] The environment map may have any of a plurality of formats, depending on, for example, the storage location of the environment map, including, for example, the local storage device and the remote storage device of the AR device. For example, the map in the remote storage device may have a higher resolution than the map in the local storage device on the wearable device when the memory is limited. To transmit the higher resolution map from the remote storage device to the local storage device, the map may be downsampled or otherwise converted to an appropriate format, such as by reducing the number of poses per area of the physical world stored in the map and / or the number of feature points stored per pose. In some embodiments, a slice or portion of the high resolution map from the remote storage device may be transmitted to the local storage device, and the slice or portion is not downsampled.
[0206] The database of environment maps may be updated as new tracking maps are created. To determine which of potentially a very large number of environment maps in the database should be updated, updating may include efficiently selecting one or more environment maps stored in the database related to the new tracking map. The selected one or more environment maps may be ranked by relevance, and one or more of the highest ranked maps may be selected for processing to merge the higher ranked selected environment maps with the new tracking map to create one or more updated environment maps. When the new tracking map represents a portion of the physical world for which there is no existing environment map to update across, that tracking map may be stored in the database as a new environment map.
[0207] View-independent display
[0208] What is described in this specification is a method and apparatus for providing virtual content using an XR system, independent of the location of the eyes viewing the virtual content. Conventionally, virtual content is re-rendered in response to any movement of the display system. For example, when a user wearing a display system views a virtual representation of a three-dimensional (3D) object on a display and walks around the area where the 3D object appears, the 3D object should be re-rendered for each viewpoint so that it has the perception of the user walking around the object as the user occupies real space. However, re-rendering consumes significant computational resources of the system and causes artifacts due to latency.
[0209] The inventors recognized and appreciated that the head pose (e.g., the location and orientation of a user wearing an XR system) can be used to render virtual content independently of eye rotations in the user's head. In some embodiments, a dynamic map of the scene, independent of eye rotations in the user's head and / or independent of sensor deformations caused by heat generated, for example, during high-computation-intensive operations, may be generated across one or more sessions so that virtual content interacting with the dynamic map can be robustly rendered based on a plurality of coordinate frames in real space. In some embodiments, the configuration of the plurality of coordinate frames may enable a first XR device worn by a first user and a second XR device worn by a second user to recognize a common location within the scene. In some embodiments, the configuration of the plurality of coordinate frames may enable a user wearing an XR device to view virtual content within the same location of the scene.
[0210] In some embodiments, a tracking map may be constructed within a world coordinate frame, which may have a world origin. The world origin may be the first pose of the XR device when the XR device is powered on. The world origin may be aligned with gravity so that the developer of the XR application can obtain gravity alignment without extra work. Different tracking maps may be constructed within different world coordinate frames because the tracking maps can be captured by the same XR device in different sessions and / or different XR devices worn by different users. In some embodiments, a session of the XR device may start when the device is powered on and continue until the device is powered off. In some embodiments, the XR device may have a head coordinate frame, which may have a head origin. The head origin may be the current pose of the XR device when an image is captured. The difference between the world coordinate frame and the head pose of the head coordinate frame may be used to estimate a tracking route.
[0211] In some embodiments, the XR device may have a camera coordinate frame, which may have a camera origin. The camera origin may be the current pose of one or more sensors of the XR device. The inventors recognized and appreciated the true value of the configuration of the camera coordinate frame enabling a robust display of virtual content independently of eye rotation in the user's head. The configuration also enables a robust display of virtual content independently of, for example, sensor deformation due to heat generated during operation.
[0212] In some embodiments, the XR device may have a head unit with a head-mountable frame that the user can attach to their head and that may include two waveguides, one in front of each of the user's eyes. The waveguides may be transparent such that ambient light from real-world objects can pass through the waveguides and the user can see the real-world objects. Each waveguide may transmit light projected from a projector to the user's individual eye. The projected light may form an image on the retina of the eye. The retina of the eye thus receives both ambient light and the projected light. The user may see both real-world objects and one or more virtual objects created by the projected light at the same time. In some embodiments, the XR device may have sensors that detect real-world objects around the user. These sensors may be, for example, cameras that capture images that can be processed to identify the location of real-world objects.
[0213] In some embodiments, the XR system may assign a coordinate frame to the virtual content, as opposed to tying the virtual content to a world coordinate frame. Such a configuration allows the virtual content to be described regardless of the location rendered for the user, but may be tied to a more persistent frame position, such as a persistent coordinate frame (PCF) described in relation to FIGS. 14-20C, and rendered at a defined location. When the location of an object changes, the XR device may detect the change in the environmental map and determine the movement of the head unit worn by the user relative to the real-world object.
[0214] FIG. 8 illustrates a user experiencing virtual content as rendered within a physical environment by an XR system 10 according to some embodiments. The XR system may include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is present within a physical environment with a real object in the form of a table 16.
[0215] In the illustrated embodiment, the first XR device 12.1 includes a head unit 22, a belt pack 24, and a cable connection 26. The first user 14.1 secures the head unit 22 to their head and secures the belt pack 24, which is remote from the head unit 22, on their waist. The cable connection 26 connects the head unit 22 to the belt pack 24. The head unit 22 is used to display virtual objects or a plurality of objects to the first user 14.1 while allowing the first user 14.1 to see real objects such as a table 16. The belt pack 24 mainly includes the processing and communication capabilities of the first XR device 12.1. In some embodiments, the processing and communication capabilities may reside wholly or partially within the head unit 22 such that the belt pack 24 can be removed or located within another device such as a backpack.
[0216] In the illustrated embodiment, the belt pack 24 is connected to the network 18 via a wireless connection. The server 20 is connected to the network 18 and holds data representing local content. The belt pack 24 downloads data representing local content from the server 20 via the network 18. The belt pack 24 provides the data to the head unit 22 via the cable connection 26. The head unit 22 may include a display having a light source, such as a laser light source or a light emitting diode (LED) light source, and a waveguide for guiding light.
[0217] In some embodiments, the first user 14.1 may mount the head unit 22 on their head and the belt pack 24 on their waist. The belt pack 24 may download image data representing virtual content from the server 20 via the network 18. The first user 14.1 may be able to see the table 16 through the display of the head unit 22. The projector, which forms part of the head unit 22, may receive the image data from the belt pack 24 and generate light based on the image data. The light may travel through one or more than one of the waveguides that form part of the display of the head unit 22. The light may then exit the waveguide and propagate onto the retina of the eyes of the first user 14.1. The projector may generate light in a pattern that is replicated on the retina of the eyes of the first user 14.1. The light hitting the retina of the eyes of the first user 14.1 may have a selected depth of field so that the first user 14.1 perceives the image at a preselected depth behind the waveguide. Additionally, both eyes of the first user 14.1 may receive slightly different images so that the brain of the first user 14.1 perceives a three-dimensional image or multiple images at a selected distance from the head unit 22. In the illustrated example, the first user 14.1 perceives the virtual content 28 above the table 16. The virtual content 28 and its location and distance ratio from the first user 14.1 are determined by the data representing the virtual content 28 and the various coordinate frames used to display the virtual content 28 to the first user 14.1.
[0218] In the illustrated embodiment, the virtual content 28 is invisible from the perspective of the drawing and is visible to the first user 14.1 through the use of the first XR device 12.1. The virtual content 28 may initially reside as a data structure within the visual data and an algorithm within the belt pack 24. The data structure may then be exposed as light when the projector of the head unit 22 generates light based on the data structure. The virtual content 28 does not exist within the three-dimensional space in front of the first user 14.1, but it should be understood that the virtual content 28 is still represented in FIG. 1 within the three-dimensional space for the purpose of exemplifying what the wearer of the head unit 22 perceives. The visualization of computer data within the three-dimensional space may be used in this description to illustrate how data structures that facilitate the rendering perceived by one or more users are interrelated among the data structures within the belt pack 24.
[0219] FIG. 9 illustrates components of the first XR device 12.1 according to some embodiments. The first XR device 12.1 may include various components that form part of the visual data and algorithms, such as a head unit 22, for example, a rendering engine 30, various coordinate systems 32, various origin and destination coordinate frames 34, and various origin / destination coordinate frame converters 36. The various coordinate systems may be based on the inherent properties of the XR device or may be determined by referring to other information such as a persistent pose or a persistent coordinate system as described herein.
[0220] The head unit 22 may include a head-mountable frame 40, a display system 42, a real object detection camera 44, a movement tracking camera 46, and an inertial measurement unit 48.
[0221] The head-mountable frame 40 may have a shape that can be fixed to the head of the first user 14.1 in FIG. 8. The display system 42, the real object detection camera 44, the movement tracking camera 46, and the inertial measurement unit 48 are mounted on the head-mountable frame 40 and can thus move together with the head-mountable frame 40.
[0222] The coordinate system 32 may include a local data system 52, a world frame system 54, a head frame system 56, and a camera frame system 58.
[0223] The local data system 52 may include a data channel 62, a local frame determination routine 64, and a local frame storage instruction 66. The data channel 62 may be a hardware component such as an internal software routine, an external cable, or a radio frequency receiver, or a hybrid component such as an open port. The data channel 62 may be configured to receive image data 68 representing virtual content.
[0224] The local frame determination routine 64 may be connected to the data channel 62. The local frame determination routine 64 may be configured to determine a local coordinate frame 70. In some embodiments, the local frame determination routine may determine the local coordinate frame based on a real-world object or a real-world location. In some embodiments, the local coordinate frame may be based on the upper edge relative to the bottom edge of the browser window, the head or feet of a character, a node on the outer surface of a prism or bounding box surrounding the virtual content, or any other suitable location for installing a coordinate frame that defines the facing direction of the virtual content and the location where the virtual content should be installed (e.g., a node such as an installation node or a PCF node).
[0225] The local frame storage instruction 66 may be connected to the local frame determination routine 64. Those skilled in the art will understand that software modules and routines may be "connected" to each other through subroutines, calls, etc. The local frame storage instruction 66 may store the local coordinate frame 70 as the local coordinate frame 72 within the origin and destination coordinate frames 34. In some embodiments, the origin and destination coordinate frames 34 may be one or more coordinate frames that can be manipulated or transformed in order for virtual content to persist across sessions. In some embodiments, a session may be the time period between the boot-up and shutdown of an XR device. Two sessions may be two startup and shutdown cycles for a single XR device, or the startup and shutdown for two different XR devices.
[0226] In some embodiments, the origin and destination coordinate frames 34 may be coordinate frames involved in one or more transformations required for the XR devices of a first user and a second user to recognize a common location. In some embodiments, the destination coordinate frame may be the output of a series of calculations and transformations applied to the target coordinate frame in order for the first and second users to visually perceive virtual content at the same location.
[0227] The rendering engine 30 may be connected to the data channel 62. The rendering engine 30 may receive the image data 68 from the data channel 62 such that the rendering engine 30 can render virtual content, at least in part, based on the image data 68.
[0228] The display system 42 may be connected to the rendering engine 30. The display system 42 may include components that convert the image data 68 into visible light. The visible light may form one or two patterns per eye. The visible light may be incident on the eye of the first user 14.1 in FIG. 8 and may be detected on the retina of the eye of the first user 14.1.
[0229] The real object detection camera 44 may include one or more cameras that can capture images from different sides of the head-mounted frame 40. The movement tracking camera 46 may include one or more cameras that capture images on the side surface of the head-mounted frame 40. One set of one or more cameras may be used instead of two sets of one or more cameras representing the real object detection camera 44 and the movement tracking camera 46. In some embodiments, the cameras 44, 46 may capture images. As described above, these cameras may collect data that is used to construct a tracking map.
[0230] The inertial measurement unit 48 may include several devices used to detect the movement of the head unit 22. The inertial measurement unit 48 may include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of the inertial measurement unit 48, in combination, track the movement of the head unit 22 in at least three orthogonal directions and about at least three orthogonal axes.
[0231] In the illustrated embodiment, the world frame system 54 includes a world surface determination routine 78, a world frame determination routine 80, and a world frame memory command 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 receives an image and / or key frames based on the image captured by the real object detection camera 44, processes the image, and identifies the surface within the image. A depth sensor (not shown) may determine the distance to the surface. The surface is thus represented by data in three dimensions, including its size, shape, and distance from the real object detection camera.
[0232] In some embodiments, the world coordinate frame 84 may be based on the origin at the initialization of the head pose session. In some embodiments, the world coordinate frame may be located at the location where the device was booted up, or may be a new location if the head pose was lost during the boot session. In some embodiments, the world coordinate frame may be the origin at the start of the head pose session.
[0233] In the illustrated embodiment, the world frame determination routine 80 is connected to the world surface determination routine 78 and determines the world coordinate frame 84 based on the location of the surface as determined by the world surface determination routine 78. The world frame memory command 82 is connected to the world frame determination routine 80 and receives the world coordinate frame 84 from the world frame determination routine 80. The world frame memory command 82 stores the world coordinate frame 84 as the world coordinate frame 86 within the origin and destination coordinate frame 34.
[0234] The head frame system 56 may include a head frame determination routine 90 and a head frame memory command 92. The head frame determination routine 90 may be connected to the motion tracking camera 46 and the inertial measurement unit 48. The head frame determination routine 90 may calculate a head coordinate frame 94 using data from the motion tracking camera 46 and the inertial measurement unit 48. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The motion tracking camera 46 may continuously capture images that are used by the head frame determination routine 90 to refine the head coordinate frame 94. The head unit 22 moves as the first user 14.1 in FIG. 8 moves their head. The motion tracking camera 46 and the inertial measurement unit 48 may continuously provide data to the head frame determination routine 90 so that the head frame determination routine 90 can update the head coordinate frame 94.
[0235] The head frame memory command 92 may be connected to the head frame determination routine 90 and may receive the head coordinate frame 94 from the head frame determination routine 90. The head frame memory command 92 may store the head coordinate frame 94 as a head coordinate frame 96 between the origin and the destination coordinate frame 34. The head frame memory command 92 may repeatedly store the updated head coordinate frame 94 as the head coordinate frame 96 when the head frame determination routine 90 recalculates the head coordinate frame 94. In some embodiments, the head coordinate frame may be the location of the wearable XR device 12.1 relative to the local coordinate frame 72.
[0236] The camera frame system 58 may include camera intrinsics 98. The camera intrinsics 98 may include the dimensions of the head unit 22, which are characteristics of its design and manufacture. The camera intrinsics 98 may be used to calculate a camera coordinate frame 100 that is stored within the origin and destination coordinate frame 34.
[0237] In some embodiments, the camera coordinate frame 100 may include all pupil positions of the left eye of the first user 14.1 in FIG. 8. When the left eye moves from left to right or vertically or horizontally, the pupil position of the left eye is located within the camera coordinate frame 100. Additionally, the pupil position of the right eye is located within the camera coordinate frame 100 for the right eye. In some embodiments, the camera coordinate frame 100 may include the location of the camera with respect to the local coordinate frame when an image is captured.
[0238] The origin / destination coordinate frame converter 36 may include a local / world coordinate converter 104, a world / head coordinate converter 106, and a head / camera coordinate converter 108. The local / world coordinate converter 104 may receive the local coordinate frame 72 and convert the local coordinate frame 72 into the world coordinate frame 86. The conversion of the local coordinate frame 72 to the world coordinate frame 86 may be represented as the local coordinate frame that is converted into the world coordinate frame 110 within the world coordinate frame 86.
[0239] The world / head coordinate converter 106 may convert from the world coordinate frame 86 to the head coordinate frame 96. The world / head coordinate converter 106 may convert the local coordinate frame that is converted into the world coordinate frame 110 into the head coordinate frame 96. The conversion may be represented as the local coordinate frame that is converted into the head coordinate frame 112 within the head coordinate frame 96.
[0240] The head / camera coordinate converter 108 may convert from the head coordinate frame 96 to the camera coordinate frame 100. The head / camera coordinate converter 108 may convert the local coordinate frame that is converted into the head coordinate frame 112 into the local coordinate frame that is converted into the camera coordinate frame 114 within the camera coordinate frame 100. The local coordinate frame that is converted into the camera coordinate frame 114 may be incorporated into the rendering engine 30. The rendering engine 30 may render the image data 68 representing the local content 28 based on the local coordinate frame that is converted into the camera coordinate frame 114.
[0241] FIG. 10 is a spatial representation of various origin and destination coordinate frames 34. Local coordinate frame 72, world coordinate frame 86, head coordinate frame 96, and camera coordinate frame 100 are represented in the figure. In some embodiments, the local coordinate frame associated with the XR content 28 may have a position and rotation relative to the local and / or world coordinate frame and / or PCF when the virtual content is installed in the real world and thus viewable by the user (e.g., may provide a node and a facing direction). Each camera may have its own camera coordinate frame 100 that encompasses all pupil positions of one eye. Reference numerals 104A and 106A represent the conversions performed by the local / world coordinate converter 104, world / head coordinate converter 106, and head / camera coordinate converter 108 in FIG. 9, respectively.
[0242] FIG. 11 depicts a camera rendering protocol for converting from a head coordinate frame to a camera coordinate frame according to some embodiments. In the illustrated example, the pupil for one eye moves from position A to position B. Virtual objects that are intended to appear stationary will be projected onto the depth plane at one of the two positions A or B depending on the position of the pupil (assuming the camera is configured to use a pupil-based coordinate frame). As a result, using a pupil coordinate frame that is converted to the head coordinate frame will introduce jitter into the stationary virtual object as the eye moves from position A to position B. This situation is referred to as view-dependent display or projection.
[0243] As depicted in FIG. 12, a camera coordinate frame (e.g., CR) is positioned to encompass all pupil positions, and the object projection will be consistent here regardless of pupil positions A and B. The head coordinate frame is transformed into the CR frame, which is referred to as a view-independent display or projection. Image reprojection may be applied to virtual content to account for changes in eye position. However, since the rendering remains at the same position, jitter is minimized.
[0244] FIG. 13 illustrates the display system 42 in more detail. The display system 42 includes a stereoscopic analyzer 144 that is connected to the rendering engine 30 and forms part of the visual data and algorithms.
[0245] The display system 42 further includes left and right projectors 166A and 166B, and left and right light pipes 170A and 170B. The left and right projectors 166A and 166B are connected to a power source. Each projector 166A and 166B has an individual input for the image data to be provided to the individual projector 166A or 166B. When powered, the individual projector 166A or 166B generates light in a two-dimensional pattern and emits the light therefrom. The left and right light pipes 170A and 170B are positioned to receive light from the left and right projectors 166A and 166B, respectively. The left and right light pipes 170A and 170B are transparent light pipes.
[0246] In use, the user mounts the head-mounted frame 40 on their head. The components of the head-mounted frame 40 may include, for example, a strap (not shown) that wraps around the perimeter of the back of the user's head. The left and right light pipes 170A and 170B are then positioned in front of the user's left and right eyes 220A and 220B.
[0247] The rendering engine 30 takes in the image data it receives into the stereoscopic analyzer 144. The image data is the three-dimensional image data of the local content 28 in FIG. 8. The image data is projected onto a plurality of virtual planes. The stereoscopic analyzer 144 analyzes the image data and determines left and right image data sets based on the image data for projection onto each depth plane. The left and right image data sets are data sets that represent two-dimensional images projected in three dimensions to give the user a perception of depth.
[0248] The stereoscopic analyzer 144 takes in the left and right image data sets into the left and right projectors 166A and 166B. The left and right projectors 166A and 166B then create left and right light patterns. Although the components of the display system 42 are shown in a plan view, it should be understood that the left and right patterns are two-dimensional patterns when shown in a front elevation view. Each light pattern includes a plurality of pixels. For illustrative purposes, light rays 224A and 226A from two of the pixels are shown exiting the left projector 166A and entering the left light guide 170A. The light rays 224A and 226A are reflected from the side of the left light guide 170A. Although the light rays 224A and 226A are shown to propagate through internal reflection from left to right within the left light guide 170A, it should be understood that the light rays 224A and 226A also propagate in a direction towards the plane of the paper using refractive and reflective systems.
[0249] The light rays 224A and 226A exit the left light guide 170A through the pupil 228A and then enter the left eye 220A through the pupil 230A of the left eye 220A. The light rays 224A and 226A then strike the retina 232A of the left eye 220A. Thus, the left light pattern strikes the retina 232A of the left eye 220A. The user is given the perception that the pixels formed on the retina 232A are pixels 234A and 236A that are at a certain distance on the side of the left light guide 170A facing the user's left eye 220A. Depth perception is created by manipulating the focal length of the light.
[0250] Similarly, the stereoscopic analyzer 144 captures the right image data set into the right projector 166B. The right projector 166B transmits a right light pattern, which is represented by pixels in the form of light rays 224B and 226B. The light rays 224B and 226B are reflected within the right waveguide 170B and exit through the pupil 228B. The light rays 224B and 226B then enter through the pupil 230B of the right eye 220B and strike the retina 232B of the right eye 220B. The pixels of the light rays 224B and 226B are perceived as pixels 134B and 236B behind the right waveguide 170B.
[0251] The patterns created on the retinas 232A and 232B are perceived individually as left and right images. The left and right images are slightly different from each other due to the function of the stereoscopic analyzer 144. The left and right images are perceived as a 3D rendering within the user's brain.
[0252] As described, the left and right waveguides 170A and 170B are transparent. Light from real objects such as the table 16 on the sides of the left and right waveguides 170A and 170B facing the eyes 220A and 220B can be projected through the left and right waveguides 170A and 170B and strike the retinas 232A and 232B.
[0253] Persistent Coordinate Frame (PCF)
[0254] What is described herein are methods and apparatuses for providing spatial persistence across user instances within a shared space. Without spatial persistence, virtual content placed by a user within the physical world within a session may not be present within the view of the user in a different session or may be mis-placed. Without spatial persistence, virtual content placed by one user within the physical world may not be present within the view of a second user, or may be displaced, even when the second user intends to share the experience of the same physical space as the first user.
[0255] The inventors recognize and understand that spatial persistence can be provided through a Persistent Coordinate Frame (PCF). The PCF may be defined based on one or more points representing features recognized within the physical world (e.g., corners, edges). The features may be selected such that they are likely to be the same from one user instance of the XR system to another.
[0256] Furthermore, drift during tracking, which can cause the calculated tracking path (e.g., camera orbit) to deviate from the actual tracking path, can cause the location of virtual content to appear shifted when rendered relative to a local map based only on the tracking map. The tracking map for the space may be refined as the XR device collects further information about the scene over time to correct for drift. However, if virtual content is placed on a physical object and stored relative to the device's world coordinate frame derived from the tracking map before map refinement, the virtual content may appear displaced as if the physical object has moved during map refinement. The PCF may be updated according to map refinement since the PCF is defined based on features and updated as the features move during map refinement.
[0257] The PCF may have six degrees of freedom involving translation and rotation relative to the map coordinate system. The PCF may be stored within local and / or remote storage media. The translation and rotation of the PCF may be calculated relative to the map coordinate system, for example, depending on the storage location. For example, the PCF used locally by a device may have translation and rotation relative to the device's world coordinate frame. The PCF within the cloud may have translation and rotation relative to the canonical coordinate frame of the canonical map.
[0258] The PCF may provide a sparse representation of the physical world such that they can be efficiently processed and transferred, and may provide less than all of the available information about the physical world. Techniques for processing persistent spatial information may create a dynamic map based on one or more coordinate systems within the physical space, across one or more sessions, and generate a persistent coordinate frame (PCF) across the sparse map that can be exposed to XR applications, for example, via an application programming interface (API).
[0259] FIG. 14 is a block diagram illustrating the creation of a persistent coordinate frame (PCF) and the binding of XR content to the PCF, according to some embodiments. Each block may represent digital information stored within a computer memory. In the case of application 1180, the data may represent computer-executable instructions. In the case of virtual content 1170, the digital information may define virtual objects, for example, as defined by application 1180. In the case of the other boxes, the digital information may characterize some aspects of the physical world.
[0260] In the illustrated embodiments, one or more PCFs are created from images captured using sensors on a wearable device. In the embodiment of FIG. 14, the sensors are visual image cameras. These cameras may be the same cameras used to form a tracking map. Thus, some of the processing proposed by FIG. 14 may be implemented as part of updating the tracking map. However, FIG. 14 illustrates that information providing persistence is generated in addition to the tracking map.
[0261] To derive a 3D PCF, two images 1110 from two cameras mounted on a wearable device in a configuration enabling stereoscopic image analysis are both processed. FIG. 14 illustrates Image 1 and Image 2, respectively derived from one of the cameras. A single image from each camera is shown for convenience. However, each camera may output a stream of image frames, and the processing illustrated in FIG. 14 may be performed for multiple image frames within the stream.
[0262] Thus, Image 1 and Image 2 may each be one frame within a sequence of image frames. The processing as depicted in FIG. 14 may be repeated on consecutive image frames in the sequence until an image frame containing feature points that form suitable images from which persistent spatial information is formed is processed. Alternatively, or in addition, the processing of FIG. 14 may be repeated as the user moves such that the user is no longer close enough to the previously identified PCF and cannot reliably use that PCF to determine a position relative to the physical world. For example, an XR system may maintain the current PCF for the user. When the distance exceeds a threshold, the system may switch to a new current PCF closer to the user that may be generated according to the process of FIG. 14 using the image frames obtained at the user's current location.
[0263] Even when generating a single PCF, a stream of image frames may be processed to identify image frames depicting content within the physical world that are likely to be stable and easily distinguishable by the device in the vicinity of the region of the physical world depicted in the image frame. In the embodiment of FIG. 14, this processing begins with the identification of features 1120 within the image. Features may be identified, for example, by finding locations of gradients or other characteristics within the image that exceed a threshold, which may correspond to corners of objects. In the illustrated embodiment, the features are points, but other recognizable features such as edges may also be used, alternatively or in addition.
[0264] In the illustrated embodiment, a fixed number N of feature points 1120 are selected for further processing. Those feature points may be selected based on one or more criteria such as the magnitude of the gradient or proximity to other feature points. Alternatively, or in addition, the feature points may be selected heuristically, such as based on characteristics that suggest the feature points are persistent. For example, the heuristic may be defined based on characteristics of feature points that are likely to correspond to corners of windows or doors or large furniture. Such a heuristic may consider the feature points themselves and what surrounds them. As a specific example, the number of feature points per image may be 100 - 500 or 150 - 250, such as 200.
[0265] Regardless of the number of feature points selected, descriptors 1130 may be calculated for the feature points. In this example, the descriptors are calculated for each selected feature point, but the descriptors may be calculated for groups of feature points, or subsets of feature points, or all features in the image. The descriptors characterize the feature points such that feature points representing the same object in the physical world are assigned similar descriptors. The descriptors may facilitate the alignment of two frames such as may occur when one map is located relative to another map. Instead of searching for the relative orientation of the two frames that minimizes the distance between feature points of two images, the initial alignment of the two frames may be done by identifying feature points with similar descriptors. The alignment of the image frames may be based on aligning points with similar descriptors, which may involve less processing than calculating the alignment of all feature points in the image.
[0266] The descriptors may be calculated as a mapping of the descriptors to the feature points, or in some embodiments, a mapping of image patches around the feature points. The descriptors may be a numerical quantity. U.S. Patent Application No. 16 / 190,948 describes calculating descriptors for feature points and is incorporated herein by reference in its entirety.
[0267] In the embodiment of FIG. 14, the descriptor 1130 is calculated for each feature point within each image frame. Based on the descriptor and / or the feature points and / or the image itself, an image frame may be identified as a key frame 1140. In the illustrated embodiment, a key frame is an image frame that meets certain criteria and is then selected for further processing. When creating a tracking map, for example, an image frame that adds meaningful information to the map may be selected as a key frame to be integrated into the map. On the other hand, image frames that substantially overlap an area where image frames have already been integrated into the map may be discarded so that they do not become key frames. Alternatively, or in addition, key frames may be selected based on the number and / or type of feature points within the image frame. In the embodiment of FIG. 14, the key frame 1150 selected for inclusion within the tracking map may also be processed as a key frame for determining the PCF, although different or additional criteria for selecting a key frame for PCF generation may be used.
[0268] FIG. 14 shows that key frames are used for further processing, but the information obtained from the images may be processed in other forms. For example, feature points such as within a key rig may be processed alternatively or in addition. Further, although key frames are described as being derived from a single image frame, there need not be a one-to-one relationship between a key frame and the image frame from which it is obtained. A key frame may be obtained from multiple image frames, for example, by stitching or aggregating the image frames together such that only features that appear in multiple images are retained within the key frame.
[0269] A keyframe may include image information and / or metadata associated with the image information. In some embodiments, an image captured by cameras 44, 46 (FIG. 9) may be calculated into one or more keyframes (e.g., keyframes 1, 2). In some embodiments, a keyframe may include a camera pose. In some embodiments, a keyframe may include one or more camera images captured in a camera pose. In some embodiments, the XR system determines that a portion of the camera image captured in the camera pose is not useful and thus may not include that portion within the keyframe. Thus, using keyframes to align new images with earlier knowledge of the scene reduces the use of the XR system's computational resources. In some embodiments, a keyframe may include an image and / or image data at a location with a certain direction / angle. In some embodiments, a keyframe may include a location and direction from which one or more map points can be observed. In some embodiments, a keyframe may include a coordinate frame with an ID. U.S. Patent Application No. 15 / 877,359 describes keyframes and is incorporated herein by reference in its entirety.
[0270] Some or all of the keyframes 1140 may be selected for further processing such as the generation of a persistent pose 1150 for the keyframes. The selection may be based on the characteristics of all or a subset of the feature points within the image frame. Those characteristics may be determined from processing descriptors, features, and / or the image frame itself. As a specific example, the selection may be based on a cluster of feature points identified as likely to be associated with a persistent object.
[0271] Each keyframe is associated with the pose of the camera at which the keyframe was obtained. With respect to the keyframes selected for processing the persistent pose, the pose information may be stored together with other metadata about the keyframe such as the WiFi fingerprint and / or GPS coordinates at the time of acquisition and / or the location of acquisition.
[0272] The persistent pose is a source of information that the device can use to orient itself with respect to previously obtained information about the physical world. For example, if the keyframe from which the persistent pose was created is incorporated into a map of the physical world, the device can orient itself with respect to that persistent pose using a sufficient number of feature points within the keyframe that are associated with the persistent pose. The device can align the obtained current image of its surroundings with the persistent pose. This alignment may be based on a matching of the current image with the image 1110, features 1120, and / or descriptors 1130 that gave rise to the persistent pose, or any subset of that image or those features or descriptors. In some embodiments, the current image frame matched to the persistent pose may be another keyframe incorporated into the tracking map of the device.
[0273] The information about the persistent pose may be stored in a format that facilitates sharing among multiple applications that may be executed on the same or different devices. In the example of FIG. 14, some or all of the persistent poses may be reflected as a persistent coordinate frame (PCF) 1160. Like the persistent pose, the PCF may also be associated with a map and may comprise a set of features or other information that the device can use to determine its orientation with respect to the PCF. The PCF may include a transformation that defines its transformation with respect to the origin of the map such that the device can determine its position with respect to any object within the physical world reflected in the map by correlating its position to the PCF.
[0274] Since the PCF provides a mechanism for determining the location of physical objects, an application such as application 1180 may define the position of a virtual object relative to one or more PCFs that serve as an anchor for virtual content 1170. FIG. 14 illustrates, for example, that application 1 associates its virtual content 2 with PCF 1.2. Similarly, application 2 associates its virtual content 3 with PCF 1.2. Application 1 is also shown to associate its virtual content 1 with PCF 4.5, and application 2 is shown to associate its virtual content 4 with PCF 3. In some embodiments, similar to the method based on image 1 and image 2 for PCF 1.2, PCF 3 may be based on image 3 (not shown), and PCF 4.5 may be based on image 4 and image 5 (not shown). When rendering this virtual content, the device may apply one or more transformations and calculate information such as the location of the virtual content relative to the device's display and / or the location of the physical object relative to the desired location of the virtual content. Using the PCF as a reference can simplify such calculations.
[0275] In some embodiments, the persistent pose may be a coordinate location and / or orientation having one or more associated keyframes. In some embodiments, the persistent pose may be automatically created after the user has traveled a certain distance, e.g., 3 meters. In some embodiments, the persistent pose may act as a reference point during localization. In some embodiments, the persistent pose may be stored within the traversable world (e.g., traversable world module 538).
[0276] In some embodiments, the new PCF may be determined based on a predefined distance allowed between adjacent PCFs. In some embodiments, one or more persistent postures may be calculated into the PCF when the user progresses a predetermined distance, for example, 5 meters. In some embodiments, the PCF may be associated with one or more world coordinate frames and / or reference coordinate frames, for example, within a traversable world. In some embodiments, the PCF may be stored in a local and / or remote database, for example, according to security settings.
[0277] FIG. 15 illustrates a method 4700 for establishing and using a persistent coordinate frame, according to some embodiments. The method 4700 may begin with capturing an image centered on a scene (e.g., image 1 and image 2 in FIG. 14) using one or more sensors of an XR device (act 4702). A plurality of cameras may be used, and one camera may generate a plurality of images, for example, in a stream.
[0278] Method 4700 may include extracting a point of interest (e.g., map point 702 in FIG. 7, feature 1120 in FIG. 14) from a captured image (4704), generating a descriptor for the extracted point of interest (e.g., descriptor 1130 in FIG. 14) (act 4706), and generating a keyframe (e.g., keyframe 1140) based on the descriptor (act 4708). In some embodiments, the method may compare the points of interest within a keyframe and form pairs of keyframes that share a predetermined amount of points of interest. The method may use individual pairs of keyframes to reconstruct a portion of the physical world. The mapped portion of the physical world may be stored as 3D features (e.g., key rig 704 in FIG. 7). In some embodiments, selected portions of pairs of keyframes may be used to construct 3D features. In some embodiments, the results of the mapping may be selectively stored. Keyframes not used to construct 3D features may be associated with 3D features through their poses, e.g., representing the distance between keyframes with a covariance matrix between the poses of the keyframes. In some embodiments, pairs of keyframes may be selected to construct 3D features such that the distance between each of the 3D features being constructed is within a predetermined distance that can balance the amount of calculation required and the level of accuracy of the resulting model. Such an approach enables providing a model of the physical world with an amount of data suitable for efficient and accurate calculations using an XR system. In some embodiments, the covariance matrix of two images may include the covariance between the poses (e.g., 6 degrees of freedom) of the two images.
[0279] Method 4700 may include generating a persistent pose based on keyframes (act 4710). In some embodiments, the method may include generating a persistent pose based on 3D features reconstructed from a pair of keyframes. In some embodiments, the persistent pose may be associated with the 3D features. In some embodiments, the persistent pose may include the poses of the keyframes used to construct the 3D features. In some embodiments, the persistent pose may include the average pose of the keyframes used to construct the 3D features. In some embodiments, the persistent pose may be generated such that the distance between neighboring persistent poses is within a predetermined value, for example, in the range of 1 meter to 5 meters, any value therebetween, or any other suitable value. In some embodiments, the distance between neighboring persistent poses may be represented by the covariance matrix of the neighboring persistent poses.
[0280] Method 4700 may include generating a PCF based on the persistent pose (act 4712). In some embodiments, the PCF may be associated with the 3D features. In some embodiments, the PCF may be associated with one or more persistent poses. In some embodiments, the PCF may include the pose of one of the associated persistent poses. In some embodiments, the PCF may include the average pose of the poses of the associated persistent poses. In some embodiments, the PCF may be generated such that the distance between neighboring PCFs is within a predetermined value, for example, in the range of 3 meters to 10 meters, any value therebetween, or any other suitable value. In some embodiments, the distance between neighboring PCFs may be represented by the covariance matrix of the neighboring PCFs. In some embodiments, the PCF may be exposed to an XR application such that the XR application can access the model of the physical world through the PCF without accessing the model itself, for example, via an application programming interface (API).
[0281] Method 4700 may include associating (act 4714) at least one of image data of a virtual object and a PCF for display by an XR device. In some embodiments, the method may include calculating a translation and orientation of the virtual object with respect to the associated PCF. It should be understood that it is not necessary to associate the virtual object with the PCF generated by the device on which the virtual object is installed. For example, the device may read the stored PCF in the canonical map in the cloud and associate the virtual object with the read PCF. It should be understood that the virtual object may move with the associated PCF as the PCF is adjusted over time.
[0282] FIG. 16 illustrates a first XR device 12.1, visual data and algorithms of a second XR device 12.2, and a server 20, according to some embodiments. The components illustrated in FIG. 16 may be operative to perform some or all of the operations associated with generating, updating, and / or using spatial information, such as a persistent pose, a persistent coordinate frame, a tracking map, or a canonical map, as described herein. Although not shown, the first XR device 12.1 may be configured identically to the second XR device 12.2. The server 20 may include a map storage routine 118, a canonical map 120, a map transmitter 122, and a map merge algorithm 124.
[0283] A second XR device 12.2 that may be in the same scene as the first XR device 12.1 may include a persistent coordinate frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that may be used to render virtual objects, and a frame embedding generator 308 (see FIG. 21). In some embodiments, the map download system 126, the PCF identification system 128, the map 2, the location module 130, the reference map incorporator 132, the reference map 133, and the map issuer 136 may be grouped within the passable world unit 1304. The PCF integration unit 1300 may be connected to the passable world unit 1304 and other components of the second XR device 12.2 and may enable reading, generating, using, uploading, and downloading of PCFs.
[0284] Maps with PCFs may enable more persistence within a changing world. In some embodiments, for example, locating a tracking map that includes matching features for an image may include selecting features representing persistent content from a map configured by a PCF, which enables fast matching and / or location. For example, in a world where people move in and out of a scene and objects such as doors move relative to the scene, less memory space and transmission rate are required, enabling the use of individual PCFs and their relationships to each other (e.g., the integrated constellation of PCFs) to map the scene.
[0285] In some embodiments, the PCF integration unit 1300 may include a previously stored PCF 1306 in the data storage on the memory unit of the second XR device 12.2, a PCF tracker 1308, a persistent pose acquirer 1310, a PCF verifier 1312, a PCF generation system 1314, a coordinate frame computer 1316, a persistent pose computer 1318, a tracking map and persistent pose converter 1320, a persistent pose and PCF converter 1322, and a PCF and image data converter 1324, including three converters.
[0286] In some embodiments, the PCF tracker 1308 may have on-prompts and off-prompts that are selectable by the application 1302. The application 1302 is executable by a processor of the second XR device 12.2 and may, for example, display virtual content. The application 1302 may have a call to switch on the PCF tracker 1308 via an on-prompt. The PCF tracker 1308 may generate a PCF when the PCF tracker 1308 is switched on. The application 1302 may have a subsequent call to switch off the PCF tracker 1308 via an off-prompt. The PCF tracker 1308 terminates PCF generation when the PCF tracker 1308 is switched off.
[0287] In some embodiments, the server 20 may include a plurality of persistent postures 1332 and a plurality of PCFs 1330 that are associated with and previously stored with the reference map 120. The map transmitter 122 may transmit the reference map 120 to the second XR device 12.2 together with the persistent postures 1332 and / or the PCFs 1330. The persistent postures 1332 and the PCFs 1330 may be stored on the second XR device 12.2 in association with the reference map 133. When Map 2 is located with respect to the reference map 133, the persistent postures 1332 and the PCFs 1330 may be stored in association with Map 2.
[0288] In some embodiments, the persistent posture acquirer 1310 may acquire a persistent posture for Map 2. The PCF verifier 1312 may be connected to the persistent posture acquirer 1310. The PCF verifier 1312 may read a PCF from the PCF 1306 based on the persistent posture read by the persistent posture acquirer 1310. The PCF read by the PCF verifier 1312 may form an initial group of PCFs for use in image display based on the PCF.
[0289] In some embodiments, application 1302 may require that an additional PCF be generated. For example, when the user moves to an area that has not been previously mapped, application 1302 may switch on the PCF tracker 1308. The PCF generation system 1314 is connected to the PCF tracker 1308 and may start generating the PCF based on map 2 as map 2 begins to expand. The PCF generated by the PCF generation system 1314 may form a second group of PCFs that can be used for PCF-based image display.
[0290] The coordinate frame computer 1316 may be connected to the PCF verifier 1312. After the PCF verifier 1312 reads the PCF, the coordinate frame computer 1316 may call the head coordinate frame 96 and determine the head pose of the second XR device 12.2. The coordinate frame computer 1316 may also call the persistent pose computer 1318. The persistent pose computer 1318 may be directly or indirectly connected to the frame embedding generator 308. In some embodiments, an image frame may be designated as a key frame after a threshold distance, for example, 3 meters, has been traveled from the previous key frame. The persistent pose computer 1318 may generate a persistent pose based on a plurality of, for example, three key frames. In some embodiments, the persistent pose may essentially be an average of the coordinate frames of a plurality of key frames.
[0291] The tracking map and persistent pose converter 1320 may be connected to map 2 and the persistent pose computer 1318. The tracking map and persistent pose converter 1320 may convert map 2 into a persistent pose and determine the persistent pose at the origin with respect to map 2.
[0292] The persistent pose and PCF converter 1322 may be connected to the tracking map and persistent pose converter 1320 and further to the PCF validator 1312 and the PCF generation system 1314. The persistent pose and PCF converter 1322 may convert the persistent pose (for which the tracking map has been converted) into a PCF from the PCF validator 1312 and the PCF generation system 1314 and determine the PCF for the persistent pose.
[0293] The PCF and image data converter 1324 may be connected to the persistent pose and PCF converter 1322 and the data channel 62. The PCF and image data converter 1324 converts the PCF into image data 68. The rendering engine 30 may be connected to the PCF and image data converter 1324 and display the image data 68 for the PCF to the user.
[0294] The PCF integration unit 1300 may store additional PCFs generated using the PCF generation system 1314 within the PCF 1306. The PCF 1306 may be stored for the persistent pose. The map publisher 136 may read the PCF 1306 and the persistent pose associated with the PCF 1306 when the map publisher 136 transmits the map 2 to the server 20 and the map publisher 136 also transmits the PCF and the persistent pose associated with the map 2 to the server 20. When the map storage routine 118 of the server 20 stores the map 2, the map storage routine 118 may also store the persistent pose and the PCF generated by the second vision device 12.2. The map merge algorithm 124 may create the reference map 120 together with the persistent pose and the PCF of the map 2, which are respectively associated with the reference map 120 and stored within the persistent pose 1332 and the PCF 1330.
[0295] The first XR device 12.1 may include a PCF integration unit similar to the PCF integration unit 1300 of the second XR device 12.2. When the map transmitter 122 transmits the reference map 120 to the first XR device 12.1, the map transmitter 122 may transmit the persistent pose 1332 and the PCF 1330 that are associated with the reference map 120 and originate from the second XR device 12.2. The first XR device 12.1 may store the PCF and the persistent pose in a data storage device on the storage device of the first XR device 12.1. The first XR device 12.1 may then utilize the persistent pose and the PCF originating from the second XR device 12.2 for image display with respect to the PCF. Additionally, or alternatively, the first XR device 12.1 may read, generate, utilize, upload, and download the PCF and the persistent pose in a manner similar to the second XR device 12.2 as described above.
[0296] In the illustrated embodiment, the first XR device 12.1 generates a local tracking map (hereinafter referred to as "Map 1"), and the map storage routine 118 receives Map 1 from the first XR device 12.1. The map storage routine 118 then stores Map 1 as the reference map 120 on the storage device of the server 20.
[0297] The second XR device 12.2 includes a map download system 126, an anchor identification system 128, a location identification module 130, a reference map incorporator 132, a local content positioning system 134, and a map publisher 136.
[0298] In use, the map transmitter 122 transmits the reference map 120 to the second XR device 12.2, and the map download system 126 downloads and stores the reference map 120 from the server 20 as the reference map 133.
[0299] The anchor identification system 128 is connected to the world surface determination routine 78. The anchor identification system 128 identifies an anchor based on an object detected by the world surface determination routine 78. The anchor identification system 128 uses the anchor to generate a second map (Map 2). As shown by cycle 138, the anchor identification system 128 continues to identify the anchor and update Map 2. The location of the anchor is recorded as three-dimensional data based on data provided by the world surface determination routine 78. The world surface determination routine 78 receives an image from the real object detection camera 44 and depth data from the depth sensor 135, and determines the location of the surface and its relative distance from the depth sensor 135.
[0300] The localization module 130 is connected to the reference map 133 and Map 2. The localization module 130 repeatedly attempts to localize Map 2 relative to the reference map 133. The reference map incorporator 132 is connected to the reference map 133 and Map 2. When the localization module 130 localizes Map 2 relative to the reference map 133, the reference map incorporator 132 incorporates the reference map 133 into the anchors of Map 2. Map 2 is then updated with the missing data contained within the reference map.
[0301] The local content positioning system 134 is connected to Map 2. The local content positioning system 134 may be a system, for example, by which a user can localize local content at a specific location within the world coordinate frame. The local content itself is then associated with one of the anchors of Map 2. The local / world coordinate converter 104 converts the local coordinate frame to the world coordinate frame based on the settings of the local content positioning system 134. The functions of the rendering engine 30, the display system 42, and the data channel 62 are described with reference to FIG. 2.
[0302] The map issuer 136 uploads map 2 to the server 20. The map storage routine 118 of the server 20 then stores map 2 in the storage medium of the server 20.
[0303] The map merge algorithm 124 merges map 2 and the reference map 120. When more than two maps, for example, three or four maps, related to the same or adjacent areas of the physical world are stored, the map merge algorithm 124 merges all the maps into the reference map 120 and renders a new reference map 120. The map transmitter 122 then transmits the new reference map 120 to any devices 12.1 and 12.2 within the area represented by the new reference map 120. When devices 12.1 and 12.2 locate their individual maps relative to the reference map 120, the reference map 120 becomes an enhanced map.
[0304] FIG. 17 illustrates an example of generating keyframes for a map of a scene according to some embodiments. In the illustrated example, the first keyframe KF1 is generated for a door on the left wall of the room. The second keyframe KF2 is generated for an area within the corner where the floor, left wall, and right wall of the room meet. The third keyframe KF3 is generated for an area of a window on the right wall of the room. The fourth keyframe KF4 is generated for an area at the edge of a rug on the floor of the wall. The fifth keyframe KF5 is generated for an area of the rug closest to the user.
[0305] FIG. 18 illustrates an example of generating a persistent pose for the map of FIG. 17 according to some embodiments. In some embodiments, a new persistent pose is created when the device measures a threshold distance traveled and / or when the application requests a new persistent pose (PP). In some embodiments, the threshold distance may be 3 meters, 5 meters, 20 meters, or any other suitable distance. Selecting a smaller threshold distance (e.g., 1m) may result in an increase in computational load because a larger number of PPs can be created and managed compared to a larger threshold distance. Selecting a larger threshold distance (e.g., 40m) results in fewer PPs being created and fewer PCFs being created, meaning that the virtual content associated with the PCF is relatively far (e.g., 30m) from the PCF, and errors can increase as the distance from the PCF to the virtual content increases, which may result in an increase in virtual content installation error.
[0306] In some embodiments, the PP may be created at the start of a new session. This initial PP can be considered zero and can be visualized as the center of a circle having a radius equal to the threshold distance. When the device reaches the perimeter of the circle and, in some embodiments, the application requests a new PP, the new PP may be placed at the current location of the device (threshold distance). In some embodiments, a new PP will not be created if the device is able to find an existing PP within the threshold distance from the new location of the device. In some embodiments, when a new PP (PP1150 in FIG. 14) is created, the device associates one or more of the closest keyframes to the PP. In some embodiments, the location of the PP relative to the keyframe may be based on the location of the device at the time the PP was created. In some embodiments, a PP will not be created unless the application requests a PP and the device travels the threshold distance.
[0307] In some embodiments, the application may request the PCF from the device when the application has virtual content to display to the user. The PCF request from the application may trigger a PP request, and a new PP will be created after the device has advanced the threshold distance. FIG. 18 illustrates, for example, a first persistent pose PP1 that can associate the closest keyframes (e.g., KF1, KF2, and KF3) by calculating the relative pose between the keyframe and the persistent pose. FIG. 18 also illustrates a second persistent pose PP2 that can associate the closest keyframes (e.g., KF4 and KF5).
[0308] FIG. 19 illustrates an example of generating a PCF for the map of FIG. 17 according to some embodiments. In the illustrated example, PCF1 may include PP1 and PP2. As described above, the PCF may be used to display the image data for the PCF. In some embodiments, each PCF may have coordinates in another coordinate frame (e.g., the world coordinate frame) and, for example, a PCF descriptor that uniquely identifies the PCF. In some embodiments, the PCF descriptor may be calculated based on the feature descriptors of the features in the frame associated with the PCF. In some embodiments, the various constellations of the PCF may be combined in a persistent mode that requires less data and less data transmission and may represent the real world.
[0309] FIGS. 20A-20C are schematic diagrams illustrating examples of establishing and using a persistent coordinate frame. FIG. 20A shows two users 4802A, 4802B with individual local tracking maps 4804A, 4804B that are not located with respect to the reference map. The origins 4806A, 4806B for the individual users are depicted by the coordinate system (e.g., the world coordinate system) within their individual areas. These origins of each tracking map can be local to each user because the origin depends on the orientation of that individual device when tracking is initiated.
[0310] As the sensors of the user device scan the environment, the device may capture images that, as described above in connection with FIG. 14, may contain features representing persistent objects such that those images can be classified as keyframes from which persistent poses can be created. In this example, the tracking map 4802A includes a persistent pose (PP) 4808A, and the tracking map 4802B includes a PP 4808B.
[0311] Also, as described above in connection with FIG. 14, some of the PPs may be classified as PCFs that are used to determine the orientation of virtual content for rendering it to the user. FIG. 20B shows that XR devices worn by individual users 4802A, 4802B can create local PCFs 4810A, 4810B based on the PPs 4808A, 4808B. FIG. 20C shows that persistent content 4812A, 4812B (e.g., virtual content) can be associated with the PCFs 4810A, 4810B by individual XR devices.
[0312] In this example, the virtual content may have a virtual content coordinate frame that can be used by the application that generates the virtual content, regardless of how the virtual content is to be displayed. The virtual content may be defined, for example, as a surface such as a triangle of a mesh at a specific location and angle relative to the virtual content coordinate frame. To render that virtual content to the user, the locations of those surfaces may be determined for the user who is to perceive the virtual content.
[0313] Associating virtual content with a PCF can simplify the calculations involved in determining the location of the virtual content for a user. The location of the virtual content for a user may be determined by applying a series of transformations. Some of those transformations may change and may be updated frequently. Others of those transformations may be stable and may not be updated as frequently or at all. Nevertheless, the transformations can be applied with a relatively low computational burden such that the location of the virtual content can be updated frequently for the user and provide a rendered virtual content with a realistic appearance.
[0314] In the examples of FIGS. 20A-20C, the device of user 1 has a coordinate system that may be related to a coordinate system that defines the origin of the map by the transformation rig1_T_w1. The device of user 2 has a similar transformation rig2_T_w2. These transformations are represented as six-degree transformations and may define translations and rotations for aligning the device coordinate system and the map coordinate system. In some embodiments, the transformation may be represented as two separate transformations, one defining a translation and the other defining a rotation. Thus, it should be understood that the transformation can be represented in a form that simplifies the calculations or otherwise provides an advantage.
[0315] The transformation between the origin of the tracking map and the PCF identified by an individual user device is represented as pcf1_T_w1 and pcf2_T_w2. In this example, the PCF and the PP are the same such that the same transformation also characterizes the PP.
[0316] The location of the user device with respect to the PCF can thus be calculated by successive application of these transformations such as rig1_T_pcf1=(rig1_T_w1) * (pcf1_T_w1).
[0317] As shown in FIG. 20C, the virtual content is positioned relative to the PCF using the transformation of obj1_T_pcf1. This transformation may be set by an application that generates virtual content and that can receive information from a world reconstruction system that describes physical objects relative to the PCF. To render the virtual content to the user, a transformation to the coordinate system of the user's device is calculated, which is the transformation obj1_t_w1=(obj1_T_pcf1) * (pcf1_T_w1) and can be calculated by relating the virtual content coordinate frame to the origin of the tracking map. That transformation can then be related to the user's device through a further transformation rig1_T_w1.
[0318] The location of the virtual content can change based on the output from the application that generates the virtual content. When it changes, the end-to-end transformation from the source coordinate system to the destination coordinate system can be recalculated. Additionally, the user's location and / or head pose can also change as the user moves. As a result, the transformation rig1_T_w1, as well as any end-to-end transformation that depends on the user's location or head pose, will change as it can change.
[0319] The transformation rig1_T_w1 may be updated with the user's movement based on tracking the user's position relative to stationary objects in the physical world. Such tracking may be performed by a headset tracking component that processes a sequence of images, or other components of the system, as described above. Such an update may be done by determining the user's pose relative to a stationary reference frame such as the PP.
[0320] In some embodiments, the location and orientation of the user device may be determined relative to the nearest persistent pose, or in this example, the PCF as PP is used as the PCF. Such determination may be made by identifying feature points that characterize the PP in the current image captured using sensors on the device. The location of the device relative to those feature points may be determined using image processing techniques such as stereoscopic image analysis. From this data, the system can calculate the change in transformation associated with the user's movement based on the relationship rig1_T_pcf1=(rig1_T_w1) * (pcf1_T_w1).
[0321] The system may determine and apply the transformation in an order that is computationally efficient. For example, the need to calculate rig1_T_w1 from the measurements that yield rig1_T_pcf1 can be avoided by both tracking the user's pose and defining the location of the virtual content relative to the PP or PCF constructed on the persistent pose. Thus, the transformation from the source coordinate system of the virtual content to the destination coordinate system of the user's device may be based on the measured transformation according to the expression (rig1_T_pcf1) * (obj1_t_pcf1), where the first transformation is measured by the system and the latter transformation is supplied by an application that defines the virtual content for rendering. In embodiments where the virtual content is positioned relative to the origin of the map, the end-to-end transformation may relate the virtual object coordinate system to the PCF coordinate system based on an additional transformation between the map coordinates and the PCF coordinates. In embodiments where the virtual content is positioned relative to a PP or PCF different from the one relative to which the user position is being tracked, a transformation between the two may be applied. Such a transformation may be fixed, for example, determined from the map where both appear.
[0322] The transformation-based approach may be implemented within a device with components that, for example, process sensor data and construct a tracking map. As part of that process, those components may identify feature points that can be used as persistent poses, which may in turn be transformed into PCFs. Those components limit the number of persistent poses generated for the map and provide a suitable spacing between persistent poses, as described above in connection with FIGS. 17-19, while allowing a user to be sufficiently close to a persistent pose location, regardless of the location within the physical environment, to accurately calculate the user's pose. As the persistent pose closest to the user is updated as a result of user movement, refinement to the tracking map, or other causes, any of the transformations used to calculate the location of virtual content for the user, which depends on the location of the PP (or PCF if used), may be updated and stored for use at least until the user moves away from that persistent pose. Note that by calculating and storing the transformations, the computational burden of calculating how often the location of virtual content is updated can be relatively low such that it can be performed with a relatively short latency.
[0323] FIGS. 20A-20C illustrate the positioning with respect to a tracking map, where each device has its own tracking map. However, the transformations may be generated for any map coordinate system. The persistence of content across a user session of an XR system can be achieved by using a persistent map. A shared experience for users can also be facilitated by using a map to which multiple user devices can be oriented.
[0324] In some embodiments, described in more detail below, the location of virtual content may be defined in relation to coordinates in a canonical map that is formatted so that any of a plurality of devices can use the map. Each device may maintain a tracking map and may determine changes in the user's pose relative to the tracking map. In this example, the transformation between the tracking map and the canonical map may be determined through a "localization" process, which may be performed by matching structures in the tracking map (such as one or more persistent poses) with one or more structures in the canonical map (such as one or more PCFs).
[0325] What is further described below are techniques for creating and using a canonical map in this way.
[0326] Deep keyframe
[0327] The techniques as described herein rely on the comparison of image frames. For example, to establish the position of a device relative to a tracking map, a new image may be captured using sensors worn by the user, and the XR system may search for an image that shares at least a predetermined amount of points of interest with the new image within the set of images used to create the tracking map. As an example of another scenario involving the comparison of image frames, a tracking map may be localized with respect to a canonical map by first finding an image frame associated with a persistent pose in the tracking map that is similar to an image frame associated with a PCF in the canonical map. Alternatively, the transformation between two canonical maps may be calculated by first finding similar image frames in the two maps.
[0328] The deep keyframe provides a method for reducing the amount of processing required to identify similar image frames. For example, in some embodiments, the comparison can be made between image features (e.g., "2D features") within a new 2D image and 3D features within the map. Such a comparison can be made in any suitable manner, such as by projecting the 3D image into a 2D plane. Conventional methods, such as the bag of words (BoW), search for the 2D features of a new image within a database that includes all of the 2D features within the map, which can require significant computational resources, particularly when the map represents a large area. The conventional method then locates images that share at least one of the 2D features with the new image, which may include images that are not useful for locating significant 3D features within the map. The conventional method then locates 3D features that are not significant with respect to the 2D features within the new image.
[0329] The inventors recognize and understand a technique for reading out images within a map that uses fewer memory resources (e.g., one quarter of the memory resources used by BoW), higher efficiency (e.g., 2.5 ms processing time per keyframe, 100 μs for comparison of 500 keyframes), and higher accuracy (e.g., 20% better recall than BoW for a 1,024-dimensional model, 5% better recall than BoW for a 256-dimensional model).
[0330] Descriptors that can be used to compare an image frame with other image frames to reduce computation may be calculated for the image frame. The descriptors may be stored instead of, or in addition to, the image frame and feature points. In a map where persistent poses and / or PCFs can be generated from the image frame, the descriptors of the image frame or frames from which each persistent pose or PCF was generated may be stored as part of the persistent pose and / or PCF.
[0331] In some embodiments, the descriptor may be calculated as a function of feature points within the image frame. In some embodiments, a neural network is configured to calculate a unique frame descriptor for representing an image. The image may have a resolution higher than 1 megabyte such that sufficient details of the 3D environment within the field of view of the device worn by the user are captured within the image. The frame descriptor can be much smaller, such as a sequence of numbers, for example, within the range of 128 bytes to 512 bytes or any number in between.
[0332] In some embodiments, the neural network is trained such that the calculated frame descriptor indicates the similarity between images. Images within the map can be located by identifying the nearest images that may have frame descriptors within a predetermined distance of the frame descriptor for the new image in a database comprising the images used to generate the map. In some embodiments, the distance between images may be represented by the difference between the frame descriptors of two images.
[0333] FIG. 21 is a block diagram illustrating a system for generating descriptors for individual images, according to some embodiments. In the illustrated example, a frame embedding generator 308 is shown. The frame embedding generator 308 may be used in combination with the server 20 in some embodiments, but alternatively, or in addition, may be executed in whole or in part within one of the XR devices 12.1 and 12.2, or any other device that processes images for comparison with other images.
[0334] In some embodiments, the frame embedding generator may be configured to generate a data representation of an image reduced from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes) that still represents the content within the image despite the reduced size. In some embodiments, the frame embedding generator may be used to generate a data representation for an image, which may be a keyframe or frame used in other methods. In some embodiments, the frame embedding generator 308 may be configured to convert an image at a particular location and orientation into a unique sequence of numbers (e.g., 256 bytes). In the illustrated embodiment, an image 320 captured by an XR device may be processed by a feature extractor 324 to detect a point of interest 322 within the image 320. The point of interest may or may not be derived from identified feature points as described above with respect to feature 1120 (FIG. 14) or otherwise described herein. In some embodiments, the point of interest may be represented by a descriptor as described above with respect to descriptor 1130 (FIG. 14), which may be generated using a deep sparse feature method. In some embodiments, each point of interest 322 may be represented by a sequence of numbers (e.g., 32 bytes). For example, there may be n features (e.g., 100), and each feature may be represented by a 32-byte sequence.
[0335] In some embodiments, the frame embedding generator 308 may include a neural network 326. The neural network 326 may include a multi-layer perceptron unit 312 and a max pooling unit 314. In some embodiments, the multi-layer perceptron (MLP) unit 312 may comprise a multi-layer perceptron, which may be trained. In some embodiments, the point of interest 322 (e.g., a descriptor for the point of interest) may be reduced by the multi-layer perceptron 312 and output as a weighted combination 310 of descriptors. For example, the MLP may reduce n features to m features, where m is less than n features.
[0336] In some embodiments, the MLP unit 312 may be configured to perform matrix multiplication. The multi-layer perceptron unit 312 receives a plurality of points of interest 322 of the image 320 and converts each point of interest into an individual column of numbers (e.g., 256). For example, there may be 100 features, and each feature may be represented by a column of 256 numbers. The matrix may be created to have 100 horizontal rows and 256 vertical columns in this example. Each row may have a series of 256 numbers that vary in size, some being smaller and some being larger. In some embodiments, the output of the MLP may be an n×256 matrix, where n represents the number of features extracted from the image. In some embodiments, the output of the MLP may be an m×256 matrix, where m is the number of points of interest reduced from n.
[0337] In some embodiments, the MLP 312 may have a training phase and a usage phase during which model parameters for the MLP are determined. In some embodiments, the MLP may be trained as illustrated in FIG. 25. The input training data may include data in three sets, the three sets comprising 1) a query image, 2) a positive sample, and 3) a negative sample. The query image may be considered a reference image.
[0338] In some embodiments, the positive sample may comprise an image that is similar to the query image. For example, in some embodiments, being similar means having the same object in both the query and positive sample images, but it can be viewed from different angles. In some embodiments, being similar means having the same object in both the query and positive sample images, but it may have an object that is offset (e.g., to the left, right, up, down) with respect to other images.
[0339] In some embodiments, the negative sample may comprise an image that does not resemble the query image. For example, in some embodiments, the non-resembling image may not contain any object that is prominent within the query image, or may contain only a small portion (e.g., <10%, 1%) of the prominent objects within the query image. In contrast, the resembling image may, for example, have a majority (e.g., >50%, or >75%) of the objects within the query image.
[0340] In some embodiments, the point of interest may be extracted from an image within the input training data and may be transformed into a feature descriptor. These descriptors may be calculated for both the training images, as shown in FIG. 25, and for the features extracted during operation of the frame embedding generator 308 of FIG. 21. In some embodiments, a deep sparse feature (DSF) process may be used to generate the descriptors (e.g., DSF descriptors), as described in U.S. Patent Application No. 16 / 190,948. In some embodiments, the DSF descriptors are of size n×32. The descriptors may then be passed through a model / MLP to create a 256-byte output. In some embodiments, the model / MLP may have the same structure as MLP312 such that once the model parameters are set through training, the resulting trained MLP can be used as MLP312.
[0341] In some embodiments, a feature descriptor (e.g., 256 bytes output from an MLP model) may then be sent to a triplet margin loss module (used only during the training phase of the MLP neural network and not during the usage phase). In some embodiments, the triplet margin loss module is configured to select parameters for the model so as to reduce the difference between 256 bytes output from a query image and 256 bytes output from a positive sample and increase the difference between 256 bytes output from the query image and 256 bytes output from a negative sample. In some embodiments, the training phase may include feeding a plurality of triplet input images into the learning process and determining model parameters. This training process may continue, for example, until the difference regarding the positive image is minimized and the difference regarding the negative image is maximized, or until other suitable termination criteria are reached.
[0342] Referring back to FIG. 21, the frame embedding generator 308 may here include a pooling layer, illustrated as a max pooling unit 314. The max pooling unit 314 may analyze each column and determine the maximum number within an individual column. The max pooling unit 314 may combine the maximum value of each column of the output matrix of the MLP 312 into a global feature column 316 of a number, e.g., 256. It should be understood that images processed within the XR system may desirably have high-resolution frames, potentially with millions of pixels. The global feature column 316 occupies relatively little memory and is relatively small in number, being easily searchable compared to an image (e.g., with a resolution higher than 1 megabyte). Thus, it is possible to search for an image without analyzing each original frame from the camera, and it is also less expensive to store 256 bytes instead of the full frame.
[0343] FIG. 22 is a flowchart illustrating a method 2200 for calculating image descriptors according to some embodiments. The method 2200 may start by receiving a plurality of images captured by an XR device worn by a user (act 2202). In some embodiments, the method 2200 may include determining one or more keyframes from the plurality of images (act 2204). In some embodiments, act 2204 may be skipped and / or may occur after step 2210 instead.
[0344] The method 2200 may include identifying one or more points of interest in the plurality of images using an artificial neural network (act 2206) and calculating feature descriptors for the individual points of interest using the artificial neural network (act 2208). The method may include calculating a frame descriptor for representing an image, at least in part, for each image, based on the calculated feature descriptors for the identified points of interest in the image, using the artificial neural network (act 2210).
[0345] FIG. 23 is a flowchart illustrating a location identification method 2300 using image descriptors according to some embodiments. In this example, a new image frame depicting the current location of the XR device may be compared with the image frames stored in relation to points in the map (such as persistent poses or PCFs as described above). Method 2300 may start by receiving a new image captured by an XR device worn by a user (act 2302). Method 2300 may include identifying one or more of the closest keyframes in a database that include keyframes used to generate one or more maps (act 2304). In some embodiments, the closest keyframes may be identified based on approximate spatial information and / or previously determined spatial information. For example, the approximate spatial information may indicate that the XR device is within a geographic area represented by a 50m x 50m area of the map. Image matching may be performed only for points within that area. As another example, based on tracking, the XR system may know that the XR device was previously close to a first persistent pose in the map and is moving in the direction of a second persistent pose in the map. That second persistent pose may be considered the closest persistent pose, and the keyframe stored with it may be considered the closest keyframe. Alternatively, or in addition, other metadata such as GPS data or WiFi fingerprints may also be used to select the closest keyframe or set of closest keyframes.
[0346] Regardless of how the nearest keyframe is selected, the frame descriptor may be used to determine whether it matches any of the frames selected as those to which the new image is associated with a neighboring persistent pose. The determination may be made by comparing the frame descriptor of the new image with the frame descriptors of the nearest keyframe or a subset of keyframes in a database selected in any other suitable manner, and selecting a keyframe with a frame descriptor within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors may be calculated by obtaining the difference between two columns of numbers that may represent the two frame descriptors. In embodiments where the columns are treated as columns of multiple quantities, the difference may be calculated as a vector difference.
[0347] Once a matching image frame is identified, the orientation of the XR device with respect to that image frame may be determined. Method 2300 may include performing feature matching (act 2306) on the 3D features in the map corresponding to the identified nearest keyframe, and calculating the pose of the device worn by the user (act 2308) based on the feature matching result. Thus, the matching, which is computationally intensive in calculating the feature points in the two images, may be performed for only the one image that has already been determined to be a likely match for the new image.
[0348] FIG. 24 is a flowchart illustrating a method 2400 for training a neural network according to some embodiments. Method 2400 may begin with generating a dataset (act 2402) comprising a plurality of image sets. Each of the plurality of image sets may include a query image, a positive sample image, and a negative sample image. In some embodiments, the plurality of image sets may include synthetic recording pairs configured, for example, to teach a neural network basic information such as shape. In some embodiments, the plurality of image sets may include real recording pairs that may be recorded from the physical world.
[0349] In some embodiments, the positive correspondence may be calculated by fitting a fundamental matrix between two images. In some embodiments, the sparse overlap may be calculated as the intersection over union (IoU) on the union of the points of interest seen in both images. In some embodiments, a positive sample may include at least 20 points of interest that are identical within the query image and serve as positive correspondences. A negative sample may include fewer than 10 positive correspondence points. A negative sample may have fewer than half sparse points that overlap with the analysis points of the query image.
[0350] Method 2400 may include calculating a loss (act 2404) by comparing a query image with positive and negative sample images for each image set. Method 2400 may include modifying the artificial neural network (act 2406) based on the calculated loss such that the distance between the frame descriptor generated by the artificial neural network for the query image and the frame descriptor for the positive sample image is less than the distance between the frame descriptor for the query image and the frame descriptor for the negative sample image.
[0351] A method and apparatus configured to generate global descriptors for individual images are described above, but it should be understood that the method and apparatus may also be configured to generate descriptors for individual maps. For example, a map may include a plurality of keyframes, each having a frame descriptor as described above. A max pooling unit may analyze the frame descriptors of the keyframes of the map and combine the frame descriptors into a unique map descriptor for the map.
[0352] Furthermore, it should be understood that other architectures may also be used for processing as described above. For example, separate neural networks are described for generating DSF descriptors and frame descriptors. Such an approach is computationally efficient. However, in some embodiments, the frame descriptor may be generated from selected feature points without first generating the DSF descriptor. Ranking and Merging of Maps
[0353] What is described herein are methods and apparatuses for ranking and merging multiple environmental maps within a cross-reality system. Map merging can enable maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking the maps can enable efficient implementation of techniques as described herein, including map merging, which involves selecting maps from a set of maps based on similarity. In some embodiments, for example, a set of reference maps, formatted in a way that can be accessed by any of several XR devices, may be maintained by the system. These reference maps may be formed by merging selected tracking maps from those devices with other tracking maps or previously stored reference maps. The reference maps may be ranked, for example, to select one or more reference maps, merge with a new tracking map, and / or select one or more reference maps from the set for use when used within the device.
[0354] To provide a realistic XR experience to the user, the XR system must understand the user's physical surroundings in order to correctly correlate the location of virtual objects with real objects. Information about the user's physical surroundings may be obtained from an environmental map regarding the user's location.
[0355] The inventors recognized and appreciated the true value that an XR system can provide an enhanced XR experience to multiple users sharing the same world with real and / or virtual content, by enabling efficient sharing of environmental maps of the real / physical world collected by multiple users, regardless of whether those users are present in the world at the same or different times. However, significant challenges exist in providing such a system. Such a system may store multiple maps generated by multiple users and / or the system may store multiple maps generated at different times. For example, for operations that may be performed using previously generated maps such as localization, substantial processing may be required to identify the relevant environmental map of the same world (e.g., the same real-world location) from all the environmental maps collected within the XR system. In some embodiments, there may be only a few environmental maps that a device can access, for example, for localization. In some embodiments, there may be a large number of environmental maps that a device can access. The inventors recognized and appreciated the true value of a technique for quickly and accurately ranking the relevance of environmental maps from any possible set of environmental maps, such as the population of all the reference maps 120 in FIG. 28. High-ranked maps may then be selected for further processing, such as rendering virtual objects to realistically interact with the physical world around the user on a user display, or merging the map data collected by that user with the stored maps to create a larger or more accurate map.
[0356] In some embodiments, a stored map associated with a task for a user at a location in the physical world may be identified by filtering the stored map based on a plurality of criteria. Those criteria may indicate a comparison of a tracking map generated by the user's wearable device at that location with candidate environment maps stored in a database. The comparison may be performed based on metadata associated with a map, such as a Wi-Fi fingerprint detected by a device generating the map, and / or a set of BSSIDs to which the device is connected while forming the map. The comparison may also be performed based on the compressed or decompressed content of the map. The comparison based on the compressed representation may be performed, for example, by comparing vectors calculated from the map content. The comparison based on the decompressed map may be performed, for example, by locating the tracking map within the stored map or vice versa. A plurality of comparisons may be performed in an order based on the calculation time required to reduce the number of candidate maps under consideration, and comparisons involving less calculation may be performed before other comparisons that require more calculation.
[0357] FIG. 26 depicts an AR system 800 configured to rank and merge one or more environmental maps according to some embodiments. The AR system may include a passable world model 802 of the AR device. Information for capturing the passable world model 802 may originate from sensors on the AR device, which may include computer-executable instructions for performing some or all of the processing for converting sensor data into a map, stored within a processor 804 (e.g., the local data processing module 570 in FIG. 4). Such a map may be a tracking map that can be constructed as sensor data is collected as the AR device operates within the area. Along with the tracking map, area attributes may be provided to indicate the area represented by the tracking map. These area attributes may be geographical location identifiers such as IDs used by the AR system to represent coordinates or locations presented as latitude and longitude. Alternatively, or in addition, the area attributes may be measured characteristics that have a high likelihood of being unique for that area. The area attributes may be derived, for example, from parameters of wireless networks detected within the area. In some embodiments, the area attributes may be associated with unique addresses of access points that the AR system is in proximity to and / or connected to. For example, the area attributes may be associated with the MAC addresses or basic service set identifiers (BSSIDs) of 5G base stations / routers, Wi-Fi routers, and the like.
[0358] In the example of FIG. 26, the tracking map may be merged with other maps of the environment. The map ranking portion 806 receives the tracking map from the device PW802, communicates with the map database 808, and selects and ranks environmental maps from the map database 808. The selected maps that are ranked higher are sent to the map merging portion 810.
[0359] The map merge part 810 may perform the merge process on the map transmitted from the map ranking part 806. The merge process may involve merging some or all of the tracking map and the ranked map and transmitting the new merged map to the passable world model 812. The map merge part may merge the maps by identifying the maps depicting the overlapping parts of the physical world. Those overlapping parts may be aligned so that the information in both maps can be aggregated into the final map. The reference map may be merged with other reference maps and / or tracking maps.
[0360] Aggregation may involve extending one map with information from another map. Alternatively, or in addition, aggregation may involve adjusting the representation of the physical world in one map based on information in another map. The latter map may represent, for example, that an object generating a feature point has moved so that the map can be updated based on the latter information. Alternatively, two maps may characterize the same area with different feature points, and aggregation may involve selecting a set of feature points from the two maps and better representing that area. Regardless of the specific processing that occurs in the process of merging, in some embodiments, the PCF from all the maps being merged may be retained so that an application positioning the content relative to them can continue to do so. In some embodiments, the merging of the maps may result in redundant persistent postures, and some of the persistent postures may be deleted. When the PCF is associated with the persistent postures to be deleted, merging the maps may involve modifying the PCF to be associated with the persistent postures remaining in the map after the merge.
[0361] In some embodiments, as maps are extended and / or updated, they may be refined. The refinement may involve calculations to reduce internal inconsistencies between feature points that are likely to represent the same object in the physical world. The inconsistencies can arise from inaccuracies within the poses associated with keyframes that supply feature points representing the same object in the physical world. Such inconsistencies can arise, for example, from an XR device calculating a pose relative to a tracking map, which in turn is constructed based on estimating poses such that errors in the pose estimation accumulate over time and create a “drift” in the pose accuracy. The map may be refined by performing bundle adjustment or other operations to reduce the inconsistencies between feature points from multiple keyframes.
[0362] In response to the refinement, the location of persistent points relative to the origin of the map may change. Thus, the transforms associated with such persistent points, such as persistent poses or PCFs, may also change. In some embodiments, the XR system may recalculate the transforms associated with any persistent points that change in relation to map refinement (whether performed as part of a merge operation or for other reasons). These transforms may be push-distributed from the component that calculates the transform to the component that uses the transform such that any use of the transform can be based on the updated location of the persistent point.
[0363] The passable world model 812 may be a cloud model, which may be shared by multiple AR devices. The passable world model 812 may store an environmental map in the map database 808 or otherwise have access thereto. In some embodiments, when a previously calculated environmental map is updated, the previous version of the map may be deleted to remove the outdated map from the database. In some embodiments, when a previously calculated environmental map is updated, the previous version of the map may be archived to enable reading / viewing of the previous version of the environment. In some embodiments, permissions may be set such that only an AR system having certain read / write access may trigger deletion / archiving of the previous version of the map.
[0364] These environmental maps created from a tracking map supplied by one or more AR devices / systems may be accessed by AR devices within the AR system. The map ranking portion 806 may also be used when supplying the environmental map to the AR device. The AR device may send a message requesting an environmental map for its current location, and the map ranking portion 806 may be used to select and rank the environmental map associated with the requesting device.
[0365] In some embodiments, the AR system 800 may include a downsampling portion 814 configured to receive the merged map from the cloud PW812. The merged map received from the cloud PW812 may be in a storage format for the cloud, which may include high-resolution information such as a large number of PCFs per square meter or multiple image frames or a large set of feature points associated with the PCFs. The downsampling portion 814 may be configured to downsample the cloud format map to a format suitable for storage on the AR device. The device format map may have less data, such as less PCFs or less data stored per PCF, and may be able to accommodate the limited local computing power and storage space of the AR device.
[0366] FIG. 27 is a simplified block diagram illustrating a plurality of reference maps 120 that may be stored in a remote storage medium, such as in the cloud. Each reference map 120 may include a plurality of reference map identifiers indicating the location of the reference map in physical space, such as any location on Earth that is a planet. These reference map identifiers may include one or more of the following identifiers: an area identifier represented by a range of longitude and latitude, a frame descriptor (e.g., the global feature column 316 in FIG. 21), a Wi-Fi fingerprint, a feature descriptor (e.g., the feature descriptor 310 in FIG. 21), and a device identifier indicating one or more devices that contributed to the map. In the illustrated embodiment, since the reference maps 120 may exist on the surface of the Earth, they are geographically arranged in a two-dimensional pattern. The reference maps 120 may be uniquely identifiable by corresponding longitude and latitude because any reference map having overlapping longitude and latitude may be merged into a new reference map. FIG. 28 is a schematic diagram illustrating a method of selecting a reference map that can be used to locate a new tracking map relative to one or more reference maps according to some embodiments. The method may start, as an example, by accessing (act 120) a parent set of reference maps 120 that may be stored in a database within a traversable world (e.g., traversable world module 538). The parent set of reference maps may include reference maps from all previously visited locations. The XR system may filter the parent set of all reference maps to a small subset or only a single map. It should be understood that in some embodiments, due to bandwidth limitations, it may not be possible to send all reference maps to the viewing device. Selecting a subset that is selected as a likely candidate for matching to the tracking map for transmission to the device may reduce the bandwidth and latency associated with accessing the remote database of maps.
[0367] This method may include filtering a parent set of reference maps based on an area with a predetermined size and shape (act 300). In the embodiment illustrated in FIG. 27, each square may represent an area. Each square may cover 50m × 50m. Each square may have six neighboring areas. In some embodiments, act 300 may select at least one matching reference map 120 that covers the longitude and latitude, including the longitude and latitude of the position identifier received from the XR device, as long as at least one map exists at that longitude and latitude. In some embodiments, act 300 may select at least one neighboring reference map that covers the longitude and latitude and is adjacent to the matching reference map. In some embodiments, act 300 may select a plurality of matching reference maps and a plurality of neighboring reference maps. Act 300 may, for example, reduce the number of reference maps by about one-tenth, for example, from thousands to hundreds, to form a first filtered selection. Alternatively, or in addition, criteria other than latitude and longitude may be used to identify neighboring maps. The XR device may have been previously located using the reference maps in the set, for example, as part of the same session. The cloud service may retain information about the XR device, including the previously located maps. In this embodiment, the maps selected in act 300 may include those that cover an area adjacent to the map at which the XR device was located.
[0368] This method may include filtering a first filtered selection of reference maps based on a Wi-Fi fingerprint (act 302). Act 302 may determine latitude and longitude based on a Wi-Fi fingerprint received as part of a location identifier from an XR device. Act 302 may compare the latitude and longitude from the Wi-Fi fingerprint with the latitude and longitude of reference map 120 and determine one or more reference maps that form a second filtered selection. Act 302 may reduce the number of reference maps to about one tenth, for example, from hundreds to dozens (e.g., 50) of reference maps that form the second selection. For example, the first filtered selection may include 130 reference maps, the second filtered selection may include 50 of the 130 reference maps, and may not include the remaining 80 of the 130 reference maps.
[0369] This method may include filtering a second filtered selection of reference maps based on keyframes (act 304). Act 304 may compare data representing an image captured by an XR device with data representing reference map 120. In some embodiments, the data representing the image and / or map may include feature descriptors (e.g., DSF descriptors in FIG. 25) and / or global feature strings (e.g., 316 in FIG. 21). Act 304 may provide a third filtered selection of reference maps. In some embodiments, the output of act 304 may be, for example, only 5 of the 50 reference maps identified following the second filtered selection. Map transmitter 122 then transmits one or more reference maps to the viewing device based on the third filtered selection. Act 304 may reduce the number of reference maps to about one tenth, forming a third selection, for example, from dozens to single-digit numbers of reference maps (e.g., 5). In some embodiments, the XR device may receive the reference maps within the third filtered selection and attempt to locate itself within the received reference maps.
[0370] For example, the act 304 may filter the reference map 120 based on the global feature sequence 316 of the reference map 120 and the global feature sequence 316 based on an image captured by the vision device (e.g., an image that may be part of a local tracking map for the user). Each of the reference maps 120 in FIG. 27 thus has one or more global feature sequences 316 associated therewith. In some embodiments, the global feature sequence 316 may be obtained when the XR device submits an image or feature details to the cloud, and the cloud processes the image or feature details and generates the global feature sequence 316 for the reference map 120.
[0371] In some embodiments, the cloud may receive feature details of a live / new / current image captured by the vision device, and the cloud may generate the global feature sequence 316 for the live image. The cloud may then filter the reference map 120 based on the live global feature sequence 316. In some embodiments, the global feature sequence may be generated on a local vision device. In some embodiments, the global feature sequence may be generated remotely, e.g., on the cloud. In some embodiments, the cloud may transmit the filtered reference map, together with the global feature sequence 316 associated with the filtered reference map, to the XR device. In some embodiments, when the vision device locates its tracking map relative to the reference map, this may be done by matching the global feature sequence 316 of the local tracking map with the global feature sequence of the reference map.
[0372] It should be understood that the operation of the XR device does not have to perform all of the actions (300, 302, 304). For example, when the parent set of the reference maps is relatively small (e.g., 500 maps), the XR device attempting to locate may filter the parent set of the reference maps based on Wi-Fi fingerprints (e.g., action 302) and key frames (e.g., action 304), but may omit filtering based on area (e.g., action 300). Further, the maps do not necessarily have to be compared as a whole. In some embodiments, for example, the comparison of two maps can result in the identification of common persistence points such as persistent poses or PCFs that appear in both the new map and the map selected from the parent set of maps. In that case, descriptors may be associated with the persistence points, and those descriptors may be compared.
[0373] FIG. 29 is a flowchart illustrating a method 900 for selecting one or more ranked environmental maps, according to some embodiments. In the illustrated embodiment, ranking is performed for the user's AR device that creates a tracking map. Thus, the tracking map is available for use in ranking the environmental maps. In embodiments where the tracking map is not available, some or all of the portion of selecting and ranking the environmental maps that does not explicitly rely on the tracking map may be used.
[0374] Method 900 may begin with act 902, where a set of maps (which may be formatted as reference maps) is accessed from a database of environmental maps in the vicinity of where the tracking map was formed and may then be filtered for ranking. Additionally, in act 902, at least one area attribute regarding the area in which the user's AR device is operating is determined. In a scenario where the user's AR device is constructing a tracking map, the area attribute may correspond to the area over which the tracking map was created. As a specific example, the area attribute may be calculated based on signals received from access points to a computer network during the time the AR device was calculating the tracking map.
[0375] FIG. 30 depicts an exemplary map ranking portion 806 of an AR system 800 according to some embodiments. The map ranking portion 806 may be executed within a cloud computing environment since it may include portions executed on an AR device and portions executed on a remote computing system such as a cloud. The map ranking portion 806 may be configured to implement at least a portion of method 900.
[0376] Figure 31A depicts an example of a tracking map (TM) 1102 and area attributes AA1 - AA8 of environmental maps CM1 - CM4 in a database according to some embodiments. As shown, the environmental map may be associated with a plurality of area attributes. Area attributes AA1 - AA8 may include parameters of a wireless network detected by an AR device that calculates the tracking map 1102, for example, the basic service set identifier (BSSID) of the network to which the AR device is connected, and / or, for example, the strength of the received signal of an access point to the wireless network through network tower 1104. The parameters of the wireless network may conform to protocols including Wi-Fi and 5G NR. In the example illustrated in Figure 32, the area attribute is a fingerprint of the area in which the user AR device collected sensor data and formed a tracking map.
[0377] Figure 31B depicts an example of a determined geographical location 1106 of a tracking map 1102 according to some embodiments. In the example shown, the determined geographical location 1106 includes a centroid point 1110 and an area 1108 surrounding the centroid point. It should be understood that the determination of the geographical location in this application is not limited to the illustrated format. The determined geographical location may have any suitable format, for example, including different area shapes. In this example, the geographical location is determined from area attributes using a database that associates area attributes with geographical locations. The database is commercially available and is, for example, a database that associates Wi-Fi fingerprints, represented as latitude and longitude, with locations and can be used for this operation.
[0378] In the embodiment of Figure 29, the map database containing the environmental map may also include location data regarding those maps, including the latitude and longitude covered by the map. The processing in act 902 may involve selecting a set of environmental maps from the database that covers the same latitude and longitude determined for the area attributes of the tracking map.
[0379] Action 904 is the first filtering of the set of environmental maps accessed in Action 902. In Action 902, the environmental maps are retained in the set based on their proximity to the geographical location of the tracking map. This filtering step may be performed by comparing the latitudes and longitudes associated with the tracking map and the environmental maps in the set.
[0380] Figure 32 depicts an example of Action 904 according to some embodiments. Each area attribute may have a corresponding geographical location 1202. The set of environmental maps may include environmental maps with at least one area attribute having a geographical location that overlaps the determined geographical location of the tracking map. In the illustrated example, the identified set of environmental maps includes environmental maps CM1, CM2, and CM4, each having at least one area attribute having a geographical location that overlaps the determined geographical location of the tracking map 1102. The environmental map CM3, associated with the area attribute AA6, is not included in the set because it is outside the determined geographical location of the tracking map.
[0381] Other filtering steps may also be performed on the set of environment maps to reduce / rank the number of environment maps in the set that are ultimately processed (e.g., for map merging or for providing passable world information to the user device). Method 900 may include filtering (act 906) the set of environment maps based on the similarity of one or more identifiers of network access points associated with the environment maps of the set of tracking maps and environment maps. During map formation, a device that collects sensor data and generates a map may be connected to the network through a network access point, such as through Wi-Fi or a similar wireless communication protocol. The access point may be identified by a BSSID. As the user device moves through an area, collects data, and forms a map, it may connect to multiple different access points. Similarly, when multiple devices supply information for forming a map, the devices may be connected through different access points, and thus, for the same reason, there may be multiple access points used in forming the map. Therefore, there may be multiple access points associated with a map, and the set of access points may be an indication of the location of the map. The strength of the signal from the access point, which may be reflected as an RSSI value, may provide additional geographic information. In some embodiments, a list of BSSIDs and RSSI values may form area attributes for the map.
[0382] In some embodiments, filtering a set of environment maps based on the similarity of one or more identifiers of network access points may include retaining, within the set of environment maps, an environment map with the highest Jaccard similarity to at least one area attribute of a tracking map based on one or more identifiers of network access points. FIG. 33 depicts an example of act 906 according to some embodiments. In the illustrated example, the network identifier associated with area attribute AA7 can be determined as the identifier for tracking map 1102. The set of environment maps after act 906 may include environment map CM2, which may have an area attribute within a higher Jaccard similarity to AA7, and environment map CM4, which also includes area attribute AA7. Environment map CM1 is not included in the set because it has the lowest Jaccard similarity to AA7.
[0383] The processing in acts 902-906 may be performed without actually accessing the content of the maps stored in the map database based on the metadata associated with the maps. Other processing may involve accessing the content of the maps. Act 908 indicates accessing the environment maps remaining in the subset after filtering based on the metadata. It should be understood that this act may be performed either earlier or later in the process if the subsequent operations can be performed using the content being accessed.
[0384] Method 900 may include filtering a set of environment maps (act 910) based on a similarity of metrics representing the content of the environment maps of a set of tracking maps and environment maps. The metrics representing the content of the tracking maps and environment maps may include a vector of values calculated from the content of the maps. For example, the deep keyframe descriptors as described above, calculated for one or more keyframes used in forming the maps, may provide a metric for comparison of maps or portions of maps. The metrics may be calculated from the maps read in act 908, or may be pre-calculated and stored as metadata associated with those maps. In some embodiments, filtering a set of environment maps based on a similarity of metrics representing the content of the environment maps of a set of tracking maps and environment maps may include retaining, within the set of environment maps, environment maps with a minimum vector distance between a vector of the characteristics of the tracking map and a vector representing the environment maps within the set of environment maps.
[0385] Method 900 may further include filtering a set of environment maps (act 912) based on a degree of matching between a portion of the tracking map and a portion of the environment maps of the set of environment maps. The degree of matching may be determined as part of a localization process. As a non-limiting example, localization may be performed by identifying salient points within the tracking map and environment map that are similar enough that they may represent the same portion of the physical world. In some embodiments, the salient points may be features, feature descriptors, keyframes, key rigs, persistent poses, and / or PCFs. The set of salient points within the tracking map may then be aligned to produce a best fit with the set of salient points within the environment map. An average squared distance between corresponding salient points may be calculated and, if it falls below a threshold for a particular region of the tracking map, is used as an indication that the tracking map and environment map represent the same region of the physical world.
[0386] In some embodiments, filtering the set of environmental maps based on the degree of matching between a portion of the tracking map and a portion of the environmental map of the set of environmental maps may include calculating the volume of the physical world represented by the tracking map, which is also represented within the environmental maps of the set of environmental maps, and retaining within the set of environmental maps an environmental map with a calculated volume larger than the filtered-out environmental maps of the set. FIG. 34 depicts an example of act 912 according to some embodiments. In the illustrated example, the set of environmental maps after act 912 includes environmental map CM4, which has an area 1402 that matches the area of tracking map 1102. Environmental map CM1 is not included in the set because it does not have an area that matches the area of tracking map 1102.
[0387] In some embodiments, the set of environmental maps may be filtered in the order of act 906, act 910, and act 912. In some embodiments, the set of environmental maps may be filtered based on act 906, act 910, and act 912, which may be performed in an order based on the processing required to perform the filtering from lowest to highest. Method 900 may include loading a set of environmental maps and data (act 914).
[0388] In the illustrated embodiment, the user database stores an area identifier indicating the area where the AR device was used. The area identifier may be an area attribute, which may include parameters of the wireless network detected by the AR device during use. The map database may store a plurality of environmental maps constructed from the data supplied by the AR device and the associated metadata. The associated metadata may include an area identifier derived from the area identifier of the AR device that supplied the data from which the environmental map was constructed. The AR device may send a message to the PW module indicating that a new tracking map is being created or is in the process of being created. The PW module may calculate an area identifier for the AR device and update the user database based on the received parameters and / or the calculated area identifier. The PW module may also determine an area identifier associated with the AR device requesting the environmental map, identify a set of environmental maps from the map database based on the area identifier, filter the set of environmental maps, and transmit the filtered set of environmental maps to the AR device. In some embodiments, the PW module may filter the set of environmental maps based on one or more criteria including, for example, the geographical location of the tracking map, the similarity of one or more identifiers of the network access points associated with the environmental maps of the set of tracking and environmental maps, the similarity of a metric representing the content of the environmental maps of the set of tracking and environmental maps, and the degree of matching between a portion of the tracking map and a portion of the environmental maps of the set of environmental maps.
[0389] Although some aspects of some embodiments have been described so far, it should be understood that various modifications, corrections, and improvements will readily occur to those skilled in the art. As an example, the embodiments are described in relation to an augmented (AR) environment. It should be understood that some or all of the techniques described herein may be applied within an MR environment, and more generally, within other XR environments and VR environments.
[0390] As another example, the embodiments are described in relation to devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), discrete applications, and / or any suitable combination of devices, networks, and discrete applications.
[0391] Further, FIG. 29 provides an example of a criterion that can be used to filter a candidate map and result in a set of highly ranked maps. Other criteria may be used instead of, or in addition to, the criteria described. For example, if multiple candidate maps have similar values of a metric used to filter out less desirable maps, the characteristics of the candidate maps may be used to determine which candidate maps are retained as candidate maps or filtered out. For example, larger or denser candidate maps may be preferred over smaller candidate maps. In some embodiments, FIGS. 27-28 may illustrate all or part of the systems and methods described in FIGS. 29-34.
[0392] FIGS. 35 and 36 are schematic diagrams illustrating an XR system configured to rank and merge multiple environmental maps according to some embodiments. In some embodiments, the passable world (PW) may determine when to trigger ranking and / or merging of maps. In some embodiments, determining which maps should be used may, according to some embodiments, be at least partially based on the deep keyframes described above in relation to FIGS. 21-25.
[0393] FIG. 37 is a block diagram illustrating a method 3700 for creating an environmental map of the physical world, according to some embodiments. The method 3700 may begin with localizing a tracking map captured by an XR device worn by a user with respect to a group of reference maps (e.g., reference maps selected by the method of FIG. 28 and / or method 900 of FIG. 29) (act 3702). Act 3702 may include localizing key rigs of the tracking map among the group of reference maps. The localization result of each key rig may include the localized pose of the key rig and a set of 2D / 3D feature correspondences.
[0394] In some embodiments, the method 3700 may include splitting the tracking map into connected components (act 3704), which may enable robust merging of the map by merging the connected fragments. Each connected component may include key rigs that are within a predetermined distance. The method 3700 may include merging one or more connected components that are larger than a predetermined threshold into one or more reference maps (act 3706) and removing the merged connected components from the tracking map.
[0395] In some embodiments, the method 3700 may include merging the reference maps of the group that are merged with the same connected component of the tracking map (act 3708). In some embodiments, the method 3700 may include promoting the remaining connected components of the tracking map that are not merged with any reference map to the reference maps (act 3710). In some embodiments, the method 3700 may include merging the reference maps that are merged with the persistent pose and / or PCF of the tracking map and at least one connected component of the tracking map (act 3712). In some embodiments, the method 3700 may include completing the reference maps, e.g., by fusing map points and pruning redundant key rigs (act 3714).
[0396] Figures 38A and 38B illustrate an environmental map 3800 created by updating a reference map 700 that can be leveled up from a tracking map 700 (Figure 7) with a new tracking map, according to some embodiments. As illustrated and described with respect to Figure 7, the reference map 700 may provide a floor plan 706 of a reconstructed physical object in the corresponding physical world, represented by points 702. In some embodiments, the map points 702 may represent features of a physical object that may include a plurality of features. The new tracking map may be captured centered on the physical world, uploaded to the cloud, and merged with map 700. The new tracking map may include map points 3802 and key rigs 3804, 3806. In the illustrated example, key rig 3804 represents a key rig that is properly positioned relative to the reference map, for example, by establishing a correspondence with key rig 704 of map 700 (as illustrated in Figure 38B). On the other hand, key rig 3806 represents a key rig that is not positioned relative to map 700. Key rig 3806 may be leveled up to a separate reference map in some embodiments.
[0397] Figures 39A-39F are schematic diagrams illustrating an example of a cloud-based persistent coordinate system that provides a shared experience for users in the same physical space. Figure 39A shows, for example, that a reference map 4814 from the cloud is received by XR devices worn by users 4802A and 4802B of Figures 20A-20C. The reference map 4814 may have a reference coordinate frame 4806C. The reference map 4814 may have a PCF 4810C with a plurality of associated PPs (e.g., 4818A, 4818B in Figure 39C).
[0398] FIG. 39B shows that the XR device has established a relationship between its individual world coordinate systems 4806A, 4806B and the reference coordinate frame 4806C. This may be done, for example, by localizing the reference map 4814 on the individual device. Localizing the tracking map with respect to the reference map may result in a transformation between the local world coordinate system of each device and the coordinate system of the reference map.
[0399] FIG. 39C shows that as a result of the localization, a transformation can be calculated between the local PCF (e.g., PCF 4810A, 4810B) on the individual device and the individual persistent poses (e.g., PP 4818A, 4818B) on the reference map (e.g., transformations 4816A, 4816B). By using these transformations, each device can use its local PCF, which processes an image detected using sensors on the device, determines a location relative to the local device, and can be locally detected on the device by displaying virtual content associated with PP 4818A, 4818B or other persistent points on the reference map. Such an approach can accurately localize virtual content for each user and may enable each user to have the same experience of the virtual content within the physical space.
[0400] FIG. 39D shows a persistent pose snapshot from a reference map to a local tracking map. As can be seen from the figure, the local tracking maps are interconnected via the persistent pose. FIG. 39E shows that the PCF4810A on the device worn by user 4802A is accessible through PP4818A within the device worn by user 4802B. FIG. 39F shows that the tracking maps 4804A, 4804B and the reference 4814 can be merged. In some embodiments, some PCFs may be removed as a result of the merge. In the illustrated embodiment, the merged map includes the PCF4810C of the reference map 4814, but does not include the PCFs 4810A, 4810B of the tracking maps 4804A, 4804B. The PPs previously associated with the PCFs 4810A, 4810B may be associated with the PCF4810C after the map merge.
Example
[0401] FIGS. 40 and 41 illustrate examples of using a tracking map by the first XR device 12.1 of FIG. 9. FIG. 40 is a two-dimensional representation of a three-dimensional first local tracking map (Map 1) according to some embodiments, which can be generated by the first XR device of FIG. 9. FIG. 41 is a block diagram illustrating uploading Map 1 from the first XR device to the server of FIG. 9 according to some embodiments.
[0402] Figure 40 illustrates Map 1 and virtual content (Content 123 and Content 456) on the first XR device 12.1. Map 1 has an origin (Origin 1). Map 1 includes several PCFs (PCFa - PCFd). From the perspective of the first XR device 12.1, PCFa is located at the origin of Map 1, for example, and has X, Y, and Z coordinates of (0,0,0), and PCFb has X, Y, and Z coordinates of (-1,0,0). Content 123 is associated with PCFa. In this embodiment, Content 123 has X, Y, and Z relationships with respect to PCFa of (1,0,0). Content 456 has a relationship with respect to PCFb. In this embodiment, Content 456 has X, Y, and Z relationships of (1,0,0) with respect to PCFb.
[0403] In Figure 41, the first XR device 12.1 uploads Map 1 to the server 20. In this embodiment, since the server does not store a reference map for the same region of the physical world represented by the tracking map, the tracking map is stored as the initial reference map. Server 20 here has a reference map based on Map 1. The first XR device 12.1 has a reference map that is empty at this stage. Server 20, for the purposes of discussion, in some embodiments, does not include other maps other than Map 1. The map is not stored on the second XR device 12.2.
[0404] The first XR device 12.1 also transmits its Wi-Fi signature data to the server 20. Server 20 may use the Wi-Fi signature data to determine the approximate location of the first XR device 12.1 based on intelligence collected from other devices that were previously connected to server 20 or other servers, along with the GPS locations of such other devices that were recorded. The first XR device 12.1 may here end the first session (see Figure 8) and disconnect from the server 20.
[0405] FIG. 42 is a schematic diagram illustrating the XR system of FIG. 16 according to some embodiments, showing that after the first user 14.1 ended the first session, the second user 14.2 started a second session using the second XR device of the XR system. FIG. 43A is a block diagram showing the start of the second session by the second user 14.2. The first user 14.1 is shown by a phantom line because the first session by the first user 14.1 has ended. The second XR device 12.2 begins to record an object. Various systems with variable granularity may be used by the server 20 to determine that the second session by the second XR device 12.2 is within the same vicinity as the first session by the first XR device 12.1. For example, Wi-Fi signature data, global positioning system (GPS) positioning data, GPS data based on Wi-Fi signature data, or any other data indicating location may be included within the first and second XR devices 12.1 and 12.2 to record that location. Alternatively, the PCF identified by the second XR device 12.2 may show similarity to the PCF of Map 1.
[0406] As shown in FIG. 43B, the second XR device boots up and begins to collect data such as image 1110 from one or more cameras 44, 46. As shown in FIG. 14, in some embodiments, an XR device (e.g., the second XR device 12.2) may collect one or more images 1110, perform image processing, and extract one or more features / keypoints 1120. Each feature may be converted into a descriptor 1130. In some embodiments, the descriptor 1130 may be used to describe a keyframe 1140 that may have the position and orientation of the associated image. One or more keyframes 1140 may correspond to a single persistent pose 1150 that may be automatically generated after a threshold distance, e.g., 3 meters, from the previous persistent pose 1150. One or more persistent poses 1150 may correspond to a single PCF 1160 that may be automatically generated after a predetermined distance, e.g., every 5 meters. Over time, as the user continues to move around the user's environment and the XR device continues to collect more data such as image 1110, additional PCFs (e.g., PCF3 and PCF4, 5) may be created. One or more applications 1180 may be launched on the XR device and provide virtual content 1170 to the XR device for presentation to the user. The virtual content may have an associated content coordinate frame, which may be established with respect to one or more PCFs. As shown in FIG. 43B, the second XR device 12.2 creates three PCFs. In some embodiments, the second XR device 12.2 may attempt to localize with respect to one or more reference maps stored on server 20.
[0407] In some embodiments, as shown in FIG. 43C, the second XR device 12.2 may download the reference map 120 from the server 20. The map 1 on the second XR device 12.2 includes PCFs a-d and the origin 1. In some embodiments, the server 20 may have multiple reference maps for various locations, and the second XR device 12.2 determines that it is in the same vicinity as the first XR device 12.1 during a first session, and the server 20 may send the reference map regarding that vicinity to the second XR device 12.2.
[0408] FIG. 44 shows that the second XR device 12.2 starts to identify the PCFs for the purpose of generating map 2. The second XR device 12.2 identifies only a single PCF, namely, PCFs 1, 2. The X, Y, and Z coordinates of PCFs 1, 2 for the second XR device 12.2 can be (1, 1, 1). Map 2 has its own origin (origin 2), which may be based on the head pose of device 2 at the start of the device for the current head pose session. In some embodiments, the second XR device 12.2 may immediately attempt to localize map 2 with respect to the reference map. In some embodiments, map 2 may be impossible to localize with respect to the reference map (map 1) (i.e., the localization may fail) because the system does not recognize any or sufficient overlap between the two maps. Localization may be performed by identifying a part of the physical world represented in the first map that is also represented in the second map and calculating the transformation required to align those parts between the first map and the second map. In some embodiments, the system may perform localization based on a PCF comparison between the local map and the reference map. In some embodiments, the system may perform localization based on a persistent pose comparison between the local map and the reference map. In some embodiments, the system may perform localization based on a keyframe comparison between the local map and the reference map.
[0409] Figure 45 shows Map 2 after the second XR device 12.2 has identified further PCFs (PCF1, 2, PCF3, PCF4, 5) of Map 2. The second XR device 12.2 again attempts to localize Map 2 relative to the reference map. Since Map 2 has been extended to overlap at least a portion of the reference map, the localization attempt will succeed. In some embodiments, the overlap between the local tracking map, Map 2, and the reference map may be represented by PCFs, persistent poses, keyframes, or any other suitable intermediate or derived constructs.
[0410] Furthermore, the second XR device 12.2 associates content 123 and content 456 with PCF1, 2, and PCF3 of Map 2. Content 123 has X, Y, and Z coordinates relative to PCF1, 2 of (1, 0, 0). Similarly, the X, Y, and Z coordinates of content 456 relative to PCF3 within Map 2 are also (1, 0, 0).
[0411] Figures 46A and 46B illustrate the successful localization of Map 2 relative to the reference map. The localization may be based on matching features within one map and the other map. Here, using an appropriate transformation that involves both translation and rotation of one map relative to the other, the overlapping area / volume / section of Map 1410 represents the common portion with Map 1 and the reference map. Since Map 2 created PCF3 and 4, 5 prior to localization and the reference map created PCFa and c prior to the creation of Map 2, different PCFs were created to represent the same volume within real space (e.g., different maps).
[0412] As shown in FIG. 47, the second XR device 12.2 expands map 2 to include PCFa-d from the reference map. The inclusion of PCFa-d represents the localization of map 2 relative to the reference map. In some embodiments, the XR system may perform an optimization step to remove from 1410 the PCFs within, i.e., PCF3 and replicated PCFs such as PCF4, 5, etc., from the overlapping areas. After map 2 is localized, the placement of virtual content such as content 456 and content 123 will be relative to the nearest updated PCF within the updated map 2. The virtual content will appear to the user within the same real-world location despite the changed PCF associations for the content and despite the updated PCF for map 2.
[0413] As shown in FIG. 48, the second XR device 12.2 continues to expand map 2 as additional PCFs (PCFe, f, g, and h) are identified by the second XR device 12.2, e.g., as the user walks around the real world. Note also that map 1 is not expanded in FIGS. 47 and 48.
[0414] Referring to FIG. 49, the second XR device 12.2 uploads map 2 to the server 20. The server 20 stores map 2 together with the reference map. In some embodiments, map 2 may be uploaded to the server 20 when the session for the second XR device 12.2 ends.
[0415] The reference map within the server 20 includes here PCFi, which is not included within map 1 on the first XR device 12.1. The reference map on the server 20 can be expanded to include PCFi when a third XR device (not shown) uploads a map to the server 20 and such a map includes PCFi.
[0416] In FIG. 50, the server 20 merges Map 2 with the reference map to form a new reference map. The server 20 determines that PCFs a-d are common to the reference map and Map 2. The server expands the reference map to include PCFs e-h and PCF1, 2 from Map 2, forming a new reference map. The reference maps on the first and second XR devices 12.1 and 12.2 become obsolete based on Map 1.
[0417] In FIG. 51, the server 20 transmits the new reference map to the first and second XR devices 12.1 and 12.2. In some embodiments, this may occur when the first XR device 12.1 and the second device 12.2 attempt to localize during a different or new or subsequent session. The first and second XR devices 12.1 and 12.2 proceed to localize their respective local maps (Map 1 and Map 2 respectively) with respect to the new reference map as described above.
[0418] As shown in FIG. 52, the head coordinate frame 96 or “head pose” is related to the PCF within Map 2. In some embodiments, the origin of the map, i.e., Origin 2, is based on the head pose of the second XR device 12.2 at the start of the session. As PCFs are created during the session, the PCFs are established with respect to the world coordinate frame, i.e., Origin 2. The PCFs of Map 2 serve as a persistent coordinate frame with respect to the reference coordinate frame, and the world coordinate frame may be the world coordinate frame of the previous session (e.g., Origin 1 of Map 1 in FIG. 40). These coordinate frames are related by the same transformation used to localize Map 2 with respect to the reference map as discussed above in relation to FIG. 46B.
[0419] The transformation from the world coordinate frame to the head coordinate frame 96 has been described above with reference to FIG. 9. The head coordinate frame 96 shown in FIG. 52 is at a specific coordinate position relative to the PCF of Map 2 and at a specific angle relative to Map 2, and has only two orthogonal axes. However, it should be understood that the head coordinate frame 96 is within a certain three-dimensional location relative to the PCF of Map 2 and has three orthogonal axes in three-dimensional space.
[0420] In FIG. 53, the head coordinate frame 96 is moving relative to the PCF of Map 2. The head coordinate frame 96 is moving because the second user 14.2 has moved their head. The user can move their head in six degrees of freedom (6dof). The head coordinate frame 96 can therefore move in 6dof, that is, in three dimensions at its previous location in FIG. 52 and in approximately three orthogonal axes relative to the PCF of Map 2. The head coordinate frame 96 is adjusted whenever the real object detection camera 44 and the inertial measurement unit 48 in FIG. 9 respectively detect the movement of the real object and the head unit 22. Further information regarding head pose tracking is disclosed in and incorporated herein by reference in its entirety in U.S. Patent Application No. 16 / 221,065, entitled “Enhanced Pose Determination for Display Device”.
[0421] Figure 54 shows that a sound may be associated with one or more PCFs. The user may wear, for example, headphones or earphones with stereo sound. The location of the sound through the headphones can be simulated using conventional techniques. When the user rotates their head to the left, the location of the sound rotates to the right, and thus may be located in a stationary position so that the user perceives sound originating from the same location in the real world. In this example, the location of the sound is represented by sound 123 and sound 456. For the purpose of discussion, Figure 54 is similar to Figure 48 in its analysis. When first and second users 14.1 and 14.2 are located in the same room at the same or different times, they perceive sound 123 and sound 456 originating from the same location in the real world.
[0422] Figures 55 and 56 illustrate further implementations of the technology described above. First user 14.1 initiated a first session as described with reference to Figure 8. As shown in Figure 55, first user 14.1 ended the first session as indicated by the imaginary line. At the end of the first session, first XR device 12.1 uploaded map 1 to server 20. First user 14.1 then initiated a second session at a time after the first session. Since map 1 is already stored on first XR device 12.1, first XR device 12.1 does not download map 1 from server 20. If map 1 is lost, first XR device 12.1 downloads map 1 from server 20. First XR device 12.1 then proceeds to construct a PCF for map 2, localize it with respect to map 1, and further develop the reference map as described above. Map 2 of first XR device 12.1 is then used to associate local content, head coordinate frames, local sound, etc. as described above.
[0423] Referring to FIGS. 57 and 58, it is also possible to consider that more than one user may interact with the server in the same session. In this embodiment, a third user 14.3 with a third XR device 12.3 is added to the first user 14.1 and the second user 14.2. Each of the XR devices 12.1, 12.2, and 12.3 begins to generate its own map, namely, Map 1, Map 2, and Map 3, respectively. As the XR devices 12.1, 12.2, and 12.3 continue to develop Maps 1, 2, and 3, the maps are gradually uploaded to the server 20. The server 20 merges Maps 1, 2, and 3 to form a standard map. The standard map is then transmitted from the server 20 to each of the XR devices 12.1, 12.2, and 12.3.
[0424] FIG. 59 illustrates aspects of a visual recognition method for restoring and / or resetting a head pose according to some embodiments. In the illustrated embodiment, in act 1400, the visual recognition device is powered on. In act 1410, in response to being powered on, a new session is started. In some embodiments, the new session may include establishing a head pose. One or more capture devices on a head-mounted frame affixed to the user's head first capture an image of the environment and then capture the surface of the environment by determining the surface from the image. In some embodiments, the surface data may also be combined with data from a gravity sensor to establish a head pose. Other suitable methods for establishing a head pose may be used.
[0425] In act 1420, the processor of the visual recognition device enters a routine for tracking the head pose. The capture device continues to capture the surface of the environment and determine the orientation of the head-mounted frame relative to the surface as the user moves their head.
[0426] In act 1430, the processor determines whether the head pose has been lost. The head pose can be lost due to "edge" cases such as too many reflective surfaces, low light levels, empty walls, outdoors, etc., which can result in low feature acquisition, or due to dynamic cases such as moving clusters that form part of the map. The routine in 1430 allows a certain amount of time, for example, 10 seconds, to elapse to allow sufficient time to determine whether the head pose has been lost. If the head pose has not been lost, the processor returns to 1420 and enters the head pose tracking again.
[0427] If the head pose is lost in act 1430, the processor enters a routine in 1440 to restore the head pose. If the head pose is lost due to low light levels, a message such as the following message is displayed to the user through the display of the visual device.
[0428] The system is detecting low light conditions. Please move to an area with more light.
[0429] The system will continue to monitor whether sufficient light is available and whether the head pose can be restored. Alternatively, the system may determine that the low texture of the surface is causing the head pose to be lost, in which case the user is given the following prompt as a suggestion to improve the capture of the surface within the display.
[0430] The system cannot detect a sufficient surface with fine texture. The surface texture is not rough. Please move to an area where the texture is more refined.
[0431] In act 1450, the processor enters a routine to determine whether head pose restoration has failed. If head pose restoration has not failed (i.e., head pose restoration has succeeded), the processor returns to act 1420 by entering head pose tracking again. If head pose restoration has failed, the processor returns to act 1410 to establish a new session. As part of the new session, all cached data is invalidated, and thereafter, a new head pose is established. Any suitable method of head tracking may be used in combination with the process described in FIG. 59. U.S. Patent Application No. 16 / 221,065 describes head tracking and is hereby incorporated by reference in its entirety. Remote location determination
[0432] Various embodiments may utilize remote resources to facilitate a persistent and consistent cross-reality experience among individual users and / or groups of users. The inventors recognize and understand that the advantages of the operation of an XR device using a canonical map as described herein, as illustrated for example in FIG. 30, may be achieved without downloading a set of canonical maps. The advantages may be achieved, for example, by transmitting feature and pose information to a remote service that maintains a set of canonical maps. A device that requires virtual content to be positioned at locations defined relative to a canonical map may receive from the remote service one or more than one transformation between the features and the canonical map. Those transformations may be used to position virtual content at locations defined relative to the canonical map on the device, while maintaining information about the location of those features within the physical world, or alternatively, to identify locations within the physical world defined relative to the canonical map.
[0433] In some embodiments, spatial information is captured by an XR device and communicated to a remote service such as a cloud-based service, which uses the spatial information to localize the XR device relative to a reference map used by an application or other components of the XR system and to define the location of virtual content relative to the physical world. Once localized, a transformation that links the tracking map maintained by the device to the reference map can be communicated to the device. The transformation, in conjunction with the tracking map, can be used to determine the positions at which virtual content defined relative to the reference map should be rendered or, alternatively, to identify locations within the physical world defined relative to the reference map.
[0434] The inventors recognize that the data required to be exchanged between the device and the remote localization service can be very small compared to communicating map data such as would occur when the device communicates a tracking map to the remote service and receives a set of reference maps for device-based localization from that service. In some embodiments, implementing the localization function on cloud resources requires that only a small amount of information be transmitted from the device to the remote service. For example, it is not a requirement that a complete tracking map be communicated to the remote service for implementing localization. In some embodiments, as described above, features and pose information that can be stored in relation to a persistent pose can be transmitted to the remote server. In embodiments where features are represented by descriptors, as described above, the information uploaded can be even smaller.
[0435] The results returned from the location-specific service to the device may be one or more transformations that associate the uploaded features with a portion of the reference map for matching. Those transformations, along with the tracking map, may be used within the XR system to identify the location of virtual content or, alternatively, to identify a location within the physical world. As described above, in embodiments where persistent spatial information such as PCF is used to define a location relative to the reference map, the location-specific service may download to the device a transformation between the feature and one or more PCFs after successful location determination.
[0436] As a result, the network bandwidth consumed by communication between the XR device and the remote service for performing location determination can be reduced. The system can thus support frequent location determination and enable each device interacting with the system to quickly obtain information for positioning virtual content or performing other location-based functions. As the device moves within the physical environment, it may repeatedly request updated location determination information. Additionally, the device may frequently obtain updates to the location determination information, expand the map, or increase its accuracy, such as through the merging of additional tracking maps when the reference map changes.
[0437] Furthermore, uploading features and downloading conversions can improve privacy in XR systems by increasing the difficulty of obtaining maps through spoofing, thereby sharing map information among multiple users. Unauthorized users can be blocked from obtaining maps from the system, for example, by sending false requests regarding a canonical map that represents a part of the physical world where the unauthorized user is not located. An unauthorized user is less likely to have access to features within a region of the physical world for which they are requesting map information if they are not physically present within that region. In embodiments where the feature information is formatted as a feature description, the difficulty of spoofing the feature information within a request regarding map information will be high. Further, when the system returns a conversion intended to be applied to a tracking map of a device operating within a region for which location information is requested, the information returned by the system is likely to be of little or no use to a fraudster.
[0438] According to one embodiment, the location service is implemented as a cloud-based microservice. In some examples, implementing a cloud-based location service can help conserve device computing resources and enable the computations required for location to be performed with very low latency. Those operations are supported by an almost infinite computing power or other computing resources made available by providing additional cloud resources, ensuring the scalability of the XR system to support a large number of devices. In one example, many canonical maps are maintained in memory for near-instantaneous access or alternatively stored in high-availability devices to reduce system latency.
[0439] Furthermore, performing location identification for multiple devices within a cloud service can enable refinement of the process. Location identification telemetry and statistics can provide information regarding which reference maps are within active memory and / or high-availability storage. Statistics regarding multiple devices may be used, for example, to identify the most frequently accessed reference maps.
[0440] Additional accuracy can also be achieved as a result of processing within a cloud environment or other remote environment using substantial processing resources for the remote device. For example, location identification can be performed on a higher density reference map within the cloud as opposed to processing performed on the local device. The map may be stored within the cloud, for example, with more PCFs or a higher density of descriptors per PCF, to increase the accuracy of matching between the set of features from the device and the reference map.
[0441] FIG. 61 is a schematic diagram of an XR system 6100. During a user session, the user device that displays cross-reality content can appear in various forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As discussed above, these devices are configured with software such as an application or other components and / or are wired and can generate local location information (e.g., a tracking map) for use in rendering virtual content on their respective displays.
[0442] The virtual content positioning information may be defined relative to global location information, which may be formatted, for example, as a reference map containing one or more PCFs. According to some embodiments, such as the embodiment shown in FIG. 61, the system 6100 is configured with a cloud-based service that supports the functionality and display of virtual content on the user device.
[0443] In one embodiment, the location function is provided as a cloud-based service 6106, which may be a microservice. The cloud-based service 6106 may be implemented on any of a plurality of computing devices, from which computing resources may be allocated to one or more services running within the cloud. Those computing devices may be interconnected to each other and to devices such as the wearable XR device 6102 and the handheld device 6104 in an accessible manner. Such connections may be provided via one or more networks.
[0444] In some embodiments, the cloud-based service 6106 is configured to receive descriptor information from individual user devices and "locate" the devices against a reference map or maps of matching criteria. For example, the cloud-based location service matches the received descriptor information to descriptor information regarding an individual reference map. The reference map may be created using techniques such as those described above of obtaining information about the physical world, merging maps provided by one or more devices having an image sensor or other sensor, creating a reference map. However, it is not a requirement that the reference maps be created by the devices accessing them, and thus the maps may be created by map developers, for example, by making the maps available to the location service 6106.
[0445] According to some embodiments, the cloud service may include operations for handling canonical map identification and filtering a repository of canonical maps against a set of potential matches. The filtering may be performed using the filtering criteria shown in FIG. 29, or instead of or in addition to the filtering criteria shown in FIG. 29, using any subset of the filtering criteria and other filtering criteria. In one embodiment, geographic data can be used to limit a search for matching canonical maps to maps representing an area proximate to a device that requires location identification. For example, area attributes such as Wi-Fi signal data, Wi-Fi fingerprint information, GPS data, and / or other device location information are used as rough filters stored on the canonical map, thereby enabling the analysis of descriptors to be limited to canonical maps that are known or likely to be in proximity to the user device. Similarly, the location history of each device may be maintained by the cloud service such that canonical maps within the vicinity of the device's most recent location are preferentially searched. In some examples, the filtering can include the functions discussed above with respect to FIGS. 31B, 32, 33, and 34.
[0446] FIG. 62 is an exemplary process flow that can be performed by a device to receive transformation information that specifies the position of the device using a reference map and defines one or more transformations between the device local coordinate system and the coordinate system of the reference map using a cloud-based service. Various embodiments and examples may describe one or more transformations as defining a transformation from a first coordinate frame to a second coordinate frame. Other embodiments may include a transformation from the second coordinate frame to the first coordinate frame. In still other embodiments, the transformation enables a transition from one coordinate frame to another, and the resulting coordinate frame depends only on the desired coordinate frame output (including, for example, the coordinate frame in which the content should be displayed). In yet further embodiments, the coordinate system transformation may enable the determination of the first coordinate frame from the second coordinate frame and the second coordinate frame from the first coordinate frame.
[0447] According to some embodiments, information reflecting the transformation for each persistent pose defined with respect to the reference map can be communicated to the device.
[0448] According to one embodiment, process 6200 can start at 6202 using a new session. Starting a new session on the device can start the capture of image information and construct a tracking map for the device. Additionally, the device may send a message, register with the server of the location service, and prompt the server to create a session for that device.
[0449] In some embodiments, starting a new session on a device may optionally include sending calibration data from the device to a location service. The location service returns one or more transformations calculated based on the features and associated set of poses for the device. If the pose of a feature is adjusted based on device-specific information before the calculation of the transformation and / or the transformation is adjusted based on device-specific information after the calculation of the transformation, rather than performing those calculations on the device, the device-specific information may be sent to the location service so that the location service can apply the adjustment. As a specific example, sending device-specific adjustment information may include capturing calibration data regarding sensors and / or displays. The calibration data may be used, for example, to adjust the location of feature points relative to the measured location. Alternatively, or in addition, the calibration data may be used to adjust the location where the display is commanded to render virtual content so that it appears accurately positioned for that particular device. The calibration data may be derived, for example, from multiple images of the same scene obtained using sensors on the device. The locations of the features detected within those images may be represented as a function of the sensor location such that the multiple images result in a set of equations that can be solved with respect to the sensor location. The calculated sensor location may be compared to a nominal location and the calibration data may be derived from any differences. In some embodiments, the unique information about the structure of the device may also, in some embodiments, enable the calibration data to be calculated for the display.
[0450] In embodiments where calibration data is generated for a sensor and / or display, the calibration data may be applied at any point in the measurement or display process. In some embodiments, the calibration data may be sent to a location server, which may store the calibration data within a data structure established for each device that is registered with the location server and is thus in session with the server. The location server may apply the calibration data to any transformation calculated as part of a location process for the device that supplied the calibration data. The computational burden of using the calibration data for higher accuracy of sensed and / or displayed information is thus borne by the calibration service, providing an additional mechanism for reducing the processing burden on the device.
[0451] Once a new session is established, process 6200 may continue to 6204, involving the capture of a new frame of the device's environment. Each frame can be processed at 6206 to generate descriptors for the captured frame (e.g., including the DSF values discussed above). These values may be calculated using some or all of the techniques described above, including techniques such as those discussed with respect to FIGS. 14, 22, and 23. As discussed, the descriptors may be calculated as a mapping of feature points to the descriptor, or in some embodiments, as a mapping of image patches around the feature points. The descriptors may have values that enable efficient matching between the newly acquired frame / image and the stored map. Further, the number of features extracted from an image may be limited to a maximum number of feature points per image, such as 200 feature points per image. Feature points may be selected to represent points of interest, as described above. Thus, acts 6204 and 6206 may be performed as part of a device process that forms a tracking map or otherwise periodically collects images of the physical world around the device, or may be performed separately for location determination, but need not be.
[0452] Feature extraction in 6206 may include adding pose information to the features extracted in 6206. The pose information may be the pose within the local coordinate system of the device. In some embodiments, the pose may be relative to a reference point in the tracking map, such as a persistent pose, as discussed above. Alternatively, or in addition, the pose may be relative to the origin of the device's tracking map. Such embodiments may enable the location service to provide location services for a wide range of devices even when they do not utilize a persistent pose, as described herein. Nevertheless, the pose information may be added to each feature or each set of features such that the location service can use the pose information for calculating a transformation that can be returned to the device in response to matching the features to the features in the stored map.
[0453] Process 6200 may continue to decision block 6207, where a decision is made as to whether to request location determination. One or more criteria may be applied to determine whether to request location determination. The criteria may include the passage of time such that the device may request location determination after a certain threshold amount of time. For example, if location determination has not been attempted within a certain threshold amount of time, the process may continue from decision block 6207 to action 6208, where location determination is requested from the cloud. The threshold amount of time may be, for example, between 10 and 30 seconds, such as 25 seconds. Alternatively, or in addition, location determination may be triggered by the movement of the device. The device executing process 6200 may use an IMU and / or its tracking map to track its movement and initiate location determination in response to detecting movement that exceeds a threshold distance from the location where the device last requested location determination. The threshold distance may be, for example, between 1 and 10 meters, such as 3 to 5 meters. Still further alternatively, location determination may be triggered in response to an event, such as when the device creates a new persistent pose or when the current persistent pose of the device changes, as described above.
[0454] In some embodiments, determination block 6207 may be implemented such that a threshold for triggering localization can be dynamically established. For example, in an environment that is primarily uniform, where there may be low confidence in matching the set of extracted features to the features of a stored map, localization may be required more frequently, increasing the chance that at least one attempt at localization will succeed. In such a scenario, the threshold applied in determination block 6207 may be decreased. Similarly, in an environment where there are relatively few features, the threshold applied in determination block 6207 may be decreased to increase the frequency of localization attempts.
[0455] Regardless of how the location determination is triggered, once triggered, process 6200 may proceed to act 6208, where the device sends a request for the location service that includes data used by the location service to perform the location determination. In some embodiments, data from multiple image frames may be provided for the location determination attempt. The location service may not consider the location determination successful, for example, unless the features in the multiple image frames yield consistent location determination results. In some embodiments, process 6200 may include storing the feature descriptors and the added pose information in a buffer. The buffer may be, for example, a circular buffer that stores a set of features extracted from the most recently captured frame. Thus, the location determination request may be sent along with some sets of features accumulated in the buffer. In some settings, the buffer size is implemented to accumulate some sets of data that are more likely to result in a successful location determination. In some embodiments, the buffer size may be set to accumulate features from, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10 frames). Optionally, the buffer size can have a baseline setting, which can be increased in response to a location determination failure. In some examples, the increasing buffer size and the corresponding number of sets of features transmitted reduce the likelihood that the subsequent location determination function fails to return a result.
[0456] Regardless of how the buffer size is set, the device may transfer the contents of the buffer to the location service as part of the location determination request. Other information may also be transmitted in conjunction with the feature points and the added pose information. For example, in some embodiments, geographic information may be transmitted. The geographic information may include, for example, GPS coordinates or a wireless signature associated with a device tracking map or the current persistent pose.
[0457] In response to a request sent at 6208, the cloud location service may analyze the descriptor and locate the device within a canonical map or other persistent map maintained by the service. For example, the descriptor may be matched to a set of features within the map against which the device is to be located. The cloud-based location service may perform the location as described above with respect to device-based location (e.g., may rely on any of the functions discussed above for location (map ranking, map filtering, place estimation, filtered map selection, as in FIGS. 44-46, and / or the embodiments discussed with respect to the location module, including PCF and / or PP identification and matching, etc.)). However, instead of communicating the identified canonical map to the device (e.g., in device location), the cloud-based location service may proceed to generate a transformation based on the relative orientation of the feature set transmitted from the device and the features that match the canonical map. The location service may return these transformations to the device, which may be received at block 6210.
[0458] In some embodiments, the reference map maintained by the location service may employ PCF, as described above. In such embodiments, the feature points of the reference map that match the feature points transmitted from the device may have positions defined for one or more PCFs. Thus, the location service may identify one or more reference maps and calculate the transformation between the coordinate frame represented within the pose transmitted with the location request and one or more PCFs. In some embodiments, the identification of one or more reference maps is assisted by filtering potential maps based on geographical data regarding the individual device. For example, once filtered into a candidate set (e.g., among other options, by GPS coordinates), the candidate set of reference maps can be analyzed in detail, as described above, to determine the matching feature points or PCFs.
[0459] The data returned to the requesting device in Act 6210 may be formatted as a table of persistent pose transformations. The table may be accompanied by one or more reference map identifiers and may indicate the reference map with respect to which the device has been located by the location service. However, it should be understood that the location information may be formatted in other ways, including as a list of transformations accompanied by the associated PCF and / or reference map identifier.
[0460] Regardless of how the transformations are formatted, in Act 6212, the device may use these transformations to calculate the location where virtual content should be rendered with respect to one of the PCFs as defined by the application or other components of the XR system for that location. This information may alternatively or additionally be used on the device to perform any location-based operations where the location is defined based on the PCF.
[0461] In some scenarios, the location service may not be able to match the features sent from the device to any stored reference map, or may not be able to match a sufficient number of sets of features communicated with the request for the location service to consider that a location determination success has occurred. In such scenarios, rather than returning the transformation to the device as described above in connection with act 6210, the location service may indicate to the device that the location determination has failed. In such scenarios, process 6200 may branch at decision block 6209 to act 6230, and the device may take one or more actions for failure handling. These actions may include increasing the size of a buffer that holds the set of features transmitted for location determination. For example, if the location service does not consider a location determination successful unless three sets of features match, the buffer size may be increased from five to six, increasing the likelihood that three of the sets of features transmitted may match the reference map maintained by the location service.
[0462] Alternatively, or in addition, failure handling may include adjusting the operating parameters of the device to trigger more frequent location determination attempts. The threshold time and / or threshold distance between location determination attempts may be decreased, for example. As another example, the number of feature points within each set of features may be increased. The matching between the set of features and the features stored in the reference map may be considered to occur when a sufficient number of features within the set transmitted from the device match the features of the map. Increasing the number of features transmitted may increase the chance of a match. As a specific example, the initial feature set size may be 50, which may be increased to 100, 150, and then 200 in response to each successive location determination failure. In response to a successful match, the set size may then be returned to its initial value.
[0463] The failure handling may also include obtaining location information from sources other than the location service. According to some embodiments, the user device can be configured to cache the reference map. The cached map enables the device to access and display the content when the cloud is unavailable. For example, the cached reference map enables device-based location determination in case of communication failure or other unavailability.
[0464] According to various embodiments, FIG. 62 illustrates a high-level flow for a device to initiate cloud-based location determination. In other embodiments, one or more of the various illustrated steps can be combined, omitted, or call other processes to perform location determination and ultimately visualize virtual content within the view of an individual device.
[0465] Furthermore, process 6200 shows that the device determines whether to initiate location determination at decision block 6207, but it should be understood that the trigger to initiate location determination may originate outside the device, including from the location service. The location service may maintain information, for example, for each device with which it is in session. That information may include, for example, an identifier of the reference map for which each device was most recently located. The location service or other components of the XR system may update the reference map, including using techniques as described above in connection with FIG. 26. When the reference map is updated, the location service may send a notification to each device that was most recently located with respect to that map. That notification may serve as a trigger for the device to request location determination and / or may include updated transforms recomputed using the most recently transmitted set of features from the device.
[0466] Figures 63A, B, and C are exemplary process flows showing the operations and communications between a device and a cloud service. What is shown in blocks 6350, 6352, 6354, and 6456 are exemplary architectures and s...
Claims
1. A method of operating a cross-reality system in which an environment map is stored in a database, the method comprising: determining whether to merge the first environment map and the second environment map; Deciding whether to merge is a determining whether a transformation applied to the second environment map to align the second environment map with the first environment map results in alignment with respect to a gravity direction; merging the first environment map and the second environment map based at least in part on alignment with the gravity direction; A method comprising:
2. The method described in claim 1, wherein the transformation applied to the second environment map to align the second environment map with the first environment map is based at least in part on a matching set of features.
3. The method described in claim 1, wherein at least one of the first environmental map and the second environmental map is accessed from the database.
4. The method of claim 1, wherein at least one of the first environmental map and the second environmental map is constructed from information collected by at least one user device.
5. The method of claim 1, further comprising storing in the database a reference map resulting from merging the first environmental map and the second environmental map.
6. The method described in claim 1, wherein at least one of the first environment map and the second environment map is a reference map resulting from a previous merging process.
7. The method of claim 1, wherein the first environment map and the second environment map represent overlapping portions of the physical world.
8. The method described in claim 1, wherein the first environment map and the second environment map represent non-overlapping portions of the physical world.
9. The method of claim 1, further comprising, based on a determination that the transformation applied to the second environment map to align the second environment map with the first environment map does not result in alignment with the gravity direction, identifying a further transformation and repeating the act of determining whether the transformation results in alignment with the gravity direction.
10. The method of claim 1, wherein determining whether the transformation results in alignment with the gravity direction includes comparing at least one error metric to a threshold value.
11. A computing device configured for use in a cross-reality system, the computing device comprising: at least one processor; a computer-readable medium coupled to the processor; a plurality of environment maps stored within the computer-readable medium; Computer executable instructions and Equipped with The computer-executable instructions, when executed by the at least one processor, determining whether to merge the first environment map and the second environment map; Deciding whether to merge is a determining whether a transformation applied to the second environment map to align the second environment map with the first environment map results in alignment with respect to a gravity direction; merging the first environment map and the second environment map based at least in part on alignment with the gravity direction; a computing device,
12. The computing device of claim 11, wherein the method further includes storing a reference map resulting from merging the first environmental map and the second environmental map within the computer-readable medium.
13. The computing device of claim 11, wherein at least one of the first environmental map and the second environmental map is constructed from information collected by at least one user device.
14. The computing device of claim 11, wherein at least one of the first environment map and the second environment map is a reference map resulting from a previous merging process.
15. The computing device of claim 11, further comprising, based on a determination that the transformation applied to the second environment map to align the second environment map with the first environment map does not result in alignment with the gravity direction, identifying a further transformation and repeating the act of determining whether the transformation results in alignment with the gravity direction.
16. A cloud computing environment for an augmented reality system configured for communication with a plurality of user devices equipped with sensors, comprising: a database storing a plurality of environment maps constructed from data provided by the plurality of user devices; a non-transitory computer storage medium storing computer-executable instructions; Equipped with The computer-executable instructions, when executed by at least one processor in the cloud computing environment, Implementing a method that includes determining whether to merge a first environment map and a second environment map; Deciding whether to merge is a determining whether a transformation applied to the second environment map to align the second environment map with the first environment map results in alignment with respect to a gravity direction; merging the first environment map and the second environment map based at least in part on alignment with the gravity direction; cloud computing environments, including
17. The cloud computing environment of claim 16, wherein the method further includes storing in the database a reference map resulting from merging the first environmental map and the second environmental map.
18. The cloud computing environment of claim 16, wherein the plurality of environmental maps represent portions of the physical world.
19. The cloud computing environment of claim 16, wherein at least one of the first environment map and the second environment map is a reference map resulting from a previous merging process.
20. The cloud computing environment of claim 16, further comprising, based on a determination that the transformation applied to the second environment map to align the second environment map with the first environment map does not result in alignment with the gravity direction, identifying a further transformation and repeating the act of determining whether the transformation results in alignment with the gravity direction.
21. A method of operating a cross-reality system, in which one or more environmental maps are stored in a database and a representation of the physical environment is calculated based on sensor data collected by a device worn by a user, the method comprising: receiving a representation of a physical environment from the device, the representation of the physical environment being aligned with respect to a direction of gravity; determining a transformation between the representation of the physical environment and an environment map; determining whether to modify the representation of the physical environment using the environment map, wherein determining whether to modify includes determining whether applying the transformation to the representation of the physical environment produces a transformed representation of the physical environment that is consistent with respect to the gravity direction; and modifying the environment map based on the representation of the physical environment based at least in part on determining that the transformed representation of the physical environment is aligned with respect to the direction of gravity; A method comprising:
22. Determining whether applying the transformation to the representation of the physical environment produces the transformed representation of the physical environment that is consistent with respect to the direction of gravity, comprising: determining whether applying the transformation to the representation of the physical environment produces the transformed representation of the physical environment that is rotated with respect to the direction of gravity by more than a threshold amount; selecting the transformation as the determined transformation in response to determining that applying the transformation to the representation of the physical environment does not produce the transformed representation of the physical environment that is rotated with respect to the direction of gravity by more than the threshold amount; and discarding the transformation in response to determining that applying the transformation to the representation of the physical environment produces the transformed representation of the physical environment that is rotated with respect to the direction of gravity by more than the threshold amount; and 22. The method of claim 21, comprising:
23. Identifying from said database a set of environment maps to be modified using said representation of said physical environment; For each environment map in the set of environment maps: determining the transformation between the representation of the physical environment and the environment map; determining whether to modify the environment map using the representation of the physical environment; and modifying the environment map with the representation of the physical environment based on determining that the transformed representation of the physical environment is consistent with the direction of gravity; To do 22. The method of claim 21 further comprising:
24. For each environment map in the set of environment maps:
24. The method of claim 23, further comprising: not modifying the environment map using the representation of the physical environment based on determining that a direction of gravity in the environment map does not match the direction of gravity in the transformed representation of the physical environment.
25. The method of claim 24, wherein identifying the set of environment maps comprises: determining an area identifier associated with the representation of the physical environment; identifying the set of environment maps from the database based at least in part on the area identifier associated with the representation of the physical environment; 24. The method of claim 23, comprising:
26. The method described in claim 23, wherein identifying the set of environmental maps from the database further comprises filtering the set of environmental maps based on the similarity of one or more metrics within the set of environmental maps associated with the representation of the physical environment.
27. A computing device configured for use in a cross-reality system in which a portable device operating in a three-dimensional (3D) environment renders virtual content, the computing device comprising: at least one processor; a computer-readable medium coupled to the processor; a plurality of environment maps stored within the computer-readable medium; Computer executable instructions and Equipped with The computer-executable instructions, when executed by the at least one processor, receiving a representation of a physical environment from the portable device, the representation of the physical environment being aligned with respect to a direction of gravity; determining whether to merge an environment map with the representation of the physical environment, wherein determining whether to merge includes retrieving a transformation of the representation of the physical environment, wherein the transformation of the representation of the physical environment aligns the transformed representation of the physical environment and the environment map in a manner that preserves alignment of the transformed representation of the physical environment with the gravity direction; merging the environment map with the transformed representation of the physical environment based on determining that a direction of gravity in the environment map matches the direction of gravity in the transformed representation of the physical environment; 10. A computing device configured to perform a method comprising:
28. The computing device of claim 27, wherein searching for a transformation includes searching for the transformation that aligns a first set of features associated with the representation of the physical environment with a second set of features associated with the environmental map with an error metric below a threshold.
29. The computing device of claim 28, wherein searching for a transformation includes searching for a transformation that does not change the orientation of the transformed representation of the physical environment relative to the direction of gravity.
30. Searching for a transformation comprises: determining whether applying the transformation to the representation of the physical environment produces the transformed representation of the physical environment that is rotated with respect to the direction of gravity by more than a threshold amount; applying the transformation to the representation of the physical environment to generate the transformed representation of the physical environment in response to determining that applying the transformation to the representation of the physical environment does not produce the transformed representation of the physical environment that is rotated with respect to the direction of gravity by more than the threshold amount; and discarding the transformation in response to determining that applying the transformation to the representation of the physical environment produces the transformed representation of the physical environment that is rotated with respect to the direction of gravity by more than the threshold amount; and 28. The computing device of claim 27, comprising:
31. The method comprising: identifying a set of environment maps from the plurality of environment maps to be merged with the representation of the physical environment; For each environment map in the set of environment maps: determining whether to merge the environment map with the representation of the physical environment; and merging the environment map with the transformed representation of the physical environment based on determining that a direction of gravity in the environment map matches the direction of gravity in the transformed representation of the physical environment. To do 28. The computing device of claim 27, further comprising:
32. The computing device of claim 31, wherein the method further includes, for each environment map in the set of environment maps, not merging the environment map with the transformed representation of the physical environment based on determining that the direction of gravity of the environment map does not align with the direction of gravity of the representation of the physical environment.
33. The method of identifying the set of environment maps, comprising: determining an area identifier associated with the representation of the physical environment; identifying the set of environment maps based, at least in part, on the area identifier associated with the representation of the physical environment; 33. The computing device of claim 32, further comprising:
34. The computing device of claim 33, wherein identifying the set of environmental maps further comprises filtering the set of environmental maps based on similarity of the representation of the physical environment and one or more metrics associated with the environmental maps in the set of environmental maps.
35. A cloud computing environment for an augmented reality system configured for communication with a plurality of user devices equipped with sensors, comprising: a map database storing a plurality of environment maps constructed from data provided by the plurality of user devices; a non-transitory computer storage medium storing computer-executable instructions; Equipped with The computer-executable instructions, when executed by at least one processor in the cloud computing environment, receiving a representation of a physical environment from a user device, the representation of the physical environment being aligned with respect to a direction of gravity; updating the map database based on the received representation of the physical environment; and carrying out a method comprising: Updating the map database includes, with respect to an environment map in a set of environment maps: determining a transformation between the representation of the physical environment and the environment map; determining whether to merge the environment map with the representation of the physical environment, wherein determining whether to merge includes determining whether applying the transformation to the representation of the physical environment produces a transformed representation of the physical environment that is consistent with respect to the gravity direction; modifying the environment map with the representation of the physical environment based on determining that the transformed representation of the physical environment is consistent with the direction of gravity; and cloud computing environments, including
36. The cloud computing environment of claim 35, wherein determining the transformation includes selecting as the determined transformation the transformation that aligns the corresponding features in the representation of the physical environment and the environmental map with an error metric that is below a threshold.
37. The cloud computing environment of claim 36, wherein the method further includes determining the corresponding features based on similarities of identifiers assigned to the corresponding features.
38. The cloud computing environment of claim 36, wherein selecting as the determined transformation further comprises selecting the transformation that, when applied, produces the transformed representation of the physical environment that is aligned with respect to the direction of gravity.
39. The cloud computing environment of claim 35, wherein determining the transformation includes applying a plurality of candidate transformations to the representation of the physical environment and selecting a candidate transformation from the plurality of candidate transformations as the determined transformation.
40. The cloud computing environment of claim 39, wherein selecting the determined transformation further comprises selecting as the determined transformation the candidate transformation that, when applied, produces the transformed representation of the physical environment that is consistent with respect to the gravity direction.
41. A method of operating a cross reality system in which an environment map is stored in a database, said method comprising: determining whether to merge the first map and the second map; the first map and the second map are aligned with respect to a gravity direction; The determining step comprises: applying a plurality of transformations to at least a portion of the first map, each of the plurality of applied transformations preserving an orientation of the first map with respect to gravity; calculating, for each applied transformation of the plurality of transformations, an error in alignment between the first map and the second map; selecting an applied transform from the plurality of transforms based on the applied transform having a low error relative to other applied transforms from the plurality of transforms, whereby the determining is based on the error of the selected applied transform; A method comprising:
42. The method of claim 41, wherein the transformation applied to at least the portion of the first map to align the first map with the second map is based at least in part on a matching set of features.
43. The method of claim 41, wherein at least one of the first map and the second map is accessed from the database.
44. The method of claim 41, wherein at least one of the first map and the second map is constructed from information collected by at least one user device.
45. The method of claim 41, further comprising storing in the database a reference map resulting from merging the first map and the second map.
46. The method of claim 41, wherein at least one of the first map and the second map is a reference map resulting from a previous merging process.
47. The method of claim 41, wherein the first map and the second map represent overlapping portions of the physical world.
48. The method of claim 41, wherein the first map and the second map represent non-overlapping portions of the physical world.
49. The method of claim 41, further comprising, based on a determination that the applied transformation does not have the low error relative to other applied transformations among the plurality of transformations, iterating to identify further transformations and determine whether the transformations result in alignment with the gravity direction.
50. The method described in claim 41, wherein determining whether the transformation results in alignment with the gravity direction includes comparing the error to a threshold value.
51. A computing device configured for use in a cross-reality system, the computing device comprising: at least one processor; a computer-readable medium coupled to the processor; a plurality of environment maps stored within the computer-readable medium; Computer executable instructions and Equipped with The computer-executable instructions, when executed by the at least one processor, determining whether to merge a first map and a second map, the first map and the second map being aligned with respect to a gravity direction; The determining step comprises: applying a plurality of transformations to at least a portion of the first map, each of the plurality of applied transformations preserving an orientation of the first map with respect to gravity; calculating, for each applied transformation of the plurality of transformations, an error in alignment between the first map and the second map; selecting an applied transform from the plurality of transforms based on the applied transform having a low error relative to other applied transforms from the plurality of transforms, whereby the determining is based on the error of the selected applied transform; a computing device, 52. The computing device of claim 51, wherein the method further comprises storing in the computer-readable medium a reference map resulting from merging the first map and the second map.
53. The computing device of claim 51, wherein at least one of the first map and the second map is constructed from information collected by at least one user device.
54. The computing device of claim 51, further comprising, based on a determination that the applied transformation does not have the low error relative to other applied transformations among the plurality of transformations, iterating to identify further transformations and determine whether the transformations result in alignment with the gravity direction.
55. The computing device of claim 51, wherein determining whether the transformation results in alignment with the gravity direction includes comparing the error to a threshold.
56. A cloud computing environment for an augmented reality system configured for communication with a plurality of user devices equipped with sensors, comprising: a database storing a plurality of environment maps constructed from data provided by the plurality of user devices; a non-transitory computer storage medium storing computer-executable instructions; Equipped with The computer-executable instructions, when executed by at least one processor in the cloud computing environment, Implementing a method including determining whether to merge a first map and a second map, the first map and the second map being aligned with respect to a gravity direction; The determining step comprises: applying a plurality of transformations to at least a portion of the first map, each of the plurality of applied transformations preserving an orientation of the first map with respect to gravity; calculating, for each applied transformation of the plurality of transformations, an error in alignment between the first map and the second map; selecting an applied transform from the plurality of transforms based on the applied transform having a low error relative to other applied transforms from the plurality of transforms, whereby the determining is based on the error of the selected applied transform; cloud computing environments, including
57. The cloud computing environment of claim 56, wherein the method further includes storing in the database a reference map resulting from merging the first environmental map and the second environmental map.
58. The cloud computing environment described in claim 56, wherein at least one of the first map and the second map is constructed from information collected by at least one user device.
59. The cloud computing environment of claim 56, further comprising, based on a determination that the applied transformation does not have the low error relative to other applied transformations of the plurality of transformations, iterating to identify further transformations and determine whether the transformations result in alignment with the gravity direction.
60. The cloud computing environment described in claim 56, wherein determining whether the transformation results in alignment with the gravity direction includes comparing the error to a threshold value.