Cross-reality system with location services

By using sensors to capture 3D environment images and aligning them with a shared coordinate frame, the system addresses the challenge of virtual content mapping in cross-reality systems, ensuring precise and immersive experiences.

JP7825024B2Active Publication Date: 2026-03-05MAGIC LEAP INC
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing cross-reality systems face challenges in accurately rendering virtual content within three-dimensional environments due to the lack of efficient methods for mapping and aligning virtual objects with real-world coordinates, leading to inconsistencies and reduced user experience.

Method used

The system employs sensors to capture 3D environment images, generates a local coordinate frame, extracts features, and transmits this information to a location service for transformation into a shared coordinate frame, allowing precise rendering of virtual content based on these transformations.

Benefits of technology

This approach enables accurate alignment of virtual content with real-world coordinates, enhancing user experience by providing consistent and immersive cross-reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007825024000001
    Figure 0007825024000001
  • Figure 0007825024000002
    Figure 0007825024000002
  • Figure 0007825024000003
    Figure 0007825024000003
Patent Text Reader

Abstract

To provide a cross reality system with localization service.SOLUTION: A cross reality system enables any of multiple devices to efficiently and accurately access previously stored maps and render a virtual content specified in relation to those maps. The cross reality system may include a cloud-based localization service that responds to requests from devices to localize with respect to a stored map. The request may include one or more sets of feature descriptors extracted from an image of a physical world around the device. Those features may be posed relative to a coordinate frame used by a local device. The localization service may identify one or more stored maps with a matching set of features.SELECTED DRAWING: Figure 61
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 915,599, filed October 15, 2019, entitled "CROSS REALITY SYSTEM WITH LOCALIZATION SERVICE," which is incorporated herein by reference in its entirety under 35 U.S.C. §119(e).

[0002] This application relates generally to cross-reality systems. [Background technology]

[0003] A computer may control a human user interface and create cross-reality (XR) environments in which part or all of the XR environment is generated by the computer as perceived by the user. These XR environments may be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments in which part or all of the XR environment may be generated by the computer using data that describes the environment. This data may, for example, describe virtual objects that may be rendered so that a user can sense or perceive and interact with the virtual objects as part of the physical world. The user may experience these virtual objects as a result of data being rendered and presented through a user interface device, such as a head-mounted display device. The data may be displayed to the user so that it can be seen or played audibly to the user, or may control audio that can be displayed to the user so that it can be heard, or may control a tactile (or haptic) interface, allowing the user to experience touch sensations that the user senses or perceives as they feel the virtual objects.

[0004] XR systems can be useful for many applications, ranging from scientific visualization, medical training, engineering design and prototyping, remote manipulation and telepresence, and personal entertainment. AR and MR, in contrast to VR, involve one or more objects in association with real objects in the physical world. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment when using XR systems and opens up possibilities for a variety of applications that present realistic and easily understandable information about how the physical world can be altered.

[0005] To realistically render virtual content, an XR system may build a representation of the physical world around a user of the system. This representation may be built, for example, by processed images obtained using sensors on a wearable device that forms part of the XR system. In such a system, a user may perform an initialization routine by looking around a room or other physical environment in which the user intends to use the XR system until the system has acquired enough information to build a representation of that environment. As the system operates and the user moves around the environment or into other environments, sensors on the wearable device may acquire additional information and expand or update the representation of the physical world. Summary of the Invention [Means for solving the problem]

[0006] Aspects of the present application relate to methods and apparatus for providing X-reality (cross-reality or XR) scenes. The techniques described herein may be used together, separately, or in any suitable combination.

[0007] According to one aspect, an electronic device is provided that is configured to operate within a cross-reality system. The electronic device comprises: one or more sensors configured to capture information about a three-dimensional (3D) environment, the captured information comprising a plurality of images of the 3D environment; at least one processor configured to execute computer-executable instructions; and the computer-executable instructions comprising instructions for generating a local coordinate frame to represent a position within the 3D environment, extracting a plurality of features from the plurality of images of the 3D environment, transmitting information about the plurality of features and position information related to the plurality of features represented in the local coordinate frame to a location service over a network, and receiving at least one transformation from the location service that relates the local coordinate frame to a second coordinate frame.

[0008] In one embodiment, the electronic device comprises a display, and the computer-executable instructions further comprise instructions for rendering virtual content on the display at a position calculated based at least in part on a transformation of the at least one transformation, the virtual content having a location defined in the second coordinate frame. In one embodiment, the computer-executable instructions further comprise instructions for generating a descriptor for the plurality of features, and transmitting information about the plurality of features comprises transmitting the descriptor for the plurality of features.

[0009] In one embodiment, the computer-executable instructions further comprise instructions for storing the information about the plurality of features and the location information for the plurality of features in a buffer, and transmitting over the network comprises transmitting the contents of the buffer together such that the information about the plurality of features and the location information for the plurality of features are transmitted together. In one embodiment, the buffer has an adjustable size, and the computer-executable instructions further comprise instructions for increasing the size of the buffer in response to a failure indication received over the network from the location service.

[0010] In one embodiment, extracting a plurality of features from the plurality of images includes extracting up to a threshold number of features from each image, and the computer-executable instructions further comprise instructions for increasing the threshold number in response to a failure indication received over the network from the location service.

[0011] In one embodiment, the computer-executable instructions further comprise instructions for transmitting, via a network to a location service, a request to initiate a session with the location service. In one embodiment, the request to initiate a session with the location service comprises an identifier of the electronic device and calibration data for the electronic device. In one embodiment, the information about the plurality of characteristics and location information related to the plurality of characteristics transmitted via the network comprises a request for the location service, and the computer-executable instructions further comprise instructions for transmitting the request for location based on one or more trigger conditions being satisfied. In one embodiment, the one or more trigger conditions comprise a distance traveled by the electronic device since the last successful location request.

[0012] According to one aspect, a method for operating an electronic device configured to operate within a cross-reality system includes receiving information about a three-dimensional (3D) environment, the information comprising a plurality of images of the 3D environment; generating a local coordinate frame to represent a position within the 3D environment; extracting a plurality of features from the plurality of images of the 3D environment; transmitting information about the plurality of features and location information related to the plurality of features represented in the local coordinate frame to a location service over a network; and receiving at least one transformation from the location service that relates the local coordinate frame to a second coordinate frame.

[0013] According to one embodiment, the method further includes rendering virtual content having a location defined in the second coordinate frame on a display of the electronic device at a position calculated based at least in part on the transformation of the at least one transformation.

[0014] In one embodiment, the method further includes generating a descriptor for the plurality of features, and transmitting information about the plurality of features includes transmitting the descriptor for the plurality of features.

[0015] In one embodiment, the method further includes storing information about the plurality of features and location information for the plurality of features in a buffer, and transmitting over the network includes transmitting the contents of the buffer together such that the information about the plurality of features and the location information for the plurality of features are transmitted together.

[0016] In one embodiment, the buffer has an adjustable size, and the method further includes increasing the size of the buffer in response to a failure indication received over the network from the location service.

[0017] In one embodiment, extracting a plurality of features from the plurality of images includes extracting up to a threshold number of features from each image, and the method further includes increasing the threshold number in response to a failure indication received over the network from the location service.

[0018] In one embodiment, the method further includes sending, via the network, to a location service a request to initiate a session with the location service, the request to initiate a session with the location service comprising an identifier of the electronic device and calibration data for the electronic device.

[0019] In one embodiment, the information about the plurality of features and location information related to the plurality of features transmitted over the network comprises a request for location services, and the method further includes transmitting the request for location based on one or more trigger conditions being satisfied.

[0020] According to one embodiment, the one or more trigger conditions comprise a distance traveled by the electronic device since the last successful location request.

[0021] According to one aspect, a computer-readable medium stores computer-executable instructions that, when executed by at least one processor, are configured to implement a method for operating an electronic device configured to operate in a cross-reality system. The method includes receiving information about a three-dimensional (3D) environment, the information comprising a plurality of images of the 3D environment, generating a local coordinate frame to represent a location within the 3D environment, extracting a plurality of features from the plurality of images of the 3D environment, transmitting information about the plurality of features and location information for the plurality of features represented in the local coordinate frame to a location service over a network, and receiving at least one transformation from the location service that relates the local coordinate frame to a second coordinate frame.

[0022] According to one aspect, an XR system is provided that supports specification of a location of virtual content relative to a stored map in a database of stored maps. The system includes one or more computing devices configured for network communication with one or more portable electronic devices, a communication component configured to receive from the portable electronic devices information about a set of features in a three-dimensional (3D) environment of the portable electronic device and positioning information related to features of the received set of features represented in a first coordinate frame, and a localization component coupled to the communication component, the localization component configured to select a stored map from the database of stored maps based on the selected map having a set of features that matches the received set of features, the selected map comprising a second coordinate frame, generate a transformation between the first coordinate frame and the second coordinate frame based on a calculated match between the set of received features in the 3D environment of the portable electronic device and the matching set of features in the selected map, and transmit the transformation to the portable electronic device.

[0023] In one embodiment, the XR system is configured to receive, from the application, a specification of a location of the virtual content relative to persistent map features. In one embodiment, the stored map comprises a plurality of persistent map features, and the localization component is configured to represent the generated transformation between the first coordinate frame and the second coordinate frame as a plurality of transformations between the first coordinate frame and each of the plurality of persistent map features in the selected map. In one embodiment, the localization component is configured to select a stored map from the database of stored maps by filtering maps in the database of stored maps based, at least in part, on location information associated with the portable electronic device. In one embodiment, the location information comprises one or more of a radio fingerprint and / or GPS coordinates received from the portable electronic device. In one embodiment, the localization component is further configured to maintain information regarding the individual portable electronic device, the maintained information comprising a location history, and the location information comprising an individual location history for the portable electronic device.

[0024] According to one embodiment, the stored map is partitioned into a plurality of areas, and the localization component is configured to select the stored map by selecting an area of ​​the stored map having a set of features that matches the set of received features. According to one embodiment, the information about the received set of features comprises a calculated descriptor for the received set of features, and the localization component is configured to select the stored map having a set of features that matches the set of received features by identifying a candidate set of maps having a set of features with a greater than threshold number of features with descriptors that match the feature descriptors of the received set of features, and selecting the stored map from the candidate set of maps based on an error metric associated with a calculated transformation between the set of features and the set of features of the selected candidate map.

[0025] In one embodiment, the localization component is configured to generate an indication of a localization failure based on a search of the database of stored maps for maps with an error metric below a threshold returning no maps with an error metric below the threshold. In one embodiment, the localization service is further configured to maintain state information for each of the plurality of portable electronic devices, comprising, for each portable electronic device, at least one or any combination of: a device ID, a tracking map ID corresponding to the first coordinate frame, a previously generated map reference, and / or a transformation of the reference map and the tracking map. In one embodiment, the received information about the received set of features comprises information about a plurality of feature sets, and the localization component is configured to select a stored map by selecting a stored map having a set of features that matches more than a threshold number of the received set of features.

[0026] According to one aspect, a method of operating a portable electronic device and rendering virtual content within a 3D environment is provided, the method including generating, on the portable electronic device, a local coordinate frame based on output of one or more sensors on the portable electronic device, generating, on the portable electronic device, a plurality of descriptors for a plurality of features sensed within the 3D environment, transmitting, via a network, the plurality of descriptors of the plurality of features and location information of the plurality of features represented in the local coordinate frame to a location service, obtaining from the location service a transformation between stored coordinate frames of stored spatial information for the 3D environment and the local coordinate frame, receiving, from the location service, a specification of a virtual object having a virtual object coordinate frame and a location of the virtual object relative to the stored coordinate frame, and rendering the virtual object on a display of the portable electronic device at a location determined at least in part based on the calculated transformation and the received location of the virtual object.

[0027] According to one embodiment, transmitting over the network includes transmitting the plurality of descriptors of the plurality of features and location information for the plurality of features expressed in a local coordinate frame over the network to a cloud-hosted location service. According to one embodiment, obtaining a transformation between the stored coordinate frame of the stored spatial information about the 3D environment and the local coordinate frame from the location service includes obtaining the stored coordinate frame through an application programming interface (API).

[0028] According to one embodiment, the portable electronic device comprises a first portable electronic device comprising a first processor, and the system further comprises a second portable electronic device comprising a second processor, wherein the first and second processors each obtain a transformation between their respective local coordinate frames and an identical stored coordinate frame, receive specifications of a virtual object, and render the virtual object on a respective display for a respective user of the first and second portable electronic devices.

[0029] In one embodiment, the method further includes executing an application to generate specifications for the virtual object and a location of the virtual object relative to the stored coordinate frame for rendering in the local coordinate frame. In one embodiment, maintaining the local coordinate frame on the portable electronic device includes, for each of the first and second portable electronic devices, capturing a plurality of images of the 3D environment from one or more sensors of the portable electronic device and generating spatial information about the 3D environment based at least in part on the calculated one or more persistent poses, the method further includes, for each of the first and second portable electronic devices, transmitting the generated spatial information to a remote server, and obtaining the transformation includes receiving the transformation from a cloud-hosted location service.

[0030] According to one embodiment, the first and second portable electronic devices each comprise a download system configured to download the stored coordinate frame from the server and enable a locally performed transformation from the local to the stored coordinate frame in a device-based localization mode.

[0031] According to one aspect, a method is provided for operating a portable electronic device to render virtual content within a three-dimensional (3D) environment comprising a framework for an XR system, the method providing a shared experience to each of a plurality of users through the use of a coordinate frame of reference. The method includes generating, on the portable electronic device, a local coordinate frame based on outputs of one or more sensors on the portable electronic device and a tracking map constructed on the portable electronic device, generating, on the portable electronic device, a plurality of features for a plurality of images of the 3D environment, transmitting, from the portable electronic device, via a network to a location service, an indication of the plurality of features and location information related to the plurality of features represented in the local coordinate frame, and receiving, from the location service, at least one transformation relating the local coordinate frame to a second coordinate frame.

[0032] In one embodiment, the method further includes capturing information about the 3D environment from one or more sensors, the captured information comprising a plurality of images, and generating a map of at least a portion of the 3D environment based on the plurality of images. In one embodiment, the portable electronic device further comprises a display, and the method further includes rendering, on the display, virtual content having a location defined in a second coordinate frame at a position calculated based at least in part on a transformation of the at least one transformation. In one embodiment, the method further includes generating a descriptor for the plurality of features, and transmitting the information about the plurality of features comprises transmitting the descriptor for the plurality of features.

[0033] In one embodiment, the method further includes storing the information about the plurality of features and location information for the plurality of features in a buffer, and transmitting over the network includes transmitting the contents of the buffer together such that the information about the plurality of features and the location information for the plurality of features are transmitted together. In one embodiment, the buffer has an adjustable size, and the method further includes increasing a size of the buffer in response to a failure indication. In one embodiment, extracting the plurality of features from the plurality of images includes extracting up to a threshold number of features from each image, and the method further includes increasing the threshold number in response to a failure indication received over the network from the location service.

[0034] In one embodiment, the method further comprises: providing a location service via a network. and transmitting, from the portable electronic device, over the network, an indication of the plurality of features and location information related to the plurality of features represented in the local coordinate frame to the location service, the request to initiate a session with the location service. In one embodiment, the method further includes receiving a request to initiate a session with the location service, the request including an identifier of the portable electronic device and calibration data for the portable electronic device. In one embodiment, transmitting, from the portable electronic device, over the network to the location service, an indication of the plurality of features and location information related to the plurality of features represented in the local coordinate frame comprises a request for the location service, and the method further includes transmitting the request for location in response to one or more trigger conditions being satisfied.

[0035] According to one embodiment, the method further comprises determining whether a location has been determined since the last successful request for location. In response to identifying a distance traveled by the portable electronic device, determining that one or more trigger conditions are met.

[0036] According to one aspect, a method is provided for operating a portable electronic device to render virtual content within a three-dimensional (3D) environment comprising an XR system framework and providing a shared experience to each of a plurality of users through the use of stored coordinate frames. The method includes receiving, with one or more processors, at a cloud-hosted location service, information about a set of features within the 3D environment of the portable electronic device and positioning information related to features of the received set of features represented within a first coordinate frame; selecting, from a database of stored maps, a stored map based on the selected map having a set of features that matches the received set of features, the selected map comprising a second coordinate frame; generating a transformation between the first coordinate frame and the second coordinate frame based on the calculated match between the received set of features within the 3D environment of the portable electronic device and the matching set of features in the selected map; and transmitting the transformation to the portable electronic device.

[0037] According to one embodiment, the method further comprises: receiving from the application a persistent map feature; and receiving a location specification of the virtual content relative to the first coordinate frame, the stored map comprising a plurality of persistent map features; and representing the generated transformation between the first coordinate frame and the second coordinate frame as a plurality of transformations between the first coordinate frame and each of the plurality of persistent map features in the selected map. In one embodiment, the method further includes selecting a stored map from the database of stored maps by filtering maps in the database of stored maps based, at least in part, on location information associated with the portable electronic device. In one embodiment, the location information comprises one or more of a radio fingerprint and / or GPS coordinates received from the portable electronic device. In one embodiment, the method further includes maintaining information regarding the individual portable electronic device, the maintained information comprising a location history, the location information comprising an individual location history for the portable electronic device.

[0038] According to one embodiment, the method further comprises partitioning the stored map into a plurality of areas. and selecting, by the location service, an area of ​​the stored map having a set of features that matches the set of features received. According to one embodiment, the information about the set of features comprises a calculated descriptor for the received set of features, and the method further includes selecting a stored map having a set of features that matches the set of features received by identifying a candidate set of maps having a set of features with a greater than threshold number of features with descriptors that match the feature descriptors of the received set of features, and selecting a stored map from the candidate set of maps based on an error metric associated with the calculated transformation between the set of features received and the set of features of the selected candidate map.

[0039] According to one embodiment, the method further comprises selecting a map with an error metric below a threshold. generating an indication of localization failure based on a search of a database of stored maps for the first coordinate frame not returning a map with an error metric below a threshold. According to one embodiment, the method further includes maintaining state information for each of the plurality of portable electronic devices comprising, for each portable electronic device, at least one or any combination of a device ID, a tracking map ID corresponding to the first coordinate frame, a previously generated map reference, and / or a transformation of the reference map and the tracking map.

[0040] According to one embodiment, the received information about the set of received features is The method further includes receiving a virtual object specification having a virtual object coordinate frame and a location of the virtual object relative to the stored coordinate frame. ...

[0041] In one embodiment, the method further comprises: displaying the virtual object on a portable electronic device. and rendering the received set of features on a display of the chair at a location determined based at least in part on the calculated transformation and the received location of the virtual object. According to one embodiment, matching the descriptor for the received set of features to the descriptor for the set of features of the stored map includes at least matching a first persistent coordinate frame (PCF) of the local coordinate frame to a first PCF of the stored map (e.g., a previously generated and stored or reference map). According to one embodiment, matching the descriptor for the received set of features to the descriptor for the set of features of the stored map includes at least matching a second PCF of the local coordinate frame to a second PCF of the stored map (e.g., a previously generated and stored or reference map). According to one embodiment, the method further includes obtaining or deriving individual feature descriptors from any one or more or any combination of a frame, a portion of a frame, a key frame, a persistent pose, or a persistent coordinate frame (PCF).

[0042] The foregoing description is provided by way of illustration and is not intended to be limiting. The present invention provides, for example, the following. (Item 1) 1. An electronic device configured to operate within a cross reality system, said electronic device comprising: one or more sensors configured to capture information about a three-dimensional (3D) environment, the captured information comprising a plurality of images of the 3D environment; At least one processor configured to execute computer-executable instructions, the computer-executable instructions comprising: generating a local coordinate frame for representing a position within the 3D environment; extracting a plurality of features from a plurality of images of the 3D environment; transmitting, via a network, information about the plurality of features and location information for the plurality of features expressed in the local coordinate frame to a location service; receiving from the location service at least one transformation relating the local coordinate frame to a second coordinate frame; at least one processor having instructions for performing An electronic device comprising: (Item 2) the electronic device comprises a display; the computer-executable instructions further comprise instructions for rendering virtual content having a location defined in the second coordinate frame on the display at a position calculated based at least in part on a transformation of the at least one transformation. Item 1. The electronic device according to item 1. (Item 3) The computer-executable instructions further comprise instructions for generating a descriptor for the plurality of features; transmitting information about the plurality of features includes transmitting descriptors related to the plurality of features. Item 1. The electronic device according to item 1. (Item 4) The computer-executable instructions further comprise instructions for storing information about the plurality of features and position information for the plurality of features in a buffer; transmitting over a network includes transmitting the contents of the buffer together such that information about the plurality of features and location information related to the plurality of features are transmitted together. Item 3. The electronic device according to item 3. (Item 5) the buffer has an adjustable size; The computer-executable instructions further comprise instructions for increasing a size of the buffer in response to a failure indication received from the location service over the network. Item 4. The electronic device according to item 4. (Item 6) extracting the plurality of features from the plurality of images includes extracting up to a threshold number of features from each image; The computer-executable instructions further comprise instructions for increasing the threshold number in response to a failure indication received from the location service via the network. Item 6. The electronic device according to item 5. (Item 7) Item 10. The electronic device of item 1, wherein the computer-executable instructions further comprise instructions for sending, via the network, to the location service a request to initiate a session with the location service. (Item 8) 8. The electronic device of claim 7, wherein the request to initiate a session with the location service comprises an identifier of the electronic device and calibration data for the electronic device. (Item 9) the information about the plurality of characteristics and location information related to the plurality of characteristics transmitted over the network comprises a request for location services; The computer-executable instructions further comprise instructions for transmitting a request for location based on one or more trigger conditions being met. Item 1. The electronic device according to item 1. (Item 10) Item 10. The electronic device of item 9, wherein the one or more trigger conditions comprise a distance traveled by the electronic device since a last successful location request. (Item 11) 1. A method for operating an electronic device configured to operate in a cross reality system, the method comprising: receiving information about a three-dimensional (3D) environment, the information comprising a plurality of images of the 3D environment; generating a local coordinate frame for representing a position within the 3D environment; extracting a plurality of features from a plurality of images of the 3D environment; transmitting, via a network, information about the plurality of features and location information for the plurality of features expressed in the local coordinate frame to a location service; receiving from the location service at least one transformation relating the local coordinate frame to a second coordinate frame; A method comprising: (Item 12) Rendering virtual content having a location defined in the second coordinate frame on a display of the electronic device at a position calculated based at least in part on a transformation of the at least one transformation. Item 12. The method of item 11, further comprising: (Item 13) The method further includes generating a descriptor for the plurality of features; transmitting information about the plurality of features includes transmitting descriptors related to the plurality of features. Item 12. The method according to item 11. (Item 14) The method further includes storing information about the plurality of features and position information related to the plurality of features in a buffer; transmitting over a network includes transmitting the contents of the buffer together such that information about the plurality of features and location information related to the plurality of features are transmitted together. Item 14. The method according to item 13. (Item 15) the buffer has an adjustable size; The method further includes increasing a size of the buffer in response to a failure indication received from the location service via the network. Item 15. The method according to item 14. (Item 16) extracting the plurality of features from the plurality of images includes extracting up to a threshold number of features from each image; The method further includes increasing the threshold number in response to a failure indication received from the location service via the network. Item 15. The method according to item 15. (Item 17) Item 12. The method of item 11, wherein the method further includes sending, via the network, to the location service a request to initiate a session with the location service, the request to initiate a session with the location service comprising an identifier of the electronic device and calibration data for the electronic device. (Item 18) the information about the plurality of characteristics and location information related to the plurality of characteristics transmitted over the network comprises a request for location services; The method further includes transmitting a request for location based on one or more trigger conditions being satisfied. Item 12. The method according to item 11. (Item 19) Item 19. The method of item 18, wherein the one or more trigger conditions comprise a distance traveled by the electronic device since a last successful location request. (Item 20) 1. A computer-readable medium having computer-executable instructions stored thereon, the computer-executable instructions being configured, when executed by at least one processor, to implement a method for operating an electronic device configured to operate in a cross-reality system, the method comprising: receiving information about a three-dimensional (3D) environment, the information comprising a plurality of images of the 3D environment; generating a local coordinate frame for representing a position within the 3D environment; extracting a plurality of features from a plurality of images of the 3D environment; transmitting, via a network, information about the plurality of features and location information for the plurality of features expressed in the local coordinate frame to a location service; receiving from the location service at least one transformation relating the local coordinate frame to a second coordinate frame; 1. A computer-readable medium comprising: (Item 21) 1. An XR system that supports specification of the location of virtual content relative to a stored map in a database of stored maps, the system comprising: one or more computing devices configured for network communication with one or more portable electronic devices; a communication component configured to receive, from a portable electronic device, information about a set of features within a three-dimensional (3D) environment of the portable electronic device and positioning information for features of the received set of features represented in a first coordinate frame; a location component coupled to the communication component, the location component comprising: selecting a stored map from the database of stored maps based on the selected map having a set of features that matches the received set of features, the selected map comprising a second coordinate frame; generating a transformation between the first coordinate frame and the second coordinate frame based on a calculated match between the received set of features in the 3D environment of the portable electronic device and the matching set of features in the selected map; transmitting the transformation to the portable electronic device; a location component configured to: one or more computing devices comprising: A system comprising: (Item 22) the stored map comprises a plurality of persistent map features; the localization component is configured to represent the generated transformation between the first coordinate frame and the second coordinate frame as a plurality of transformations between the first coordinate frame and each of a plurality of persistent map features in the selected map. 22. The XR system according to item 21. (Item 23) Item 22. The system of item 21, wherein the location component is configured to select the stored map from the database of stored maps by filtering maps in the database of stored maps based, at least in part, on location information associated with the portable electronic device. (Item 24) 24. The system of claim 23, wherein the location information comprises one or more of a radio fingerprint and / or GPS coordinates received from the portable electronic device. (Item 25) The location component further comprises: maintaining information about individual portable electronic devices, the maintained information comprising location history; configured to: Item 24. The system of item 23, wherein the location information comprises an individual location history for the portable electronic device. (Item 26) The stored map is partitioned into a plurality of areas; the location component is configured to select a stored map by selecting an area of ​​the stored map having a set of features that matches the set of received features. Item 22. The system according to item 21. (Item 27) the information about the received feature set comprises a descriptor calculated for the received feature set; The location component: identifying a candidate set of maps having feature sets with a greater than threshold number of features with descriptors that match feature descriptors of the received feature set; selecting the stored map from the candidate set of maps based on an error metric associated with a calculated transformation between the received set of features and the selected candidate map's set of features; and selecting the stored map having a set of features that matches the set of received features by Item 22. The system according to item 21. (Item 28) 28. The system of claim 27, wherein the localization component is configured to generate an indication of localization failure based on a search of the database of stored maps for maps with the error metric below a threshold not returning a map with the error metric below the threshold. (Item 29) Item 22. The system of item 21, wherein the location service is further configured to maintain state information for each of the plurality of portable electronic devices comprising, for each portable electronic device, at least one or any combination of a device ID, a tracking map ID corresponding to the first coordinate frame, a previously generated map reference, and / or a transformation of a reference map and a tracking map. (Item 30) the received information about the received feature set comprises information about a plurality of feature sets; the location component is configured to select a stored map by selecting a stored map having a set of features that matches more than a threshold number of the received set of features. Item 22. The system according to item 21. (Item 31) 1. A method of operating a portable electronic device to render virtual content in a 3D environment, the method comprising: generating, on the portable electronic device, a local coordinate frame based on outputs of one or more sensors on the portable electronic device; generating, on the portable electronic device, a plurality of descriptors relating to a plurality of features sensed in the 3D environment; transmitting, via a network to a location service, a plurality of descriptors of the plurality of features and location information of the plurality of features expressed in the local coordinate frame; obtaining, from the location service, a transformation between the stored coordinate frame of stored spatial information about the 3D environment and the local coordinate frame; receiving a specification of a virtual object having a virtual object coordinate frame and a location of the virtual object relative to the stored coordinate frame; rendering the virtual object on a display of the portable electronic device at a location determined based at least in part on the calculated transformation and the received location of the virtual object; A method comprising: (Item 32) 32. The method of claim 31, wherein transmitting over the network includes transmitting a plurality of descriptors of the plurality of features and location information of the plurality of features expressed in the local coordinate frame over the network to a cloud-hosted location service. (Item 33) Item 32. The method of item 31, wherein obtaining from the location service a transformation between the stored coordinate frame of stored spatial information about the 3D environment and the local coordinate frame includes obtaining the stored coordinate frame through an application programming interface (API). (Item 34) the portable electronic device comprises a first portable electronic device comprising a first processor; the system further comprises a second portable electronic device comprising a second processor; The first and second processors each include: Obtaining a transformation between the individual local coordinate frame and the same stored coordinate frame; receiving a specification of the virtual object; rendering the virtual object on a separate display for each user of the first and second portable electronic devices; Item 32. The method according to Item 31, wherein (Item 35) Executing an application to generate a specification of the virtual object and its location relative to the stored coordinate frame for rendering in the local coordinate frame. Item 35. The method of item 34, further comprising: (Item 36) Maintaining a local coordinate frame on the portable electronic device includes, for each of the first and second portable electronic devices: capturing a plurality of images of the 3D environment from one or more sensors of the portable electronic device; generating spatial information about the 3D environment based at least in part on the calculated one or more sustained poses; and Including, The method further includes transmitting, for each of the first and second portable electronic devices, the generated spatial information to a remote server; obtaining the transformation includes receiving the transformation from a cloud-hosted location service; Item 35. The method according to item 34. (Item 37) The first and second portable electronic devices each include: a download system configured to download the stored coordinate frame from a server and enable a locally performed transformation from the local to the stored coordinate frame in a device-based localization mode; Item 35. The method of item 34, comprising: (Item 38) 1. A method of operating a portable electronic device to render virtual content within a three-dimensional (3D) environment comprising an XR system framework and providing a shared experience to each of a plurality of users through the use of a coordinate frame of reference, the method comprising: generating a local coordinate frame on the portable electronic device based on outputs of one or more sensors on the portable electronic device and a tracking map constructed on the portable electronic device; generating, on the portable electronic device, a plurality of features for a plurality of images of the 3D environment; transmitting from the portable electronic device, via a network, to a location service an indication of the plurality of features and location information for the plurality of features expressed in the local coordinate frame; receiving from the location service at least one transformation relating the local coordinate frame to a second coordinate frame; A method comprising: (Item 39) capturing information about the 3D environment from one or more sensors, the captured information comprising the plurality of images; and generating a map of at least a portion of the 3D environment based on the plurality of images; Item 39. The method of item 38, further comprising: (Item 40) Item 39. The method of item 38, wherein the portable electronic device further comprises a display, and the method further comprises rendering virtual content having a location defined in the second coordinate frame on the display at a position calculated based at least in part on a transformation of the at least one transformation. (Item 41) The method further comprises: generating a descriptor for the plurality of features; transmitting information about the plurality of features includes transmitting descriptors related to the plurality of features. Item 39. The method according to item 39. (Item 42) The method further comprises: storing information about the plurality of features and position information relating to the plurality of features in a buffer; transmitting over the network includes transmitting the contents of the buffer together such that information about the plurality of features and location information related to the plurality of features are transmitted together. Item 39. The method according to item 38. (Item 43) 43. The method of claim 42, wherein the buffer has an adjustable size, the method further comprising increasing the size of the buffer in response to a failure indication. (Item 44) extracting the plurality of features from the plurality of images includes extracting up to a threshold number of features from each image; The method further includes increasing the threshold number in response to a failure indication received from the location service via the network. Item 44. The method according to item 43. (Item 45) 39. The method of claim 38, further comprising sending, via the network, to the location service a request to initiate a session with the location service. (Item 46) 46. ​​The method of claim 45, further comprising receiving a request to initiate a session with the location service including an identifier of the portable electronic device and calibration data for the portable electronic device. (Item 47) transmitting, from the portable electronic device, via a network, to a location service, an indication of the plurality of features and location information related to the plurality of features expressed in the local coordinate frame, comprising a request for a location service; The method further includes transmitting a request for location in response to one or more trigger conditions being satisfied. Item 39. The method according to item 38. (Item 48) Item 48. The method of item 47, further comprising determining that the one or more trigger conditions have been met in response to identifying a distance traveled by the portable electronic device since a last successful request for location. (Item 49) 1. A method of operating a portable electronic device to render virtual content within a three-dimensional (3D) environment comprising an XR system framework and providing a shared experience to each of a plurality of users through the use of stored coordinate frames, the method comprising: receiving, at a cloud-hosted location service, information about a set of features within the 3D environment of the portable electronic device and positioning information related to features of the received set of features represented in a first coordinate frame; selecting a stored map from a database of stored maps based on the selected map having a set of features that matches the received set of features, the selected map comprising a second coordinate frame; generating a transformation between the first coordinate frame and the second coordinate frame based on a calculated match between the received set of features in the 3D environment of the portable electronic device and the matching set of features in the selected map; transmitting the transformation to the portable electronic device; A method comprising: (Item 50) The method further comprises: receiving, from an application, a specification of a location of virtual content relative to a persistent map feature, the stored map comprising a plurality of persistent map features; expressing the generated transformation between the first coordinate frame and the second coordinate frame as a plurality of transformations between the first coordinate frame and each of a plurality of persistent map features in the selected map; Item 49. The method according to Item 49, comprising: (Item 51) 50. The method of claim 49, further comprising selecting the stored map from the database of stored maps by filtering maps in the database of stored maps based, at least in part, on location information associated with the portable electronic device. (Item 52) 50. The method of claim 49, wherein the location information comprises one or more of a radio fingerprint and / or GPS coordinates received from the portable electronic device. (Item 53) The method further comprises: maintaining information about individual portable electronic devices, the maintained information comprising location history; Including, the location information comprises an individual location history for the portable electronic device. Item 49. The method according to item 49. (Item 54) The method further comprises: Partitioning the stored map into a plurality of areas; selecting, by the location service, an area of ​​the stored map having a set of characteristics that matches the set of received characteristics; Item 49. The method according to Item 49, comprising: (Item 55) The information about the received feature set comprises a descriptor calculated for the received feature set, and the method further comprises: identifying a candidate set of maps having feature sets with a greater than threshold number of features with descriptors that match feature descriptors of the received feature set; selecting the stored map from the candidate set of maps based on an error metric associated with a calculated transformation between the received set of features and the selected candidate map's set of features; 50. The method of claim 49, comprising selecting the stored map having a set of features that matches the set of received features by performing (Item 56) 50. The method of claim 49, further comprising generating an indication of a location failure based on a search of the database of stored maps for maps with the error metric below a threshold not returning a map with the error metric below the threshold. (Item 57) 50. The method of claim 49, wherein the method further includes maintaining, for each of the portable electronic devices, state information for each of the plurality of portable electronic devices, the state information comprising at least one or any combination of a device ID, a tracking map ID corresponding to the first coordinate frame, a previously generated map reference, and / or a transformation of a reference map and a tracking map. (Item 58) 58. The method of claim 57, wherein the received information about the received feature set comprises information about a plurality of feature sets, and the method further includes selecting a stored map by selecting a stored map having a feature set that matches more than a threshold number of the received feature set. (Item 59) Item 49. The method of item 49, further comprising obtaining a second transformation from the cloud-hosted location service, the second transformation defining a transformation between a stored coordinate frame of stored spatial information about the 3D environment and a local coordinate frame of the device. (Item 60) 60. The method of claim 59, further comprising receiving a specification of a virtual object having a virtual object coordinate frame and a location of the virtual object relative to the stored coordinate frame. (Item 61) Item 61. The method of item 60, further comprising rendering the virtual object on a display of the portable electronic device at a location determined based at least in part on the calculated transformation and the received location of the virtual object. (Item 62) Item 59. The method of item 59, wherein matching a descriptor for the set of features of the received feature set to a descriptor for the set of features of the stored map includes matching at least a first persistent coordinate frame (PCF) of the local coordinate frame to a first PCF of the stored map. (Item 63) Item 63. The method of item 62, wherein matching a descriptor for the received set of features to a descriptor for the set of features of the stored map includes at least matching a second PCF of the local coordinate frame to a second PCF of the stored map. (Item 64) Item 56. The method of item 55, wherein the method further includes obtaining or deriving individual feature descriptors from any one or more or any combination of a frame, a portion of a frame, a keyframe, a persistent pose, or a persistent coordinate frame (PCF). [Brief explanation of the drawings]

[0043] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component illustrated in various figures is represented by a like numeral. For purposes of clarity, not every component is labeled in every drawing.

[0044] [Figure 1] FIG. 1 is a sketch illustrating an example of a simplified augmented reality (AR) scene, according to some embodiments.

[0045] [Figure 2] FIG. 2 is a sketch of an exemplary simplified AR scene illustrating an exemplary use case of an XR system, according to some embodiments.

[0046] [Figure 3] FIG. 3 is a schematic diagram illustrating data flow for a single user in an AR system configured to provide the user with an experience of AR content that interacts with the physical world, according to some embodiments.

[0047] [Figure 4] FIG. 4 is a schematic diagram illustrating an exemplary AR display system displaying virtual content for a single user, according to some embodiments.

[0048] [Figure 5A] FIG. 5A is a schematic diagram illustrating a user wearing an AR display system that renders AR content as the user moves through a physical world environment, according to some embodiments.

[0049] [Figure 5B] FIG. 5B is a schematic diagram illustrating a viewing optics assembly and associated components, according to some embodiments.

[0050] [Figure 6A] FIG. 6A is a schematic diagram illustrating an AR system using a world reconstruction system, according to some embodiments.

[0051] [Figure 6B] FIG. 6B is a schematic diagram illustrating components of an AR system that maintains a model of a passable world, according to some embodiments.

[0052] [Figure 7] FIG. 7 is a schematic diagram of a tracking map formed by a device traversing a path through the physical world.

[0053] [Figure 8] FIG. 8 is a schematic diagram illustrating a user of a cross-reality (XR) system perceiving virtual content, according to some embodiments.

[0054] [Figure 9] FIG. 9 is a block diagram of components of the first XR device of the XR system of FIG. 8 that convert between coordinate systems, according to some embodiments.

[0055] [Figure 10] FIG. 10 is a schematic diagram illustrating an example transformation of an origin coordinate frame into a destination coordinate frame for correctly rendering local XR content, according to some embodiments.

[0056] [Figure 11] FIG. 11 is a top plan view illustrating a pupil-based coordinate frame, according to some embodiments.

[0057] [Figure 12] FIG. 12 is a top plan view illustrating a camera coordinate frame including all pupil positions, according to some embodiments.

[0058] [Figure 13] FIG. 13 is a schematic diagram of the display system of FIG. 9, according to some embodiments.

[0059] [Figure 14] FIG. 14 is a block diagram illustrating the creation of a persistent coordinate frame (PCF) and binding XR content to the PCF, according to some embodiments.

[0060] [Figure 15] FIG. 15 is a flowchart illustrating a method for establishing and using a PCF, according to some embodiments.

[0061] [Figure 16] FIG. 16 is a block diagram of the XR system of FIG. 8 including a second XR device, according to some embodiments.

[0062] [Figure 17] FIG. 17 is a schematic diagram illustrating a room and key frames established for various areas within the room, according to some embodiments.

[0063] [Figure 18] FIG. 18 is a schematic diagram illustrating the establishment of a sustained pose based on key frames, according to some embodiments.

[0064] [Figure 19] FIG. 19 is a schematic diagram illustrating the establishment of a persistent coordinate frame (PCF) based on a persistent attitude, according to some embodiments.

[0065] [Figure 20]20A-20C are schematic diagrams illustrating examples of creating a PCF, according to some embodiments.

[0066] [Figure 21] FIG. 21 is a block diagram illustrating a system for generating global descriptors for individual images and / or maps, according to some embodiments.

[0067] [Figure 22] FIG. 22 is a flowchart illustrating a method for calculating image descriptors, according to some embodiments.

[0068] [Figure 23] FIG. 23 is a flowchart illustrating a method for localization using image descriptors, according to some embodiments.

[0069] [Figure 24] FIG. 24 is a flowchart illustrating a method for training a neural network, according to some embodiments.

[0070] [Figure 25] FIG. 25 is a block diagram illustrating a method for training a neural network, according to some embodiments.

[0071] [Figure 26] FIG. 26 is a schematic diagram illustrating an AR system configured to rank and merge multiple environment maps, according to some embodiments.

[0072] [Figure 27] FIG. 27 is a simplified block diagram illustrating multiple criteria maps stored on a remote storage medium, according to some embodiments.

[0073] [Figure 28]FIG. 28 is a schematic diagram illustrating a method for selecting a reference map, for example, locating a new tracking map within one or more reference maps, and / or obtaining a PCF from a reference map, according to some embodiments.

[0074] [Figure 29] FIG. 29 is a flowchart illustrating a method for selecting a plurality of ranked environment maps, according to some embodiments.

[0075] [Figure 30] FIG. 30 is a schematic diagram illustrating an example map ranking portion of the AR system of FIG. 26, according to some embodiments.

[0076] [Figure 31A] FIG. 31A is a schematic diagram illustrating an example of area attributes of a tracking map (TM) and an environment map in a database, according to some embodiments.

[0077] [Figure 31B] FIG. 31B is a schematic diagram illustrating an example of determining the geographic location of a tracking map (TM) for the geographic location filtering of FIG. 29, according to some embodiments.

[0078] [Figure 32] FIG. 32 is a schematic diagram illustrating an example of the geographic location filtering of FIG. 29, according to some embodiments.

[0079] [Figure 33] FIG. 33 is a schematic diagram illustrating an example of the Wi-Fi BSSID filtering of FIG. 29, according to some embodiments.

[0080] [Figure 34] FIG. 34 is a schematic diagram illustrating an example of the use of the location determination of FIG. 29, according to some embodiments.

[0081] [Figure 35] 35 and 36 are block diagrams of an XR system configured to rank and merge multiple environment maps, according to some embodiments. [Figure 36] 35 and 36 are block diagrams of an XR system configured to rank and merge multiple environment maps, according to some embodiments.

[0082] [Figure 37] FIG. 37 is a block diagram illustrating a method for creating an environmental map of the physical world in a canonical form, according to some embodiments.

[0083] [Figure 38A] 38A and 38B are schematic diagrams illustrating an environment map created in a reference configuration by updating the tracking map of FIG. 7 with a new tracking map, according to some embodiments. [Figure 38B] 38A and 38B are schematic diagrams illustrating an environment map created in a reference configuration by updating the tracking map of FIG. 7 with a new tracking map, according to some embodiments.

[0084] [Figure 39-1] 39A-39F are schematic diagrams illustrating examples of merging maps, according to some embodiments. [Figure 39-2] 39A-39F are schematic diagrams illustrating examples of merging maps, according to some embodiments.

[0085] [Figure 40] FIG. 40 is a two-dimensional representation of a three-dimensional first local tracking map (Map 1), according to some embodiments, which may be generated by the first XR device of FIG.

[0086] [Figure 41]FIG. 41 is a block diagram illustrating uploading map 1 from a first XR device to the server of FIG. 9, according to some embodiments.

[0087] [Figure 42] FIG. 42 is a schematic diagram illustrating the XR system of FIG. 16 , according to some embodiments, showing a second user starting a second session using a second XR device of the XR system after the first user has finished the first session.

[0088] [Figure 43A] FIG. 43A is a block diagram illustrating a new session for the second XR device of FIG. 42, according to some embodiments.

[0089] [Figure 43B] FIG. 43B is a block diagram illustrating the creation of a tracking map for the second XR device of FIG. 42, according to some embodiments.

[0090] [Figure 43C] FIG. 43C is a block diagram illustrating downloading a reference map from a server to the second XR device of FIG. 42, according to some embodiments.

[0091] [Figure 44] FIG. 44 is a schematic diagram illustrating localization that attempts to locate a second tracking map (Map 2), which may be generated by the second XR device of FIG. 42, relative to a reference map, according to some embodiments.

[0092] [Figure 45] FIG. 45 is a schematic diagram illustrating localization attempting to locate the second tracking map (Map 2) of FIG. 44 with XR content associated with the PCF of Map 2, which may be further deployed against a reference map, according to some embodiments.

[0093] [Figure 46A] 46A-46B are schematic diagrams illustrating successful localization of map 2 of FIG. 45 relative to a reference map, according to some embodiments. [Figure 46B] 46A-46B are schematic diagrams illustrating successful localization of map 2 of FIG. 45 relative to a reference map, according to some embodiments.

[0094] [Figure 47] FIG. 47 is a schematic diagram illustrating a reference map generated by including one or more PCFs from the reference map of FIG. 46A into Map 2 of FIG. 45, according to some embodiments.

[0095] [Figure 48] FIG. 48 is a schematic diagram illustrating the reference map of FIG. 47 with a further extension of Map 2 on a second XR device, according to some embodiments.

[0096] [Figure 49] FIG. 49 is a block diagram illustrating uploading map 2 from a second XR device to a server, according to some embodiments.

[0097] [Figure 50] FIG. 50 is a block diagram illustrating merging Map 2 and a reference map, according to some embodiments.

[0098] [Figure 51] FIG. 51 is a block diagram illustrating the transmission of a new reference map from a server to a first and second XR device according to some embodiments.

[0099] [Figure 52] FIG. 52 is a block diagram illustrating a two-dimensional representation of map 2 and the head coordinate frame of a second XR device referenced to map 2, according to some embodiments.

[0100] [Figure 53] FIG. 53 is a block diagram illustrating adjustments of the head coordinate frame that can occur in two dimensions and six degrees of freedom, according to some embodiments.

[0101] [Figure 54] FIG. 54 is a block diagram illustrating a reference map on a second XR device where sounds are localized relative to the PCFs of Map 2, according to some embodiments.

[0102] [Figure 55] 55 and 56 are perspective and block diagrams illustrating use of an XR system when a first user has finished a first session and the first user has started a second session using the XR system, according to some embodiments. [Figure 56] 55 and 56 are perspective and block diagrams illustrating use of an XR system when a first user has finished a first session and the first user has started a second session using the XR system, according to some embodiments.

[0103] [Figure 57] 57 and 58 are perspective and block diagrams illustrating the use of an XR system when three users use the XR system simultaneously within the same session, according to some embodiments. [Figure 58] 57 and 58 are perspective and block diagrams illustrating the use of an XR system when three users use the XR system simultaneously within the same session, according to some embodiments.

[0104] [Figure 59] FIG. 59 is a flowchart illustrating a method for restoring and resetting head pose, according to some embodiments.

[0105] [Figure 60]FIG. 60 is a block diagram of a machine in the form of a computer that may find application within the system of the present invention, according to some embodiments.

[0106] [Figure 61] FIG. 61 is a schematic diagram of an exemplary XR system in which any of multiple devices may access location services, according to some embodiments.

[0107] [Figure 62] FIG. 62 is an exemplary process flow for operation of a portable device as part of an XR system that provides cloud-based localization, according to some embodiments.

[0108] [Figure 63A] 63A, B, and C are exemplary process flows for cloud-based location determination, according to some embodiments. [Figure 63B] 63A, B, and C are exemplary process flows for cloud-based location determination, according to some embodiments. [Figure 63C] 63A, B, and C are exemplary process flows for cloud-based location determination, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0109] Detailed Description Described herein are methods and apparatus for providing a cross-reality (XR) scene. To provide a realistic XR experience to multiple users, the XR system must understand the users' physical surroundings in order to correctly correlate the locations of virtual objects relative to real objects. The XR system may build an environmental map of the scene, which may be created from imagery and / or depth information collected using the sensor portion of an XR device worn by a user of the XR system.

[0110] The inventors have recognized and appreciated that it would be beneficial to have an XR system in which each XR device develops a local map of its physical environment by integrating information from one or more images collected during a scan at a given time. In some embodiments, the coordinate system of that map is tied to the device's orientation when the scan began. That orientation can change from session to session as the user interacts with the XR system, whether the different sessions are associated with different users, each with its own wearable device, with sensors scanning the environment, or the same user using the same device at different times. The inventors have recognized and appreciated a technique for operating an XR system based solely on persistent spatial information that overcomes the limitations of XR systems that rely solely on spatial information that each user device collects for different orientations for different user instances (e.g., snapshots in time) or sessions of the system (e.g., time between on and off). The techniques may provide XR scenes for more computationally efficient and immersive experiences for single or multiple users, for example, by enabling persistent spatial information to be created, stored, and retrieved by any of multiple users of an XR system.

[0111] The persistent spatial information may be represented by a persistent map, which may enable one or more functions that enhance the XR experience. The persistent map may be stored in a remote storage medium (e.g., the cloud). For example, after being turned on, a wearable device worn by a user may retrieve an appropriate stored map that was previously created and stored from a persistent storage device, such as a cloud storage device. The previously stored map may be based on data about the environment collected using sensors on the user's wearable device during a previous session. Retrieving the stored map may enable use of the wearable device without scanning the physical world using sensors on the wearable device. Alternatively, or in addition, the system / device may similarly retrieve an appropriate stored map in response to entering a new area in the physical world.

[0112] The stored map may be represented in a canonical form that may be relative to a local frame of reference on each XR device. In a multi-device XR system, a stored map accessed by one device may have been created and stored by another device and / or may have been constructed by aggregating data about the physical world collected by sensors on multiple wearable devices that previously existed within at least a portion of the physical world represented by the stored map.

[0113] The relationship between the reference map and each device's local map may be determined through a localization process. The localization process may be performed on each XR device based on a set of reference maps that are selected and transmitted to the device. However, the inventors have recognized and appreciated that network bandwidth and computational resources on the XR device may be reduced by providing a localization service that may be performed on a remote processor, such as may be implemented in the cloud. Battery drain and heat generation on the XR device may be reduced as a result, enabling the device to devote resources such as computation time, network bandwidth, battery life, and thermal budget to providing a more immersive user experience. Nevertheless, by appropriate selection of the information passed between each XR device and the localization service, localization may be performed with the latency and accuracy required to support such an immersive experience.

[0114] Sharing data about the physical world between multiple devices can enable a shared user experience of virtual content. For example, two XR devices with access to the same stored map can both be localized relative to the map. Once localized, the user devices may render virtual content having a location defined by reference to the stored map by transforming that location into a frame or reference maintained by the user devices. The user devices may use this local frame of reference to control their display and render the virtual content at the defined location.

[0115] To support these and other functions, an XR system may include components that develop, maintain, and use persistent spatial information, including one or more stored maps, based on data about the physical world collected using sensors on the user device. These components may be distributed across the XR system, with some operating, for example, on the head-mounted portion of the user device. Other components may operate on a computer associated with the user that is coupled to the head-mounted portion via a local or personal area network. Still others may operate at remote locations, such as one or more servers accessible via a wide area network.

[0116] These components may include, for example, a component that may identify, from information about the physical world collected by one or more user devices, information that is of sufficient quality to be stored as or within a persistent map. An example of such a component, described in more detail below, is a map merge component. Such a component may, for example, receive input from a user device and determine the suitability of portions of the input to be used to update the persistent map. The map merge component may, for example, split a local map created by a user device into portions, determine the mergeability of one or more of the portions with the persistent map, and merge portions that meet the identified mergeability criteria into the persistent map. The map merge component may also, for example, promote portions that are not merged with the persistent map to become separate persistent maps.

[0117] As another example, these components may include components that can help determine appropriate persistent maps that can be retrieved and used by a user device. An example of such a component, described in more detail below, is a map ranking component. Such a component may, for example, receive input from a user device and identify one or more persistent maps that are likely to represent an area of ​​the physical world in which the device is operating. The map ranking component may, for example, help select a persistent map to be used by the local device when rendering virtual content, gathering data about the environment, or performing other actions. Alternatively, or in addition, the map ranking component may help identify persistent maps that should be updated as additional information about the physical world is gathered by one or more user devices.

[0118] Still other components may determine transformations that transform information captured or described relative to one frame of reference into another frame of reference. For example, a sensor may be attached to a head-mounted display such that data read from the sensor indicates the location of an object in the physical world relative to the wearer's head pose. One or more transformations may be applied to relate that location information to a coordinate frame associated with the persistent environment map. Similarly, data that indicates where a virtual object should be rendered when represented in the coordinate frame of the persistent environment map may undergo one or more transformations to be in the frame of reference of the display on the user's head. As described in more detail below, there may be multiple such transformations. These transformations may be partitioned across components of the XR system so that they can be efficiently updated and / or applied in a distributed system.

[0119] In some embodiments, a persistent map may be constructed from information collected by multiple user devices. The XR devices may capture local spatial information at various locations and times and construct separate tracking maps using information collected by their respective sensors. Each tracking map may contain points that may be associated with features of a real object, which may include multiple features. Potentially, in addition to providing input for creating and maintaining the persistent map, the tracking maps may be used to track users' movements within a scene and enable the XR system to estimate individual users' head poses based on the tracking maps.

[0120] This codependency between map creation and head pose estimation constitutes a significant challenge. Substantial processing may be required to simultaneously create the map and estimate head pose. Because latency makes the XR experience less realistic for the user, processing must be performed quickly as objects move within the scene (e.g., moving a cup across a table) and as the user moves within the scene. On the other hand, XR devices may offer limited computational resources because their weight should be light for a user to wear comfortably. A lack of computational resources cannot be compensated for with more sensors, as adding sensors would also undesirably add weight. Furthermore, either more sensors or more computational resources leads to heat, which may cause deformation of the XR device.

[0121] The inventors have realized and appreciated techniques for operating XR systems and providing XR scenes for a more immersive user experience, such as estimating head pose at a frequency of 1 kHz, low usage of computational resources associated with an XR device that may be configured with, for example, four video graphics array (VGA) cameras operating at 30 Hz, one inertial measurement unit (IMU) operating at 1 kHz, the computational power of a single advanced RISC machine (ARM) core, less than 1 GB of memory, and less than 100 Mbps of network bandwidth. These techniques involve generating and maintaining maps, reducing the processing required to estimate head pose, and providing and consuming data with low computational overhead. The XR system may calculate its pose based on matched visual features. U.S. Patent Application No. 16 / 221,065 describes hybrid tracking and is incorporated herein by reference in its entirety.

[0122] These techniques may include reducing the amount of data processed when constructing a map, such as by constructing a sparse map using a set of mapped points and keyframes, and / or by dividing the map into blocks and enabling block-wise updates. Mapped points may be associated with points of interest in the environment. Keyframes may include information selected from camera-captured data. U.S. Patent Application No. 16 / 520,582 describes determining and / or evaluating a localization map and is incorporated herein by reference in its entirety.

[0123] In some embodiments, persistent spatial information may be represented in a way that can be easily shared among users and distributed components, including applications. Information about the physical world may be represented, for example, as a persistent coordinate frame (PCF). A PCF may be defined based on one or more points that represent recognized features in the physical world. The features may be selected so that they are likely to be the same for each user session of the XR system. PCFs may be sparse and provide less than all of the available information about the physical world so that they can be efficiently processed and transferred. Techniques for processing persistent spatial information may include creating dynamic maps based on one or more coordinate systems in real space across one or more sessions and generating persistent coordinate frames (PCFs) over the sparse maps that can be exposed to XR applications, for example, via an application programming interface (API). These capabilities may be supported by techniques for ranking and merging multiple maps created by one or more XR devices. Persistent spatial information may also enable quickly restoring and resetting head pose on each of one or more XR devices in a computationally efficient manner.

[0124] Furthermore, the technique may enable efficient comparison of spatial information. In some embodiments, image frames may be represented by numerical descriptors. The descriptors may be computed via a transformation that maps a set of features identified in the image to the descriptors. The transformation may be implemented within a trained neural network. In some embodiments, the set of features provided as input to the neural network may be a filtered set of features extracted from the image using a technique that, for example, preferentially selects features that are likely to be persistent.

[0125] Representing image frames as descriptors enables, for example, efficient matching of new image information with stored image information. The XR system may store one or more frame descriptors beneath the persistent map, along with the persistent map. Local image frames acquired by the user device may similarly be converted to such descriptors. By selecting stored maps with descriptors similar to those of the local image frames, one or more persistent maps likely to represent the same physical space as the user device may be selected with a relatively small amount of processing. In some embodiments, descriptors may be calculated for key frames in the local map and the persistent map, further reducing processing when comparing the maps. Such efficient comparison may be used, for example, to simplify finding a persistent map to load or update into the local device based on image information acquired using the local device.

[0126] The techniques described herein may be used together or separately with and for many types of devices, including wearable or portable devices with limited computing resources, that provide augmented or mixed reality scenes. In some embodiments, the techniques may be implemented by one or more services that form part of an XR system.

[0127] AR system overview

[0128] 1 and 2 illustrate scenes with virtual content displayed in conjunction with a portion of the physical world. For illustrative purposes, an AR system is used as an example of an XR system. Figures 3-6B illustrate example AR systems including one or more processors, memory, sensors, and a user interface that can operate in accordance with the techniques described herein.

[0129] 1 , an outdoor AR scene 354 is depicted in which a user of the AR technology sees a physical-world park-like setting 356 featuring people, trees, a building in the background, and a concrete platform 358. In addition to these items, the user of the AR technology also perceives as "seeing" a robotic figure 357 standing on the physical-world concrete platform 358 and a flying, cartoon-like avatar character 352, which thereby appears to be an anthropomorphic bumblebee, although these elements (e.g., avatar character 352 and robotic figure 357) do not exist within the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is difficult to produce AR technology that facilitates a comfortable, natural-feeling, and rich presentation of virtual image elements among other virtual or physical-world image elements.

[0130] Such AR scenes can be achieved using a system that allows a user to place AR content within a physical world, determines the location within a map of the physical world where the AR content is placed, saves the AR scene so that the placed AR content can be reloaded for display within the physical world, for example, between different AR experience sessions, and builds a map of the physical world based on tracking information, allowing multiple users to share the AR experience. The system can build and update a digital representation of the physical world surfaces around the user. This representation can be used to render virtual content so that it appears occluded, fully or partially, by physical objects between the user and the rendered location of the virtual content, for placing virtual objects, in physics-based interactions, and for virtual character path planning and navigation, or for other operations in which information about the physical world is used.

[0131] 2 depicts another example of an indoor AR scene 400, illustrating an exemplary use case of an XR system, according to some embodiments. The exemplary scene 400 is a living room with a wall, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of the AR technology also perceives virtual objects, such as an image on the wall behind the sofa, a bird flying through the door, a deer peeking out from the bookshelf, and a decoration in the form of a windmill placed on the coffee table.

[0132] For an image on a wall, the AR technology requires information not only about the surface of the wall, but also about objects and surfaces in the room, such as lamp shapes, that occlude the image to properly render the virtual object. For a flying bird, the AR technology requires information about all objects and surfaces around the room to render the bird with realistic physics, such as avoiding objects and surfaces or bouncing off if it collides. For a deer, the AR technology requires information about surfaces, such as the floor or coffee table, to calculate where the deer should be placed. For a windmill, the system may identify it as a separate object from the table and determine that it is movable, while the corners of the shelf or wall may be determined to be stationary. Such specificity may be used in determining which parts of the scene to use or update in each of the various actions.

[0133] The virtual object may be placed within a previous AR experience session. When a new AR experience session begins in a living room, the AR technology requires that the virtual object be displayed exactly where it was previously placed and be realistically visible from a different perspective. For example, a windmill should appear to be standing on a book, rather than drifting above the table, even in a different location without the book. Such drifting may occur if the user's location in the new AR experience session is not accurately located within the living room. As another example, if the user is viewing a windmill from a different perspective than the perspective from which the windmill was placed, the AR technology requires the corresponding side of the windmill to be displayed.

[0134] The scene may be presented to the user via a system that includes multiple components, including a user interface that may stimulate one or more user senses, such as vision, hearing, and / or touch. In addition, the system may include one or more sensors that may measure parameters of the physical portion of the scene, including the user's position and / or movement within the physical portion of the scene. Furthermore, the system may include one or more computing devices with associated computer hardware, such as memory. These components may be integrated into a single device or distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated into a wearable device.

[0135] 3 depicts an AR system 502 configured to provide an experience of AR content that interacts with a physical world 506, according to some embodiments. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by a user as part of a headset such that the user may wear the display over their eyes, like a pair of goggles or glasses. At least a portion of the display may be transparent such that the user may observe a see-through reality 510. The see-through reality 510 may correspond to a portion of the physical world 506 within the current viewpoint of the AR system 502, which may correspond to the user's viewpoint when the user is wearing a headset incorporating both the display and sensors of the AR system and obtaining information about the physical world.

[0136] AR content may also be presented on the display 508, overlaid on the see-through reality 510. To provide accurate interaction between the AR content and the see-through reality 510 on the display 508, the AR system 502 may include a sensor 522 configured to capture information about the physical world 506.

[0137] The sensors 522 may include one or more depth sensors that output depth maps 512. Each depth map 512 may have multiple pixels, each of which may represent a distance to a surface in the physical world 506 in a particular direction relative to the depth sensor. Raw depth data may originate from the depth sensors to create depth maps. Such depth maps may be updated as fast as the depth sensors can form new images, which may be hundreds or thousands of times per second. However, the data may be noisy and incomplete and may have holes, shown as black pixels on the illustrated depth maps.

[0138] The system may include other sensors, such as image sensors. The image sensors may obtain monocular or stereoscopic information that may be processed to represent the physical world in other ways. For example, images may be processed in the world reconstruction component 516 to create a mesh that represents connected portions of objects in the physical world. Metadata about such objects, including, for example, color and surface texture, may also be obtained using the sensors and stored as part of the world reconstruction.

[0139] The system may also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, a head pose tracking component of the system may be used to calculate head pose in real time. The head pose tracking component may represent the user's head pose in a coordinate frame with six degrees of freedom, including, for example, translation in three perpendicular axes (e.g., forward / back, up / down, left / right) and rotation about three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, the sensors 522 may include an inertial measurement unit ("IMU"), which may be used to calculate and / or determine the head pose 514. The head pose 514 for a depth map may indicate, for example, the current viewpoint of the sensor capturing the depth map with six degrees of freedom, although the head pose 514 may also be used for other purposes, such as relating image information to a particular portion of the physical world or relating the position of a display worn on the user's head to the physical world.

[0140] In some embodiments, head pose information may be derived in a manner other than by an IMU, such as from analysis of objects in images. For example, the head pose tracking component may calculate the relative position and orientation of the AR device with respect to a physical object based on visual information captured by a camera and inertial information captured by an IMU. The head pose tracking component may then calculate the head pose of the AR device, for example, by comparing the calculated relative position and orientation of the AR device with features of the physical object. In some embodiments, the comparison may be made by identifying features in images captured using one or more of sensors 522 that are stable over time, such that changes in the position of these features in images captured over time can be associated with changes in the user's head pose.

[0141] In some embodiments, the AR device may build a map from feature points recognized in successive images in a series of image frames captured as the user moves through the physical world with the AR device. Although each image frame may be obtained from a different pose as the user moves, the system may adjust the orientation of features in each successive image frame by matching features in the successive image frames with previously captured image frames to match the orientation of the initial image frame. Translation of the successive image frames can be used to align each successive image frame and match the orientation of the previously processed image frame so that points representing the same feature will match corresponding feature points from the previously collected image frame. The frames in the resulting map may have a common orientation established when the first image frame was added to the map. This map may be used to determine the user's pose in the physical world by matching features from the current image frame to the map, along with a set of feature points in a common reference frame. In some embodiments, this map may be referred to as a tracking map.

[0142] In addition to enabling tracking of the user's pose within the environment, this map may enable other components of the system, such as a world reconstruction component 516, to determine the location of physical objects relative to the user. The world reconstruction component 516 may receive the depth map 512 and head pose 514 and any other data from the sensors and integrate the data into a reconstruction 518. The reconstruction 518 may be more complete and less noisy than the sensor data. The world reconstruction component 516 may update the reconstruction 518 using spatial and temporal averages of the sensor data from multiple viewpoints over time.

[0143] Reconstruction 518 may include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats may represent alternative representations of the same portion of the physical world or may represent different portions of the physical world. In the illustrated example, on the left side of reconstruction 518, a portion of the physical world is presented as a global surface, and on the right side of reconstruction 518, a portion of the physical world is presented as a mesh.

[0144] In some embodiments, the map maintained by the head pose component 514 may be sparse with respect to other maps that may be maintained of the physical world. Rather than providing information about the location and possibly other characteristics of surfaces, the sparse map may indicate the location of points of interest and / or structures, such as corners or edges. In some embodiments, the map may include image frames as captured by the sensors 522. These frames may be reduced to features that may represent points of interest and / or structures. Along with each frame, information about the user's pose from which the frame was obtained may also be stored as part of the map. In some embodiments, all images acquired by the sensors may or may not be stored. In some embodiments, the system may process images as they are collected by the sensors and select a subset of image frames for further calculation. The selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may add new image frames to the map based on overlap with previous image frames already added to the map, for example, or based on image frames containing a sufficient number of features determined to likely represent stationary objects. In some embodiments, a selected image frame or a group of features from a selected image frame may serve as a keyframe for a map, which is used to provide spatial information.

[0145] The AR system 502 may integrate sensor data from multiple perspectives of the physical world over time. The pose (e.g., position and orientation) of the sensor may be tracked as the device containing the sensor is moved. As the frame pose of the sensor and how it relates to other poses is understood, each of these multiple perspectives of the physical world may be fused together into a single combined reconstruction of the physical world, which may serve as an abstraction layer for the map and provide spatial information. The reconstruction may be more complete and less noisy than the original sensor data by using spatial and temporal averaging (i.e., averaging data from multiple perspectives over time) or any other suitable method.

[0146] 3, the map represents a portion of the physical world in which a user of a single wearable device resides. In that scenario, head poses associated with frames in the map may be represented as local head poses, indicating orientation relative to an initial orientation for the single device at the start of a session. For example, head poses may be tracked relative to an initial head pose when the device is turned on or otherwise operated to scan the environment and build a representation of that environment.

[0147] In combination with content characterizing that portion of the physical world, the map may include metadata. The metadata may, for example, indicate the capture time of the sensor information used to form the map. Alternatively, or in addition, the metadata may indicate the location of the sensor at the capture time of the information used to form the map. Location may be represented directly, such as using information from a GPS chip, or indirectly, such as using a wireless (e.g., Wi-Fi) signature indicating the strength of signals received from one or more wireless access points while the sensor data was being collected, and / or an identifier such as the BSSID of a wireless access point to which the user device connected while the sensor data was being collected.

[0148] Reconstruction 518 may be used for AR functions, such as producing a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation may change as the user moves or as objects in the physical world change. Aspects of reconstruction 518 may be used by component 520, for example, to produce a changing global surface representation in world coordinates that can be used by other components.

[0149] AR content may be generated, such as by an AR application 504, based on this information. The AR application 504 may be, for example, a game program that performs one or more functions based on information about the physical world, such as visual occlusion, physics-based interactions, and environmental inference. It may perform these functions by querying data in different formats from the reconstruction 518 produced by the world reconstruction component 516. In some embodiments, the component 520 may be configured to output updates as the representation of the physical world within a region of interest changes. The region of interest may be set to approximate a portion of the physical world in the vicinity of a user of the system, such as a portion within the user's field of view, or may be projected (predicted / determined) to be within the user's field of view.

[0150] The AR application 504 may use this information to generate and update AR content, which virtual portions may be presented on the display 508 in combination with see-through reality 510 to create a realistic user experience.

[0151] In some embodiments, the AR experience may be provided to the user through an XR device, which may be a wearable display device that may be part of a system that may include remote processing and / or remote data storage, and / or, in some embodiments, other wearable display devices worn by other users. FIG. 4 illustrates an example of a system 580 (hereinafter referred to as “system 580”) that includes a single wearable device for ease of illustration. System 580 includes a head-mounted display device 562 (hereinafter referred to as “display device 562”) and various mechanical and electronic modules and systems that support the functionality of display device 562. Display device 562 may be coupled to a frame 564, which is wearable by a user or viewer 560 of the display system (hereinafter referred to as “user 560”) and configured to position display device 562 directly in front of the eyes of user 560. According to various embodiments, display device 562 may be a sequential display. Display device 562 may be monocular or binocular. In some embodiments, display device 562 may be an example of display 508 in FIG.

[0152] In some embodiments, a speaker 566 is coupled to the frame 564 and positioned proximate the ear canal of the user 560. In some embodiments, another speaker, not shown, is positioned adjacent another ear canal of the user 560 to provide stereo / adjustable sound control. The display device 562 is operably coupled, such as by wired leads or wireless connectivity 568, to a local data processing module 570, which may be mounted in a variety of configurations, such as fixedly attached to the frame 564, fixedly attached to a helmet or hat worn by the user 560, integrated into headphones, or otherwise removably attached to the user 560 (e.g., in a backpack configuration, in a belt-coupled configuration).

[0153] The local data processing module 570 may include a processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storing data. The data includes a) data captured from sensors (e.g., which may be operatively coupled to the frame 564 or otherwise attached to the user 560), such as image capture devices (e.g., cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, wireless devices, and / or gyroscopes, and / or b) data obtained and / or processed using the remote processing module 572 and / or the remote data repository 574, possibly for passing to the display device 562 after processing or retrieval.

[0154] In some embodiments, the wearable device may communicate with remote components. The local data processing module 570 may be operably coupled to a remote processing module 572 and a remote data repository 574 by communication links 576, 578, respectively, such as via wired or wireless communication links, such that these remote modules 572, 574 are operably coupled to each other and available as resources to the local data processing module 570. In further embodiments, in addition to or as an alternative to the remote data repository 574, the wearable device may access cloud-based remote data repositories and / or services. In some embodiments, the head pose tracking components described above may be implemented at least in part within the local data processing module 570. In some embodiments, the world reconstruction component 516 in FIG. 3 may be implemented at least in part within the local data processing module 570. For example, the local data processing module 570 may be configured to execute computer-executable instructions to generate a map and / or a physical world representation based, at least in part, on at least a portion of the data.

[0155] In some embodiments, processing may be distributed across local and remote processors. For example, local processing may be used to build a map (e.g., a tracking map) on the user device based on sensor data collected using sensors on the user's device. Such a map may be used by applications on the user's device. In addition, a previously created map (e.g., a reference map) may be stored in the remote data repository 574. If a suitable stored or persistent map is available, it may be used instead of or in addition to a tracking map created locally on the device. In some embodiments, the tracking map may be located relative to a stored map such that a correspondence is established between the tracking map, which may be oriented relative to the position of the wearable device at the time the user turned on the system, and the reference map, which may be oriented relative to one or more persistent features. In some embodiments, the persistent map may be loaded onto the user device, enabling the user device to render virtual content without the delay associated with scanning a location to build a tracking map of the user's complete environment from sensor data obtained during the scan. In some embodiments, a user device may access a remote persistent map (eg, stored on the cloud) without having to download the persistent map onto the user device.

[0156] In some embodiments, spatial information may be communicated from the wearable device to a remote service, such as a cloud service, configured to locate the device and store it in a map maintained on the cloud service. According to one embodiment, the localization process may occur in the cloud, matching the device location to an existing map, such as a reference map, and returning a transformation that links virtual content to the wearable device location. In such an embodiment, the system may avoid communicating a map from a remote resource to the wearable device. Other embodiments are configured for both device-based and cloud-based localization and may enable functionality when, for example, network connectivity is unavailable or the user chooses not to enable cloud-based localization.

[0157] Alternatively, or in addition, the tracking map may be merged with previously stored maps to enhance or improve the quality of those maps. Processing to determine whether suitable previously created environment maps are available and / or whether to merge the tracking map with one or more stored environment maps may occur within the local data processing module 570 or the remote processing module 572.

[0158] In some embodiments, the local data processing module 570 may include one or more processors (e.g., graphics processing units (GPUs)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which would limit the computational budget of the local data processing module 570 but enable smaller devices. In some embodiments, the world reconstruction component 516 may generate the physical world representation in real time over a non-predetermined space using a computational budget of less than a single Advanced RISC Machine (ARM) core, such that the remaining computational budget of the single ARM core can be accessed for other uses, such as, for example, extracting meshes.

[0159] In some embodiments, the remote data repository 574 may include a digital data storage facility, which may be available through the Internet or other networking configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, allowing for fully autonomous use from the remote module. In some embodiments, all data is stored and all or most computations are performed in the remote data repository 574, allowing for smaller devices. World reconstructions may be stored in whole or in part in this repository 574, for example.

[0160] In embodiments in which data is stored remotely and accessible via a network, the data may be shared by multiple users of the augmented reality system. For example, a user device may upload its tracking map and augment it into a database of environmental maps. In some embodiments, uploading of the tracking map occurs at the end of a user session with the wearable device. In some embodiments, uploading of the tracking map may occur continuously, semi-continuously, intermittently, at predefined times, after a predefined period since the previous upload, or when triggered by an event. A tracking map uploaded by any user device may be used to augment or refine a previously stored map, whether based on data from that user device or any other user device. Similarly, a persistent map downloaded to a user device may be based on data from that user device or any other user device. In this way, high-quality environmental maps may be readily available to users to refine their experience with the AR system.

[0161] In further embodiments, downloading of persistent maps may be limited and / or avoided based on localization performed on remote resources (e.g., in the cloud). In such a configuration, a wearable device or other XR device communicates feature information (e.g., positioning information about the device at the time the feature represented in the feature information was sensed) combined with pose information to a cloud service. One or more components of the cloud service may match the feature information with a separate stored map (e.g., a reference map) and generate a transformation between the tracking map maintained by the XR device and the coordinate system of the reference map. Each XR device, having its tracking map localized relative to the reference map, can accurately render virtual content at a location defined relative to the reference map based on its own tracking.

[0162] In some embodiments, local data processing module 570 is operably coupled to battery 582. In some embodiments, battery 582 is a removable power source, such as a commercially available battery. In other embodiments, battery 582 is a lithium-ion battery. In some embodiments, battery 582 includes both an internal lithium-ion battery that is rechargeable by user 560 during periods of non-operation of system 580, and a removable battery, so that user 560 can operate system 580 for longer periods of time without being plugged in and having to charge the lithium-ion battery or shut off system 580 and replace the battery.

[0163] FIG. 5A illustrates a user 530 wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as “environment 532”). Information captured by the AR system along the user's path of movement may be processed into one or more tracking maps. The user 530 positions the AR display system at a location 534, and the AR display system records ambient information of the passable world (e.g., digital representations of real objects in the physical world that may be stored and updated as the real objects change) relative to the location 534. That information may be stored as a pose in combination with images, features, directional audio input, or other desired data. The location 534 is aggregated with data input 536, e.g., as part of a tracking map, and processed by at least a passable world module 538, which may be implemented, for example, by processing on the remote processing module 572 of FIG. 4. In some embodiments, the passable world module 538 may include a head pose component 514 and a world reconstruction component 516 so that the processed information, in combination with other information about physical objects used in the rendered virtual content, may indicate the location of the objects in the physical world.

[0164] The passable world module 538 determines, at least in part, where and how the AR content 540 can be placed in the physical world, as determined from the data input 536. The AR content is “placed” in the physical world by presenting both a representation of the physical world and the AR content via a user interface, where the AR content is rendered as if interacting with objects in the physical world, and the objects in the physical world are presented as if the AR content obscures the user's view of those objects, when appropriate. In some embodiments, the AR content may be placed by appropriately selecting a portion of a fixed element 542 (e.g., a table) from the reconstruction (e.g., reconstruction 518) and determining the shape and position of the AR content 540. As an example, the fixed element may be a table, and the virtual content may be positioned to appear on the table. In some embodiments, the AR content may be placed within a structure in the field of view 544, which may be the current field of view or an estimated future field of view. In some embodiments, the AR content may be persisted to a model 546 (e.g., a mesh) of the physical world.

[0165] As depicted, fixed element 542 serves as a proxy (e.g., a digital copy) for any fixed element in the physical world, which may be stored within passable world module 538 so that user 530 may perceive content on fixed element 542 without the system having to map it to fixed element 542 each time user 530 sees it. Fixed element 542 may thus be a mesh model, determined from a previous modeling session or from a separate user, but stored by passable world module 538 for future reference by multiple users. Thus, passable world module 538 may recognize environment 532 from previously mapped environments and display AR content without user 530's device having to first map all or part of environment 532, saving computational processes and cycles and avoiding latency for any rendered AR content.

[0166] A mesh model 546 of the physical world may be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying the AR content 540 can be stored by the passable world module 538 for future retrieval by the user 530 or other users, without having to recreate the model, in whole or in part. In some embodiments, the data inputs 536 are inputs such as geographic location, user identification, and current activity to indicate to the passable world module 538 which fixed elements 542 of one or more fixed elements are available, the AR content 540 last placed on the fixed elements 542, and whether that same content should be displayed (such AR content is “persistent” content regardless of whether the user is viewing a particular passable world model).

[0167] Even in embodiments where objects are considered fixed (e.g., a kitchen table), the passable world module 538 may update those objects in the model of the physical world from time to time to account for possible changes in the physical world. Models of fixed objects may be updated very infrequently. Other objects in the physical world may be moving or otherwise not considered fixed (e.g., a kitchen chair). To render the AR scene with a realistic feel, the AR system may update the positions of these non-fixed objects much more frequently than those used to update fixed objects. To enable accurate tracking of all of the objects in the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.

[0168] 5B is a schematic illustration of a viewing optics assembly 548 and associated components. In some embodiments, two eye tracking cameras 550 are pointed towards the user's eye 549 to detect metrics of the user's eye 549, such as eye shape, eyelid occlusion, pupil direction, and phosphenes on the user's eye 549.

[0169] In some embodiments, one of the sensors may be a depth sensor 551, such as a time-of-flight sensor, that emits signals into the world, detects reflections of those signals from nearby objects, and determines the distance to a given object. The depth sensor may, for example, quickly determine whether objects have entered the user's field of view, either as a result of the objects' movement or a change in the user's posture. However, information about the location of objects within the user's field of view may alternatively or additionally be collected using other sensors. Depth information may be obtained, for example, from a stereoscopic image sensor or a plenoptic sensor.

[0170] In some embodiments, world camera 552 records, maps, and / or otherwise models a wider-than-surrounding view of environment 532 and detects inputs that may affect the AR content. In some embodiments, world camera 552 and / or camera 553 may be grayscale and / or color image sensors that output grayscale and / or color image frames at fixed time intervals. Camera 553 may also capture physical world images within the user's field of view at specific times. Pixels of a frame-based image sensor may be repeatedly sampled even if their values ​​remain constant. World camera 552, camera 553, and depth sensor 551 each have separate fields of view 554, 555, and 556, and collect and record data from a physical world scene, such as physical world environment 532 depicted in FIG. 34A .

[0171] An inertial measurement unit 557 may determine the movement and orientation of the viewing optics assembly 548. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 551 is operably coupled to the eye tracking camera 550 as a check of the measured accommodation against the actual distance seen by the user's eyes 549.

[0172] It should be understood that viewing optics assembly 548 may include some of the components illustrated in FIG. 34B , or may include components instead of or in addition to the components illustrated. In some embodiments, for example, viewing optics assembly 548 may include two world cameras 552 instead of four. Alternatively, or in addition, cameras 552 and 553 need not capture visible light images of their full field of view. Viewing optics assembly 548 may include other types of components. In some embodiments, viewing optics assembly 548 may include one or more dynamic vision sensors (DVS), whose pixels may asynchronously respond to relative changes in light intensity exceeding a threshold.

[0173] In some embodiments, the viewing optics assembly 548 may not include a depth sensor 551 based on time-of-flight information. In some embodiments, for example, the viewing optics assembly 548 may include one or more plenoptic cameras, the pixels of which may capture the light intensity and angle of incident light, from which depth information can be determined. For example, the plenoptic camera may include an image sensor overlaid with a transmissive diffractive mask (TDM). Alternatively, or in addition, the plenoptic camera may include an image sensor containing angle-sensing pixels and / or phase-detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors may serve as a depth information source instead of, or in addition to, the depth sensor 551.

[0174] 5B is provided as an example. Viewing optics assembly 548 may include components in any suitable configuration, which may be configured to provide the user with the largest field of view practical for a particular set of components. For example, if viewing optics assembly 548 has one world camera 552, the world camera may be located within a central region of the viewing optics assembly instead of on the side.

[0175] Information from sensors in the viewing optics assembly 548 may be coupled to one or more processors in the system. The processor may generate data that can be rendered to cause the user to perceive the virtual content as interacting with objects in the physical world. The rendering may be implemented in any suitable manner, including generating image data that depicts both physical and virtual objects. In other embodiments, physical and virtual content may be depicted in a single scene by modulating the opacity of a display device through which the user sees the physical world. The opacity may be controlled to create the appearance of virtual objects and block the user from seeing objects in the physical world that are occluded by the virtual objects. In some embodiments, the image data may include only the virtual content, which may be modified (e.g., clipping content and accounting for occlusion) so that the virtual content is perceived by the user to interact realistically with the physical world when viewed through the user interface.

[0176] The locations on the viewing optics assembly 548 where content can be displayed to create the impression of an object at a particular location may depend on the physics of the viewing optics assembly. Additionally, the user's head posture relative to the physical world and the direction the user's eyes are looking will affect the location within the physical world content displayed at a particular location on the viewing optics assembly where the content will appear. Sensors such as those described above may provide information from which this information can be collected and / or calculated so that a processor receiving the sensor input can calculate where objects should be rendered on the viewing optics assembly 548 to create a desired appearance for the user.

[0177] Regardless of how content is presented to a user, a model of the physical world may be used so that the characteristics of virtual objects that may be affected by physical objects, including the shape, position, movement, and visibility of the virtual objects, may be correctly calculated. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 518.

[0178] The model may be created from data collected from sensors on the user's wearable device, although in some embodiments the model may be created from data collected by multiple users, which may be aggregated in a computing device remote from all users (and may be "in the cloud").

[0179] The model may be created, at least in part, by a world reconstruction system, such as the world reconstruction component 516 of FIG. 3 , which is depicted in further detail in FIG. 6A . The world reconstruction component 516 may include a perception module 660, which may generate, update, and store representations for portions of the physical world. In some embodiments, the perception module 660 may represent a portion of the physical world within a sensor's reconstruction range as a plurality of voxels. Each voxel corresponds to a 3D cube of a predetermined volume within the physical world and may include surface information, indicating whether a surface exists within the volume represented by the voxel. A voxel may be assigned a value indicating whether its corresponding volume has been determined to contain a surface of a physical object, has been determined to be empty, or has not yet been measured with a sensor, and therefore its value is unknown. It should be understood that values ​​indicating voxels determined to be empty or unknown need not be explicitly stored, and that voxel values ​​may be stored in computer memory in any suitable manner, including not storing information about voxels determined to be empty or unknown.

[0180] In addition to generating information for the persisted world representation, perception module 660 may identify and output indications of changes in the area surrounding the user of the AR system. Such indications of changes may trigger updates to the stereoscopic data stored as part of the persisted world, or trigger other functionality, such as triggering the generate AR content and update AR content component 604.

[0181] In some embodiments, the perception module 660 may identify changes based on a signed distance function (SDF) model. The perception module 660 may be configured to receive sensor data, such as a depth map 660a and a head pose 660b, and then fuse the sensor data into an SDF model 660c. The depth map 660a may provide SDF information directly, or images may be processed to arrive at the SDF information. The SDF information represents distances from sensors used to capture that information. Because those sensors may be part of the wearable unit, the SDF information may represent the physical world from the perspective of the wearable unit and, therefore, the user. The head pose 660b may allow the SDF information to be associated with voxels in the physical world.

[0182] In some embodiments, perception module 660 may generate, update, and store representations for portions of the physical world that are within its perception range. The perception range may be determined, at least in part, based on the reconstruction range of the sensor, which may be determined, at least in part, based on the limits of the sensor's observation range. As a specific example, an active depth sensor that operates using active IR pulses may operate reliably over a range of distances, creating a sensor's observation range that may be a few centimeters or tens of centimeters to several meters.

[0183] The world reconstruction component 516 may include additional modules that may interact with the perception module 660. In some embodiments, the persisted world module 662 may receive a representation for the physical world based on data obtained by the perception module 660. The persisted world module 662 may also include representations of the physical world in various formats. For example, volumetric metadata 662b, such as voxels, may be stored along with the meshes 662c and planes 662d. In some embodiments, other information, such as depth maps, may also be saved.

[0184] In some embodiments, a representation of the physical world such as that illustrated in FIG. 6A may provide relatively dense information about the physical world compared to a sparse map such as a tracking map based on feature points, as described above.

[0185] In some embodiments, perception module 660 may include modules that generate representations for the physical world in various formats, including, for example, meshes 660d, planes, and semantics 660e. Representations for the physical world may be stored across local and remote storage media. Representations for the physical world may be described in different coordinate frames, for example, depending on the location of the storage media. For example, a representation for the physical world stored in a device may be described in a coordinate frame local to the device. A representation for the physical world may have a counterpart stored in the cloud. The counterpart in the cloud may be described in a coordinate frame shared by all devices in the XR system.

[0186] In some embodiments, these modules may generate representations based on data within the perception range of one or more sensors at the time the representation is generated, data captured at earlier times, and information in the persisted world module 662. In some embodiments, these components may operate on depth information captured using a depth sensor. However, AR systems may also include vision sensors and generate such representations by analyzing monocular or binocular visual information.

[0187] In some embodiments, these modules may operate on regions of the physical world. They may be triggered to update subregions of the physical world when perception module 660 detects a change in the physical world within that subregion. Such a change may be detected, for example, by detecting a new surface in SDF model 660c or by other criteria, such as a change in the values ​​of a sufficient number of voxels representing the subregion.

[0188] The world reconstruction component 516 may include components 664 that may receive representations of the physical world from the perception module 660. Information about the physical world may be pulled by these components, for example, according to usage requests from applications. In some embodiments, information may be pushed to the usage components, such as via indications of changes in pre-identified areas or changes in the physical world representation within the perception range. The components 664 may include, for example, game programs and other components that implement processing for visual occlusion, physics-based interactions, and environmental inference.

[0189] In response to a query from component 664, perception module 660 may transmit a representation for the physical world in one or more formats. For example, when component 664 indicates the use is for visual occlusion or physics-based interaction, perception module 660 may transmit a representation of a surface. When component 664 indicates the use is for environmental inference, perception module 660 may transmit meshes, planes, and semantics of the physical world.

[0190] In some embodiments, perception module 660 may include a component that provides formatting information to component 664. An example of such a component may be ray casting component 660f. A usage component (e.g., component 664) may, for example, query information about the physical world from a particular viewpoint. Ray casting component 660f may select from one or more representations of the physical world data within the field of view from that viewpoint.

[0191] As should be understood from the foregoing description, the perception module 660 or another component of the AR system may process data and create a 3D representation of a portion of the physical world. The data to be processed may be reduced, at least in part, by thinning a portion of the 3D reconstruction volume based on the camera frustum and / or depth image; extracting and persisting planar data; capturing, persisting, and updating the 3D reconstruction data in blocks that enable local updates while maintaining neighborhood consistency; deriving occlusion data from a combination of one or more depth data sources; providing occlusion data to an application that generates such a scene; and / or performing multi-stage mesh simplification. The reconstruction may contain data of different levels of sophistication, including, for example, raw data such as live depth data, fused volumetric data such as voxels, and calculated data such as meshes.

[0192] In some embodiments, components of the passable world model may be distributed, with some parts running locally on the XR device and some parts running remotely, such as on a network connected to a server or otherwise in the cloud. The distribution of information processing and storage between the local XR device and the cloud can affect the functionality and user experience of the XR system. For example, reducing processing on the local device by distributing processing to the cloud may enable longer battery life and reduce heat generated on the local device. However, distributing much more processing to the cloud may create undesirable latency that causes an unacceptable user experience.

[0193] FIG. 6B depicts a distributed component architecture 600 configured for spatial computing, according to some embodiments. The distributed component architecture 600 may include a passable world component 602 (e.g., PW 538 in FIG. 5A ), a Lumin OS 604, an API 606, an SDK 608, and applications 610. The Lumin OS 604 may include a Linux-based kernel with custom drivers compatible with XR devices. The API 606 may include an application programming interface that gives XR applications (e.g., applications 610) access to the spatial computing features of the XR device. The SDK 608 may include a software development kit that enables the creation of XR applications.

[0194] One or more components in architecture 600 may create and maintain a model of the passable world. In this example, sensor data is collected on the local device. Processing of that sensor data may be performed partially locally on the XR device and partially in the cloud. PW 538 may include an environment map created at least in part based on data captured by AR devices worn by multiple users. During a session of an AR experience, individual AR devices (such as the wearable devices described above in connection with FIG. 4) may create a tracking map, which is one type of map.

[0195] In some embodiments, a device may include components that build both a sparse map and a dense map. The tracking map may serve as a sparse map and may include information about the head pose of the AR device scanning the environment and the objects detected in the environment at each head pose. These head poses may be maintained locally per device. For example, the head pose on each device may be relative to the initial head pose when the device is turned on for the session. As a result, each tracking map may be local to the device that creates it. The dense map may include surface information, which may be represented by mesh or depth information. Alternatively, or in addition, the dense map may include higher-level information derived from surface or depth information, such as the location and / or properties of planes and / or other objects.

[0196] The creation of the dense map may be independent from the creation of the sparse map in some embodiments. The creation of the dense map and the sparse map may be performed, for example, in separate processing pipelines within the AR system. Separating the processing may allow, for example, the generation or processing of different types of maps to be performed at different rates. The sparse map may, for example, be refreshed at a faster rate than the dense map. However, in some embodiments, the processing of the dense and sparse maps may be related even when performed in different pipelines. Changes in the physical world revealed in the sparse map may, for example, trigger an update of the dense map, or vice versa. Furthermore, even when created independently, the maps may be used together. For example, a coordinate system derived from the sparse map may be used to define the position and / or orientation of objects in the dense map.

[0197] The sparse map and / or the dense map may be persisted for reuse by the same device and / or for sharing with other devices. Such persistence may be achieved by storing the information in a cloud. The AR device may send the tracking map to the cloud and merge it with an environment map selected from persisted maps previously stored in the cloud, for example. In some embodiments, the selected persisted map may be sent from the cloud to the AR device for merging. In some embodiments, the persisted map may be oriented with respect to one or more persistent coordinate frames. Such maps may serve as reference maps because they may be used by any of multiple devices. In some embodiments, a model of the passable world may include or be created with one or more reference maps. A device may perform some operations based on a coordinate frame local to the device, but may use the reference map by determining a transformation between its coordinate frame local to the device and the reference map.

[0198] The reference map may originate as a tracking map (TM) (e.g., TM 1102 in FIG. 31A ), which may be promoted to a reference map. The reference map may be persisted so that a device accessing the reference map, once it has determined the transformation between its local coordinate system and the coordinate system of the reference map, can use the information in the reference map to determine the location of objects represented in the reference map in the physical world around the device. In some embodiments, the TM may be a head pose sparse map created by an XR device. In some embodiments, the reference map may be created when an XR device sends one or more TMs to a cloud server to be merged with additional TMs captured by the XR device at different times or by other XR devices.

[0199] A reference map or other map may provide information about the portion of the physical world represented by the processed data to create the individual map. FIG. 7 depicts an example tracking map 700, according to some embodiments. The tracking map 700 may provide a top view 706 of a physical object in the corresponding physical world represented by points 702. In some embodiments, the map points 702 may represent features of the physical object, which may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. Features may be derived from processed images, such as those obtained using sensors in a wearable device in an augmented reality system. Features may be derived, for example, by processing image frames output by the sensors and identifying features based on large gradients in the images or other suitable criteria. Further processing may limit the number of features in each frame. For example, processing may select features that are likely to represent persistent objects. One or more heuristics may be applied for this selection.

[0200] The tracking map 700 may include data about points 702 collected by the device. A pose may be stored for each image frame with data points included in the tracking map. The pose may represent the orientation from which the image frame was captured so that feature points in each image frame can be spatially correlated. The pose may be determined by positioning information, such as may be derived from a sensor on the wearable device, such as an IMU sensor. Alternatively, or in addition, the pose may be determined from matching an image frame with another image frame that depicts an overlapping portion of the physical world. By finding such positional correlation, which may be accomplished by matching subsets of feature points in the two frames, a relative pose between the two frames may be calculated. Relative pose may be appropriate for the tracking map because the map may be relative to a coordinate system local to the device, which is established based on the device's initial pose when construction of the tracking map began.

[0201] Because much of the information collected using the sensors is likely redundant, not all of the feature points and image frames collected by the device may be retained as part of the tracking map. Rather, only certain frames may be added to the map. These frames may be selected based on one or more criteria, such as the degree of overlap with image frames already in the map, the number of new features they contain, or a quality metric for the features in the frame. Image frames that are not added to the tracking map may be discarded or used to revise feature locations. As a further alternative, all or most of the image frames, represented as a set of features, may be retained, but a subset of those frames may be designated as key frames, which are used for further processing.

[0202] The keyframes may be processed to produce a keyrig 704. The keyframes may be processed to produce a three-dimensional set of feature points and stored as the keyrig 704. Such processing may involve, for example, comparing image frames derived simultaneously from two cameras and stereoscopically determining the 3D positions of the feature points. Metadata such as pose may be associated with these keyframes and / or keyrigs.

[0203] The environment map may have any of a number of formats, depending on, for example, the storage location of the environment map, including, for example, local storage and remote storage of the AR device. For example, a map in remote storage may have a higher resolution than a map in local storage on the wearable device when memory is limited. To transmit a higher resolution map from remote storage to local storage, the map may be downsampled or otherwise converted to a suitable format, such as by reducing the number of poses per area of ​​the physical world stored in the map and / or the number of feature points stored per pose. In some embodiments, a slice or portion of the high-resolution map from remote storage may be transmitted to local storage, and the slice or portion is not downsampled.

[0204] The database of environment maps may be updated as new tracking maps are created. To determine which of the potentially numerous environment maps in the database should be updated, updating may include efficiently selecting one or more environment maps stored in the database associated with the new tracking map. The selected one or more environment maps may be ranked by relevance, and one or more of the highest-ranked maps may be selected for processing to merge the new tracking map with the higher-ranked selected environment map and create one or more updated environment maps. When a new tracking map represents a portion of the physical world over which there is no existing environment map to update, the tracking map may be stored in the database as a new environment map.

[0205] View Independent Display

[0206] Described herein are methods and apparatus for providing virtual content using an XR system, independent of the location of the eyes viewing the virtual content. Traditionally, virtual content is re-rendered in response to any movement of the display system. For example, when a user wearing a display system views a virtual representation of a three-dimensional (3D) object on the display and walks around the area in which the 3D object appears, the 3D object should be re-rendered for each viewpoint so that the user has the perception that they are walking around the object occupying real space. However, re-rendering consumes significant computing resources of the system and introduces artifacts due to latency.

[0207] The inventors have recognized and appreciated that head pose (e.g., the location and orientation of a user wearing an XR system) can be used to render virtual content independent of eye rotations of the user's head. In some embodiments, a dynamic map of a scene may be generated based on multiple coordinate frames in real space across one or more sessions so that virtual content interacting with the dynamic map can be robustly rendered independent of eye rotations of the user's head and / or independent of sensor deformations caused by, for example, heat generated during high-speed, computationally intensive operations. In some embodiments, the configuration of multiple coordinate frames may enable a first XR device worn by a first user and a second XR device worn by a second user to recognize a common location within a scene. In some embodiments, the configuration of multiple coordinate frames may enable users wearing the XR devices to view virtual content within the same location within a scene.

[0208] In some embodiments, the tracking map may be constructed in a world coordinate frame, which may have a world origin. The world origin may be the initial pose of the XR device when the XR device is powered on. The world origin may be aligned to gravity, allowing XR application developers to obtain gravity alignment without extra work. Different tracking maps may be constructed in different world coordinate frames because the tracking maps may be captured by the same XR device in different sessions and / or by different XR devices worn by different users. In some embodiments, an XR device session may start when the device is powered on and continue until it is powered off. In some embodiments, the XR device may have a head coordinate frame, which may have a head origin. The head origin may be the current pose of the XR device when the image is captured. The difference between the head poses in the world coordinate frame and the head coordinate frame may be used to estimate the tracking route.

[0209] In some embodiments, the XR device may have a camera coordinate frame, which may have a camera origin. The camera origin may be the current pose of one or more sensors of the XR device. The inventors have recognized and appreciated that configuring the camera coordinate frame enables robust display of virtual content independent of eye rotation in the user's head. This configuration also enables robust display of virtual content independent of sensor deformations due to, for example, heat generated during movement.

[0210] In some embodiments, the XR device may have a head unit with a head-mountable frame that a user can fasten to their head and that may include two waveguides, one in front of each of the user's eyes. The waveguides may be transparent so that ambient light from real-world objects can pass through the waveguides, allowing the user to see the real-world objects. Each waveguide may transmit projected light from a projector to a respective eye of the user. The projected light may form an image on the retina of the eye. The retina of the eye therefore receives the ambient light and the projected light. The user may simultaneously see the real-world object and one or more virtual objects created by the projected light. In some embodiments, the XR device may have sensors that detect real-world objects around the user. These sensors may be, for example, cameras that capture images that can be processed to identify the location of the real-world objects.

[0211] In some embodiments, the XR system may assign coordinate frames to virtual content, as opposed to tying the virtual content to a world coordinate frame. Such a configuration allows virtual content to be described regardless of where it is rendered for the user, but may be tied to a more persistent frame position and rendered at a defined location, such as the persistent coordinate frame (PCF) described in connection with FIGS. 14-20C. As the location of an object changes, the XR device may detect the change in the environment map and determine the movement of a head unit worn by the user relative to the real-world object.

[0212] 8 illustrates a user experiencing virtual content as rendered in a physical environment by an XR system 10, according to some embodiments. The XR system may include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is present in the physical environment with a real object in the form of a table 16.

[0213] In the illustrated example, the first XR device 12.1 includes a head unit 22, a belt pack 24, and a cable connection 26. The first user 14.1 wears the head unit 22 on their head and the belt pack 24, remote from the head unit 22, on their waist. The cable connection 26 connects the head unit 22 to the belt pack 24. The head unit 22 includes technology used to display a virtual object or objects to the first user 14.1 while still allowing the first user 14.1 to see real objects, such as a table 16. The belt pack 24 primarily contains the processing and communication capabilities of the first XR device 12.1. In some embodiments, the processing and communication capabilities may reside, in whole or in part, within the head unit 22, such that the belt pack 24 may be removed or located in another device, such as a backpack.

[0214] In the illustrated embodiment, the belt pack 24 is connected to the network 18 via a wireless connection. A server 20 is connected to the network 18 and holds data representing local content. The belt pack 24 downloads the data representing the local content from the server 20 via the network 18. The belt pack 24 provides the data to the head unit 22 via a cable connection 26. The head unit 22 may include a display having a light source, for example, a laser light source or a light emitting diode (LED) light source, and a waveguide for guiding the light.

[0215] In some embodiments, the first user 14.1 may wear a head unit 22 on his or her head and a belt pack 24 on his or her waist. The belt pack 24 may download image data representing virtual content from a server 20 via a network 18. The first user 14.1 may view a table 16 through a display on the head unit 22. A projector forming part of the head unit 22 may receive the image data from the belt pack 24 and generate light based on the image data. The light may travel through one or more waveguides forming part of the display of the head unit 22. The light may then exit the waveguide and propagate onto the retina of the first user's 14.1 eye. The projector may generate light in a pattern that is replicated on the retina of the first user's 14.1 eye. The light that strikes the retina of the first user's 14.1 eye may have a selected depth of field so that the first user 14.1 perceives an image at a preselected depth behind the waveguide. Additionally, the first user's 14.1 eyes may receive slightly different images such that the first user's 14.1 brain perceives a three-dimensional image or multiple images at a selected distance from the head unit 22. In the illustrated embodiment, the first user 14.1 perceives the virtual content 28 above the table 16. The virtual content 28 and its location and relative distance from the first user 14.1 are determined by data representing the virtual content 28 and the various coordinate frames used to display the virtual content 28 to the first user 14.1.

[0216] In the illustrated example, the virtual content 28 is invisible from the perspective of the drawing and is visible to the first user 14.1 through use of the first XR device 12.1. The virtual content 28 may initially reside as a data structure within the visual data and algorithms within the belt pack 24. The data structure may then be revealed as light when the projector of the head unit 22 generates light based on the data structure. It should be understood that although the virtual content 28 does not exist in three-dimensional space in front of the first user 14.1, the virtual content 28 is still represented in FIG. 1 in three-dimensional space for illustration of what would be perceived by a wearer of the head unit 22. Visualization of computer data in three-dimensional space may be used in this description to illustrate how the data structures that facilitate the rendering perceived by one or more users are interrelated among the data structures within the belt pack 24.

[0217] 9 illustrates components of a first XR device 12.1, according to some embodiments. The first XR device 12.1 may include a head unit 22 and various components that form part of the visual data and algorithms, including, for example, a rendering engine 30, various coordinate systems 32, various origin and destination coordinate frames 34, and various origin / destination coordinate frame converters 36. The various coordinate systems may be based on intrinsic properties of the XR device or may be determined by reference to other information, such as a persistent pose or persistent coordinate system as described herein.

[0218] The head unit 22 may include a head-mountable frame 40 , a display system 42 , a real object detection camera 44 , a motion tracking camera 46 , and an inertial measurement unit 48 .

[0219] The head-mountable frame 40 may have a shape that is attachable to the head of the first user 14.1 in Figure 8. The display system 42, the real object detection camera 44, the mobile tracking camera 46, and the inertial measurement unit 48 are mounted on the head-mountable frame 40 and may therefore move with the head-mountable frame 40.

[0220] The coordinate system 32 may include a local data system 52 , a world frame system 54 , a head frame system 56 , and a camera frame system 58 .

[0221] Local data system 52 may include a data channel 62, a local frame determination routine 64, and local frame storage instructions 66. Data channel 62 may be an internal software routine, a hardware component such as an external cable or radio frequency receiver, or a hybrid component such as an open port. Data channel 62 may be configured to receive image data 68 representing virtual content.

[0222] A local frame determination routine 64 may be connected to the data channel 62. The local frame determination routine 64 may be configured to determine a local coordinate frame 70. In some embodiments, the local frame determination routine may determine the local coordinate frame based on a real-world object or real-world location. In some embodiments, the local coordinate frame may be based on the top edge relative to the bottom edge of the browser window, the head or feet of a character, a node on the outer surface of a prism or bounding box that encloses the virtual content, or any other suitable location for placing a coordinate frame that defines the facing direction of the virtual content and where the virtual content should be placed (e.g., a node such as a placement node or PCF node).

[0223] The local frame storage instructions 66 may be connected to the local frame determination routine 64. Those skilled in the art will understand that software modules and routines are "connected" to one another through subroutines, calls, etc. The local frame storage instructions 66 may store the local coordinate frame 70 as the local coordinate frame 72 within the origin and destination coordinate frame 34. In some embodiments, the origin and destination coordinate frame 34 may be one or more coordinate frames that can be manipulated or transformed for virtual content to persist between sessions. In some embodiments, a session may be the time period between boot-up and shutdown of an XR device. Two sessions may be two startup and shutdown cycles for a single XR device, or startup and shutdown for two different XR devices.

[0224] In some embodiments, the origin and destination coordinate frames 34 may be coordinate frames involved in one or more transformations required for the first user's XR device and the second user's XR device to perceive a common location. In some embodiments, the destination coordinate frame may be the output of a series of calculations and transformations applied to the target coordinate frames in order for the first and second users to view virtual content at the same location.

[0225] Rendering engine 30 may be connected to data channel 62. Rendering engine 30 may receive image data 68 from data channel 62 such that rendering engine 30 may render virtual content based, at least in part, on image data 68.

[0226] A display system 42 may be connected to the rendering engine 30. The display system 42 may include components that convert image data 68 into visible light. The visible light may form two patterns, one for each eye. The visible light may be incident on the eye of the first user 14.1 in FIG. 8 and may be detected on the retina of the eye of the first user 14.1.

[0227] The real object detection camera 44 may include one or more cameras that may capture images from different sides of the head-mountable frame 40. The mobile tracking camera 46 may include one or more cameras that capture images on a side of the head-mountable frame 40. One set of one or more cameras may be used instead of two sets of one or more cameras representing the real object detection camera 44 and the mobile tracking camera 46. In some embodiments, the cameras 44, 46 may capture images. As explained above, these cameras may collect data that is used to build a tracking map.

[0228] The inertial measurement unit 48 may include several devices used to detect movement of the head unit 22. The inertial measurement unit 48 may include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of the inertial measurement unit 48 combine to track movement of the head unit 22 in at least three orthogonal directions and approximately at least three orthogonal axes.

[0229] In the illustrated embodiment, the world frame system 54 includes a world surface determination routine 78, a world frame determination routine 80, and world frame storage instructions 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 receives images and / or keyframes based on images captured by the real object detection camera 44, processes the images, and identifies surfaces within the images. A depth sensor (not shown) may determine the distance to the surface. The surface is therefore represented by data in three dimensions, including its size, shape, and distance from the real object detection camera.

[0230] In some embodiments, the world coordinate frame 84 may be based on the origin at the time of initialization of the head pose session. In some embodiments, the world coordinate frame may be located where the device was booted up, or may be some new location if the head pose is lost during the boot session. In some embodiments, the world coordinate frame may be the origin at the start of the head pose session.

[0231] In the illustrated embodiment, the world frame determination routine 80 is connected to the world surface determination routine 78 and determines a world coordinate frame 84 based on the location of the surface as determined by the world surface determination routine 78. The world frame storage instruction 82 is connected to the world frame determination routine 80 and receives the world coordinate frame 84 from the world frame determination routine 80. The world frame storage instruction 82 stores the world coordinate frame 84 as a world coordinate frame 86 in the origin and destination coordinate frame 34.

[0232] The head frame system 56 may include a head frame determination routine 90 and head frame storage instructions 92. The head frame determination routine 90 may be connected to the mobile tracking camera 46 and the inertial measurement unit 48. The head frame determination routine 90 may use data from the mobile tracking camera 46 and the inertial measurement unit 48 to calculate a head coordinate frame 94. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The mobile tracking camera 46 may continuously capture images that are used by the head frame determination routine 90 to refine the head coordinate frame 94. The head unit 22 moves as the first user 14.1 in FIG. 8 moves his or her head. The mobile tracking camera 46 and the inertial measurement unit 48 may continuously provide data to the head frame determination routine 90 so that the head frame determination routine 90 may update the head coordinate frame 94.

[0233] The head frame storage instruction 92 may be coupled to the head frame determination routine 90 and may receive a head coordinate frame 94 from the head frame determination routine 90. The head frame storage instruction 92 may store the head coordinate frame 94 as a head coordinate frame 96 between the origin and destination coordinate frames 34. The head frame storage instruction 92 may repeatedly store the updated head coordinate frame 94 as the head coordinate frame 96 when the head frame determination routine 90 recalculates the head coordinate frame 94. In some embodiments, the head coordinate frame may be the location of the wearable XR device 12.1 relative to the local coordinate frame 72.

[0234] The camera frame system 58 may include camera specific properties 98. The camera specific properties 98 may include the dimensions of the head unit 22, which are characteristic of its design and manufacture. The camera specific properties 98 may be used to calculate a camera coordinate frame 100, which is stored in the origin and destination coordinate frame 34.

[0235] In some embodiments, camera coordinate frame 100 may include all pupil positions for the left eye of first user 14.1 in Figure 8. As the left eye moves from left to right or up and down, the pupil positions for the left eye are located within camera coordinate frame 100. Additionally, the pupil positions for the right eye are located within camera coordinate frame 100 for the right eye. In some embodiments, camera coordinate frame 100 may include the location of the camera relative to a local coordinate frame when an image is taken.

[0236] The origin / destination coordinate frame converter 36 may include a local / world coordinate converter 104, a world / head coordinate converter 106, and a head / camera coordinate converter 108. The local / world coordinate converter 104 may receive the local coordinate frame 72 and transform the local coordinate frame 72 to the world coordinate frame 86. The transformation of the local coordinate frame 72 to the world coordinate frame 86 may be expressed as a local coordinate frame that is transformed within the world coordinate frame 86 to the world coordinate frame 110.

[0237] The world / head coordinate converter 106 may transform from the world coordinate frame 86 to the head coordinate frame 96. The world / head coordinate converter 106 may transform a local coordinate frame that is transformed into the world coordinate frame 110 to the head coordinate frame 96. The transformation may be expressed as a local coordinate frame that is transformed into the head coordinate frame 112 within the head coordinate frame 96.

[0238] The head / camera coordinate converter 108 may convert from the head coordinate frame 96 to the camera coordinate frame 100. The head / camera coordinate converter 108 may convert the local coordinate frame, which is converted to the head coordinate frame 112, to a local coordinate frame, which is converted to the camera coordinate frame 114 in the camera coordinate frame 100. The local coordinate frame, which is converted to the camera coordinate frame 114, may be imported into the rendering engine 30. The rendering engine 30 may render image data 68 representing the local content 28 based on the local coordinate frame, which is converted to the camera coordinate frame 114.

[0239] FIG. 10 is a spatial representation of various origin and destination coordinate frames 34. Represented in the figure are local coordinate frame 72, world coordinate frame 86, head coordinate frame 96, and camera coordinate frame 100. In some embodiments, the local coordinate frame associated with XR content 28 may have a position and rotation relative to the local and / or world coordinate frame and / or PCF (e.g., may provide nodal and facing direction) when the virtual content is placed in the real world and thus may be viewed by a user. Each camera may have its own camera coordinate frame 100 that encompasses all pupil positions for one eye. Reference numerals 104A and 106A represent the transformations performed by local-to-world coordinate converter 104, world-to-head coordinate converter 106, and head-to-camera coordinate converter 108 in FIG. 9, respectively.

[0240] 11 depicts a camera rendering protocol for converting from a head coordinate frame to a camera coordinate frame, according to some embodiments. In the illustrated example, the pupil for one eye moves from position A to B. A virtual object intended to appear stationary would be projected onto the depth plane at one of two positions A or B, depending on the position of the pupil (assuming the camera is configured to use as a pupil-based coordinate frame). As a result, using a pupil coordinate frame that is converted to the head coordinate frame would cause jitter in the stationary virtual object as the eye moves from position A to position B. This situation is referred to as view-dependent display or projection.

[0241] As depicted in Figure 12, the camera coordinate frame (e.g., CR) is positioned to encompass all pupil positions, but the object projection will now be consistent regardless of pupil positions A and B. The head coordinate frame transforms to the CR frame, which is referred to as view-independent display or projection. Image reprojection may be applied to the virtual content to account for changes in eye position; however, because the rendering is still in the same position, jitter is minimized.

[0242] 13 illustrates in more detail the display system 42. The display system 42 includes a stereoscopic analyzer 144 that is connected to the rendering engine 30 and forms part of the visual data and algorithms.

[0243] Display system 42 further includes left and right projectors 166A and 166B and left and right waveguides 170A and 170B. Left and right projectors 166A and 166B are connected to a power supply. Each projector 166A and 166B has a separate input through which image data is provided to the individual projector 166A or 166B. When powered, the individual projector 166A or 166B generates and emits light in a two-dimensional pattern. Left and right waveguides 170A and 170B are positioned to receive light from left and right projectors 166A and 166B, respectively. Left and right waveguides 170A and 170B are transparent waveguides.

[0244] In use, a user mounts head-mountable frame 40 on their head. Components of head-mountable frame 40 may include, for example, a strap (not shown) that wraps around the back of the user's head. Left and right waveguides 170A and 170B are then positioned in front of the user's left and right eyes 220A and 220B.

[0245] The rendering engine 30 feeds the image data it receives into the stereoscopic analyzer 144. The image data is three-dimensional image data of the local content 28 in FIG. 8. The image data is projected onto multiple virtual planes. The stereoscopic analyzer 144 analyzes the image data and determines left and right image data sets based on the image data for projection onto each depth plane. The left and right image data sets are data sets representing two-dimensional images that are projected in three dimensions to give the user the perception of depth.

[0246] Stereoscopic analyzer 144 captures left and right image data sets into left and right projectors 166A and 166B. Left and right projectors 166A and 166B then create left and right light patterns. While the components of display system 42 are shown in plan view, it should be understood that the left and right patterns are two-dimensional patterns when shown in a front elevation view. Each light pattern includes multiple pixels. For illustrative purposes, light rays 224A and 226A from two of the pixels are shown exiting left projector 166A and entering left waveguide 170A. Light rays 224A and 226A reflect from the sides of left waveguide 170A. Although rays 224A and 226A are shown propagating from left to right within left waveguide 170A through internal reflection, it should be understood that rays 224A and 226A also propagate into the page using refractive and reflective systems.

[0247] Light rays 224A and 226A exit left optical waveguide 170A through pupil 228A and then enter left eye 220A through pupil 230A of left eye 220A. Light rays 224A and 226A then impinge on retina 232A of left eye 220A. In this manner, the left light pattern impinges on retina 232A of left eye 220A. The user is given the perception that the pixels formed on retina 232A are pixels 234A and 236A, which the user perceives as being at a distance on the side of left waveguide 170A facing left eye 220A. Depth perception is created by manipulating the focal length of the light.

[0248] Similarly, stereoscopic analyzer 144 captures a right image data set into right projector 166B. Right projector 166B transmits a right light pattern, which is represented by pixels in the form of light rays 224B and 226B. Light rays 224B and 226B reflect within right waveguide 170B and exit through pupil 228B. Light rays 224B and 226B then enter through pupil 230B of right eye 220B and impinge on retina 232B of right eye 220B. The pixels of light rays 224B and 226B are perceived as pixels 134B and 236B behind right waveguide 170B.

[0249] The patterns created on the retinas 232A and 232B are perceived as left and right images, respectively, which differ slightly from each other due to the function of the stereoscopic analyzer 144. The left and right images are perceived as three-dimensional renderings in the user's brain.

[0250] As mentioned, the left and right waveguides 170A and 170B are transparent. Light from a real object, such as a table 16, on the side of the left and right waveguides 170A and 170B facing the eyes 220A and 220B can be projected through the left and right waveguides 170A and 170B and impinge on the retinas 232A and 232B.

[0251] Persistent Coordinate Frame (PCF)

[0252] Described herein are methods and apparatus for providing spatial persistence across user instances in a shared space. Without spatial persistence, virtual content placed in the physical world by a user in a session may not be present or may be misplaced in the view of a user in a different session. Without spatial persistence, virtual content placed in the physical world by one user may not be present or may be misplaced in the view of a second user, even if the second user intends to share the experience of the same physical space with the first user.

[0253] The inventors recognize and appreciate that spatial persistence may be provided through a Persistent Coordinate Frame (PCF). A PCF may be defined based on one or more points that represent recognized features (e.g., corners, edges) in the physical world. Features may be selected such that they are likely to be identical from one user instance to another user instance of the XR system.

[0254] Furthermore, drift during tracking, which may cause the calculated tracking path (e.g., camera trajectory) to deviate from the actual tracking path, may cause the location of virtual content to appear out of place when rendered relative to a local map based solely on the tracking map. The tracking map for the space may be refined to correct for drift as the XR device gathers more information about the scene over time. However, if virtual content is placed on real objects and stored relative to the device's world coordinate frame derived from the tracking map before map refinement, the virtual content may appear displaced as if the real objects had moved during map refinement. The PCF may be updated according to map refinement because the PCF is defined based on features and is updated as features move during map refinement.

[0255] A PCF may have six degrees of freedom, including translation and rotation relative to a map coordinate system. The PCF may be stored in a local and / or remote storage medium. The translation and rotation of the PCF may be calculated relative to the map coordinate system, for example, depending on the storage location. For example, a PCF used locally by a device may have translation and rotation relative to the device's world coordinate frame. A PCF in the cloud may have translation and rotation relative to the reference coordinate frame of a reference map.

[0256] PCFs provide a sparse representation of the physical world and may provide less than all of the available information about the physical world so that they can be efficiently processed and transferred. Techniques for processing persistent spatial information may include creating a dynamic map based on one or more coordinate systems in real space across one or more sessions and generating a persistent coordinate frame (PCF) over the sparse map that can be exposed to XR applications, for example, via an application programming interface (API).

[0257] 14 is a block diagram illustrating the creation of a Persistent Coordinate Frame (PCF) and the association of XR content with the PCF, according to some embodiments. Each block may represent digital information stored in computer memory. In the case of an application 1180, the data may represent computer-executable instructions. In the case of virtual content 1170, the digital information may, for example, define a virtual object as defined by the application 1180. In the case of other boxes, the digital information may characterize some aspect of the physical world.

[0258] In the illustrated embodiment, one or more PCFs are created from images captured using sensors on the wearable device. In the embodiment of FIG. 14, the sensors are visual image cameras. These cameras may be the same cameras used to form the tracking map. Thus, some of the processing suggested by FIG. 14 may be performed as part of updating the tracking map. However, FIG. 14 illustrates that information providing persistence is generated in addition to the tracking map.

[0259] To derive a 3D PCF, two images 1110 from two cameras mounted on a wearable device in a configuration that enables stereoscopic image analysis are processed together. Figure 14 illustrates image 1 and image 2, each derived from one of the cameras. A single image from each camera is shown for convenience. However, each camera may output a stream of image frames, and the processing illustrated in Figure 14 may be performed for multiple image frames in the stream.

[0260] Thus, Image 1 and Image 2 may each be one frame in a sequence of image frames. The process as depicted in FIG. 14 may be repeated on successive image frames in the sequence until an image frame containing feature points that provide a suitable image from which to form persistent spatial information is processed. Alternatively, or in addition, the process of FIG. 14 may be repeated as the user moves such that the user is no longer sufficiently close to a previously identified PCF to reliably use that PCF to determine location relative to the physical world. For example, the XR system may maintain a current PCF for the user. If the distance exceeds a threshold, the system may switch to a new current PCF closer to the user, which may be generated according to the process of FIG. 14 using image frames acquired at the user's current location.

[0261] Even when generating a single PCF, the stream of image frames may be processed to identify image frames that depict content in the physical world that is likely to be stable and that can be easily identified by devices in the vicinity of the area of ​​the physical world depicted in the image frames. In the embodiment of Figure 14, the process begins with identifying features 1120 in the image. Features may be identified, for example, by finding locations of gradients or other characteristics in the image that exceed a threshold, which may correspond, for example, to corners of an object. In the illustrated embodiment, the features are points, although other recognizable features, such as edges, may alternatively or additionally be used.

[0262] In the illustrated embodiment, a fixed number N of features 1120 are selected for further processing. These feature points may be selected based on one or more criteria, such as gradient magnitude or proximity to other feature points. Alternatively, or in addition, feature points may be selected heuristically, such as based on characteristics that suggest the feature point is persistent. For example, a heuristic may be defined based on a feature point's characteristics that are likely to correspond to a window or door or the corner of a large piece of furniture. Such a heuristic may consider the feature point itself and its surroundings. As a specific example, the number of feature points per image may be 100-500, or 150-250, such as 200.

[0263] Regardless of the number of selected feature points, a descriptor 1130 may be calculated for the feature points. In this embodiment, a descriptor is calculated for each selected feature point, but descriptors may also be calculated for a group of feature points, a subset of feature points, or all features in an image. The descriptors characterize the feature points so that feature points that represent the same object in the physical world are assigned similar descriptors. The descriptors may facilitate matching two frames, such as may occur when one map is localized relative to another. Rather than searching for a relative orientation of the frames that minimizes the distance between feature points in the two images, initial matching of the two frames may be performed by identifying feature points with similar descriptors. Matching of image frames may be based on matching points with similar descriptors, which may involve less processing than calculating matching for all feature points in an image.

[0264] The descriptor may be calculated as a mapping of the descriptor to the feature point, or in some embodiments, a mapping of a patch of the image surrounding the feature point. The descriptor may be a numerical quantity. U.S. Patent Application No. 16 / 190,948 describes calculating descriptors for feature points and is incorporated herein by reference in its entirety.

[0265] In the example of FIG. 14 , a descriptor 1130 is calculated for each feature point in each image frame. Based on the descriptors and / or the feature points and / or the image itself, image frames may be identified as keyframes 1140. In the illustrated embodiment, keyframes are image frames that meet certain criteria and are then selected for further processing. When creating a tracking map, for example, image frames that add meaningful information to the map may be selected as keyframes to be integrated into the map. On the other hand, image frames that substantially overlap regions over which image frames have already been integrated into the map may be discarded so that they do not become keyframes. Alternatively, or in addition, keyframes may be selected based on the number and / or type of feature points in the image frames. In the embodiment of FIG. 14 , keyframes 1150 selected for inclusion in the tracking map may also be processed as keyframes for determining PCFs, although different or additional criteria for selecting keyframes for generation of PCFs may be used.

[0266] While Figure 14 shows that keyframes are used for further processing, information obtained from images may be processed in other forms. For example, feature points, such as in a key rig, may alternatively or additionally be processed. Furthermore, while keyframes are described as being derived from a single image frame, there need not be a one-to-one relationship between the keyframes and the obtained image frames. Keyframes may also be obtained from multiple image frames, such as by stitching or aggregating the image frames together, such that only features that appear in multiple images are retained in the keyframe.

[0267] A keyframe may include image information and / or metadata associated with the image information. In some embodiments, images captured by cameras 44, 46 (FIG. 9) may be computed into one or more keyframes (e.g., keyframe 1, 2). In some embodiments, a keyframe may include a camera pose. In some embodiments, a keyframe may include one or more camera images captured at a camera pose. In some embodiments, the XR system may determine that a portion of a camera image captured at a camera pose is not useful and therefore not include that portion in a keyframe. Thus, using keyframes to match new images with earlier knowledge of the scene reduces the use of computational resources of the XR system. In some embodiments, a keyframe may include images and / or image data at a location with a certain direction / angle. In some embodiments, a keyframe may include a location and direction from which one or more map points can be observed. In some embodiments, a keyframe may include a coordinate frame with an ID. US Patent Application No. 15 / 877,359 describes keyframes and is incorporated herein by reference in its entirety.

[0268] Some or all of the keyframes 1140 may be selected for further processing, such as generating persistent poses 1150 for the keyframes. The selection may be based on characteristics of all or a subset of the feature points in the image frames. These characteristics may be determined from processing the descriptors, features, and / or the image frames themselves. As a specific example, the selection may be based on a cluster of feature points identified as likely to be associated with a persistent object.

[0269] Each keyframe is associated with the pose of the camera at which the keyframe was acquired. For keyframes selected for processing into persistent poses, that pose information may be stored along with other metadata about the keyframe, such as a WiFi fingerprint and / or GPS coordinates at the time and / or location of acquisition. In some embodiments, metadata, e.g., GPS coordinates, may be used individually or in combination as part of the localization process.

[0270] A persistent pose is a source of information that a device can use to orient itself relative to previously obtained information about the physical world. For example, if a keyframe from which a persistent pose was created is incorporated into a map of the physical world, the device may orient itself relative to that persistent pose using a sufficient number of feature points in the keyframe associated with the persistent pose. The device may match the persistent pose with a current image obtained of its surroundings. This matching may be based on matching the current image with the image 1110, features 1120, and / or descriptors 1130 that gave rise to the persistent pose, or any subset of that image or their features or descriptors. In some embodiments, the current image frame matched to the persistent pose may be another keyframe incorporated into the device's tracking map.

[0271] Information about persistent attitudes may be stored in a format that facilitates sharing between multiple applications, which may be running on the same or different devices. In the example of FIG. 14, some or all of the persistent attitudes may be reflected as persistent coordinate frames (PCFs) 1160. Like persistent attitudes, PCFs may also be associated with a map and may comprise a set of features or other information that a device may use to determine its orientation relative to the PCF. A PCF may include a transform that defines its transformation relative to the origin of the map, so that a device may determine its position relative to any object in the physical world reflected in the map by correlating its position with the PCF.

[0272] Because PCFs provide a mechanism for determining location relative to physical objects, an application, such as application 1180, may define the location of a virtual object relative to one or more PCFs that serve as anchors for virtual content 1170. FIG. 14 illustrates, for example, app 1 associating PCF 1.2 with its virtual content 2. Similarly, app 2 associates PCF 1.2 with its virtual content 3. App 1 is also shown associating PCF 4.5 with its virtual content 1, and app 2 is shown associating PCF 3 with its virtual content 4. In some embodiments, PCF 3 may be based on image 3 (not shown), and PCF 4.5 may be based on image 4 and image 5 (not shown), similar to how PCF 1.2 is based on image 1 and image 2. When rendering this virtual content, a device may apply one or more transforms to calculate information such as the location of the virtual content relative to the device's display and / or the location of the physical object relative to the virtual content's desired location. Using PCFs as a reference may simplify such calculations.

[0273] In some embodiments, a persistent pose may be a coordinate location and / or orientation with one or more associated keyframes. In some embodiments, a persistent pose may be created automatically after a user travels a certain distance, for example, 3 meters. In some embodiments, a persistent pose may act as a reference point during localization. In some embodiments, a persistent pose may be stored in the passable world (e.g., passable world module 538).

[0274] In some embodiments, a new PCF may be determined based on a predefined distance allowed between adjacent PCFs. In some embodiments, one or more sustained postures may be calculated into the PCF once the user has traveled a predetermined distance, e.g., 5 meters. In some embodiments, the PCF may be associated with one or more world coordinate frames and / or reference coordinate frames, e.g., within a passable world. In some embodiments, the PCF may be stored in a local and / or remote database, e.g., depending on security settings.

[0275] 15 illustrates a method 4700 of establishing and using a persistent coordinate frame, according to some embodiments. Method 4700 may begin with capturing images (e.g., image 1 and image 2 in FIG. 14) centered on a scene using one or more sensors of an XR device (act 4702). Multiple cameras may be used, and one camera may generate multiple images, for example, in a stream.

[0276] Method 4700 may include extracting points of interest (e.g., map points 702 in FIG. 7 , features 1120 in FIG. 14 ) from the captured image (4704), generating descriptors (e.g., descriptors 1130 in FIG. 14 ) for the extracted points of interest (act 4706), and generating keyframes (e.g., keyframe 1140) based on the descriptors (act 4708). In some embodiments, the method may compare points of interest in the keyframes and form pairs of keyframes that share a predetermined amount of points of interest. The method may use each pair of keyframes to reconstruct a portion of the physical world. The mapped portions of the physical world may be saved as 3D features (e.g., keyrig 704 in FIG. 7 ). In some embodiments, selected portions of the pair of keyframes may be used to construct the 3D features. In some embodiments, the results of the mapping may be selectively saved. Keyframes not used to construct 3D features may be associated with the 3D features through poses, e.g., representing the distance between the keyframes, along with a covariance matrix between the poses of the keyframes. In some embodiments, pairs of keyframes may be selected to construct 3D features such that the distance between each of the constructed 3D features is within a predetermined distance, which may be determined to balance the amount of computation required and the level of accuracy of the resulting model. Such an approach may provide a model of the physical world with a suitable amount of data for efficient and accurate computation using an XR system. In some embodiments, the covariance matrix of two images may include the covariance between the poses (e.g., six degrees of freedom) of the two images.

[0277] Method 4700 may include generating persistent poses based on the keyframes (act 4710). In some embodiments, the method may include generating persistent poses based on 3D features reconstructed from the pair of keyframes. In some embodiments, the persistent poses may be bound to the 3D features. In some embodiments, the persistent poses may include poses of keyframes used to construct the 3D features. In some embodiments, the persistent poses may include an average pose of keyframes used to construct the 3D features. In some embodiments, the persistent poses may be generated such that the distance between neighboring persistent poses is within a predetermined value, for example, in the range of 1 meter to 5 meters, any value therebetween, or any other suitable value. In some embodiments, the distance between neighboring persistent poses may be represented by a covariance matrix of the neighboring persistent poses.

[0278] Method 4700 may include generating a PCF based on the sustained pose (act 4712). In some embodiments, the PCF may be bound to a 3D feature. In some embodiments, the PCF may be associated with one or more sustained poses. In some embodiments, the PCF may include one pose of the associated sustained poses. In some embodiments, the PCF may include an average pose of the poses of the associated sustained poses. In some embodiments, the PCFs may be generated such that the distance between neighboring PCFs is within a predetermined value, e.g., a range of 3 meters to 10 meters, any value therebetween, or any other suitable value. In some embodiments, the distance between neighboring PCFs may be represented by a covariance matrix of the neighboring PCFs. In some embodiments, the PCF may be exposed to an XR application, for example, via an application programming interface (API), so that the XR application may access a model of the physical world through the PCF without accessing the model itself.

[0279] Method 4700 may include associating at least one of the PCFs with image data of the virtual object for display by the XR device (act 4714). In some embodiments, the method may include calculating a translation and orientation of the virtual object relative to the associated PCF. It should be understood that associating a virtual object with a PCF generated by a device that places the virtual object is not required. For example, the device may retrieve a stored PCF in a reference map in the cloud and associate the virtual object with the retrieved PCF. It should be understood that the virtual object may move with the associated PCF as the PCF is adjusted over time.

[0280] 16 illustrates the visual data and algorithms of a first XR device 12.1, a second XR device 12.2, and a server 20, according to some embodiments. The components illustrated in FIG. 16 may operate to perform some or all of the operations associated with generating, updating, and / or using spatial information, such as a persistent pose, a persistent coordinate frame, a tracking map, or a reference map, as described herein. Although not shown, the first XR device 12.1 may be configured identically to the second XR device 12.2. The server 20 may include a map storage routine 118, a reference map 120, a map transmitter 122, and a map merging algorithm 124.

[0281] A second XR device 12.2, which may be in the same scene as the first XR device 12.1, may include a persistent coordinate frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that can be used to render virtual objects, and a frame embedding generator 308 (see FIG. 21 ). In some embodiments, the map download system 126, PCF identification system 128, map 2, localization module 130, reference map embedder 132, reference map 133, and map publisher 136 may be grouped into a passable world unit 1304. The PCF integration unit 1300 may be connected to the passable world unit 1304 and other components of the second XR device 12.2 to enable reading, creating, using, uploading, and downloading PCFs.

[0282] A map comprising PCFs may enable more persistence in a changing world. In some embodiments, locating a tracking map, including, for example, matching features for images, may include selecting features representing persistent content from a map constructed by PCFs, which enables fast matching and / or locating. For example, a world in which people move in and out of a scene and objects such as doors move relative to a scene requires less storage space and transmission rates, and enables the use of individual PCFs and their relationships to one another (e.g., a unified constellation of PCFs) to map a scene.

[0283] In some embodiments, the PCF integration unit 1300 may include a PCF 1306 previously stored in a data storage on a storage unit of the second XR device 12.2, a PCF tracker 1308, a persistent attitude obtainer 1310, a PCF verifier 1312, a PCF generation system 1314, a coordinate frame calculator 1316, a persistent attitude calculator 1318, a tracking map and persistent attitude converter 1320, a persistent attitude and PCF converter 1322, and a PCF and image data converter 1324.

[0284] In some embodiments, the PCF tracker 1308 may have an on prompt and an off prompt that are selectable by the application 1302. The application 1302 may be executable by a processor of the second XR device 12.2 and may, for example, display virtual content. The application 1302 may have a call, via an on prompt, to turn the PCF tracker 1308 on. The PCF tracker 1308 may generate a PCF when the PCF tracker 1308 is turned on. The application 1302 may have a subsequent call, via an off prompt, to turn the PCF tracker 1308 off. The PCF tracker 1308 terminates PCF generation when the PCF tracker 1308 is turned off.

[0285] In some embodiments, the server 20 may include a plurality of sustained poses 1332 and a plurality of PCFs 1330 that have been previously stored in association with the reference map 120. The map transmitter 122 may transmit the reference map 120 along with the sustained poses 1332 and / or PCFs 1330 to the second XR device 12.2. The sustained poses 1332 and PCFs 1330 may be stored on the second XR device 12.2 in association with the reference map 133. Once map 2 is located relative to the reference map 133, the sustained poses 1332 and PCFs 1330 may be stored in association with map 2.

[0286] In some embodiments, persistent attitude obtainer 1310 may obtain a persistent attitude for map 2. PCF verifyer 1312 may be coupled to persistent attitude obtainer 1310. PCF verifyer 1312 may read PCFs from PCF 1306 based on the persistent attitude read by persistent attitude obtainer 1310. The PCFs read by PCF verifyer 1312 may form an initial group of PCFs to be used for image display based on the PCFs.

[0287] In some embodiments, application 1302 may request that additional PCFs be generated. For example, if a user moves into a previously unmapped area, application 1302 may turn on PCF tracker 1308. PCF generation system 1314 may be connected to PCF tracker 1308 and begin generating PCFs based on map 2 as map 2 begins to expand. The PCFs generated by PCF generation system 1314 may form a second group of PCFs that can be used for PCF-based image display.

[0288] The coordinate frame calculator 1316 may be connected to the PCF verifier 1312. After the PCF verifier 1312 reads the PCF, the coordinate frame calculator 1316 may retrieve the head coordinate frame 96 and determine the head pose of the second XR device 12.2. The coordinate frame calculator 1316 may also retrieve a sustained pose calculator 1318. The sustained pose calculator 1318 may be directly or indirectly connected to the frame embedding generator 308. In some embodiments, an image frame may be designated a keyframe after a threshold distance, e.g., 3 meters, from the previous keyframe has been advanced. The sustained pose calculator 1318 may generate a sustained pose based on multiple, e.g., three, keyframes. In some embodiments, the sustained pose may essentially be an average of the coordinate frames of multiple keyframes.

[0289] A tracking map and sustained attitude converter 1320 may be coupled to the map 2 and sustained attitude calculator 1318. The tracking map and sustained attitude converter 1320 may convert map 2 to a sustained attitude and determine a sustained attitude at the origin for map 2.

[0290] The persistent attitude and PCF converter 1322 may be coupled to the tracking map and persistent attitude converter 1320 and further coupled to the PCF verifier 1312 and the PCF generation system 1314. The persistent attitude and PCF converter 1322 may convert the persistent attitude (to which the tracking map has been transformed) into a PCF from the PCF verifier 1312 and the PCF generation system 1314 and determine the PCF for the persistent attitude.

[0291] A PCF and image data converter 1324 may be connected to the persistent attitude and PCF converter 1322 and the data channel 62. The PCF and image data converter 1324 converts the PCF into image data 68. A rendering engine 30 may be connected to the PCF and image data converter 1324 and may display the image data 68 for the PCF to a user.

[0292] The PCF integration unit 1300 may store additional PCFs generated using the PCF generation system 1314 in the PCF 1306. The PCF 1306 may be stored for a sustained attitude. The map publisher 136 may retrieve the PCF 1306 and the sustained attitude associated with the PCF 1306 when the map publisher 136 transmits Map 2 to the server 20 and also transmits the PCF and sustained attitude associated with Map 2 to the server 20. When the map storage routine 118 of the server 20 stores Map 2, the map storage routine 118 may also store the sustained attitude and PCF generated by the second viewing device 12.2. The map merge algorithm 124 may create the reference map 120 with the sustained attitude and PCF of Map 2 associated with the reference map 120 and stored in the sustained attitude 1332 and the PCF 1330, respectively.

[0293] The first XR device 12.1 may include a PCF integration unit similar to the PCF integration unit 1300 of the second XR device 12.2. When the map transmitter 122 transmits the reference map 120 to the first XR device 12.1, the map transmitter 122 may transmit a sustained posture 1332 and a PCF 1330 associated with the reference map 120 and resulting from the second XR device 12.2. The first XR device 12.1 may store the PCF and sustained posture in a data storage device on the storage device of the first XR device 12.1. The first XR device 12.1 may then utilize the sustained posture and PCF resulting from the second XR device 12.2 for image display for the PCF. Additionally or alternatively, the first XR device 12.1 may read, generate, utilize, upload and download PCFs and persistent postures in a manner similar to the second XR device 12.2, as described above.

[0294] In the illustrated example, the first XR device 12.1 generates a local tracking map (hereinafter referred to as "Map 1"), and the map storage routine 118 receives Map 1 from the first XR device 12.1. The map storage routine 118 then stores Map 1 as a reference map 120 on the storage device of the server 20.

[0295] The second XR device 12.2 includes a map download system 126, an anchor identification system 128, a location determination module 130, a reference map embedder 132, a local content location system 134, and a map publisher 136.

[0296] In use, the map transmitter 122 transmits the reference map 120 to the second XR device 12.2, and the map download system 126 downloads and stores the reference map 120 from the server 20 as the reference map 133.

[0297] The anchor identification system 128 is connected to the world surface determination routine 78. The anchor identification system 128 identifies anchors based on objects detected by the world surface determination routine 78. The anchor identification system 128 uses the anchors to generate a second map (Map 2). As shown by cycle 138, the anchor identification system 128 continues to identify anchors and update Map 2. The locations of the anchors are recorded as three-dimensional data based on the data provided by the world surface determination routine 78. The world surface determination routine 78 receives images from the real object detection camera 44 and depth data from the depth sensor 135 and determines the location of the surface and its relative distance from the depth sensor 135.

[0298] Location module 130 is connected to reference map 133 and map 2. Location module 130 repeatedly attempts to locate map 2 relative to reference map 133. Reference map embedder 132 is connected to reference map 133 and map 2. Once location module 130 locates map 2 relative to reference map 133, reference map embedder 132 embeds reference map 133 into the anchor of map 2. Map 2 is then updated with the missing data contained in the reference map.

[0299] A local content location system 134 is connected to map 2. The local content location system 134 may be, for example, a system in which a user can locate local content at a specific location in a world coordinate frame. The local content itself is then tied to an anchor in map 2. A local-to-world coordinate converter 104 converts the local coordinate frame to the world coordinate frame based on the settings of the local content location system 134. The functionality of the rendering engine 30, the display system 42, and the data channel 62 is described with reference to FIG. 2.

[0300] Map publisher 136 uploads Map 2 to server 20. Map storage routine 118 of server 20 then stores Map 2 in the server's 20 storage medium.

[0301] A map merge algorithm 124 merges map 2 with the reference map 120. When more than two maps, for example, three or four maps, relating to the same or adjacent regions of the physical world are stored, the map merge algorithm 124 merges all maps into the reference map 120 and renders the new reference map 120. A map transmitter 122 then transmits the new reference map 120 to every device 12.1 and 12.2 within the area represented by the new reference map 120. Once devices 12.1 and 12.2 locate their respective maps relative to the reference map 120, the reference map 120 becomes a promoted map.

[0302] 17 illustrates an example of generating keyframes for a map of a scene, according to some embodiments. In the example shown, a first keyframe KF1 is generated for a door on the left wall of a room. A second keyframe KF2 is generated for the area in the corner where the room floor, left wall, and right wall meet. A third keyframe KF3 is generated for the window area on the right wall of the room. A fourth keyframe KF4 is generated for the area at the edge of a rug on the floor of the wall. A fifth keyframe KF5 is generated for the area of ​​the rug closest to the user.

[0303] FIG. 18 illustrates an example of generating a persistent posture for the map of FIG. 17 , according to some embodiments. In some embodiments, a new persistent posture is created when the device measures a threshold distance traveled and / or when an application requests a new persistent posture (PP). In some embodiments, the threshold distance may be 3 meters, 5 meters, 20 meters, or any other suitable distance. Selecting a smaller threshold distance (e.g., 1 m) may result in an increased computational load because a larger number of PPs may be created and managed compared to a larger threshold distance. Selecting a larger threshold distance (e.g., 40 m) may result in fewer PPs being created and fewer PCFs being created, which may result in an increased virtual content placement error because virtual content bound to a PCF may be a relatively large distance (e.g., 30 m) from the PCF, meaning that error may increase with increasing distance from the PCF to the virtual content.

[0304] In some embodiments, a PP may be created at the start of a new session. This initial PP may be considered zero and visualized as the center of a circle with a radius equal to the threshold distance. When the device reaches the circumference of the circle and, in some embodiments, an application requests a new PP, the new PP may be placed at the device's current location (threshold distance). In some embodiments, a new PP will not be created at the threshold distance if the device can find an existing PP within the threshold distance from the device's new location. In some embodiments, when a new PP (PP 1150 in FIG. 14) is created, the device associates one or more of the closest key frames with the PP. In some embodiments, the location of the PP relative to the key frame may be based on the device's location at the time the PP is created. In some embodiments, a PP will not be created even if the device progresses the threshold distance unless an application requests the PP.

[0305] In some embodiments, an application may request a PCF from a device when the application has virtual content to display to the user. A PCF request from the application may trigger a PP request, and a new PP will be created after the device has progressed a threshold distance. Figure 18 illustrates a first sustained posture PP1, which may cause the closest keyframes (e.g., KF1, KF2, and KF3) to be linked, for example, by calculating the relative posture between the keyframes and the sustained posture. Figure 18 also illustrates a second sustained posture PP2, which may cause the closest keyframes (e.g., KF4 and KF5) to be linked.

[0306] FIG. 19 illustrates an example of generating a PCF for the map of FIG. 17 , according to some embodiments. In the illustrated example, PCF1 may include PP1 and PP2. As described above, PCFs may be used to display image data for the PCF. In some embodiments, each PCF may have coordinates in another coordinate frame (e.g., a world coordinate frame) and a PCF descriptor, e.g., that uniquely identifies the PCF. In some embodiments, the PCF descriptor may be calculated based on feature descriptors of features in the frame associated with the PCF. In some embodiments, various constellations of PCFs may be combined to represent the real world in a persistent manner, requiring less data and less data transmission.

[0307] 20A-20C are schematic diagrams illustrating examples of establishing and using persistent coordinate frames. FIG. 20A shows two users 4802A, 4802B with their respective local tracking maps 4804A, 4804B, which are not localized relative to a reference map. Origins 4806A, 4806B for each user are described by a coordinate system within their respective area (e.g., a world coordinate system). These origins for each tracking map may be local to each user, since the origin depends on the orientation of that individual device when tracking begins.

[0308] As the user device's sensors scan the environment, the device may capture images that may contain features representative of persistent objects so that those images may be classified as keyframes from which persistent poses may be created, as described above in connection with Figure 14. In this example, tracking map 4802A includes persistent pose (PP) 4808A, and tracking map 4802B includes PP 4808B.

[0309] Also, as described above in connection with Figure 14, some of the PPs may be classified as PCFs, which are used to determine the orientation of virtual content for rendering it to the user. Figure 20B shows that XR devices worn by individual users 4802A, 4802B can create local PCFs 4810A, 4810B based on PPs 4808A, 4808B. Figure 20C shows that persistent content 4812A, 4812B (e.g., virtual content) can be bound to the PCFs 4810A, 4810B by the individual XR devices.

[0310] In this example, virtual content may have a virtual content coordinate frame that can be used by an application that generates the virtual content regardless of how the virtual content is to be displayed. The virtual content may be defined as surfaces, such as triangles of a mesh, at particular locations and angles relative to the virtual content coordinate frame. To render that virtual content to a user, the locations of those surfaces may be determined relative to the user who will perceive the virtual content.

[0311] Binding virtual content to a PCF may simplify the computations involved in determining the location of the virtual content relative to a user. The location of the virtual content relative to a user may be determined by applying a series of transformations. Some of those transformations may change and be updated frequently. Others of those transformations may be stable and updated less frequently or not at all. Nevertheless, the transformations may be applied with a relatively low computational burden such that the location of the virtual content is updated frequently to the user and may provide a realistic appearance to the rendered virtual content.

[0312] In the example of Figures 20A-20C, User 1's device has a coordinate system that may be related to the coordinate system that defines the map origin by a transformation rig1_T_w1. User 2's device has a similar transformation rig2_T_w2. These transformations may be expressed as six-degree transformations and may define a translation and rotation to align the device coordinate system with the map coordinate system. In some embodiments, the transformation may be expressed as two separate transformations, one defining the translation and the other defining the rotation. It should be understood, therefore, that the transformations may be expressed in a form that simplifies calculations or otherwise provides advantages.

[0313] The transformations between the origin of the tracking map and the PCFs identified by individual user devices are denoted as pcf1_T_w1 and pcf2_T_w2. In this example, the PCFs and PPs are the same, so that the same transformation also characterizes the PP.

[0314] The location of the user device relative to the PCF is therefore rig1_T_pcf1=(rig1_T_w1) * It can be calculated by successive applications of these transformations such as (pcf1_T_w1).

[0315] As shown in Figure 20C, the virtual content is located relative to the PCF using the transformation obj1_T_pcf1. This transformation may be set by the application generating the virtual content, which may receive information from a world reconstruction system that describes the physical objects relative to the PCF. To render the virtual content to the user, a transformation to the coordinate system of the user's device is calculated, which is the transformation obj1_t_w1 = (obj1_T_pcf1) * The virtual content coordinate frame can be calculated by relating it to the origin of the tracking map through (pcf1_T_w1), which can then be related to the user's device through a further transformation rig1_T_w1.

[0316] The location of the virtual content may change based on the output from the application generating the virtual content. As it changes, the end-to-end transformation from the source coordinate system to the destination coordinate system may be recalculated. In addition, the user's location and / or head pose may also change as the user moves. As a result, just as the transformation rig1_T_w1 may change, any end-to-end transformation that depends on the user's location or head pose will also change.

[0317] The transformation rig1_T_w1 may be updated with the user's movement based on tracking the user's position relative to stationary objects in the physical world. Such tracking may be performed by a headphone tracking component that processes a sequence of images, as described above, or by other components of the system. Such updates may be performed by determining the user's pose relative to a stationary reference frame, such as PP.

[0318] In some embodiments, the location and orientation of the user device may be determined relative to the nearest sustained pose, or in this example, the PCF as PP is used as the PCF. Such a determination may be made by identifying feature points that characterize the PP in a current image captured using sensors on the device. Using image processing techniques such as stereoscopic image analysis, the location of the device relative to those feature points may be determined. From this data, the system calculates the relationship rig1_T_pcf1=(rig1_T_w1) * Based on (pcf1_T_w1), the change in transformation associated with the user's movement may be calculated.

[0319] The system may determine and apply transformations in an order that is computationally efficient. For example, the need to calculate rig1_T_w1 from measurements that result in rig1_T_pcf1 may be avoided by both tracking the user pose and defining the location of the virtual content relative to a PP or PCF that is built on a persistent pose. In this way, the transformation from the source coordinate system of the virtual content to the destination coordinate system of the user's device is expressed as (rig1_T_pcf1) * The end-to-end transformation may be based on a measured transformation according to (obj1_t_pcf1), where the first transformation is measured by the system and the latter transformation is supplied by the application that defines the virtual content for rendering. In embodiments where the virtual content is positioned relative to the map origin, the end-to-end transformation may relate the virtual object coordinate system to the PCF coordinate system based on a further transformation between map coordinates and PCF coordinates. In embodiments where the virtual content is positioned relative to a different PP or PCF than the one for which the user position is tracked, a transformation between the two may be applied. Such a transformation may be fixed or, for example, determined from the map in which both appear.

[0320] A transform-based approach may be implemented within a device, for example, with components that process sensor data and build a tracking map. As part of that process, those components may identify feature points that can be used as persistent poses, which may then be turned into PCFs. Those components may limit the number of persistent poses generated for the map and provide suitable spacing between persistent poses, as described above in connection with Figures 17-19, while allowing the user to be close enough to the persistent pose location to accurately calculate the user's pose, regardless of their location in the physical environment. As the persistent pose closest to the user is updated as a result of user movement, refinements to the tracking map, or other causes, any of the transforms used to calculate the location of virtual content relative to the user, which depend on the location of the PP (or PCF, if used), may be updated and stored for use at least until the user moves away from that persistent pose. Note that by calculating and storing the transforms, the computational burden each time the location of virtual content is updated may be relatively low, such that it can be implemented with relatively low latency.

[0321] 20A-20C illustrate positioning relative to a tracking map, with each device having its own tracking map. However, transformations may be generated relative to any map coordinate system. Persistence of content across user sessions in an XR system can be achieved through the use of a persistent map. Shared user experiences can also be facilitated through the use of a map to which multiple user devices can be oriented.

[0322] In some embodiments, described in more detail below, the location of virtual content may be defined relative to coordinates in a reference map, which is formatted so that any of multiple devices may use the map. Each device may maintain a tracking map and may determine changes in the user's posture relative to the tracking map. In this example, the transformation between the tracking map and the reference map may be determined through a process of "localization," which may be performed by matching structures in the tracking map (e.g., one or more persistent postures) with one or more structures in the reference map (e.g., one or more PCFs).

[0323] Described further below are techniques for creating and using such reference maps.

[0324] Deep Keyframes

[0325] Techniques as described herein rely on comparison of image frames. For example, to establish the device's position relative to the tracking map, a new image may be captured using a sensor worn by the user, and the XR system may search for images in the set of images used to create the tracking map that share at least a predetermined amount of points of interest with the new image. As an example of another scenario involving comparison of image frames, the tracking map may be located relative to the reference map by first finding an image frame associated with a sustained pose in the tracking map that is similar to an image frame associated with a PCF in the reference map. Alternatively, the transformation between two reference maps may be calculated by first finding similar image frames in the two maps.

[0326] Deep keyframes provide a method for reducing the amount of processing required to identify similar image frames. For example, in some embodiments, a comparison may be made between image features (e.g., "2D features") in a new 2D image and 3D features in the map. Such a comparison may be made in any suitable manner, such as by projecting the 3D image into a 2D plane. Conventional methods, such as bag of words (BoW), search for the 2D features of the new image in a database that includes all 2D features in the map, which may require significant computational resources, especially when the map represents a large area. Conventional methods then locate images that share at least one of the 2D features with the new image, which may include images that are not useful for locating meaningful 3D features in the map. Conventional methods then locate 3D features that are not meaningful relative to the 2D features in the new image.

[0327] The inventors have recognized and understood techniques for retrieving images in a map using fewer memory resources (e.g., one-quarter of the memory resources used by BoW), greater efficiency (e.g., 2.5 ms processing time per keyframe, 100 μs for comparison against 500 keyframes), and greater accuracy (e.g., 20% better read-back than BoW for 1,024-dimensional models, 5% better read-back than BoW for 256-dimensional models).

[0328] To reduce computation, descriptors may be computed for an image frame that can be used to compare the image frame to other image frames. The descriptors may be stored instead of or in addition to the image frame and feature points. In maps where sustained poses and / or PCFs may be generated from image frames, the descriptors for the image frame or frames from which each sustained pose or PCF was generated may be stored as part of the sustained pose and / or PCF.

[0329] In some embodiments, the descriptor may be calculated as a function of feature points within an image frame. In some embodiments, a neural network is configured to calculate a unique frame descriptor to represent the image. The image may have a resolution greater than 1 megabyte so that sufficient detail of the 3D environment within the field of view of the device worn by the user is captured in the image. The frame descriptor may be much smaller, such as a string of numbers, for example, any number in the range of 128 bytes to 512 bytes or therebetween.

[0330] In some embodiments, the neural network is trained so that the calculated frame descriptor indicates the similarity between images. Images in the map may be located by identifying the nearest image in a database comprising the images used to generate the map, which may have a frame descriptor within a predetermined distance to the frame descriptor for the new image. In some embodiments, the distance between images may be represented by the difference between the frame descriptors of the two images.

[0331] 21 is a block diagram illustrating a system for generating descriptors for individual images, according to some embodiments. In the illustrated example, a frame embedding generator 308 is shown. The frame embedding generator 308 may in some embodiments be used in conjunction with the server 20, but may alternatively or additionally be executed, in whole or in part, within one of the XR devices 12.1 and 12.2, or any other device that processes images for comparison with other images.

[0332] In some embodiments, the frame embedding generator may be configured to generate a data representation of an image reduced from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes) that, despite the reduced size, still represents the content within the image. In some embodiments, the frame embedding generator may be used to generate a data representation for an image, which may be a keyframe or frame used in other methods. In some embodiments, the frame embedding generator 308 may be configured to convert an image at a particular location and orientation into a unique string of numbers (e.g., 256 bytes). In the illustrated example, an image 320 captured by an XR device may be processed by a feature extractor 324 to detect points of interest 322 within the image 320. The points of interest may or may not be derived from identified feature points, as described above with respect to features 1120 ( FIG. 14 ) or as otherwise described herein. In some embodiments, the points of interest may be represented by descriptors, as described above with respect to descriptors 1130 ( FIG. 14 ), which may be generated using deep sparse feature methods. In some embodiments, each point of interest 322 may be represented by a string of numbers (e.g., 32 bytes). For example, there may be n features (e.g., 100), each represented by a string of 32 bytes.

[0333] In some embodiments, the frame embedding generator 308 may include a neural network 326. The neural network 326 may include a multilayer perceptron unit 312 and a max pooling unit 314. In some embodiments, the multilayer perceptron (MLP) unit 312 may comprise a multilayer perceptron, which may be trained. In some embodiments, points of interest 322 (e.g., descriptors for the points of interest) may be reduced by the multilayer perceptron 312 and output as a weighted combination of descriptors 310. For example, the MLP may reduce n features to m features, which is less than n features.

[0334] In some embodiments, the MLP unit 312 may be configured to perform matrix multiplication. The multilayer perceptron unit 312 receives multiple points of interest 322 of the image 320 and converts each point of interest into a separate string of numbers (e.g., 256). For example, there may be 100 features, and each feature may be represented by a string of 256 numbers. The matrix may be created to have 100 horizontal rows and 256 vertical columns in this example. Each row may have a series of 256 numbers of varying magnitude, some smaller and some larger. In some embodiments, the output of the MLP may be an n x 256 matrix, where n represents the number of features extracted from the image. In some embodiments, the output of the MLP may be an m x 256 matrix, where m is the number of points of interest reduced from n.

[0335] In some embodiments, the MLP 312 may have a training phase during which model parameters for the MLP are determined, and an application phase. In some embodiments, the MLP may be trained as illustrated in Figure 25. The input training data may comprise data in three sets, the three sets comprising: 1) a query image, 2) positive samples, and 3) negative samples. The query image may be considered a reference image.

[0336] In some embodiments, the positive sample may comprise an image that is similar to the query image. For example, in some embodiments, similar may mean having the same object in both the query and positive sample images, but viewed from different angles. In some embodiments, similar may mean having the same object in both the query and positive sample images, but having an object that is shifted (e.g., left, right, up, down) relative to the other image.

[0337] In some embodiments, negative samples may comprise images that are dissimilar to the query image. For example, in some embodiments, dissimilar images may not contain any objects salient in the query image, or may contain only a small fraction (e.g., <10%, 1%) of objects salient in the query image. Similar images, in contrast, may have, for example, a large fraction (e.g., >50%, or >75%) of the objects in the query image.

[0338] In some embodiments, points of interest may be extracted from images in the input training data and converted into feature descriptors. These descriptors may be calculated for both the training images, as shown in FIG. 25, and for features extracted during operation of the frame embedding generator 308 of FIG. 21. In some embodiments, a deep sparse feature (DSF) process may be used to generate the descriptors (e.g., DSF descriptors), as described in U.S. patent application Ser. No. 16 / 190,948. In some embodiments, the DSF descriptors are n×32 in size. The descriptors may then be passed through the model / MLP to create a 256-byte output. In some embodiments, the model / MLP may have the same structure as the MLP 312, so that once the model parameters are set through training, the resulting trained MLP can be used as the MLP 312.

[0339] In some embodiments, the feature descriptors (e.g., the 256 bytes output from the MLP model) may then be sent to a triplet margin loss module (used only during the training phase of the MLP neural network and may not be used during the use phase). In some embodiments, the triplet margin loss module may be configured to select parameters for the model to reduce the difference between the 256 bytes output from the query image and the 256 bytes output from the positive samples and to increase the difference between the 256 bytes output from the query image and the 256 bytes output from the negative samples. In some embodiments, the training phase may include feeding multiple triplet input images into a learning process to determine model parameters. This training process may continue, for example, until the difference for positive images is minimized and the difference for negative images is maximized, or until some other suitable termination criterion is reached.

[0340] Referring back to FIG. 21 , the frame embedding generator 308 may include a pooling layer, here illustrated as a max pooling unit 314. The max pooling unit 314 may analyze each column and determine the maximum number in each column. The max pooling unit 314 may combine the maximum values ​​of each column of the MLP 312 output matrix into a global feature column 316, for example, a number of 256. It should be understood that images processed in an XR system may desirably have high-resolution frames, potentially with millions of pixels. The global feature column 316 is a relatively small number that occupies relatively little memory and is easily searchable compared to images (e.g., with a resolution greater than 1 megabyte). Therefore, it is possible to search images without analyzing each original frame from the camera, and it is also cheaper to store 256 bytes instead of a full frame.

[0341] 22 is a flowchart illustrating a method 2200 of calculating image descriptors, according to some embodiments. Method 2200 may begin with receiving a plurality of images captured by an XR device worn by a user (act 2202). In some embodiments, method 2200 may include determining one or more key frames from the plurality of images (act 2204). In some embodiments, act 2204 may be skipped and / or may instead occur after step 2210.

[0342] Method 2200 may include identifying one or more points of interest in the plurality of images using an artificial neural network (act 2206) and calculating feature descriptors for each point of interest using the artificial neural network (act 2208). The method may also include, for each image, calculating, at least in part, using the artificial neural network, a frame descriptor to represent the image based on the calculated feature descriptors for the identified points of interest in the image (act 2210).

[0343] FIG. 23 is a flowchart illustrating a method 2300 of localization using image descriptors, according to some embodiments. In this example, a new image frame depicting the current location of the XR device may be compared to stored image frames associated with points in a map (such as a persistent pose or PCF, as described above). Method 2300 may begin with receiving a new image captured by an XR device worn by a user (act 2302). Method 2300 may include identifying one or more nearest keyframes in a database (act 2304), which comprise keyframes used to generate one or more maps. In some embodiments, the nearest keyframes may be identified based on coarse spatial information and / or previously determined spatial information. For example, coarse spatial information may indicate that the XR device is located within a geographic region represented by a 50 m x 50 m area of ​​the map. Image matching may be performed only with respect to points within that area. As another example, based on tracking, the XR system may know that the XR device was previously proximate to a first sustained pose in the map and was moving in the direction of a second sustained pose in the map. That second sustained pose may be considered the nearest sustained pose, and the keyframe stored therewith may be considered the nearest keyframe. Alternatively or additionally, other metadata, such as GPS data or WiFi fingerprints, may also be used to select the nearest keyframe or set of nearest keyframes.

[0344] Regardless of how the nearest keyframe is selected, the frame descriptor may be used to determine whether the new image matches any of the frames selected as being associated with nearby persistent poses. The determination may be made by comparing the frame descriptor of the new image with the frame descriptors of the nearest keyframes or a subset of keyframes in the database selected in any other suitable manner, and selecting keyframes with frame descriptors within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors may be calculated by taking the difference between two numeric strings that may represent the two frame descriptors. In embodiments where the strings are treated as a string of multiple quantities, the difference may be calculated as a vector difference.

[0345] Once a matching image frame is identified, the orientation of the XR device relative to that image frame can be determined. Method 2300 may include performing feature matching on 3D features in the map corresponding to the identified nearest key frame (act 2306) and calculating a pose of the device worn by the user based on the feature matching results (act 2308). In this way, computationally intensive matching of feature points in two images may be performed with respect to as few as one image that has already been determined to be a likely match for the new image.

[0346] 24 is a flowchart illustrating a method 2400 of training a neural network, according to some embodiments. Method 2400 may begin with generating a dataset (act 2402) comprising a plurality of image sets. Each of the plurality of image sets may include a query image, a positive sample image, and a negative sample image. In some embodiments, the plurality of image sets may include synthetic record pairs configured to teach the neural network basic information such as shape, for example. In some embodiments, the plurality of image sets may include real record pairs, which may be recorded from the physical world.

[0347] In some embodiments, positive correspondences may be calculated by matching fundamental matrices between the two images. In some embodiments, sparse overlaps may be calculated as the intersection over union (IoU) of points of interest found in both images. In some embodiments, a positive sample may contain at least 20 points of interest that are identical in the query image and serve as positive correspondences. A negative sample may contain fewer than 10 positive correspondences. A negative sample may have less than half of the sparse points overlapping with analysis points in the query image.

[0348] Method 2400 may include, for each set of images, calculating a loss by comparing the query image with the positive and negative sample images (act 2404). Method 2400 may also include modifying the artificial neural network based on the calculated loss (act 2406) such that the distance between the frame descriptor generated by the artificial neural network for the query image and the frame descriptor for the positive sample images is less than the distance between the frame descriptor for the query image and the frame descriptor for the negative sample images.

[0349] While methods and apparatus configured to generate a global descriptor for an individual image are described above, it should be understood that the methods and apparatus may also be configured to generate a descriptor for an individual map. For example, a map may include multiple keyframes, each having a frame descriptor as described above. The max pooling unit may analyze the frame descriptors of the map's keyframes and combine the frame descriptors into a unique map descriptor for the map.

[0350] Furthermore, it should be understood that other architectures may be used for processing as described above. For example, separate neural networks are described for generating the DSF descriptor and the frame descriptor. Such an approach is computationally efficient. However, in some embodiments, the frame descriptor may be generated from selected feature points without first generating the DSF descriptor. Ranking and merging maps

[0351] Described herein are methods and apparatus for ranking and merging multiple environment maps within an XReality (XR) system. Map merging may allow maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking maps may enable efficient implementation of techniques such as those described herein, including map merging, which involves selecting a map from a set of maps based on similarity. In some embodiments, a set of reference maps may be maintained by the system, formatted in a manner that can be accessed by any of several XR devices, for example. These reference maps may be formed by merging selected tracking maps from those devices with other tracking maps or previously stored reference maps. The reference maps may be ranked, for example, to select one or more reference maps and merge them with new tracking maps and / or select one or more reference maps from the set and use them in a device.

[0352] To provide a user with a realistic XR experience, the XR system must understand the user's physical surroundings in order to correctly correlate the location of virtual objects in relation to real objects. Information about the user's physical surroundings may be obtained from an environmental map of the user's location.

[0353] The inventors have recognized and appreciated that an XR system can provide an enhanced XR experience to multiple users sharing the same world, comprising real and / or virtual content, regardless of whether those users are present in the world at the same or different times, by enabling efficient sharing of real / physical world environment maps collected by multiple users. However, significant challenges exist in providing such a system. Such a system may store multiple maps generated by multiple users and / or the system may store multiple maps generated at different times. For example, as described above, with respect to operations that may be performed using previously generated maps, such as localization, substantial processing may be required to identify relevant environment maps of the same world (e.g., the same real-world location) from all environment maps collected within the XR system. In some embodiments, there may be only a small number of environment maps that a device can access, for example, for localization. In some embodiments, there may be a large number of environment maps that a device can access. The present inventors have recognized and appreciated techniques for quickly and accurately ranking the relevance of environmental maps from all possible environmental maps, such as the population of all reference maps 120 in Figure 28. Highly ranked maps may then be selected for further processing, such as rendering virtual objects on a user display to interact realistically with the physical world around the user, or merging stored maps with map data collected by the user to create larger or more accurate maps.

[0354] In some embodiments, stored maps relevant to a task for a user at a location in the physical world may be identified by filtering the stored maps based on multiple criteria. These criteria may refer to a comparison of a tracking map generated by the user's wearable device at the location with candidate environment maps stored in a database. The comparison may be performed based on metadata associated with the map, such as a Wi-Fi fingerprint detected by the device generating the map, and / or the set of BSSIDs to which the device connected while forming the map. The comparison may also be performed based on the compressed or decompressed contents of the map. A comparison based on a compressed representation may be performed, for example, by comparing vectors calculated from the map contents. A comparison based on an decompressed map may be performed, for example, by locating the tracking map within the stored map, or vice versa. Multiple comparisons may be performed in an order based on the calculation time required to reduce the number of candidate maps for consideration, with comparisons involving fewer calculations being performed before other comparisons requiring more calculations.

[0355] FIG. 26 depicts an AR system 800 configured to rank and merge one or more environment maps, according to some embodiments. The AR system may include a passable world model 802 of the AR device. Information for populating the passable world model 802 may originate from sensors on the AR device, which may include computer-executable instructions stored in a processor 804 (e.g., local data processing module 570 in FIG. 4 ) that may perform some or all of the processing for converting sensor data into a map. Such a map may be a tracking map that may be constructed as sensor data is collected as the AR device operates within an area. Along with the tracking map, area attributes may be provided to indicate the area that the tracking map represents. These area attributes may be coordinates presented as latitude and longitude or geographic location identifiers, such as IDs used by the AR system to represent locations. Alternatively, or in addition, area attributes may be measured characteristics that have a high likelihood of being unique for the area. Area attributes may be derived, for example, from parameters of wireless networks detected within the area. In some embodiments, the area attribute may be associated with the unique address of an access point that the AR system is near and / or connected to. For example, the area attribute may be associated with the MAC address or basic service set identifier (BSSID) of a 5G base station / router, Wi-Fi router, and the like.

[0356] 26, the tracking map may be merged with other maps of the environment. Map ranking portion 806 receives the tracking map from device PW 802, communicates with map database 808, and selects and ranks environment maps from map database 808. The selected maps that are higher ranked are sent to map merging portion 810.

[0357] The map merging portion 810 may perform a merging process on the maps sent from the map ranking portion 806. The merging process may involve merging some or all of the tracking maps with the ranked maps and transmitting the new merged map to the passable world model 812. The map merging portion may merge maps by identifying maps that depict overlapping portions of the physical world. Those overlapping portions may be matched so that information in both maps can be aggregated in a final map. Reference maps may be merged with other reference maps and / or tracking maps.

[0358] Aggregation may involve extending one map with information from another map. Alternatively, or in addition, aggregation may involve adjusting the representation of the physical world in one map based on information in another map. The latter map may, for example, represent that an object causing a feature point has moved so that the map can be updated based on the latter information. Alternatively, two maps may characterize the same area with different feature points, and aggregation may involve selecting a set of feature points from the two maps to better represent the area. Regardless of the specific processing that occurs in the merging process, in some embodiments, PCFs from all maps being merged may be retained so that applications that position content relative to them can continue to do so. In some embodiments, merging maps may result in redundant persistent attitudes, and some of the persistent attitudes may be deleted. When a PCF is associated with a persistent attitude to be deleted, merging maps may involve modifying the PCF to be associated with the persistent attitude that remains in the map after the merge.

[0359] In some embodiments, maps may be refined as they are expanded and / or updated. Refinement may involve calculations to reduce internal discrepancies between feature points that are likely to represent the same object in the physical world. Discrepancies may arise from inaccuracies in the poses associated with keyframes that provide feature points that represent the same object in the physical world. Such discrepancies may arise, for example, from an XR device calculating a pose for a tracking map on which pose estimation is built, such that errors in the pose estimate accumulate and create a "drift" in pose accuracy over time. Maps may be refined by performing bundle adjustment or other operations to reduce discrepancies in feature points from multiple keyframes.

[0360] Depending on the refinement, the location of a persistence point relative to the map origin may change. Thus, the transformation associated with that persistence point, such as a persistence pose or PCF, may also change. In some embodiments, the XR system may recalculate the transformations associated with any persistence point that has changed in connection with a map refinement (whether performed as part of a merge operation or for other reasons). These transformations may be pushed from the component that computes the transforms to the component that uses the transforms, so that any use of the transforms can be based on the updated location of the persistence point.

[0361] The passable world model 812 may be a cloud model, which may be shared by multiple AR devices. The passable world model 812 may store or otherwise have access to the environment map in a map database 808. In some embodiments, when a previously calculated environment map is updated, the previous version of the map may be deleted, removing the outdated map from the database. In some embodiments, when a previously calculated environment map is updated, the previous version of the map may be archived, enabling reading / viewing of the previous version of the environment. In some embodiments, permissions may be set so that only AR systems with certain read / write access can trigger the deletion / archiving of the previous version of the map.

[0362] These environment maps created from tracking maps provided by one or more AR devices / systems may be accessed by AR devices in the AR system. The map ranking portion 806 may also be used in providing environment maps to AR devices. An AR device may send a message requesting an environment map for its current location, and the map ranking portion 806 may be used to select and rank the environment map associated with the requesting device.

[0363] In some embodiments, the AR system 800 may include a downsampling portion 814 configured to receive the merged map from the cloud PW 812. The merged map received from the cloud PW 812 may be in a storage format for the cloud, which may include high-resolution information, such as multiple image frames or a large set of feature points associated with a large number of PCFs per square meter or PCFs. The downsampling portion 814 may be configured to downsample the cloud-format map to a format suitable for storage on an AR device. The device-format map may have less data, such as fewer PCFs or less data stored per PCF, to accommodate the limited local computing power and storage space of the AR device.

[0364] 27 is a simplified block diagram illustrating multiple reference maps 120, which may be stored in a remote storage medium, e.g., the cloud. Each reference map 120 may include multiple reference map identifiers, which indicate the reference map's location in physical space, such as anywhere on planet Earth. These reference map identifiers may include one or more of the following identifiers: an area identifier, represented by a longitude and latitude range; a frame descriptor (e.g., global feature column 316 in FIG. 21); a Wi-Fi fingerprint; a feature descriptor (e.g., feature descriptor 310 in FIG. 21); and a device identification, which indicates one or more devices that contributed to the map.

[0365] In the illustrated embodiment, reference maps 120 may reside on the surface of the Earth and thus be geographically arranged in a two-dimensional pattern. Reference maps 120 may be uniquely identifiable by their corresponding longitudes and latitudes, such that any reference maps with overlapping longitudes and latitudes may be merged into a new reference map. FIG. 28 is a schematic diagram illustrating a method for selecting a reference map that can be used to locate a new tracking map against one or more reference maps, according to some embodiments. The method may begin by accessing (act 120) a universe of reference maps 120, which, by way of example, may be stored in a database within the passable world (e.g., passable world module 538). The universe of reference maps may include reference maps from all previously visited locations. The XR system may filter the universe of all reference maps to a small subset or only a single map. It should be understood that in some embodiments, it may be impossible to transmit all reference maps to the viewing device due to bandwidth limitations. Selecting a subset selected as likely candidates for matching a tracking map to transmit to the device may reduce bandwidth and latency associated with accessing a remote database of maps.

[0366] The method may include filtering the population of reference maps (act 300) based on areas with a predetermined size and shape. In the example illustrated in FIG. 27, each square may represent an area. Each square may cover 50 m by 50 m. Each square may have six neighboring areas. In some embodiments, act 300 may select at least one matching reference map 120 that covers a longitude and latitude that includes the longitude and latitude of the location identifier received from the XR device, as long as at least one map exists at that longitude and latitude. In some embodiments, act 300 may select at least one neighboring reference map that covers a longitude and latitude that is adjacent to the matching reference map. In some embodiments, act 300 may select multiple matching reference maps and multiple neighboring reference maps. Act 300 may, for example, reduce the number of reference maps by about a factor of 10, e.g., from thousands to hundreds, to form a first filtered selection. Alternatively, or in addition, criteria other than latitude and longitude may be used to identify nearby maps. The XR device may have previously been located using a reference map in the set, for example, as part of the same session. The cloud service may retain information about the XR device, including previously located maps. In this example, the maps selected in act 300 may include those covering an area adjacent to the map for which the XR device was located.

[0367] The method may include filtering a first filtered selection of reference maps based on the Wi-Fi fingerprint (act 302). Act 302 may determine a latitude and longitude based on the Wi-Fi fingerprint received as part of the location identifier from the XR device. Act 302 may compare the latitude and longitude from the Wi-Fi fingerprint with the latitude and longitude of reference map 120 to determine one or more reference maps, which form a second filtered selection. Act 302 may reduce the number of reference maps by approximately a factor of 10, e.g., from hundreds to tens of reference maps (e.g., 50), which form the second selection. For example, the first filtered selection may include 130 reference maps, and the second filtered selection may include 50 of the 130 reference maps and not the remaining 80 of the 130 reference maps.

[0368] The method may include filtering a second filtered selection of reference maps based on the keyframes (act 304). Act 304 may compare data representing the image captured by the XR device with data representing the reference map 120. In some embodiments, the data representing the image and / or map may include feature descriptors (e.g., DSF descriptors in FIG. 25) and / or global feature columns (e.g., 316 in FIG. 21). Act 304 may provide a third filtered selection of reference maps. In some embodiments, the output of act 304 may be, for example, only five of the 50 reference maps identified following the second filtered selection. Map transmitter 122 then transmits one or more reference maps to the viewing device based on the third filtered selection. Act 304 may reduce the number of reference maps by approximately a factor of ten, for example, from tens to a single-digit number of reference maps (e.g., five), forming the third selection. In some embodiments, the XR device may receive a reference map within the third filtered selection and attempt to locate within the received reference map.

[0369] For example, act 304 may filter the reference map 120 based on the global feature columns 316 of the reference map 120 and the global feature columns 316 based on an image captured by the viewing device (e.g., an image that may be part of a local tracking map for the user). Each of the reference maps 120 in FIG. 27 thus has one or more global feature columns 316 associated with it. In some embodiments, the global feature columns 316 may be obtained when the XR device submits the image or feature details to the cloud, which processes the image or feature details and generates the global feature columns 316 for the reference map 120.

[0370] In some embodiments, the cloud may receive feature details of a live / new / current image captured by the viewing device, and the cloud may generate a global feature sequence 316 for the live image. The cloud may then filter the reference map 120 based on the live global feature sequence 316. In some embodiments, the global feature sequence may be generated on the local viewing device. In some embodiments, the global feature sequence may be generated remotely, for example, on the cloud. In some embodiments, the cloud may transmit the filtered reference map along with the global feature sequence 316 associated with the filtered reference map to the XR device. In some embodiments, once the viewing device locates its tracking map relative to the reference map, it may do so by matching the global feature sequence 316 of the local tracking map with the global feature sequence of the reference map.

[0371] It should be understood that an XR device's operation need not perform all of the acts (300, 302, 304). For example, if the universe of reference maps is relatively small (e.g., 500 maps), an XR device attempting to locate may filter the universe of reference maps based on Wi-Fi fingerprints (e.g., act 302) and keyframes (e.g., act 304), but omit filtering based on area (e.g., act 300). Furthermore, maps need not be compared in their entirety. In some embodiments, for example, a comparison of two maps may result in the identification of common persistence points, such as persistence poses or PCFs, that appear in both the new map and the map selected from the universe of maps. Descriptors may then be associated with the persistence points, and those descriptors may be compared.

[0372] 29 is a flowchart illustrating a method 900 of selecting one or more ranked environment maps according to some embodiments. In the illustrated embodiment, the ranking is performed for the user's AR device, which creates the tracking map. Thus, the tracking map is available for use in ranking the environment maps. In embodiments where a tracking map is not available, some or all of the portions of the environment map selection and ranking that do not explicitly rely on the tracking map may be used.

[0373] Method 900 may begin at act 902, where a set of maps (which may be formatted as reference maps) from a database of environmental maps near where the tracking map was formed is accessed and then filtered for ranking. Additionally, act 902 determines at least one area attribute for an area in which the user's AR device is operating. In a scenario in which the user's AR device is building a tracking map, the area attribute may correspond to the area over which the tracking map is created. As a specific example, the area attribute may be calculated based on signals received from an access point on a computer network while the AR device was calculating the tracking map.

[0374] 30 depicts an example map ranking portion 806 of an AR system 800, according to some embodiments. The map ranking portion 806 may include portions that execute on the AR device and portions that execute on a remote computing system, such as a cloud, and thus may execute within a cloud computing environment. The map ranking portion 806 may be configured to implement at least a portion of the method 900.

[0375] FIG. 31A depicts an example of area attributes AA1-AA8 of the tracking map (TM) 1102 and environment maps CM1-CM4 in the database, according to some embodiments. As shown, the environment map may be associated with multiple area attributes. The area attributes AA1-AA8 may include parameters of the wireless network detected by the AR device calculating the tracking map 1102, such as the basic service set identifier (BSSID) of the network to which the AR device is connected and / or the received signal strength of an access point to the wireless network, for example, through a network tower 1104. The wireless network parameters may conform to protocols including Wi-Fi and 5G NR. In the example illustrated in FIG. 32, the area attributes are fingerprints of the areas in which the user AR device collected sensor data and formed the tracking map.

[0376] FIG. 31B depicts an example of a determined geographic location 1106 of a tracking map 1102, according to some embodiments. In the example shown, the determined geographic location 1106 includes a centroid point 1110 and an area 1108 surrounding the centroid point. It should be understood that the determination of geographic locations herein is not limited to the format shown. The determined geographic location may have any suitable format, including, for example, different area shapes. In this example, the geographic location is determined from area attributes using a database that associates area attributes with geographic locations. The database is commercially available, for example, a database that associates Wi-Fi fingerprints with locations, expressed as latitude and longitude, that can be used for the present operations.

[0377] 29 embodiment, the map database containing the environmental maps may also include location data for those maps, including the latitude and longitude covered by the maps. Processing in act 902 may involve selecting from that database a set of environmental maps that cover the same latitude and longitude determined for the area attribute of the tracking map.

[0378] Act 904 is a first filtering of the set of environment maps accessed in act 902. In act 902, environment maps are retained in the set based on their proximity to the geographic location of the tracking map. This filtering step may be performed by comparing the latitude and longitude associated with the tracking map and the environment maps in the set.

[0379] FIG. 32 depicts an example of act 904, according to some embodiments. Each area attribute may have a corresponding geographic location 1202. The set of environment maps may include environment maps with at least one area attribute that have a geographic location that overlaps with the determined geographic location of the tracking map. In the illustrated example, the set of identified environment maps includes environment maps CM1, CM2, and CM4, each having at least one area attribute that has a geographic location that overlaps with the determined geographic location of the tracking map 1102. Environment map CM3, associated with area attribute AA6, is not included in the set because it is outside the determined geographic location of the tracking map.

[0380] Other filtering steps may also be performed on the set of environmental maps to reduce / rank the number of environmental maps in the set that are ultimately processed (such as for map merging or for providing passable world information to a user device). Method 900 may include filtering the set of environmental maps (act 906) based on similarity of identifiers of one or more network access points associated with the environmental maps of the set of tracking maps and environmental maps. During map formation, devices that collect sensor data and generate the map may be connected to a network through a network access point, such as through Wi-Fi or a similar wireless communication protocol. The access point may be identified by a BSSID. A user device may connect to multiple different access points as it moves through an area, collects data, and forms the map. Similarly, when multiple devices provide information to form a map, the devices may be connected through different access points, and therefore, for similar reasons, there may be multiple access points used in forming the map. Thus, there may be multiple access points associated with a map, and the set of access points may be an indication of the map's location. The strength of the signal from the access point, which may be reflected as an RSSI value, may provide further geographic information. In some embodiments, the list of BSSIDs and RSSI values ​​may form an area attribute for a map.

[0381] In some embodiments, filtering the set of environment maps based on similarity of the one or more identifiers of the network access points may include retaining within the set of environment maps the environment map with the highest Jaccard similarity to at least one area attribute of the tracking map based on the one or more identifiers of the network access points. FIG. 33 depicts an example of act 906, according to some embodiments. In the example shown, a network identifier associated with area attribute AA7 may be determined as the identifier for tracking map 1102. The set of environment maps after act 906 includes environment map CM2, which may have an area attribute within a higher Jaccard similarity with AA7, and environment map CM4, which also includes area attribute AA7. Environment map CM1 is not included in the set because it has the lowest Jaccard similarity with AA7.

[0382] The processing in acts 902-906 may be performed without actually accessing the contents of the map stored in the map database based on metadata associated with the map. Other processing may involve accessing the contents of the map. Act 908 illustrates accessing the environment maps remaining in the subset after filtering based on the metadata. It should be understood that this act may be performed either earlier or later in the process, where subsequent operations may be performed using the accessed content.

[0383] Method 900 may include filtering the set of environment maps (act 910) based on similarity of metrics representing the content of the environment maps of the set of tracking maps and environment maps. The metrics representing the content of the tracking maps and environment maps may include vectors of values ​​calculated from the content of the maps. For example, deep keyframe descriptors as described above, calculated for one or more keyframes used in forming the maps, may provide metrics for comparison of maps or portions of maps. The metrics may be calculated from the maps retrieved in act 908, or may be pre-calculated and stored as metadata associated with the maps. In some embodiments, filtering the set of environment maps based on similarity of metrics representing the content of the environment maps of the set of tracking maps and environment maps may include retaining in the set of environment maps an environment map with a minimum vector distance between a vector of characteristics of the tracking map and a vector representing an environment map in the set of environment maps.

[0384] Method 900 may further include filtering the set of environment maps (act 912) based on the degree of match between the portion of the tracking map and the portion of the environment map of the set of environment maps. The degree of match may be determined as part of the localization process. As a non-limiting example, localization may be performed by identifying interest points in the tracking map and the environment map that are similar enough that they may represent the same portion of the physical world. In some embodiments, the interest points may be features, feature descriptors, keyframes, keyrigs, sustained poses, and / or PCFs. The set of interest points in the tracking map may then be matched to produce a best match with the set of interest points in the environment map. The mean squared distance between corresponding interest points may be calculated, and if it is below a threshold for a particular region of the tracking map, it is used as an indication that the tracking map and the environment map represent the same region of the physical world.

[0385] In some embodiments, filtering the set of environment maps based on the degree of match between the portion of the tracking map and the portion of the environment map of the set of environment maps may include calculating a volume of the physical world represented by the tracking map that is also represented in the environment map of the set of environment maps, and retaining in the set of environment maps environment maps with a calculated volume that is larger than the filtered-out environment map of the set. Figure 34 depicts an example of act 912, according to some embodiments. In the example shown, the set of environment maps after act 912 includes environment map CM4, which has area 1402 that matches an area of ​​tracking map 1102. Environment map CM1 is not included in the set because it does not have an area that matches an area of ​​tracking map 1102.

[0386] In some embodiments, the set of environment maps may be filtered in the order of act 906, act 910, and act 912. In some embodiments, the set of environment maps may be filtered based on act 906, act 910, and act 912, which may be performed in an order based on the processing required to perform the filtering from lowest to highest. Method 900 may include loading the set of environment maps and data (act 914).

[0387] In the illustrated example, the user database stores an area identification indicating an area in which the AR device was used. The area identification may be an area attribute, which may include parameters of a wireless network detected by the AR device during use. The map database may store multiple environment maps constructed from data provided by the AR device and associated metadata. The associated metadata may include an area identification derived from the area identification of the AR device that provided the data from which the environment map was constructed. The AR device may send a message to the PW module indicating that a new tracking map has been created or is being created. The PW module may calculate an area identifier for the AR device and update the user database based on the received parameters and / or the calculated area identifier. The PW module may also determine an area identifier associated with the AR device requesting the environment map, identify a set of environment maps from the map database based on the area identifier, filter the set of environment maps, and transmit the filtered set of environment maps to the AR device. In some embodiments, the PW module may filter the set of environmental maps based on one or more criteria including, for example, the geographic location of the tracking map, the similarity of one or more identifiers of network access points associated with the tracking map and the environmental maps of the set of environmental maps, the similarity of metrics representing the content of the tracking map and the environmental maps of the set of environmental maps, and the degree of matching between portions of the tracking map and portions of the environmental maps of the set of environmental maps.

[0388] While several aspects of some embodiments have been described above, it should be understood that various alterations, modifications, and improvements will readily occur to those skilled in the art. As an example, the embodiments are described in the context of an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein may be applied in MR environments, and more generally in other XR and VR environments.

[0389] As another example, embodiments are described in connection with devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), a discrete application, and / or any suitable combination of devices, networks, and discrete applications.

[0390] Additionally, Figure 29 provides examples of criteria that may be used to filter candidate maps and result in a set of highly ranked maps. Other criteria may be used instead of or in addition to the described criteria. For example, if multiple candidate maps have similar values ​​of a metric used to filter out less desirable maps, characteristics of the candidate maps may be used to determine which maps are retained as candidate maps or filtered out. For example, larger or denser candidate maps may be prioritized over smaller candidate maps. In some embodiments, Figures 27-28 may describe all or part of the systems and methods described in Figures 29-34.

[0391] 35 and 36 are schematic diagrams illustrating an XR system configured to rank and merge multiple environment maps, according to some embodiments. In some embodiments, a passable world (PW) may determine when to trigger ranking and / or merging of maps. In some embodiments, determining the map to be used may be based, at least in part, on deep keyframes, as described above in connection with FIGS. 21-25, according to some embodiments.

[0392] 37 is a block diagram illustrating a method 3700 of creating an environment map of a physical world, according to some embodiments. Method 3700 may begin with locating (act 3702) a tracking map captured by an XR device worn by a user relative to a set of reference maps (e.g., a reference map selected by the method of FIG. 28 and / or method 900 of FIG. 29). Act 3702 may include locating key rigs of the tracking map within the set of reference maps. The localization results for each key rig may include a located pose of the key rig and a set of 2D / 3D feature correspondences.

[0393] In some embodiments, method 3700 may include splitting the tracking map into connected components (act 3704), which may allow for robust merging of maps by merging connected fragments. Each connected component may include key features that are within a predetermined distance. Method 3700 may include merging connected components greater than a predetermined threshold into one or more reference maps (act 3706) and removing the merged connected components from the tracking map.

[0394] In some embodiments, method 3700 may include merging a group of reference maps that are merged with identical connected components of the tracking map (act 3708). In some embodiments, method 3700 may include promoting remaining connected components of the tracking map that are not merged with any reference map to the reference map (act 3710). In some embodiments, method 3700 may include merging a reference map that is merged with at least one connected component of the tracking map with a persistent attitude and / or PCF of the tracking map (act 3712). In some embodiments, method 3700 may include completing the reference map (act 3714), for example, by fusing map points and pruning redundant keyrings.

[0395] 38A and 38B illustrate an environment map 3800 created by updating a reference map 700, which may be promoted from the tracking map 700 (FIG. 7) with a new tracking map, according to some embodiments. As shown and described with respect to FIG. 7, the reference map 700 may provide a floor plan 706 of reconstructed physical objects in a corresponding physical world, represented by points 702. In some embodiments, the map points 702 may represent features of the physical objects, which may include multiple features. A new tracking map may be captured centered on the physical world, uploaded to the cloud, and merged with the map 700. The new tracking map may include the map point 3802 and key rigs 3804, 3806. In the illustrated example, the key rig 3804 represents a key rig that has been successfully located relative to the reference map, for example, by establishing a correspondence with the key rig 704 of the map 700 (as illustrated in FIG. 38B). Key rig 3806, on the other hand, represents a key rig that is not located relative to map 700. Key rig 3806 may, in some embodiments, be promoted to a separate reference map.

[0396] 39A-39F are schematic diagrams illustrating an example of a cloud-based persistent coordinate system that provides a shared experience for users in the same physical space. FIG. 39A shows, for example, a reference map 4814 from the cloud being received by XR devices worn by users 4802A and 4802B of FIGS. 20A-20C. The reference map 4814 may have a reference coordinate frame 4806C. The reference map 4814 may have a PCF 4810C with multiple associated PPs (e.g., 4818A, 4818B in FIG. 39C).

[0397] 39B shows that the XR device has established a relationship between its individual world coordinate system 4806A, 4806B and the reference coordinate frame 4806C. This may be done, for example, by locating the reference map 4814 on the individual device. Locating the tracking map relative to the reference map may result in a transformation for each device between its local world coordinate system and the coordinate system of the reference map.

[0398] 39C shows that as a result of the localization, a transformation can be calculated between the local PCFs (e.g., PCFs 4810A, 4810B) on each device and the respective persistent poses (e.g., PPs 4818A, 4818B) on the reference map (e.g., transformations 4816A, 4816B). Using these transformations, each device can use its local PCF, which can be detected locally on the device by processing images detected with sensors on the device, determining a location relative to the local device, and displaying virtual content tied to PPs 4818A, 4818B or other persistent points on the reference map. Such an approach can accurately localize virtual content for each user, allowing each user to have the same experience of the virtual content within physical space.

[0399] FIG. 39D illustrates a persistent posture snapshot from the reference map to the local tracking map. As can be seen, the local tracking maps are interconnected via persistent postures. FIG. 39E illustrates that a PCF 4810A on a device worn by user 4802A is accessible in a device worn by user 4802B through a PP 4818A. FIG. 39F illustrates that tracking maps 4804A, 4804B and the reference 4814 may be merged. In some embodiments, some PCFs may be removed as a result of the merging. In the illustrated example, the merged map includes PCF 4810C of the reference map 4814 but does not include PCFs 4810A, 4810B of tracking maps 4804A, 4804B. PPs previously associated with PCFs 4810A, 4810B may be associated with PCF 4810C after the map merging.

[0400] [Example]

[0401] Figures 40 and 41 illustrate examples of using a tracking map by the first XR device 12.1 of Figure 9. Figure 40 is a two-dimensional representation of a three-dimensional first local tracking map (Map 1), according to some embodiments, which may be generated by the first XR device of Figure 9. Figure 41 is a block diagram illustrating uploading Map 1 from the first XR device to the server of Figure 9, according to some embodiments.

[0402] FIG. 40 illustrates map 1 and virtual content (content 123 and content 456) on a first XR device 12.1. Map 1 has an origin (origin 1). Map 1 includes several PCFs (PCFa-PCFd). From the perspective of the first XR device 12.1, PCFa is, for example, located at the origin of map 1 and has X, Y, and Z coordinates of (0,0,0), and PCFb has X, Y, and Z coordinates of (-1,0,0). Content 123 is associated with PCFa. In this example, content 123 has an X, Y, and Z relationship to PCFa of (1,0,0). Content 456 has a relationship to PCFb. In this example, content 456 has an X, Y, and Z relationship to PCFb of (1,0,0).

[0403] In Figure 41, the first XR device 12.1 uploads Map 1 to the server 20. In this example, the tracking map was stored as an initial reference map because the server does not store a reference map for the same area of ​​the physical world represented by the tracking map. The server 20 now has a reference map based on Map 1. The first XR device 12.1 has a reference map that is empty at this stage. The server 20, for purposes of discussion, in some embodiments does not contain any other maps other than Map 1. No maps are stored on the second XR device 12.2.

[0404] The first XR device 12.1 also transmits its Wi-Fi signature data to the server 20. The server 20 may use the Wi-Fi signature data to determine the general location of the first XR device 12.1 based on intelligence gathered from other devices that have previously connected to the server 20 or other servers, along with recorded GPS locations of such other devices. The first XR device 12.1 may now end its first session (see FIG. 8) and disconnect from the server 20.

[0405] FIG. 42 is a schematic diagram illustrating the XR system of FIG. 16 according to some embodiments, showing that after a first user 14.1 ends a first session, a second user 14.2 starts a second session using a second XR device of the XR system. FIG. 43A is a block diagram showing the start of a second session by a second user 14.2. The first user 14.1 is shown in phantom because the first session by the first user 14.1 has ended. The second XR device 12.2 begins recording objects. Various systems with varying granularity may be used by the server 20 to determine that the second session by the second XR device 12.2 is within the same vicinity of the first session by the first XR device 12.1. For example, Wi-Fi signature data, Global Positioning System (GPS) positioning data, GPS data based on Wi-Fi signature data, or any other data indicative of a location may be included in the first and second XR devices 12.1 and 12.2 to record the location. Alternatively, the PCF identified by the second XR device 12.2 may show similarity to the PCF of map 1.

[0406] As shown in FIG. 43B, the second XR device boots up and begins collecting data, such as images 1110, from one or more cameras 44, 46. As shown in FIG. 14, in some embodiments, an XR device (e.g., second XR device 12.2) may collect one or more images 1110, perform image processing, and extract one or more features / points of interest 1120. Each feature may be converted into a descriptor 1130. In some embodiments, the descriptor 1130 may be used to describe a key frame 1140, which may have an associated image position and orientation tied to it. One or more key frames 1140 may correspond to a single sustained pose 1150, which may be automatically generated after a threshold distance, e.g., 3 meters, from a previous sustained pose 1150. One or more sustained poses 1150 may correspond to a single PCF 1160, which may be automatically generated after a predetermined distance, e.g., every 5 meters. Over time, as the user continues to move around their environment and the XR device continues to collect more data, such as images 1110, additional PCFs (e.g., PCF3 and PCFs 4, 5) may be created. One or more applications 1180 may be launched on the XR device and provide virtual content 1170 to the XR device for presentation to the user. The virtual content may have an associated content coordinate frame, which may be placed relative to one or more PCFs. As shown in FIG. 43B, the second XR device 12.2 creates three PCFs. In some embodiments, the second XR device 12.2 may attempt to locate against one or more reference maps stored on the server 20.

[0407] In some embodiments, as shown in Figure 43C, the second XR device 12.2 may download the reference map 120 from the server 20. Map 1 on the second XR device 12.2 includes PCFa-d and origin 1. In some embodiments, the server 20 may have multiple reference maps for various locations and may determine that the second XR device 12.2 is within the same vicinity as the first XR device 12.1 during the first session and send the reference map for that vicinity to the second XR device 12.2.

[0408] FIG. 44 shows the second XR device 12.2 beginning to identify PCFs for the purpose of generating Map 2. The second XR device 12.2 has identified only a single PCF, namely, PCF1 and PCF2. The X, Y, and Z coordinates of PCF1 and PCF2 for the second XR device 12.2 may be (1,1,1). Map 2 has its own origin (Origin 2), which may be based on the head pose of Device 2 at device startup for the current head pose session. In some embodiments, the second XR device 12.2 may immediately attempt to locate Map 2 relative to the reference map. In some embodiments, Map 2 may be impossible to locate relative to the reference map (Map 1) (i.e., localization may fail) because the system does not recognize any or sufficient overlap between the two maps. Localization may be performed by identifying portions of the physical world represented in a first map that are also represented in a second map and calculating the transformation required to align those portions between the first and second maps. In some embodiments, the system may locate based on a PCF comparison between the local map and a reference map. In some embodiments, the system may locate based on a sustained pose comparison between the local map and a reference map. In some embodiments, the system may locate based on a keyframe comparison between the local map and a reference map.

[0409] FIG. 45 shows Map 2 after the second XR device 12.2 has identified additional PCFs (PCF1, PCF2, PCF3, PCF4, PCF5) for Map 2. The second XR device 12.2 again attempts to locate Map 2 relative to the reference map. Because Map 2 has now been extended to overlap with at least a portion of the reference map, the localization attempt will be successful. In some embodiments, the overlap between the local tracking map, Map 2, and the reference map may be represented by PCFs, persistent poses, keyframes, or any other suitable intermediate or derived constructs.

[0410] Additionally, the second XR device 12.2 associates content 123 and content 456 with PCFs 1, 2, and PCF3 of map 2. Content 123 has X, Y, and Z coordinates for PCFs 1 and 2 of (1, 0, 0). Similarly, the X, Y, and Z coordinates of content 456 for PCF3 in map 2 are also (1, 0, 0).

[0411] 46A and 46B illustrate the successful localization of map 2 relative to the reference map. Localization may be based on matching features in one map and another. Here, using an appropriate transformation, involving both a translation and a rotation of one map relative to the other, the overlapping area / volume / section of map 1410 represents the intersection with map 1 and the reference map. Because map 2 created PCFs 3, 4, and 5 before localization, and the reference map created PCFs a and c before map 2 was created, different PCFs were created to represent the same volume in real space (e.g., different maps).

[0412] As shown in Figure 47, the second XR device 12.2 expands Map 2 to include PCFs a-d from the reference map. The inclusion of PCFs a-d represents the localization of Map 2 relative to the reference map. In some embodiments, the XR system may perform an optimization step to remove duplicate PCFs in 1410, i.e., PCF3 and PCFs 4, 5, etc., from the overlapping area. After Map 2 is localized, placement of virtual content, such as content 456 and content 123, will be relative to the nearest updated PCF in updated Map 2. The virtual content will appear in the same real-world location to the user despite the changed PCF binding for the content and despite the updated PCF for Map 2.

[0413] As shown in Figure 48, the second XR device 12.2 continues to expand Map 2 as additional PCFs (PCFs f, g, and h) are identified by the second XR device 12.2, e.g., as the user walks around the real world. Also, note that Map 1 is not expanded in Figures 47 and 48.

[0414] 49, the second XR device 12.2 uploads Map 2 to the server 20. The server 20 stores Map 2 along with the reference map. In some embodiments, Map 2 may be uploaded to the server 20 when the session for the second XR device 12.2 ends.

[0415] The reference map in the server 20 now includes PCFi, which is not included in map 1 on the first XR device 12.1. The reference map on the server 20 can be expanded to include PCFi when a third XR device (not shown) uploads a map to the server 20 and such map includes PCFi.

[0416] In Figure 50, the server 20 merges Map 2 with the reference map to form a new reference map. The server 20 determines that PCFa-d are common to the reference map and Map 2. The server extends the reference map to include PCFe-h and PCF1, 2 from Map 2 to form the new reference map. The reference maps on the first and second XR devices 12.1 and 12.2 are based on Map 1 and become outdated.

[0417] In Figure 51, the server 20 transmits a new reference map to the first and second XR devices 12.1 and 12.2. In some embodiments, this may occur when the first XR device 12.1 and the second device 12.2 attempt to locate during a different or new or subsequent session. The first and second XR devices 12.1 and 12.2 proceed to locate their individual local maps (Map 1 and Map 2, respectively) against the new reference map, as described above.

[0418] As shown in FIG. 52, a head coordinate frame 96 or "head pose" is related to the PCF in map 2. In some embodiments, the map's origin, i.e., origin 2, is based on the head pose of the second XR device 12.2 at the start of the session. As PCFs are created during the session, they are placed relative to the world coordinate frame, i.e., origin 2. The PCF for map 2 serves as a persistent coordinate frame relative to the reference coordinate frame, which may be the world coordinate frame of the previous session (e.g., origin 1 of map 1 in FIG. 40). These coordinate frames are related by the same transformation used to locate map 2 relative to the reference map, as discussed above in connection with FIG. 46B.

[0419] The transformation from the world coordinate frame to the head coordinate frame 96 was described above with reference to Figure 9. The head coordinate frame 96 shown in Figure 52 has only two orthogonal axes at a particular coordinate location relative to the PCF of map 2 and at a particular angle relative to map 2. However, it should be understood that the head coordinate frame 96 is in a three-dimensional location relative to the PCF of map 2 and has three orthogonal axes in three-dimensional space.

[0420] In FIG. 53, the head coordinate frame 96 is moving relative to the PCF of map 2. The head coordinate frame 96 is moving because the second user 14.2 is moving their head. A user can move their head in six degrees of freedom (6 dof). The head coordinate frame 96 can therefore move in 6 dof, i.e., in three dimensions from its previous location in FIG. 52, and in approximately three orthogonal axes relative to the PCF of map 2. The head coordinate frame 96 is adjusted as the real object detection camera 44 and inertial measurement unit 48 in FIG. 9 detect movement of the real object and head unit 22, respectively. For more information regarding head pose tracking, see "Enhanced Pose Tracking." No. 16 / 221,065, entitled "Determination for Display Device," which is incorporated herein by reference in its entirety.

[0421] FIG. 54 shows that a sound may be associated with one or more PCFs. A user may, for example, wear headphones or earphones with stereophonic sound. The location of the sound through the headphones can be simulated using conventional techniques. The location of the sound may be located at a constant position such that as the user rotates their head to the left, the location of the sound rotates to the right, thus causing the user to perceive the sound as originating from the same location in the real world. In this example, the location of the sound is represented by sound 123 and sound 456. For purposes of discussion, FIG. 54 is similar in its analysis to FIG. 48. When first and second users 14.1 and 14.2 are located in the same room, at the same or different times, they perceive sound 123 and sound 456 as originating from the same location in the real world.

[0422] 55 and 56 illustrate further implementations of the techniques described above. A first user 14.1 started a first session as described with reference to FIG. 8. As shown in FIG. 55, the first user 14.1 ended the first session, as indicated by the phantom line. At the end of the first session, the first XR device 12.1 uploaded Map 1 to the server 20. The first user 14.1 then started a second session, this time at a time after the first session. The first XR device 12.1 does not download Map 1 from the server 20 because Map 1 is already stored on the first XR device 12.1. If Map 1 is lost, the first XR device 12.1 downloads Map 1 from the server 20. The first XR device 12.1 then proceeds to build a PCF for Map 2, locate relative to Map 1, and further develop a reference map, as described above. Map 2 of the first XR device 12.1 is then used to associate local content, head coordinate frames, local sounds, etc. as explained above.

[0423] 57 and 58, it is also possible for more than one user to interact with the server in the same session. In this example, a first user 14.1 and a second user 14.2 are joined by a third user 14.3 with a third XR device 12.3. Each XR device 12.1, 12.2, and 12.3 begins to generate its own map, namely, Map 1, Map 2, and Map 3, respectively. As XR devices 12.1, 12.2, and 12.3 continue to develop Maps 1, 2, and 3, the maps are progressively uploaded to the server 20. The server 20 merges Maps 1, 2, and 3 to form a reference map. The reference map is then transmitted from the server 20 to each of the XR devices 12.1, 12.2, and 12.3.

[0424] FIG. 59 illustrates aspects of a viewing method for restoring and / or resetting head pose, according to some embodiments. In the illustrated example, in act 1400, the viewing device is powered on. In act 1410, in response to being powered on, a new session is started. In some embodiments, the new session may include establishing a head pose. One or more capture devices on a head-mounted frame affixed to the user's head capture the surfaces of the environment by first capturing images of the environment and then determining the surfaces from the images. In some embodiments, the surface data may also be combined with data from a gravity sensor to establish the head pose. Other suitable methods of establishing head pose may be used.

[0425] In act 1420, the viewing device's processor enters a routine for head pose tracking. The capture device continues to capture surfaces in the environment and determine the orientation of the head-mounted frame relative to the surfaces as the user moves their head.

[0426] At act 1430, the processor determines whether the head pose has been lost. A head pose may be lost due to "edge" cases such as too many reflective surfaces, low light, bare walls, outdoors, etc., which may result in low feature acquisition, or due to dynamic cases such as a crowd that moves and forms part of the map. The routine at 1430 allows a certain amount of time to pass, for example, 10 seconds, to allow sufficient time to determine whether the head pose has been lost. If the head pose has not been lost, the processor returns to 1420 and begins tracking the head pose again.

[0427] If the head pose is lost in act 1430, the processor enters a routine to restore the head pose in 1440. If the head pose is lost due to low light, a message such as the following message is displayed to the user through the display of the viewing device:

[0428] The system is detecting low light conditions. Move to an area with more light.

[0429] The system will continue to monitor whether sufficient light is available and whether head pose can be restored. Alternatively, the system may determine that low texture on the surface is causing head pose to be lost, in which case the user will be given the following prompt in the display as a suggestion to improve surface capture:

[0430] The system is unable to detect sufficient surfaces with fine texture. Move to an area where the surface is less rough and has a more refined texture.

[0431] In act 1450, the processor enters a routine to determine whether head pose reconstruction failed. If head pose reconstruction did not fail (i.e., head pose reconstruction was successful), the processor returns to act 1420 by again entering head pose tracking. If head pose reconstruction failed, the processor returns to act 1410 and establishes a new session. As part of the new session, all cached data is invalidated and the head pose is newly established thereafter. Any suitable method of head tracking may be used in combination with the process described in FIG. 59. U.S. Patent Application No. 16 / 221,065 describes head tracking and is incorporated herein by reference in its entirety. Remote Location

[0432] Various embodiments may utilize remote resources to facilitate persistent and consistent cross-reality experiences between individual users and / or groups of users. The inventors recognize and appreciate that the benefits of operating an XR device with reference maps as described herein may be achieved without downloading a set of reference maps. FIG. 30, discussed above, illustrates an example implementation in which reference maps would be downloaded to the device. The benefits of not downloading maps may be achieved, for example, by transmitting feature and pose information to a remote service that maintains a set of reference maps. According to one embodiment, a device seeking to use reference maps to position virtual content at a defined location relative to the reference maps may receive one or more transformations between the features and the reference maps from the remote service. Those transformations may be used on a device that maintains information about the location of those features in the physical world to position virtual content at a defined location relative to the reference maps or otherwise identify a location in the physical world defined relative to the reference maps.

[0433] In some embodiments, spatial information is captured by the XR device and communicated to a remote service, such as a cloud-based service, which uses the spatial information to locate the XR device relative to a reference map used by applications or other components of the XR system and to define the location of virtual content relative to the physical world. Once located, a transform can be communicated to the device that links a tracking map maintained by the device to the reference map. The transform, in conjunction with the tracking map, may be used to determine a location into which to render virtual content defined relative to the reference map, or otherwise identify a location in the physical world defined relative to the reference map.

[0434] The inventors recognize that the data required to be exchanged between a device and a remote location service may be very minimal compared to communicating map data, such as may occur when a device communicates tracking maps to a remote service and receives a set of reference maps for device-based location from that service. In some embodiments, performing location functions on cloud resources requires only a small amount of information to be transmitted from the device to the remote service. For example, it is not a requirement that a complete tracking map be communicated to the remote service to perform location. In some embodiments, feature and attitude information, such as may be stored in association with a persistent attitude, may be transmitted to a remote server, as described above. In embodiments in which features are represented by descriptors, as described above, the uploaded information may be even more minimal.

[0435] The results returned from the location service to the device may be one or more transformations that relate the uploaded features to matching portions of the reference map. These transformations, in conjunction with the tracking map, may be used within the XR system to identify the location of the virtual content or otherwise identify a location in the physical world. As described above, in embodiments where persistent spatial information such as PCFs is used to define a location relative to the reference map, the location service may download to the device the transformations between the features and one or more PCFs after successful location.

[0436] As a result, the network bandwidth consumed by communications between the XR device and the remote service for performing localization can be small. The system can therefore support frequent localization, enabling each device interacting with the system to quickly obtain information for locating virtual content or performing other location-based functions. As a device moves through the physical environment, it may repeat requests for updated localization information. In addition, a device may frequently obtain updates to its localization information, such as through the merging of additional tracking maps, to expand the map or increase its accuracy, such as when the reference map changes.

[0437] Furthermore, uploading features and downloading transformations may improve privacy in XR systems that share map information among multiple users by increasing the difficulty of retrieving a map through deception. Unauthorized users can be prevented from retrieving maps from the system, for example, by sending a false request for a reference map representing a portion of the physical world in which the unauthorized user is not located. Unauthorized users will likely not have access to features in an area of ​​the physical world about which they are requesting map information if they are not physically present in that area. In embodiments in which feature information is formatted as feature descriptions, the difficulty of spoofing feature information in a request for map information will be increased. Furthermore, when the system returns transformations intended to be applied to a tracking map of a device operating in an area for which location information is requested, the information returned by the system is likely to be of little or no use to an imposter.

[0438] According to one embodiment, the location service is implemented as a cloud-based microservice. In some examples, implementing a cloud-based location service can help conserve device computational resources and enable the calculations required for location to be performed with very low latency. These operations can be supported by the nearly infinite computational power or other computing resources available by providing additional cloud resources, ensuring the scalability of the XR system to support a large number of devices. In one example, many reference maps can be maintained in memory for near-instant access, or alternatively, stored in a highly available device, reducing system latency.

[0439] Additionally, implementing location determination for multiple devices within a cloud service can enable refinements to the process. Location determination telemetry and statistics can provide information about which reference maps are in active memory and / or high-availability storage. Statistics for multiple devices may be used, for example, to identify the most frequently accessed reference maps.

[0440] Additional accuracy may also be achieved as a result of processing in a cloud environment or other remote environment, using substantial processing resources relative to the remote device. For example, localization may be performed on a denser reference map in the cloud relative to processing performed on the local device. The map may be stored in the cloud, for example, with more PCFs or a denser feature descriptor, increasing the accuracy of the match between the set of features from the device and the reference map.

[0441] FIG. 61 is a schematic diagram of an XR system 6100. During a user session, the user device displaying the cross-reality content can appear in a variety of forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As discussed above, these devices can be configured with software, such as applications or other components, and / or wired to generate local location information (e.g., a tracking map) that can be used to render virtual content on its respective display.

[0442] The virtual content positioning information may be defined relative to global location information, which may be formatted, for example, as a reference map containing one or more PCFs. According to some embodiments, the system 6100 is configured with cloud-based services that support the functionality and display of virtual content on user devices.

[0443] In one example, the location functionality is provided as a cloud-based service 6106, which may be a microservice. The cloud-based service 6106 may be implemented on any of a number of computing devices, from which computing resources may be allocated to one or more services running in the cloud. The computing devices may be accessible and interconnected to each other and to devices such as the wearable XR dev...

Claims

1. 1. An XR system that supports specification of the location of virtual content relative to a stored map in a database of stored maps, the system comprising: one or more computing devices configured for network communication with one or more portable electronic devices; a communication component configured to receive, from a portable electronic device, information about a set of features within a three-dimensional (3D) environment of the portable electronic device and positioning information for features of the received set of features represented in a first coordinate frame; a location component coupled to the communication component, the location component comprising: selecting a stored map from the database of stored maps based on the selected map having a set of features that matches the received set of features, the selected map comprising a second coordinate frame; generating transformation data between the first coordinate frame and the second coordinate frame based on a calculated match between the received set of features in the 3D environment of the portable electronic device and the matching set of features in the selected map; transmitting the transformed data to the portable electronic device; a location component configured to: one or more computing devices, A system comprising:

2. the stored map comprises a plurality of persistent map features; The XR system of claim 1 , configured to express the generated transformation data in a form that defines a translation and a form that defines a rotation.

3. 10. The system of claim 1, wherein the location component is configured to select the stored map from the database of stored maps by filtering maps in the database of stored maps based at least in part on location information associated with the portable electronic device.

4. The system of claim 3 , wherein the location information comprises one or more of a radio fingerprint and / or GPS coordinates received from the portable electronic device.

5. The location component: maintaining information about individual portable electronic devices, the maintained information comprising location history; further configured to: The system of claim 3 , wherein the location information comprises an individual location history for the portable electronic device.

6. The stored map is partitioned into a plurality of areas; 2. The system of claim 1, wherein the location component is configured to select a stored map by selecting an area of ​​the stored map having a set of features that matches the received set of features.

7. the information about the received feature set comprises a descriptor calculated for the received feature set; The location component: identifying a candidate set of maps having feature sets with a greater than threshold number of features with descriptors that match descriptors of features in the received feature set; selecting the stored map from the candidate set of maps based on an error metric associated with a calculated transformation between the received set of features and a set of features of a selected candidate map; The system of claim 1 , configured to select the stored map having a set of features that matches the received set of features by:

8. 8. The system of claim 7, wherein the localization component is configured to generate an indication of localization failure based on a search of the database of stored maps for maps with the error metric below a threshold returning no maps with the error metric below the threshold.

9. 2. The system of claim 1, wherein the localization component is further configured to maintain state information for each of a plurality of portable electronic devices, the state information comprising, for each portable electronic device, at least one or any combination of a device ID, a tracking map ID corresponding to the first coordinate frame, a previously generated map reference, and / or a transformation of a reference map and a tracking map.

10. the received information about the received feature set comprises information about a plurality of feature sets; 10. The system of claim 1, wherein the location component is configured to select stored maps by selecting stored maps having a set of features that matches more than a threshold number of the received set of features.

11. 1. A method of operating a portable electronic device for rendering virtual content in a 3D environment, said method comprising: generating, on the portable electronic device, a local coordinate frame based on outputs of one or more sensors on the portable electronic device; generating, on the portable electronic device, a plurality of descriptors for a plurality of features sensed in the 3D environment; transmitting the descriptors of the features and location information of the features expressed in the local coordinate frame to a location service over a network; obtaining, from the location service, transformation data between the stored coordinate frame of stored spatial information about the 3D environment and the local coordinate frame; receiving a specification of a virtual object having a virtual object coordinate frame and a location of the virtual object relative to the stored coordinate frame; rendering the virtual object on a display of the portable electronic device at a location determined based at least in part on the transformation data and the received location of the virtual object; A method comprising:

12. 12. The method of claim 11 , wherein transmitting over the network comprises transmitting the plurality of descriptors of the plurality of features and location information of the plurality of features expressed in the local coordinate frame over the network to a cloud-hosted location service.

13. 12. The method of claim 11, wherein obtaining transformation data between the stored coordinate frame and the local coordinate frame of stored spatial information about the 3D environment from the location service includes obtaining the stored coordinate frame through an application programming interface (API).

14. the portable electronic device comprises a first portable electronic device that is part of a system and that comprises a first processor; the system further comprising a second portable electronic device comprising a second processor; The first and second processors each include: obtaining transformation data between the individual local coordinate frames and the same stored coordinate frame; receiving a specification of the virtual object; rendering the virtual object on a separate display for each user of the first and second portable electronic devices; The method according to any one of claims 11 to 13, wherein

15. Executing an application to generate a specification of the virtual object and its location relative to the stored coordinate frame for rendering in the local coordinate frame.

15. The method of claim 14, further comprising:

16. Maintaining a local coordinate frame on the portable electronic device includes, for each of the first and second portable electronic devices: capturing a plurality of images of the 3D environment from the one or more sensors of the portable electronic device; generating spatial information about the 3D environment based at least in part on the calculated one or more sustained poses; and Including, The method further includes transmitting, for each of the first and second portable electronic devices, the generated spatial information to a remote server; The method of claim 14 , wherein obtaining the transformation data comprises receiving the transformation data from a cloud-hosted location service.

17. The first and second portable electronic devices each include: a download system configured to download the stored coordinate frame from a server and enable a locally performed transformation from the local to the stored coordinate frame in a device-based localization mode; The method of claim 14, comprising:

18. 1. A method of operating a portable electronic device for rendering virtual content within a three-dimensional (3D) environment comprising a framework for an XR system that provides a shared experience to each of a plurality of users through the use of a coordinate frame of reference, the method comprising: generating a local coordinate frame on the portable electronic device based on outputs of one or more sensors on the portable electronic device and a tracking map constructed on the portable electronic device; generating, on the portable electronic device, a plurality of features for a plurality of images of the 3D environment; transmitting from the portable electronic device, via a network, to a location service an indication of the plurality of features and location information for the plurality of features expressed in the local coordinate frame; receiving transformation data from the location service, the transformation data indicating a relationship between the local coordinate frame and a second coordinate frame; A method comprising:

19. capturing information about the 3D environment from one or more sensors, the captured information comprising the plurality of images; and generating a map of at least a portion of the 3D environment based on the plurality of images; 20. The method of claim 18, further comprising:

20. 20. The method of claim 18, wherein the portable electronic device further comprises a display, the method further comprising rendering virtual content having a location defined in the second coordinate frame on the display at a position calculated based on the transformation data.

21. The method further includes generating a descriptor for the plurality of features; The method of claim 19 , wherein transmitting information about the plurality of features comprises transmitting the descriptors for the plurality of features.

22. 20. The method of claim 18, further comprising sending, via the network, to the location service a request to initiate a session with the location service.

23. 23. The method of claim 22, wherein the method further comprises receiving a request to initiate a session with the location service, the request being sent along with an identifier of the portable electronic device and calibration data for the portable electronic device.

24. transmitting, from the portable electronic device, via a network, to a location service, an indication of the plurality of features and location information related to the plurality of features expressed in the local coordinate frame, comprising a request for a location service; The method of claim 18 , further comprising transmitting a request for location in response to one or more types of trigger conditions being met.

25. The method further includes storing information about the plurality of features and position information related to the plurality of features in a buffer; 25. A method according to any one of claims 18 to 24, wherein transmitting over a network comprises transmitting the contents of the buffer together such that information about the plurality of features and location information relating to the plurality of features are transmitted together.

26. 26. The method of claim 25, wherein the buffer comprises an adjustable size, the method further comprising increasing the size of the buffer in response to a failure indication.

27. extracting the plurality of features includes extracting a threshold number of features from each image of the plurality of images; 27. The method of claim 26, further comprising increasing the threshold number in response to a failure indication received from the location service via the network.

28. 25. The method of claim 24, further comprising determining that the one or more types of trigger conditions have been met in response to identifying a distance traveled by the portable electronic device since a last successful request for location.

29. 1. A method of operating a system for rendering virtual content within a three-dimensional (3D) environment comprising an XR system framework that provides a shared experience to each of a plurality of users through the use of stored coordinate frames, the method comprising: receiving, at a cloud-hosted location service of the system, information about a set of features within the 3D environment of a portable electronic device and positioning information related to features of the received set of features represented in a first coordinate frame; selecting a stored map from a database of stored maps based on the selected map having a set of features that matches the received set of features, the selected map comprising a second coordinate frame; generating transformation data between the first coordinate frame and the second coordinate frame based on a calculated match between the received set of features in the 3D environment of the portable electronic device and the matching set of features in the selected map; transmitting the transformed data to the portable electronic device; A method comprising:

30. The method comprises: receiving, from an application, a specification of a location of virtual content relative to a persistent map feature, the stored map comprising a plurality of persistent map features; expressing the generated transformation data in a form that defines a translation and a form that defines a rotation; 30. The method of claim 29, further comprising:

31. 30. The method of claim 29, wherein the method further comprises selecting the stored map from the database of stored maps by filtering maps in the database of stored maps based at least in part on location information associated with the portable electronic device.

32. 32. The method of claim 31 , wherein the location information comprises one or more of a radio fingerprint and / or GPS coordinates received from the portable electronic device.

33. The method comprises: maintaining information about individual portable electronic devices, the maintained information comprising location history; further comprising 32. The method of claim 31, wherein the location information comprises an individual location history for the portable electronic device.

34. The method comprises: Partitioning the stored map into a plurality of areas; selecting, by the location service, a stored map by selecting an area of ​​the stored map having a set of characteristics that matches the set of received characteristics; 30. The method of claim 29, further comprising:

35. The information about the received feature set comprises a descriptor calculated for the received feature set, and the method further comprises: identifying a candidate set of maps having feature sets with a greater than threshold number of features with descriptors that match descriptors of features in the received feature set; selecting the stored map from the candidate set of maps based on an error metric associated with a calculated transformation between the received set of features and a set of features of a selected candidate map; 30. The method of claim 29, further comprising selecting the stored map having a set of features that matches the received set of features by performing:

36. 36. The method of claim 35, further comprising generating an indication of a location failure based on a search of the database of stored maps for maps with the error metric below a threshold returning no maps with the error metric below the threshold.

37. 30. The method of claim 29, wherein the method further includes maintaining state information for each of a plurality of portable electronic devices, the state information comprising, for each portable electronic device, at least one or any combination of a device ID, a tracking map ID corresponding to the first coordinate frame, a previously generated map reference, and / or a transformation of a reference map and a tracking map.

38. 38. The method of claim 37, wherein the received information about the received feature set comprises information about a plurality of feature sets, the method further comprising selecting a stored map by selecting a stored map having feature sets that match more than a threshold number of the received feature sets.

39. 36. The method of claim 35, wherein the method further comprises obtaining or deriving individual feature descriptors from any one or more or any combination of frames, portions of frames, key frames, sustained poses, or sustained coordinate frames (PCFs).

40. 40. The method of claim 29, further comprising obtaining second transformation data from the cloud-hosted location service, the second transformation data defining a transformation between a stored coordinate frame of stored spatial information about the 3D environment and a local coordinate frame of the device.

41. 41. The method of claim 40, further comprising receiving a specification of a virtual object having a virtual object coordinate frame and a location of the virtual object relative to the stored coordinate frame.

42. 42. The method of claim 41 , further comprising rendering the virtual object on a display of the portable electronic device at a location determined based at least in part on the calculated transformation and the received location of the virtual object.

43. 41. The method of claim 40, wherein matching a descriptor for the received set of features to a descriptor for the set of features of the stored map includes matching at least a first persistent coordinate frame (PCF) of the local coordinate frame to a first PCF of the stored map.

44. 44. The method of claim 43, wherein matching the descriptors for the received set of features to the descriptors for the set of features of the stored map includes at least matching a second PCF of the local coordinate frame to a second PCF of the stored map.

Citation Information

Patent Citations

  • Retrieval image registration device, retrieval image display system, retrieval image registration method and program

    JP2013210974A

  • Image processing device, image processing method, and program

    JP2014200097A

  • Range calibration of binocular optical augmented reality system

    JP2015142383A

  • Enable augmented reality using eye tracking

    JP2016509705A

  • Virtual map display system, method and program

    JP2018200358A