Cross Reality System with Location Services and Location-Based Shared Content

Through network resources in a distributed computing environment, a method for multiple portable electronic devices to render virtual content in a 3D environment is provided, which solves the synchronization problem of virtual content rendering and device positioning in the prior art, and realizes an efficient cross-reality system.

CN114730546BActive Publication Date: 2025-05-30MAGIC LEAP INC
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202080078518.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-12
Filing Date
2020-11-11
Publication Date
2025-05-30
Estimated Expiration
2040-11-11

AI Technical Summary

Technical Problem

Existing cross-reality systems have difficulty effectively rendering shared location-based virtual content in a 3D environment, and it is difficult to achieve synchronous positioning and content rendering between multiple portable electronic devices.

Method used

Through network resources in a distributed computing environment, a method of providing multiple portable electronic devices to render virtual content in a 3D environment. The method includes using a processor and a computer-readable medium to generate storage maps and data structures for representing virtual content areas in a 3D environment, and providing links to positioning information and virtual content through a positioning service.

Benefits of technology

Multiple portable electronic devices are implemented to synchronously render and share location-based virtual content in a 3D environment, improving the real-time and synergy of cross-reality systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114730546B_ABST
    Figure CN114730546B_ABST
Patent Text Reader

Abstract

A cross-reality system enables any one of a plurality of devices to effectively render shared location-based content. The cross-reality system can include a cloud-based service that responds to requests from the devices to locate themselves relative to a stored map. The service can return information to the devices that locates the devices relative to the stored map. In combination with the location information, the service can provide information related to locations in the physical world that are near the location of a device at which virtual content has been provided. Based on the information received from the service, the devices can render or cease rendering virtual content to each of a plurality of users based on the location of the user and the specified location of the virtual content.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 934,485, filed on November 12, 2019, entitled "CROSS REALITY SYSTEM WITH LOCALIZATION SERVICE AND SHARED LOCATION - BASED CONTENT", under 35 U.S.C. § 119(e). The entire content of this application is hereby incorporated by reference. Field of the technology

[0003] This application generally relates to cross - reality systems. Background art

[0004] Computers can control human - user interfaces to create cross - reality (XR) environments, where some or all of the XR environment perceived by the user is generated by the computer. These XR environments can be virtual - reality (VR) environments, augmented - reality (AR) environments, and mixed - reality (MR) environments, where some or all of the XR environment can be generated by the computer using data that describes the environment in part. This data can, for example, describe virtual objects that can be rendered in such a way that the user feels or perceives the virtual objects as part of the physical world and can interact with the virtual objects. Due to the rendering and presentation of data through a user - interface device (such as a head - mounted display device), the user can experience these virtual objects. The data can be displayed for the user to view, or can control the audio played for the user to listen to, or can control a haptic (or tactile) interface, enabling the user to experience the tactile sensations felt or perceived when the user experiences the virtual objects.

[0005] XR systems can be used in many applications, covering fields such as scientific visualization, medical training, engineering design and prototyping, teleoperation and telepresence, and personal entertainment. Compared with VR, AR and MR include one or more virtual objects related to real objects in the physical world. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment of using XR systems and also opens the door to various applications for presenting realistic and understandable information about how the physical world might be changed.

[0006] To realistically render virtual content, an XR system can construct a representation of the physical world around the user of the system. For example, the representation can be constructed by processing images obtained using sensors on wearable devices that form part of the XR system. In such a system, the user can perform an initialization routine by looking at the room or other physical environment in which the user intends to use the XR system until the system has obtained enough information to construct a representation of the environment. When the system is running and the user moves around in the environment or moves to another environment, sensors on the wearable device may obtain additional information to extend or update the representation of the physical world. SUMMARY OF THE INVENTION

[0007] Some aspects of the present application relate to methods and apparatuses for providing cross-reality (XR) scenarios. The techniques described herein can be used together, used individually, or used in any suitable combination.

[0008] According to some aspects, there is provided a network resource within a distributed computing environment for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a 3D environment. The resource includes: one or more processors and at least one computer-readable medium, the at least one computer-readable medium including a plurality of stored maps of the 3D environment and a plurality of data structures, each data structure of the plurality of data structures representing a corresponding area in the 3D environment in which virtual content is to be displayed. Each data structure of the plurality of data structures includes: information associating the data structure with a location in the plurality of stored maps; and a link to virtual content for rendering in the corresponding area in the 3D environment. The computer-readable medium further includes computer-executable instructions. When executed by one of the one or more processors, these instructions implement a service for providing location information to a portable electronic device among the plurality of portable electronic devices, wherein the location information indicates the location of the plurality of portable electronic devices relative to one or more shared maps; and selectively providing a copy of at least one of the plurality of data structures to the portable electronic device based on the location of the portable electronic device relative to an area represented by the plurality of data structures.

[0009] According to some embodiments, when executed by the processor, the computer-executable instructions further implement an authentication service for determining the access rights of the portable electronic device. Additionally, the computer-executable instructions for selectively providing the at least one data structure to the portable electronic device can determine whether to send the at least one data structure, at least in part, based on the access rights of the portable electronic device and access attributes associated with the at least one data structure.

[0010] According to some embodiments, in the network resources, each of the plurality of data structures further includes a common attribute. Additionally, the computer-executable instructions selectively provided to the portable electronic device for the at least one data structure can determine whether to send the at least one data structure, at least in part, based on the common attribute of the at least one data structure.

[0011] According to some embodiments, for a portion of the plurality of data structures, the link to the virtual content includes a link to an application that provides the virtual content.

[0012] According to some embodiments, each of the plurality of data structures further includes display characteristics of a prism on the portable electronic device. The prism is a volume in which the virtual content linked to the data structure is displayed.

[0013] According to some embodiments, the display characteristics can include the size of the prism.

[0014] According to some embodiments, the display characteristics include the behavior of the virtual content rendered within the prism relative to a physical surface.

[0015] According to some embodiments, the display characteristics include one or more of the following: an offset of the prism relative to a persistent location associated with a map, a spatial orientation of the prism, the behavior of the virtual content rendered within the prism relative to the position of the portable electronic device, and the behavior of the virtual content rendered within the prism relative to the direction the portable electronic device is facing.

[0016] According to some aspects, a method of operating a portable electronic device to render virtual content in a 3D environment is provided. The method includes using one or more processors: generating a local coordinate system on the portable electronic device based on an output of one or more sensors on the portable electronic device; generating information indicative of a position in the 3D environment on the portable electronic device based on the output of the one or more sensors and an indication of a position in the local coordinate system; sending, via a network, the information indicative of the position and the indication of the position in the local coordinate system to a positioning service; obtaining, from the positioning service, a transformation between a coordinate system of storage space information about the 3D environment and the local coordinate system; obtaining, from the positioning service, one or more data structures, each data structure representing a corresponding region in the 3D environment and virtual content for display in the corresponding region; and rendering the virtual content represented in the one or more data structures in the corresponding regions of the one or more data structures.

[0017] According to some embodiments, rendering virtual content in the corresponding region includes: creating a prism with a set of parameters based on the data structure representing the corresponding region.

[0018] According to some embodiments, the virtual content is represented as an indicator of the location of the virtual content on the network in at least one of the one or more data structures.

[0019] According to some embodiments, rendering the virtual content includes: executing an application that generates the virtual content on the portable electronic device.

[0020] According to some embodiments, rendering the virtual content further includes: determining whether the application is currently installed on the portable electronic device, and based on determining that the application is not currently installed, downloading the application to the portable electronic device.

[0021] According to some embodiments, the method further includes: detecting that the portable electronic device has left the region represented by the data structure in the one or more data structures; and based on the detection, deleting the virtual content represented in the data structure.

[0022] According to some embodiments, the one or more received data structures include a first set of data structures, and the first set of data structures is received at a first time. Additionally, the method may further include: storing rendering information associated with a first data structure in the first set of data structures; receiving a second set of data structures at a second time after the first time; and based on determining that the first data structure is not included in the second set, deleting the rendering information associated with the first data structure.

[0023] According to some aspects, an electronic device configured to operate within a cross-reality system is provided. The electronic device includes: one or more sensors configured to capture information about a three-dimensional (3D) environment, the captured information including a plurality of images; at least one processor; and at least one computer-readable medium storing computer-executable instructions. The computer instructions, when executed on a processor in the at least one processor, cause the instructions to: maintain a local coordinate system for representing a position in the 3D environment based on at least a first portion of the plurality of images; manage a prism associated with one or more applications such that virtual content generated by an application in the one or more applications is rendered within the prism; and send information derived from outputs of the one or more sensors to a service via a network. The instructions also receive from the service: location information; and a data structure representing corresponding virtual content and a region in the 3D environment for rendering the virtual content. The instructions also include creating a prism associated with the data structure to render the corresponding virtual content within the prism.

[0024] According to some embodiments, the computer-executable instructions further include computer-executable instructions for performing the following operations: obtaining the corresponding virtual content based on information in the data structure; and rendering the obtained virtual content within the prism.

[0025] According to some embodiments, obtaining the corresponding virtual content based on information in the data structure includes: accessing the corresponding virtual content via a network based on a virtual content location indicator in the data structure.

[0026] According to some embodiments, obtaining the corresponding virtual content based on information in the data structure includes: accessing the corresponding virtual content via the network based on the virtual content location indicator in the data structure.

[0027] According to some embodiments, the computer-executable instructions further include instructions for performing the following operations: detecting that the electronic device has left the region represented by the data structure; and based on the detection, deleting the prism associated with the data structure.

[0028] According to some embodiments, rendering the obtained virtual content within the prism further includes: using the coordinate system of the electronic device to determine a set of coordinates for rendering the virtual content in the 3D environment.

[0029] According to some aspects, a method for planning location-based virtual content for a cross-reality system is provided, the cross-display system being operable with a plurality of portable electronic devices capable of rendering virtual content in a 3D environment. The method includes: using one or more processors to generate a data structure representing an area in the 3D environment in which virtual content is to be displayed; using one or more processors to store in the data structure information indicating the virtual content to be rendered in the area in the 3D environment; and using one or more processors to associate the data structure with a location in a map for positioning the plurality of portable electronic devices into a shared coordinate system.

[0030] According to some embodiments, the method further includes: setting access rights to the data structure.

[0031] According to some embodiments, setting access rights to the data structure includes: indicating that the data structure can be accessed by one or more specific categories of users of the plurality of portable electronic devices.

[0032] According to some embodiments, storing in the data structure information indicating the virtual content to be rendered in the area in the 3D environment includes: specifying an application executable on the portable electronic device to generate the virtual content.

[0033] According to some embodiments, the method further includes: storing the data structure in association with a location service that uses the map to locate the plurality of portable electronic devices.

[0034] According to some embodiments, the method further includes: receiving a specification of the area in the 3D environment and the virtual content via a user interface.

[0035] According to some embodiments, the method further includes: receiving a specification of the area in the 3D environment and the virtual content from an application via a programming interface.

[0036] The above summary of the invention is provided by way of illustration and not by way of limitation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings are not necessarily to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by the same reference numeral. For clarity, not every component is labeled in every figure. In the drawings:

[0038] Figure 1 is a sketch showing an example of a simplified augmented reality (AR) scene according to some embodiments;

[0039] Figure 2is a sketch showing an exemplary simplified AR scenario demonstrating an exemplary use case of an XR system according to some embodiments;

[0040] Figure 3 is a schematic diagram showing the data flow of a single user in an AR system configured to provide a user with an experience of interacting with the physical world by AR content according to some embodiments;

[0041] Figure 4 is a schematic diagram showing an exemplary AR display system that displays virtual content to a single user according to some embodiments;

[0042] Figure 5A is a schematic diagram showing a user wearing an AR display system that renders AR content when the user moves in a physical world environment according to some embodiments;

[0043] Figure 5B is a schematic diagram showing an observation optical component and its accessories according to some embodiments;

[0044] Fig. 6A is a schematic diagram showing an AR system using a world reconstruction system according to some embodiments;

[0045] Figure 6B is a schematic diagram showing the components of an AR system that maintains a model of a connected world according to some embodiments;

[0046] Figure 7 is a schematic diagram of a tracking map formed by a device traversing a path through the physical world.

[0047] Figure 8 is a schematic diagram showing a user of a cross-reality (XR) system that perceives virtual content according to some embodiments;

[0048] Fig. 9 is a Figure 8 block diagram of the components of a first XR device of an XR system that transforms between coordinate systems according to some embodiments;

[0049] Fig.10 is a schematic diagram showing an exemplary transformation from an origin coordinate system to a destination coordinate system for correctly rendering local XR content according to some embodiments;

[0050] Fig.11 is a top view showing a pupil-based coordinate system according to some embodiments;

[0051] Fig.12 is a top view showing a camera coordinate system including all pupil positions according to some embodiments;

[0052] Fig.13 is according to some embodiments Fig. 9 Schematic diagram of a display system;

[0053] Fig.14 Is a block diagram showing the creation of a Persistent Coordinate Frame (PCF) and the attachment of XR content to the PCF according to some embodiments;

[0054] Fig.15 Is a flowchart showing a method of establishing and using a PCF according to some embodiments;

[0055] Fig.16 Is an XR system according to some embodiments including a second XR device Figure 8 Block diagram of;

[0056] Fig.17 Is a schematic diagram showing a room and key frames established for various areas in the room;

[0057] Fig.18 Is a schematic diagram showing the establishment of a key-frame-based persistent pose according to some embodiments;

[0058] Fig.19 Is a schematic diagram showing the establishment of a persistent coordinate frame (PCF) based on a persistent pose according to some embodiments;

[0059] Figures 20A to 20C Is a schematic diagram showing an example of creating a PCF according to some embodiments;

[0060] Fig.21 Is a block diagram showing a system for generating global descriptors for individual images and / or maps according to some embodiments;

[0061] Fig. 22 Is a flowchart showing a method of calculating an image descriptor according to some embodiments;

[0062] Fig.23 Is a flowchart showing a localization method using an image descriptor according to some embodiments;

[0063] Fig.24 Is a flowchart showing a method of training a neural network according to some embodiments;

[0064] Fig.25 Is a block diagram showing a method of training a neural network according to some embodiments;

[0065] Fig.26 Is a schematic diagram showing an AR system configured to rank and merge multiple environmental maps according to some embodiments;

[0066] Fig. 27is a simplified block diagram showing a plurality of canonical maps stored on a remote storage medium according to some embodiments;

[0067] Fig.28 is a schematic diagram showing a method of selecting a canonical map to locate a new tracking map and / or obtain a PCF from a canonical map, for example, in one or more canonical maps;

[0068] Fig.29 is a flowchart showing a method of selecting a plurality of ranked environmental maps according to some embodiments;

[0069] Fig.30 is showing according to some embodiments Fig.26 schematic diagram of an exemplary map ranking section of an AR system of;

[0070] Fig.31A is a schematic diagram showing an example of area attributes of a tracking map (TM) and an environmental map in a database according to some embodiments;

[0071] Fig.31B is showing according to some embodiments the determination for Fig.29 example of the geographical location of a tracking map (TM) for geographical location filtering of;

[0072] Fig.32 is showing according to some embodiments Fig.29 example of geographical location filtering of;

[0073] Fig.33 is showing according to some embodiments Fig.29 example of Wi-Fi BSSID filtering of;

[0074] Fig.34 is showing according to some embodiments the use of Fig.29 example of positioning of;

[0075] Fig.35 and Fig.36 is a block diagram of an XR system configured to rank and merge a plurality of environmental maps according to some embodiments;

[0076] Fig.37 is a block diagram showing a method of creating an environmental map of the physical world in a canonical form according to some embodiments;

[0077] Fig.38A and Fig.38B is a schematic diagram showing an environmental map created in a canonical form by updating a tracking map of Figure 7 with a new tracking map according to some embodiments.

[0078] FIG. 39A to FIG. 39F is a schematic diagram showing an example of a merged map according to some embodiments;

[0079] Fig.40 is a two-dimensional representation of a three-dimensional first local tracking map (ground Fig. 9 ) that can be generated by a first XR device of Figure 1 ;

[0080] Fig.41 is a block diagram showing uploading of the ground Figure 1 from the first XR device to a server of Fig. 9 according to some embodiments;

[0081] Fig.42 is a schematic diagram of an XR system showing that after a first user terminated a first session, a second user initiated a second session using a second XR device of the XR system according to some embodiments; Fig.16 ;

[0082] Fig.43A is a block diagram showing a new session of a second XR device of Fig.42 according to some embodiments;

[0083] Fig.43B is a block diagram showing creation of a tracking map for a second XR device of Fig.42 according to some embodiments;

[0084] Fig.43C is a block diagram showing downloading of a canonical map from the server to a second XR device of Fig.42 according to some embodiments;

[0085] Fig.44 is a schematic diagram showing an attempt to localize a second tracking map (ground Fig.42 ) that can be generated by a second XR device of Figure 2 to the canonical map according to some embodiments;

[0086] Fig.45 is a schematic diagram showing an attempt to localize a second tracking map (ground Figure 2 ) that can be further developed and has XR content associated with a PCF of ground Fig.44 to the canonical map according to some embodiments; Figure 2 ;

[0087] Figure 46A-46B is a schematic diagram showing successful localization of the ground Fig.45 of Figure 2 to the canonical map according to some embodiments;

[0088] Fig.47is a schematic diagram of a canonical map generated by incorporating one or more PCFs from a canonical map of Fig.46A into Fig.45 the Figure 2 ;

[0089] Fig.48 is a schematic diagram of a canonical map of Figure 2 further extended on a second XR device Fig.47 ;

[0090] Fig.49 is a block diagram showing uploading Figure 2 from a second XR device to a server according to some embodiments;

[0091] Fig.50 is a block diagram showing merging Figure 2 with a canonical map according to some embodiments;

[0092] Fig.51 is a block diagram showing sending a new canonical map from a server to a first XR device and a second XR device according to some embodiments;

[0093] Fig.52 is a block diagram showing a two - dimensional representation of Figure 2 and the head coordinate system of a second XR device with reference to Figure 2 ;

[0094] Fig.53 is a block diagram showing, in two - dimensional form, adjustments to the head coordinate system that can occur in six degrees of freedom according to some embodiments;

[0095] Fig.54 is a block diagram of a canonical map on a second XR device showing the positioning of sound relative to the PCF of Figure 2 ;

[0096] Fig.55 and Fig.56 are a perspective view and a block diagram showing use cases of an XR system when a first user terminates a first session and the first user initiates a second session using the XR system according to some embodiments;

[0097] Fig.57 and Fig.58 are a perspective view and a block diagram showing use cases of an XR system when three users use the XR system simultaneously in the same session according to some embodiments;

[0098] Fig.59 is a flowchart showing a method for restoring and resetting a head pose according to some embodiments;

[0099] Fig.60 is a block diagram of a machine in the form of a computer that can find application in the system of the present invention according to some embodiments;

[0100] Fig.61 is a schematic diagram of an example XR system in which any one of a plurality of devices can access a positioning service according to some embodiments;

[0101] Fig.62 is an example processing flow for operating a portable device as part of an XR system that provides cloud-based positioning according to some embodiments; and

[0102] Fig.63A 、 Fig.63B and Fig.63C is an example processing flow of cloud-based positioning according to some embodiments.

[0103] Fig.64 is a schematic diagram of a system for managing and displaying location-based shared virtual content in a physical environment, and a sketch of how exemplary content is presented to a user in the physical environment.

[0104] Fig.65 is a schematic diagram of an exemplary volumetric data structure and related data.

[0105] Fig.66 is a block diagram of an exemplary software architecture of a cross-reality device configured to acquire and display virtual content based on the position of the cross-reality device relative to a physical environment.

[0106] Fig.67 is a flowchart showing the interaction between system components for acquiring and displaying location-based shared content.

[0107] Fig.68 is an exemplary software architecture for configuring a device to work with a cross-reality system so that the device can acquire and render content from the cross-reality system.

[0108] Fig.69 is a sketch of an exemplary physical environment with positioning content shared so as to be perceivable by any one of a plurality of users in the physical environment. DETAILED DESCRIPTION

[0109] Methods and apparatus for providing a cross-reality (XR) scenario to any one of a plurality of users who may traverse the physical world are described herein. The system can enable a content manager to specify virtual content associated with a location in the physical world such that when a user wearing an XR device passes near the location, the XR device can render content for the user, causing the content to appear at the specified location in the physical world.

[0110] Such a system can be implemented to effectively operate services that provide virtual content and XR devices that interact with these services to render the virtual content to users. In some embodiments, the virtual content can be provided by a localization service that enables each of a plurality of XR devices to determine its position relative to a shared map. According to some embodiments, volumes in which location-based virtual content will be rendered can be defined in the shared map. Due to the localization process, when the service determines that an XR device is located at a position within the map that has such a volume associated with it, the service can provide the XR device with an indication of the volume and the content that will be displayed within the volume. The device can then use these indications to render the virtual content.

[0111] Information generated during localization can be used similarly for removing content. When an XR device has moved away from a position where a volume will be rendered, the device can delete information indicating the location and nature of the virtual content. The movement of the XR device can be determined by a service that specifies the content or a service on the device. Since localization can be performed repeatedly to support other functions of the XR system, the computational burden and network bandwidth for identifying and / or removing location-based virtual content can be less.

[0112] In some embodiments, information about location-based virtual content can be represented in an efficient format when passed from the service to the XR device. For example, the virtual content can be represented as a link to the virtual content or a link to the application that generates the virtual content. Thus, less network bandwidth is consumed by the communication of the content between the relevant service and the XR device, enabling the content information to be updated frequently. Additionally, the volume in which the virtual content will be displayed can correspond to the constructs used by the XR device to manage the rendering of content specified by applications executed on the XR device. An example of an application executed on the XR device is a prism. The XR device can use utilities that otherwise manage the prism to manage the rendering of location-based virtual content in combination with other virtual content. For example, when multiple applications specify virtual content for the same volume in the physical world, these utilities can determine the content to be rendered, associate user actions with the specific application providing the virtual content to be rendered, and, when the virtual content is no longer being rendered, remove the virtual content and associated data from the XR device.

[0113] The localization process (which can be used to identify location-based virtual content) can be used for some functions of the XR system, such as providing a realistic shared experience for multiple users. To provide a realistic XR experience for multiple users, the XR system must understand the physical environment of the users in order to correctly associate the positions of virtual objects with real objects. The XR system can build an environmental map of the scene, which can be created by leveraging the images and / or depth information collected by sensors that are part of the XR devices worn by the users of the XR system.

[0114] In an XR system, each XR device can develop a local map of its physical environment by integrating information from one or more images collected at some point during a scan. In some embodiments, the coordinate system of the map is associated with the device orientation at the start of the scan. As the user interacts with the XR system, the orientation changes over the session, whether different sessions are associated with different users (each with their own wearable device and sensors for the scanned environment), or the same user uses the same device at different times. The inventors have recognized and understood techniques for operating an XR system based on persistent spatial information, which overcome the limitations of XR systems where each user device relies only on spatial information collected relative to an orientation that varies for different user instances (e.g., time snapshots) or sessions (e.g., the time between power on and off) of the system. For example, these techniques can provide an XR scenario that is more computationally efficient and immersive for a single or multiple users by allowing any of the multiple users of the XR system to create, store, and retrieve persistent spatial information.

[0115] The persistent spatial information can be represented by a persistent map, enabling one or more functions that enhance the XR experience. The persistent map can be stored in a remote storage medium (e.g., the cloud). For example, a wearable device worn by a user can retrieve a previously created and stored appropriate stored map from persistent storage such as cloud storage after being powered on. The previously stored map may have been based on data about the environment collected using sensors on the user's wearable device during a previous session. Retrieving the stored map can allow the wearable device to be used without scanning the physical world through the sensors on the wearable device. Alternatively or additionally, the system / device can similarly retrieve an appropriate stored map when entering a new area of the physical world.

[0116] The stored map can be represented in a canonical form that is related to the local reference frame on each XR device. In a multi-device XR system, a stored map accessed by one device may have been created and stored by another device, and / or may have been constructed by aggregating data about the physical world collected by sensors on multiple wearable devices that previously existed in a portion of the physical world represented by the stored map.

[0117] The relationship between the canonical map and the local map of each device can be determined through a localization process. This localization process can be performed on each XR device based on a set of selected canonical maps sent to the device. However, the inventors have recognized and understood that the network bandwidth and computing resources on XR devices can be reduced by providing a localization service that can be executed on a remote processor (such as, can be implemented in the cloud). Therefore, the battery consumption and heat generation on XR devices may be reduced, enabling the device to allocate resources such as computing time, network bandwidth, battery life, and heat budget to provide a more immersive user experience. However, by appropriately selecting the information transmitted between each XR device and the localization service, the localization can be performed with the latency and accuracy required to support such an immersive experience.

[0118] Sharing data about the physical world between multiple devices can enable a shared user experience of virtual content. For example, two XR devices that can access the same stored map can both be localized relative to the stored map. Once localized, the user device can render virtual content at that location by transforming the position specified by the reference stored map into the reference frame maintained by the user device. The user device can use this local reference frame to control the display of the user device to render virtual content at the specified location.

[0119] To support these and other functions, the XR system can include components for developing, maintaining, and using persistent spatial information (including one or more stored maps) based on data about the physical world collected using sensors on the user device. These components can be distributed in the XR system, where some components, for example, can operate on the head-mounted part of the user device. Other components can operate on a computer associated with the user that is coupled to the head-mounted part via a local area network or a personal area network. Still other components can operate at a remote location (such as, one or more servers accessible via a wide area network).

[0120] For example, these components can include components capable of identifying, from information about the physical world collected by one or more user devices, information of sufficient quality to be stored as a persistent map or stored in a persistent map. An example of such a component, described in more detail below, is a map merging component. For example, such a component can receive input from a user device and determine the applicability of a portion of the input for updating the persistent map. For example, the map merging component can divide a local map created by a user device into multiple parts, determine the mergability of one or more parts with the persistent map, and merge the parts that meet the merge criteria into the persistent map. For example, the map merging component can also promote the parts that are not merged with the persistent map to separate persistent maps.

[0121] As another example, the components can include components that help determine an appropriate persistent map that can be retrieved and used by a user device. An example of such a component, described in more detail below, is a map ranking component. For example, such a component can receive input from a user device and identify one or more persistent maps that may represent the physical world area in which the device is operating. For example, the map ranking component can help select a persistent map to be used by a local device when rendering virtual content, collecting data about the environment, or performing other actions. Alternatively or additionally, the map ranking component can help identify a persistent map to be updated as additional information about the physical world is collected by one or more user devices.

[0122] Additional components can determine a transformation that transforms information captured or described with respect to one reference frame into another reference frame. For example, sensors can be attached to a head-mounted display so that the data read from the sensors indicates the position of an object in the physical world relative to the pose of the wearer's head. One or more transformations can be applied to associate the position information with the coordinate system of a persistent environment map. Similarly, data indicating the position at which a virtual object will be rendered, represented in the coordinate system of a persistent environment map, can be transformed one or more times to be in the reference frame of a display at the user's head. As described in more detail below, there may be multiple such transformations. These transformations can be partitioned across XR system components so that they can be updated and / or applied efficiently to a distributed system.

[0123] In some embodiments, a persistent map can be constructed based on information collected by multiple user devices. An XR device can capture local space information and construct a separate tracking map using the information collected by sensors in each of the XR devices at different locations and times. Each tracking map can include points, each of which can be associated with a feature of a real object that may include multiple features. In addition to potentially providing input for creating and maintaining a persistent map, the tracking map can also be used to track the movement of a user in a scene, enabling the XR system to estimate the head pose of the corresponding user based on the tracking map.

[0124] XR systems can be operated using techniques that provide XR scenes to obtain a highly immersive user experience, such as estimating head pose at a frequency of 1 kHz, where the usage rate of computing resources associated with the XR device is low. The XR device can be configured with, for example, four video graphics array (VGA) cameras operating at 30 Hz, an inertial measurement unit (IMU) operating at 1 kHz, the computing power of a single advanced RISC machine (ARM) core, less than 1 GB of memory, and less than 100 Mbp of network bandwidth. These techniques involve reducing the processing required to generate and maintain maps and estimate head pose, and involve providing and using data with low computational overhead. The XR system can calculate its pose based on matching visual features. U.S. Patent Application No. 16 / 221,065 describes hybrid tracking, the entire content of which is incorporated herein by reference.

[0125] These techniques can include reducing the amount of data processed when building a map, such as by building a sparse map with a set of drawing points and key frames and / or dividing the map into chunks for per-chunk updates. The drawing points can be associated with points of interest in the environment. The key frames include selected information from the data captured by the camera. U.S. Patent Application No. 16 / 520,582 describes determining and / or evaluating a localization map, the entire content of which is incorporated herein by reference.

[0126] In some embodiments, persistent spatial information can be represented in a manner that is easily shared among users and among distributed components including applications. For example, information about the physical world can be represented as a persistent coordinate frame (PCF). The PCF can be defined based on one or more points representing features identified in the physical world. The features can be selected to be the same across XR system user sessions. The PCF can exist sparsely, providing less than all the available information about the physical world, thus effectively processing and transmitting this information. Techniques for processing persistent spatial information can include creating a dynamic map based on one or more coordinate systems in the real space across one or more sessions, and also include generating a persistent coordinate frame (PCF) on the sparse map, which can be exposed to XR applications, for example, via an application programming interface (API). These capabilities are supported by techniques for ranking and merging multiple maps created by one or more XR devices. Persistent spatial information can also quickly restore and reset the head pose on each of one or more XR devices in a computationally efficient manner.

[0127] In addition, these techniques can enable effective comparison of spatial information. In some embodiments, an image frame can be represented by a digital descriptor. The descriptor can be calculated via a transformation that maps a set of identified features in the image to the descriptor. The transformation can be performed in a trained neural network. In some embodiments, the set of features provided as input to the neural network can be a filtered set of features, which are extracted from the image using techniques that preferentially select features that are likely to be persistent.

[0128] Representing an image frame as a descriptor, for example, allows for effective matching of new image information with stored image information. An XR system can combine a persistent map with descriptors of one or more frames under the persistent map. A local image frame acquired by a user device can similarly be transformed into such a descriptor. By selecting a stored map having a descriptor similar to that of the local image frame, one or more persistent maps that may represent the same physical space as the user device can be selected with a relatively small amount of processing. In some embodiments, descriptors of key frames in the local map and the persistent map can be calculated, thus further reducing processing when comparing the maps. For example, such effective comparison can be used to simplify finding a persistent map to be loaded into a local device or for finding a persistent map to be updated based on image information acquired with the local device.

[0129] The techniques described herein can be used together or separately with a variety of types of devices, as well as for a variety of types of scenarios, including wearable or portable devices with limited computing resources that provide augmented or mixed reality scenarios. In some embodiments, these techniques can be implemented by one or more services that form part of an XR system.

[0130] AR System Overview

[0131] Figure 1 and Figure 2 shows a scenario with virtual content displayed in conjunction with a portion of the physical world. For illustrative purposes, an AR system is used as an example of an XR system. Figure 3-6B shows an exemplary AR system, including one or more processors, a memory, sensors, and a user interface that can operate according to the techniques described herein.

[0132] Reference Figure 1, shows an outdoor AR scene 354, where a user of AR technology sees a physical world park-like environment 356 that includes people, trees, buildings in the background, and a concrete platform 358 feature. In addition to these items, the user of AR technology also perceives that they "see" a robotic statue 357 standing on the physical world concrete platform 358 and a flying cartoon avatar character 352, seemingly a humanoid bumblebee, even though these elements (e.g., avatar character 352 and robotic statue 357) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is challenging to produce an AR technology that promotes the comfortable, natural-feeling, and rich presentation of virtual image elements among other virtual or physical world image elements.

[0133] Such an AR scene can be implemented with a system that constructs a map of the physical world based on tracking information, allows a user to place AR content in the physical direct, determines the location in the map of the physical world to place the AR content, maintains the AR scene so that the placed AR content can be reloaded to be displayed in the physical world during, for example, different AR experience sessions, and allows multiple users to share the AR experience. The system can construct and update a digital representation of the surfaces of the physical world around the user. This representation can be used to render virtual content to appear to be occluded in whole or in part by physical objects between the user and the rendering location of the virtual content, for placing virtual objects in physics-based interactions, and for virtual character path planning and navigation, or for other operations that use information about the physical world.

[0134] Figure 2 Shows another example of an indoor AR scene 400 according to some embodiments, where an exemplary XR system use case is shown. The exemplary scene 400 is a living room that has walls, a bookshelf on one side of the wall, a floor lamp in a corner of the room, a floor, a sofa and a coffee table on the floor. In addition to these physical objects, a user of AR technology can also perceive virtual objects, such as an image on the wall behind the sofa, a bird flying through the door, a deer sticking its head out from the bookshelf, an ornament in the form of a windmill placed on the coffee table, etc.

[0135] For an image on a wall, AR technology requires not only information about the wall surface but also information about objects in the room (such as the shape of a lamp) and surfaces that occlude the image to correctly render virtual objects. For a flying bird, AR technology requires information about all objects and surfaces in the corners of the room to render the bird avoiding objects and surfaces realistically or bouncing off them when it collides. For a deer, AR technology requires information about surfaces such as the floor or a coffee table to calculate the placement of the deer. For a windmill, the system can identify it as an object separate from the table and determine that it is movable, while the corners of a bookshelf or a wall corner are determined to be stationary. This distinction can be used to determine which parts of the scene are used or updated in each of various operations.

[0136] Virtual objects can be placed in a previous AR experience session. When a new AR experience session starts in a living room, AR technology requires the virtual objects to be accurately displayed in the previously placed positions and be realistically visible from different angles. For example, the windmill should be shown standing on a book, rather than floating in another position above the table without the book. Such floating can occur if the user's position in the new AR experience session is not accurately located in the living room. As another example, if the user's viewing angle of the windmill is different from the angle at which the windmill was placed, AR technology needs to display the corresponding side of the windmill.

[0137] A scene can be presented to a user via a system including multiple components, which include a user interface that can stimulate one or more of the user's senses (such as vision, sound, and / or touch). Additionally, the system can include one or more sensors that can measure parameters of the physical parts of the scene, including the user's position and / or movement within the physical parts of the scene. Additionally, the system can include one or more computing devices having associated computer hardware (such as memory). These components can be integrated into a single device or can be distributed among multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.

[0138] Figure 3An AR system 502 configured to provide an experience of interacting with the physical world 506 with AR content is shown. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by a user as part of a head-mounted device such that the user can wear the display over their eyes, such as a pair of goggles or glasses. At least a portion of the display may be transparent such that the user can observe the see-through reality 510. The see-through reality 510 may correspond to a portion of the physical world 506 that is within the current viewpoint of the AR system 502. In the case where the user wears a head-mounted device that includes both the display and sensors of the AR system to obtain information about the physical world, the current viewpoint of the AR system 502 corresponds to the user's viewpoint.

[0139] AR content may also be presented on the display 508, overlaying the see-through reality 510. To provide accurate interaction between the AR content and the see-through reality 510 on the display 508, the AR system 502 may include sensors 522 configured to capture information about the physical world 506.

[0140] The sensors 522 may include one or more depth sensors that output depth maps 512. Each depth map 512 may have a plurality of pixels, and each pixel may represent the distance to a surface in the physical world 506 in a particular direction relative to the depth sensor. Raw depth data comes from the depth sensors to create the depth maps. The update rate of such depth maps can be as fast as the depth sensors can form new images, hundreds or thousands of times per second. However, the data may be noisy and incomplete, and there are holes shown as black pixels on the illustrated depth maps.

[0141] The system may include other sensors, such as image sensors. The image sensors may acquire monocular or stereoscopic information that is processed to otherwise represent the physical world. For example, the images may be processed in the world reconstruction component 516 to create a mesh representing connected parts of objects in the physical world. Metadata about these objects (e.g., including color and surface texture) may similarly be acquired using sensors and stored as part of the world reconstruction.

[0142] The system can also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, the head pose of the system can be calculated in real time using the head pose tracking component. The head pose tracking component can represent the user's head pose in a coordinate system with six degrees of freedom, including, for example, translation along three perpendicular axes (e.g., front / back, up / down, left / right) and rotation about three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, sensor 522 can include an inertial measurement unit ("IMU") that can be used to calculate and / or determine the head pose 514. For example, the head pose 514 for the depth map can indicate the current viewing point of the sensor that captured the depth map in six degrees of freedom, but the head pose 514 can also be used for other purposes, such as associating image information with a particular part of the physical world or associating the position of a display worn on the user's head with the physical world.

[0143] In some embodiments, the head pose information can be derived in other ways in addition to from the IMU, such as by analyzing objects in the image. For example, the head pose tracking component can calculate the relative position and orientation of the AR device and the physical object based on the visual information captured by the camera and the inertial information captured by the IMU. The head pose tracking component can then calculate the head pose of the AR device by, for example, comparing the calculated relative position and orientation of the AR device and the physical object with the characteristics of the physical object. In some embodiments, this comparison can be performed by identifying features in the images captured using one or more sensors 522 that remain stable over a period of time such that the change in the position of these features in the images captured over a period of time can be correlated with the change in the user's head pose.

[0144] In some embodiments, when the user moves around in the physical world while wearing the AR device, the AR device can construct a map based on the feature points identified in the consecutive images within a captured sequence of image frames. Although each image frame can be taken from a different pose when the user moves, the system can adjust the orientation of the features of each consecutive image frame by matching the features of the consecutive image frames with the previously captured features to match the orientation of the initial image frame. The translation of the consecutive image frames (such that the points representing the same feature will match the corresponding feature points from the previously collected image frames) can be used to align each consecutive image frame to match the orientation of the previously processed image frame. The frames in the resulting map can have a common orientation established when the first image frame was added to the map. The map has a set of feature points located in a common reference frame that can be used to determine the user's pose in the physical world by matching the features from the current image frame with the map. In some embodiments, this map can be referred to as a tracking map.

[0145] In addition to being able to track the user's pose in the environment, the map can also enable other components of the system, such as the world reconstruction component 516, to determine the position of physical objects relative to the user. The world reconstruction component 516 can receive the depth map 512 and the head pose 514, as well as any other data from the sensors, and integrate this data into the reconstruction 518. The reconstruction 518 can be more complete and less noisy than the sensor data. The world reconstruction component 516 can use the spatial and temporal averaging of sensor data from multiple viewpoints over a period of time to update the reconstruction 518.

[0146] The reconstruction 518 can include a physical world representation in one or more data formats, such as including voxels, meshes, planes, etc. Different formats can represent alternative representations of the same part of the physical world, or can represent different parts of the physical world. In the example shown, on the left side of the reconstruction 518, a part of the physical world is presented as a global surface; on the right side of the reconstruction 518, a part of the physical world is presented as a mesh.

[0147] In some embodiments, the map maintained by the head pose component 514 may be sparse relative to other maps of the physical world that may be maintained. A sparse map can indicate the location of points of interest and / or structures, such as corners or edges, rather than providing information about the location of surfaces and other possible features. In some embodiments, the map can include image frames captured by the sensor 522. These frames can be reduced to features that can represent points of interest and / or structures. In conjunction with each frame, information about the user pose from which the frame was acquired can also be stored as part of the map. In some embodiments, each image acquired by the sensor may or may not be stored. In some embodiments, the system can process the image when it is collected by the sensor and select a subset of the image frames for further computation. This selection can be based on one or more criteria that limit information addition but ensure that the map contains useful information. The system can add new image frames to the map, for example, based on the overlap with previous image frames that have already been added to the map or based on image frames that contain a sufficient number of features that are determined to likely represent stationary objects. In some embodiments, the selected image frames or groups of features from the selected image frames can be used as key frames of the map, which are used to provide spatial information.

[0148] The AR system 502 can integrate sensor data from multiple viewpoints of the physical world over a period of time. When a device including sensors moves, the pose (e.g., position and orientation) of the sensors can be tracked. Since the frame pose of the known sensors and its relationship with other poses are known, each of these multiple viewpoints of the physical world can be fused together to form a single combined reconstruction of the physical world, which can serve as an abstract layer of a map and provide spatial information. By using spatial and temporal averaging (i.e., averaging data from multiple viewpoints over a period of time) or any other suitable method, the reconstruction may be more complete and less noisy than the original sensor data.

[0149] In Figure 3 In the illustrated embodiment, the map represents the portion of the physical world where the user of a single wearable device is located. In this case, the head pose associated with a frame in the map can be represented as a local head pose, indicating the orientation relative to the initial orientation of the single device at the start of a session. For example, when the device is turned on or otherwise operated to scan the environment to build a representation of that environment, the head pose can be tracked relative to the initial head pose.

[0150] In combination with the content characterizing the portion of the physical world, the map can include metadata. For example, the metadata can indicate the capture time of the sensor information used to form the map. Alternatively or additionally, the metadata can indicate the location of the sensors when the information for forming the map was captured. The location can be expressed directly, such as using information from a GPS chip, or indirectly, such as using a wireless (e.g., Wi-Fi) signature (indicating the signal strength received from one or more wireless access points when collecting sensor data), and / or using the identifier of the wireless access point to which the user device was connected when collecting sensor data, such as a BSSID.

[0151] The reconstruction 518 can be used for AR functions, such as generating a surface representation of the physical world for occlusion processing or processing based on physical phenomena. The surface representation can change as the user moves or objects in the physical world change. Some aspects of the reconstruction 518 can be used, for example, by component 320, which generates a changing global surface representation in world coordinates that can be used by other components.

[0152] AR content is generated, for example, by an AR application 504 based on this information. The AR application 504 can be a game program that, for example, based on information about the physical world, performs one or more functions such as visual occlusion, interaction based on physical phenomena, and environmental reasoning. These functions can be performed by querying data in different formats from a reconstruction 518 generated by a world reconstruction component 516. In some embodiments, a component 520 can be configured to output an update when a representation in an area of interest in the physical world changes. For example, the area of interest can be set to a part of the physical world near the user of the system, such as a part within the user's field of view, or projected (predicted / determined) to enter the user's field of view.

[0153] The AR application 504 can use this information to generate and update AR content. The virtual part of the AR content can be presented on a display 508 in combination with the see-through reality 510, creating a realistic user experience.

[0154] In some embodiments, an AR experience can be provided to a user via an XR device, which can be a wearable display device and can be part of a system that can include remote processing and / or remote data storage and / or, in some embodiments, other wearable display devices worn by other users.

[0155] For simplicity of illustration, Figure 4 an example of a system 580 (hereinafter referred to as "system 580") is shown that includes a single wearable device. The system 580 includes a head-mounted display device 562 (hereinafter referred to as "display device 562"), and various mechanical and electronic modules and systems that support the functions of the display device 562. The display device 562 can be coupled to a frame 564 that can be worn by a user or viewer 560 of the display system (hereinafter referred to as "user 560") and is configured to position the display device 562 in front of the eyes of the user 560. According to various embodiments, the display device 562 can be a sequential display. The display device 562 can be a monocular or binocular display. In some embodiments, the display device 562 can be Figure 3 an example of the display 508 in

[0156] In some embodiments, speaker 566 is coupled to frame 564 and is positioned near the ear canal of user 560. In some embodiments, another speaker (not shown) is positioned near the other ear canal of user 560 to provide stereo / plastic sound control. Display device 562 is operatively coupled to local data processing module 570, such as via a wired wire or wireless connection 568, and this module can be installed in various configurations, such as being fixedly attached to frame 564, being fixedly attached to a helmet or hat worn by user 560, being embedded in the earphone, or otherwise being detachably attached to user 560 (e.g., taking a backpack configuration, taking a belt-coupled configuration).

[0157] Local data processing module 570 may include a processor and a digital memory, such as a non-volatile memory (e.g., flash memory), both of which can be used to assist in the processing, caching, and storage of data. The data includes: a) data captured from sensors, which can be operatively coupled to frame 564 or otherwise attached to user 560, such as an image capture device (such as a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a radio device, and / or a gyroscope; and / or b) data obtained and / or processed using remote processing module 572 and / or remote data repository 574, and this data can be transmitted to display device 562 after such processing or retrieval.

[0158] In some embodiments, the wearable device can communicate with remote components. Local data processing module 570 can be operatively coupled to remote processing module 572 and remote data repository 574 via communication links 576, 578, such as via a wired or wireless communication link, respectively, such that these remote modules 572, 574 are operatively coupled to each other and can be used as resources of local data processing module 570. In a further embodiment, as a supplement or alternative to remote data repository 574, the wearable device can access a cloud-based remote data repository and / or service. In some embodiments, the above-mentioned head pose tracking component can be at least partially implemented in local data processing module 570. In some embodiments, Figure 3 the world reconstruction component 516 in can be at least partially implemented in local data processing module 570. For example, local data processing module 570 can be configured to execute computer-executable instructions to generate a map and / or a physical world representation based at least in part on at least a portion of the data.

[0159] In some embodiments, processing may be distributed across local and remote processors. For example, local processing may be used to build a map (e.g., a tracking map) on a user device based on sensor data collected using sensors on the user device. Such a map may be used by an application on the user device. Additionally, previously created maps (e.g., canonical maps) may be stored in a remote data repository 574. Where appropriate stored maps or persistent maps are available, these maps may be used in place of or to supplement the tracking map created locally on the device. In some embodiments, the tracking map may be aligned to the stored map to establish a correspondence between the tracking map (which may be oriented relative to the position of the wearable device when the user powers on the system) and the canonical map (which may be oriented relative to one or more persistent features). In some embodiments, a persistent map may be loaded on the user device to allow the user device to render virtual content without latency associated with the scanned location and to build a tracking map of the user's complete environment based on sensor data acquired during the scan. In some embodiments, the user device may access a remote persistent map (e.g., stored in the cloud) without downloading the persistent map on the user device.

[0160] In some embodiments, spatial information may be transmitted from the wearable device to a remote service, such as a cloud service configured to locate the device to a stored map maintained on the cloud service. According to some embodiments, the location processing may be performed in the cloud matching the device location to an existing map (such as a canonical map) and returning a transformation that links virtual content to the wearable device location. In such an embodiment, the system may avoid transmitting the map from the remote resource to the wearable device. Other embodiments may be configured for both device-based location and cloud-based location, such as to enable functionality where the network connection is unavailable or the user chooses not to enable cloud-based location.

[0161] Alternatively or additionally, the tracking map may be merged with previously stored maps to extend or enhance the quality of those maps. The process of determining whether a suitable previously created environmental map is available and / or merging the tracking map with one or more stored environmental maps may be done in the local data processing module 570 or the remote processing module 572.

[0162] In some embodiments, the local data processing module 570 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which would limit the computational budget of the local data processing module 570 but enable a smaller device. In some embodiments, the world reconstruction component 516 may use less than the computational budget of a single advanced RISC machine (ARM) core to generate a physical world representation in real time on an undefined space, such that the remaining computational budget of a single ARM core can be used for other purposes, such as mesh extraction.

[0163] In some embodiments, the remote data repository 574 may include a digital data storage facility accessible via the Internet or other network configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, allowing for fully autonomous use from a remote module. In some embodiments, all data is stored and all computations are performed in the remote data repository 574, allowing for the use of a smaller device. For example, world reconstruction may be stored in whole or in part in the repository 574.

[0164] In embodiments where data is remotely stored and network-accessible, the data may be shared by multiple users of the augmented reality system. For example, user devices may upload their tracking maps to augment a database of environmental maps. In some embodiments, tracking map upload occurs at the end of a user session with the wearable device. In some embodiments, tracking map upload may occur continuously, semi-continuously, intermittently, at predefined times, after a predefined period from a previous upload, or when triggered by an event. Tracking maps uploaded by any user device may be used to extend or enhance previously stored maps, whether based on data from that user device or any other user device. Similarly, the persistent maps downloaded to a user device may be based on data from that user device or any other user device. In this way, users can easily obtain high-quality environmental maps to improve their experience of the AR system.

[0165] In a further embodiment, the persistent map download can be restricted and / or avoided based on a localization performed on a remote resource (e.g., in the cloud). In such a configuration, a wearable device or other XR device transmits feature information coupled with pose information (e.g., the device's localization information when sensing features represented by the feature information) to a cloud service. One or more components of the cloud service can match the feature information with a corresponding stored map (e.g., a canonical map) and generate a transformation between the coordinate systems of the tracking map and the canonical map maintained by the XR device. Each XR device that localizes its tracking map relative to the canonical map can accurately render virtual content at a position specified relative to the canonical map based on its own tracking.

[0166] In some embodiments, the local data processing module 570 is operatively coupled to a battery 582. In some embodiments, the battery 582 is a removable power source, such as a battery that can be purchased over the counter. In other embodiments, the battery 582 is a lithium-ion battery. In some embodiments, the battery 582 includes a built-in lithium-ion battery that can be charged by the user 560 during non-operating times of the system 580, as well as a removable battery, such that the user 560 does not have to be tethered to a power source to charge the lithium-ion battery and can operate the system 580 for a longer period of time without having to turn off the system 580 to replace the battery.

[0167] Figure 5A A user 530 wearing an AR display system that renders AR content is shown as the user 530 moves in a physical world environment 532 (hereinafter referred to as "environment 532"). Information captured by the AR system along the user's movement path can be processed into one or more tracking maps. The user 530 positions the AR display system at a location 534, and the AR display system records background information of the connected world relative to the location 534 (e.g., a digital representation of real objects in the physical world that can be stored and updated as real objects in the physical world change). This information can be stored as a pose in combination with images, features, directional audio inputs, or other desired data. The location 534 is integrated into the data input 536, for example, as part of a tracking map, and is processed by at least the connected world module 538, which can be implemented, for example, by Figure 4 processing on the remote processing module 572. In some embodiments, the connected world module 538 can include a head pose component 514 and a world reconstruction component 516 such that the processed information can combine information about physical objects used in rendering virtual content to indicate the positions of objects in the physical world.

[0168] As determined based on data input 536, the connected world module 538 at least partially determines where and how the AR content 540 can be placed in the physical world. The AR content is "placed" in the physical world by presenting a representation of the physical world and the AR content via the user interface, and the AR content is rendered as if interacting with the objects in the physical world, while the objects in the physical world are presented as if the AR content occludes the user's view of those objects when appropriate. In some embodiments, the AR content can be placed by appropriately selecting a portion of a fixed element 542 (e.g., a table) from the reconstruction (e.g., reconstruction 518) to determine the shape and position of the AR content 540. As an example, the fixed element can be a table, and the virtual content can be positioned such that it appears to be on the table. In some embodiments, the AR content can be placed within a structure in the field of view 544, which can be the current field of view or an estimated future field of view. In some embodiments, the AR content can persist relative to a model 546 (mesh) of the physical world.

[0169] As shown, the fixed element 542 serves as an indicator (e.g., a digital copy) of any fixed element within the physical world, which can be stored in the connected world module 538 such that the user 530 can perceive the content on the fixed element 542 without the system having to map to the fixed element 542 every time the user 530 views the content. Thus, the fixed element 542 can be a mesh model from a previous modeling session, or determined by a separate user, but still stored by the connected world module 538 for future reference by multiple users. Accordingly, the connected world module 538 can identify the environment 532 from a previously drawn environment and display the AR content without the device of the user 530 having to first draw all or part of the environment 532, thereby saving computational processes and cycles and avoiding latency of any rendered AR content.

[0170] The mesh model 546 of the physical world can be created by the AR display system, and the appropriate surfaces and metrics for interacting with and displaying the AR content 540 can be stored by the connected world module 538 for future retrieval by the user 530 or other users without having to fully or partially reconstruct the model. In some embodiments, the data input 536 is an input such as a geographical location, user identification, and current activity, used to indicate to the connected world module 538 which fixed element 542 among one or more fixed elements is available, which AR content 540 was last placed on the fixed element 542, and whether to display the same content (such AR content is "persistent" content regardless of whether the user views a particular connected world model).

[0171] Even in embodiments where an object is considered stationary (e.g., a kitchen table), the conjunctive world module 538 can update those objects in the physical world model from time to time while taking into account possible changes in the physical world. The model of a stationary object may be updated at a very low frequency. Other objects in the physical world may be moving or not considered stationary (e.g., a kitchen chair). To render an AR scene with a sense of realism, the AR system can update the positions of these non-stationary objects at a much higher frequency than that used to update stationary objects. To be able to accurately track all objects in the physical world, the AR system can obtain information from multiple sensors, including one or more image sensors.

[0172] Figure 5B is a schematic diagram of the viewing optical component 548 and its accessories. In some embodiments, two eye-tracking cameras 550 pointing at the user's eyes 549 detect metrics of the user's eyes 549, such as eye shape, eyelid closure, pupil direction, and bright spots on the user's eyes 549.

[0173] In some embodiments, one of the sensors can be a depth sensor 551, such as a time-of-flight sensor, which emits signals into the world and detects the reflections of these signals from nearby objects to determine the distance to a given object. For example, the depth sensor can quickly determine whether an object has entered the user's field of view due to the movement of those objects or a change in the user's posture. However, information about the position of an object in the user's field of view can alternatively or additionally be collected using other sensors. For example, depth information can be obtained from a stereo vision image sensor or a plenoptic sensor.

[0174] In some embodiments, the world camera 552 records a view larger than the periphery to map and / or otherwise create a model of the environment 532 and detect inputs that can affect the AR content. In some embodiments, the world camera 552 and / or the camera 553 can be grayscale and / or color image sensors, which can output grayscale and / or color image frames at fixed time intervals. The camera 553 can also capture an image of the physical world within the user's field of view at a specific time. The pixels of a frame-based image sensor can be resampled repeatedly even if their values are not changed. Each of the world camera 552, the camera 553, and the depth sensor 551 has its own field of view 554, 555, and 556 to collect data from and record the physical world scene, such as Figure 5A the physical world environment 532 depicted in

[0175] The inertial measurement unit 557 can determine the movement and orientation of the viewing optical assembly 548. In some embodiments, each component is operatively coupled to at least one other component. For example, the depth sensor 551 is operatively coupled to the eye tracking camera 550 as a confirmation for adjusting the measurement of the actual distance at which the user's eye 549 is gazing.

[0176] It should be understood that the viewing optical assembly 548 may include Figure 5B some of the components shown in, and may include components in place of or in addition to the shown components. In some embodiments, for example, the viewing optical component 548 may include two world cameras 552 instead of four. Alternatively or additionally, cameras 552 and 553 do not need to capture visible light images of their full fields of view. The viewing optical assembly 548 may include other types of components. In some embodiments, the viewing optical assembly 548 may include one or more dynamic vision sensors (DVSs), the pixels of which may respond asynchronously to relative changes in light intensity above a threshold.

[0177] In some embodiments, the viewing optical component 548 may not include the depth sensor 551 based on time-of-flight information. In some embodiments, for example, the viewing optical assembly 548 may include one or more plenoptic cameras, the pixels of which may capture light intensity and the angle of incident light, from which depth information can be determined. For example, a plenoptic camera may include an image sensor covered with a transmissive diffraction mask (TDM). Alternatively or additionally, a plenoptic camera may include an image sensor that includes angle-sensitive pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such a sensor may be used as a source of depth information in addition to or in place of the depth sensor 551.

[0178] It should also be understood that Figure 5B the configuration of the components in is provided as an example. The viewing optical assembly 548 may include components having any suitable configuration, which may be arranged to provide the user with the maximum field of view practical for a particular set of components. For example, if the viewing optical assembly 548 has one world camera 552, that world camera may be placed in the central region of the viewing optical assembly rather than on the side.

[0179] Information from these sensors in the viewing optics 548 can be coupled to one or more processors in the system. The processor(s) can generate data that, through rendering, enables a user to perceive interacting with virtual content in the physical world. The rendering can be implemented in any suitable manner, including generating image data depicting physical and virtual objects. In other embodiments, physical and virtual content can be shown in a scene by modulating the opacity of a display device through which the user views the physical world. The opacity can be controlled to create the appearance of virtual objects and also to prevent the user from seeing objects in the physical world that are occluded by the virtual objects. In some embodiments, the image data can include only virtual content that can be modified such that when viewed through a user interface, the virtual content is perceived by the user as interacting realistically with the physical world (e.g., clipped content that addresses occlusion issues).

[0180] The location on the viewing optics 548 at which content is displayed to create an impression of an object at a particular location can depend on the physical characteristics of the viewing optics. Additionally, the pose of the user's head relative to the physical world and the direction of the user's eye gaze can affect the location in the physical world at which the content displayed at a particular location on the viewing optics appears. The sensors described above can collect this information, and / or provide information from which this information can be calculated, such that a processor receiving sensor input can calculate where an object should be rendered on the viewing optics 548 to create the desired appearance for the user.

[0181] Regardless of how the content is presented to the user, a physical world model can be used so that the characteristics of virtual objects that can be affected by physical objects can be correctly calculated, including the shape, location, motion, and visibility of the virtual objects. In some embodiments, the model can include a reconstruction of the physical world, such as reconstruction 518.

[0182] The model can be created based on data collected from sensors on the user's wearable device. However, in some embodiments, the model can be created based on data collected from multiple users, which can be aggregated in a computing device remote from all users (and can be "in the cloud").

[0183] The model can be created at least in part by a world reconstruction system, such as Figure 3 world reconstruction component 516, which is in Fig. 6Ais depicted in more detail. The world reconstruction component 516 can include a perception module 660 that can generate, update, and store a representation of a portion of the physical world. In some embodiments, the perception module 660 can represent a portion of the physical world within the sensor reconstruction range as a plurality of voxels. Each voxel can correspond to a 3D cube of a predetermined volume in the physical world and include surface information indicating whether a surface exists within the volume represented by the voxel. The voxels can be assigned a value indicating whether the volume corresponding to them has been determined to include a surface of a physical object, has been determined to be empty, or has not yet been measured by the sensor, in which case the value is unknown. It should be understood that it is not necessary to explicitly store the values of the voxels determined to be empty or unknown, as the values of these voxels can be stored in computer memory in any suitable manner, including not storing any information for the voxels determined to be empty or unknown.

[0184] In addition to generating information for the persistent world representation, the perception module 660 can identify and output indications of changes in the area around the AR system user. These indications of changes can trigger an update to the volume data stored as part of the persistent world or trigger other functions, such as triggering the component 304 that generates AR content to update the AR content.

[0185] In some embodiments, the perception module 660 can identify changes based on a signed distance function (SDF) model. The perception module 660 can be configured to receive sensor data, such as a depth map 660a and a head pose 660b, and then fuse the sensor data into the SDF model 660c. The depth map 660a can directly provide SDF information, and the image can be processed to obtain SDF information. The SDF information represents the distance from the sensor used to capture the information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and thus from the user's perspective. The head pose 660b can correlate the SDF information with the voxels in the physical world.

[0186] In some embodiments, the perception module 660 can generate, update, and store a representation of a portion of the physical world within the perception range. The perception range can be determined at least in part based on the sensor reconstruction range, which can be determined at least in part based on the limits of the sensor's observation range. As a specific example, an active depth sensor operating with active IR pulses can operate reliably within a certain distance range, creating an observation range for the sensor that can range from a few centimeters or tens of centimeters to several meters.

[0187] The world reconstruction component 516 may include additional modules that can interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data acquired by the perception module 660. The persistent world module 662 may also include various representation formats of the physical world. For example, volumetric metadata 662b such as voxels, as well as meshes 662c and planes 662d, may be stored. In some embodiments, other information such as depth maps may be saved.

[0188] In some embodiments, a representation of the physical world such as Fig. 6A shown may provide relatively dense information about the physical world for comparison with a sparse map (e.g., the feature point-based tracking map described above).

[0189] In some embodiments, the perception module 660 may include modules that generate a representation of the physical world in various formats (e.g., including mesh 660d, plane, and semantic 660e). The representation of the physical world may be stored on local and remote storage media. Depending on, for example, the location of the storage medium, the representation of the physical world may be described in different coordinate systems. For example, the representation of the physical world stored in the device may be described in a coordinate system local to the device. The representation of the physical world may have a counterpart stored in the cloud. The counterpart in the cloud may be described in a coordinate system shared by all devices in the XR system.

[0190] In some embodiments, these modules may generate a representation based on data within the perception range of one or more sensors at the time of generating the representation, as well as data captured at a previous time and information in the persistent world module 662. In some embodiments, these components may operate on depth information captured using a depth sensor. However, the AR system may include visual sensors and may generate such a representation by analyzing monocular or binocular visual information.

[0191] In some embodiments, these modules may operate on regions of the physical world. When the perception module 660 detects a change in the physical world in a sub-region, those modules may be triggered to update the sub-region of the physical world. For example, such a change may be detected by detecting a new surface in the SDF model 660c or other criteria (such as changing the values of a sufficient number of voxels representing the sub-region).

[0192] The world reconstruction component 516 may include a component 664 that may receive a representation of the physical world from the perception module 660. Information about the physical world may be pulled by these components based on, for example, usage requests from applications. In some embodiments, the information may be pushed to the usage components, such as via an indication of a change in a pre-identified area or a change in the representation of the physical world within the perception range. The component 664 may include, for example, game programs and other components that perform processing for visual occlusion, physics-based interactions, and environmental reasoning.

[0193] In response to a query from the component 664, the perception module 660 may send a representation of the physical world in one or more formats. For example, when the component 664 indicates that the usage is for visual occlusion or physics-based interaction, the perception module 660 may send a surface representation. When the component 664 indicates that the usage is for environmental reasoning, the perception module 660 may send a mesh, plane, and semantics of the physical world.

[0194] In some embodiments, the perception module 660 may include a component that sets the format of the information to provide to the component 664. An example of such a component may be the ray casting component 660f. A usage component (e.g., the component 664) may query, for example, information about the physical world from a particular perspective. The ray casting component 660f may select from one or more representations of the physical world data within the field of view from that viewpoint.

[0195] From the foregoing description, it should be understood that the perception module 660 or another component of the AR system may process data to create a 3D representation of a portion of the physical world. This may be done by at least partially culling a portion of the 3D reconstruction volume based on the camera frustum and / or depth image, extracting and persisting plane data, capturing, persisting, and updating 3D reconstruction data in blocks that allow local updates while maintaining neighbor consistency, providing occlusion data (where the occlusion data is derived from a combination of one or more depth data sources) to applications that generate such scenes, and / or performing multi-stage mesh simplification to simplify the data to be processed. The reconstruction may include data of different levels of complexity, such as including raw data such as real-time depth data, fused volume data such as voxels, and computational data such as meshes.

[0196] In some embodiments, the components of the connected world model can be distributed, with some parts executed locally on the XR device and some parts executed remotely, such as on a network-connected server or otherwise in the cloud. The distribution of information processing and storage between the local XR device and the cloud affects the functionality and user experience of the XR system. For example, reducing the processing on the local device by offloading processing to the cloud can extend battery life and reduce the heat generated on the local device. However, offloading too much processing to the cloud can introduce undesirable latency, resulting in an unacceptable user experience.

[0197] Figure 6B FIG. 600 shows a distributed component architecture configured for spatial computing according to some embodiments. The distributed component architecture 600 can include a connected world component 602 (e.g., Figure 5A PW 538 in ), Lumin OS 604, API 606, SDK 608, and application 610. Lumin OS 604 can include a Linux-based kernel with custom drivers compatible with the XR device. API 606 can include an application programming interface that allows XR applications (e.g., application 610) to access the spatial computing features of the XR device. SDK 608 can include a software development kit that allows for the creation of XR applications.

[0198] One or more components in architecture 600 can create and maintain a model of the connected world. In this example, sensor data is collected on the local device. The processing of this sensor data can be executed partially locally on the XR device and partially in the cloud. PW 538 can include an environmental map created at least in part based on data captured by AR devices worn by multiple users. During a session of an AR experience, a single AR device (such as the wearable device described above in connection with Figure 4 can create a tracking map, which is a type of map.

[0199] In some embodiments, the device can include components for building sparse and dense maps. The tracking map can be used as a sparse map and can include the head pose of the AR device scanning the environment and information about objects detected in the environment at each head pose. These head poses can be maintained locally for each device. For example, the head pose on each device is related to the initial head pose when the device is turned on for a session. Thus, each tracking map is local to the device that created it. The dense map can include surface information representable by a mesh or depth information. Alternatively or additionally, the dense map can include higher-level information derived from the surface or depth information, such as the location and / or properties of planes and / or other objects.

[0200] In some embodiments, the creation of the dense map can be independent of the creation of the sparse map. For example, the creation of the dense map and the sparse map can be performed in separate processing pipelines within the AR system. For example, the separate processing can enable the generation or processing of different types of maps to be performed at different rates. For example, the sparse map may be refreshed at a faster rate than the dense map. However, in some embodiments, the processing of the dense map and the sparse map may be related, even when performed in different pipelines. For example, changes in the physical world shown in the sparse map may trigger an update of the dense map, and vice versa. Additionally, even if the maps are created independently, the maps can still be used together. For example, the coordinate system derived from the sparse map can be used to define the position and / or orientation of objects in the dense map.

[0201] The sparse map and / or the dense map can be persisted for reuse by the same device and / or shared with other devices. Such persistence can be achieved by storing the information in the cloud. The AR device can send the tracking map to the cloud to, for example, merge with an environmental map selected from the persistent maps previously stored in the cloud. In some embodiments, the selected persistent map can be sent from the cloud to the AR device for merging. In some embodiments, the persistent map can be oriented relative to one or more persistent coordinate systems. Such maps can be used as canonical maps because they can be used by any of the multiple devices. In some embodiments, the model of the connected world can include or be created based on one or more canonical maps. Even if the devices perform some operations based on their local coordinate systems, they can still use the canonical maps by determining the transformation between their device local coordinate systems and the coordinate systems of the canonical maps.

[0202] The canonical map starts with a tracking map (TM) (e.g., Fig.31A TM 1102 in ) which can be promoted to a canonical map. The canonical map can be persisted so that a device accessing the canonical map can use the information in the canonical map to determine the position of the objects represented in the canonical map in the physical world around the device once the transformation between its local coordinate system and the coordinate system of the canonical map is determined. In some embodiments, the TM can be a head pose sparse map created by the XR device. In some embodiments, the canonical map can be created when the XR device sends one or more TMs to the cloud server to merge with additional TMs captured by the XR device at different times or captured by other XR devices.

[0203] The canonical map or other maps can provide information about the part of the physical world represented by the data processed to create the corresponding map. Figure 7An exemplary tracking map 700 according to some embodiments is shown. The tracking map 700 may provide a floor plan 706 of a physical object in the corresponding physical world represented by points 702. In some embodiments, the map points 702 may represent features of a physical object that may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. These features may be derived from processing images, such as may be obtained using sensors of a wearable device in an augmented reality system. For example, features may be derived by processing image frames output by the sensors to identify features based on large gradients or other suitable criteria in the images. Further processing may limit the number of features in each frame. For example, the processing may select features that may represent persistent objects. One or more heuristics may be used to make this selection.

[0204] The tracking map 700 may include data about the points 702 collected by the device. For each image frame having a data point included in the tracking map, a pose may be stored. The pose may represent the orientation at which the image frame was captured such that the feature points within each image frame may be spatially related. The pose may be determined by positioning information, such as may be derived from sensors (such as IMU sensors) on the wearable device. Alternatively or additionally, the pose may be determined by matching the image frame with other image frames depicting overlapping portions of the physical world. By finding such position correlations, which may be achieved by matching a subset of the feature points in the two frames, the relative pose between the two frames may be calculated. The relative pose is sufficient for the tracking map since the map may be relative to a device-local coordinate system established based on the initial pose of the device when starting to build the tracking map.

[0205] Not all of the feature points and image frames collected by the device may be retained as part of the tracking map since much of the information collected using the sensors may be redundant. Instead, only certain frames may be added to the map. These frames may be selected based on one or more criteria, such as including the degree of overlap with image frames already present in the map, the number of new features they contain, or a quality metric of the features in the frame. Image frames not added to the tracking map may be discarded or may be used to modify the positions of the features. As a further alternative, all or most of the image frames represented as a feature set may be retained, but a subset of these frames may be designated as key frames for further processing.

[0206] The key frames may be processed to produce a keyrig 704. The key frames may be processed to produce a set of three-dimensional feature points and saved as a keyrig 704. Such processing may require, for example, comparing image frames simultaneously derived from two cameras to stereoscopically determine the 3D positions of the feature points. Metadata may be associated with these key frames and / or the keyrig, such as the pose.

[0207] The environmental map can be in any of a variety of formats, e.g., depending on where the environmental map is stored, such as including the local memory and remote memory of the AR device. For example, the resolution of the map in the remote memory may be higher than that in the local memory on a wearable device with limited memory. To send a higher-resolution map from the remote memory to the local memory, the map can be downsampled or otherwise converted to an appropriate format, such as by reducing the number of poses per unit area of the physical world stored in the map and / or the number of feature points stored for each pose. In some embodiments, slices or portions of the high-resolution map from the remote memory can be sent to the local memory where the slices or portions are not downsampled.

[0208] The database of environmental maps can be updated when a new tracking map is created. To determine which one of the potentially very large number of environmental maps in the database to be updated, the update can include efficiently selecting one or more environmental maps stored in the database associated with the new tracking map. The selected one or more environmental maps can be ranked by relevance, and one or more of the highest-ranked maps can be selected for processing to merge the higher-ranked selected environmental maps with the new tracking map to create one or more updated environmental maps. When the new tracking map represents a portion of the physical world for which there is no pre-existing environmental map to update, the tracking map can be stored in the database as a new environmental map.

[0209] Viewing a separate display

[0210] Methods and apparatuses are described herein for providing virtual content independent of the position of the eyes viewing the virtual content using an XR system. Conventionally, virtual content is re-rendered whenever the display system makes any movement. For example, if a user wearing a display system views a virtual representation of a three-dimensional (3D) object on a display and walks around the area where the 3D object appears, the 3D object should be re-rendered for each viewpoint so that the user feels as if he or she is walking around an object that occupies real space. However, re-rendering consumes a large amount of computational resources of the system and causes artifacts due to latency.

[0211] The inventors have recognized and understood that head pose (e.g., the position and orientation of a user wearing an XR system) can be used to render virtual content independently of eye rotation within the user's head. In some embodiments, a dynamic map of a scene can be generated based on multiple coordinate systems in the real space spanning one or more sessions, such that virtual content interacting with the dynamic map can be robustly rendered independently of eye rotation within the user's head and / or independently of sensor deformation caused by, for example, heat generated during high-speed computationally intensive operations. In some embodiments, configuring multiple coordinate systems can enable a first XR device worn by a first user and a second XR device worn by a second user to identify a common position in the scene. In some embodiments, configuring multiple coordinate systems can enable a user wearing an XR device to view virtual content at the same position in the scene.

[0212] In some embodiments, a tracking map can be constructed in a world coordinate system having a world origin. The world origin can be the first pose when the XR device is powered on. The world origin can be aligned with gravity so that developers of XR applications can obtain gravity alignment without additional work. Different tracking maps can be established in different world coordinate systems because the tracking maps may be captured by the same XR device in different sessions and / or by different XR devices worn by different users. In some embodiments, a session of an XR device can extend from when the device is powered on until it is powered off. In some embodiments, an XR device can have a head coordinate system including a head origin. The head origin may be the pose of the XR device when an image is captured. The difference between the head pose in the world coordinate system and the head pose in the head coordinate system can be used to estimate a tracking route.

[0213] In some embodiments, an XR device can have a camera coordinate system including a camera origin. The camera origin can be the current pose of one or more sensors of the XR device. The inventors have recognized and understood that the configuration of the camera coordinate system enables robust display of virtual content independently of eye rotation within the user's head. This configuration also enables robust display of virtual content independently of sensor deformation caused by, for example, heat generated during operation.

[0214] In some embodiments, an XR device can have a head unit with a head-mounted frame (which a user can fasten to their head) and can include two waveguides, one in front of each of the user's eyes. The waveguides can be transparent such that background light from real-world objects can pass through the waveguides, and the user can see the real-world objects. Each waveguide can transmit projection light from a projector to the user's corresponding eye. The projection light can form an image on the retina of the eye. Thus, the retina of the eye receives the background light and the projection light. The user can simultaneously see real-world objects and one or more virtual objects created by the projection light. In some embodiments, the XR device can have sensors that detect real-world objects around the user. For example, these sensors can be cameras that capture images that can be processed to identify the locations of real-world objects.

[0215] In some embodiments, an XR system can assign a coordinate system to virtual content, as opposed to attaching the virtual content in a world coordinate system. Such a configuration enables the description of virtual content without regard to where the virtual content is rendered to the user, but the virtual content can be attached to a more persistent coordinate system location, such as with respect to, for example, Figure 14-20C the persistent coordinate frame (PCF) described, so as to be rendered at a specified location. When the position of an object changes, the XR device can detect changes in the environmental map and determine the movement of the head unit worn by the user relative to real-world objects.

[0216] Figure 8 Shown is a user experiencing virtual content rendered by an XR system 10 in a physical environment. The XR system can include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is in a physical environment with a real object in the form of a table 16.

[0217] In the example shown, the first XR device 12.1 includes a head unit 22, a hip pack 24, and a cable connection 26. The first user 14.1 fastens the head unit 22 to their head and fastens the hip pack 24 around the waist away from the head unit 22. The cable connection 26 connects the head unit 22 to the hip pack 24. The head unit 22 includes technology for displaying one or more virtual objects to the first user 14.1 while allowing the first user 14.1 to see real objects such as the table 16. The hip pack 24 primarily includes the processing and communication functions of the first XR device 12.1. In some embodiments, the processing and communication functions can reside fully or partially in the head unit 22, such that the hip pack 24 can be removed or the hip pack 24 can be located in another device such as a backpack.

[0218] In the example shown, the waist pack 24 is connected to the network 18 via a wireless connection. The server 20 is connected to the network 18 and stores data representing local content. The waist pack 24 downloads the data representing local content from the server 20 via the network 18. The waist pack 24 provides the data to the head unit 22 via a cable connection 26. The head unit 22 may include a display having a light source, such as a laser light source or a light emitting diode (LED), and a waveguide to guide the light.

[0219] In some embodiments, the first user 14.1 may mount the head unit 22 to their head and the waist pack 24 to their waist. The waist pack 24 may download image data representing virtual content from the server 20 via the network 18. The first user 14.1 may see the table 16 through the display of the head unit 22. A projector forming part of the head unit 22 may receive the image data from the waist pack 24 and generate light based on the image data. The light may travel through one or more of the waveguides forming part of the display of the head unit 22. The light may then leave the waveguide and propagate onto the retina of the eye of the first user 14.1. The projector generates light in a pattern that is replicated on the retina of the eye of the first user 14.1. The light that falls on the retina of the eye of the first user 14.1 may have a selected depth field so that the first user 14.1 perceives an image at a preselected depth behind the waveguide. In addition, the two eyes of the first user 14.1 may receive slightly different images so that the brain of the first user 14.1 perceives one or more three-dimensional images at a selected distance from the head unit 22. In the example shown, the first user 14.1 perceives virtual content 28 above the desktop 16. The scale of the virtual content 28 and its position and distance from the first user 14.1 are determined by the data representing the virtual content 28 and the various coordinate systems used to display the virtual content 28 to the first user 14.1.

[0220] In the example shown, the virtual content 28 is not visible from the perspective of the drawing and is visible to the first user 14.1 using the first XR device 12.1. The virtual content 28 may initially exist as a data structure within the visual data and algorithms in the belt pack 24. The data structures may then manifest themselves as light as the projector of the head unit 22 generates light based on the data structures. It should be understood that although the virtual content 28 does not exist in the three-dimensional space in front of the first user 14.1, the virtual content 28 is still visible in the three-dimensional space. Figure 1 24 to illustrate what is perceived by a wearer of head unit 22. Visualization of computer data in three-dimensional space may be used in this description to illustrate how data structures that facilitate rendering are perceived by one or more users as being related to each other within data structures in waist pack 24.

[0221] Fig. 9Shows components of a first XR device 12.1 according to some embodiments. The first XR device 12.1 may include a head unit 22, and various components that form part of the visual data and algorithms, such as including a rendering engine 30, various coordinate systems 32, various origin and destination coordinate systems 34, and various origin-to-destination coordinate system transformers 36. The various coordinate systems may be based on intrinsic factors of the XR device, or may be determined by referring to other information (such as, a persistent pose or a persistent coordinate system), as described herein.

[0222] The head unit 22 may include a head-mounted frame 40, a display system 42, a real object detection camera 44, a motion tracking camera 46, and an inertial measurement unit 48.

[0223] The head-mounted frame 40 may have a shape that can be fixed to Figure 8 the head of a first user 14.1. The display system 42, the real object detection camera 44, the motion tracking camera 46, and the inertial measurement unit 48 may be mounted to the head-mounted frame 40 and thus move with the head-mounted frame 40.

[0224] The coordinate systems 32 may include a local data system 52, a world coordinate system 54, a head coordinate system 56, and a camera coordinate system 58.

[0225] The local data system 52 may include a data channel 62, a local coordinate system determination routine 64, and local coordinate system storage instructions 66. The data channel 62 may be an internal software routine, a hardware component such as an external cable or a radio frequency receiver, or a hybrid component such as an open port. The data channel 62 may be configured to receive image data 68 representing virtual content.

[0226] The local coordinate system determination routine 64 may be connected to the data channel 62. The local coordinate system determination routine 64 may be configured to determine a local coordinate system 70. In some embodiments, the local coordinate system determination routine may determine the local coordinate system based on a real-world object or a real-world location. In some embodiments, the local coordinate system may be based on the top edge relative to the bottom edge of a browser window, the head or feet of a character, a node on the outer surface of a prism, or a bounding box enclosing virtual content, or any other suitable location for placing a coordinate system that defines the orientation of the virtual content and the location for placing the virtual content (e.g., a node such as a placement node or a PCF node).

[0227] The local coordinate system storage instruction 66 can be connected to the local coordinate system determination routine 64. Those skilled in the art will understand that software modules and routines are "connected" to each other through subroutines, calls, etc. The local coordinate system storage instruction 66 can store the local coordinate system 70 as the local coordinate system 72 in the origin and destination coordinate system 34. In some embodiments, the origin and destination coordinate system 34 can be one or more coordinate systems that can be manipulated or transformed so that virtual content persists between sessions. In some embodiments, a session can be the time period between the startup and shutdown of an XR device. Two sessions can be two startup and shutdown cycles of a single XR device, or can be one startup and shutdown of two different XR devices.

[0228] In some embodiments, the origin and destination coordinate system 34 can be a coordinate system involving one or more transformations required for the XR device of a first user and the XR device of a second user to identify a common location. In some embodiments, the destination coordinate system can be the output of a series of calculations and transformations applied to a target coordinate system so that the first user and the second user view virtual content at the same location.

[0229] The rendering engine 30 can be connected to the data channel 62. The rendering engine 30 can receive the image data 68 from the data channel 62, such that the rendering engine 30 can render virtual content at least partially based on the image data 68.

[0230] The display system 42 can be connected to the rendering engine 30. The display system 42 can include components that transform the image data 68 into visible light. The visible light can form two patterns, one for each eye. The visible light can enter Figure 8 the eyes of the first user 14.1 and can be detected on the retina of the eyes of the first user 14.1.

[0231] The real object detection camera 44 can include one or more cameras that can capture images from different sides of the head-mounted frame 40. The motion tracking camera 46 can include one or more cameras that capture image frames on the sides of the head-mounted frame 40. A single set of one or more cameras can be used instead of two sets of one or more cameras representing the one or more real object detection cameras 44 and the one or more motion tracking cameras 46. In some embodiments, the cameras 44, 46 can capture images. As described above, these cameras can collect data for constructing a tracking map.

[0232] The inertial measurement unit 48 can include multiple devices for detecting the movement of the head unit 22. The inertial measurement unit 48 can include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of the inertial measurement unit 48 collectively track the movement of the head unit 22 in at least three orthogonal directions and around at least three orthogonal axes.

[0233] In the example shown, the world coordinate system 54 includes a world surface determination routine 78, a world coordinate system determination routine 80, and world coordinate system storage instructions 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 receives images and / or key frames based on the images captured by the real object detection camera 44 and processes the images to identify the surfaces in the images. A depth sensor (not shown) may determine the distance to the surface. Thus, the surface is represented by three-dimensional data that includes their size, shape, and distance from the real object detection camera.

[0234] In some embodiments, the world coordinate system 84 may be based on the origin at the initialization of the head pose session. In some embodiments, the world coordinate system may be located at the position where the device is started, or if the head pose is lost during the startup session, it may be located at a new position. In some embodiments, the world coordinate system may be the origin at the start of the head pose session.

[0235] In the example shown, the world coordinate system determination routine 80 is connected to the world surface determination routine 78 and determines the world coordinate system 84 based on the position of the surface determined by the world surface determination routine 78. The world coordinate system storage instructions 82 are connected to the world coordinate system determination routine 80 to receive the world coordinate system 84 from the world coordinate system determination routine 80. The world coordinate system storage instructions 82 store the world coordinate system 84 as the world coordinate system 86 within the origin and the destination coordinate system 34.

[0236] The head coordinate system 56 may include a head coordinate system determination routine 90 and head coordinate system storage instructions 92. The head coordinate system determination routine 90 may be connected to the motion tracking camera 46 and the inertial measurement unit 48. The head coordinate system determination routine 90 may use the data from the motion tracking camera 46 and the inertial measurement unit 48 to calculate the head coordinate system 94. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The motion tracking camera 46 may continuously capture images that the head coordinate system determination routine 90 uses to refine the head coordinate system 94. When Figure 8 the first user 14.1 in moves their head, the head unit 22 also moves. The motion tracking camera 46 and the inertial measurement unit 48 may continuously provide data to the head coordinate system determination routine 90 such that the head coordinate system determination routine 90 can update the head coordinate system 94.

[0237] The head coordinate system storage instruction 92 can be connected to the head coordinate system determination routine 90 to receive the head coordinate system 94 from the head coordinate system determination routine 90. The head coordinate system storage instruction 92 can store the head coordinate system 94 as the head coordinate system 96 in the origin and destination coordinate system 34. When the head coordinate system determination routine 90 recalculates the head coordinate system 94, the head coordinate system storage instruction 92 can repeatedly store the updated head coordinate system 94 as the head coordinate system 96. In some embodiments, the head coordinate system can be the position of the wearable XR device 12.1 relative to the local coordinate system 72.

[0238] The camera coordinate system 58 can include camera intrinsics 98. The camera intrinsics 98 can include the dimensions of the head unit 22 as its design and manufacturing features. The camera intrinsics 98 can be used to calculate the camera coordinate system 100 stored within the origin and destination coordinate system 34.

[0239] In some embodiments, the camera coordinate system 100 can include Figure 8 all pupil positions of the left eye of the first user 14.1. When the left eye moves from left to right or from top to bottom, the pupil position of the left eye is positioned within the camera coordinate system 100. Additionally, the pupil position of the right eye is positioned within the camera coordinate system 100 of the right eye. In some embodiments, the camera coordinate system 100 can include the position of the camera relative to the local coordinate system when taking an image.

[0240] The origin to destination coordinate system converter 36 can include a local to world coordinate converter 104, a world to head coordinate converter 106, and a head to camera coordinate converter 108. The local to world coordinate converter 104 can receive the local coordinate system 72 and transform the local coordinate system 72 to the world coordinate system 86. The transformation of the local coordinate system 72 to the world coordinate system 86 can be represented as the local coordinate system 110 transformed to the world coordinate system within the world coordinate system 86.

[0241] The world to head coordinate converter 106 can transform from the world coordinate system 86 to the head coordinate system 96. The world to head coordinate converter 106 can transform the local coordinate system 110 transformed to the world coordinate system to the head coordinate system 96. This transformation can be represented as the local coordinate system 112 transformed to the head coordinate system within the head coordinate system 96.

[0242] The head-to-camera coordinate transformer 108 can transform from the head coordinate system 96 to the camera coordinate system 100. The head-to-camera coordinate transformer 108 can transform the local coordinate system 112 transformed to the head coordinate system to the local coordinate system 114 transformed to the camera coordinate system within the camera coordinate system 100. The local coordinate system 114 transformed to the camera coordinate system can be input into the rendering engine 30. The rendering engine 30 can render the image data 68 representing the local content 28 based on the local coordinate system 114 transformed to the camera coordinate system.

[0243] Fig.10 is a spatial representation of various origin and destination coordinate systems 34. The local coordinate system 72, the world coordinate system 86, the head coordinate system 96, and the camera coordinate system 100 are shown in the figure. In some embodiments, when virtual content is placed in the real world so that the user can view the virtual content, the local coordinate system associated with the XR content 28 can have a position and rotation relative to the local and / or world coordinate systems and / or the PCF (e.g., a node and orientation can be provided). Each camera can have its own camera coordinate system 100, which contains all the pupil positions of one eye. Reference numerals 104A and 106A respectively represent the transformations performed by Fig. 9 the local-to-world coordinate transformer 104, the world-to-head coordinate transformer 106, and the head-to-camera coordinate transformer 108 in

[0244] Fig.11 illustrates a camera rendering protocol for transforming from a head coordinate system to a camera coordinate system according to some embodiments. In the example shown, the pupil of a single eye moves from position A to B. A virtual object intended to be presented as stationary will be projected onto the depth plane at one of the two positions A or B, depending on the position of the pupil (assuming the camera is configured to use a pupil-based coordinate system). Therefore, using the pupil coordinate system transformed to the head coordinate system will cause jitter in the stationary virtual object when the eye moves from position A to position B. This situation is called viewing-related display or projection.

[0245] As Fig.12 shown, the camera coordinate system (e.g., CR) is positioned and contains all pupil positions, and regardless of the pupil positions A and B, the object projection is now consistent. The head coordinate system is transformed to the CR coordinate system, which is called viewing-independent display or projection. Image reprojection can be applied to the virtual content to cope with changes in the eye position. However, when the rendering remains at the same position, the jitter is minimized.

[0246] Fig.13 The display system 42 is shown in more detail. The display system 42 includes a stereoscopic analyzer 144 that is connected to the rendering engine 30 and forms part of the visual data and algorithms.

[0247] The display system 42 also includes a left projector 166A and a right projector 166B, as well as a left waveguide 170A and a right waveguide 170B. The left projector 166A and the right projector 166B are connected to a power source. Each projector 166A and 166B has a corresponding input so that image data is provided to the corresponding projector 166A or 166B. The corresponding projector 166A or 166B generates and emits light of a two-dimensional pattern when powered. The left waveguide 170A and the right waveguide 170B are positioned to receive light from the left projector 166A and the right projector 166B, respectively. The left waveguide 170A and the right waveguide 170B are transparent waveguides.

[0248] In use, the user mounts the head-mounted frame 40 on their head. Components of the head-mounted frame 40 can include, for example, a strap (not shown) that wraps around the back of the user's head. The left waveguide 170A and the right waveguide 170B are then positioned in front of the user's left eye 220A and right eye 220B.

[0249] The rendering engine 30 inputs the image data it receives into the stereo analyzer 144. The image data is Figure 8 three-dimensional image data of the local content 28 in. The image data is projected onto a plurality of virtual planes. The stereo analyzer 144 analyzes the image data to determine left and right image data sets based on the image data projected onto each depth plane. The left and right image data sets are data sets representing two-dimensional images that are projected in three dimensions to give the user a sense of depth.

[0250] The stereo analyzer 144 inputs the left and right image data sets into the left projector 166A and the right projector 166B. The left projector 166A and the right projector 166B then create left and right light patterns. The components of the display system 42 are shown in a plan view, but it should be understood that when shown in a front view, the left and right patterns are two-dimensional patterns. Each light pattern includes a plurality of pixels. For illustrative purposes, light rays 224A and 226A from two pixels are shown leaving the left projector 166A and entering the left waveguide 170A. The light rays 224A and 226A are reflected from the side of the left waveguide 170A. The light rays 224A and 226A are shown propagating within the left waveguide 170A by total internal reflection from left to right, although it should be understood that the light rays 224A and 226A also propagate in one direction into the paper using a refraction and reflection system.

[0251] Light rays 224A and 226A exit the left light waveguide 170A through the pupil 228A and then enter the left eye 220A through the pupil 230A of the left eye 220A. The light rays 224A and 226A then fall on the retina 232A of the left eye 220A. In this way, the left light pattern falls on the retina 232A of the left eye 220A. The user perceives the pixels formed on the retina 232A as pixels 234A and 236A, and the user feels that pixels 234A and 236A are located at a certain distance on the side of the left light waveguide 170A opposite to the left eye 220A. The sense of depth is created by manipulating the focal length of the light.

[0252] In a similar manner, the stereo analyzer 144 inputs the right image dataset into the right projector 166B. The right projector 166B emits a right light pattern, which is represented by pixels in the form of light rays 224B and 226B. The light rays 224B and 226B are reflected within the right light waveguide 170B and exit through the pupil 228B. The light rays 224B and 226B then enter through the pupil 230B of the right eye 220B and fall on the retina 232B of the right eye 220B. The pixels of the light rays 224B and 226B are perceived as pixels 134B and 236B behind the right light waveguide 170B.

[0253] The patterns created on the retinas 232A and 232B are perceived as the left image and the right image, respectively. Due to the function of the stereo analyzer 144, the left and right images are slightly different from each other. The left and right images are perceived as a three-dimensional rendering in the user's mind.

[0254] As described above, the left light waveguide 170A and the right light waveguide 170B are transparent. Light from real-life objects such as a table 16 on the side surfaces of the left light waveguide 170A and the right light waveguide 170B opposite to the eyes 220A and 220B can be projected through the left light waveguide 170A and the right light waveguide 170B and fall on the retinas 232A and 232B.

[0255] Persistent Coordinate Frame (PCF)

[0256] This document describes methods and apparatuses for providing spatial persistence across user instances within a shared space. Without spatial persistence, virtual content placed by a user in the physical world during a session may not exist or may be misplaced in the user's view during different sessions. Without spatial persistence, virtual content placed by one user in the physical world may not exist or may be inappropriately placed in the view of a second user, even if the second user intends to share the experience of the same physical space with the first user.

[0257] The inventors have recognized and understood that spatial persistence can be provided through a Persistent Coordinate Frame (PCF). The PCF can be defined based on one or more points representing features (e.g., corners, edges) identified in the physical world. The features can be selected such that they appear to be the same between one user instance of the XR system and another.

[0258] In addition, when rendering relative to a local map based only on a tracking map, tracking drift that causes the calculated tracking path (e.g., camera trajectory) to deviate from the actual tracking path can cause the position of virtual content to appear in an inappropriate location. As the XR device collects more information about the scene over time, the tracking map of the space can be refined to correct the drift. However, if the virtual content is placed on a real object before the map refinement and saved relative to the device's world coordinate system derived from the tracking map, the virtual content may appear displaced as if the real object has moved during the map refinement. The PCF can be updated according to the map refinement because the PCF is defined based on features and is updated as the features move during the map refinement.

[0259] The PCF can include six degrees of freedom including translation and rotation relative to the map coordinate system. The PCF can be stored in a local storage medium and / or a remote storage medium. The translation and rotation of the PCF can be calculated depending on, for example, the storage location relative to the map coordinate system. For example, the PCF used locally by a device has translation and rotation relative to the device's world coordinate system. The PCF in the cloud has translation and rotation relative to the canonical coordinate system of the canonical map.

[0260] The PCF can provide a sparse representation of the physical world, providing less information than all the available information about the physical world in order to efficiently process and transmit this information. Techniques for processing persistent space information can include creating a dynamic map across one or more sessions, based on one or more coordinate systems in the real space, generating a Persistent Coordinate Frame (PCF) above the sparse map, which can be exposed to XR applications via, for example, an Application Programming Interface (API).

[0261] Fig.14 is a block diagram showing the creation of a Persistent Coordinate Frame (PCF) and the attachment of XR content to the PCF according to some embodiments. Each block can represent digital information stored in a computer memory. In the case of application 1180, the data can represent computer-executable instructions. For example, in the case of virtual content 1170, the digital information can define a virtual object specified by, for example, application 1180. In the case of other blocks, the digital information can characterize some aspect of the physical world.

[0262] In the illustrated embodiment, one or more PCFs are created based on images captured by sensors on a wearable device. In Fig.14 the embodiment, the sensors are visual image cameras. These cameras can be the same cameras used to form a tracking map. Thus, Fig.14 some of the processing proposed can be performed as part of updating the tracking map. However, Fig.14 it is shown that information providing persistence is generated in addition to the tracking map.

[0263] To derive a 3D PCF, in a configuration that allows stereoscopic image analysis, two images 1110 from two cameras mounted on the wearable device are processed together. Fig.14 Images 1 and 2 are shown, each originating from one of the cameras. For simplicity, a single image from each camera is shown. However, each camera can output a stream of image frames, and Fig.14 the processing shown can be performed for multiple image frames in the stream.

[0264] Thus, each of Images 1 and 2 is one frame in a sequence of image frames. Fig.14 The processing shown can be repeated for consecutive image frames in the sequence until an image frame containing feature points that provide suitable images for forming persistent spatial information is processed. Alternatively or additionally, Fig.14 the processing can be repeated when the user moves such that the user is no longer close enough to a previously identified PCF to reliably use that PCF to determine a position relative to the physical world. For example, an XR system can maintain a current PCF for the user. When that distance exceeds a threshold, the system can switch to a new current PCF that is closer to the user and that can be generated according to Fig.14 the process using image frames acquired at the user's current position.

[0265] Even when generating a single PCF, a stream of image frames can be processed to identify image frames depicting content in the physical world that may be stable and easily recognizable by a device in the vicinity of the region of the physical world described in the image frame. In Fig.14 the embodiment, the processing starts with identifying features 1120 in the image. For example, features can be identified by finding gradient positions or other characteristics above a threshold in the image, which can, for example, correspond to corners of an object. In the illustrated embodiment, the features are points, but other recognizable features, such as edges, can be used alternatively or additionally.

[0266] In the illustrated embodiment, a fixed number N of feature points 1120 are selected for further processing. These feature points can be selected based on one or more criteria, such as gradient magnitude or proximity to other feature points. Alternatively or additionally, feature points can be selected heuristically, such as based on characteristics that imply the persistence of the feature points. For example, a heuristic can be defined based on the characteristics of feature points that may correspond to the corners of windows or doors or large pieces of furniture. Such a heuristic takes into account the feature points themselves and their surroundings. As a specific example, the number of feature points per image can be between 100 and 500, or between 150 and 250, such as 200.

[0267] Regardless of the number of feature points selected, descriptors 1130 can be calculated for the feature points. In this example, descriptors are calculated for each selected feature point, but descriptors can be calculated for multiple sets of feature points or subsets of feature points or all features within an image. The descriptors characterize the feature points such that feature points representing the same object in the physical world are assigned similar descriptors. The descriptors can facilitate the alignment of two frames, such as the alignment that may occur when one map is positioned relative to another map. The initial alignment of two frames can be performed by identifying feature points with similar descriptors, rather than searching for a relative frame orientation that minimizes the distance between the feature points of two images. The alignment of the image frames can be based on the alignment of points with similar descriptors, and doing so requires less processing than calculating the alignment of all feature points in the image.

[0268] The descriptors can be calculated as a mapping from the feature points to the descriptors, or in some embodiments, as a mapping from image patches around the feature points to the descriptors. The descriptors can be numerical quantities. U.S. Patent Application 16 / 190,948 describes calculating descriptors for feature points, and its entire content is incorporated herein by reference.

[0269] In Fig.14 the example, descriptors 1130 are calculated for each feature point in each image frame. Based on the descriptors and / or the feature points and / or the image itself, an image frame can be identified as a key frame 1140. In the illustrated embodiment, a key frame is an image frame that meets a specific criterion and is then selected for further processing. For example, when creating a tracking map, an image frame that adds meaningful information to the map can be selected as a key frame to be integrated into the map. On the other hand, an image frame that substantially overlaps with an area for which an image frame has already been integrated into the map can be discarded so that these image frames do not become key frames. Alternatively or additionally, key frames can be selected based on the number and / or type of feature points in the image frame. In Fig.14 the embodiment, the key frames 1150 selected to be included in the tracking map can also be considered key frames for determining the PCF, but different or additional criteria can also be used to select key frames for generating the PCF.

[0270] Although Fig.14 key frames are shown for further processing, the information obtained from the images can also be processed in other forms. For example, feature points in Keyrig can be processed alternatively or additionally. Furthermore, although key frames are described as originating from a single image frame, there does not have to be a one-to-one relationship between the key frames and the acquired image frames. For example, key frames can be obtained from multiple image frames, such as by stitching the image frames together or aggregating the image frames so that only features that appear in multiple images are retained in the key frames.

[0271] A key frame can include image information and / or metadata associated with the image information. In some embodiments, the images captured by cameras 44, 46 ( Fig. 9 ) can be computed into one or more key frames (e.g., key frames 1, 2). In some embodiments, a key frame can include a camera pose. In some embodiments, a key frame can include one or more camera images captured at the camera pose. In some embodiments, the XR system can determine that a portion of the camera image captured at the camera pose is useless, and thus that portion is not included in the key frame. Thus, using key frames to align new images with early scene knowledge can reduce the use of computing resources of the XR system. In some embodiments, a key frame can include an image and / or image data located at a position with a direction / angle. In some embodiments, a key frame can include a position and a direction from which one or more map points can be observed. In some embodiments, a key frame can include a coordinate system with an ID. U.S. Patent Application No. 15 / 877,359 describes key frames, and the entire content is hereby incorporated herein by reference.

[0272] Some or all of the key frames 1140 can be selected for further processing, such as generating a persistent pose 1150 for the key frame. This selection can be based on the characteristics of all feature points or a subset of feature points in the image frame. These characteristics can be determined by processing descriptors, features, and / or the image frame itself. As a specific example, the selection can be based on clusters of feature points identified as likely related to persistent objects.

[0273] Each key frame is associated with the camera pose at which the key frame was acquired. For key frames selected to be processed into a persistent pose, this pose information can be saved together with other metadata about the key frame, such as acquisition time and / or WiFi fingerprint and / or GPS coordinates at the acquisition location. In some embodiments, metadata such as GPS coordinates can be used, either alone or in combination, as part of a positioning process.

[0274] A persistent pose is an information source that the device uses to self-orient relative to previously acquired information about the physical world. For example, if the keyframes that create the persistent pose are incorporated into a map of the physical world, the device can use a sufficient number of feature points associated with the persistent pose in the keyframes to self-orient relative to that persistent pose. The device can align the current image of its surrounding environment with the persistent pose. This alignment can be based on matching the current image with the image 1110, features 1120, and / or descriptors 1130 that caused the persistent pose, or any subset of that image or those features or descriptors. In some embodiments, the current image frame that matches the persistent pose may be another keyframe that has been incorporated into the device's tracking map.

[0275] Information about the persistent pose can be stored in a format that facilitates sharing among multiple applications executed on the same or different devices. In Fig.14 the example of, some or all of the persistent poses can be reflected as a Persistent Coordinate Frame (PCF) 1160. Like the persistent pose, the PCF can be associated with the map and can include a set of features or other information that the device uses to determine its orientation relative to the PCF. The PCF can include a transformation that defines its transformation relative to the origin of its map, such that the device determines its position relative to any object in the physical world reflected in the map by associating its position with the PCF.

[0276] Since the PCF provides a mechanism for determining position relative to physical objects, applications (such as application 1180) can define the position of virtual objects relative to one or more PCFs that serve as anchors for virtual content 1170. For example, Fig.14 shows App 1 associating its virtual content 2 with PCF 1.2. Similarly, App 2 associates its virtual content 3 with PCF 1.2. App 1 is also shown associating its virtual content 1 with PCF 4.5, and App 2 is shown associating its virtual content 4 with PCF 3. In some embodiments, PCF 3 can be based on image 3 (not shown), and PCF 4.5 can be based on image 4 and image 5 (not shown), similar to how PCF 1.2 is based on image 1 and image 2. When rendering the virtual content, the device can apply one or more transformations to calculate information such as the position of the virtual content relative to the device's display and / or the position of the physical object relative to the desired position of the virtual content. Using the PCF as a reference can simplify such calculations.

[0277] In some embodiments, a persistent pose can be a coordinate position and / or orientation with one or more associated keyframes. In some embodiments, a persistent pose can be automatically created after the user has traveled a certain distance (e.g., three meters). In some embodiments, a persistent pose can act as a reference point during localization. In some embodiments, a persistent pose can be stored in a bonded world (e.g., bonded world module 538).

[0278] In some embodiments, a new PCF can be determined based on a predefined distance allowed between adjacent PCFs. In some embodiments, when the user travels a predetermined distance (e.g., five meters), one or more persistent poses can be calculated into PCFs. In some embodiments, a PCF can be associated with one or more world coordinate systems and / or canonical coordinate systems, e.g., in a bonded world. In some embodiments, depending on, e.g., security settings, a PCF can be stored in a local and / or remote database.

[0279] Fig.15 Method 4700 for establishing and using a persistent coordinate system according to some embodiments is shown. Method 4700 can start by capturing (act 4702) an image of a scene (e.g., Fig.14 Images 1 and 2 therein) using one or more sensors of an XR device. Multiple cameras can be used and one camera can generate multiple images, e.g., in one stream.

[0280] Method 4700 can include extracting (4704) interest points (e.g., Figure 7 Map point 702 therein, Fig.14 Feature 1120 therein) from the captured image, generating (act 4706) a descriptor of the extracted interest points (e.g., Fig.14 Descriptor 1130 therein), and generating (act 4708) keyframes (e.g., keyframe 1140) based on the descriptor. In some embodiments, the method can compare the interest points in the keyframes and form pairs of keyframes that share a predetermined number of interest points. The method can use individual keyframe pairs to reconstruct a portion of the physical world. The mapped portion of the physical world can be saved as 3D features (e.g., Figure 7in Keyrig 704). In some embodiments, a selection portion of keyframe pairs can be used to construct 3D features. In some embodiments, the result of the drawing can be selectively saved. Keyframes not used for constructing 3D features can be associated with the 3D features by poses. For example, the distance between keyframes can be represented by the covariance matrix between the poses of the keyframes. In some embodiments, keyframe pairs can be selected to construct 3D features such that the distance between every two constructed 3D features is within a predetermined distance, and the predetermined distance can be determined to balance the required computational amount and the accuracy of the resulting model. This method can provide an appropriate amount of data for the physical world model to facilitate efficient and accurate calculations using the XR system. In some embodiments, the covariance matrix of two images can include the covariance matrix between the poses (e.g., six degrees of freedom) of the two images.

[0281] Method 4700 can include generating (act 4710) a persistent pose based on keyframes. In some embodiments, the method can include generating a persistent pose based on 3D features reconstructed from keyframe pairs. In some embodiments, the persistent pose can be attached to the 3D features. In some embodiments, the persistent pose can include the poses of the keyframes used to construct the 3D features. In some embodiments, the persistent pose can include the average pose of the keyframes used to construct the 3D features. In some embodiments, a persistent pose can be generated such that the distance between adjacent persistent poses is within a predetermined value, such as in the range of one meter to five meters, any value therebetween, or any other suitable value. In some embodiments, the distance between adjacent persistent poses can be represented by the covariance matrix of the adjacent persistent poses.

[0282] Method 4700 can include generating (act 4712) a PCF based on the persistent pose. In some embodiments, the PCF can be attached to the 3D features. In some embodiments, the PCF can be associated with one or more persistent poses. In some embodiments, the PCF can include the pose of one of the associated persistent poses. In some embodiments, the PCF can include the average pose of the poses of the associated persistent poses. In some embodiments, a PCF can be generated such that the distance between adjacent PCFs is within a predetermined value, such as in the range of three meters to ten meters, any value therebetween, or any other suitable value. In some embodiments, the distance between adjacent PCFs can be represented by the covariance matrix of the adjacent PCFs. In some embodiments, the PCF can be exposed to the XR application, for example, via an application programming interface (API), such that the XR application can access the model of the physical world through the PCF without accessing the model itself.

[0283] Method 4700 may include associating image data of a virtual object to be displayed by an XR device with at least one PCF (act 4714). In some embodiments, the method may include calculating a translation and an orientation of the virtual object relative to the associated PCF. It should be understood that the virtual object need not be associated with a PCF generated by the device on which the virtual object is placed. For example, the device may retrieve a PCF stored in a canonical map in the cloud and associate the virtual object with the retrieved PCF. It should be understood that as the PCF is adjusted over time, the virtual object may move with the associated PCF.

[0284] Fig.16 Visual data and algorithms of a first XR device 12.1, a second XR device 12.2, and a server 20 are shown, according to some embodiments. Fig.16 The components shown may operate to perform some or all of the operations associated with generating, updating, and / or using spatial information, such as persistent poses, persistent coordinate systems, tracking maps, or canonical maps, as described herein. Although not shown, the first XR device 12.1 may be configured to be the same as the second XR device 12.2. The server 20 may have a map storage routine 118, a canonical map 120, a map transmitter 122, and a map merging algorithm 124.

[0285] The second XR device 12.2 that may be in the same scene as the first XR device 12.1 may include a Persistent Coordinate Frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that can be used to render virtual content, and a frame embedding generator 308 (see Fig.21 ). In some embodiments, a map download system 126, a PCF identification system 128, the Figure 2 , a positioning module 130, a canonical map merger 132, a canonical map 133, and a map publisher 136 may be grouped into a connected world unit 1304. The PCF integration unit 1300 may be connected to the connected world unit 1304 and other components of the second XR device 12.2 to allow retrieval, generation, use, upload, and download of the PCF.

[0286] Maps that include PCFs may enable more persistence in a changing world. In some embodiments, localizing a tracking map that includes, for example, image matching features may include selecting features representing persistent content from a map constituted by PCFs, which enables fast matching and / or localization. For example, a world in which people enter and exit a scene and objects such as doors move relative to the scene requires less storage space and a lower transmission rate and allows the use of a single PCF and its relationships relative to each other (e.g., an integrated constellation of PCFs) to map the scene.

[0287] In some embodiments, the PCF integration unit 1300 may include a PCF 1306, a PCF tracker 1308, a persistent pose acquirer 1310, a PCF inspector 1312, a PCF generation system 1314, a coordinate system calculator 1316, a persistent pose calculator 1318, and three transformers, which are stored in the data storage of the storage unit of the second XR device 12.2 in advance. The three transformers include a tracking map and persistent pose transformer 1320, a persistent pose and PCF transformer 1322, and a PCF and image data transformer 1324.

[0288] In some embodiments, the PCF tracker 1308 may have an on cue and an off cue that can be selected by the application 1302. The application 1302 may be executed by the processor of the second XR device 12.2 to, for example, display virtual content. The application 1302 may have an invocation to turn on the PCF tracker 1308 via the on cue. The PCF tracker 1308 may generate a PCF when the PCF tracker 1308 is turned on. The application 1302 may have a subsequent invocation to turn off the PCF tracker 1308 via the off cue. When the PCF tracker 1308 is turned off, the PCF tracker 1308 terminates PCF generation.

[0289] In some embodiments, the server 20 may include a plurality of persistent poses 1332 and a plurality of PCFs 1330 that have been previously saved in association with the canonical map 120. The map transmitter 122 may send the canonical map 120 together with the persistent poses 1332 and / or PCFs 1330 to the second XR device 12.2. The persistent poses 1332 and PCFs 1330 may be stored on the second XR device 12.2 in association with the canonical map 133. When Figure 2 localized to the canonical map 133, the persistent poses 1332 and PCFs 1330 can be stored in association with Figure 2 the ground.

[0290] In some embodiments, the persistent pose acquirer 1310 may acquire the Figure 2 persistent pose of the ground. The PCF inspector 1312 may be connected to the persistent pose acquirer 1310. The PCF inspector 1312 may retrieve a PCF from the PCF 1306 based on the persistent pose retrieved by the persistent pose acquirer 1310. The PCF retrieved by the PCF inspector 1312 may form an initial PCF group for PCF-based image display.

[0291] In some embodiments, the application 1302 may need to generate additional PCFs. For example, if the user moves to an area that has not been previously mapped, the application 1302 can turn on the PCF tracker 1308. The PCF generation system 1314 can be connected to the PCF tracker 1308 and start generating PCFs as the area begins to expand. The PCFs generated by the PCF generation system 1314 can form a second set of PCFs that can be used for PCF-based image display. Figure 2 begin to expand and begin to generate PCFs based on the area Figure 2 The PCFs generated by the PCF generation system 1314 can form a second set of PCFs that can be used for PCF-based image display.

[0292] The coordinate system calculator 1316 can be connected to the PCF checker 1312. After the PCF checker 1312 retrieves a PCF, the coordinate system calculator 1316 can call the head coordinate system 96 to determine the head pose of the second XR device 12.2. The coordinate system calculator 1316 can also call the persistent pose calculator 1318. The persistent pose calculator 1318 can be directly or indirectly connected to the frame embedding generator 308. In some embodiments, an image / frame can be designated as a key frame after traveling a threshold distance (e.g., 3 meters) from the previous key frame. The persistent pose calculator 1318 can generate a persistent pose based on multiple (e.g., three) key frames. In some embodiments, the persistent pose can be substantially the average of the coordinate systems of multiple key frames.

[0293] The tracking map and persistent pose transformer 1320 can be connected to the area Figure 2 and the persistent pose calculator 1318. The tracking map and persistent pose transformer 1320 can transform the area Figure 2 into a persistent pose to determine the persistent pose at the origin relative to the area Figure 2 The persistent pose and PCF transformer 1322 can be connected to the tracking map and persistent pose transformer 1320 and further connected to the PCF checker 1312 and the PCF generation system 1314. The persistent pose and PCF transformer 1322 can transform the persistent pose (the persistent pose that the tracking map has been transformed into) into the PCFs from the PCF checker 1312 and the PCF generation system 1314 to determine the PCFs relative to the persistent pose.

[0294] The PCF and image data transformer 1324 can be connected to the persistent pose and PCF transformer 1322 and the data channel 62. The PCF and image data transformer 1324 transforms the PCFs into image data 68. The rendering engine 30 can be connected to the PCF and image data transformer 1324 to display the image data 68 relative to the PCFs to the user.

[0295] The PCF and image data transformer 1324 can be connected to the persistent pose and PCF transformer 1322 and the data channel 62. The PCF and image data transformer 1324 transforms the PCFs into image data 68. The rendering engine 30 can be connected to the PCF and image data transformer 1324 to display the image data 68 relative to the PCFs to the user.

[0296] The PCF integration unit 1300 can store additional PCFs generated using the PCF generation system 1314 within the PCF 1306. The PCF 1306 can be stored relative to the persistent pose. When the map publisher 136 sends the map Figure 2 to the server 20, the map publisher 136 can retrieve the PCF 1306 and the persistent pose associated with the PCF 1306, and the map publisher 136 also sends the PCF and the persistent pose Figure 2 associated with the map to the server 20. When the map storage routine 118 of the server 20 stores the map Figure 2 , the map storage routine 118 can also store the persistent pose and the PCF generated by the second viewing device 12.2. The map merging algorithm 124 can create a canonical map 120, where the persistent pose and the PCF of the map Figure 2 are associated with the canonical map 120 and are respectively stored in the persistent pose 1332 and the PCF 1330.

[0297] The first XR device 12.1 can include a PCF integration unit similar to the PCF integration unit 1300 of the second XR device 12.2. When the map sender 122 sends the canonical map 120 to the first XR device 12.1, the map sender 122 can send the persistent pose 1332 and the PCF 1330 associated with the canonical map 120 and originating from the second XR device 12.2. The first XR device 12.1 can store the PCF and the persistent pose in the data storage of the storage device of the first XR device 12.1. The first XR device 12.1 can then use the persistent pose and the PCF originating from the second XR device 12.2 for image display relative to the PCF. Additionally or alternatively, the first XR device 12.1 can retrieve, generate, use, upload, and download PCFs and persistent poses in a manner similar to the above-described second XR device 12.2.

[0298] In the illustrated example, the first XR device 12.1 generates a local tracking map (hereinafter referred to as "the map Figure 1 ") and the map storage routine 118 receives the map Figure 1 from the first XR device 12.1. Then the map storage routine 118 stores the map Figure 1 on the storage device of the server 20 as the canonical map 120.

[0299] The second XR device 12.2 includes a map download system 126, an anchor point recognition system 128, a positioning module 130, a canonical map merger 132, a local content localization system 134, and a map publisher 136.

[0300] In use, the map transmitter 122 sends the canonical map 120 to the second XR device 12.2, and the map download system 126 downloads the canonical map 120 from the server 20 and stores it as the canonical map 133.

[0301] The anchor point recognition system 128 is connected to the world surface determination routine 78. The anchor point recognition system 128 identifies anchor points based on objects detected by the world surface determination routine 78. The anchor point recognition system 128 uses the anchor points to generate a second map (the Figure 2 ). As shown in loop 138, the anchor point recognition system 128 continues to identify anchor points and continues to update the Figure 2 . The anchor point positions are recorded as three-dimensional data based on the data provided by the world surface determination routine 78. The world surface determination routine 78 receives images from the real object detection camera 44 and depth data from the depth sensor 135 to determine the positions of the surfaces and their relative distances from the depth sensor 135.

[0302] The localization module 130 is connected to the canonical map 133 and the Figure 2 . The localization module 130 repeatedly attempts to localize the Figure 2 to the canonical map 133. The canonical map merger 132 is connected to the canonical map 133 and the Figure 2 . When the localization module 130 localizes the Figure 2 to the canonical map 133, the canonical map merger 132 merges the canonical map 133 into the anchor points of the Figure 2 . The is then updated with the missing data included in the canonical map. Figure 2

[0303] The local content localization system 134 is connected to the Figure 2 . For example, the local content localization system 134 can be a system for a user to localize local content at a specific position within the world coordinate system. The local content then attaches itself to an anchor point of the Figure 2 . The local-to-world coordinate transformer 104 transforms the local coordinate system to the world coordinate system based on the settings of the local content localization system 134. The functions of the rendering engine 30, the display system 42, and the data channel 62 have been described with reference to Figure 2 .

[0304] The map publisher 136 uploads the Figure 2 to the server 20. The map storage routine 118 of the server 20 then stores the Figure 2 in the storage medium of the server 20.

[0305] The map merging algorithm 124 merges the Figure 2Merge with the canonical map 120. When more than two maps (e.g., three or four maps related to the same or adjacent regions of the physical world) have been stored, the map merging algorithm 124 merges all maps into the canonical map 120 to render a new canonical map 120. The map transmitter 122 then sends the new canonical map 120 to any and all devices 12.1 and 12.2 in the region represented by the new canonical map 120. When devices 12.1 and 12.2 align their respective maps to the canonical map 120, the canonical map 120 becomes the elevated map.

[0306] Fig.17 Shows examples of generating key frames of a scene map according to some embodiments. In the example shown, a first key frame KF1 is generated for the door on the left wall of the room. A second key frame KF2 is generated for the corner area where the floor, left wall, and right wall of the room intersect. A third key frame KF3 is generated for the window area on the right wall of the room. A fourth key frame KF4 is generated for the distal area of the carpet on the floor of the room. A fifth key frame KF5 is generated for the area of the carpet closest to the user.

[0307] Fig.18 Shows the generation according to some embodiments Fig.17 Examples of persistent poses of the map. In some embodiments, a new persistent pose is created when the device measures a threshold distance traveled and / or when the application requests a new persistent pose (PP). In some embodiments, the threshold distance can be 3 meters, 5 meters, 20 meters, or any other suitable distance. Selecting a smaller threshold distance (e.g., 1m) can result in an increased computational load because more PP will be created and managed compared to a larger threshold distance. Selecting a larger threshold distance (e.g., 40m) can result in an increased virtual content placement error because fewer PP are created, which will result in fewer PCF being created, meaning the virtual content attached to the PCF is at a larger distance from the PCF (e.g., 30m), and the error increases as the distance from the PCF to the virtual content increases.

[0308] In some embodiments, a PP can be created at the start of a new session. This initial PP can be considered zero and can be visualized as the center of a circle with a radius equal to the threshold distance. When the device reaches the perimeter of the circle and, in some embodiments, the application requests a new PP, the new PP can be placed at the device's current location (at the threshold distance). In some embodiments, if the device can find an existing PP within the threshold distance of the device's new location, a new PP will not be created at the threshold distance. In some embodiments, when a new PP is created (e.g., Fig.14When the device is at a PP 1150), the device attaches one or more of the most recent key frames to the PP. In some embodiments, the position of the PP relative to the key frames may be based on the position of the device when the PP is created. In some embodiments, a PP is not created when the device travels a threshold distance unless the application requests the PP.

[0309] In some embodiments, when the application has virtual content to display to the user, the application may request a PCF from the device. A PCF request from the application may trigger a PP request, and a new PP is created after the device travels a threshold distance. Fig.18 Shows a first persistent pose PP1, which may have attached thereto the most recent key frames (e.g., KF1, KF2, and KF3) by calculating the relative pose between the key frames and the persistent pose. Fig.18 Also shown is a second persistent pose PP2, which may have attached thereto the most recent key frames (e.g., KF4 and KF5).

[0310] Fig.19 Shows an example of generating Fig.17 a PCF of a map according to some embodiments. In the example shown, PCF1 may include PP1 and PP2. As described above, the PCF can be used to display image data relative to the PCF. In some embodiments, each PCF may have coordinates in another coordinate system (e.g., a world coordinate system), and a PCF descriptor that uniquely identifies the PCF, for example. In some embodiments, the PCF descriptor may be calculated based on the feature descriptors of the features in the frames associated with the PCF. In some embodiments, various PCF constellations can be combined to represent the real world in a persistent manner, which requires less data and less data transmission.

[0311] FIG. 20A to FIG. 20C is a schematic diagram showing an example of establishing and using a persistent coordinate system. Fig. 20A Shows two users 4802A, 4802B, whose respective local tracking maps 4804A, 4804B have not been localized to the canonical map. The origins 4806A, 4806B of the respective users are depicted by the coordinate systems (e.g., world coordinate systems) in their respective regions. These origins of each tracking map are local origins for each user because these origins depend on the orientation of their respective devices at the start of tracking.

[0312] When the sensors of the user device scan the environment, the device can capture images, as described above in connection with Fig.14 which the images can contain features representing persistent objects, and thus these images can be classified as key frames for creating persistent poses. In this example, the tracking map 4804A includes a persistent pose (PP) 4808A; the tracking map 4804B includes a PP 4808B.

[0313] As described above in connection with Fig.14 some PPs can be classified as PCFs, and these PCFs are used to determine the orientation of virtual content for rendering it to a user. Fig. 20B It is shown that XR devices worn by respective users 4802A, 4802B can create local PCFs 4810A, 4810B based on PPs 4808A, 4808B. Fig. 20C It is shown that persistent content 4812A, 4812B (e.g., virtual content) can be attached to PCFs 4810A, 4810B via respective XR devices.

[0314] In this example, the virtual content can have a virtual content coordinate system that can be used by the application generating the virtual content, regardless of how the virtual content will be displayed. For example, the virtual content can be specified as a surface at a particular position and angle relative to the virtual content coordinate system, such as a triangle of a mesh. To render the virtual content to the user, the positions of these surfaces can be determined relative to the user perceiving the virtual content.

[0315] Attaching the virtual content to the PCF can simplify the calculations involved in determining the position of the virtual content relative to the user. The position of the virtual content relative to the user can be determined by applying a series of transformations. Some of these transformations may change and may be updated frequently. Other of these transformations may be stable and may be updated frequently or not at all. In any case, the transformations can be applied with a relatively low computational burden so that the position of the virtual content can be updated frequently relative to the user, providing a realistic appearance for the rendered virtual content.

[0316] In Figure 20A-20C the example, the device of user 1 has a coordinate system related to the coordinate system that defines the origin of the map via the transformation rig1_T_w1. The device of user 2 has a similar transformation rig2_T_w2. These transformations can be represented as 6 - degree - of - freedom transformations, specifying translation and rotation to align the device coordinate system with the map coordinate system. In some embodiments, the transformation can be represented as two separate transformations, one specifying translation and the other specifying rotation. Thus, it should be understood that the transformation can be expressed in a form that simplifies calculations or otherwise provides an advantage.

[0317] The transformations between the origin of the tracked map and the PCF recognized by respective user devices are represented as pcf1_T_w1 and pcf2_T_w2. In this example, the PCF and the PP are the same, so the same transformation also characterizes the PP.

[0318] The position of the user device relative to the PCF can thus be calculated by a series of applications of these transformations, such as rig1_T_pcf1 = (rig1_T_w1)*(pcf1_T_w1).

[0319] As Fig. 20C shown, the virtual content is positioned relative to the PCF and has the obj1_T_pcf1 transformation. This transformation can be set by the application that generates the virtual content, which can receive information from a world reconstruction system that describes the physical objects with respect to the PCF. To render the virtual content to the user, the transformation to the coordinate system of the user device is calculated, which can be calculated by associating the virtual content coordinate system with the origin of the tracking map via the transformation obj1_t_w1 = (obj1_T_pcf1)*(pcf1_T_w1). Then, this transformation is associated with the user device via the further transformation rig1_T_w1.

[0320] The position of the virtual content can be changed based on the output from the application that generates the virtual content. When a change occurs, the end-to-end transformation from the source coordinate system to the target coordinate system can be recalculated. Additionally, the position and / or head pose of the user can change as the user moves. Thus, the transformation rig1_T_w1 may change, as will any end-to-end transformation that depends on the position or head pose of the user.

[0321] The transformation rigl_T_wl can be updated based on the user's movement by tracking the position of the user relative to stationary objects in the physical world. Such tracking can be performed by the headset tracking component (described above) that processes the image sequence or other components of the system. Such an update can be made by determining the pose of the user relative to a stationary reference frame (such as the PP).

[0322] In some embodiments, the position and orientation of the user device can be determined relative to the most recent persistent pose or the PCF in this example (since the PP is used as the PCF). Such a determination can be made by identifying the feature points that characterize the PP in the current image captured by the sensors on the device. Using image processing techniques such as stereo image analysis, the position of the device relative to these feature points can be determined. Based on this data, the system can calculate the change in the transformation associated with the user's movement based on the relationship rig1_T_pcf1 = (rig1_T_w1)*(pcf1_T_w1).

[0323] The system can determine and apply the transformation in a computationally efficient order. For example, by tracking the user's pose and defining the position of the virtual content relative to the PP or the PCF established on the persistent pose, the need to compute rig1_T_w1 from the measurements that yield rig1_T_pcf1 can be avoided. Thus, the transformation from the source coordinate system of the virtual content to the target coordinate system of the user device can be based on the transformation measured according to the expression (rig1_T_pcf1)*(obj1_t_pcf1), where the first transformation is measured by the system and the subsequent transformation is provided by the application that specifies the virtual content to be rendered. In an embodiment where the virtual content is positioned relative to the origin of the map, the end-to-end transformation can be based on a further transformation between the map coordinates and the PCF coordinates to associate the virtual content coordinate system with the PCF coordinate system. In an embodiment where the virtual content is positioned relative to a different PP or PCF than the PP or PCF used to track the user's position, the transformation between the two can be applied. Such a transformation can be fixed and can be determined, for example, from a map in which both are present.

[0324] For example, a transformation-based method can be implemented in a device having components that process sensor data to build a tracking map. As part of this process, these components can identify feature points that serve as persistent poses, which in turn can be translated into PCFs. These components can limit the number of persistent poses for map generation to provide a suitable spacing between persistent poses while allowing the user, regardless of their position in the physical environment, to be close enough to the persistent pose positions to accurately compute the user's pose, as described above in connection with Figure 17-Figure 19 When the closest persistent pose to the user is updated due to user movement, refinement of the tracking map, or other reasons, any transformation used to compute the position of the virtual content relative to the user (depending on the position of the PP or, if using a PCF, depending on the position of the PCF) is updated and stored for use, at least until the user moves away from the persistent pose. Nevertheless, by computing and storing the transformation, the computational burden each time the virtual content position is updated can be relatively low, allowing for execution with relatively low latency.

[0325] Figure 20A-20C The positioning relative to the tracking map is shown, and each device has its own tracking map. However, the transformation can be generated relative to any map coordinate system. Content persistence across user sessions in an XR system can be achieved by using a persistent map. A shared experience for the user can also be facilitated by using a map to which multiple user devices can be oriented.

[0326] In some embodiments described in more detail below, the position of virtual content can be specified relative to coordinates in a canonical map, the format of which can be set to allow any of a plurality of devices to use the map. Each device can maintain a tracking map and can determine changes in the user's pose relative to the tracking map. In this example, the transformation between the tracking map and the canonical map can be determined by a "localization" process, which can be performed by matching structures in the tracking map (such as one or more persistent poses) with one or more structures in the canonical map (such as one or more PCFs).

[0327] Techniques for creating and using a canonical map in this manner are described in more detail below.

[0328] Depth keyframes

[0329] The techniques described herein rely on the comparison of image frames. For example, to establish the position of a device relative to a tracking map, new images can be captured using sensors worn by the user, and the XR system can search an image set used to create the tracking map for images that share at least a predetermined number of interest points with the new images. As an example of another scenario involving image frame comparison, the tracking map can be localized to the canonical map by first finding the image frame associated with a persistent pose in the tracking map (similar to the image frame associated with a PCF in the canonical map). Alternatively, the transformation between two canonical maps can be calculated by first finding similar image frames in both maps.

[0330] Depth keyframes provide a way to reduce the amount of processing required to identify similar image frames. For example, in some embodiments, the comparison can be made between image features in a new 2D image (e.g., "2D features") and 3D features in the map. This comparison can be made in any suitable manner, such as by projecting the 3D image onto a 2D plane. Traditional methods such as bag-of-words (BoW) search a database of all 2D features in the map for the 2D features of the new image, which can require a large amount of computational resources, especially when the map represents a large area. The traditional method then locates images that share at least one 2D feature with the new image, which includes images that are not useful for locating meaningful 3D features in the map. The traditional method then locates 3D features that are not meaningful relative to the 2D features in the new image.

[0331] The inventors have recognized and understood a technique for retrieving images in a map using fewer memory resources (e.g., a quarter of the memory resources used by BoW), higher efficiency (2.5 ms of processing time per key frame, 100 μs compared to 500 key frames), and higher accuracy (e.g., 20% higher retrieval recall than BoW for a 1024 - dimensional model, 5% higher retrieval recall than BoW for a 256 - dimensional model).

[0332] To reduce computation, a descriptor of an image frame can be computed, which is used to compare the image frame with other image frames. The descriptor can be stored as an alternative or supplement to the image frame and feature points. In a map where persistent poses and / or PCFs can be generated based on the image frame, the descriptors of one or more image frames that generate each persistent pose or PCF can be stored as part of the persistent pose and / or PCF.

[0333] In some embodiments, the descriptor can be computed based on feature points in the image frame. In some embodiments, a neural network is configured to compute a unique frame descriptor representing the image. The image can have a resolution higher than 1 megabyte to capture sufficient details of the 3D environment within the field of view of a device worn by a user. The frame descriptor can be much shorter, such as a string of numbers, e.g., in the range of 128 bytes to 512 bytes or any number in - between.

[0334] In some embodiments, the neural network is trained such that the computed frame descriptor indicates the similarity between images. An image in the map can be located by identifying the nearest images in a database including images used to generate the map, and these nearest images can have descriptors within a predetermined distance from the frame descriptor of the new image. In some embodiments, the distance between images can be represented by the difference between the frame descriptors of two images.

[0335] Fig.21 is a block diagram showing a system for generating a descriptor of a single image according to some embodiments. In the example shown, a frame embedding generator 308 is shown. In some embodiments, the frame embedding generator 308 can be used within the server 20, but alternatively or additionally, it can be executed in whole or in part in one of the XR devices 12.1 and 12.2 or any other device that processes images for comparison with other images.

[0336] In some embodiments, the frame embedding generator can be configured to generate a compact image data representation from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes), which still indicates the content in the image despite the size reduction. In some embodiments, the frame embedding generator can be used to generate an image data representation, which can be a key frame or a frame used otherwise. In some embodiments, the frame embedding generator 308 can be configured to convert an image at a specific location and orientation into a unique digital string (e.g., 256 bytes). In the example shown, the image 320 captured by the XR device can be processed by the feature extractor 324 to detect the points of interest 322 in the image 320. The points of interest can be based on or derived from the identified feature points, as described above for feature 1120 ( Fig.14 ) or as described elsewhere herein. In some embodiments, the points of interest can be represented by descriptors, as described above with reference to descriptor 1130 ( Fig.14 ), and the descriptors can be generated using a depth sparse feature method. In some embodiments, each point of interest 322 can be represented by a digital string (e.g., 32 bytes). For example, there may be n features (e.g., 100), and each feature is represented by a 32-byte string.

[0337] In some embodiments, the frame embedding generator 308 can include a neural network 326. The neural network 326 can include a multi-layer perceptron unit 312 and a max pool unit 314. In some embodiments, the multi-layer perceptron (MLP) unit 312 can include a multi-layer perceptron that can be trained. In some embodiments, the points of interest 322 (e.g., descriptors of the points of interest) can be compacted by the multi-layer perceptron 312 and output as a weighted combination 310 of the descriptors. For example, the MLP can reduce n features to m features that are less than n features.

[0338] In some embodiments, the MLP unit 312 can be configured to perform matrix multiplication. The multi-layer perceptron unit 312 receives multiple points of interest 322 of the image 320 and converts each point of interest into a corresponding digital string (e.g., 256). For example, there may be 100 features, and each feature can be represented by a string of 256 digits. In this example, a matrix with 100 horizontal rows and 256 vertical columns can be created. Each row has a series of 256 digits, and the sizes of these digits are different, some smaller and some larger. In some embodiments, the output of the MLP can be an n×256 matrix, where n represents the number of points of interest extracted from the image. In some embodiments, the output of the MLP can be an m×256 matrix, where m is the reduced number of points of interest from n.

[0339] In some embodiments, the MLP 312 can have a training phase and a usage phase, during which the model parameters of the MLP are determined. In some embodiments, the MLP can be trained as Fig.25 shown. The input training data can include triplets of data, where the three data in a triplet include 1) a query image, 2) a positive sample, and 3) a negative sample. The query image can be considered a reference image.

[0340] In some embodiments, the positive sample can include an image similar to the query image. For example, in some embodiments, similar means having the same object in both the query image and the positive sample image but viewing the object from different angles. In some embodiments, similar means having the same object in both the query image and the positive sample image, but the object is moved relative to the other image (e.g., left, right, up, down).

[0341] In some embodiments, the negative sample can include an image dissimilar to the query image. For example, in some embodiments, a dissimilar image does not contain any object prominent in the query image or contains only a small portion (e.g., <10%, 1%) of the objects prominent in the query image. In contrast, for example, a similar image has a majority (e.g., >50% or >75%) of the objects in the query image.

[0342] In some embodiments, interest points can be extracted from the images in the input training data and can be transformed into feature descriptors. These descriptors can be calculated for both Fig.25 the training images shown, as well as Fig.21 the features extracted in the operation of the frame embedding generator 308. In some embodiments, as described in U.S. Patent Application 16 / 190,948, a deep sparse feature (DSF) process can be used to generate the descriptors (e.g., DSF descriptors). In some embodiments, the DSF descriptor is n×32-dimensional. The descriptors can then be passed through the model / MLP to create a 256-byte output. In some embodiments, the model / MLP can have the same structure as the MLP 312, such that once the model parameters are set through training, the resulting trained MLP can be used as the MLP 312.

[0343] In some embodiments, the feature descriptor (e.g., the 256-byte output from the MLP model) can then be sent to a triplet margin loss module (which can be used only during the training phase of the MLP neural network and not during the usage phase of the MLP neural network). In some embodiments, the triplet margin loss module can be configured to select model parameters to reduce the difference between the 256-byte output of the query image and the 256-byte output of the positive sample, and to increase the difference between the 256-byte output of the query image and the 256-byte output of the negative sample. In some embodiments, the training phase can include feeding multiple triplet input images into the learning process to determine the model parameters. For example, the training process can continue until the difference for the positive images is minimized and the difference for the negative images is maximized, or until other suitable exit criteria are reached.

[0344] Return reference Fig.21 , the frame embedding generator 308 can include a pooling layer, shown here as a max pooling unit 314. The max pooling unit 314 can analyze each column to determine the maximum number in the corresponding column. The max pooling unit 314 can combine the maximum values of each column of numbers in the output matrix of the MLP 312 into a global feature string 316 of, for example, 256 numbers. It should be understood that the images processed in the XR system preferably have high-resolution frames, possibly reaching millions of pixels. The global feature string 316 is a relatively small number, which occupies relatively little memory and is easier to search than an image (e.g., with a resolution higher than 1 megabyte). Thus, images can be searched without analyzing each raw frame from the camera, and the cost of storing 256 bytes instead of the full frame is also lower.

[0345] Fig. 22 is a flowchart showing a method 2200 for computing an image descriptor according to some embodiments. The method 2200 can start with receiving (act 2202) multiple images captured by an XR device worn by a user. In some embodiments, the method 2200 can include determining (act 2204) one or more key frames from the multiple images. In some embodiments, act 2204 can be skipped and / or can occur after step 2210.

[0346] The method 2200 can include identifying (act 2206) one or more interest points in the multiple images with an artificial neural network, and computing (act 2208) feature descriptors for the respective interest points with an artificial neural network. The method can include computing (act 2210) a frame descriptor to represent an image for each image, at least partially based on the feature descriptors of the identified interest points in the image computed with the artificial neural network.

[0347] Fig.23It is a flowchart showing a method 2300 for localization using image descriptors according to some embodiments. In this example, a new image frame depicting the current position of the XR device can be compared with image frames stored in association with points in the map, such as the aforementioned persistent poses or PCFs. Method 2300 can start with receiving (act 2302) a new image captured by an XR device worn by a user. Method 2300 can include identifying (act 2304) one or more of the most recent key frames in a database that includes key frames for generating one or more maps. In some embodiments, the most recent key frames can be identified based on rough spatial information and / or previously determined spatial information. For example, the rough spatial information can indicate that the XR device is located in a geographic area represented by a 50m x 50m area of the map. Image matching can be performed only for points within this area. As another example, based on tracking, the XR system can know that the XR device was previously near a first persistent pose in the map and is moving in the direction of a second persistent pose in the map. The second persistent pose can be considered the most recent persistent pose, and the key frames stored with it can be considered the most recent key frames. Alternatively or additionally, other metadata such as GPS data or WiFi fingerprints can be used to select the most recent key frame or set of most recent key frames.

[0348] Regardless of how the most recent key frames are selected, the frame descriptors can be used to determine whether the new image matches any of the frames selected as being associated with nearby persistent poses. This determination can be made by comparing the frame descriptor of the new image with the frame descriptors of the most recent key frames or a subset of key frames in the database selected in any other suitable manner, and selecting the key frames that have frame descriptors within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors can be calculated by obtaining the difference between two strings of numbers representing the two frame descriptors. In embodiments where the strings are treated as strings of multiple quantities, the difference can be calculated as a vector difference.

[0349] Once a matching image frame is identified, the orientation of the XR device relative to that image frame can be determined. Method 2300 can include performing (act 2306) feature matching for 3D features in the map corresponding to the identified most recent key frame, and calculating (act 2308) the pose of the device worn by the user based on the feature matching results. In this way, the computationally intensive matching of feature points in two images can be performed for as few as one image that has been determined to potentially match the new image.

[0350] Fig.24is a flowchart showing a method 2400 for training a neural network according to some embodiments. Method 2400 may start by generating (act 2402) a data set including a plurality of image sets. Each of the plurality of image sets may include a query image, a positive sample image, and a negative sample image. In some embodiments, the plurality of image sets may include synthetic record pairs that are configured to inform the neural network of basic information, such as shape, for example. In some embodiments, the plurality of image sets may include real record pairs that may be recorded from the physical world.

[0351] In some embodiments, inliers may be calculated by fitting an essential matrix between two images. In some embodiments, sparse overlap may be calculated as the intersection over union (IoU) of the interest points seen in two images. In some embodiments, the positive sample may include at least twenty interest points that serve as inliers and are the same as the interest points in the query image. The negative sample may include fewer than ten inlier points. Less than half of the sparse points in the negative sample overlap with the sparse points in the query image.

[0352] Method 2400 may include, for each image set, calculating (act 2404) a loss by comparing the query image with the positive sample image and the negative sample image. Method 2400 may include modifying (act 2406) the artificial neural network based on the calculated loss such that the distance between the frame descriptor generated by the artificial neural network for the query image and the frame descriptor generated for the positive sample image is less than the distance between the frame descriptor generated for the query image and the frame descriptor generated for the negative sample image.

[0353] It should be understood that although methods and apparatuses for generating global descriptors for individual images have been described above, these methods and apparatuses may be configured to generate descriptors for individual maps. For example, a map may include a plurality of key frames, and each key frame may have a frame descriptor as described above. A max pooling unit may analyze the frame descriptors of the map key frames and combine the frame descriptors into a unique map descriptor for the map.

[0354] In addition, it should be understood that other architectures may be used for the processing described above. For example, separate neural networks for generating DSF descriptors and frame descriptors have been described. This method saves computation. However, in some embodiments, frame descriptors may be generated from selected feature points without first generating DSF descriptors.

[0355] Map Ranking and Merging

[0356] This disclosure describes methods and apparatus for ranking and merging multiple environmental maps in a cross-reality (XR) system. Map merging enables maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking the maps enables techniques as described herein, including map merging, which involves selecting maps from a map set based on similarity. In some embodiments, for example, a canonical map set may be maintained by the system, and the format of this set of canonical maps may be set in a manner accessible to any of a plurality of XR devices. These canonical maps may be formed by merging selection-tracked maps from these devices with other tracked maps or previously stored canonical maps. For example, the canonical maps may be ranked for selecting one or more canonical maps to merge with a new tracked map and / or for selecting one or more canonical maps from the set for use in a device.

[0357] To provide a realistic XR experience to a user, an XR system must understand the user's physical environment in order to correctly associate the positions of virtual objects with real objects. Information about the user's physical environment can be obtained from an environmental map of the user's location.

[0358] The inventors have recognized and understood that an XR system can provide an enhanced XR experience including real and / or virtual content to multiple users sharing the same world, whether those users are present in the world at the same time or different times, by enabling efficient sharing of environmental maps of the real / physical world collected by multiple users. However, there are significant challenges in providing such a system. Such a system may store multiple maps generated by multiple users and / or the system may store multiple maps generated at different times. For operations that may be performed using previously generated maps, such as, for example, localization as described above, significant processing may be required to identify relevant environmental maps of the same world (e.g., the same real-world location) from all of the environmental maps collected in the XR system. In some embodiments, a device may only have access to a small number of environmental maps, for example for localization. In some embodiments, a device may have access to a large number of environmental maps. For example, the inventors have recognized and understood techniques for quickly and accurately ranking the relevance of environmental maps from all possible environmental maps (such as, Fig.28 the universe of all canonical maps 120 in ). High-ranked maps can then be selected for further processing, such as rendering virtual objects on a user display, interacting realistically with the physical world around the user, or merging map data collected by that user with stored maps to create a larger or more accurate map.

[0359] In some embodiments, stored maps can be filtered based on multiple criteria to identify stored maps relevant to a user task at a location in the physical world. These criteria can indicate a comparison of a tracking map generated by a user's wearable device at that location with candidate environmental maps stored in a database. The comparison can be performed based on metadata associated with the maps, such as Wi-Fi fingerprints detected by the device that generated the map and / or a set of BSSIDs to which the device was connected when forming the map. The comparison can also be performed based on the compressed or uncompressed content of the map. For example, a comparison based on the compressed representation can be performed by comparing vectors calculated from the map content. For example, a comparison based on the uncompressed map can be performed by locating the tracking map within the stored map, and vice versa. Multiple comparisons can be performed in sequence based on the computational time required to reduce the number of candidate maps considered, where comparisons involving less computation are performed earlier in the sequence than other comparisons that require more computation.

[0360] Fig.26 An AR system 800 configured to rank and merge one or more environmental maps is shown according to some embodiments. The AR system can include a connected world model 802 of the AR device. Information populating the connected world model 802 can come from sensors on the AR device, and the AR device can include computer-executable instructions stored in a processor 804 (e.g., Figure 4 the local data processing module 570 therein), and the processor can perform some or all of the processing to convert the sensor data into a map. Such a map can be a tracking map as it can be constructed while the AR device collects sensor data during operation in an area. Area attributes can be provided along with the tracking map to indicate the area represented by the tracking map. These area attributes can be geographical location identifiers, such as coordinates represented as latitude and longitude or an ID used by the AR system to represent a location. Alternatively or additionally, the area attributes can be measurement characteristics that are most likely unique to the area. For example, the area attributes can be derived from parameters of a wireless network detected in the area. In some embodiments, the area attributes can be associated with the unique address of an access point near and / or connected to the AR system. For example, the area attributes can be associated with the MAC address or basic service set identifier (BSSID) of a 5G base station / router, Wi-Fi router, etc. Figure 1 In an example of

[0361] In Fig.26 the example, the tracking map can be merged with other maps of the environment. The map ranking unit 806 receives the tracking map from the device PW802 and communicates with the map database 808 to select environmental maps from the map database 808 and rank the environmental maps. The selected maps with higher rankings are sent to the map merging unit 810.

[0362] The map merger 810 can perform a merging process on the maps sent from the map ranker 806. The merging process requires merging the tracking map with some or all of the ranked maps and sending the new merged map to the connected world model 812. The map merger can merge maps by identifying maps depicting overlapping portions of the physical world. These overlapping portions can be aligned so that the information in the two maps can be aggregated into a final map. A canonical map can be merged with other canonical maps and / or tracking maps.

[0363] Aggregation requires extending one map with information from another map. Alternatively or additionally, aggregation requires adjusting the representation of the physical world in one map based on information in another map. For example, a subsequent map can reveal that an object generating a feature point has moved, and thus the map can be updated based on the subsequent information. Or, two maps can characterize the same area with different feature points, and aggregation requires selecting a set of feature points from the two maps to better represent the area. Regardless of the specific processing that occurs during the merging process, in some embodiments, the PCFs from all the maps being merged can be preserved so that applications that localize content relative to these PCFs can continue to operate. In some embodiments, the merging of maps can result in redundant persistent poses, and some persistent poses can be deleted. When a PCF is associated with a persistent pose to be deleted, the merged map may need to modify the PCF to be associated with the persistent poses remaining in the map after merging.

[0364] In some embodiments, a map can be refined when it is extended and / or updated. Refinement requires calculations to reduce internal inconsistencies between feature points that may represent the same object in the physical world. The inconsistencies may be caused by inaccuracies in the poses associated with keyframes that provide feature points representing the same object in the physical world. For example, such inconsistencies may be due to the XR device calculating a pose relative to a tracking map that is itself based on an estimated pose, such that errors in the pose estimation accumulate, resulting in a “drift” in pose accuracy over time. The map can be refined by performing bundle adjustment or other operations to reduce the inconsistencies of feature points from multiple keyframes.

[0365] During refinement, the position of a persistent point relative to the map origin can change. Accordingly, the transformation associated with that persistent point (e.g., a persistent pose or PCF) may change. In some embodiments, the XR system can recompute the transformations associated with any persistent points that have changed in conjunction with map refinement (whether as part of a merging operation or for other reasons). These transformations can be pushed from the component computing the transformation to the components using the transformation so that any use of the transformation is based on the updated position of the persistent point.

[0366] The connected world model 812 can be a cloud model, which can be shared by multiple AR devices. The connected world model 812 can store or otherwise access the environmental map in the map database 808. In some embodiments, when updating a previously computed environmental map, the previous version of the map can be deleted to remove expired data from the database. In some embodiments, when a previously computed environmental map is updated, the previous version of the map can be archived, allowing retrieval / viewing of the previous version of the environment. In some embodiments, permissions can be set such that only AR systems with specific read / write access rights can trigger the deletion / archiving of the previous version of the map.

[0367] AR devices in the AR system can access these environmental maps created based on the tracking maps provided by one or more AR devices / systems. The map ranking unit 806 can also be used to provide environmental maps to the AR devices. The AR device can send a message requesting its current location in the environmental map, and the map ranking unit 806 can be used to select the environmental map relevant to the requesting device and rank the environmental map relevant to the requesting device.

[0368] In some embodiments, the AR system 800 can include a downsampling unit 814 configured to receive a merged map from the cloud PW 812. The merged map received from the cloud PW 812 can be in a storage format suitable for the cloud, which can include high-resolution information, such as a large number of PCFs per square meter or multiple image frames or a huge set of feature points associated with the PCF. The downsampling unit 814 can be configured to downsample the cloud-format map into a format suitable for storage on the AR device. The device-format map has less data, such as fewer PCFs or less data stored for each PCF, to accommodate the limited local computing power and storage space of the AR device.

[0369] Fig. 27 is a simplified block diagram showing a plurality of canonical maps 120 that can be stored in a remote storage medium (e.g., the cloud). Each canonical map 120 can include a plurality of canonical map identifiers indicating the location of the canonical map within physical space, such as somewhere on the earth. These canonical map identifiers can include one or more of the following identifiers: a region identifier represented by a series of longitudes and latitudes, a frame descriptor (e.g., Fig.21 the global feature string 316 in Fig.21feature descriptors 310), and device identifiers indicating one or more devices contributing to the map. In the example shown, the canonical maps 120 are arranged geographically in a two-dimensional pattern as they might exist on the Earth's surface. The canonical maps 120 can be uniquely identified by corresponding longitude and latitude because any canonical maps with overlapping longitude and latitude can be merged into a new canonical map.

[0370] Fig.28 is a schematic diagram showing a method of selecting a canonical map according to some embodiments, which can be used to position a new tracking map to one or more canonical maps. The method can start from accessing (action 120) the universe of the canonical maps 120. As an example, the universe of the canonical maps can be stored in a database in a connected world (e.g., the connected world module 538). The universe of the canonical maps can include canonical maps from all previously visited locations. The XR system can filter the universe of all canonical maps into a small subset or just filter it into a single map. It should be understood that in some embodiments, due to bandwidth limitations, it is not possible to send all canonical maps to the viewing device. Selecting a subset of possible candidates that are chosen to match the tracking map to be sent to the device can reduce the bandwidth and latency associated with accessing the database of remote maps.

[0371] The method can include filtering (action 300) the universe of the canonical maps based on a region having a predetermined size and shape. In Fig. 27 the example shown, each square represents a region. Each square can cover 50m x 50m. Each square has six adjacent regions. In some embodiments, action 300 can select at least one matching canonical map 120 that covers the longitude and latitude including the location identifier received from the XR device, as long as at least one map exists at that longitude and latitude. In some embodiments, action 300 can select at least one adjacent canonical map that covers the longitude and latitude adjacent to the matching canonical map. In some embodiments, action 300 can select multiple matching canonical maps and multiple adjacent canonical maps. For example, action 300 can reduce the number of canonical maps by about ten times, e.g., from thousands to hundreds to form a first filtered selection. Alternatively or additionally, criteria other than longitude and latitude can be used to identify adjacent maps. For example, the XR device may have previously been positioned using a canonical map in the set as part of the same session. The cloud service may retain information about the XR device, including the previously positioned map. In this example, the maps selected in action 300 can include those that cover regions adjacent to the map to which the XR device was positioned.

[0372] The method may include a first filtering selection of the canonical map based on Wi-Fi fingerprinting (action 302). Action 302 may determine latitude and longitude based on the Wi-Fi fingerprint received from the XR device as part of the location identifier. Action 302 may compare the latitude and longitude from the Wi-Fi fingerprint with the latitude and longitude of the canonical map 120 to determine one or more canonical maps that form a second filtering selection. Action 302 may reduce the number of canonical maps by approximately tenfold, e.g., from hundreds to dozens (e.g., 50) of canonical maps that form the second selection. For example, the first filtering selection may include 130 canonical maps, the second filtering selection may be 50 out of the 130 canonical maps, and may exclude the other 80 out of the 130 canonical maps.

[0373] The method may include a second filtering selection of the canonical map based on key frames (action 304). Action 304 may compare the data representing the image captured by the XR device with the data representing the canonical map 120. In some embodiments, the data representing the image and / or the map may include feature descriptors (e.g., Fig.25 the DSF descriptor in Fig.21 and / or global feature strings (e.g.,

[0374] the 316 in Fig. 27 ). Action 304 may provide a third filtering selection of the canonical map. In some embodiments, for example, the output of action 304 may be only five out of the 50 canonical maps identified after the second filtering selection. The map transmitter 122 then sends one or more canonical maps based on the third filtering selection to the viewing device. Action 304 may reduce the number of canonical maps by approximately tenfold, e.g., from dozens to single digits (e.g., 5) of canonical maps that form the third selection. In some embodiments, the XR device may receive the canonical maps in the third filtering selection and attempt to localize to the received canonical maps.

[0375] In some embodiments, the cloud can receive the feature details of the live / new / current image captured by the viewing device, and the cloud can generate the global feature string 316 of the live image. The cloud can then filter the canonical map 120 based on the live global feature string 316. In some embodiments, the global feature string can be generated on the local viewing device. In some embodiments, the global feature string can be generated remotely, such as on the cloud. In some embodiments, the cloud can send the filtered canonical map together with the global feature string 316 associated with the filtered canonical map to the XR device. In some embodiments, when the viewing device aligns its tracking map to the canonical map, it can do so by matching the global feature string 316 of the local tracking map with the global feature string of the canonical map.

[0376] It should be understood that the operations of the XR device may not perform all of the actions (300, 302, 304). For example, if the universe of the canonical map is relatively small (e.g., 500 maps), the XR device attempting to align may filter the universe of the canonical map based on Wi-Fi fingerprints (e.g., action 302) and keyframes (e.g., action 304), but omit region-based filtering (e.g., action 300). Additionally, it is not necessary to compare the entire maps. In some embodiments, for example, the comparison of two maps can result in the identification of common persistent points, such as persistent poses or PCFs that appear in the new map and a selected map from the map universe. In such cases, descriptors can be associated with the persistent points, and these descriptors can be compared.

[0377] Fig.29 is a flowchart showing a method 900 for selecting one or more ranked environmental maps according to some embodiments. In the illustrated embodiment, the user AR device creating the tracking map is ranked. Thus, the tracking map can be used to rank the environmental maps. In embodiments where the tracking map is not available, some or all of the selection and ranking of the environmental maps that do not explicitly rely on the tracking map can be used.

[0378] Method 900 can start at action 902, where an atlas of maps from a database of environmental maps (whose format can be set to a canonical map) near the location where the tracking map is being formed can be accessed, and these maps can then be filtered for ranking. Additionally, at action 902, at least one region attribute of the region in which the user AR device is operating is determined. In the scenario where the user AR device is building a tracking map, the region attribute can correspond to the region where the tracking map is being created. As a specific example, the region attribute can be calculated based on the received signal from the access point to the computer network when the AR device calculates the tracking map.

[0379] Fig.30Illustrates an exemplary map ranking unit 806 of an AR system 800 according to some embodiments. The map ranking unit 806 can be executed in a cloud computing environment as it can include parts executed on an AR device and parts executed on a remote computing system such as a cloud. The map ranking unit 806 can be configured to execute at least a part of method 900.

[0380] Fig.31A Illustrates an example of area attributes AA1 - AA8 of a tracking map (TM) 1102 and environment maps CM1 - CM4 in a database according to some embodiments. As shown, an environment map can be associated with multiple area attributes. The area attributes AA1 - AA8 can include parameters of a wireless network detected by an AR device that calculates the tracking map 1102, e.g., the basic service set identifier (BSSID) of the network to which the AR device is connected and / or the strength of the received signal from the access point to the wireless network via, e.g., a network tower 1104. The parameters of the wireless network can conform to protocols including Wi-Fi and 5G NR. In Fig.32 the example shown, the area attribute is a fingerprint of the area where the user's AR device collects sensor data to form a tracking map.

[0381] Fig.31B Illustrates an example of a determined geographical location 1106 of a tracking map 1102 according to some embodiments. In the example shown, the determined geographical location 1106 includes a centroid point 1110 and an area 1108 that is circular around the centroid point. It should be understood that the determination of the geographical location in this application is not limited to the format shown. The determined geographical location can have any suitable format, e.g., including different area shapes. In this example, the geographical location is determined from the area attributes using a database that associates area attributes with geographical locations. The database is commercially available, e.g., this operation can be performed using a database that associates Wi-Fi fingerprints with locations represented as latitude and longitude.

[0382] In Fig.29 an embodiment, the map database containing environment maps can also include location data of these maps, including the latitude and longitude covered by the map. The processing at action 902 requires selecting a set of environment maps from this database that cover the same latitude and longitude determined for the area attributes of the tracking map.

[0383] Action 904 is the first filtering of the set of environment maps accessed in action 902. In action 902, the environment maps are retained in the set based on their proximity to the geographical location of the tracking map. This filtering step can be performed by comparing the latitude and longitude associated with the tracking map and the latitude and longitude associated with the environment maps in the set.

[0384] Fig.32 An example of operation 904 according to some embodiments is shown. Each area attribute may have a corresponding geographical location 1202. The environmental map set may include an environmental map having at least one area attribute, the geographical location of which overlaps with the geographical location of the determined tracking map. In the example shown, the identified environmental map set includes environmental maps CM1, CM2, and CM4, where each environmental map has at least one area attribute, the geographical location of which overlaps with the determined geographical location of the tracking map 1102. The environmental map CM3 associated with the area attribute AA6 is not included in the set because it is located outside the determined geographical location of the tracking map.

[0385] Other filtering steps may also be performed on the environmental map set to reduce / arrange the number of environmental maps finally processed in the set (such as for map merging or providing connected world information to the user device). Method 900 may include filtering (operation 906) the environmental map set based on the similarity of one or more identifiers of network access points associated with the tracking map. During map formation, a device that collects sensor data to generate a map may be connected to the network through a network access point, such as through Wi-Fi or a similar wireless communication protocol. The access point may be identified by a BSSID. When the user device moves through the area where data is collected to form a map, the user device may connect to multiple different access points. Similarly, when multiple devices provide information to form a map, these devices may have been connected through different access points, so there may also be multiple access points for map formation. Therefore, there may be multiple access points associated with the map, and the set of access points may be an indication of the location of the map. The strength of the signal from the access point (which may be reflected as an RSSI value) may provide further geographical information. In some embodiments, a list of BSSID and RSSI values may form an area attribute of the map.

[0386] In some embodiments, filtering the environmental map set based on the similarity of one or more identifiers of network access points may include retaining, in the environmental map set, the environmental map having the highest Jaccard similarity with respect to at least one area attribute of the tracking map, based on one or more identifiers of the network access point.

[0387] Fig.33An example of operation 906 according to some embodiments is shown. In the example shown, the network identifier associated with the region attribute AA7 can be determined as the identifier of the tracking map 1102. The environmental map set after operation 906 includes environmental map CM2 and environmental map CM4. The region attribute of environmental map CM2 has a relatively high Jaccard similarity with AA7, and environmental map CM2 also includes region attribute AA7. Environmental map CM1 is not included in this environmental map set because it has the lowest Jaccard similarity with AA7.

[0388] The processing at operations 902 - 906 can be performed based on the metadata associated with the maps without actually accessing the content of the maps stored in the map database. Other processing may involve accessing the content of the maps. Operation 908 indicates accessing the environmental maps remaining in the subset after filtering based on the metadata. It should be understood that if subsequent operations can be performed on the accessed content, this operation can be performed earlier or later in the process.

[0389] Method 900 may include filtering (operation 910) the environmental map set based on the similarity of a metric representing the content of the tracking map and the environmental maps in the environmental map set. The metric representing the content of the tracking map and the environmental maps may include a vector of values calculated based on the content of the maps. For example, as described above, the depth key - frame descriptors calculated for one or more key frames used to form the map can provide a metric for comparing the map or a part of the map. The metric can be calculated from the maps retrieved at operation 908, or can be pre - calculated and stored as metadata associated with these maps. In some embodiments, filtering the environmental map set based on the similarity of a metric representing the content of the tracking map and the environmental maps in the environmental map set may include: retaining in the environmental map set the environmental maps having the smallest vector distance between the vector of the characteristics of the tracking map and the vector representing the environmental maps in the environmental map set.

[0390] Method 900 may include further filtering (operation 912) the environmental map set based on the degree of match between a part of the tracking map and multiple parts of the environmental maps in the environmental map set. The degree of match can be determined as part of the localization process. As a non - limiting example, localization can be performed by identifying key points in the tracking map and the environmental maps that are similar enough (because they represent the same part of the physical world). In some embodiments, the key points can be features, feature descriptors, key frames / key rigs, persistent poses, and / or PCFs. Then the set of key points in the tracking map can be aligned to produce the best fit with the set of key points in the environmental map. The mean - square distance between the corresponding key points can be calculated and, if below a threshold for a specific region of the tracking map, used as an indication that the tracking map and the environmental map represent the same region of the physical world.

[0391] In some embodiments, filtering the environmental map set based on a degree of match between a portion of a tracking map and multiple portions of environmental maps in an environmental map set may include calculating a volume of the physical world represented by the tracking map (the volume being also represented in the environmental maps in the environmental map set), and retaining in the environmental map set environmental maps having a greater calculated volume than the environmental maps filtered out from the set. Fig.34 An example of operation 912 according to some embodiments is shown. In the example shown, the environmental map set after operation 912 includes environmental map CM4, which has a region 1402 that matches a region of tracking map 1102. Environmental map CM1 is not included in the set because it does not have a region that matches a region of tracking map 1102.

[0392] In some embodiments, the environmental map set may be filtered in the order of operation 906, operation 910, and operation 912. In some embodiments, the environmental map set may be filtered based on operations 906, 910, and 912, which may be performed in ascending order based on the processing required to perform the filtering. Method 900 may include loading (operation 914) the environmental map set and data.

[0393] In the example shown, the user database stores region identifiers indicating regions where the AR device is used. The region identifier may be a region attribute, which may include parameters of a wireless network detected when the AR device is in use. The map database may store multiple environmental maps constructed from data and associated metadata provided by the AR device. The associated metadata may include region identifiers derived from the region identifiers of the AR device that provided the data required to construct the environmental map. The AR device may send a message to the PW module indicating that a new tracking map is being created or has been created. The PW module may calculate the region identifier of the AR device and update the user database based on the received parameters and / or the calculated region identifier. The PW module may also determine the region identifier associated with the AR device requesting the environmental map, identify multiple environmental map sets from the map database based on the region identifier, filter these environmental map sets, and send the filtered multiple environmental map sets to the AR device. In some embodiments, the PW module may filter the multiple environmental map sets based on one or more criteria, such as the geographical location of the tracking map, the similarity of one or more identifiers of network access points associated with the tracking map and the environmental maps in the environmental map set, the similarity of a metric representing the content of the tracking map and the environmental maps in the environmental map set, and the degree of match between a portion of the tracking map and a portion of the environmental maps in the environmental map set.

[0394] Several aspects of some embodiments have been described as such. It should be understood that those skilled in the art will readily conceive of various changes, modifications, and improvements. As an example, embodiments are described in the context of an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein can be applied in a mixed reality (MR) environment or more generally in other XR environments and virtual reality (VR) environments.

[0395] As another example, embodiments are described in the context of a device such as a wearable device. It should be understood that some or all of the techniques described herein can be implemented via a network (such as the cloud), discrete applications, and / or devices, or any suitable combination of a network and discrete applications.

[0396] In addition, Fig.29 Examples of criteria that can be used to filter candidate maps to produce a high-ranked map set are provided. Other criteria can be used instead of or in addition to the described criteria. For example, if multiple candidate maps have similar values of a metric for filtering out less desirable maps, the characteristics of the candidate maps can be used to determine which maps are retained as candidates or filtered out. For example, larger or denser candidate maps can be prioritized over smaller candidate maps. In some embodiments, Figure 27-Figure 28 can describe Figure 29-Figure 34 all or part of the systems and methods described in

[0397] Fig.35 and Fig.36 is a schematic diagram showing an XR system configured to rank and merge multiple environmental maps according to some embodiments. In some embodiments, a persistent world (PW) can determine when to trigger ranking and / or merging of maps. In some embodiments, determining which maps to use can be at least partially based on the depth keyframes described above according to some embodiments with respect to Figure 21-25 description.

[0398] Fig.37 is a block diagram showing a method 3700 for creating an environmental map of a physical world according to some embodiments. Method 3700 can begin by positioning a tracked map (act 3702) captured by an XR device worn by a user into a canonical map set (e.g., a canonical map selected by the method of Fig.28 method and / or Fig.29 method 900). Act 3702 can include positioning the Keyrig of the tracked map into the canonical map set. The positioning result of each Keyrig may include the positioning pose of the Keyrig and a set of 2D-to-3D feature correspondences.

[0399] In some embodiments, method 3700 may include splitting (act 3704) a tracking map into connected components, which may be robustly merged by merging connected patches. Each connected component may include Keyrigs within a predetermined distance. Method 3700 may include merging (act 3706) connected components larger than a predetermined threshold into one or more canonical maps and removing the merged connected components from the tracking map.

[0400] In some embodiments, method 3700 may include merging (act 3708) canonical maps that are merged with the same connected components of the tracking map in a group. In some embodiments, method 3700 may include promoting (act 3710) the remaining connected components of the tracking map that have not been merged with any canonical map to canonical maps. In some embodiments, method 3700 may include merging (act 3712) the persistent poses and / or PCFs of the tracking map and canonical maps merged with at least one connected component of the tracking map. In some embodiments, method 3700 may include finalizing (act 3714) the canonical maps, for example, by fusing map points and pruning redundant Keyrigs.

[0401] Fig.38A and Fig.38B shows an environmental map 3800 created by updating a canonical map 700 according to some embodiments, where the canonical map 700 may be promoted from a tracking map 700 ( Figure 7 ) by a new tracking map. As shown and described with respect to Figure 7 , the canonical map 700 may provide a floor plan 706 of a reconstructed physical object in the corresponding physical world represented by points 702. In some embodiments, map points 702 may represent features of a physical object that may include multiple features. A new tracking map of the physical world may be captured and uploaded to the cloud for merging with map 700. The new tracking map may include map points 3802 and Keyrigs 3804, 3806. In the example shown, Keyrig 3804 represents a Keyrig that has been successfully located in the canonical map, for example, by establishing a correspondence with Keyrig 704 of map 700 (as Fig.38B shown). On the other hand, Keyrig 3806 represents a Keyrig that has not been located in map 700. In some embodiments, Keyrig 3806 may be promoted to a separate canonical map.

[0402] FIG. 39A to FIG. 39F is a schematic diagram showing an example of a cloud-based persistent coordinate system that provides a shared experience for users in the same physical space. Fig.39A shows, for example, a canonical map 4814 from the cloud being Figure 20A-20CThe XR devices worn by users 4802A and 4802B receive. The canonical map 4814 may have a canonical coordinate system 4806C. The canonical map 4814 may have a PCF 4810C, and the PCF 4810C has a plurality of associated PPs (e.g., 4818A, 4818B in 39C).

[0403] Fig.39B Illustrates that the XR device establishes a relationship between its respective world coordinate systems 4806A, 4806B and the canonical coordinate system 4806C. This can be done, for example, by locating the canonical map 4814 on the corresponding device. For each device, aligning the tracking map with the canonical map can result in a transformation between its local world coordinate system and the coordinate system of the canonical map.

[0404] Fig.39C Illustrates that due to the alignment, the transformation (e.g., transformations 4816A, 4816B) between the local PCF (e.g., PCF 4810A, 4810B) on the corresponding device and the corresponding persistent poses (e.g., PPs 4818A, 4818B) on the canonical map can be calculated. Through these transformations, each device can use its local PCF (which can be locally detected on the device by processing the images detected by the sensors on the device) to determine where on the local device to display the virtual content attached to the PPs 4818A, 4818B or other persistent points on the canonical map. This method can accurately position the virtual content relative to each user and allows each user to have the same experience of the virtual content in the physical space.

[0405] Fig.39D Illustrates a snapshot of the persistent poses from the canonical map to the local tracking map. It can be seen that the local tracking maps are interconnected via the persistent poses. Fig.39E Illustrates that the PCF 4810A on the device worn by user 4802A can be accessed in the device worn by user 4802B via the PP 4818A. Fig.39F Illustrates that the tracking maps 4804A, 4804B and the canonical map 4814 can be merged. In some embodiments, some PCFs may be removed due to the merge. In the example shown, the merged map includes the PCF 4810C of the canonical map 4814, but does not include the PCFs 4810A, 4810B of the tracking maps 4804A, 4804B. The PPs previously associated with the PCFs 4810A, 4810B may be associated with the PCF 4810C after the map merge.

[0406] Example

[0407] Fig.40 and Fig.41 Illustrates Fig. 9Example of the first XR device 12.1 using a tracking map. Fig.40 is a three-dimensional first local tracking map (ground Fig. 9 ) that can be generated by the first XR device according to some embodiments Figure 1 Two-dimensional representation. Fig.41 Shows the ground according to some embodiments Figure 1 Uploaded from the first XR device to Fig. 9 Server block diagram.

[0408] Fig.40 Shows the ground on the first XR device 12.1 Figure 1 And virtual content (Content 123 and Content 456). The ground Figure 1 Has an origin (Origin 1). The ground Figure 1 Includes many PCFs (PCF a to PCF d). From the perspective of the first XR device 12.1, as an example, PCF a is located at the origin of the ground Figure 1 And has X, Y, and Z coordinates (0,0,0), and PCF b has X, Y, and Z coordinates (-1,0,0). Content 123 is associated with PCF a. In this example, Content 123 has X, Y, and Z relationships relative to PCF a (1,0,0). Content 456 has a relationship relative to PCF b. In this example, Content 456 has X, Y, and Z relationships (1,0,0) relative to PCF b.

[0409] In Fig.41 The first XR device 12.1 uploads the ground Figure 1 To the server 20. In this example, the server does not store a canonical map of the same area of the physical world represented by the tracking map, and the tracking map is stored as an initial canonical map. The server 20 now has a canonical map based on the ground Figure 1 . The first XR device 12.1 has a canonical map that is empty at this stage. For the purpose of discussion, and in some embodiments, the server 20 does not include maps other than the ground Figure 1 . The second XR device 12.2 does not store any maps.

[0410] The first XR device 12.1 also sends its Wi-Fi signature data to the server 20. The server 20 can use the Wi-Fi signature data to determine the approximate location of the first XR device 12.1 based on intelligence collected from other devices that have connected to the server 20 or other servers in the past and the GPS locations recorded by such other devices. The first XR device 12.1 can now end the first session (see 8) and can disconnect from the server 20.

[0411] Fig.42 Shows according to some embodiments Fig.16 Schematic diagram of an XR system, showing that after the first user 14.1 terminates the first session, the second user 14.2 initiates a second session using the second XR device of the XR system. Fig.43A It is a block diagram showing the initiation of the second session by the second user 14.2. The first user 14.1 is shown in dashed lines because the first session of the first user 14.1 has ended. The second XR device 12.2 starts recording objects. The server 20 can use various systems with different granularities to determine that the second session of the second XR device 12.2 and the first session of the first XR device 12.1 are in the same vicinity. For example, the first XR device 12.1 and the second XR device 12.2 may include Wi-Fi signature data, Global Positioning System (GPS) positioning data, GPS data based on Wi-Fi signature data, or any other data indicating location to record their locations. Alternatively, the PCF identified by the second XR device 12.2 can display the similarity with the Figure 1 PCF of the ground.

[0412] As Fig.43B shown, the second XR device starts and begins to collect data, such as images 1110 from one or more cameras 44, 46. As Fig.14 shown, in some embodiments, the XR device (e.g., the second XR device 12.2) can collect one or more images 1110 and perform image processing to extract one or more feature / interest points 1120. Each feature can be converted into a descriptor 1130. In some embodiments, the descriptor 1130 can be used to describe a key frame 1140, and the key frame 1140 can have the position and orientation of an attached associated image. One or more key frames 1140 can correspond to a single persistent pose 1150, which can be automatically generated after a threshold distance (e.g., 3 meters) from the previous persistent pose 1150. One or more persistent poses 1150 can correspond to a single PCF 1160, which can be automatically generated after a predetermined distance (e.g., every 5 meters). Over time, as the user continues to move around the user environment and the XR device continues to collect more data, such as when collecting images 1110, additional PCFs (e.g., PCF 3 and PCF4, 5) can be created. One or more applications 1180 can run on the XR device and provide virtual content 1170 to the XR device to be presented to the user. The virtual content can have an associated content coordinate system, which can be placed relative to one or more PCFs. As Fig.43B shown, the second XR device 12.2 creates three PCFs. In some embodiments, the second XR device 12.2 can attempt to locate one or more canonical maps stored on the server 20.

[0413] In some embodiments, as Fig.43CAs shown, the second XR device 12.2 can download the canonical map 120 from the server 20. The ground on the second XR device 12.2 Figure 1 includes PCFs a to d and the origin 1. In some embodiments, the server 20 may have multiple canonical maps for different locations and may determine that the second XR device 12.2 is in the same vicinity as the first XR device 12.1 during the first session and send the canonical map of that vicinity to the second XR device 12.2.

[0414] Fig.44 shows that the second XR device 12.2 starts to identify PCFs for generating the ground Figure 2 . The second XR device 12.2 only identifies a single PCF, i.e., PCF 1,2. The X, Y, and Z coordinates of PCF 1,2 of the second XR device 12.2 can be (1, 1, 1). The ground Figure 2 has its own origin (origin 2), which can be based on the head pose of the device 2 when the device starts for the current head pose session. In some embodiments, the second XR device 12.2 can immediately attempt to map the ground Figure 2 to the canonical map. In some embodiments, the ground Figure 2 may not be mappable to the canonical map (the ground Figure 1 ) (i.e., the mapping may fail) because the system cannot identify any or sufficient overlap between the two maps. Mapping can be performed by identifying a portion of the physical world represented in the first map that is also represented in the second map and calculating the transformation between the first map and the second map required to align these portions. In some embodiments, the system can perform mapping based on a comparison of PCFs between the local map and the canonical map. In some embodiments, the system can perform mapping based on a comparison of persistent poses between the local map and the canonical map. In some embodiments, the system can perform mapping based on a comparison of key frames between the local map and the canonical map.

[0415] Fig.45 shows the ground Figure 2 after the second XR device 12.2 has identified other PCFs (PCF 1,2, PCF 3, PCF 4,5) of the ground Figure 2 . The second XR device 12.2 attempts to map the ground Figure 2 to the canonical map again. Since the ground Figure 2 has been extended to overlap at least a portion of the canonical map, the mapping attempt will succeed. In some embodiments, the overlap between the local tracking map, the ground Figure 2 and the canonical map can be represented by PCFs, persistent poses, key frames, or any other suitable intermediate or derived constructs.

[0416] In addition, the second XR device 12.2 has associated content 123 and content 456 with the Figure 2 PCFs 1, 2 and PCF 3 of. Content 123 has X, Y, and Z coordinates (1, 0, 0) relative to PCFs 1, 2. Similarly, content 456 has Figure 2 X, Y, and Z coordinates of (1, 0, 0) relative to PCF 3 in the

[0417] Fig.46A and Fig.46B shows a successful localization to the canonical map. The localization can be based on matching features in one map to another. Through appropriate transformations, here involving translation and rotation of one map relative to another, the overlapping region / volume / portion of map 1410 represents the Figure 2 common part of the Figure 1 and the canonical map. Since the Figure 2 created PCFs 3 and 4, 5 before localization, while the canonical map created PCFs a and c before creating the Figure 2 different PCFs are created to represent the same volume in real space (e.g., in different maps).

[0418] As Fig.47 shown, the second XR device 12.2 extends the Figure 2 to include PCFs a - d from the canonical map. The inclusion of PCFs a - d represents the Figure 2 localization to the canonical map. In some embodiments, the XR system can perform an optimization step to remove duplicate PCFs from the overlapping region, such as the PCFs in 1410, PCF 3, and PCF 4, 5. After Figure 2 localization, the placement of virtual content (such as content 456 and content 123) will be relative to the most recently updated PCF in the updated Figure 2 . Although the attachment of the PCF for the content has changed and although the PCFs of the Figure 2 have been updated, the virtual content appears in the same real - world position relative to the user.

[0419] As Fig.48 shown, the second XR device 12.2 continues to extend the Figure 2 because, for example, when the user walks back and forth in the real world, the second XR device 12.2 identifies more PCFs (e.g., PCFs e, f, g, and h). It can also be noted that the Figure 1 does not extend in the Fig.47 and Fig.48 .

[0420] Referring to Fig.49 , the second XR device 12.2 associates the Figure 2 Uploaded to server 20. Server 20 will Figure 2 in accordance with the specification Figure 1 start storing. In some embodiments, Figure 2 It can be uploaded to server 20 at the end of the session of the second XR device 12.2.

[0421] The canonical map within server 20 now includes PCF i, which was not included in the Figure 1 on the first XR device 12.1. When a third XR device (not shown) uploads a map to server 20 and the map includes PCF i, the canonical map on server 20 is extended to include PCF i.

[0422] In Fig.50 Server 20 will Figure 2 merge with the canonical map to form a new canonical map. Server 20 determines that PCF a to d are common to the canonical map and Figure 2 Server extends the canonical map to include PCF e to h and PCF 1, 2 from Figure 2 to form a new canonical map. The canonical maps on the first XR device 12.1 and the second XR device 12.2 are based on Figure 1 and are out of date.

[0423] In Figure 51 Server 20 sends the new canonical map to the first XR device 12.1 and the second XR device 12.2. In some embodiments, this occurs when the first XR device 12.1 and the second device 12.2 attempt to localize during different sessions or a new session or a subsequent session. The first XR device 12.1 and the second device 12.2 continue to align their respective local maps ( Figure 1 and Figure 2 respectively) to the new canonical map as described above.

[0424] As Figure 52 shown, the head coordinate system 96 or "head pose" is related to the Figure 2 PCF in. In some embodiments, the origin of the map (origin 2) is based on the head pose of the second XR device 12.2 at the start of the session. Since the PCF is created during the session, the PCF is placed relative to the origin 2 of the world coordinate system. The Figure 2 PCF of the map is used as a persistent coordinate system relative to the canonical coordinate system, where the world coordinate system is the world coordinate system of the previous session (e.g., Figure 40 the origin 1 of the Figure 1 in). These coordinate systems are associated by the same transformation used to align the Figure 2 to the canonical map, as described above in connection with Figure 46Bas discussed

[0425] The transformation from the world coordinate system to the head coordinate system 96 has been previously referenced Figure 9 and discussed Figure 52 The head coordinate system 96 shown has only two orthogonal axes, which are at specific coordinate positions relative to the Figure 2 PCF of the ground and at a specific angle relative to the Figure 2 ground. However, it should be understood that the head coordinate system 96 is in a three-dimensional position relative to the Figure 2 PCF of the ground and has three orthogonal axes in three-dimensional space

[0426] In Figure 53 the head coordinate system 96 has been moved relative to the Figure 2 PCF of the ground. The head coordinate system 96 has been moved because the second user 14.2 has moved their head. The user can move their head within six degrees of freedom (6dof). The head coordinate system 96 can thus move in 6dof, i.e., starting from its previous position in Figure 52 it moves in three-dimensional space around three orthogonal axes relative to the Figure 2 PCF of the ground. When Figure 9 the real object detection camera 44 and the inertial measurement unit 48 in

[0427] Figure 54 respectively detect the movement of the real object and the head unit 22, the head coordinate system 96 is adjusted. More information about head pose tracking is disclosed in U.S. Patent Application No. 16 / 221,065, entitled "Enhanced Pose Determination for Display Device", the entire content of which is incorporated herein by reference Figure 54 It is shown that sound can be associated with one or more PCFs. For example, a user can wear a headset or earphone with stereo sound. Conventional techniques can be used to simulate the position of sound through the earphone. The position of the sound can be at a fixed position such that when the user rotates their head to the left, the position of the sound rotates to the right, so that the user perceives the sound as coming from the same position in the real world. In this example, the positions of the sound are represented by sound 123 and sound 456. For the purposes of discussion Figure 48 the analysis of

[0428] Figure 55 and Figure 56 show further implementations of the above techniques. As referenced in Figure 8As described, the first user 14.1 has initiated a first session. As Figure 55 shown, the first user 14.1 has terminated the first session, as indicated by the dashed line. At the end of the first session, the first XR device 12.1 uploads the ground Figure 1 to the server 20. The first user 14.1 has now initiated a second session at a time later than the first session. The first XR device 12.1 does not download the ground Figure 1 from the server 20 because the ground Figure 1 has already been stored on the first XR device 12.1. If the ground Figure 1 is lost, the first XR device 12.1 downloads the ground Figure 1 from the server 20. Then the first XR device 12.1 continues to build the PCF of the ground Figure 2 , locates to the ground Figure 1 , and further develops the canonical map as described above. The ground Figure 2 of the first XR device 12.1 is then used to associate the above local content, head coordinate system, local sound, etc.

[0429] Reference Figure 57 and Figure 58 , it is also possible that more than one user interacts with the server in the same session. In this example, the third user 14.3 with the third XR device 12.3 joins the first user 14.1 and the second user 14.2. Each of the XR devices 12.1, 12.2, and 12.3 starts to generate its own map, namely the ground Figure 1 , the ground Figure 2 and the ground Figure 3 . As the XR devices 12.1, 12.2, and 12.3 continue to develop the ground Figure 1 , the ground Figure 2 and the ground Figure 3 , these maps are continuously uploaded to the server 20. The server 20 merges the ground Figure 1 , the ground Figure 2 and the ground Figure 3 to form the canonical map. Then the canonical map is sent from the server 20 to each of the XR device 12.1, XR device 12.2, and XR device 12.3.

[0430] Figure 59Aspects of a viewing method for restoring and / or resetting a head pose according to some embodiments are shown. In the example shown, at action 1400, a viewing device is powered on. At action 1410, in response to being powered on, a new session is started. In some embodiments, the new session may include establishing a head pose. One or more capture devices fixed to a head-mounted frame worn on a user's head capture the surface of the environment by first capturing an image of the environment and then determining the surface based on the image. In some embodiments, the surface data may be combined with data from a gravity sensor to establish the head pose. Other suitable methods for establishing the head pose may be used.

[0431] At action 1420, the processor of the viewing device enters a routine for tracking the head pose. As the user moves their head to determine the orientation of the head-mounted frame relative to the surface, the capture device continues to capture the surface of the environment.

[0432] At action 1430, the processor determines whether the head pose has been lost. The head pose may be lost due to "edge" cases that result in insufficient feature acquisition (such as too many reflective surfaces, insufficient light, blank walls, outdoors, etc.), or due to dynamic situations of a crowd that moves and forms part of the map. The routine at 1430 allows a certain amount of time (e.g., 10 seconds), so that there is sufficient time to determine whether the head pose has been lost. If the head pose has not been lost, the processor returns to 1420 and enters head pose tracking again.

[0433] If the head pose has been lost at action 1430, the processor enters a routine at 1440 to restore the head pose. If the head pose is lost due to insufficient light, the following message is displayed to the user via the display of the viewing device:

[0434] The system is detecting a situation of insufficient light. Please move to a place with more sufficient light.

[0435] The system will continue to monitor whether sufficient light is available and whether the head pose can be restored. Alternatively, the system may determine that the low texture of the surface has caused the head pose to be lost. In this case, the following prompt is provided to the user in the display as a suggestion for improving surface capture:

[0436] The system cannot detect enough surfaces with fine texture. Please move to an area with less rough and more finely textured surfaces.

[0437] At operation 1450, the processor enters a routine to determine if head pose recovery has failed. If head pose recovery has not failed (i.e., head pose recovery has been successful), the processor returns to operation 1420 by entering head pose tracking again. If head pose recovery fails, the processor returns to operation 1410 to establish a new session. As part of the new session, all cached data will be invalidated and then a head pose will be re-established. Any suitable method of head tracking can be used in conjunction with Figure 59 the process described. U.S. Patent Application No. 16 / 221,065 describes head tracking, the entire content of which is hereby incorporated by reference herein.

[0438] Remote localization

[0439] Various embodiments can utilize remote resources to facilitate persistent and consistent cross-reality experiences between individuals and / or groups of users. The inventors have recognized and understood that the operational benefits of an XR device having a canonical map as described herein can be achieved without downloading a canonical map set. The example implementation Figure 30 illustrated above shows downloading a canonical map to a device. For example, the benefits of not downloading the map can be achieved by sending feature and pose information to a remote service that maintains a canonical map set. According to some embodiments, a device seeking to use a canonical map to localize virtual content at a location specified relative to the canonical map can receive one or more transforms between the features and the canonical map from the remote service. These transforms can be used on a device that maintains information about the location of these features in the physical world to localize virtual content at a location specified relative to the canonical map, or to otherwise identify a location in the physical world specified relative to the canonical map.

[0440] In some embodiments, spatial information is captured by the XR device and transmitted to a remote service, such as a cloud-based service, that uses the spatial information to localize the XR device to a canonical map used by an application or other components of the XR system to specify the location of virtual content relative to the physical world. Once localized, the transform that links the tracking map maintained by the device to the canonical map can be transmitted to the device. The transform can be used in conjunction with the tracking map to determine the location of the rendered virtual content specified relative to the canonical map, or to otherwise identify a location in the physical world specified relative to the canonical map.

[0441] The inventors have realized that the data that needs to be exchanged between the device and the remote location service may be very small compared to transmitting map data, which may occur when the device transmits a tracking map to the remote service and receives a canonical map set from the remote service to enable device-based positioning. In some embodiments, performing the positioning function on cloud resources only requires sending a small amount of information from the device to the remote service. For example, it is not necessary to transmit the complete tracking map to the remote service to perform positioning. In some embodiments, features and pose information such as those described above in connection with persistent pose storage may be sent to the remote server. In embodiments where features are represented by descriptors, as described above, the uploaded information may be even smaller.

[0442] The result returned from the location service to the device can be one or more transformations that associate the uploaded features with multiple parts of the matching canonical map. These transformations can be used in the XR system in conjunction with its tracking map to identify the location of virtual content or otherwise identify locations in the physical world. In embodiments where persistent spatial information such as the above-described PCF is used to specify a location relative to the canonical map, the location service can download the transformation between the features and one or more PCFs to the device after successful positioning.

[0443] Therefore, the network bandwidth consumed by the communication between the XR device and the remote service for performing positioning may be very low. Thus, the system can support frequent positioning, allowing each device interacting with the system to quickly obtain information for positioning virtual content or performing other location-based functions. When the device moves in the physical environment, requests for updated location information may be repeated. In addition, the device may frequently obtain updates to the location information, such as when the canonical map changes, such as by incorporating additional tracking maps to expand the map or improve its accuracy.

[0444] In addition, uploading features and downloading transformations can enhance privacy in an XR system that shares map information among multiple users by increasing the difficulty of obtaining maps through spoofing. For example, unauthorized users can be prevented from obtaining maps from the system by sending false requests for canonical maps representing parts of the physical world where they are not located. When an unauthorized user is not actually in the physical world region for which they request map information, it is less likely that they will have access to the features in that region. In embodiments where the feature information is set in a feature description format, it will be more complex to obtain the feature information by spoofing requests for map information. In addition, when the system returns a transformation of the tracking map of a device intended to operate in the region for which the requested location information applies, the information returned by the system is of little or no use to an imposter.

[0445] According to some embodiments, the location service is implemented as a cloud-based microservice. In some examples, implementing a cloud-based location service can help save device computing resources and enable the computations required for location to be performed with very low latency. These operations can be supported by nearly unlimited computing power or other computing resources available by provisioning additional cloud resources, ensuring that the XR system can scale to support numerous devices. In one example, many canonical maps can be maintained in memory for near-instant access or alternatively stored in highly available devices to reduce system latency.

[0446] In addition, performing location for multiple devices in a cloud service can enable process improvements. Location telemetry and statistics can provide information about which canonical maps are in active memory and / or highly available storage. For example, statistics from multiple devices can be used to identify the most frequently accessed canonical maps.

[0447] Additional accuracy can also be obtained due to processing in a cloud environment or other remote environments with a large amount of processing resources relative to the remote device. For example, location can be performed on a higher density of canonical maps in the cloud relative to processing performed on a local device. The maps can be stored in the cloud, such as having more PCFs or a higher density of feature descriptors per PCF, thus improving the accuracy of the match between the device's feature set and the canonical map.

[0448] Figure 61 is a schematic diagram of an XR system 6100. The user device that displays cross-reality content during a user session can take many forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As described above, these devices can be configured with software, such as an application or other components, and / or hardwired to generate local location information (e.g., a tracking map) that can be used to render virtual content on their respective displays.

[0449] The virtual content location information can be specified relative to global location information. For example, the format of the global location information can be set to include a canonical map with one or more PCFs. According to some embodiments, for example Figure 61 in the illustrated embodiment, the system 6100 is configured with a cloud-based service that supports the running and display of virtual content on the user device.

[0450] In one example, the positioning function is provided as a cloud-based service 6106, which can be a microservice. The cloud-based service 6106 can be implemented on any of a plurality of computing devices from which computing resources can be allocated to one or more services executing in the cloud. These computing devices can be interconnected with each other and accessible by devices such as the wearable XR device 6102 and the handheld device 6104. These connections can be provided via one or more networks.

[0451] In some embodiments, the cloud-based service 6106 is configured to receive descriptor information from respective user devices and "locate" the devices to one or more matching canonical maps. For example, the cloud-based positioning service matches the received descriptor information with the descriptor information of the corresponding canonical map. The techniques described above can be used to create the canonical maps, which are created by merging maps provided by one or more devices having image sensors or other sensors that acquire information about the physical world. However, the canonical maps do not have to be created by the devices accessing them, as these maps can be created by map developers, e.g., developers who make the maps available to the positioning service 6106.

[0452] According to some embodiments, the cloud service processes canonical map identification and can include an operation of filtering a repository of canonical maps into a set of potential matches. The filtering can be performed as shown in Figure 29 or, as an alternative or addition to the filtering criteria shown in Figure 29 , using any subset of the filtering criteria and other filtering criteria. In one embodiment, geographic data can be used to limit the search for matching canonical maps to maps representing areas close to the device requesting the location. For example, regional attributes such as Wi-Fi signal data, Wi-Fi fingerprint information, GPS data, and / or other device location information can be used as a coarse filter on the stored canonical maps, thereby limiting the analysis of descriptors to canonical maps known or likely to be near the user device. Similarly, the location history of each device can be maintained by the cloud service to prioritize the search for canonical maps near the device's last location. In some examples, the filtering can include the functions discussed above with respect to Figure 31B , Figure 32 , Figure 33 and Figure 34 .

[0453] Figure 62This is an example process flow that can be executed by a device to use cloud-based services to utilize one or more canonical maps to locate the device's position and receive transformation information specifying one or more transformations between the device's local coordinate system and the canonical map coordinate system. Various embodiments and examples may describe one or more transformations as specifying a transformation from a first coordinate system to a second coordinate system. Other embodiments include transformations from a second coordinate system to a first coordinate system. In yet another embodiment, the transformation implements a transformation from one coordinate system to another, and the resulting coordinate system depends only on the desired coordinate system output, such as the coordinate system for displaying content, the output. In still other embodiments, the coordinate system transformation allows the determination of the first coordinate system based on the second coordinate system and the determination of the second coordinate system based on the first coordinate system.

[0454] According to some embodiments, information reflecting the transformation for each persistent pose defined relative to the canonical map may be sent to the device.

[0455] According to some embodiments, process 6200 may start at 6202 with a new session. Starting a new session on the device initiates the capture of image information to construct a tracking map of the device. Additionally, the device may send a message to register the location service with the server, prompting the server to create a session for the device.

[0456] In some embodiments, starting a new session on the device optionally includes sending adjustment data from the device to the location service. The location service returns to the device one or more transformations calculated based on the feature and the associated set of poses. If the feature poses are adjusted based on device-specified information before calculating the transformation and / or the transformation is adjusted based on device-specific information after calculating the transformation, rather than performing these calculations on the device, the device-specified information may be sent to the location service so that the location service applies the adjustment. As a specific example, sending device-specified adjustment information may include capturing calibration data for sensors and / or displays. For example, the calibration data can be used to adjust the position of feature points relative to the measured position. Alternatively or additionally, the calibration data can be used to adjust the position where the command display renders virtual content so that it appears accurately positioned for the specified device. For example, the calibration data can be obtained from multiple images of the same scene taken using sensors on the device. The positions of the features detected in these images can be represented as a function of the sensor positions, such that the multiple images result in a system of equations that can be solved for the sensor positions. The calculated sensor positions can be compared to the nominal positions, and the calibration data can be derived from any differences. In some embodiments, the calibration data for the display can also be calculated using intrinsic information about the device's construction.

[0457] In embodiments for generating calibration data for a sensor and / or a display, the calibration data can be applied at any point during the measurement or display process. In some embodiments, the calibration data can be sent to a location server, which can store the calibration data in a data structure established for each device that has registered with the location server and thus has a session with the server. The location server can apply the calibration data to any transformation computed as part of the location process of the device providing the calibration data. Thus, the computational burden of using the calibration data to improve the accuracy of sensed and / or displayed information is borne by the calibration service, providing a further mechanism to reduce the processing burden on the device.

[0458] Once a new session is established, process 6200 can continue at 6204 to capture a new frame of the device environment. At 6206, each frame can be processed to generate a descriptor of the captured frame (e.g., including the DSF values discussed above). These values can be computed using some or all of the techniques described above, including those described above with respect to Figure 14 , Figure 22 and Figure 23 discussed techniques. As discussed, the descriptor can be computed as a mapping of feature points to the descriptor, or in some embodiments, as a mapping of image patches around the feature points to the descriptor. The descriptor can have values that enable an effective match between the newly acquired frame / image and the stored map. Additionally, the number of features extracted from the image can be limited to a maximum number of feature points per image, such as 200 feature points per image. As described above, the feature points can be selected to represent points of interest. Thus, actions 6204 and 6206 are performed as part of a device process for forming a tracking map or otherwise periodically collecting images of the physical world around the device, or can be but need not be performed separately for localization.

[0459] Feature extraction at 6206 can include attaching pose information to the features extracted at 6206. The pose information can be a pose in the device's local coordinate system. In some embodiments, the pose can be relative to a reference point in the tracking map, such as the persistent pose described above. Alternatively or additionally, the pose can be relative to the origin of the device's tracking map. Such embodiments can enable the localization services described herein to provide localization services for a wide range of devices, even if they do not use a persistent pose. In any case, the pose information can be attached to each feature or each set of features such that the localization service can use the pose information to compute a transformation that can be returned to the device when matching the features with features in the stored map.

[0460] Process 6200 can continue to decision block 6207, where a decision is made whether to request a localization. One or more criteria can be applied to determine whether to request a localization. The criteria can include the passage of time such that the device can request a localization after a certain threshold amount of time. For example, if no localization attempt has been made within the threshold amount of time, the process can continue from decision block 6207 to action 6208, where action 6208 requests a localization from the cloud. The threshold amount of time can be between 10 and 30 seconds, such as 25 seconds. Alternatively or additionally, the localization can be triggered by the movement of the device. The device executing process 6200 can use an IMU and / or its tracking map to track its movement and initiate a localization when it detects movement that is more than a threshold distance from the location where the device last requested a localization. For example, the threshold distance can be between 1 and 10 meters, such as between 3 and 5 meters. As yet another alternative, the localization can be triggered in response to an event, such as when the device creates a new persistent pose or the current persistent pose of the device changes as described above.

[0461] In some embodiments, decision block 6207 can be implemented such that the threshold for triggering a localization can be established dynamically. For example, where the features are substantially consistent such that the confidence in matching the set of features to be extracted with the features of the stored map is low, localization may be requested more frequently to increase the chance of at least one localization attempt being successful. In such a case, the threshold applied at decision block 6207 can be lowered. Similarly, in an environment with relatively few features, the threshold applied at decision block 6207 can be lowered to increase the frequency of localization attempts.

[0462] Regardless of how the localization is triggered, when triggered, process 6200 can proceed to action 6208, at which time the device sends a request to the localization service, which includes the data that the localization service uses to perform the localization. In some embodiments, data from multiple image frames can be provided for the localization attempt. For example, the localization service will not consider the localization successful unless the features in the multiple image frames produce consistent localization results. In some embodiments, process 6200 can include saving the feature descriptors and additional pose information to a buffer. For example, the buffer can be a circular buffer that stores the set of features extracted from the most recently captured frames. Thus, the localization request can be sent along with multiple sets of features accumulated in the buffer. In certain settings, the buffer size is implemented to accumulate multiple data sets, which will be more likely to result in a successful localization. In some embodiments, the buffer size can be set to accumulate features from, for example, two, three, four, five, six, seven, eight, nine, or ten frames. Optionally, the buffer size can have a baseline setting that increases in response to a localization failure. In some examples, increasing the buffer size and the corresponding number of feature sets transmitted reduces the likelihood that subsequent localization functions will not return a result.

[0463] Regardless of how the buffer size is set, the device can transmit the contents of the buffer to the location service as part of a location request. Other information can be sent along with the feature points and additional pose information. For example, in some embodiments, geographic information can be sent. The geographic information can include, for example, GPS coordinates or wireless signatures associated with the device's tracking map or the current persistent pose.

[0464] In response to the request sent at 6208, the cloud location service can analyze the feature descriptors to locate the device to a canonical map or other persistent maps maintained by the service. For example, the descriptors match a set of features in the map to which the device is located. The cloud-based location service can perform the location as described above with respect to the location of the device (e.g., can rely on any of the functions discussed above for location, including map ranking, map filtering, position estimation, filtered map selection, Figure 44 - Figure 4 the examples in 6, and / or those discussed with respect to location module, PCF, and / or PP identification and matching, etc.). However, instead of sending the identified canonical map to the device (e.g., in device location), the cloud-based location service can continue to generate a transformation based on the relative orientation of the feature set sent from the device and the matching features of the canonical map. The location service can return these transformations to the device, and the transformations can be received at block 6210.

[0465] In some embodiments, the canonical map maintained by the location service can employ PCFs, as described above. In such embodiments, the feature points of the canonical map that match the feature points sent from the device can have positions specified relative to one or more PCFs. Thus, the location service can identify one or more canonical maps and can calculate the transformation between the coordinate system represented in the pose sent with the location request and the one or more PCFs. In some embodiments, identifying one or more canonical maps is aided by filtering potential maps based on the geographic data of the corresponding device. For example, once filtered into a candidate set (e.g., by GPS coordinates and other options), the candidate set of canonical maps can be analyzed in detail to determine the matching feature points or PCFs described above.

[0466] The format of the data returned to the requesting device at action 6210 can be set to a persistent pose transformation table. This table can contain one or more canonical map identifiers indicating the canonical map to which the device is located by the location service. However, it should be understood that the location information can be formatted in other ways, including as a list of transformations with associated PCFs and / or canonical map identifiers.

[0467] Regardless of how the transformation format is set, at action 6212, the device can use these transformations to calculate the position for rendering virtual content, the position of which has been specified by an application or other component of the XR system relative to any one of the PCFs. This information can alternatively or additionally be used on the device to perform any location-based operations where the location is specified based on the PCF.

[0468] In some scenarios, the positioning service may not be able to match the features sent from the device with any stored canonical map, or may not be able to match a sufficient number of feature sets communicating with the positioning service request to determine that positioning has been successfully performed. In such scenarios, the positioning service can indicate a positioning failure to the device, rather than returning the transformation to the device as described above in connection with action 6210. In such scenarios, process 6200 can branch to action 6230 at decision box 6209, at which time the device can take one or more actions for failure handling. These actions include increasing the size of the buffer that stores the feature sets sent for positioning. For example, if the positioning service does not consider positioning successful unless three features actively match, the buffer size may increase from 5 to 6, thereby increasing the chance that the three feature sets sent match the canonical map maintained by the positioning service.

[0469] Alternatively or additionally, failure handling can include adjusting the device operation parameters to trigger more frequent positioning attempts. For example, the threshold time and / or threshold distance between positioning attempts can be reduced. As another example, the number of feature points in each feature set can be increased. A match can be considered to exist between the feature set and the features stored in the canonical map when a sufficient number of the features from the feature set sent from the device match the features of the map. Increasing the number of features sent increases the chance of a match. As a specific example, the initial feature set size can be 50, and at each successive positioning failure, it can be increased to 100, 150, and then 200. Upon successful matching, the size of the feature set can be returned to its initial value.

[0470] Failure handling can also include obtaining positioning information from sources other than the positioning service. According to some embodiments, the user device can be configured to cache the canonical map. Caching the map allows the device to access and display content when the cloud is unavailable. For example, the cached canonical map allows for device-based positioning in the event of communication failure or other unavailability.

[0471] According to various embodiments, Figure 62 a high-level process for a device to initiate cloud-based positioning is described. In other embodiments, one or more of the various steps shown can be combined, omitted, or other processes can be invoked to complete the positioning and ultimately achieve the visualization of virtual content in the view of the corresponding device.

[0472] In addition, it should be understood that although process 6200 shows that the device determines whether to initiate positioning at decision block 6207, the trigger for initiating positioning can come from outside the device, including from a positioning service. For example, the positioning service can maintain information about each of the devices in the session with it. For example, the information may include an identifier of the canonical map to which each device was most recently located. The positioning service or other components of the XR system can update the canonical map, including using the above combined Figure 26 When a canonical map is updated, the location service can send a notification to each device that has recently been located in the map. The notification can serve as a trigger for the device to request a location and / or can include an updated transform that is recalculated using the feature set most recently sent from the device.

[0473] Figure 63A , Figure 63B and Figure 63C 6350, 6352, 6354, and 6456 illustrate an example architecture and separation between components involved in a cloud-based positioning process. For example, a module, component, and / or software configured to process perception on a user device is shown at 6350 (e.g., Figure 6A Device functionality for persistent world operations is shown at 6352 (e.g., including the above description of the persistent world module (e.g., Figure 6A In other embodiments, separation between 6350 and 6352 is not required and the communication shown may exist between processes executing on the device.

[0474] Similarly, block 6354 shows an example of a method configured to process data associated with a connected world / connected world modeling (e.g., Figure 26 Block 6356 shows a cloud process configured to process functions associated with locating a device to one or more maps of a repository of stored canonical maps based on information sent from the device.

[0475] In the illustrated embodiment, when a new session begins, process 6300 begins at 6302. Sensor calibration data is obtained at 6304. The calibration data obtained may depend on the device (e.g., multiple cameras, sensors, positioning devices, etc.) represented at 6350. Once the sensor calibration of the device is obtained, the calibration may be cached at 6306. If device operation results in a change in frequency parameters (e.g., collection frequency, sampling frequency, matching frequency, etc. options), the frequency parameters are reset to a baseline at 6308.

[0476] Once the new session functionality is complete (e.g., calibration, steps 6302 - 6306), process 6300 can proceed to capture a new frame 6312. Features and their corresponding descriptors are extracted from the frame at 6314. In some examples, as described above, the descriptor can include a DSF. According to some embodiments, the descriptor can have spatial information attached to it to facilitate subsequent processing (e.g., transform generation). Pose information generated on the device (e.g., the information discussed above for positioning features in the physical world relative to the tracking map of the device) can be attached to the extracted descriptor at 6316.

[0477] At 6318, the descriptors and pose information are added to a buffer. The new frame capture and addition to the buffer shown in steps 6312 - 6318 are performed in a loop until the buffer size threshold is exceeded at 6319. In response to determining that the buffer size is met, a localization request is sent from the device to the cloud at 6320. According to some embodiments, the request can be processed by a connected-world service instantiated in the cloud (e.g., 6354). In other embodiments, the functional operations for identifying candidate canonical maps can be separated from the operations for actual matching (e.g., shown as boxes 6354 and 6356). In one embodiment, a cloud service for map filtering and / or map ranking can be executed at 6354 and process the localization request received at 6320. According to some embodiments, the map ranking operation is configured to determine a set of candidate maps that may include the device location at 6322.

[0478] In one example, the map ranking function includes an operation of identifying candidate canonical maps based on geographical attributes or other location data (e.g., observed or inferred location information). For example, other location data can include Wi-Fi signatures or GPS information.

[0479] According to other embodiments, location data can be captured during a cross-reality session with the device and the user. Process 6300 can include additional operations to populate locations for a given device and / or session (not shown). For example, the location data can be stored as device area attribute values and attribute values for selecting candidate canonical maps close to the device location.

[0480] Any one or more location options can be used to filter a plurality of sets of canonical maps into a set of canonical maps that may represent the area including the user device location. In some embodiments, the canonical maps can cover a relatively large physical world area. The canonical maps can be segmented into multiple regions such that the selection of a map requires the selection of a map region. For example, the map regions can be on the order of dozens of square meters. Thus, the filtered set of canonical maps can be a set of map regions.

[0481] According to some embodiments, a localization snapshot can be constructed based on a candidate canonical map, pose features, and sensor calibration data. For example, an array of the candidate canonical map, pose features, and sensor calibration information can be sent along with a request to determine a specific matching canonical map. Matching with the canonical map can be performed based on descriptors received from the device and stored PCF data associated with the canonical map.

[0482] In some embodiments, a feature set from the device is compared with a feature set stored as part of the canonical map. The comparison can be based on feature descriptors and pose. For example, a candidate feature set of the canonical map can be selected based on multiple features in a candidate set, where descriptors of the multiple features are similar enough to descriptors of the feature set from the device, and these features can be the same features. For example, the candidate set can be features derived from image frames used to form the canonical map.

[0483] In some embodiments, if the number of similar features exceeds a threshold, further processing can be performed on the candidate feature set. The further processing can determine the degree to which the pose feature set from the device is aligned with the features in the candidate feature set. Features from the canonical map, such as features from the device, can be posed.

[0484] In some embodiments, the features are formatted into high-dimensional embeddings (e.g., DSF, etc.) and comparison can be performed using nearest neighbor search. In one example, the system is configured (e.g., by executing processes 6200 and / or 6300) to find the top two nearest neighbors using Euclidean distance, and a ratio test can be performed. If the nearest neighbor is closer than the second nearest neighbor, the system considers the nearest neighbor to be a match. For example, "closer" in this context can be determined based on the ratio of the Euclidean distance to the second nearest neighbor being greater than a threshold multiplied by the ratio of the Euclidean distance to the nearest neighbor. Once the features from the device are considered "matched" with the features in the canonical map, the system can be configured to calculate a relative transformation using the pose of the matched features. The transformation derived from the pose information can be used to indicate the transformation required to localize the device to the canonical map.

[0485] The number of correct data can be used as an indication of the matching quality. For example, in the case of DSF matching, the number of correct data reflects the number of features that match between the received descriptor information and the stored map / canonical map. In other embodiments, the correct data can be determined in this embodiment by calculating the number of "matched" features in each set.

[0486] Alternatively or additionally, an indication of the matching quality can be determined in other ways. In some embodiments, for example, when calculating a transformation to localize a map from a device containing multiple features to a canonical map based on the relative poses of the matching features, the transformation statistics calculated for each of the multiple matching features can be used as an indication of quality. For example, a large variance can indicate poor matching quality. Alternatively or additionally, for a determined transformation, the system can calculate the average error between features having matching descriptors. The average error of the transformation can be calculated to reflect the degree of position mismatch. Mean squared error is a specific example of an error metric. Regardless of the specific error metric, if the error is below a threshold, the transformation can be determined to be usable for the features received from the device, and the calculated transformation is used to localize the device. Alternatively or additionally, the number of correct data can also be used to determine whether there is a map that matches the device location information and / or descriptors received from the device.

[0487] As described above, in some embodiments, a device can send multiple feature sets for localization. When at least a threshold number of feature sets match the feature sets from the canonical map, the error is below a threshold, and / or the number of correct data is above a threshold, the localization can be considered successful. For example, the threshold number can be three feature sets. However, it should be understood that the threshold for determining whether a sufficient number of feature sets have a suitable value can be determined empirically or in other suitable ways. Similarly, other thresholds or parameters of the matching process, such as the similarity between feature descriptors considered to be a match, the number of correct data for selecting candidate feature sets, and / or the magnitude of the mismatch error, can be determined empirically or in other suitable ways.

[0488] Once a match is determined, a set of persistent map features associated with one or more canonical maps that match is identified. In embodiments where the match is based on a map region, the persistent map features can be the map features in the matching region. The persistent map features can be the persistent poses or PCFs as described above. In the example of FIG. 63, the persistent map features are persistent poses.

[0489] Regardless of the format of the persistent map features, each persistent map feature can have a predetermined orientation relative to the canonical map to which it belongs. This relative orientation is applied to the transformation that is calculated to align the feature set from the device with the feature set from the canonical map to determine the transformation between the feature set from the device and the persistent map feature. Any adjustments that can be derived from calibration data, for example, can then be applied to this calculated transformation. The resulting transformation can be the transformation between the device local coordinate system and the persistent map feature. This calculation can be performed for each persistent map feature of the matching map region, and the results can be stored in a table, which is represented at 6326 as the persistent_pose_table.

[0490] In one example, the box 6326 returns the persistent pose transformation table, the canonical map identifier, and the number of correct data. According to some embodiments, the canonical map ID is an identifier used to uniquely identify the canonical map and the canonical map version (or map region, in embodiments where localization is based on map regions).

[0491] In various embodiments, the computed localization data can be used to populate the localization statistics and telemetry maintained by the localization service at 6328. This information can be stored for each device and updated for each localization attempt, and can be cleared at the end of a device session. For example, which maps the device matches can be used to improve map ranking operations. For example, maps that cover the same area previously matched by the device can be given priority in the ranking. Similarly, maps that cover adjacent areas can have a higher priority than more distant areas. Additionally, adjacent maps can be prioritized based on the detected trajectory of the device over time, where map regions in the direction of movement are given a higher priority than other map regions. The localization service can use this information, for example, when subsequent localization requests from the device limit the search for candidate feature sets in stored canonical maps to maps or map regions. If a match with a low error metric and / or a large amount of correct data or a large percentage of correct data is identified in that limited area, processing of maps outside that area can be avoided.

[0492] Process 6300 continues to transmit the information cloud (e.g., 6354) to the user device (e.g., 6352). According to some embodiments, the persistent pose table and the canonical map identifier are transmitted to the user device at 6330. In one example, the persistent pose table can be composed of multiple elements, which include at least one string identifying the persistent pose ID and the transformation that links the device's tracking map to that persistent pose. In embodiments where the persistent map feature is a PCF, the table can instead indicate the transformation to the PCF of the matching map.

[0493] If positioning fails at 6336, process 6300 continues by adjusting parameters that can increase the amount of data sent from the device to the positioning service, thereby increasing the chance of successful positioning. For example, a failure can be indicated when a feature set with more than a threshold number of similar descriptors cannot be found in the canonical map, or when the error metric associated with all transformed candidate feature sets is above a threshold. As an example of an adjustable parameter, the size constraint of the descriptor buffer can be increased (6319). For example, in the case of a descriptor buffer size of 5, a positioning failure triggers an increase to at least six feature sets extracted from at least six image frames. In some embodiments, process 6300 may include a descriptor buffer increment value. In one example, the increment value can be used to control the rate of increase of the buffer size, for example, in response to a positioning failure. Other parameters, such as parameters that control the rate of positioning requests, can be changed when a matching canonical map cannot be found.

[0494] In some embodiments, execution 6300 may generate an error condition at 6340, which includes execution in the event that a positioning request fails to work instead of returning a mismatched result. For example, an error may occur when a network error causes the memory holding the canonical map database to be unavailable to a server performing the positioning service or a request for the positioning service is received containing incorrectly formatted information. In this example, in the event of an error condition, process 6300 schedules a retry of the request at 6342.

[0495] When the positioning request succeeds, any parameters adjusted in response to the failure may be reset. At 6332, process 6300 may continue to operate to reset the frequency parameters to any default values ​​or baselines. In some embodiments, 6332 is performed regardless of any changes, thereby ensuring that a baseline frequency is always established.

[0496] At 6334, the device may use the received information to update the cached positioning snapshot. According to various embodiments, the corresponding transformation, canonical map identifier, and other positioning data may be stored by the device and used to associate the position specified relative to the canonical map or its persistent map features (such as a persistent gesture or PCF) with the position determined by the device relative to its local coordinate system (such as may be determined from its tracking map).

[0497] Various embodiments of the process for positioning in the cloud can implement any one or more of the aforementioned steps and be based on the aforementioned architecture. Other embodiments can combine one or more of the aforementioned steps, perform the steps simultaneously, in parallel, or in another order.

[0498] According to some embodiments, the cloud-based positioning service in the context of cross-reality experiences may include additional features. For example, a canonical map cache may be executed to address connectivity issues. In some embodiments, the device may periodically download and cache the canonical maps it localizes to. If the cloud-based positioning service is unavailable, the device may perform localization on its own (e.g., as described above - including with respect to Figure 26 as described). In other embodiments, the transforms returned from a positioning request may be chained together and applied to subsequent sessions. For example, the device may cache a series of transforms and use the sequence of transforms to establish localization.

[0499] Various embodiments of the system may use the positioning operation results to update the transform information. For example, the positioning service and / or the device may be configured to maintain state information regarding the tracking map to canonical map transform. The transforms received over a period of time may be averaged. According to some embodiments, the averaging operation may be restricted to occur after a threshold number of positioning successes (e.g., three, four, five, or more). In other embodiments, other state information may be tracked in the cloud, such as via the Connected World module. In one example, the state information may include device identifiers, tracking map IDs, canonical map references (e.g., version and ID), and the canonical map to tracking map transform. In some examples, by performing the cloud-based positioning function each time, the system may use the state information to continuously update and obtain a more accurate canonical map to tracking map transform.

[0500] Addit...

Claims

1. A network resource within a distributed computing environment for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a 3D environment, the network resource comprises: one or more processors; at least one computer-readable medium comprising: a plurality of stored maps of the 3D environment; a plurality of data structures, each data structure of the plurality of data structures representing a corresponding area in the 3D environment in which virtual content is to be displayed, wherein each data structure of the plurality of data structures comprises: information associating the data structure with a location in the plurality of stored maps; and a link to virtual content for rendering in the corresponding area in the 3D environment; and computer-executable instructions that, when executed by a processor of the one or more processors: implement a service for providing location information to a portable electronic device of the plurality of portable electronic devices, wherein the location information indicates the location of the plurality of portable electronic devices relative to one or more shared maps; and selectively provide a copy of at least one data structure of the plurality of data structures to the portable electronic device based on the location of the portable electronic device relative to the areas represented by the plurality of data structures; detect that the portable electronic device has left an area represented by a data structure of the one or more data structures; and delete the virtual content represented in the data structure based on the detection.

2. The network resource according to claim 1, wherein: the computer-executable instructions further implement an authentication service for determining the access rights of the portable electronic device when executed by the processor; and the computer-executable instructions for selectively providing the at least one data structure to the portable electronic device determine whether to send the at least one data structure based in part on the access rights of the portable electronic device and access attributes associated with the at least one data structure.

3. The network resource according to claim 1, wherein: each data structure of the plurality of data structures further comprises common attributes; and the computer-executable instructions for selectively providing the at least one data structure to the portable electronic device determine whether to send the at least one data structure based in part on the common attributes of the at least one data structure.

4. The network resource according to claim 1, wherein: for a portion of the plurality of data structures, the link to the virtual content comprises a link to an application that provides the virtual content.

5. The network resource according to claim 1, wherein: each data structure of the plurality of data structures further comprises display characteristics of a prism on the portable electronic device, wherein the prism is a volume in which the virtual content linked to the data structure is displayed.

6. The network resource according to claim 5, wherein: the display characteristics include the size of the prism.

7. The network resource according to claim 5, wherein: The display characteristics include the behavior of virtual content rendered within the prism relative to a physical surface.

8. The network resource according to claim 5, wherein: The display characteristics include one or more of the following: the offset of the prism relative to a persistent location associated with a map, the spatial orientation of the prism, the behavior of virtual content rendered within the prism relative to the position of the portable electronic device, and the behavior of virtual content rendered within the prism relative to the direction the portable electronic device is facing.

9. The network resource according to claim 1, wherein, The at least one data structure is associated with the prism such that corresponding virtual content is rendered within the prism.

10. A method of operating a portable electronic device to render virtual content in a 3D environment, the method comprising using one or more processors: Generate a local coordinate system on the portable electronic device based on the output of one or more sensors on the portable electronic device; Generate information indicating a position in the 3D environment on the portable electronic device based on the output of the one or more sensors and an indication of a position in the local coordinate system; Send, over a network, the information indicating the position and the indication of the position in the local coordinate system to a location service; Obtain a transformation between a coordinate system of the stored spatial information about the 3D environment from the location service and the local coordinate system; Obtain from the location service one or more data structures, each data structure representing a corresponding region in the 3D environment and virtual content for display in the corresponding region; and Render the virtual content represented in the one or more data structures in the corresponding regions of the one or more data structures; Detect that the portable electronic device has left a region represented by a data structure among the one or more data structures; and Based on the detection, delete the virtual content represented in the data structure.

11. The method according to claim 10, wherein: Rendering virtual content in the corresponding region includes: creating a prism with a set of parameters based on the data structure representing the corresponding region.

12. The method according to claim 10, wherein, The virtual content is represented as an indicator of the location of the virtual content on the network in at least one of the one or more data structures.

13. The method according to claim 10, wherein: Rendering the virtual content includes: executing an application that generates the virtual content on the portable electronic device.

14. The method according to claim 13, wherein, Rendering the virtual content further includes: Determining whether the application is currently installed on the portable electronic device; and Based on determining that the application is not currently installed, downloading the application to the portable electronic device.

15. The method according to claim 10, wherein: The one or more received data structures include a first set of data structures; The first set of data structures is received at a first time; The method further includes: Store rendering information associated with a first data structure among the first set of data structures; Receive a second set of data structures at a second time after the first time; Based on determining that the first data structure is not included in the second set, delete the rendering information associated with the first data structure.

16. The method according to claim 10, further comprising: Obtain corresponding prisms associated with the one or more data structures such that corresponding virtual content is rendered within the corresponding prisms, and wherein rendering the virtual content represented in the one or more data structures comprises: rendering the virtual content within the corresponding prisms.

17. An electronic device configured to operate within a cross-reality system, the electronic device comprising: One or more sensors configured to capture information about a three-dimensional (3D) environment, the captured information including a plurality of images; At least one processor; At least one computer-readable medium storing computer-executable instructions that, when executed on a processor in the at least one processor: Maintain a local coordinate system for representing positions in the 3D environment based on at least a first portion of the plurality of images; Manage prisms associated with one or more applications such that virtual content generated by an application in the one or more applications is rendered within the prisms; Send information derived from outputs of the one or more sensors to a service via a network; Receive from the service: Location information; and Data structures representing corresponding virtual content and regions in the 3D environment for rendering the virtual content; and Create a prism associated with the data structure such that the corresponding virtual content is rendered within the prism, wherein the computer-executable instructions further include instructions for performing the following operations: Detect that the electronic device has left the region represented by the data structure; and Based on the detection, delete the prism associated with the data structure.

18. The electronic device according to claim 17, wherein, The computer-executable instructions further include computer-executable instructions for performing the following operations: Obtain the corresponding virtual content based on information in the data structure; and Render the obtained virtual content within the prism.

19. The electronic device according to claim 18, wherein, Obtaining the corresponding virtual content based on information in the data structure includes: accessing the corresponding virtual content via a network based on an indicator of a virtual content position in the data structure.

20. The electronic device according to claim 19, wherein, Obtaining the corresponding virtual content based on information in the data structure includes: downloading an application that generates the corresponding virtual content via a network based on the indicator of the virtual content position in the data structure.

21. The electronic device according to claim 18, wherein, Rendering the obtained virtual content within the prism further includes: using the coordinate system of the electronic device to determine a set of coordinates for rendering the virtual content in the 3D environment.

22. The electronic device according to claim 18, wherein, obtaining the corresponding virtual content based on the information in the data structure includes: based on an indicator of the virtual content position in the data structure, downloading an application for generating the corresponding virtual content through a network, and wherein rendering the obtained virtual content within the prism further includes: using the coordinate system of the electronic device to determine a set of coordinates for rendering the virtual content in the 3D environment.

23. A method for planning location-based virtual content for a cross-reality system operable with a plurality of portable electronic devices capable of rendering virtual content in a 3D environment, the method comprising using one or more processors: generating a data structure representing an area in the 3D environment in which virtual content is to be displayed; storing in the data structure information indicating the virtual content to be rendered in the area in the 3D environment ; and associating the data structure with a location in a map for positioning the plurality of portable electronic devices in a shared coordinate system; detecting that the portable electronic device has left the area represented by the data structure in the one or more data structures; and based on the detection, deleting the virtual content represented in the data structure.

24. The method according to claim 23, further comprising: setting access rights to the data structure.

25. The method according to claim 24, wherein, setting access rights to the data structure includes: indicating that the data structure can be accessed by one or more specific categories of users of the plurality of portable electronic devices.

26. The method according to claim 23, wherein, storing in the data structure information indicating the virtual content to be rendered in the area in the 3D environment includes: specifying an application executable on the portable electronic device to generate the virtual content.

27. The method according to claim 23, further comprising: storing the data structure in association with a positioning service for positioning the plurality of portable electronic devices using the map.

28. The method according to claim 23, further comprising: receiving a specification of the area in the 3D environment and the virtual content through a user interface.

29. The method according to claim 23, further comprising: receiving a specification of the area in the 3D environment and the virtual content from an application through a programming interface.

Citation Information

Patent Citations

  • Localization determination for mixed reality systems

    US10812936B2

  • System and method for augmented and virtual reality

    US20140306866A1

  • Centralized rendering

    US20180286116A1

  • Matching content to a spatial 3D environment

    US20180315248A1

  • Fully convolutional interest point detection and description via homographic adaptation

    US20190147341A1