Cross-reality system with precise shared map

CN119984235APending Publication Date: 2025-05-13MAGIC LEAP INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411977117.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-02-13
Filing Date
2021-02-11
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing cross-reality systems have difficulty accurately combining environmental maps with sensor data collected by user equipment, resulting in inaccurate positioning of virtual content in the physical world.

Method used

By receiving the tracking map of the user equipment, determine the transformation between it and the environment map, and determine whether to merge the environment map with the tracking map to ensure that the merged map is aligned with the gravity direction.

Benefits of technology

The accurate merging of environmental maps and user equipment sensor data is achieved, the position accuracy of virtual content in the physical world is improved, and the immersion and user experience of the cross-reality system is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119984235A_ABST
    Figure CN119984235A_ABST
Patent Text Reader

Abstract

A cross-reality system enables any of a plurality of devices to efficiently and accurately access previous persistent maps of a very large scale environment and render virtual content related to those maps. A cross-reality system may construct a persistent map, which may be in a canonical form, by merging tracking maps from multiple devices. The map merging process determines a mergability of the tracking map with the specification map and merges the tracking map with the specification map according to a mergability criterion, e.g., when a gravity direction of the tracking map is aligned with a gravity direction of the specification map. If the orientation of the tracking map relative to gravity is not preserved, merging of the map is avoided to avoid distortions in persistent maps and results in a more realistic and immersive experience to their users that can use the map to determine their location.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese patent application “Cross-reality system with accurate shared maps” with application number 202180027861.7 (filing date February 11, 2021).

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Patent Application Serial No. 62 / 975,983, filed on February 13, 2020, and entitled “CROSS REALITY SYSTEM WITH ACCURATE SHARED MAPS,” which is incorporated herein by reference in its entirety. Technical Field

[0004] The present application relates generally to cross-reality systems. Background Art

[0005] A computer can control a human user interface to create a cross-reality (XR) environment in which some or all of the XR environment perceived by the user is generated by the computer. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments, some or all of which can be generated in part by a computer using data describing the environment. For example, the data can describe a virtual object that can be rendered in a way that the user feels or perceives as part of the physical world, and can interact with the virtual object. Because the data is rendered and presented by a user interface device (such as, for example, a head-mounted display device), the user can experience these virtual objects. The data can be displayed to the user, or can control audio that is played to the user, or can control a tactile (or haptic) interface, allowing the user to experience the touch sensation that the user feels or perceives as feeling the virtual object.

[0006] XR systems can be used for many applications across scientific visualization, medical training, engineering design and prototyping, telemanipulation and telepresence, and personal entertainment. Compared to VR, AR and MR include one or more virtual objects related to real objects in the physical world. The experience of virtual objects interacting with real objects significantly enhances the user's enjoyment of using XR systems, and also opens the door to a variety of applications that present realistic and easy-to-understand information about how to change the physical world.

[0007] To render virtual content realistically, an XR system may build a representation of the physical world surrounding a user of the system. For example, this representation may be built by processing images acquired using sensors on a wearable device, where the wearable device forms part of the XR system. In such a system, a user may perform an initialization routine by looking around a room or other physical environment in which the user intends to use the XR system until the system obtains enough information to build a representation of that environment. As the system runs and the user moves through the environment or to other environments, sensors on the wearable device may acquire additional information to expand or update the representation of the physical world. Summary of the invention

[0008] Aspects of the present application relate to methods and devices for providing a cross-reality (XR) scene. The techniques described herein may be used together, individually, or in any suitable combination.

[0009] According to one embodiment, a method is provided for merging one or more environment maps stored in a database with a tracking map calculated based on sensor data collected by a device worn by a user. The method may include: receiving the tracking map from the device, wherein the tracking map is aligned with respect to a gravity direction; determining a transformation between the tracking map and the environment map; determining whether to merge the environment map with the tracking map, wherein determining whether to merge includes: determining whether applying the transformation to the tracking map produces a transformed tracking map aligned with respect to the gravity direction; and merging the environment map with the tracking map based on determining that the transformed tracking map is aligned with the gravity direction.

[0010] According to one embodiment, determining the transformation may include: for corresponding features in the tracking map and the environment map, selecting a transformation that is aligned with the corresponding feature with an error metric below a threshold as the determined transformation.

[0011] According to one embodiment, the method may further include determining the corresponding features based on similarities of identifiers assigned to the features.

[0012] According to one embodiment, selecting as the determined transformation may further include: selecting a transformation that, when applied, results in the transformed tracking map being aligned with respect to the gravity direction.

[0013] According to one embodiment, determining the transformation may include: applying a plurality of candidate transformations to the tracking map, and selecting a candidate transformation from among the plurality of candidate transformations as the determined transformation.

[0014] According to one embodiment, selecting the determined transformation may further include: selecting the candidate transformation that, when applied, produces the transformed tracking map aligned with respect to the gravity direction as the determined transformation.

[0015] According to one embodiment, determining whether applying the transformation to the tracking map produces the transformed tracking map aligned with respect to the gravity direction includes: determining whether applying the transformation to the tracking map produces the transformed tracking map rotated by more than a threshold amount relative to the gravity direction; in response to determining that applying the transformation to the tracking map does not produce the transformed tracking map rotated by more than the threshold amount relative to the gravity direction, selecting the transformation as the determined transformation; and in response to determining that applying the transformation to the tracking map produces the transformed tracking map rotated by more than the threshold amount relative to the gravity direction, discarding the transformation.

[0016] According to one embodiment, the method may further include: identifying a set of environmental maps to be merged with the tracking map from the database; and for each environmental map in the set of environmental maps: determining the transformation between the tracking map and the environmental map; determining whether to merge the environmental map with the tracking map; and merging the environmental map with the tracking map based on determining that the transformed tracking map is aligned with the gravity direction.

[0017] According to one embodiment, the method may further include: for each environment map in the set of environment maps, avoiding merging the environment map with the tracking map based on determining that a gravity direction of the environment map is not aligned with a gravity direction of the transformed tracking map.

[0018] According to one embodiment, identifying the set of environment maps may include: determining a region identifier associated with the tracking map; and identifying the set of environment maps from the database based at least in part on the region identifier associated with the tracking map.

[0019] According to one embodiment, identifying the set of environment maps from the database may further include: identifying the set of environment maps based on the environment map associated with the tracking map and in the set of environment maps. Figure 1 The environment map set is filtered based on the similarity of one or more metrics.

[0020] According to one embodiment, a computing device configured for use in a cross-reality system in which a portable device operating in a three-dimensional (3D) environment renders virtual content. The computing device may include: at least one processor; a computer-readable medium connected to the processor; a plurality of environment maps stored in the computer-readable medium; and computer-executable instructions configured to perform a method when executed by the at least one processor. The method performed by the computer-executable instructions may include: receiving a tracking map from the portable device, wherein the tracking map is aligned with respect to a gravity direction; determining whether to merge an environment map with the tracking map, wherein determining whether to merge includes: searching for a transformation of the tracking map, the transformation aligning the transformed tracking map and the environment map in a manner that maintains the alignment of the transformed tracking map with respect to the gravity direction; and merging the environment map with the transformed tracking map based on determining that the gravity direction of the environment map is aligned with the center of gravity direction of the transformed tracking map.

[0021] According to one embodiment, searching for a transformation may include searching for a transformation that aligns a first set of features associated with the tracking map and a second set of features associated with the environment map with an error metric below a threshold.

[0022] According to one embodiment, searching for a transformation may include searching for a transformation that does not change the orientation of the transformed tracking map relative to the gravity direction.

[0023] According to one embodiment, searching for a transformation may include: determining whether applying the transformation to the tracking map produces the transformed tracking map rotated by more than a threshold amount relative to the gravity direction; in response to determining that applying the transformation to the tracking map does not produce the transformed tracking map rotated by more than the threshold amount relative to the gravity direction, applying the transformation to the tracking map to generate the transformed tracking map; and in response to determining that applying the transformation to the tracking map produces the transformed tracking map rotated by more than the threshold amount relative to the gravity direction, discarding the transformation.

[0024] According to one embodiment, the method may further include: identifying a set of environmental maps to be merged with the tracking map from the multiple environmental maps; and for each environmental map in the set of environmental maps: determining whether to merge the environmental map with the tracking map; and merging the environmental map with the transformed tracking map based on determining that the gravity direction of the environmental map is aligned with the gravity direction of the transformed tracking map.

[0025] According to one embodiment, the method may further include: for each environment map in the set of environment maps, based on determining that a gravity direction of the environment map is not aligned with a gravity direction of the tracking map, avoiding merging the environment map with the transformed tracking map.

[0026] According to one embodiment, identifying the set of environment maps may further include: determining a region identifier associated with the tracking map; and identifying the set of environment maps based at least in part on the region identifier associated with the tracking map.

[0027] According to one embodiment, identifying the set of environment maps may further include filtering the set of environment maps based on similarity of one or more metrics associated with the tracking map and the environment maps in the set of environment maps.

[0028] According to one embodiment, a cloud computing environment for an augmented reality system is configured to communicate with a plurality of user devices including sensors. The cloud computing environment for an augmented reality system may include: a map database storing a plurality of environment maps constructed from data provided by the plurality of user devices; and a non-transitory computer storage medium storing computer executable instructions, wherein the computer executable instructions perform a method when executed by at least one processor in the cloud computing environment. The method may include: receiving a tracking map from a user device, wherein the tracking map is aligned with respect to a gravity direction; updating the map database based on the received tracking map, wherein updating the map database includes: for each environment map in a set of environment maps, determining a transformation between the tracking map and the environment map; determining whether to merge the environment map with the tracking map, wherein determining whether to merge includes: determining whether applying the transformation to the tracking map produces a transformed tracking map aligned with respect to the gravity direction; and merging the environment map with the tracking map based on determining that the transformed tracking map is aligned with the gravity direction.

[0029] According to one embodiment, determining the transformation may include: for corresponding features in the tracking map and the environment map, selecting as the determined transformation a transformation that aligns the corresponding features with an error metric below a threshold.

[0030] According to an embodiment, the method may further include determining the corresponding features based on similarities of identifiers assigned to the features.

[0031] According to one embodiment, selecting as the determined transformation may further include: selecting a transformation that, when applied, results in the transformed tracking map being aligned with respect to the gravity direction.

[0032] According to one embodiment, determining the transformation may include: applying a plurality of candidate transformations to the tracking map, and selecting a candidate transformation from among the plurality of candidate transformations as the determined transformation.

[0033] According to one embodiment, selecting as the determined transformation may further include: selecting as the determined transformation the candidate transformation that, when applied, produces the transformed tracking map aligned with respect to the gravity direction.

[0034] According to one embodiment, determining whether applying the transformation to the tracking map produces the transformed tracking map aligned with respect to the gravity direction may include: determining whether applying the transformation to the tracking map produces the transformed tracking map rotated by more than a threshold amount relative to the gravity direction; in response to determining that applying the transformation to the tracking map does not produce the transformed tracking map rotated by more than the threshold amount relative to the gravity direction, selecting the transformation as the determined transformation; and in response to determining that applying the transformation to the tracking map produces the transformed tracking map rotated by more than the threshold amount relative to the gravity direction, discarding the transformation.

[0035] The foregoing summary is provided by way of illustration and is not intended to be limiting. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in various figures is represented by a like numeral. For clarity, not every component may be labeled in every figure. In the drawings:

[0037] Figure 1 is a schematic diagram illustrating an example of a simplified augmented reality (AR) scene according to some embodiments;

[0038] Figure 2 is a schematic diagram of an exemplary simplified AR scene, illustrating an exemplary use case of an XR system, according to some embodiments;

[0039] Figure 3 is a schematic diagram illustrating data flow for a single user in an AR system configured to provide the user with an experience of AR content that interacts with the physical world, according to some embodiments;

[0040] Figure 4 is a schematic diagram illustrating an exemplary AR display system that displays virtual content for a single user according to some embodiments;

[0041] Figure 5Ais a schematic diagram showing an AR display system rendering AR content as the user moves through a physical world environment when the user is wearing the AR display system according to some embodiments;

[0042] Figure 5B is a schematic diagram illustrating a viewing optics assembly and accompanying components according to some embodiments;

[0043] Fig. 6A is a schematic diagram illustrating an AR system using a world reconstruction system according to some embodiments;

[0044] Figure 6B is a schematic diagram illustrating components of an AR system that maintains a model of a navigable world according to some embodiments;

[0045] Figure 7 A schematic diagram of a tracking map formed by the path traversed by a device through the physical world.

[0046] Figure 8 is a schematic diagram illustrating a user of a cross reality (XR) system perceiving virtual content according to some embodiments;

[0047] Fig. 9 is a method for transforming between coordinate systems according to some embodiments Figure 8 A block diagram of components of a first XR device of an XR system;

[0048] Fig.10 is a schematic diagram illustrating an exemplary transformation of an origin coordinate frame into a destination coordinate frame in order to correctly render local XR content according to some embodiments;

[0049] Fig.11 is a top plan view illustrating a pupil-based coordinate framework according to some embodiments;

[0050] Fig.12 is a top plan view showing a camera coordinate frame including all pupil locations according to some embodiments;

[0051] Fig.13 According to some embodiments Fig. 9 A schematic diagram of a display system;

[0052] Fig.14 is a block diagram illustrating the creation of a persistent coordinate system (PCF) and the attachment of XR content to the PCF according to some embodiments;

[0053] Fig.15 is a flow chart illustrating a method of establishing and using a PCF according to some embodiments;

[0054] Fig.16 is according to some embodiments including a second XR device Figure 8 Block diagram of the XR system;

[0055] Fig.17 is a schematic diagram showing a room and key frames established for various areas in the room according to some embodiments;

[0056] Fig.18 is a schematic diagram illustrating the establishment of a keyframe-based persistent gesture according to some embodiments;

[0057] Fig.19 is a schematic diagram illustrating the establishment of a persistent coordinate system (PCF) based on a persistent gesture according to some embodiments;

[0058] FIG. 20A to FIG. 20C is a schematic diagram illustrating an example of creating a PCF according to some embodiments;

[0059] Fig.21 is a block diagram illustrating a system for generating a global descriptor for a single image and / or map according to some embodiments;

[0060] Fig. 22 is a flow chart illustrating a method of computing an image descriptor according to some embodiments;

[0061] Fig.23 is a flow chart illustrating a localization method using image descriptors according to some embodiments;

[0062] Fig.24 is a flow chart illustrating a method of training a neural network according to some embodiments;

[0063] Fig.25 is a block diagram illustrating a method of training a neural network according to some embodiments;

[0064] Fig.26 is a schematic diagram illustrating an AR system configured to rank and merge multiple environment maps according to some embodiments;

[0065] Fig. 27 is a simplified block diagram illustrating a plurality of canonical maps stored on a remote storage medium according to some embodiments;

[0066] Fig.28 is a schematic diagram illustrating a method of selecting a canonical map, for example, to locate a new tracking map in one or more canonical maps and / or to obtain a PCF from a canonical map, according to some embodiments;

[0067] Fig.29 is a flow chart illustrating a method of selecting a plurality of ranked environment maps according to some embodiments;

[0068] Fig.30is a diagram showing some embodiments of the present invention. Fig.26 A schematic diagram of an exemplary map ranking portion of an AR system;

[0069] Fig.31A is a schematic diagram showing an example of area attributes of a tracking map (TM) and an environment map in a database according to some embodiments;

[0070] Fig.31B is a diagram showing the determination of a Fig.29 A schematic diagram of an example of a geo-location filtered Tracking Map(TM);

[0071] Fig.32 is a diagram showing some embodiments of the present invention. Fig.29 A schematic diagram of an example of geographic location filtering;

[0072] Fig.33 is a diagram showing some embodiments of the present invention. Fig.29 A schematic diagram of an example of Wi-Fi BSSID filtering;

[0073] Fig.34 is a diagram showing the use of some embodiments Fig.29 A schematic diagram of an example of positioning;

[0074] Fig.35 and 36 is a block diagram of an XR system configured to rank and merge multiple environment maps, according to some embodiments.

[0075] Fig.37 is a block diagram illustrating a method of creating an environment map of the physical world in a canonical form according to some embodiments;

[0076] Fig.38A and 38B is shown in accordance with some embodiments by updating the Figure 7 A schematic diagram of an environment map created in a canonical form.

[0077] Figures 39A to 39F is a schematic diagram illustrating an example of a merged map according to some embodiments;

[0078] Fig.40 According to some embodiments, Fig. 9 The first XR device to generate the first 3D local tracking map (ground Figure 1 ) two-dimensional representation;

[0079] Fig.41 is an illustration of a flow from a first XR device to a Fig. 9 Server upload location Figure 1 Block diagram of

[0080] Fig.42 is a diagram showing some embodiments of the present invention. Fig.16 A schematic diagram of an XR system showing that a second user has initiated a second session using a second XR device of the XR system after the first user has terminated the first session;

[0081] Fig.43A is a diagram showing a method for Fig.42 A block diagram of a new session for a second XR device;

[0082] Fig.43B is a diagram showing a method for Fig.42 A block diagram of creation of a tracking map for a second XR device;

[0083] Fig.43C is a diagram showing a method for transmitting data from a server to a server according to some embodiments. Fig.42 A block diagram of a second XR device downloading a canonical map;

[0084] Fig.44 is to show that according to some embodiments, Fig.42 A second tracking map (map) generated by a second XR device Figure 2 ) A schematic diagram of a positioning attempt to locate the canonical map;

[0085] Fig.45 is a diagram showing the Fig.44 Second tracking map (map Figure 2 ) is a schematic diagram of a positioning attempt to locate to a canonical map, the second tracking map can be further developed and has Figure 2 The XR content associated with the PCF;

[0086] FIG. 46A to FIG. 46B is a diagram showing the Fig.45 land Figure 2 Schematic diagram of successful positioning on the normative map;

[0087] Fig.47 is a diagram showing that according to some embodiments, Fig.46A The canonical map of one or more PCFs includes Fig.45 land Figure 2 Schematic diagram of the normative map generated by

[0088] Fig.48 is a diagram showing some embodiments of the present invention. Fig.47 The canonical map and the map on the second XR device Figure 2 A further extended schematic diagram of

[0089] Fig.49FIG. 2 is a diagram illustrating uploading a map from a second XR device to a server according to some embodiments. Figure 2 Block diagram of

[0090] Fig.50 is a diagram showing how to convert the ground Figure 2 box plot merged with normative map;

[0091] Fig.51 is a block diagram illustrating transmission of a new canonical map from a server to a first XR device and a second XR device according to some embodiments;

[0092] Fig.52 is a diagram showing a method according to some embodiments Figure 2 Two-dimensional representation and reference ground Figure 2 A block diagram of a head coordinate frame of a second XR device;

[0093] Fig.53 is a block diagram illustrating in two dimensions adjustments of a head coordinate frame that may occur in six degrees of freedom according to some embodiments;

[0094] Fig.54 is a block diagram illustrating a canonical map on a second XR device according to some embodiments, wherein sound is relative to ground Figure 2 The PCF is located;

[0095] Fig.55 and Fig.56 are perspective and block diagrams illustrating use of an XR system when a first user has terminated a first session and the first user has initiated a second session using the XR system in accordance with some embodiments;

[0096] Fig.57 and Fig.58 are perspective views and block diagrams illustrating use of an XR system when three users are using the XR system simultaneously in the same session, according to some embodiments;

[0097] Fig.59 is a flow chart illustrating a method of recovering and resetting head posture according to some embodiments;

[0098] Fig.60 is a block diagram of a machine in the form of a computer that may find application in the system of the present invention according to some embodiments;

[0099] Fig.61 is a schematic diagram of an example XR system in which any of a plurality of devices may access a location service, according to some embodiments;

[0100] Fig.62is an example process flow for operating a portable device as part of an XR system providing cloud-based positioning according to some embodiments;

[0101] Fig.63A , Fig.63B and Fig.63C is an example process flow for cloud-based positioning according to some embodiments;

[0102] Fig.64 is a block diagram of an XR system providing large-scale positioning according to some embodiments;

[0103] Fig.65 is a diagram showing the Fig.64 Schematic diagram of the information of the physical world processed by the XR system;

[0104] Fig.66 According to some embodiments Fig.64 A block diagram of the subsystems of an XR system including a correspondence quality prediction component and a pose estimation component;

[0105] Fig.67 is a diagram showing the generation of a training Fig.66 A flowchart of a method for generating a data set of a subsystem;

[0106] Fig.68 is a diagram showing training according to some embodiments Fig.66 A flowchart of a method for a subsystem of the invention; and

[0107] Fig.69 A gravity-preserving map merging process is shown in accordance with some embodiments. DETAILED DESCRIPTION

[0108] Methods and apparatus for providing XR scenes are described herein. In order to provide realistic XR experiences to multiple users, the XR system must know the users' locations within the physical world in order to correctly associate the locations of virtual objects with real objects. The inventors have recognized and appreciated methods and apparatus for locating XR devices in large-scale and ultra-large-scale environments (e.g., neighborhood, city, country, global) with reduced time and improved accuracy.

[0109] The XR system can build an environmental map of the scene, which can be created based on images and / or depth information collected by sensors that are part of an XR device worn by a user of the XR system. Each XR device can develop a local map of its physical environment by integrating information from one or more images collected while the device is running. In some embodiments, the coordinate system of the map is associated with the orientation of the device when the device begins scanning the physical world. As the user interacts with the XR system, this orientation may change from session to session, whether different sessions are associated with different users, each user has their own wearable device with sensors that scan the environment, or the same user uses the same device at different times.

[0110] The XR system may implement one or more techniques to enable operations based on persistent spatial information. For example, these techniques may provide a more computationally efficient and immersive XR scene for a single or multiple users by allowing any of multiple users of the XR system to create, store, and retrieve persistent spatial information. The persistent spatial information may also enable rapid recovery and reset of head poses on each of one or more XR devices in a computationally efficient manner.

[0111] Persistent spatial information can be represented by a persistent map. The persistent map can be stored in a remote storage medium (e.g., the cloud). For example, after being turned on, a wearable device worn by a user can retrieve a previously created and stored appropriate map from a persistent storage such as cloud storage. The previously stored map may be based on data about the environment collected by sensors on the user's wearable device during a previous session. Retrieving the stored map can enable the use of the wearable device without having to complete a scan of the physical world with sensors on the wearable device. Alternatively or additionally, the system / device can similarly retrieve an appropriate stored map when entering a new area of ​​the physical world.

[0112] The stored map may be represented in a canonical form to which a local reference frame on each XR device may be related. In a multi-device XR system, a stored map accessed by one device may have been created and stored by another device and / or may have been constructed by aggregating data about the physical world collected by sensors on multiple wearable devices that previously existed in at least a portion of the physical world represented by the stored map.

[0113] In some embodiments, persistent spatial information may be represented in a manner that can be easily shared between users and between distributed components including applications. A canonical map may provide information about the physical world, for example, as a persistent coordinate system (PCF). The PCF may be defined based on a set of features identified in the physical world. Features may be selected so that they may be the same between user sessions of the XR system. The PCF may exist sparsely, providing less available information about the physical world than all available information so that they can be efficiently processed and transmitted. Techniques for processing persistent spatial information may include creating dynamic maps based on the local coordinate system of one or more devices across one or more sessions. These maps may be sparse maps that represent the physical world based on a subset of feature points detected in an image used to form a map. The persistent coordinate system (PCF) may be generated from a sparse map and may be exposed to an XR application, for example, through an application programming interface (API). These capabilities may be supported by techniques for forming canonical maps by merging multiple maps created by one or more XR devices.

[0114] The relationship between each device's local map and the canonical map may be determined through a positioning process. The positioning process may be performed on each XR device based on a set of canonical maps selected and sent to the device. Alternatively or additionally, a positioning service may be provided on a remote processor, such as a positioning service that may be implemented in the cloud.

[0115] Sharing data about the physical world between multiple devices can enable a shared user experience of virtual content. For example, two XR devices accessing the same stored map can both be positioned relative to the stored map. Once positioned, the user device can render virtual content with the location specified by referencing the stored map into a reference frame maintained by the user device. The user device can use this local reference frame to control the display of the user device to render virtual content in the specified location.

[0116] To support these and other functions, the XR system may include components that develop, maintain, and use persistent spatial information (including one or more stored maps) based on data about the physical world collected from sensors on the user device. These components may be distributed across the XR system, for example, with some operating on a head-mounted portion of the user device. Other components may operate on a computer associated with a user coupled to the head-mounted portion via a local area network or a personal area network. Still others may operate at a remote location, such as at one or more servers accessible via a wide area network.

[0117] For example, these components may include components that are capable of identifying information of sufficient quality to be stored as a persistent map or stored in a persistent map from information about the physical world collected by one or more user devices. An example of such a component, described in more detail below, is a map merging component. For example, such a component may receive input from a user device and determine a suitable portion of the input to be used to update a persistent map, which may be in a canonical form. For example, the map merging component may also promote a local map from a user device that is not merged with a persistent map to a separate persistent map.

[0118] As another example, these components can include components that can help select an appropriate set of one or more persistent maps, which may represent the same area of ​​the physical world, as represented by the location information provided by the user device. Examples of such components, described in more detail below, are map ranking and map selection components. For example, such a component can receive input from a user device and identify one or more persistent maps that may represent the area in the physical world in which the device is operating. For example, a map ranking component can help select a persistent map to be used by a local device when rendering virtual content, collecting data about the environment, or performing other actions. Alternatively or additionally, a map ranking component can help identify persistent maps to be updated as additional information about the physical world is collected by one or more user devices.

[0119] Other components can determine a transformation that transforms information captured or described with respect to one reference frame to another reference frame. For example, sensors may be attached to a head-mounted display so that data read from these sensors indicates the location of objects in the physical world relative to the wearer's head posture. One or more transformations may be applied to relate this position information to a coordinate system associated with a persistent environment map. Similarly, data indicating where a virtual object will be rendered when represented in the coordinate system of a persistent environment map may be transformed one or more to be located in the reference frame of a display on the user's head. As described in more detail below, there may be multiple such transformations. These transformations may be partitioned across components of the XR system so that they can be efficiently updated and / or applied to a distributed system.

[0120] In some embodiments, a persistent map may be constructed based on information collected by multiple user devices. The XR devices may each capture local spatial information and construct a separate tracking map using information collected by sensors of each XR device at different locations and times. Each tracking map may include points, each of which may be associated with a feature of a real object that may include multiple features. In addition to potentially providing input for creating and maintaining persistent maps, the tracking map may also be used to track the movement of a user in a scene, enabling the XR system to estimate the head pose of the corresponding user relative to the reference frame established by the tracking map on the user device.

[0121] This interdependence between the creation of the map and the estimation of the head pose poses a significant challenge. A lot of processing may be required to simultaneously create the map and estimate the head pose. When objects move in the scene (for example, moving a cup on a table) and when the user moves in the scene, the processing must be done quickly because latency makes the XR experience less realistic for the user. On the other hand, XR devices can offer limited computing resources because XR devices should be lightweight so that users can wear them comfortably. Using more sensors cannot make up for the lack of computing resources because adding sensors also increases weight. In addition, more sensors or more computing resources will cause heat, which may cause deformation of the XR device.

[0122] XR systems may be configured to create, share, and use persistent spatial information with low computational resource usage and / or low latency to provide a more immersive user experience. Some such techniques may enable efficient comparison of spatial information. For example, such comparison may occur as part of localization, where a set of features from a local device is matched to a set of features in a canonical map.

[0123] Similarly, during a map merging process, an attempt can be made to match one or more sets of features in a tracking map from the device with corresponding features in a canonical map, and to determine a transformation between the sets of corresponding features that provides a suitably low error between the positions of the transformed features in a first set of features derived from the tracking map and a second set of features derived from the canonical map. Subsequent processing to merge the tracking maps into a set of canonical maps can be based on the results of this comparison. For example, determining a transformation with a suitably low error can indicate that the area represented by the second set of features derived from the canonical map corresponds to the same area represented by the first set of features derived from the tracking map, and the two maps can be merged.

[0124] The inventors have recognized that, despite this, even when there is a low error in the alignment of the set of features, errors may be introduced in the merging process. The inventors have further recognized and appreciated that such errors may be detected by a transformation that changes the orientation of the tracking map relative to the direction of gravity, and that merging such a tracking map with the canonical map may result in a skewed merged map.

[0125] By inhibiting the merging of transformed tracking maps that have changed orientations relative to gravity, the canonical map can start and remain aligned relative to gravity. By ensuring that the gravity direction of the transformed tracking map (i.e., after applying the determined transformation) is aligned with the gravity direction of the canonical map with which the tracking map is to be merged, merging errors can be reduced.

[0126] The technology described herein can be used with or alone on many types of devices and for many types of scenarios, including wearable or portable devices that provide augmented or mixed reality scenarios with limited computing resources. In some embodiments, the technology can be implemented by one or more services that form part of an XR system.

[0127] AR System Overview

[0128] Figure 1 and Figure 2 A scene with virtual content is shown, which is displayed together with a portion of the physical world. For illustration purposes, an AR system is used as an example of an XR system. Figure 3-6B An exemplary AR system is shown that includes one or more processors, memory, sensors, and a user interface that can operate according to the techniques described herein.

[0129] refer to Figure 1 , depicting an outdoor AR scene 354 in which the user of the AR technology sees a park-like setting 356 of the physical world, which features people, trees, buildings in the background, and a concrete platform 358. In addition to these items, the user of the AR technology also perceives that they "see" a robot statue 357 standing on the physical world concrete platform 358, and a flying cartoon-like avatar character 352 that appears to be the head of a bumblebee, even though these elements (e.g., avatar character 352 and robot statue 357) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is challenging to produce an AR technology that promotes a comfortable, natural feeling, rich presentation of virtual image elements among other virtual or physical world image elements.

[0130] Such an AR scene can be implemented by a system that builds a map of the physical world based on tracking information, enables users to place AR content in the physical world, determines where to place AR content in the map of the physical world, preserves the AR scene so that the placed AR content can be reloaded during, for example, different AR experience sessions to be displayed in the physical world, and enables multiple users to share AR experiences. The system can build and update a digital representation of the physical world surface around the user. The representation can be used to render virtual content to appear to be fully or partially occluded by physical objects between the user and the rendered location of the virtual content, to place virtual objects in physics-based interactions, and for virtual character path planning and navigation, or for other operations in which information about the physical world is used.

[0131] Figure 2 Another example of an indoor AR scene 400 is depicted, showing an exemplary use case of an XR system, in accordance with some embodiments. The exemplary scene 400 is a living room with a wall, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, the user of the AR technology may also perceive virtual objects, such as an image on the wall behind the sofa, a bird flying through the door, a deer peeking from the bookshelf, and a decoration in the form of a windmill placed on the coffee table.

[0132] For an image on a wall, AR technology needs information not only about the surface of the wall, but also about objects and surfaces in the room (such as the shape of a lamp) that occlude the image to correctly render the virtual object. For a flying bird, AR technology needs information about all the objects and surfaces around the room in order to render the bird with realistic physics to avoid objects and surfaces or avoid bouncing when the bird collides. For a deer, AR technology needs information about surfaces (such as the floor or coffee table) to calculate where to place the deer. For a windmill, the system can recognize that it is a separate object from the table and can determine that it is movable, while the corner of a shelf or the corner of a wall can be determined to be stationary. This distinction can be used to determine which parts of the scene are used or updated in each of the various operations.

[0133] Virtual objects may be placed in a previous AR experience session. When a new AR experience session starts in the living room, AR technology needs to display the virtual objects exactly where they were previously placed and actually visible from a different perspective. For example, a windmill should be displayed standing on a book, rather than floating above the table in a different location without a book. This floating may occur if the user of the new AR experience session is not accurately positioned in the living room. As another example, if the user views the windmill from a different perspective than when the windmill was placed, AR technology needs to display the corresponding side of the windmill.

[0134] The scene can be presented to the user via a system including multiple components, including a user interface that can stimulate one or more user senses (such as vision, sound and / or touch). In addition, the system can include one or more sensors that can measure parameters of the physical part of the scene, including the position and / or movement of the user within the physical part of the scene. In addition, the system can include one or more computing devices, and associated computer hardware, such as memory. These components can be integrated into a single device, or can be distributed across multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.

[0135] Figure 3 An AR system 502 is depicted that is configured to provide an experience of AR content that interacts with a physical world 506 in accordance with some embodiments. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by a user as part of a headset so that the user may wear the display over their eyes like a pair of goggles or glasses. At least a portion of the display may be transparent so that the user may observe a see-through reality 510. The see-through reality 510 may correspond to a portion of the physical world 506 that is within a current viewpoint of the AR system 502, which may correspond to the user's viewpoint if the user wears a headset that incorporates a display and sensors of the AR system to obtain information about the physical world.

[0136] AR content may also be presented on display 508, overlaid on see-through reality 510. To provide accurate interaction between AR content and see-through reality 510 on display 508, AR system 502 may include sensors 522 configured to capture information about physical world 506.

[0137] Sensors 522 may include one or more depth sensors that output depth maps 512. Each depth map 512 may have a plurality of pixels, each of which may represent a distance from a surface in the physical world 506 in a particular direction relative to the depth sensor. Raw depth data may come from the depth sensor to create a depth map. The depth map may be updated as fast as the depth sensor can form a new image, which may be hundreds or thousands of times per second. However, the data may be noisy and incomplete, and have holes shown as black pixels on the depth map shown.

[0138] The system may include other sensors, such as image sensors. Image sensors may acquire monocular or stereo information, which may be processed to represent the physical world in other ways. For example, images may be processed in world reconstruction component 516 to create a mesh that represents connected parts of objects in the physical world. Metadata about such objects, including, for example, color and surface texture, may similarly be acquired using sensors and stored as part of the world reconstruction.

[0139] The system may also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, a head pose tracking component of the system may be used to calculate the head pose in real time. The head pose tracking component may represent the user's head pose in a coordinate system having six degrees of freedom, including, for example, translation of three vertical axes (e.g., forward / backward, up / down, left / right) and rotation about the three vertical axes (e.g., pitch, yaw, and roll). In some embodiments, the sensor 522 may include an inertial measurement unit that may be used to calculate and / or determine the head pose 514. The head pose 514 for the depth map may indicate, for example, the current viewpoint of the sensor that captured the depth map in six degrees of freedom, but the headset 514 may be used for other purposes, such as associating image information with a specific part of the physical world or associating the position of a display worn on the user's head with the physical world.

[0140] In some embodiments, head pose information can be derived in other ways than an IMU (such as analyzing objects in an image). For example, a head pose tracking component can calculate the relative position and orientation of the AR device relative to a physical object based on visual information captured by a camera and inertial information captured by an IMU. The head pose tracking component can then calculate the head pose of the AR device, for example, by comparing the calculated relative position and orientation of the AR device relative to the physical object with features of the physical object. In some embodiments, the comparison can be performed by identifying features in images captured using one or more sensors 522, which are stable over time so that changes in the positions of these features in images captured over time can be associated with changes in the user's head pose.

[0141] The inventors have recognized and appreciated techniques for operating an XR system to provide an XR scene for a more immersive user experience, such as estimating head pose at a frequency of 1kHz, with low usage of computing resources associated with an XR device, which may be configured with, for example, four video graphics array (VGA) cameras operating at 30Hz, one inertial measurement unit (IMU) operating at 1kHz, the computing power of a single Advanced RISC Machine (ARM) core, less than 1GB of memory, and less than 100Mbp of network bandwidth. These techniques relate to reducing the processing required to generate and maintain maps and estimate head pose, and providing and using data with low computational overhead. The XR system can compute its pose based on matching visual features. U.S. patent application Ser. No. 16 / 221,065 describes hybrid tracking, and is hereby incorporated by reference in its entirety.

[0142] In some embodiments, the AR device can build a map based on feature points identified in consecutive images in a series of image frames captured as the user moves throughout the physical world with the AR device. Although each image frame can be taken from a different posture of the user as they move, the system can adjust the orientation of the features of each consecutive image frame to match the orientation of the initial image frame by matching the features of the consecutive image frames with the previously captured image frames. The translation of consecutive image frames causes points representing the same features to match corresponding feature points in previously collected image frames, which can be used to align each consecutive image frame to match the orientation of the previously processed image frame. The frames in the generated map can have a common orientation established when the first image frame is added to the map. The map has multiple sets of feature points in a common reference frame, which can be used to determine the user's posture in the physical world by matching the features in the current image frame with the map. In some embodiments, the map can be referred to as a tracking map.

[0143] In addition to being able to track the user's posture in the environment, the map can also enable other components of the system, such as a world reconstruction component 516, to determine the position of physical objects relative to the user. The world reconstruction component 516 can receive the depth map 512 and head posture 514 and any other data from the sensor and integrate the data into the reconstruction 518. The reconstruction 518 can be more complete and less noisy than the sensor data. The world reconstruction component 516 can update the reconstruction 518 using spatial and temporal averages of sensor data from multiple viewpoints over time.

[0144] Reconstruction 518 may include a representation of the physical world in one or more data formats, including, for example, voxels, grids, planes, etc. Different formats may represent alternative representations of the same portion of the physical world or may represent different portions of the physical world. In the example shown, on the left side of reconstruction 518, portions of the physical world are presented as global surfaces; on the right side of reconstruction 518, portions of the physical world are presented as grids.

[0145] In some embodiments, the map maintained by the head posture component 514 may be sparse relative to other maps of the physical world that may be maintained. A sparse map may indicate the location of points of interest and / or structures (e.g., corners or edges) rather than providing information about the location of surfaces and possible other features. In some embodiments, the map may include image frames captured by the sensor 522. These frames may be simplified to features that may represent points of interest and / or structures. In conjunction with each frame, information about the posture of the user from which the frame was acquired may also be stored as part of the map. In some embodiments, each image acquired by the sensor may or may not be stored. In some embodiments, when images are collected by the sensor, the system may process the images and select a subset of image frames for further calculations. The selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may, for example, add new image frames to the map based on overlap with previous image frames that have been added to the map or based on image frames that contain a sufficient number of features that are determined to be likely to represent stationary objects. In some embodiments, the selected image frames or feature groups from the selected image frames may be used as key frames for the map, which are used to provide spatial information.

[0146] In some embodiments, the amount of data processed when building a map can be reduced, such as by building a sparse map with a set of map points and keyframes and / or dividing the map into blocks to enable block-by-block updates. Map points can be associated with points of interest in the environment. Keyframes may include information selected from data captured by a camera. U.S. Patent Application No. 16 / 520,582 describes determining and / or evaluating a positioning map and is hereby incorporated by reference in its entirety.

[0147] The AR system 502 can integrate sensor data from multiple perspectives of the physical world over time. As the device including the sensor moves, the pose (e.g., position and orientation) of the sensor can be tracked. Since the frame pose of the sensor and its relationship to other poses are known, each of these multiple viewpoints of the physical world can be fused together to form a single combined reconstruction of the physical world, which can be used as an abstract layer of the map and provide spatial information. By using spatial and temporal averaging (i.e., averaging data from multiple viewpoints over time) or any other appropriate method, the reconstruction can be more complete and less noisy than the original sensor data.

[0148] exist Figure 3 In the illustrated embodiment, the map represents a portion of the user's physical world in which a single wearable device is present. In that case, the head pose associated with a frame in the map can be represented as a local head pose, indicating an orientation relative to an initial orientation of the single device at the start of the session. For example, the head pose can be tracked relative to the initial head pose when the device is turned on, or otherwise operated to scan the environment to build a representation of the environment.

[0149] In conjunction with the content representing that portion of the physical world, a map may include metadata. The metadata may, for example, indicate the time at which sensor information used to form the map was captured. Alternatively or in addition, the metadata may indicate the location of the sensor when the information used to form the map was captured. The location may be represented directly, such as using information from a GPS chip, or indirectly, such as using a wireless (e.g., Wi-Fi) signature that indicates the strength of the signal received from one or more wireless access points while the sensor data was being collected, and / or using an identifier, such as a BSSID, of a wireless access point to which the user device was connected while the sensor data was being collected.

[0150] Reconstruction 518 can be used for AR functions, such as generating a surface representation of the physical world for occlusion handling or physics-based processing. The surface representation may change as the user moves or objects in the real world change. Aspects of reconstruction 518 can be used, for example, by a component 520 that generates a changing global surface representation in world coordinates, which can be used by other components.

[0151] AR content can be generated based on this information, such as by an AR application 504. The AR application 504 can be, for example, a game program that performs one or more functions, such as visual occlusion, physics-based interaction, and environmental reasoning, based on information about the physical world. It can perform these functions by querying data in different formats from the reconstruction 518 generated by the world reconstruction component 516. In some embodiments, component 520 can be configured to output updates when the representation in the focus area of ​​the physical world changes. For example, the focus area can be set to approximate a portion of the physical world near the system user, such as a portion within the user's field of view, or projected (predicted / determined) as entering the user's field of view.

[0152] The AR application 504 can use this information to generate and update the AR content. The virtual portion of the AR content can be presented on the display 508 in combination with the see-through reality 510, thereby creating a realistic user experience.

[0153] In some embodiments, an AR experience may be provided to a user via an XR device, which may be a wearable display device, which may be part of a system that may include remote processing and / or remote data storage and / or, in some embodiments, other wearable display devices worn by other users. Figure 4 An example of a system 580 (hereinafter “system 580”) including a single wearable device is shown. System 580 includes a head mounted display device 562 (hereinafter “display device 562”), and various mechanical and electronic modules and systems that support the functionality of display device 562. Display device 562 can be coupled to a frame 564 that can be worn by a display system user or viewer 560 (hereinafter “user 560”) and is configured to position display device 562 in front of the eyes of user 560. According to various embodiments, display device 562 can be displayed sequentially. Display device 562 can be monocular or binocular. In some embodiments, display device 562 can be Figure 3 An example of display 508 in FIG.

[0154] In some embodiments, a speaker 566 is coupled to the frame 564 and positioned near the ear canal of the user 560. In some embodiments, another speaker, not shown, is positioned near the other ear canal of the user 560 to provide stereo / plastic sound control. The display device 562 is operably coupled to a local data processing module 570, such as by a wired conductor or wireless connection 568, which can be mounted in a variety of configurations, such as fixedly attached to the frame 564, fixedly attached to a helmet or hat worn by the user 560, embedded in headphones, or otherwise removably attached to the user 560 (e.g., in a backpack configuration, in a belt-coupled configuration).

[0155] The local data processing module 570 may include a processor and digital storage such as non-volatile memory (e.g., flash memory), both of which may be used to facilitate processing, caching, and storage of data. The data may include: a) data captured from a sensor (e.g., which may be operably coupled to the frame 564) or otherwise attached to the user 560, such as an image capture device (such as a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a radio device, and / or a gyroscope; and / or b) data acquired and / or processed using the remote processing module 572 and / or remote data repository 574, and possibly transferred to the display device 562 after such processing or acquisition.

[0156] In some embodiments, the wearable device can communicate with remote components. The local data processing module 570 can be operably coupled to the remote processing module 572 and the remote data repository 574 respectively via communication links 576, 578 (such as via wired or wireless communication links), so that these remote modules 572, 574 are operably coupled to each other and can be used as resources for the local data processing module 570. In further embodiments, in addition to or instead of the remote data repository 574, the wearable device can access a cloud-based remote data repository and / or service. In some embodiments, the above-mentioned head posture tracking component can be implemented at least in part in the local data processing module 570. In some embodiments, Figure 3 The world reconstruction component 516 in can be implemented at least in part in the local data processing module 570. For example, the local data processing module 570 can be configured to execute computer-executable instructions to generate a map and / or a physical world representation based at least in part on at least a portion of the data.

[0157] In some embodiments, processing can be distributed on a local processor and a remote processor. For example, local processing can be used to construct a map (e.g., a tracking map) on the user device based on sensor data collected using sensors on the user device. Such maps can be used by applications on the user device. In addition, previously created maps (e.g., specification maps) can be stored in a remote data repository 574. In the case where an appropriate stored or persistent map is available, it can be used in place of a tracking map created locally on the device or outside of a tracking map created locally on the device. In some embodiments, the tracking map can be located to a stored map so that a correspondence is established between the tracking map and the specification map, wherein the tracking map may be oriented relative to the position of the wearable device when the user turns on the system, and the specification map may be oriented relative to one or more persistent features. In some embodiments, a persistent map can be loaded on the user device to allow the user device to render virtual content without the delay associated with scanning the location, thereby constructing a tracking map of the user's entire environment based on the sensor data acquired during the scan. In some embodiments, the user device can access a remote persistent map (e.g., stored in the cloud) without downloading the persistent map on the user device.

[0158] In some embodiments, spatial information can be transmitted from the wearable device to a remote service, such as a cloud service configured to locate the device to a stored map maintained on the cloud service. According to one embodiment, the positioning process can be performed in the cloud, matching the device location to an existing map (e.g., a canonical map) and returning a transformation that links the virtual content to the wearable device location. In such an embodiment, the system can avoid transmitting the map from the remote resource to the wearable device. Other embodiments can be configured for device-based and cloud-based positioning, for example, to enable functionality where a network connection is not available or the user chooses not to enable cloud-based positioning.

[0159] Alternatively or additionally, the tracking map may be merged with previously stored maps to extend or improve the quality of those maps. The process of determining whether a suitable previously created environment map is available and / or merging the tracking map with one or more stored environment maps may be completed in the local data processing module 570 or the remote processing module 572.

[0160] In some embodiments, the local data processing module 570 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the computational budget of the local data processing module 570, but enable a smaller device. In some embodiments, the world reconstruction component 516 can use a computational budget less than a single Advanced RISC Machine (ARM) core to generate a physical world representation in real time on a non-predefined space, so that the remaining computational budget of the single ARM core can be accessed for other uses, such as, for example, extracting a mesh.

[0161] In some embodiments, the remote data repository 574 may include a digital data storage facility that may be available via the Internet or other networking configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, allowing for fully autonomous use from a remote module. In some embodiments, all data is stored and all or most computations are performed in the remote data repository 574, allowing for smaller devices. For example, a world reconstruction may be stored in whole or in part in this repository 574.

[0162] In embodiments where the data is stored remotely and can be accessed over a network, the data can be shared by multiple users of the augmented reality system. For example, user devices can upload their tracking maps to enhance the environmental map database. In some embodiments, the tracking map upload occurs at the end of the user session with the wearable device. In some embodiments, the tracking map upload can occur continuously, semi-continuously, or intermittently at a predefined time, after a predefined time period from a previous upload, or when triggered by an event. Whether based on data from the user device or any other user device, the tracking map uploaded by any user device can be used to expand or improve the previously stored map. Similarly, the persistent map downloaded to the user device can be based on data from the user device or any other user device. In this way, users can easily obtain high-quality environmental maps to improve their experience in the AR system.

[0163] In further embodiments, persistent map downloads may be limited and / or avoided based on positioning performed on a remote resource (e.g., in the cloud). In such a configuration, a wearable device or other XR device transmits feature information (e.g., positioning information of the device when a feature represented in the feature information is sensed) in combination with pose information to a cloud service. One or more components of the cloud service may match the feature information with a corresponding stored map (e.g., a canonical map) and generate a transformation between the coordinate systems of a tracking map maintained by the XR device and the canonical map. Each XR device whose tracking map is positioned relative to the canonical map can accurately render virtual content at a location specified relative to the canonical map based on its own tracking.

[0164] In some embodiments, the local data processing module 570 is operably coupled to a battery 582. In some embodiments, the battery 582 is a removable power source, such as above a counter battery. In other embodiments, the battery 582 is a lithium-ion battery. In some embodiments, the battery 582 includes both an internal lithium-ion battery that can be charged by the user 560 during the non-operating time of the system 580, and a removable battery so that the user 560 can operate the system 580 for a longer period of time without having to connect to a power source to charge the lithium-ion battery, or without having to shut down the system 580 to replace the battery.

[0165] Figure 5A A user 530 is shown wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as "environment 532"). Information captured by the AR system along the user's path of movement can be processed into one or more tracking maps. The user 530 positions the AR display system at a location 534, and the AR display system records environmental information of the traversable world relative to the location 534 (e.g., digital representations of real objects in the physical world, which can be stored and updated as changes are made to the real objects in the physical world). This information can be combined with images, features, directional audio input, or other desired data and stored as a gesture. The location 534 is aggregated to a data input 536, for example as part of a tracking map, and is processed by at least a traversable world module 538, which can, for example, be processed by Figure 4 The traversable world module 538 may include a head pose component 514 and a world reconstruction component 516 in some embodiments so that the processed information can be combined with other information related to physical objects used in rendering virtual content to indicate the location of the object in the physical world.

[0166] The traversable world module 538 determines, at least in part, where and how the AR content 540 as determined from the data input 536 can be placed in the physical world. The AR content is "placed" in the physical world by presenting both the physical world presentation and the AR content via a user interface, the AR content is rendered as if interacting with objects in the physical world, and the objects in the physical world are presented as if the AR content obscures the user's view of these objects when appropriate. In some embodiments, the AR content can be placed by appropriately selecting a portion of a fixed element 542 (e.g., a table) from a reconstruction (e.g., reconstruction 518) to determine the shape and position of the AR content 540. As an example, the fixed element can be a table, and the virtual content can be positioned so that it appears to be on the table. In some embodiments, the AR content can be placed within a structure in a field of view 544, which can be a current field of view or an estimated future field of view. In some embodiments, the AR content can be continuously relative to a model 546 (e.g., a grid) of the physical world.

[0167] As depicted, fixed element 542 serves as a proxy (e.g., a digital copy) for any fixed element within the physical world that may be stored in traversable world module 538, such that user 530 may perceive content on fixed element 542 without the system having to map to fixed element 542 each time user 530 sees fixed element 542. Thus, fixed element 542 may be a mesh model from a previous modeling session, or may be determined by a separate user but still stored by traversable world module 538 for future reference by multiple users. Thus, traversable world module 538 may recognize environment 532 from a previously mapped environment and display AR content without requiring user 530's device to first map all or a portion of environment 532, thereby saving computational processes and cycles and avoiding latency of any rendered AR content.

[0168] A mesh model 546 of the physical world may be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying AR content 540 may be stored by the traversable world module 538 for future retrieval by the user 530 or other users without having to recreate the model in whole or in part. In some embodiments, data input 536 is input such as geographic location, user identification, and current activity to indicate to the traversable world module 538 which of one or more fixed elements 542 is available, which AR content 540 was last placed on the fixed element 542, and whether to display the same content (such AR content is "persistent" content regardless of how the user views a particular traversable world model).

[0169] Even in embodiments where objects are considered fixed (e.g., a kitchen table), the traversable world module 538 may update those objects in the physical world model from time to time to account for the possibility of changes in the physical world. Models of fixed objects may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not considered fixed (e.g., a kitchen chair). In order to render a realistic AR scene, the AR system may update the positions of these non-fixed objects at a much higher frequency than that used to update fixed objects. In order to be able to accurately track all objects in the physical world, the AR system may obtain information from multiple sensors (including one or more image sensors).

[0170] Figure 5B is a schematic diagram of viewing optics assembly 548 and accompanying components. In some embodiments, two eye tracking cameras 550 directed toward user eyes 549 detect metrics of user eyes 549 such as eye shape, eyelid occlusion, pupil direction, and glint on user eyes 549.

[0171] In some embodiments, one of the sensors may be a depth sensor 551, such as a time-of-flight sensor, which transmits signals to the world and detects reflections of those signals from nearby objects to determine the distance to a given object. A depth sensor may, for example, quickly determine whether an object has entered the user's field of view due to the motion of those objects or a change in the user's posture. However, information about the location of an object in the user's field of view may alternatively or additionally be collected by other sensors. Depth information may, for example, be obtained from a stereoscopic image sensor or a plenoptic sensor.

[0172] In some embodiments, world camera 552 records a view larger than the periphery to map environment 532 and / or otherwise create a model of environment 532, and detect inputs that can affect AR content. In some embodiments, world camera 552 and / or camera 553 can be grayscale and / or color image sensors that can output grayscale and / or color image frames at fixed time intervals. Camera 553 can further capture an image of the physical world within the user's field of view at a particular time. Even if the value of a pixel of a frame-based image sensor does not change, its pixels can be repeatedly sampled. Each of world camera 552, camera 553, and depth sensor 551 has a corresponding field of view 554, 555, and 556 to obtain information such as the user's surroundings. Figure 5A Collect data in the physical world scene of the physical world environment 532 depicted in and record the physical world scene.

[0173] Inertial measurement unit 557 can determine the motion and orientation of viewing optical assembly 548. In some embodiments, inertial measurement unit 557 can provide an output indicating the direction of gravity. In some embodiments, each component is operably coupled to at least one other component. For example, depth sensor 551 can be operably coupled to eye tracking camera 550 to confirm the measured accommodation relative to the actual distance at which the user's eye 549 is looking.

[0174] It should be appreciated that viewing optics assembly 548 may include Figure 5B , and may include components instead of or in addition to the components shown. For example, in some embodiments, viewing optical assembly 548 may include two world cameras 552 instead of four. Alternatively or in addition, cameras 552 and 553 do not need to capture visible light images of their entire fields of view. Viewing optical assembly 548 may include other types of components. In some embodiments, viewing optical assembly 548 may include one or more dynamic vision sensors (DVS) whose pixels may asynchronously respond to relative changes in light intensity exceeding a threshold.

[0175] In some embodiments, based on the time-of-flight information, the viewing optical assembly 548 may not include a depth sensor 551. For example, in some embodiments, the viewing optical assembly 548 may include one or more plenoptic cameras whose pixels may capture light intensity and the angle of incident light, from which depth information may be determined. For example, a plenoptic camera may include an image sensor covered with a transmissive diffraction mask (TDM). Alternatively or in addition, the plenoptic camera may include an image sensor that includes angle-sensitive pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such a sensor may be used as a source of depth information instead of or in addition to the depth sensor 551.

[0176] It should also be understood that Figure 5B The configuration of components in is provided as an example. The viewing optics assembly 548 may include components having any suitable configuration that may be configured to provide the user with the maximum field of view practicable for a particular set of components. For example, if the viewing optics assembly 548 has a world camera 552, the world camera may be placed in a central region of the viewing optics assembly rather than on the side.

[0177] Information from the sensors in the viewing optical assembly 548 can be coupled to one or more processors in the system. The processor can generate data that can be rendered so that the user perceives virtual content that interacts with objects in the physical world. The rendering can be implemented in any suitable manner, including generating image data that depicts both physical and virtual objects. In other embodiments, physical and virtual content can be depicted in a scene by modulating the opacity of a display device that the user browses in the physical world. The opacity can be controlled to create the appearance of a virtual object and also prevent the user from seeing objects in the physical world that are obscured by the virtual object. In some embodiments, when viewed through a user interface, the image data can include only virtual content, which can be modified so that the virtual content is perceived by the user as realistically interacting with the physical world (e.g., clipping the content to take into account the obstruction).

[0178] The location of displayed content on viewing optical assembly 548 to create the impression that an object is located at a particular location can depend on the physical properties of the viewing optical assembly. In addition, the posture of the user's head relative to the physical world and the direction of the user's eye gaze can affect the location of displayed content in the physical world will appear at a particular location on the viewing optical assembly. Sensors as described above can collect this information, and / or provide information from which this information can be calculated, so that a processor receiving the sensor input can calculate the location where an object should be rendered on viewing optical assembly 548 to create a desired appearance for the user.

[0179] Regardless of how content is presented to the user, a model of the physical world can be used so that properties of virtual objects that can be affected by physical objects can be correctly calculated, including the shape, position, motion, and visibility of virtual objects. In some embodiments, the model can include a reconstruction of the physical world, such as reconstruction 518.

[0180] The model may be created from data collected from sensors on the user's wearable device. However, in some embodiments, the model may be created from data collected from multiple users, which may be aggregated in a computing device remote from all users (and the data may be in the "cloud").

[0181] The model may be created at least in part by a world reconstruction system, such as, for example, Fig. 6A Described in more detail in Figure 3The world reconstruction component 516 may include a perception module 660 that can generate, update, and store a representation of a portion of the physical world. In some embodiments, the perception module 660 may represent the portion of the physical world within the reconstruction range of the sensor as a plurality of voxels. Each voxel may correspond to a 3D cube of a predetermined volume in the physical world and include surface information indicating whether a surface exists in the volume represented by the voxel. Voxels may be assigned values ​​indicating whether their corresponding volumes have been determined to include the surface of the physical object, determined to be empty, or have not yet been measured with the sensor and therefore their values ​​are unknown. It should be understood that there is no need to explicitly store values ​​indicating voxels that are determined to be empty or unknown, as the values ​​of the voxels may be stored in the computer memory in any suitable manner, including not storing information for voxels that are determined to be empty or unknown.

[0182] In addition to generating information for the persistent world representation, the perception module 660 can also identify and output indications of changes in the area surrounding the user of the AR system. Such indications of changes can trigger updates to volumetric data stored as part of the persistent world, or trigger other functions, such as triggering the generation of AR content to update the trigger component 604 of the AR content.

[0183] In some embodiments, the perception module 660 can identify changes based on a signed distance function (SDF) model. The perception module 660 can be configured to receive sensor data such as, for example, a depth map 660a and a head pose 660b, and then fuse the sensor data into an SDF model 660c. The depth map 660a can directly provide SDF information, and the image can be processed to obtain the SDF information. The SDF information represents the distance from the sensor used to capture the information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and therefore from the perspective of the user. The head pose 660b can enable the SDF information to be related to voxels in the physical world.

[0184] In some embodiments, the perception module 660 can generate, update, and store representations of portions of the physical world within a perception range. The perception range can be determined at least in part based on a reconstruction range of the sensor, which can be determined at least in part based on a limit of an observation range of the sensor. As a specific example, an active depth sensor operating using active IR pulses can reliably operate within a range of distances, thereby creating an observation range of the sensor, which can be from a few centimeters or tens of centimeters to several meters.

[0185] The world reconstruction component 516 may include additional modules that may interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data acquired by the perception module 660. The persistent world module 662 may also include representations of the physical world in various formats. For example, volume metadata 662b such as voxels may be stored, as well as meshes 662c and planes 662d. In some embodiments, other information may be saved, such as a depth map.

[0186] In some embodiments, a representation of the physical world (such as Fig. 6A The representation shown in ) can provide relatively dense information about the physical world compared to sparse maps (such as the feature point-based tracking maps described above).

[0187] In some embodiments, the perception module 660 may include modules that generate representations of the physical world in various formats, including, for example, grids 660d, planes, and semantics 660e. Representations of the physical world can be stored across local storage media and remote storage media. Depending on, for example, the location of the storage media, representations of the physical world can be described in different coordinate frames. For example, a representation of the physical world stored in a device can be described in a coordinate frame that is local to the device. The representation of the physical world can have a corresponding representation stored in the cloud. The corresponding representation in the cloud can be described in a coordinate frame shared by all devices in the XR system.

[0188] In some embodiments, these modules can generate representations based on data within the perception range of one or more sensors when generating the representation, as well as data captured at previous times and information in the persistent world module 662. In some embodiments, these components can operate on depth information captured using a depth sensor. However, the AR system can include a visual sensor and can generate such a representation by analyzing monocular or binocular visual information.

[0189] In some embodiments, these modules may operate on regions of the physical world. When perception module 660 detects a change in the physical world in a sub-region of the physical world, those modules may be triggered to update the sub-region of the physical world. For example, such a change may be detected by detecting a new surface or other criteria in SDF model 660c (e.g., changing the value of a sufficient number of voxels representing the sub-region).

[0190] The world reconstruction component 516 may include a component 664 that may receive a representation of the physical world from the perception module 660. Information about the physical world may be extracted by these components based on, for example, a usage request from an application. In some embodiments, information may be pushed to the usage component, such as via an indication of a change in a pre-identified area or a change in the representation of the physical world within the perception range. The component 664 may include, for example, a game program and other components that perform processing for visual occlusion, physics-based interaction, and environmental reasoning.

[0191] In response to a query from component 664, perception module 660 can send representations for the physical world in one or more formats. For example, when component 664 indicates that the use is for visual occlusion or physics-based interaction, perception module 660 can send representations of surfaces. When component 664 indicates that the use is for environmental reasoning, perception module 660 can send meshes, planes, and semantics of the physical world.

[0192] In some embodiments, perception module 660 may include a component that formats information to provide component 664. An example of such a component may be ray casting component 660f. Using a component (e.g., component 664), for example, information about the physical world may be queried from a particular viewpoint. Ray casting component 660f may select from one or more representations of physical world data within the field of view from the viewpoint.

[0193] It should be understood from the above description that the perception module 660 or another component of the AR system can process data to create a 3D representation of a portion of the physical world. The data to be processed can be reduced by: culling portions of the 3D reconstruction volume based at least in part on the camera cone and / or depth image; extracting and retaining planar data; capturing, retaining, and updating 3D reconstruction data in blocks that allow local updates while maintaining neighbor consistency; providing occlusion data to an application that generates such a scene, wherein the occlusion data is derived from a combination of one or more depth data sources; and / or performing multi-stage mesh simplification. The reconstruction can contain data of varying complexity, including, for example, raw data (e.g., real-time depth data), fused volumetric data (e.g., voxels), and computed data (e.g., meshes).

[0194] In some embodiments, components of the traversable world model may be distributed, with some portions executing locally on the XR device and some portions executing remotely, such as on a network-connected server, or in the cloud. The allocation of information processing and storage between the local XR device and the cloud can affect the functionality and user experience of the XR system. For example, reducing processing on the local device by allocating processing to the cloud can extend battery life and reduce heat generated on the local device. However, allocating too much processing to the cloud may create undesirable latency, which results in an unacceptable user experience.

[0195] Figure 6B A distributed component architecture 600 configured for spatial computing according to some embodiments is depicted. The distributed component architecture 600 may include a traversable world component 602 (e.g., Figure 5A 608, LuminOS 604, API 606, SDK 608, and applications 610. LuminOS 604 may include a Linux-based kernel with custom drivers compatible with XR devices. API 606 may include an application programming interface that allows XR applications (e.g., application 610) to access the spatial computing features of the XR device. SDK 608 may include a software development kit that allows the creation of XR applications.

[0196] One or more components in architecture 600 may create and maintain a model of the navigable world. In this example, sensor data is collected on a local device. Processing of this sensor data may be performed partially locally on the XR device and partially in the cloud. PW 538 may include a map of the environment created based at least in part on data captured by AR devices worn by multiple users. During a session of an AR experience, each AR device (such as the one described above in combination) may be used to create a map of the environment. Figure 4 The wearable device described herein can create a tracking map, which is a type of map.

[0197] In some embodiments, the device may include components for building sparse maps and dense maps. The tracking map can be used as a sparse map and can include the head pose of the AR device scanning the environment and information about the objects detected within the environment at each head pose. Those head poses can be maintained locally for each device. For example, the head pose on each device can be an initial head pose relative to when the device opens its session. As a result, each tracking map can be local to the device that created it and may have its own reference frame defined by its own local coordinate system. However, in some embodiments, the tracking map on each device can be formed so that one coordinate of its local coordinate system is aligned with the direction of gravity measured by its sensor (e.g., inertial measurement unit 557).

[0198] The dense map may include surface information, which may be represented by a mesh or depth information. Alternatively or additionally, the dense map may include higher-level information derived from the surface or depth information, such as the location and / or features of planes and / or other objects.

[0199] In some embodiments, the creation of dense maps can be independent of the creation of sparse maps. For example, the creation of dense maps and sparse maps can be performed in separate processing pipelines within the AR system. For example, separate processing can enable the generation or processing of different types of maps to be performed at different rates. For example, the refresh rate of a sparse map may be faster than the refresh rate of a dense map. However, in some embodiments, the processing of dense maps and sparse maps may be related even if performed in different pipelines. For example, changes in the physical world revealed in a sparse map can trigger an update of a dense map, and vice versa. In addition, even if created independently, these maps can be used together. For example, a coordinate system derived from a sparse map can be used to define the position and / or orientation of objects in a dense map.

[0200] Sparse maps and / or dense maps can be persisted for reuse by the same device and / or shared with other devices. Such persistence can be achieved by storing information in the cloud. The AR device can send the tracking map to the cloud to merge, for example, with an environment map selected from a persistent map previously stored in the cloud. In some embodiments, the selected persistent map can be sent from the cloud to the AR device for merging. In some embodiments, the persistent map can be oriented relative to one or more persistent coordinate systems. Such maps can be used as canonical maps because they can be used by any of multiple devices. In some embodiments, the model of the traversable world can include or be created by or based on one or more canonical maps. Even if some operations are performed based on the coordinate frame local to the device, the device can use the canonical map by determining the transformation between the coordinate frame local to the device and the canonical map.

[0201] A canonical map can be derived from a tracking map (TM) (e.g. Fig.31A 1102 in the canonical map), which can be promoted to a canonical map. The canonical map can be persisted so that a device accessing the canonical map, once it determines the transformation between its local coordinate system and the coordinate system of the canonical map, can use the information in the canonical map to determine the location of objects represented in the canonical map in the physical world around the device. In some embodiments, the TM can be a sparse map of head pose created by the XR device. In some embodiments, the canonical map can be created when the XR device sends one or more TMs to a cloud server to be merged with additional TMs captured by the XR device at a different time or by other XR devices.

[0202] In embodiments where a tracking map is formed on a local device and one of the coordinates of the local coordinate system is aligned with gravity, that orientation relative to gravity may be preserved when creating a canonical map. For example, a tracking map submitted for merging may be promoted to a canonical map when the tracking map does not overlap with any previously stored maps. Other tracking maps may also have orientations relative to gravity and may subsequently be merged with the canonical map. The merge may be performed to ensure that the resulting canonical map maintains its orientation relative to gravity. For example, if the gravity-aligned coordinates of each map are not aligned with each other to a close enough tolerance, the two maps may not be merged, regardless of the correspondence of the feature points in those maps.

[0203] A canonical map or other map may provide information about the portions of the physical world represented by the data that is processed to create the corresponding map. Figure 7 An exemplary tracking map 700 is depicted in accordance with some embodiments. The tracking map 700 may provide a floor plan 706 of physical objects in the corresponding physical world represented by the points 702. In some embodiments, the map points 702 may represent features of a physical object that may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. These features may be derived by processing an image, such as an image that may be acquired by a sensor of a wearable device in an augmented reality system. For example, features may be derived by processing image frames output by a sensor to identify features based on large gradients in the image or other appropriate criteria. Further processing may limit the number of features in each frame. For example, the processing may select features that may represent persistent objects. One or more heuristics may be applied to the selection.

[0204] The tracking map 700 may include data about points 702 collected by the device. For each image frame having data points included in the tracking map, a pose may be stored. The pose may represent the orientation from which the image frame was captured, so that feature points within each image frame may be spatially correlated. The pose may be determined by positioning information, such as may be derived by sensors on the wearable device (such as an IMU sensor). Alternatively or additionally, the pose may be determined by matching the image frame to other image frames that depict overlapping portions of the physical world. By looking for such positional correlations, which may be achieved by matching subsets of feature points in two frames, a relative pose between the two frames may be calculated. Relative poses may be sufficient for a tracking map because the map may be relative to a coordinate system local to the device established based on an initial pose of the device when construction of the tracking map began.

[0205] Not all feature points and image frames collected by the device can be retained as part of the tracking map, as much of the information collected with the sensors is likely to be redundant. Instead, only certain frames can be added to the map. Those frames can be selected based on one or more criteria, such as the degree of overlap with image frames already in the map, the number of new features they contain, or a quality measure of the features in the frame. Image frames that are not added to the tracking map can be discarded or can be used to modify the location of the features. As another alternative, all or most of the image frames represented as a set of features can be retained, but a subset of these frames can be designated as key frames for further processing.

[0206] The keyframes may be processed to produce a keyrig 704. The keyframes may be processed to produce a three-dimensional set of feature points and saved as a keyrig 704. For example, such processing may require comparing image frames obtained simultaneously from two cameras to stereoscopically determine the 3D positions of the feature points. Metadata may be associated with these keyframes and / or keyrigs (e.g., poses).

[0207] The environment map can have any of a variety of formats depending on, for example, where the environment map is stored, including, for example, local storage of the AR device and remote storage. For example, on a wearable device with limited memory, the map in the remote storage may have a higher resolution than the map in the local storage. In order to send the higher resolution map from the remote storage to the local storage, the map can be downsampled or otherwise converted to an appropriate format, such as by reducing the number of poses for each area of ​​the physical world stored in the map and / or the number of feature points stored for each pose. In some embodiments, slices or portions of the high-resolution map from the remote storage can be sent to the local storage, where the slices or portions are not downsampled.

[0208] When a new tracking map is created, the environment map database may be updated. To determine which of a potentially very large number of environment maps in the database is to be updated, the updating may include effectively selecting one or more environment maps stored in the database that are related to the new tracking map. The selected one or more environment maps may be ranked by relevance, and one or more of the highest ranked maps may be selected for processing to merge the higher ranked selected environment maps with the new tracking map to create one or more updated environment maps. When the new tracking map represents a portion of the physical world for which no pre-existing environment map is to be updated, the tracking map may be stored in the database as a new environment map.

[0209] Watch standalone display

[0210] Methods and apparatus are described herein for providing virtual content using an XR system that is independent of the position of the eyes viewing the virtual content. Traditionally, virtual content is re-rendered upon any movement of the display system. For example, if a user wearing a display system views a virtual representation of a three-dimensional (3D) object on a display and walks around an area where the 3D object appears, the 3D object should be re-rendered for each viewpoint so that the user has the perception that he or she is walking around the object occupying real space. However, re-rendering consumes a large amount of the system's computing resources and causes artifacts due to latency.

[0211] The inventors have recognized and appreciated that head pose (e.g., position and orientation of a user wearing an XR system) can be used to render virtual content that is independent of eye rotation within the user's head. In some embodiments, a dynamic map of a scene can be generated based on multiple coordinate frames in real space across one or more sessions, so that virtual content that interacts with the dynamic map can be robustly rendered, independent of eye rotation within the user's head and / or independent of sensor deformation caused by, for example, heat generated during high-speed, computationally intensive operations. In some embodiments, the configuration of multiple coordinate frames can enable a first XR device worn by a first user and a second XR device worn by a second user to identify a common location in a scene. In some embodiments, the configuration of multiple coordinate frames can enable users wearing XR devices to view virtual content at the same location of a scene.

[0212] In some embodiments, a tracking map may be constructed in a world coordinate frame, which may have a world origin. When the XR device is powered on, the world origin may be the first pose of the XR device. The world origin may be aligned with gravity, allowing developers of XR applications to perform gravity alignment without additional work. Different tracking maps may be constructed in different world coordinate frames, because the tracking map may be captured by the same XR device in different sessions and / or by different XR devices worn by different users. In some embodiments, a session of an XR device may start from when the device is powered on to when the device is turned off. In some embodiments, the XR device may have a head coordinate frame, which may have a head origin. The head origin may be the current pose of the XR device when the image is taken. The difference between the head pose of the world coordinate frame and the head pose of the head coordinate frame may be used to estimate the tracking route.

[0213] In some embodiments, the XR device may have a camera coordinate frame that may have a camera origin. The camera origin may be the current pose of one or more sensors of the XR device. The inventors have recognized and appreciated that the configuration of the camera coordinate frame enables robust display of virtual content that is independent of eye rotation within the user's head. The configuration also enables robust display of virtual content that is independent of sensor deformation, such as due to heat generated during operation.

[0214] In some embodiments, the XR device may have a head unit having a head-mounted frame that a user can fix to their head and may include two waveguides, one in front of each eye of the user. The waveguides may be transparent so that ambient light from real-world objects can be transmitted through the waveguides and the user can see the real-world objects. Each waveguide may send projection light from a projector to a corresponding eye of the user. The projection light may form an image on the retina of the eye. Thus, the retina of the eye receives both the ambient light and the projection light. The user may simultaneously see real-world objects and one or more virtual objects created by the projection light. In some embodiments, the XR device may have sensors that detect real-world objects around the user. These sensors may, for example, be cameras that capture images that can be processed to identify the location of real-world objects.

[0215] In some embodiments, instead of attaching virtual content to a world coordinate frame, the XR system can assign a coordinate frame to the virtual content. Such a configuration enables the description of virtual content without regard to where the virtual content is rendered to the user, but the virtual content can be attached to a more persistent frame location, such as a coordinate frame that is rendered at a specified location. Figures 14 to 20C When the position of an object changes, the XR device can detect the change in the environment map and determine the movement of the head unit worn by the user relative to the real-world object.

[0216] Figure 8 A user is shown in a physical environment experiencing virtual content rendered by an XR system 10 according to some embodiments. The XR system may include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is in a physical environment with real objects in the form of a table 16.

[0217] In the example shown, the first XR device 12.1 includes a head unit 22, a waist pack 24, and a cable connection 26. The first user 14.1 secures the head unit 22 to his head and secures the waist pack 24, which is remote from the head unit 22, to his waist. The cable connection 26 connects the head unit 22 to the waist pack 24. The head unit 22 includes technology for displaying one or more virtual objects to the first user 14.1 while allowing the first user 14.1 to see real objects such as a table 16. The waist pack 24 primarily includes the processing and communication capabilities of the first XR device 12.1. In some embodiments, the processing and communication capabilities may reside in whole or in part in the head unit 22, so that the waist pack 24 may be removed or may be located in another device such as a backpack.

[0218] In the example shown, the waist pack 24 is connected to the network 18 via a wireless connection. The server 20 is connected to the network 18 and maintains data representing local content. The waist pack 24 downloads the data representing the local content from the server 20 via the network 18. The waist pack 24 provides the data to the head unit 22 via a cable connection 26. The head unit 22 may include a display having a light source, such as a laser light source or a light emitting diode (LED) light source, and a waveguide to guide the light.

[0219] In some embodiments, the first user 14.1 may mount the head unit 22 to his head and the waist pack 24 to his waist. The waist pack 24 may download image data from the server 20 via the network 18. The first user 14.1 may see the table 16 through the display of the head unit 22. A projector forming part of the head unit 22 may receive image data from the waist pack 24 and generate light based on the image data. The light may travel through one or more waveguides forming part of the display of the head unit 22. The light may then leave the waveguide and propagate onto the retina of the eye of the first user 14.1. The projector may generate light in a pattern that is replicated on the retina of the eye of the first user 14.1. The light that falls on the retina of the eye of the first user 14.1 may have a selected depth of field so that the first user 14.1 perceives an image at a preselected depth behind the waveguide. In addition, the two eyes of the first user 14.1 may receive slightly different images so that the brain of the first user 14.1 perceives one or more three-dimensional images at a selected distance from the head unit 22. In the example shown, the first user 14.1 perceives virtual content 28 above the table 16. The scale of the virtual content 28 and its position and distance from the first user 14.1 are determined by the data representing the virtual content 28 and the various coordinate frames used to display the virtual content 28 to the first user 14.1.

[0220] In the example shown, the virtual content 28 is invisible from the perspective of the drawing and is visible to the first user 14.1 using the first XR device 12.1. The virtual content 28 may initially reside as a data structure within the visual data and algorithms in the waist pack 24. The data structure may then manifest itself as light when the projector of the head unit 22 generates light based on the data structure. It should be understood that although the virtual content 28 does not exist in the three-dimensional space in front of the first user 14.1, the virtual content 28 still represents the three-dimensional space. Figure 1 , to illustrate the wearer perception of the head unit 22. Visualization of computer data in three-dimensional space may be used in this description to show how data structures that contribute to the rendering by one or more users relate to each other within the data structures in the waist pack 24.

[0221] Fig. 9 Components of a first XR device 12.1 are shown in accordance with some embodiments. The first XR device 12.1 may include a head unit 22, and various components that form part of the visual data and algorithms, including, for example, a rendering engine 30, various coordinate frames 32, various origin and destination coordinate frames 34, and various origin to destination coordinate frame transformers 36. The various coordinate systems may be based on intrinsic properties of the XR device, or may be determined by reference to other information, such as a persistent pose or a persistent coordinate system as described herein.

[0222] Head unit 22 may include a head mounted frame 40 , a display system 42 , a real object detection camera 44 , a motion tracking camera 46 , and an inertial measurement unit 48 .

[0223] The head-mounted frame 40 may have a Figure 8 The display system 42, real object detection camera 44, motion tracking camera 46 and inertial measurement unit 48 can be mounted to the head mounted frame 40 and thus move with the head mounted frame 40.

[0224] Coordinate system 32 may include local data system 52 , world frame system 54 , head frame system 56 , and camera frame system 58 .

[0225] The local data system 52 may include a data channel 62, a local frame determination routine 64, and local frame storage instructions 66. The data channel 62 may be an internal software routine, a hardware component such as an external cable or radio frequency receiver, or a hybrid component such as an open port. The data channel 62 may be configured to receive image data 68 representing virtual content.

[0226] A local frame determination routine 64 may be connected to the data channel 62. The local frame determination routine 64 may be configured to determine a local coordinate frame 70. In some embodiments, the local frame determination routine may determine the local coordinate frame based on a real-world object or real-world location. In some embodiments, the local coordinate frame may be based on a top edge relative to a bottom edge of the browser window, a head or foot of a character, a node on an outer surface of a prism or bounding box surrounding the virtual content, or any other suitable location of a coordinate frame that defines a facing direction of the virtual content and a location where the virtual content is placed (e.g., a node, such as a placement node or a PCF node), etc.

[0227] The local frame storage instructions 66 may be connected to the local frame determination routine 64. Those skilled in the art will appreciate that software modules and routines are "connected" to each other through subroutines, calls, and the like. The local frame storage instructions 66 may store the local coordinate frame 70 as a local coordinate frame 72 within the origin and destination coordinate frame 34. In some embodiments, the origin and destination coordinate frames 34 may be one or more coordinate frames that may be manipulated or transformed to allow virtual content to persist between sessions. In some embodiments, a session may be a time period between startup and shutdown of an XR device. Two sessions may be two startup and shutdown time periods of a single XR device, or startup and shutdown time periods of two different XR devices.

[0228] In some embodiments, the origin and destination coordinate frames 34 may be coordinate frames involved in one or more transformations required to enable the XR device of the first user and the XR device of the second user to recognize a common location. In some embodiments, the destination coordinate frame may be the output of a series of calculations and transformations applied to the target coordinate frame so that the first and second users view virtual content in the same location.

[0229] The rendering engine 30 may be connected to the data channel 62. The rendering engine 30 may receive image data 68 from the data channel 62 so that the rendering engine 30 may render virtual content based at least in part on the image data 68.

[0230] Display system 42 may be connected to rendering engine 30. Display system 42 may include components that transform image data 68 into visible light. The visible light may be formed into two patterns, one for each eye. The visible light may enter Figure 8 The eye of the first user 14.1 may be detected on the retina of the eye of the first user 14.1.

[0231] The real object detection camera 44 may include one or more cameras that can capture images from different sides of the head mounted frame 40. The motion tracking camera 46 may include one or more cameras that can capture images on the side of the head mounted frame 40. Instead of two sets of one or more cameras representing the real object detection camera 44 and the motion tracking camera 46, one set of one or more cameras may be used. In some embodiments, the cameras 44, 46 may capture images. As described above, these cameras may collect data for constructing a tracking map.

[0232] Inertial measurement unit 48 may include multiple devices for detecting motion of head unit 22. Inertial measurement unit 48 may include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of inertial measurement unit 48, in combination, track motion of head unit 22 in at least three orthogonal directions and around at least three orthogonal axes.

[0233] In the example shown, the world frame system 54 includes a world surface determination routine 78, a world frame determination routine 80, and a world frame storage instruction 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 accepts images and / or keyframes based on images captured by the real object detection camera 44 and processes the images to identify surfaces in the images. A depth sensor (not shown) can determine the distance to the surface. Therefore, these surfaces are represented by data in three dimensions including their size, shape, and distance from the real object detection camera.

[0234] In some embodiments, the world coordinate frame 84 can be based on the origin when the head gesture session was initialized. In some embodiments, the world coordinate frame can be located at the location where the device was started, or if the head gesture was lost during the startup session, the world coordinate frame can be located in a new place. In some embodiments, the world coordinate frame can be the origin when the head gesture session started.

[0235] In the example shown, a world frame determination routine 80 is connected to the world surface determination routine 78 and determines a world coordinate frame 84 based on the position of the surface determined by the world surface determination routine 78. World frame storage instructions 82 are connected to the world frame determination routine 80 to receive the world coordinate frame 84 from the world frame determination routine 80. The world frame storage instructions 82 store the world coordinate frame 84 as a world coordinate frame 86 within the origin and destination coordinate frame 34.

[0236] The head frame system 56 may include a head frame determination routine 90 and head frame storage instructions 92. The head frame determination routine 90 may be connected to the motion tracking camera 46 and the inertial measurement unit 48. The head frame determination routine 90 may use data from the motion tracking camera 46 and the inertial measurement unit 48 to calculate a head coordinate frame 94. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The motion tracking camera 46 may continuously capture images that are used by the head frame determination routine 90 to refine the head coordinate frame 94. When Figure 8 When the first user 14.1 in the image moves their head, the head unit 22 moves. The motion tracking camera 46 and the inertial measurement unit 48 may continuously provide data to the head frame determination routine 90 so that the head frame determination routine 90 may update the head coordinate frame 94.

[0237] The head frame storage instructions 92 may be connected to the head frame determination routine 90 to receive the head coordinate frame 94 from the head frame determination routine 90. The head frame storage instructions 92 may store the head coordinate frame 94 as a head coordinate frame 96 in the origin and destination coordinate frames 34. The head frame storage instructions 92 may repeatedly store the updated head coordinate frame 94 as the head coordinate frame 96 when the head frame determination routine 90 recalculates the head coordinate frame 94. In some embodiments, the head coordinate frame may be the position of the wearable XR device 12.1 relative to the local coordinate frame 72.

[0238] The camera frame system 58 may include camera intrinsics 98. The camera intrinsics 98 may include the dimensions of the head unit 22 as a feature of its design and manufacture. The camera intrinsics 98 may be used to calculate a camera coordinate frame 100 stored within the origin and destination coordinate frame 34.

[0239] In some embodiments, the camera coordinate frame 100 may include Figure 8 1. When the left eye moves from left to right or up and down, the pupil position of the left eye is located in the camera coordinate frame 100. In addition, the pupil position of the right eye is located in the camera coordinate frame 100 of the right eye. In some embodiments, the camera coordinate frame 100 may include the position of the camera relative to the local coordinate frame when the image is captured.

[0240] The origin-to-destination coordinate frame transformer 36 may include a local-to-world coordinate transformer 104, a world-to-head coordinate transformer 106, and a head-to-camera coordinate transformer 108. The local-to-world coordinate transformer 104 may receive the local coordinate frame 72 and transform the local coordinate frame 72 to the world coordinate frame 86. The transformation of the local coordinate frame 72 to the world coordinate frame 86 may be represented as a transformation of the local coordinate frame within the world coordinate frame 86 to the world coordinate frame 110.

[0241] The world to head coordinate transformer 106 may transform from the world coordinate frame 86 to the head coordinate frame 96. The world to head coordinate transformer 106 may transform the local coordinate frame transformed to the world coordinate frame 110 to the head coordinate frame 96. The transformation may be represented as a local coordinate frame transformed within the head coordinate frame 96 to the head coordinate frame 112.

[0242] The head-to-camera coordinate transformer 108 may transform from the head coordinate frame 96 to the camera coordinate frame 100. The head-to-camera coordinate transformer 108 may transform the local coordinate frame transformed to the head coordinate frame 112 to the local coordinate frame transformed to the camera coordinate frame 114 within the camera coordinate frame 100. The local coordinate frame transformed to the camera coordinate frame 114 may be input to the rendering engine 30. The rendering engine 30 may render the image data 68 representing the local content 28 based on the local coordinate frame transformed to the camera coordinate frame 114.

[0243] Fig.10 is a spatial representation of the various origin and destination coordinate frames 34. In the figure, a local coordinate frame 72, a world coordinate frame 86, a head coordinate frame 96, and a camera coordinate frame 100 are represented. In some embodiments, when virtual content is placed in the real world so that a user can view the virtual content, the local coordinate frame associated with the XR content 28 can have a position and rotation relative to the local and / or world coordinate frames and / or PCF (e.g., a node and facing direction can be provided). Each camera can have its own camera coordinate frame 100 that contains all pupil locations of one eye. Reference numerals 104A and 106A respectively represent the camera coordinate frames 104A and 106A represented by Fig. 9 The transformations are performed by the local to world coordinate transformer 104, the world to head coordinate transformer 106 and the head to camera coordinate transformer 108 in FIG.

[0244] Fig.11A camera rendering protocol for transforming from a head coordinate frame to a camera coordinate frame according to some embodiments is depicted. In the example shown, the pupil of a single eye moves from position A to position B. Virtual objects that are to appear stationary will be projected onto a depth plane at one of two positions A or B, depending on the position of the pupil (assuming the camera is configured to use a pupil-based coordinate frame). As a result, using a pupil coordinate frame transformed to a head coordinate frame will result in jitter of stationary virtual objects as the eye moves from position A to position B. This situation is referred to as view-dependent display or projection.

[0245] like Fig.12 As shown, a camera coordinate frame (e.g., CR) is placed and contains all pupil positions, and the object projection will now be consistent regardless of the pupil positions A and B. The head coordinate frame is transformed into the CR frame, which is called a view-independent display or projection. Image re-projection can be applied to the virtual content to account for changes in eye position, however, since the rendering is still in the same position, jitter can be minimized.

[0246] Fig.13 The display system 42 is shown in more detail. The display system 42 comprises a stereo analyser 144 which is connected to the rendering engine 30 and forms part of the visual data and algorithms.

[0247] The display system 42 further includes a left projector 166A and a right projector 166B and a left waveguide 170A and a right waveguide 170B. The left projector 166A and the right projector 166B are connected to a power supply. Each projector 166A and 166B has a corresponding input for image data to be provided to the corresponding projector 166A or 166B. The corresponding projector 166A or 166B generates a two-dimensional pattern of light and emits light therefrom when powered on. The left waveguide 170A and the right waveguide 170B are positioned to receive light from the left projector 166A and the right projector 166B, respectively. The left waveguide 170A and the right waveguide 170B are transparent waveguides.

[0248] In use, the user mounts the head mounted frame 40 to their head. The components of the head mounted frame 40 may, for example, include a strap (not shown) that wraps around the back of the user's head. The left and right waveguides 170A, 170B are then located in front of the user's left and right eyes 220A, 220B.

[0249] The rendering engine 30 inputs the received image data into the stereo analyzer 144. The image data is Figure 8The three-dimensional image data of the local content 28 in the image data is projected onto a plurality of virtual planes. The stereo analyzer 144 analyzes the image data to determine a left image data set and a right image data set based on the image data for projection onto each depth plane. The left image data set and the right image data set are data sets representing a two-dimensional image that is projected in three dimensions to give a user a sense of depth.

[0250] The stereo analyzer 144 inputs the left image data set and the right image data set to the left projector 166A and the right projector 166B. Then, the left projector 166A and the right projector 166B create the left illumination pattern and the right illumination pattern. The components of the display system 42 are shown in a plan view, but it should be understood that when shown in a front view, the left pattern and the right pattern are two-dimensional patterns. Each light pattern includes a plurality of pixels. For the purpose of illustration, it is shown that the light 224A and 226A from two pixels leave the left projector 166A and enter the left waveguide 170A. The light 224A and 226A are reflected from the side of the left waveguide 170A. It is shown that the light 224A and 226A propagate from left to right in the left waveguide 170A by internal reflection, but it should be understood that the light 224A and 226A also propagate into the paper in a certain direction using a refraction and reflection system.

[0251] Light rays 224A and 226A exit the left optical waveguide 170A through the pupil 228A and then enter the left eye 220A through the pupil 230A of the left eye 220A. Light rays 224A and 226A then fall on the retina 232A of the left eye 220A. In this manner, the left light pattern falls on the retina 232A of the left eye 220A. The perception to the user is that the pixels formed on the retina 232A are what the user perceives as pixels 234A and 236A at a certain distance on the side of the left waveguide 170A opposite the left eye 220A. Depth perception is created by manipulating the focal length of the light.

[0252] In a similar manner, the stereo analyzer 144 inputs the right image data set into the right projector 166B. The right projector 166B transmits a right light pattern, which is represented by pixels in the form of light rays 224B and 226B. Light rays 224B and 226B are reflected within the right waveguide 170B and exit through the pupil 228B. Light rays 224B and 226B then enter through the pupil 230B of the right eye 220B and fall on the retina 232B of the right eye 220B. The pixels of light rays 224B and 226B are perceived as pixels 134B and 236B behind the right waveguide 170B.

[0253] The patterns created on the retinas 232A and 232B are perceived as left and right images, respectively. The left and right images are slightly different from each other due to the function of the stereo analyzer 144. The left and right images are perceived as three-dimensional renderings in the user's mind.

[0254] As mentioned, the left waveguide 170A and the right waveguide 170B are transparent. Light from real objects on the table 16 on the side opposite to the eyes 220A and 220B, such as the left waveguide 170A and the right waveguide 170B, can be projected through the left waveguide 170A and the right waveguide 170B and fall on the retinas 232A and 232B.

[0255] Persistent Coordinate Framework (PCF)

[0256] Methods and apparatus are described herein for providing spatial persistence between user instances within a shared space. Without spatial persistence, virtual content placed in the physical world by a user in one session may not exist or may be misplaced in the view of a user in a different session. Without spatial persistence, virtual content placed in the physical world by one user may not exist or may be misplaced in the view of a second user, even if the second user intends to share the same physical space experience as the first user.

[0257] The inventors have recognized and appreciated that spatial persistence can be provided through a persistent coordinate system (PCF). A PCF can be defined based on one or more points that represent features (e.g., corners, edges) recognized in the physical world. Features can be selected so that they appear the same from one user instance of the XR system to another.

[0258] Additionally, drift during tracking that causes the calculated tracking path (e.g., camera trajectory) to deviate from the actual tracking path can cause the position of virtual content to appear misaligned when rendering relative to a local map that is based solely on a tracking map. As the XR device collects more information about the scene over time, the tracking map of the space can be refined to correct for drift. However, if virtual content is placed on real objects before map refinement and saved relative to the device's world coordinate frame derived from the tracking map, the virtual content may appear displaced as if the real objects had moved during map refinement. The PCF can be updated based on map refinement because the PCF is defined based on features and is updated as features move during map refinement.

[0259] The PCF may include six degrees of freedom, translation and rotation relative to a map coordinate system. The PCF may be stored in a local storage medium and / or a remote storage medium. Depending on, for example, the storage location, the translation and rotation of the PCF may be calculated relative to a map coordinate system. For example, a PCF used locally on a device may have a translation and rotation relative to the device's world coordinate frame. A PCF in the cloud may have a translation and rotation relative to a canonical coordinate frame of a canonical map.

[0260] PCF can provide a sparse representation of the physical world, providing less information about the physical world than all available information so that they can be efficiently processed and transferred. Techniques for processing persistent spatial information may include creating a dynamic map based on one or more coordinate systems in real space across one or more sessions, generating a persistent coordinate system (PCF) on the sparse map, which can be exposed to XR applications through, for example, an application programming interface (API).

[0261] Fig.14 11 is a block diagram illustrating the creation of a persistent coordinate system (PCF) and the attachment of XR content to the PCF according to some embodiments. Each box may represent digital information stored in a computer memory. In the case of application 1180, the data may represent computer executable instructions. In the case of virtual content 1170, the digital information may define, for example, a virtual object specified by application 1180. In the case of other boxes, the digital information may represent certain aspects of the physical world.

[0262] In the illustrated embodiment, one or more PCFs are created based on images captured by sensors on the wearable device. Fig.14 In an embodiment of the invention, the sensors are visual image cameras. These cameras may be the same cameras used to form the tracking map. Thus, Fig.14 Some of the suggested processing can be performed as part of updating the tracking map. However, Fig.14 It is shown that in addition to the tracking map, information providing persistence is also generated.

[0263] To derive a 3D PCF, two images 1110 from two cameras mounted to the wearable device in a configuration enabling stereo image analysis are processed together. Fig.14 An image 1 and an image 2 are shown, each of which is from one of the cameras. For simplicity, a single image from each camera is shown. However, each camera may output a stream of image frames, and the operation may be performed for multiple image frames in the stream. Fig.14 processing.

[0264] Thus, image 1 and image 2 may each be a frame in a sequence of image frames. Fig.14The process shown is repeated until the image frame containing the feature point provides a suitable image from which to form persistent spatial information. Alternatively or additionally, the process may be repeated when the user moves such that the user is no longer close enough to a previously identified PCF to reliably use the PCF to determine the position relative to the physical world. Fig.14 For example, the XR system can maintain the current PCF for the user. When the distance exceeds a threshold, the system can switch to a new current PCF that is closer to the user, which can be based on Fig.14 The process uses image frames acquired at the user's current location to generate.

[0265] Even when a single PCF is generated, the stream of image frames can be processed to identify image frames that depict content in the physical world that is likely to be stable and easily recognizable by devices near the region of the physical world depicted in the image frames. Fig.14 In the embodiment of the present invention, the process begins with the identification of features 1120 in the image. For example, features can be identified by finding locations in the image where a gradient exceeds a threshold or other feature, which can correspond to corners of an object, for example. In the embodiment shown, the features are points, but other identifiable features, such as edges, may be used instead or in addition.

[0266] In the illustrated embodiment, a fixed number N of features 1120 are selected for further processing. Those feature points may be selected based on one or more criteria, such as the magnitude of a gradient or proximity to other feature points. Alternatively or additionally, feature points may be selected tentatively, for example, based on properties that suggest the feature points are persistent. For example, a heuristic may be defined based on properties of feature points that may correspond to the corners of windows or doors or large pieces of furniture. Such a heuristic may take into account the feature points themselves and their surroundings. As a specific example, the number of feature points per image may be between 100 and 500 or between 150 and 250, for example 200.

[0267] Regardless of the number of feature points selected, descriptors may be calculated 1130 for the feature points. In this example, a descriptor is calculated for each selected feature point, but descriptors may be calculated for groups of feature points or subsets of feature points or for all features within an image. The descriptors characterize the feature points so that feature points that represent the same object in the physical world are assigned similar descriptors. The descriptors may enable alignment of two frames, such as may occur when one map is positioned relative to another. Instead of searching for a relative orientation of the frames that minimizes the distance between feature points of the two images, an initial alignment of the two frames may be performed by identifying feature points with similar descriptors. Alignment of image frames may be based on alignment points with similar descriptors, which may require less processing than calculating alignment of all feature points in the images.

[0268] Descriptors may be computed as a mapping of feature points to descriptors, or in some embodiments, as a mapping of patches of an image around a feature point to descriptors. Descriptors may be numerical quantities. U.S. Patent Application 16 / 190,948 describes computing descriptors for feature points, and is incorporated herein by reference in its entirety.

[0269] exist Fig.14 In the example of , descriptors are calculated for each feature point in each image frame 1130. Based on the descriptors and / or the feature points and / or the image itself, image frames may be identified as key frames 1140. In the illustrated embodiment, a key frame is an image frame that meets a certain criterion, which is then selected for further processing. For example, when making a tracking map, image frames that add meaningful information to the map may be selected as key frames for integration into the map. On the other hand, image frames that substantially overlap areas where image frames have already been integrated into the map may be discarded so that they do not become key frames. Alternatively or additionally, key frames may be selected based on the number and / or type of feature points in the image frames. In Fig.14 In embodiments of the present invention, keyframes 1150 selected for inclusion in the tracking map may also be considered keyframes for determining the PCF, although different or additional criteria for selecting keyframes for generating the PCF may be used.

[0270] although Fig.14 The keyframes are shown to be used for further processing, but the information obtained from the images may be processed in other forms. For example, feature points such as in a key assembly may be processed alternatively or additionally. Moreover, although the keyframes are described as being derived from a single image frame, there need not be a one-to-one relationship between the keyframes and the image frames obtained. For example, the keyframes may be obtained from multiple image frames, such as by splicing or aggregating the image frames together, so that only features that appear in multiple images are retained in the keyframe.

[0271] The key frame may include image information and / or metadata associated with the image information. In some embodiments, the key frame may be captured by cameras 44, 46 ( Fig. 9) is calculated as one or more key frames (e.g., key frames 1, 2). In some embodiments, a key frame may include a camera pose. In some embodiments, a key frame may include one or more camera images captured at a camera pose. In some embodiments, the XR system may determine that a portion of a camera image captured at a camera pose is useless and therefore does not include that portion in the key frame. Therefore, using key frames to align new images with early cognition of the scene can reduce the use of XR system computing resources. In some embodiments, a key frame may include an image and / or image data at a location with a direction / angle. In some embodiments, a key frame may include a location and direction in which one or more map points can be observed. In some embodiments, a key frame may include a coordinate frame with an ID. U.S. Patent Application No. 15 / 877,359 describes key frames and is incorporated herein by reference in its entirety.

[0272] Some or all of the keyframes 1140 may be selected for further processing, such as generating a persistent gesture for the keyframes 1150. The selection may be based on characteristics of all or a subset of the feature points in the image frame. These characteristics may be determined based on processing the descriptors, features, and / or the image frame itself. As a specific example, the selection may be based on clustering of feature points identified as being likely to be associated with a persistent object.

[0273] Each keyframe is associated with the pose of the camera that captured it. For keyframes selected for processing into persistent poses, this pose information can be saved along with other metadata about the keyframe, such as WiFi fingerprint and / or GPS coordinates at the time of capture and / or at the location of capture.

[0274] The persistent gesture is a source of information that the device can use to orient itself relative to previously acquired information about the physical world. For example, if the keyframe from which the persistent gesture was created is incorporated into a map of the physical world, the device can use a sufficient number of feature points in the keyframe associated with the persistent gesture to orient itself relative to the persistent gesture. The device can align a current image it takes of the surrounding environment with the persistent gesture. The alignment can be based on matching the current image with the image 1110, features 1120 and / or descriptors 1130 that caused the persistent gesture, or any subset of the image or those features or descriptors. In some embodiments, the current image frame matched with the persistent gesture can be another keyframe that has been incorporated into the tracking map of the device.

[0275] Information about persistent gestures can be stored in a format that facilitates sharing among multiple applications that may be executing on the same or different devices. Fig.14In an example of , some or all persistent gestures may be reflected as a persistent coordinate system (PCF) 1160. Like persistent gestures, a PCF may be associated with a map and may include a set of features or other information that a device may use to determine its orientation relative to the PCF. A PCF may include a transform that defines a transformation relative to the origin of its map, such that by associating its position with the PCF, a device may determine its position relative to any object in the physical world reflected in the map.

[0276] Because PCFs provide a mechanism for determining positions relative to physical objects, an application (e.g., application 1180) can define the positions of virtual objects relative to one or more PCFs, which serve as anchor points for virtual content 1170. For example, Fig.14 App 1 is shown to have associated its virtual content 2 with PCF 1.2. Similarly, App 2 has associated its virtual content 3 with PCF 1.2. App 1 is also shown to have associated its virtual content 1 with PCF 4.5, and App 2 is shown to have associated its virtual content 4 with PCF 3. In some embodiments, PCF 3 may be based on image 3 (not shown), and PCF 4.5 may be based on image 4 and image 5 (not shown), similar to how PCF 1.2 is based on image 1 and image 2. When rendering this virtual content, the device may apply one or more transformations to calculate information, such as the position of the virtual content relative to the device's display and / or the position of physical objects relative to the desired position of the virtual content. Using the PCF as a reference may simplify such calculations.

[0277] In some embodiments, a persistent gesture can be a coordinate position and / or direction with one or more associated keyframes. In some embodiments, a persistent gesture can be automatically created after the user has traveled a certain distance (e.g., three meters). In some embodiments, a persistent gesture can be used as a reference point during positioning. In some embodiments, a persistent gesture can be stored in a traversable world (e.g., traversable world module 538).

[0278] In some embodiments, a new PCF may be determined based on a predetermined distance allowed between adjacent PCFs. In some embodiments, when the user travels a predetermined distance (e.g., five meters), one or more persistent gestures may be calculated into the PCF. In some embodiments, the PCF may be associated with one or more world coordinate frames and / or canonical coordinate frames, such as in a navigable world. In some embodiments, the PCF may be stored in a local database and / or a remote database, depending on, for example, security settings.

[0279] Fig.15Method 4700 of establishing and using a persistent coordinate system according to some embodiments is shown. Method 4700 may begin by capturing (act 4702) an image of a scene (e.g., Fig.14 1 and 2 in FIG. 1 ). Multiple cameras may be used, and one camera may generate multiple images, for example in the form of a stream.

[0280] Method 4700 may include extracting (4704) points of interest (e.g., Figure 7 Map point 702 in Fig.14 1120), generating (action 4706) a descriptor of the extracted point of interest (e.g., Fig.14 1130 in ), and generates (act 4708) a keyframe (e.g., keyframe 1140) based on the descriptor. In some embodiments, the method may compare the points of interest in the keyframes and form pairs of keyframes that share a predetermined amount of points of interest. The method may use the respective keyframe pairs to reconstruct a portion of the physical world. The mapped portion of the physical world may be saved as a 3D feature (e.g., Figure 7 704 in the key assembly). In some embodiments, selected portions of the key frame pairs can be used to construct 3D features. In some embodiments, the results of the mapping can be selectively saved. Key frames that are not used to construct 3D features can be associated with 3D features by poses, for example, using a covariance matrix between the poses of the key frames to represent the distance between the key frames. In some embodiments, key frame pairs can be selected to construct 3D features so that the distance between each two of the constructed 3D features is within a predetermined distance, which can be determined to balance the amount of computation required and the level of accuracy of the resulting model. Such a method can provide the XR system with a model of the physical world with an amount of data suitable for efficient and accurate computation. In some embodiments, the covariance matrix of two images may include the covariance between the poses (e.g., six degrees of freedom) of the two images.

[0281] Method 4700 may include generating (action 4710) a persistent gesture based on the keyframes. In some embodiments, the method may include generating a persistent gesture based on a 3D feature reconstructed from a pair of keyframes. In some embodiments, the persistent gesture may be attached to the 3D feature. In some embodiments, the persistent gesture may include a gesture of a keyframe used to construct the 3D feature. In some embodiments, the persistent gesture may include an average gesture of the keyframes used to construct the 3D feature. In some embodiments, the persistent gesture may be generated so that the distance between adjacent persistent gestures is within a predetermined value, such as within a range of one meter to five meters, any value therebetween, or any other appropriate value. In some embodiments, the distance between adjacent persistent gestures may be represented by a covariance matrix of adjacent persistent gestures.

[0282] Method 4700 may include generating (action 4712) a PCF based on the persistent gesture. In some embodiments, the PCF may be attached to the 3D feature. In some embodiments, the PCF may be associated with one or more persistent gestures. In some embodiments, the PCF may include a pose of one of the associated persistent gestures. In some embodiments, the PCF may include an average pose of the poses of the associated persistent gestures. In some embodiments, the PCF may be generated so that the distance between adjacent PCFs is within a predetermined value, such as within a range of three meters to ten meters, any value therebetween, or any other appropriate value. In some embodiments, the distance between adjacent PCFs may be represented by a covariance matrix of adjacent PCFs. In some embodiments, the PCF may be exposed to an XR application via, for example, an application programming interface (API) so that the XR application can access a model of the physical world through the PCF without accessing the model itself.

[0283] Method 4700 may include associating image data of a virtual object to be displayed by the XR device with at least one of the PCFs (action 4714). In some embodiments, the method may include calculating a translation and orientation of the virtual object relative to the associated PCF. It should be understood that it is not necessary to associate the virtual object with a PCF generated by the device where the virtual object is placed. For example, the device may obtain a saved PCF in a canonical map in the cloud and associate the virtual object with the obtained PCF. It should be understood that as the PCF is adjusted over time, the virtual object may move with the associated PCF.

[0284] Fig.16 Visual data and algorithms of a first XR device 12 . 1 and a second XR device 12 . 2 and a server 20 are shown according to some embodiments. Fig.16 The components shown in may be operable to perform some or all of the operations associated with generating, updating, and / or using spatial information (such as, persistent poses, persistent coordinate systems, tracking maps, or canonical maps) as described herein. Although not shown, the first XR device 12.1 may be configured identically to the second XR device 12.2. The server 20 may have a map storage routine 118, a canonical map 120, a map transmitter 122, and a map merging algorithm 124.

[0285] The second XR device 12.2, which may be in the same scene as the first XR device 12.1, may include a persistent coordinate frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that may be used to render a virtual object, and a frame embedding generator 308 (see Fig.21 ). In some embodiments, the map download system 126, the PCF identification system 128, the map Figure 2 , positioning module 130, canonical map merger 132, canonical map 133, and map publisher 136 are collected as a traversable world unit 1304. The PCF integration unit 1300 can be connected to the traversable world unit 1304 and other components of the second XR device 12.2 to allow the acquisition, generation, use, upload, and download of PCF.

[0286] Maps including PCFs can achieve more persistence in a changing world. In some embodiments, locating a tracking map including matching features such as images can include selecting features representing persistent content from a map composed of PCFs, which enables fast matching and / or positioning. For example, in a world where people enter and exit a scene and objects such as doors move relative to the scene, less storage space and transmission rates are required, and the scene can be mapped using separate PCFs and their relationships to each other (e.g., an integrated constellation of PCFs).

[0287] In some embodiments, the PCF integrated unit 1300 may include a PCF 1306 previously stored in a data storage on a storage unit of the second XR device 12.2, a PCF tracker 1308, a persistent pose acquirer 1310, a PCF checker 1312, a PCF generation system 1314, a coordinate frame calculator 1316, a persistent pose calculator 1318, and three transformers including a tracking map and persistent pose transformer 1320, a persistent pose and PCF transformer 1322, and a PCF and image data transformer 1324.

[0288] In some embodiments, the PCF tracker 1308 may have an on prompt and a off prompt selectable by the application 1302. The application 1302 may be executable by a processor of the second XR device 12.2 to, for example, display virtual content. The application 1302 may have a call to turn on the PCF tracker 1308 via an on prompt. When the PCF tracker 1308 is on, the PCF tracker 1308 may generate a PCF. The application 1302 may have a subsequent call to turn off the PCF tracker 1308 via a off prompt. When the PCF tracker 1308 is off, the PCF tracker 1308 terminates the PCF generation.

[0289] In some embodiments, the server 20 may include a plurality of persistent gestures 1332 and a plurality of PCFs 1330 that have been previously saved in association with the canonical map 120. The map sender 122 may send the canonical map 120 together with the persistent gestures 1332 and / or the PCFs 1330 to the second XR device 12.2. The persistent gestures 1332 and the PCFs 1330 may be stored on the second XR device 12.2 in association with the canonical map 133. Figure 2When locating to the standard map 133, you can Figure 2 The persistent gesture 1332 and the PCF 1330 are stored in association.

[0290] In some embodiments, the persistent posture acquirer 1310 may acquire the Figure 2 The PCF checker 1312 may be connected to the persistent gesture acquirer 1310. The PCF checker 1312 may acquire a PCF from the PCF 1306 based on the persistent gesture acquired by the persistent gesture acquirer 1310. The PCF acquired by the PCF checker 1312 may form an initial set of PCFs for PCF-based image display.

[0291] In some embodiments, application 1302 may need to generate additional PCFs. For example, if the user moves to an area that has not been previously mapped, application 1302 may open PCF tracker 1308. PCF generation system 1314 may be connected to PCF tracker 1308 and generate additional PCFs as the map is created. Figure 2 Start expanding and start based on the ground Figure 2 Generate PCFs. The PCFs generated by the PCF generation system 1314 may form a second set of PCFs, which may be used for PCF-based image display.

[0292] The coordinate frame calculator 1316 may be connected to the PCF checker 1312. After the PCF checker 1312 obtains the PCF, the coordinate frame calculator 1316 may call the head coordinate frame 96 to determine the head pose of the second XR device 12.2. The coordinate frame calculator 1316 may also call the persistent pose calculator 1318. The persistent pose calculator 1318 may be directly or indirectly connected to the frame embedding generator 308. In some embodiments, an image / frame may be designated as a keyframe after traveling a threshold distance (e.g., 3 meters) from a previous keyframe. The persistent pose calculator 1318 may generate a persistent pose based on multiple (e.g., three) keyframes. In some embodiments, the persistent pose may be substantially an average of the coordinate frames of the multiple keyframes.

[0293] Tracking map and persistent posture changer 1320 can be connected to the ground Figure 2 and persistent posture calculator 1318. Tracking map and persistent posture converter 1320 can Figure 2 Transform to a persistent pose to determine relative to the ground Figure 2 The persistent pose at the origin of .

[0294] Persistent gesture and PCF transformer 1322 may be connected to tracking map and persistent gesture transformer 1320, and further connected to PCF checker 1312 and PCF generation system 1314. Persistent gesture and PCF transformer 1322 may transform the persistent gesture (to which the tracking map has been transformed) from PCF checker 1312 and PCF generation system 1314 into a PCF to determine a PCF relative to the persistent gesture.

[0295] PCF and image data transformer 1324 may be connected to persistent gesture and PCF transformer 1322 and data channel 62. PCF and image data transformer 1324 transforms the PCF into image data 68. Rendering engine 30 may be connected to PCF and image data transformer 1324 to display image data 68 to a user relative to the PCF.

[0296] The PCF integration unit 1300 may store additional PCFs generated using the PCF generation system 1314 in the PCF 1306. The PCF 1306 may be stored relative to the persistent posture. Figure 2 When the map publisher 136 obtains the PCF 1306 and the persistent gesture associated with the PCF 1306, the map publisher 136 also sends the map publisher 136 to the server 20. Figure 2 The associated PCF and persistent gesture. When the map storage routine 118 of the server 20 stores the map Figure 2 When the map storage routine 118 is used, the persistent gesture and PCF generated by the second viewing device 12.2 may also be stored. The map merging algorithm 124 may use the map associated with the canonical map 120 and stored in the persistent gesture 1332 and PCF 1330, respectively. Figure 2 The persistent pose and PCF are used to create a canonical map 120 .

[0297] The first XR device 12.1 may include a PCF integrated unit similar to the PCF integrated unit 1300 of the second XR device 12.2. When the map transmitter 122 sends the canonical map 120 to the first XR device 12.1, the map transmitter 122 may send a persistent gesture 1332 and a PCF 1330 associated with the canonical map 120 and originating from the second XR device 12.2. The first XR device 12.1 may store the PCF and the persistent gesture in a data store on a storage device of the first XR device 12.1. The first XR device 12.1 may then utilize the persistent gesture and PCF originating from the second XR device 12.2 for image display relative to the PCF. Additionally or alternatively, the first XR device 12.1 may obtain, generate, use, upload, and download the PCF and the persistent gesture in a manner similar to the second XR device 12.2 as described above.

[0298] In the example shown, the first XR device 12.1 generates a local tracking map (hereinafter referred to as a “map”). Figure 1 ”), and the map storage routine 118 receives the map from the first XR device 12.1 Figure 1 The map storage routine 118 then stores the map Figure 1 The specification map 120 is stored on the storage device of the server 20 .

[0299] The second XR device 12 . 2 includes a map download system 126 , an anchor point identification system 128 , a positioning module 130 , a canonical map merger 132 , a local content positioning system 134 , and a map publisher 136 .

[0300] In use, the map sender 122 sends the normative map 120 to the second XR device 12 . 2 , and the map download system 126 downloads and stores the normative map 120 from the server 20 as the normative map 133 .

[0301] The anchor point identification system 128 is connected to the world surface determination routine 78. The anchor point identification system 128 identifies anchor points based on objects detected by the world surface determination routine 78. The anchor point identification system 128 generates a second map (map) using the anchor points. Figure 2 As shown in loop 138, the anchor point identification system 128 continues to identify anchor points and continues to update the ground Figure 2 The positions of the anchor points are recorded as three-dimensional data based on data provided by the world surface determination routine 78 . The world surface determination routine 78 receives images from the real object detection camera 44 and depth data from the depth sensor 135 to determine the positions of surfaces and their relative distances from the depth sensor 135 .

[0302] The positioning module 130 is connected to the specification map 133 and the Figure 2 The positioning module 130 repeatedly attempts to Figure 2 The standard map 133 is located. The standard map merger 132 is connected to the standard map 133 and the map Figure 2 When the positioning module 130 sets the ground Figure 2 When the standard map 133 is located, the standard map merger 132 merges the standard map 133 into the map. Figure 2 Then, the map is updated with the missing data included in the canonical map. Figure 2 .

[0303] The local content location system 134 is connected to the ground Figure 2 The local content location system 134 may be, for example, a system in which a user can locate local content at a specific location within a world coordinate frame. The local content then attaches itself to the local content. Figure 2The local to world coordinate converter 104 converts the local coordinate frame into the world coordinate frame based on the settings of the local content positioning system 134. Figure 2 The functionality of the rendering engine 30, display system 42, and data pipeline 62 are described.

[0304] Map publisher 136 will Figure 2 Upload to the server 20. The map storage routine 118 of the server 20 then Figure 2 Stored in the storage medium of the server 20.

[0305] Map merging algorithm 124 Figure 2 Merge with the canonical map 120. When more than two maps have been stored (e.g., three or four maps relating to the same or adjacent areas of the physical world), a map merging algorithm 124 merges all of the maps into the canonical map 120 to render a new canonical map 120. The map sender 122 then sends the new canonical map 120 to any and all devices 12.1 and 12.2 located in the area represented by the new canonical map 120. When the devices 12.1 and 12.2 locate their respective maps to the canonical map 120, the canonical map 120 becomes the upgraded map.

[0306] Fig.17 An example of generating keyframes for a map of a scene according to some embodiments is illustrated. In the example shown, a first keyframe KF1 is generated for a door on the left wall of a room. A second keyframe KF2 is generated for a corner area where the floor, left wall, and right wall of the room intersect. A third keyframe KF3 is generated for a window area on the right wall of the room. A fourth keyframe KF4 is generated on the floor of the wall, at the far end of the carpet. A fifth keyframe KF5 is generated for an area of ​​the carpet closest to the user.

[0307] Fig.18 According to some embodiments, Fig.17 An example of a persistent gesture generated for a map of the device. In some embodiments, a new persistent gesture is created when the device measures a threshold distance traveled, and / or when an application requests a new persistent gesture (PP). In some embodiments, the threshold distance can be 3 meters, 5 meters, 20 meters, or any other suitable distance. Selecting a smaller threshold distance (e.g., 1m) can result in an increase in computational load because a larger number of PPs can be created and managed compared to a larger threshold distance. Selecting a larger threshold distance (e.g., 40m) can result in increased virtual content placement errors because a smaller number of PPs will be created, which will result in a smaller number of PCFs created, meaning that virtual content attached to a PCF may be a relatively large distance away from the PCF (e.g., 30m), and the error increases with increasing distance from the PCF to the virtual content.

[0308] In some embodiments, a PP can be created when a new session begins. This initial PP can be considered to be zero and can be visualized as the center of a circle with a radius equal to the threshold distance. When the device reaches the circumference of the circle, and in some embodiments, an application requests a new PP, the new PP can be placed at the device's current location (at the threshold distance). In some embodiments, if the device is able to find an existing PP within the threshold distance from the device's new location, a new PP is not created at the threshold distance. In some embodiments, when a new PP is created (e.g., Fig.14 In some embodiments, the device may create a PP 1150 in the image, and the device may attach one or more closest key frames to the PP. In some embodiments, the location of the PP relative to the key frames may be based on the location of the device when the PP is created. In some embodiments, a PP will not be created when the device travels a threshold distance unless the application requests a PP.

[0309] In some embodiments, when an application has virtual content to display to the user, the application can request a PCF from the device. A PCF request from the application can trigger a PP request, and a new PP will be created after the device travels a threshold distance. Fig.18 A first persistent pose PP1 is shown, which may have closest keyframes (eg, KF1, KF2, and KF3) attached by, for example, calculating relative poses between the keyframes and the persistent pose. Fig.18 Also shown is a second permanent pose PP2, which may have additional proximal keyframes (eg, KF4 and KF5).

[0310] Fig.19 According to some embodiments, Fig.17 An example of a map generating PCF. In the example shown, PCF1 may include PP1 and PP2. As described above, the PCFs may be used to display image data associated with the PCFs. In some embodiments, each PCF may have coordinates in another coordinate frame (e.g., a world coordinate frame) and a PCF descriptor, e.g., to uniquely identify the PCF. In some embodiments, the PCF descriptor may be calculated based on feature descriptors of features in a frame associated with the PCF. In some embodiments, various constellations of PCFs may be combined to represent the real world in a persistent manner that requires less data and less data transmission.

[0311] Figures 20A to 20C is a schematic diagram showing an example of establishing and using a persistent coordinate system. Fig. 20ATwo users 4802A, 4802B are shown with respective local tracking maps 4804A, 4804B that have not yet been localized to a canonical map. The origin 4806A, 4806B of each user is depicted by a coordinate system (e.g., a world coordinate system) in their respective regions. These origins of each tracking map may be local to each user, as they depend on the orientation of their respective devices when tracking is initiated.

[0312] When the user device's sensors scan the environment, the device can capture Fig.14 The images described may contain features representing persistent objects, such that those images may be classified as keyframes from which persistent gestures may be created. In this example, track map 4802A includes persistent gestures (PP) 4808A; track 4802B includes PP 4808B.

[0313] Similarly, as above combined Fig.14 As described, some PPs may be classified as PCFs, which are used to determine the orientation of virtual content to render it to the user. Fig. 20B It is shown that the XR devices worn by the respective users 4802A, 4802B can create local PCFs 4810A, 4810B based on PPs 4808A, 4808B. Fig. 20C It is shown that persistent content 4812A, 4812B (e.g., virtual content) can be attached to the PCF 4810A, 4810B through corresponding XR devices.

[0314] In this example, the virtual content may have a virtual content coordinate frame that may be used by the application generating the virtual content regardless of how the virtual content should be displayed. For example, the virtual content may be specified as surfaces, such as triangles of a mesh, at specific positions and angles relative to the virtual content coordinate frame. In order to render the virtual content to the user, the positions of those surfaces may be determined relative to the user who is to perceive the virtual content.

[0315] Attaching virtual content to a PCF can simplify the computations involved in determining the position of the virtual content relative to the user. The position of the virtual content relative to the user can be determined by applying a series of transforms. Some of these transforms may change and may be updated frequently. Others of these transforms may be stable and may be updated frequently or not updated at all. Regardless, the transforms can be applied with a relatively low computational burden so that the position of the virtual content can be updated frequently relative to the user, thereby providing a realistic appearance to the rendered virtual content.

[0316] exist FIG. 20A to FIG. 20CIn the example of , user 1's device has a coordinate system that is related to a coordinate system that defines a map origin via a transformation rig1_T_w1. User 2's device has a similar transformation rig2_T_w2. These transformations can be expressed as 6 degrees of transformation, specifying a translation and a rotation to align the device coordinate system with the map coordinate system. In some embodiments, the transformation can be expressed as two separate transformations, one specifying a translation and the other specifying a rotation. Therefore, it should be understood that the transformations can be expressed in a form that simplifies computations or otherwise provides advantages.

[0317] The transformations between the origin of the tracking map and the PCF identified by the corresponding user equipment are denoted as pcf1_T_w1 and pcf2_T_w2. In this example, the PCF and the PP are the same, so the same transformation also characterizes the PP.

[0318] Therefore, the position of the user equipment relative to the PCF can be calculated by serial application of these transformations, for example rig1_T_pcf1=(rig1_T_w1)*(pcf1_T_w1).

[0319] like Fig. 20C As shown, the virtual content is positioned relative to the PCF through the transformation of obj1_T_pcf1. This transformation can be set by the application generating the virtual content, which can receive information describing physical objects relative to the PCF from the world reconstruction system. In order to render the virtual content to the user, the transformation to the coordinate system of the user's device is calculated, which can be calculated by associating the virtual content coordinate frame to the origin of the tracking map through the transformation obj1_t_w1 = (obj1_T_pcf1) * (pcf1_T_w1). This transformation can then be related to the user's device through a further transformation rig1_T_w1.

[0320] Based on the output from the application generating the virtual content, the position of the virtual content may change. When it does, the end-to-end transformation from the source coordinate system to the destination coordinate system may be recalculated. Additionally, the user's position and / or head pose may change as the user moves. As a result, the transformation rig1_T_w1 may change, as may any end-to-end transformations that depend on the user's position or head pose.

[0321] The transformation rig1_T_w1 may be updated with the user's movement based on tracking the user's position relative to stationary objects in the physical world. Such tracking may be performed by a head pose tracking component or other component of the system that processes the image sequence as described above. Such updating may be performed by determining the user's pose relative to a fixed reference frame (e.g., PP).

[0322] In some embodiments, since the PP is used as the PCF, the position and orientation of the user device can be determined relative to the most recent persistent gesture, or in this example, the PCF. Such a determination can be made by identifying feature points that characterize the PP in the current image captured using sensors on the device. Using image processing techniques such as stereo image analysis, the position of the device relative to those feature points can be determined. From this data, the system can calculate the change in the transform associated with the user's motion based on the relationship rig1_T_pcf1=(rig1_T_w1)*(pcf1_T_w1).

[0323] The system can determine and apply transformations in a computationally efficient order. For example, the need to calculate rig1_T_w1 from measurements that generate rig1_T_pcf1 can be avoided by tracking user gestures and defining the position of virtual content relative to a PP or PCF constructed based on persistent gestures. In this way, the transformation from the source coordinate system of the virtual content to the destination coordinate system of the user device can be based on a measured transformation according to the expression (rig1_T_pcf1)*(obj1_t_pcf1), where the first transformation is measured by the system and the latter transformation is provided by the application that specifies the virtual content to be rendered. In embodiments where the virtual content is positioned relative to the origin of the map, the end-to-end transformation can relate the virtual object coordinate system to the PCF coordinate system based on a further transformation between map coordinates and PCF coordinates. In embodiments where the virtual content is positioned relative to a PP or PCF different from the PP or PCF for which the user's position is tracked, a transformation can be performed between the two. Such a transformation can be fixed and can be determined, for example, from a map where both appear.

[0324] For example, a transformation-based approach can be implemented in a device having components that process sensor data to build a tracking map. As part of this process, these components can identify feature points that can be used as persistent gestures, which in turn can be turned into PCFs. These components can limit the number of persistent gestures generated for the map to provide appropriate spacing between persistent gestures while allowing the user to be close enough to the persistent gesture location regardless of the location in the physical environment to accurately calculate the user's gesture, such as the above combined with Figures 17 to 19 As the persistent gesture closest to the user is updated, as the user moves, any transformations used to calculate the position of the virtual content relative to the user that depend on the position of the PP (or PCF, if used) can be updated and stored for use, at least until the user leaves the persistent gesture. Nevertheless, by calculating and storing the transformations, the computational burden of each update of the position of the virtual content can be relatively low, so that it can be performed with relatively low latency.

[0325] FIG. 20A to FIG. 20C Positioning is shown relative to a tracking map, and each device has its own tracking map. However, transforms can be generated relative to any map coordinate system. Content persistence between user sessions of an XR system can be achieved through the use of a persistent map. Shared experiences for users can also be achieved through the use of a map that multiple user devices can be directed to.

[0326] In some embodiments described in more detail below, the location of virtual content can be specified relative to coordinates in a canonical map, which is formatted so that any of multiple devices can use the map. Each device may maintain a tracking map, and changes in the user's posture relative to the tracking map can be determined. In this example, the transformation between the tracking map and the canonical map can be determined by a "localization" process, which can be performed by matching structures in the tracking map (such as one or more persistent postures) to one or more structures of the canonical map (e.g., one or more PCFs).

[0327] Techniques for creating and using canonical maps in this manner are described in more detail below.

[0328] Depth Keyframes

[0329] The techniques described herein rely on comparison of image frames. For example, to establish the position of a device relative to a tracking map, a new image may be captured using a sensor worn by a user, and the XR system may search the set of images used to create the tracking map for an image that shares at least a predetermined number of points of interest with the new image. As an example of another scenario involving image frame comparison, a tracking map can be localized to a canonical map by first finding an image frame in the tracking map associated with a persistent gesture that is similar to an image frame in the canonical map associated with a PCF. Alternatively, a transformation between two canonical maps can be computed by first finding similar image frames in the two maps.

[0330] Depth keyframes provide a method to reduce the amount of processing required to identify similar image frames. For example, in some embodiments, the comparison can be between image features (e.g., "2D features") in a new 2D image and 3D features in a map. This comparison can be performed in any appropriate manner, such as by projecting the 3D image into a 2D plane. Conventional methods such as Bag of Words (BoW) search for 2D features of the new image in a database that includes all 2D features in the map, which can require a lot of computing resources, especially when the map represents a large area. Conventional methods then locate images that share at least one 2D feature with the new image, which may include images that are not useful for locating meaningful 3D features in the map. Conventional methods then locate 3D features that are meaningless relative to the 2D features in the new image.

[0331] The inventors have recognized and understood techniques for retrieving images in a map using fewer memory resources (e.g., one-quarter of the memory resources used by BoW), higher efficiency (e.g., 2.5 ms processing time per keyframe and 100 μs for comparison of 500 keyframes), and higher accuracy (e.g., 20% better retrieval recall than BoW for a 1024-dimensional model and 5% better retrieval recall than BoW for a 256-dimensional model).

[0332] To reduce computation, a descriptor may be computed for an image frame that may be used to compare the image frame to other image frames. The descriptors may be stored instead of or in addition to the image frames and feature points. In a map in which persistent gestures and / or PCFs may be generated from image frames, descriptors of one or more image frames from which each persistent gesture or PCF was generated may be stored as part of the persistent gesture and / or PCF.

[0333] In some embodiments, the descriptor may be calculated from feature points in an image frame. In some embodiments, the neural network is configured to calculate a unique frame descriptor representing an image. The image may have a resolution greater than 1 megabyte, thereby capturing sufficient detail in the image of a 3D environment within the field of view of a device worn by a user. The frame descriptor may be much shorter, such as a string of numbers, for example, in the range of 128 bytes to 512 bytes or any number in between.

[0334] In some embodiments, the neural network is trained to compute frame descriptors that indicate similarities between images. Images in the map may be located by identifying, in a database including images used to generate the map, the nearest image that may have a frame descriptor within a predetermined distance from the frame descriptor of the new image. In some embodiments, the distance between images may be represented by a difference between the frame descriptors of the two images.

[0335] Fig.21 is a block diagram illustrating a system for generating descriptors for individual images according to some embodiments. In the example shown, a frame embedding generator 308 is shown. In some embodiments, frame embedding generator 308 may be used within server 20, but may alternatively or additionally be executed in whole or in part in one of XR devices 12.1 and 12.2, or any other device that processes images for comparison with other images.

[0336] In some embodiments, the frame embedding generator can be configured to generate a reduced data representation of an image from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes) that is indicative of the content in the image despite the reduced size. In some embodiments, the frame embedding generator can be used to generate a data representation of an image, which can be a keyframe or a frame used in other ways. In some embodiments, the frame embedding generator 308 can be configured to convert an image at a particular location and orientation into a unique string of numbers (e.g., 256 bytes). In the example shown, an image 320 captured by an XR device can be processed by a feature extractor 324 to detect points of interest 322 in the image 320. The points of interest may or may not be derived from feature points identified as described above for features 1120 ( Fig.14 ) or as otherwise described herein. In some embodiments, the point of interest may be represented by a descriptor 1130 ( Fig.14 ) are represented by the descriptors described in the above description, and these points of interest can be generated using the deep sparse feature method. In some embodiments, each point of interest 322 can be represented by a digital string (e.g., 32 bytes). For example, there can be n features (e.g., 100), and each feature is represented by a 32-byte string.

[0337] In some embodiments, the frame embedding generator 308 may include a neural network 326. The neural network 326 may include a multi-layer perceptron unit 312 and a max pooling unit 314. In some embodiments, the multi-layer perceptron (MLP) unit 312 may include a multi-layer perceptron, which may be trained. In some embodiments, the points of interest 322 (e.g., descriptors for the points of interest) may be reduced by the multi-layer perceptron 312 and may be output as a weighted combination of descriptors 310. For example, the MLP may reduce n features to m features, which is less than n features.

[0338] In some embodiments, the MLP unit 312 can be configured to perform matrix multiplication. The multilayer perceptron unit 312 receives a plurality of points of interest 322 of an image 320 and converts each point of interest into a corresponding string of numbers (e.g., 256). For example, there may be 100 features, and each feature may be represented by a string of 256 numbers. In this example, a matrix with 100 horizontal rows and 256 vertical columns may be created. Each row may have a series of 256 numbers that vary in size, some being smaller and others being larger. In some embodiments, the output of the MLP may be an n×256 matrix, where n represents the number of points of interest extracted from the image. In some embodiments, the output of the MLP may be an m×256 matrix, where m is the number of points of interest reduced from n.

[0339] In some embodiments, the MLP 312 may have a training phase during which model parameters for the MLP are determined and a use phase. Fig.25 The MLP can be trained as shown in . The input training data may include a set of three data, each set of three including 1) query image, 2) positive sample, and 3) negative sample. The query image can be considered as a reference image.

[0340] In some embodiments, the positive sample may include an image that is similar to the query image. For example, in some embodiments, the similarity may be having the same object in the query image and the positive sample image, but viewed from different angles. In some embodiments, the similarity may be having the same object in the query image and the positive sample image, but the object is shifted relative to the other image (e.g., left, right, up, down).

[0341] In some embodiments, negative samples may include images that are dissimilar to the query image. For example, in some embodiments, a dissimilar image may not contain any objects that are prominent in the query image, or may contain only a small portion (e.g., <10%, 1%) of the prominent objects in the query image. In contrast, for example, a similar image may have a large portion (e.g., >50% or >75%) of the objects in the query image.

[0342] In some embodiments, the focus points can be extracted from the images in the input training data and converted into feature descriptors. Fig.25 The training images and Fig.21 These descriptors are calculated using both features extracted from the operation of the frame embedding generator 308. In some embodiments, as described in U.S. Patent Application 16 / 190,948, deep sparse feature (DSF) processing can be used to generate descriptors (e.g., DSF descriptors). In some embodiments, the DSF descriptors are n×32 dimensions. The descriptors can then be passed through the model / MLP to create a 256-byte output. In some embodiments, the model / MLP can have the same structure as the MLP 312, so that once the model parameters are set through training, the resulting trained MLP can be used as the MLP 312.

[0343] In some embodiments, the feature descriptors (e.g., 256 bytes output from the MLP model) can then be sent to a triple boundary loss module (which can be used only during the training phase and not during the use phase of the MLP neural network). In some embodiments, the triple boundary loss module can be configured to select parameters of the model to reduce the difference between the 256-byte output from the query image and the 256-byte output from the positive samples, and to increase the 256-byte output from the query image and the 256-byte output from the negative samples. In some embodiments, the training phase can include feeding multiple triple input images into the learning process to determine the model parameters. The training process can continue, for example, until the difference for positive images is minimized and the difference for negative images is maximized, or until other appropriate exit criteria are met.

[0344] Reference again Fig.21, the frame embedding generator 308 may include a pooling layer, shown here as a max pooling unit 314. The max pooling unit 314 may analyze each column to determine the maximum number in the corresponding column. The max pooling unit 314 may combine the maximum values ​​of the numbers in each column of the output matrix of the MLP 312 into a global feature string 316 of, for example, 256 numbers. It should be understood that images processed in an XR system may be expected to have high-resolution frames, potentially with millions of pixels. The global feature string 316 is a relatively small number that takes up relatively little memory and is easy to search compared to an image (e.g., having a resolution of more than 1 megabyte). The image can therefore be searched without analyzing each raw frame from the camera, and it is also cheaper to store 256 bytes instead of a full frame.

[0345] Fig. 22 2200 . The method 2200 may begin by receiving (act 2202 ) a plurality of images captured by an XR device worn by a user. In some embodiments, the method 2200 may include determining (act 2204 ) one or more keyframes from the plurality of images. In some embodiments, act 2204 may be skipped and / or may instead occur after step 2210 .

[0346] Method 2200 may include: identifying (act 2206) one or more points of interest in a plurality of images using an artificial neural network; and computing (act 2208) feature descriptors for the respective points of interest using the artificial neural network. The method may include computing (act 2210) a frame descriptor for each image, thereby representing the image based at least in part on the feature descriptors computed for the points of interest identified in the image using the artificial neural network.

[0347] Fig.232300 is a flowchart illustrating a method 2300 for positioning using image descriptors according to some embodiments. In this example, a new image frame describing the current location of the XR device can be compared to an image frame stored in conjunction with a point in a map (e.g., a persistent gesture or PCF as described above). Method 2300 can begin by receiving (action 2302) a new image captured by an XR device worn by a user. Method 2300 may include identifying (action 2304) one or more recent keyframes in a database that includes keyframes for generating one or more maps. In some embodiments, the most recent keyframe can be identified based on coarse spatial information and / or previously determined spatial information. For example, the coarse spatial information may indicate that the XR device is located in a geographic area represented by a 50m×50m area of ​​the map. Image matching can be performed only for points within the area. As another example, based on tracking, the XR system may know that the XR device was previously close to a first persistent gesture in the map and was moving in the direction of a second persistent gesture in the map at the time. The second persistent gesture can be considered to be the most recent persistent gesture, and the keyframe stored with it can be considered to be the most recent keyframe. Alternatively or additionally, other metadata such as GPS data or WiFi fingerprints may be used to select the most recent keyframe or a set of most recent keyframes.

[0348] Regardless of how the nearest keyframe is selected, the frame descriptors may be used to determine whether the new image matches any of the frames selected to be associated with a nearby persistent gesture. This determination may be made by comparing the frame descriptor of the new image to the frame descriptors of the nearest keyframe, or to the frame descriptors of a subset of keyframes in a database selected in any other suitable manner, and selecting a keyframe having a frame descriptor within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors may be calculated by taking the difference between two numeric strings that may represent the two frame descriptors. In embodiments where the strings are processed as multiple strings, the difference may be calculated as a vector difference.

[0349] Once a matching image frame is identified, the orientation of the XR device relative to the image frame can be determined. Method 2300 may include performing (action 2306) feature matching on 3D features in the map corresponding to the identified nearest keyframe, and calculating (action 2308) a pose of a device worn by the user based on the feature matching results. In this way, computationally dense matching of feature points in two images can be performed for as few as one image that has been determined to be a possible match to the new image.

[0350] Fig.242400 . The method 2400 may begin by generating (act 2402) a data set including a plurality of image sets. Each of the plurality of image sets may include a query image, a positive sample image, and a negative sample image. In some embodiments, the plurality of image sets may include synthetic record pairs configured to, for example, teach the neural network basic information (such as shape). In some embodiments, the plurality of image sets may include real record pairs, which may be based on physical world records.

[0351] In some embodiments, inliers may be computed by fitting a fundamental matrix between the two images. In some embodiments, sparse overlap may be computed as the intersection over union (IoU) of the points of interest seen in the two images. In some embodiments, a positive sample may include at least twenty of the same points of interest as in the query image as inliers. A negative sample may include less than ten inliers. A negative sample may have less than half of the sparse points that overlap with the sparse points of the query image.

[0352] The method 2400 may include calculating (act 2404) a loss for each image set by comparing the query image with the positive sample images and the negative sample images. The method 2400 may include modifying (act 2406) the artificial neural network based on the calculated loss so that the distance between the frame descriptor generated by the artificial neural network for the query image and the frame descriptor for the positive sample images is less than the distance between the frame descriptor for the query image and the frame descriptor for the negative sample images.

[0353] It should be understood that although the above describes methods and apparatus configured to generate global descriptors for individual images, the methods and apparatus may be configured to generate descriptors for individual maps. For example, a map may include multiple keyframes, each of which may have a frame descriptor as described above. The maximum pooling unit may analyze the frame descriptors of the keyframes of a map and combine the frame descriptors into a unique map descriptor for that map.

[0354] Furthermore, it should be appreciated that other architectures may be used for the processing described above. For example, separate neural networks for generating DSF descriptors and frame descriptors are described. This approach is computationally efficient. However, in some embodiments, frame descriptors may be generated from selected feature points without first generating DSF descriptors.

[0355] Ranking and merging maps

[0356] Described herein are methods and apparatus for ranking and merging multiple environmental maps in a cross-reality system. Map merging can enable maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking maps can enable efficient execution of the techniques described herein, including map merging, which involves selecting maps from a set of maps based on similarity. In some embodiments, for example, a system may maintain a set of canonical maps that are formatted in a manner that any of a number of XR devices can access them. These canonical maps can be formed by merging selected tracking maps from those devices with other tracking maps or previously stored canonical maps. For example, canonical maps can be ranked for selecting one or more canonical maps to merge with a new tracking map and / or selecting one or more canonical maps from a set for use within a device.

[0357] In order to provide a realistic XR experience to the user, the XR system must understand the user's physical environment in order to correctly relate the positions of virtual objects to real objects. Information about the user's physical environment can be obtained from an environment map of the user's location.

[0358] The inventors have recognized and appreciated that an XR system can provide an enhanced XR experience to multiple users sharing the same world including real and / or virtual content by enabling efficient sharing of environmental maps of the real / physical world collected by multiple users, whether the users are present in the world at the same time or at different times. However, there are significant challenges in providing such a system. Such a system may store multiple maps generated by multiple users and / or the system may store multiple maps generated at different times. For operations that may be performed using previously generated maps (such as, for example, positioning, such as described above), a significant amount of processing may be required to identify relevant environmental maps of the same world (e.g., the same real-world location) from all environmental maps collected by the XR system. In some embodiments, there may be a small number of environmental maps that a device can access, such as for positioning. In some embodiments, there may be a large number of environmental maps that a device can access. The inventors have recognized and appreciated techniques for quickly and accurately ranking the relevance of environmental maps from all possible environmental maps, such as, for example Fig.28 The full domain of all canonical maps 120 in the user's image. Highly ranked maps can then be selected for further processing, such as rendering virtual objects on the user's display to realistically interact with the physical world around the user, or merging map data collected by the user with stored maps to create a larger or more accurate map.

[0359] In some embodiments, a stored map related to a user's task at a location in the physical world can be identified by filtering the stored maps based on multiple criteria. These criteria can indicate a comparison of a tracking map generated by the user's wearable device in the location with a candidate environment map stored in a database. The comparison can be performed based on metadata associated with the map, such as a Wi-Fi fingerprint detected by the device generating the map and / or a set of BSSIDs to which the device is connected while forming the map. The comparison can also be performed based on compressed or uncompressed content of the map. Comparisons based on compressed representations can be performed by comparing vectors calculated from the map content. For example, a comparison based on an uncompressed map can be performed by locating a tracking map within a stored map, or vice versa. Multiple comparisons can be performed in sequence based on the computation time required to reduce the number of candidate maps to be considered, where comparisons involving less computation will be performed earlier in the sequence than other comparisons that require more computation.

[0360] Fig.26 An AR system 800 configured to rank and merge one or more environment maps according to some embodiments is depicted. The AR system may include a traversable world model 802 of the AR device. Information to populate the traversable world model 802 may come from sensors on the AR device, which may include data stored in a processor 804 (e.g., Figure 4 The processor may perform some or all of the processing to convert the sensor data into a map. Such a map may be a tracking map, because the AR device may build a tracking map while collecting sensor data as it operates in an area. Figure 1 In addition, regional attributes may be provided to indicate the area represented by the tracking map. These regional attributes may be geographic location identifiers, such as coordinates expressed as latitude and longitude, or IDs used by the AR system to represent locations. Alternatively or in addition, regional attributes may be measured characteristics that have a high probability of being unique to the area. Regional attributes may, for example, be derived from parameters of wireless networks detected in the area. In some embodiments, regional attributes may be associated with a unique address of an access point that the AR system is nearby and / or connected to. For example, regional attributes may be associated with a MAC address or basic service set identifier (BSSID) of a 5G base station / router, a Wi-Fi router, or the like.

[0361] exist Fig.26 In the example of FIG. 8 , the tracking map can be merged with other maps of the environment. The map ranking portion 806 receives the tracking map from the device PW 802 and communicates with the map database 808 to select and rank the environment map from the map database 808. The selected map with a higher ranking is sent to the map merging portion 810.

[0362] The map merging section 810 may perform a merging process on the maps sent from the map ranking section 806. The merging process may entail merging the tracking map with some or all of the ranking maps and sending the new merged map to the traversable world model 812. The map merging section may merge the maps by identifying maps that depict overlapping portions of the physical world. Those overlapping portions may be aligned so that the information in the two maps may be aggregated into the final map. The canonical map may be merged with other canonical maps and / or tracking maps.

[0363] Aggregation may require extending one map with information from another map. Alternatively or additionally, aggregation may require adjusting the representation of the physical world in one map based on information in another map. For example, the later map may reveal objects that caused feature points to move, so that the map can be updated based on the later information. Alternatively, two maps may characterize the same area with different feature points, and aggregation may require selecting a set of feature points from both maps to better represent the area. Regardless of the specific processing that occurs during the merge, in some embodiments, PCFs from all maps being merged may be retained so that applications that locate content relative to them can continue to do so. In some embodiments, the merging of maps may result in redundant persistent gestures, and some persistent gestures may be deleted. When a PCF is associated with a persistent gesture that is to be deleted, merging the maps may require modifying the PCF to be associated with the persistent gesture that remains in the map after the merge.

[0364] In some embodiments, as maps are expanded and / or updated, they may be refined. Refinement may require calculations to reduce internal inconsistencies between feature points that may represent the same object in the physical world. Such inconsistencies may arise from pose inaccuracies associated with keyframes that provide feature points that represent the same object in the physical world. For example, such inconsistencies may arise from the XR device calculating a pose relative to a tracking map, which is in turn built based on an estimated pose, so that errors in the pose estimate accumulate, causing "drift" in pose accuracy over time. Maps can be refined by performing bundle adjustment or other operations to reduce inconsistencies in feature points from multiple keyframes.

[0365] During refinement, the position of a persistent point relative to the map origin may change. As a result, the transforms associated with that persistent point, such as a persistent pose or PCF, may change. In some embodiments, an XR system in conjunction with map refinement (whether performed as part of a merge operation or for other reasons) may recalculate the transforms associated with any persistent points that have changed. These transforms may be pushed from the component that computes the transform to the component that uses the transform so that any use of the transform may be based on the updated position of the persistent point.

[0366] The traversable world model 812 may be a cloud model that may be shared by multiple AR devices. The traversable world model 812 may store or otherwise access a map of the environment in the map database 808. In some embodiments, when a previously calculated map of the environment is updated, a previous version of the map may be deleted so that outdated maps are removed from the database. In some embodiments, when a previously calculated map of the environment is updated, a previous version of the map may be archived so that a previous version of the environment can be retrieved / viewed. In some embodiments, permissions may be set so that only an AR system with certain read / write access rights can trigger the deletion / archiving of a previous version of the map.

[0367] These environment maps created from tracking maps provided by one or more AR devices / systems can be accessed by AR devices in the AR system. The map ranking portion 806 can also be used to provide environment maps to AR devices. An AR device can send a message requesting an environment map for its current location, and the map ranking portion 806 can be used to select and rank environment maps relevant to the requesting device.

[0368] In some embodiments, the AR system 800 may include a downsampling unit 814 configured to receive a merged map from the cloud PW 812. The merged map received from the cloud PW 812 may be a storage format for the cloud, which may include high-resolution information, such as a large number of PCFs or multiple image frames per square meter or a large number of feature point sets associated with the PCF. The downsampling unit 814 may be configured to downsample the cloud format map to a format suitable for storage on the AR device. The device-formatted map may contain less data, such as fewer PCFs or less data stored for each PCF, to accommodate the limited local computing power and storage space of the AR device.

[0369] Fig. 271 is a simplified block diagram showing a plurality of canonical maps 120 that may be stored in a remote storage medium, such as a cloud. Each canonical map 120 may include a plurality of canonical map identifiers that indicate the location of the canonical map in a physical space, such as somewhere on the earth. The canonical map identifiers may include one or more of the following identifiers: a region identifier represented by a longitude and latitude range, a frame descriptor (e.g., Fig.21 ), Wi-Fi fingerprints, feature descriptors (e.g., Fig.21 ), and device identifiers indicating one or more devices contributing to the map.

[0370] In the example shown, the canonical maps 120 are geographically arranged in a two-dimensional pattern as they may exist on the surface of the earth. The canonical maps 120 may be uniquely identifiable by corresponding longitudes and latitudes, as any canonical maps with overlapping longitudes and latitudes may be merged into a new canonical map.

[0371] Fig.28 is a schematic diagram illustrating a method of selecting a canonical map according to some embodiments, which method can be used to locate a new tracking map to one or more canonical maps. The method can begin by accessing (act 120) a world of canonical maps 120, which, as an example, can be stored in a database of traversable worlds (e.g., traversable world module 538). The world of canonical maps can include canonical maps from all previously visited locations. The XR system can filter the world of all canonical maps to a small subset or just one map. It should be understood that in some embodiments, due to bandwidth limitations, all canonical maps may not be sent to the viewing device. Selecting a subset that is selected as a possible candidate for matching tracking maps to send to the device can reduce bandwidth and latency associated with accessing a remote map database.

[0372] The method may include filtering (act 300) the universe of the canonical map based on regions having a predetermined size and shape. Fig. 27In the example shown, each square can represent an area. Each square can cover 50m×50m. Each square can have six adjacent areas. In some embodiments, action 300 can select at least one matching specification map 120 covering longitude and latitude, where the longitude and latitude include the longitude and latitude of the location identifier received from the XR device, as long as there is at least one map at the longitude and latitude. In some embodiments, action 300 can select at least one adjacent specification map covering longitude and latitude adjacent to the matching specification map. In some embodiments, action 300 can select multiple matching specification maps and multiple adjacent specification maps. Action 300 can, for example, reduce the number of specification maps by approximately ten times, such as from thousands to hundreds, to form a first filtering selection. Alternatively or additionally, adjacent maps can be identified using criteria other than latitude and longitude. For example, the XR device may have previously been positioned using a specification map in the collection as part of the same session. The cloud service can retain information about the XR device, including previously positioned maps. In this example, the maps selected at act 300 may include those maps that cover an area adjacent to the map at which the XR device is positioned.

[0373] The method may include filtering (act 302) a first filtered selection of canonical maps based on a Wi-Fi fingerprint. Act 302 may determine a latitude and longitude based on a Wi-Fi fingerprint received from the XR device as part of a location identifier. Act 302 may compare the latitude and longitude from the Wi-Fi fingerprint with the latitude and longitude of the canonical map 120 to determine one or more canonical maps that form a second filtered selection. Act 302 may reduce the number of canonical maps by approximately ten times, for example, from hundreds of canonical maps to dozens (e.g., 50) of canonical maps that form a second selection. For example, the first filtered selection may include 130 canonical maps, the second filtered selection may include 50 of the 130 canonical maps, and may not include the other 80 of the 130 canonical maps.

[0374] The method may include filtering (act 304) a second filter selection of the canonical map based on the keyframes. Act 304 may compare data representing an image captured by the XR device with data representing the canonical map 120. In some embodiments, the data representing the image and / or the map may include feature descriptors (e.g., Fig.25 DSF descriptor in ) and / or global feature string (e.g. Fig.21316 in). Action 304 may provide a third filtered selection of canonical maps. In some embodiments, for example, the output of action 304 may be only five canonical maps out of 50 canonical maps identified after the second filtered selection. The map transmitter 122 then sends one or more canonical maps based on the third filtered selection to the viewing device. Action 304 may reduce the number of canonical maps by approximately ten times, for example, from dozens of canonical maps to a single-digit number of canonical maps (e.g., 5) forming the third selection. In some embodiments, the XR device may receive a canonical map in the third filtered selection and attempt to locate within the received canonical map.

[0375] For example, action 304 may filter canonical map 120 based on global feature string 316 of canonical map 120 and based on global feature string 316 of an image captured by a viewing device (e.g., an image that may be part of a user's local tracking map). Fig. 27 Each canonical map 120 in has one or more global feature strings 316 associated therewith. In some embodiments, the global feature strings 316 may be obtained when the XR device submits images or feature details to the cloud and the images or feature details are processed in the cloud to generate global feature strings 316 for the canonical map 120.

[0376] In some embodiments, the cloud can receive feature details of a live / new / current image captured by the viewing device, and the cloud can generate a global feature string 316 of the live image. The cloud can then filter the canonical map 120 based on the real-time global feature string 316. In some embodiments, the global feature string can be generated on the local viewing device. In some embodiments, the global feature string can be generated remotely, for example, in the cloud. In some embodiments, the cloud can send the filtered canonical map to the XR device along with the global feature string 316 associated with the filtered canonical map. In some embodiments, when the viewing device positions its tracking map to the canonical map, it can do this by matching the global feature string 316 of the local tracking map with the global feature string of the canonical map.

[0377] It should be understood that the operation of the XR device may not perform all of the actions (300, 302, 304). For example, if the world of the canonical map is relatively small (e.g., 500 maps), the XR device attempting to perform positioning may filter the world of the canonical map based on Wi-Fi fingerprints (e.g., action 302) and keyframes (e.g., action 304), but omit the region-based filtering (e.g., action 300). Moreover, it is not necessary to compare the entire map. For example, in some embodiments, comparison of the two maps may result in identification of common persistent points, such as persistent gestures or PCFs that appear in both the new map and in a map selected from the map world. In that case, descriptors may be associated with the persistent points, and those descriptors may be compared.

[0378] Fig.29 9 is a flow chart illustrating a method 900 of selecting one or more ranked environment maps according to some embodiments. In the illustrated embodiment, the ranking is performed on the AR device of the user who is creating the tracking map. Thus, the tracking map can be used to rank the environment maps. In embodiments where the tracking map is not available, some or all of the selection and ranking of the environment maps that do not explicitly rely on the tracking map can be used.

[0379] Method 900 may begin with action 902, where a set of maps in a database of environmental maps (which may be formatted as canonical maps) located near the location where the tracking map is formed may be accessed and then filtered for ranking. Additionally, at action 902, at least one area attribute of the area in which the user's AR device is operating is determined. In a scenario where the user's AR device is constructing a tracking map, the area attribute may correspond to the area on which the tracking map is created. As a specific example, the area attribute may be calculated based on a received signal from an access point to a computer network while the AR device is calculating the tracking map.

[0380] Fig.30 An exemplary map ranking portion 806 of the AR system 800 is depicted in accordance with some embodiments. The map ranking portion 806 can be executed in a cloud computing environment, as it can include a portion executed on the AR device and a portion executed on a remote computing system such as the cloud. The map ranking portion 806 can be configured to perform at least a portion of the method 900.

[0381] Fig.31AExamples of area attributes AA1-AA8 of a tracking map (TM) 1102 and environment maps CM1-CM4 in a database according to some embodiments are depicted. As shown, the environment map can be associated with multiple area attributes. The area attributes AA1-AA8 may include parameters of a wireless network detected by the AR device computing the tracking map 1102, such as a basic service set identifier (BSSID) of the network to which the AR device is connected and / or the strength of a received signal to an access point of the wireless network through, for example, a network tower 1104. The parameters of the wireless network may conform to protocols including Wi-Fi and 5G NR. In Fig.32 In the example shown in , the region attribute is a fingerprint of the area in which the user AR device collects sensor data to form a tracking map.

[0382] Fig.31B An example of a determined geographic location 1106 of a tracking map 1102 according to some embodiments is depicted. In the example shown, the determined geographic location 1106 includes a centroid point 1110 and an area 1108 surrounding the centroid point. It should be understood that the determination of the geographic location of the present application is not limited to the format shown. The determined geographic location can have any suitable format, including, for example, different area shapes. In this example, the geographic location is determined from the area attributes using a database that associates area attributes with the geographic location. Databases are commercially available, for example, a database that associates Wi-Fi fingerprints with locations expressed as latitude and longitude and can be used for this operation.

[0383] exist Fig.29 In an embodiment, the map database containing environmental maps may also include location data for those maps, including the latitudes and longitudes covered by the maps. Processing at action 902 may require selecting a set of environmental maps from the database that cover the same latitudes and longitudes determined for the regional attributes of the tracking map.

[0384] Action 904 is a first filtering of the set of environment maps accessed in action 902. In action 902, environment maps are retained in the set based on proximity to the geographic location of the tracking map. This filtering step may be performed by comparing the latitude and longitude associated with the tracking map and the environment map in the set.

[0385] Fig.32An example of action 904 according to some embodiments is depicted. Each area attribute may have a corresponding geographic location 1202. The set of environment maps may include an environment map having at least one area attribute having a geographic location that overlaps with the determined geographic location of the tracking map. In the example shown, a set of identified environment maps includes environment maps CM1, CM2, and CM4, each of which has at least one area attribute having a geographic location that overlaps with the determined geographic location of the tracking map 1102. CM3, which is associated with area attribute AA6, is not included in the set because it is outside the determined geographic location of the tracking map.

[0386] Other filtering steps may also be performed on the group of environmental maps to reduce / rank the number of environmental maps that are ultimately processed in the group (such as for map merging or providing navigable world information to a user device). Method 900 may include filtering (action 906) the group of environmental maps based on the similarity of one or more identifiers of network access points associated with the tracking map and the environmental map of the group of environmental maps. During the formation of the map, the device that collects sensor data to generate the map may be connected to the network through a network access point (such as through Wi-Fi or a similar wireless communication protocol). The access point may be identified by a BSSID. When the user device moves through the area where the data is collected to form a map, the user device may be connected to a plurality of different access points. Similarly, when multiple devices provide information to form a map, the device may have been connected through different access points, so for this reason, multiple access points may also be used when forming a map. Therefore, there may be multiple access points associated with the map, and the group of access points may be an indication of the map location. The signal strength from the access point may be reflected as an RSSI value, which may provide further geographic information. In some embodiments, a list of BSSID and RSSI values ​​may form a regional attribute for a map.

[0387] In some embodiments, filtering the set of environment maps based on similarity of one or more identifiers of the network access points may include retaining in the set of environment maps an environment map having a highest Jaccard similarity to at least one area attribute of the tracking map based on the one or more identifiers of the network access points. Fig.33 An example of action 906 according to some embodiments is depicted. In the example shown, a network identifier associated with area attribute AA7 may be determined as an identifier of tracking map 1102. The set of environment maps after action 906 includes: environment map CM2, which may have an area attribute within a higher Jaccard similarity to AA7; and environment map CM4, which also includes area attribute AA7. Environment map CM1 is not included in the set because it has the lowest Jaccard similarity to AA7.

[0388] The processing of actions 902-906 can be performed based on metadata associated with the map without actually accessing the content of the map stored in the map database. Other processing may involve accessing the content of the map. Action 908 indicates accessing the environment map retained in the subset after filtering based on the metadata. It should be understood that if subsequent operations can be performed on the accessed content, this action can be performed earlier or later in the process.

[0389] Method 900 may include filtering (action 910) a set of environmental maps based on similarities of metrics representing the content of the tracking map and the environmental maps of the set of environmental maps. The metrics representing the content of the tracking map and the environmental maps may include vectors of values ​​calculated from the content of the maps. For example, as described above, a deep keyframe descriptor calculated for one or more keyframes used to form a map may provide a metric for comparing maps or portions of maps. The metric may be calculated from the map obtained at action 908, or may be pre-calculated and stored as metadata associated with those maps. In some embodiments, filtering a set of environmental maps based on similarities of metrics representing the content of the tracking map and the environmental maps of the set of environmental maps may include retaining an environmental map in the set of environmental maps that has a minimum vector distance between a feature vector of the tracking map and a vector representing an environmental map in the set of environmental maps.

[0390] Method 900 may include further filtering (action 912) a set of environment maps based on a degree of match between a portion of the tracking map and a portion of an environment map of the set of environment maps. The degree of match may be determined as part of a positioning process. As a non-limiting example, positioning may be performed by identifying critical points in the tracking map and the environment map that are sufficiently similar to the same portion of the physical world that they may represent. In some embodiments, the key points may be features, feature descriptors, keyframes, key assemblies, persistent poses, and / or PCFs. Then, a set of critical points in the tracking map may be aligned to produce an optimal fit with the set of critical points in the environment map. The mean square distance between corresponding critical points may be calculated and, if below a threshold for a particular area of ​​the tracking map, used as an indication that the tracking map and the environment map represent the same area of ​​the physical world.

[0391] In some embodiments, filtering a set of environment maps based on a degree of match between a portion of a tracking map and a portion of an environment map of the set of environment maps may include: calculating a volume of the physical world represented by the tracking map, which tracking map is also represented in an environment map of the set of environment maps; and retaining in the set of environment maps an environment map having a larger calculated volume than an environment map filtered from the set. Fig.34An example of action 912 is depicted in accordance with some embodiments. In the example shown, a set of environment maps following action 912 includes environment map CM4 having regions 1402 that match regions of tracking map 1102. Environment map CM1 is not included in the set because it does not have regions that match regions of tracking map 1102.

[0392] In some embodiments, the set of environment maps may be filtered in the order of act 906, act 910, and act 912. In some embodiments, the set of environment maps may be filtered based on act 906, act 910, and act 912, which may be performed in the order of processing required to perform the filtering, from lowest to highest. The method 900 may include loading (act 914) the set of environment maps and data.

[0393] In the example shown, the user database stores a region identifier indicating the region where the AR device is used. The region identifier may be a region attribute that may include parameters of a wireless network detected by the AR device during use. The map database may store multiple environment maps constructed from data provided by the AR device and associated metadata. The associated metadata may include a region identifier derived from the region identifier of the AR device providing the data, and the environment map is constructed from the data. The AR device may send a message to the PW module indicating that a new tracking map has been created or is being created. The PW module may calculate a region identifier for the AR device and update the user database based on the received parameters and / or the calculated region identifier. The PW module may also determine a region identifier associated with the AR device requesting the environment map, identify the group of environment maps from the map database based on the region identifier, filter the group of environment maps, and send a filtered set of environment maps to the AR device. In some embodiments, the PW module can filter the set of environmental maps based on one or more criteria including, for example, the geographic location of the tracking map, the similarity of one or more identifiers of network access points associated with the tracking map and the environmental maps of the set of environmental maps, the similarity of metrics representing the contents of the tracking map and the environmental maps of the set of environmental maps, and the degree of match between a portion of the tracking map and a portion of the environmental maps of the set of environmental maps.

[0394] Having described several aspects of some embodiments, it will be appreciated that various changes, modifications, and improvements will readily occur to those skilled in the art. As an example, embodiments are described in conjunction with an augmented (AR) environment. It will be appreciated that some or all of the techniques described herein may be applied in an MR environment or more generally in other XR environments and VR environments.

[0395] As another example, embodiments are described in conjunction with devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via any suitable combination of a network (such as the cloud), a discrete application, and / or a device, network, and a discrete application.

[0396] also, Fig.29 Examples of criteria that can be used to filter candidate maps to produce a set of highly ranked maps are provided. Other criteria can be used instead of or in addition to the criteria described. For example, if multiple candidate maps have similar values ​​for a metric used to filter out less desirable maps, the characteristics of the candidate maps can be used to determine which maps are retained as candidate maps or filtered out. For example, larger or more densely populated candidate maps can be prioritized over smaller candidate maps. In some embodiments, Figure 27-28 Can describe Figure 29-34 All or part of the systems and methods described in.

[0397] Fig.35 and 36 is a schematic diagram illustrating an XR system configured to rank and merge multiple environment maps according to some embodiments. In some embodiments, the traversable world (PW) can determine when to trigger ranking and / or merging of maps. In some embodiments, determining which map to use can be based at least in part on the above Figure 21 to Figure 25 Depth keyframes described.

[0398] Fig.37 3700 is a block diagram illustrating a method 3700 for creating an environment map of the physical world according to some embodiments. The method 3700 may be performed from locating a tracking map captured by an XR device worn by a user (act 3702) to a canonical map (e.g., by Fig.28 Action 3702 may include positioning the key assemblies of the tracking map into the group of the canonical map. The positioning result of each key assembly may include a localized pose of the key assembly and a set of 2D to 3D feature correspondences.

[0399] In some embodiments, method 3700 may include splitting (act 3704) the tracking map into connected parts, which may robustly merge maps by merging connected fragments. Each connected part may include a critical assembly within a predetermined distance. Method 3700 may include: merging (act 3706) connected parts greater than a predetermined threshold into one or more canonical maps; and removing the merged connected parts from the tracking map.

[0400] In some embodiments, method 3700 may include merging (action 3708) canonical maps in a group that are merged with the same connected portion of the tracking map. In some embodiments, method 3700 may include promoting (action 3710) the remaining connected portions of the tracking map that have not been merged with any canonical map to the canonical map. In some embodiments, method 3700 may include merging (action 3712) the persistent poses and / or PCFs of the tracking map and the canonical map, wherein the canonical map is merged with at least one connected portion of the tracking map. In some embodiments, method 3700 may include finalizing (action 3714) the canonical map, for example by fusing map points and pruning redundant critical assemblies.

[0401] Fig.38A and 38B 3800 is shown, according to some embodiments, by updating the canonical map 700, which can be updated from the tracking map 700 ( Figure 7 ) to upgrade. Figure 7 As illustrated and described, the canonical map 700 may provide a floor plan 706 of a reconstructed physical object represented by point 702 in the corresponding physical world. In some embodiments, the map point 702 may represent a feature of the physical object, which may include multiple features. A new tracking map of the physical world may be captured and uploaded to the cloud to be merged with the map 700. The new tracking map may include map point 3802 and key assemblies 3804, 3806. In the example shown, key assembly 3804 represents a corresponding relationship with key assembly 704 of map 700 (e.g., Fig.38B On the other hand, key assembly 3806 represents a key assembly that has not yet been located to map 700. In some embodiments, key assembly 3806 can be promoted to a separate canonical map.

[0402] Figures 39A to 39F is a diagram illustrating an example of a cloud-based persistent coordinate system that provides a shared experience for users in the same physical space. Fig.39A For example, a canonical map 4814 from the cloud is shown by FIG. 20A to FIG. 20C The canonical map 4814 may have a canonical coordinate frame 4806C. The canonical map 4814 may have a PCF 4810C with multiple associated PPs (e.g., Fig.39C 4818A, 4818B).

[0403] Fig.39BThe relationship established between the XR devices' respective world coordinate systems 4806A, 4806B and the canonical coordinate frame 4806C is shown. For example, this can be accomplished by positioning to a canonical map 4814 on the respective devices. For each device, positioning the tracking map to the canonical map can result in a transformation between its local world coordinate system and the coordinate system of the canonical map for each device.

[0404] Fig.39C It is shown that a transformation (e.g., transformation 4816A, transformation 4816B) between a local PCF (e.g., PCF 4810A, PCF 4810B) on a respective device to a respective persistent pose (e.g., PP 4818A, PP 4818B) on a canonical map can be calculated as a result of positioning. Using these transformations, each device can use its local PCF to determine where to display virtual content attached to PP 4818A, PP 4818B, or other persistent points of the canonical map relative to the local device, where the local PCF can be detected locally on the device by processing images detected using sensors on the device. Such an approach can accurately position virtual content relative to each user and can enable each user to have the same experience of virtual content in physical space.

[0405] Fig.39D A snapshot of the persistent pose from the canonical map to the local tracking map is shown. It can be seen that the local tracking maps are connected to each other through the persistent pose. Fig.39E It is shown that PCF 4810A on the device worn by user 4802A can be accessed in the device worn by user 4802B through PP 4818A. Fig.39F It is shown that tracking maps 4804A, 4804B and canonical map 4814 can be merged. In some embodiments, some PCFs can be removed due to the merge. In the example shown, the merged map includes PCF 4810C of canonical map 4814, but does not include PCFs 4810A, PCF 4810B of tracking maps 4804A, 4804B. After the map merge, PPs previously associated with PCFs 4810A, PCF 4810B can be associated with PCF 4810C.

[0406] Example

[0407] Fig.40 and Fig.41 Shown by Fig. 9 An example of a first XR device 12.1 using a tracking map. Fig.40 is a three-dimensional first local tracking map (ground Figure 1 ), which can be represented by Fig. 9The first XR device generation. Fig.41 is a diagram showing the Fig. 9 The first XR device uploads the address to the server Figure 1 Block diagram of the .

[0408] Fig.40 The first XR device 12.1 shows the Figure 1 and virtual content (content 123 and content 456). Figure 1 Has an origin (origin 1). Figure 1 The PCF a is located at the ground level. Figure 1 , and PCF b has X, Y, and Z coordinates of (0, 0, 0). Content 123 is associated with PCF a. In this example, content 123 has an X, Y, and Z relationship with respect to PCF a of (1, 0, 0). Content 456 has a relationship with respect to PCF b. In this example, content 456 has an X, Y, and Z relationship with respect to PCF b of (1, 0, 0).

[0409] exist Fig.41 In the first XR device 12.1, Figure 1 Uploaded to the server 20. In this example, since the server does not store a canonical map for the same area of ​​the physical world represented by the tracking map, and the tracking map is stored as the initial canonical map. The server 20 now has a map based on the map. Figure 1 The first XR device 12.1 has a canonical map which is empty at this stage. For the purposes of discussion, and in some embodiments, the server 20 has a canonical map in addition to the map. Figure 1 No other maps are included. No maps are stored on the second XR device 12.2.

[0410] The first XR device 12.1 also sends its Wi-Fi signature data to the server 20. The server 20 can use the Wi-Fi signature data to determine the approximate location of the first XR device 12.1 based on intelligence gathered from other devices that have connected to the server 20 or other servers in the past along with the recorded GPS locations of such other devices. The first XR device 12.1 can now end the first session (see Figure 8 ), and can disconnect from the server 20.

[0411] Fig.42 is a diagram showing a method according to some embodiments Fig.16Schematic diagram of an XR system showing that after the first user 14 . 1 terminated the first session, the second user 14 . 2 has initiated a second session using a second XR device of the XR system. Fig.43A A block diagram showing a second user 14.2 initiating a second session. Because the first session of the first user 14.1 has ended, the first user 14.1 is shown in dashed lines. The second XR device 12.2 begins recording the object. The server 20 may use various systems with different granularity to determine that the second session of the second XR device 12.2 is in the same vicinity as the first session of the first XR device 12.1. For example, the first XR device 12.1 and the second XR device 12.2 may include Wi-Fi signature data, global positioning system (GPS) location data, GPS data based on the Wi-Fi signature data, or any other data indicating location to record their locations. Alternatively, the PCF identified by the second XR device 12.2 may be displayed in a manner similar to that of the ground. Figure 1 The similarity of PCF.

[0412] like Fig.43B As shown in , the second XR device starts and begins collecting data, such as images 1110 from one or more cameras 44, 46. Fig.14 As shown in , in some embodiments, an XR device (e.g., a second XR device 12.2) may collect one or more images 1110 and perform image processing to extract one or more features / points of interest 1120. Each feature may be converted into a descriptor 1130. In some embodiments, the descriptor 1130 may be used to describe a keyframe 1140, which may have a position and orientation of an associated image attached. One or more keyframes 1140 may correspond to a single persistent pose 1150, which may be automatically generated after a threshold distance (e.g., 3 meters) from a previous persistent pose 1150. One or more persistent poses 1150 may correspond to a single PCF 1160, which may be automatically generated after a predetermined distance (e.g., every 5 meters). Over time, as the user continues to move around the user's environment and the XR device continues to collect more data (such as images 1110), additional PCFs (e.g., PCF 3 and PCFs 4, 5) may be created. One or more applications 1180 may run on the XR device and provide virtual content 1170 to the XR device for presentation to the user. The virtual content may have an associated content coordinate frame that may be placed relative to one or more PCFs. Fig.43B As shown in , the second XR device 12 . 2 creates three PCFs. In some embodiments, the second XR device 12 . 2 may attempt to locate one or more canonical maps stored on the server 20 .

[0413] In some embodiments, Fig.43C As shown in FIG. 1 , the second XR device 12.2 can download the canonical map 120 from the server 20. Figure 1 Includes PCFs a to d and origin 1. In some embodiments, server 20 may have multiple canonical maps for various locations and may determine that second XR device 12.2 is located in the same vicinity as first XR device 12.1 during the first session and send a canonical map of that vicinity to second XR device 12.2.

[0414] Fig.44 The second XR device 12.2 begins to identify the PCF for generating the address Figure 2 The second XR device 12.2 recognizes only a single PCF, namely PCF 1,2. The X, Y, and Z coordinates of PCF 1,2 of the second XR device 12.2 may be (1, 1, 1). Figure 2 The second XR device 12.2 may have its own origin (origin 2), which may be based on the head pose of device 2 at the start of the device's current head pose session. In some embodiments, the second XR device 12.2 may immediately attempt to place the ground Figure 2 In some embodiments, because the system cannot identify any or sufficient overlap between the two maps, the map Figure 2 It may not be possible to locate the standard map (map Figure 1 ) (i.e., localization may fail). Localization can be performed by identifying a portion of the physical world represented in the first map that is also represented in the second map, and computing the transformation between the first map and the second map required to align these portions. In some embodiments, the system can localize based on a PCF comparison between the local map and the canonical map. In some embodiments, the system can localize based on a persistent pose comparison between the local map and the canonical map. In some embodiments, the system can localize based on a keyframe comparison between the local map and the canonical map.

[0415] Fig.45 The second XR device 12.2 identifies the Figure 2 The other PCFs (PCF 1, 2, PCF 3, PCF 4, 5) after Figure 2 The second XR device 12.2 tries again to Figure 2 Locate to the standard map. Figure 2 has been extended to overlap at least a portion of the canonical map, so the positioning attempt will succeed. Figure 2 The overlap between the and canonical maps can be represented by PCFs, persistent poses, keyframes, or any other suitable intermediate or derived construct.

[0416] In addition, the second XR device 12.2 has linked content 123 and content 456 to the local Figure 2 Content 123 has X, Y, and Z coordinates (1, 0, 0) relative to PCF 1, 2, and PCF 3. Similarly, content 123 has X, Y, and Z coordinates (1, 0, 0) relative to PCF 1, 2. Figure 2 In PCF 3, the X, Y, and Z coordinates of content 456 are (1, 0, 0).

[0417] Fig.46A and Fig.46B Shown Figure 2 The successful localization to the canonical map. Localization can be based on matching features in one map to another map. By appropriate transformations, which involve translation and rotation of one map relative to the other, the overlapping area / volume / section of the maps 1410 represents the localization of the map. Figure 1 and the common parts of the normative map. Figure 2 PCFs 3, 4, and 5 were created before positioning, and the normative map was created before the location Figure 2 PCFs a and c were created previously, so different PCFs are created to represent the same volume in real space (e.g., in different maps).

[0418] like Fig.47 As shown in FIG. 1 , the second XR device 12.2 extends the ground Figure 2 , to include the PCF ad from the canonical map. Including the PCF ad indicates that the Figure 2 In some embodiments, the XR system may perform an optimization step to remove duplicate PCFs from overlapping regions, such as PCF 1410, PCF 3, and PCF 4, 5. Figure 2 After positioning, the placement of virtual content (such as content 456 and content 123) will be updated. Figure 2 The virtual content appears in the same real-world location relative to the user, despite the change to the content’s PCF attachment, and despite the update to the map. Figure 2 PCF.

[0419] like Fig.48 As shown in FIG. , the second XR device 12.2 continues to expand Figure 2 , for example, as the user moves around the real world, the second XR device 12.2 will identify other PCFs (PCF e, f, g, and h). Figure 1 exist Fig.47 and Fig.48 There are no extensions in .

[0420] refer to Fig.49 , the second XR device 12.2 will be Figure 2Upload to server 20. Server 20 will Figure 2 With normative Figure 1 In some embodiments, when the session for the second XR device 12.2 ends, the address Figure 2 It can be uploaded to the server 20.

[0421] The canonical map in the server 20 now includes PCF i, which does not include the map on the first XR device 12.1. Figure 1 When a third XR device (not shown) uploads a map to the server 20 and the map includes PCFi, the canonical map on the server 20 may have been extended to include PCFi.

[0422] exist Fig.50 In the example, the server 20 will Figure 2 The server 20 determines the PCFs a to d for the normative map and the map. Figure 2 is common. The server extends the specification map to include PCF e to h and from ground Figure 2 The PCFs 1 and 2 of the first XR device 12.1 and the second XR device 12.2 are based on the ground. Figure 1 , and is outdated.

[0423] exist Fig.51 In some embodiments, this may occur when the first XR device 12.1 and the second XR device 12.2 attempt to locate during a different or new or subsequent session. The first XR device 12.1 and the second XR device 12.2 proceed as described above to locate their respective local maps (respectively, local Figure 1 peacefully Figure 2 ) to locate the new canonical map.

[0424] like Fig.52 As shown in FIG. 1 , the head coordinate frame 96 or “head pose” is relative to the ground. Figure 2 In some embodiments, the origin of the map, origin 2, is based on the head pose of the second XR device 12.2 at the start of the session. When a PCF is created during a session, the PCF is placed relative to the world coordinate frame origin 2. Figure 2 The PCF is used as a persistent coordinate system relative to the canonical coordinate frame, where the world coordinate frame can be the world coordinate frame of the previous session (e.g., Fig.40 The land in Figure 1 These coordinate frames are used to Figure 2 Positioning to the same transformation related to the canonical map, as above combined Fig.46Bdiscussed.

[0425] Previously referenced Fig. 9 The transformation from the world coordinate frame to the head coordinate frame 96 is discussed. Fig.52 The head coordinate frame 96 shown in FIG. 1 has only two orthogonal axes relative to the ground. Figure 2 The PCF is at a specific coordinate position and relative to the ground Figure 2 However, it should be understood that the head coordinate frame 96 is relative to the ground. Figure 2 The PCF is located in three dimensions and has three orthogonal axes in the three-dimensional space.

[0426] exist Fig.53 In the figure, the head coordinate frame 96 is relative to the ground Figure 2 Because the second user 14.2 has moved his head, the head coordinate frame 96 has moved. The user can move his head with six degrees of freedom (6dof). The head coordinate frame 96 can therefore be moved in 6dof (i.e., from its Fig.52 The previous position in three dimensions, and relative to the ground Figure 2 The PCF moves around three orthogonal axes). Fig. 9 The head coordinate frame 96 is adjusted as the real object detection camera 44 and the inertial measurement unit 48 in the head unit 22 detect real objects and motion, respectively. More information about head pose tracking is disclosed in U.S. Patent Application No. 16 / 221,065, entitled "Enhanced Pose Determination for Display Device," which is incorporated herein by reference in its entirety.

[0427] Fig.54 Sounds can be associated with one or more PCFs. A user can, for example, wear headphones or earphones with stereo sound. The location of the sound through the earphones can be simulated using conventional techniques. The location of the sound can be located at a fixed position so that when the user rotates his head to the left, the location of the sound rotates to the right so that the user perceives the sound as coming from the same location in the real world. In this example, the location of the sound is represented by sound 123 and sound 456. For ease of discussion, Fig.54 In terms of analysis Fig.48 When the first user 14.1 and the second user 14.2 are in the same room at the same or different times, they perceive the sound 123 and the sound 456 to come from the same location in the real world.

[0428] Fig.55 and Fig.56 Another implementation of the above technology is shown. Figure 8As described above, the first user 14.1 has initiated a first session. Fig.55 As shown in FIG. 1 , the first user 14.1 has terminated the first session, as shown by the dashed line. At the end of the first session, the first XR device 12.1 Figure 1 Uploaded to the server 20. The first user 14.1 has now initiated a second session at a later time than the first session. Figure 1 The first XR device 12.1 is already stored in the first XR device 12.1, so the first XR device 12.1 does not download the address from the server 20. Figure 1 If you lose your place Figure 1 , the first XR device 12.1 downloads the address from the server 20 Figure 1 Then, the first XR device 12.1 continues to construct Figure 2 PCF, located at the ground Figure 1 , and further develops the canonical map as described above. Then, as described above, the map of the first XR device 12.1 Figure 2 Used to associate local content, head coordinate frame, local sound, etc.

[0429] refer to Fig.57 and Fig.58 It is also possible that more than one user interacts with the server in the same session. In this example, the first user 14.1 and the second user 14.2 are joined together by a third user 14.3 and a third XR device 12.3. Each XR device 12.1, 12.2 and 12.3 starts generating its own map, i.e., the map Figure 1 ,land Figure 2 peacefully Figure 3 As XR devices 12.1, 12.2 and 12.3 continue to develop Figure 1 , 2 and 3, the map is incrementally uploaded to the server 20. The server 20 merges the Figure 1 , 2 and 3 to form a canonical map. The canonical map is then sent from the server 20 to each of the XR devices 12.1, 12.2 and 12.3.

[0430] Fig.59Aspects of a viewing method for recovering and / or resetting a head pose according to some embodiments are shown. In the example shown, at action 1400, the viewing device is powered on. At action 1410, in response to the power on, a new session is initiated. In some embodiments, the new session can include establishing a head pose. One or more capture devices on a head-mounted frame fixed to the user's head capture the surface of the environment by first capturing an image of the environment and then determining the surface from the image. In some embodiments, the surface data can be combined with data from a gravity sensor to establish a head pose. Other suitable methods of establishing a head pose can be used.

[0431] At action 1420, the processor of the viewing device enters a routine for tracking head pose. As the user moves their head to determine the orientation of the head mounted frame relative to the surface, the capture device continues to capture the surface of the environment.

[0432] At action 1430, the processor determines whether the head pose has been lost. The head pose may be lost due to "edge" situations, such as too many reflective surfaces, low light, blank walls, outdoors, etc. that can result in low feature acquisition; or due to dynamic situations such as a crowd of people moving and forming part of a map. The routine at 1430 allows a certain amount of time to pass, such as 10 seconds, to allow sufficient time to determine whether the head pose has been lost. If the head pose has not been lost, the processor returns to 1420 and once again enters tracking of the head pose.

[0433] If the head pose has been lost at action 1430, the processor enters a routine to recover the head pose at 1440. If the head pose is lost due to low light, a message such as the following message will be displayed to the user via the display of the viewing device:

[0434] The system is detecting low light conditions. Please move to a better lit area.

[0435] The system will continue to monitor whether sufficient light is available and whether the head pose can be recovered. The system may alternatively determine that low texture of the surface is causing the head pose to be lost, in which case the following prompt is given to the user in the display as a suggestion to improve surface capture:

[0436] The system cannot detect enough surfaces with fine textures. Please move to an area with less rough surface textures and finer textures.

[0437] At action 1450, the processor enters a routine to determine whether head posture recovery has failed. If head posture recovery has not failed (i.e., head posture recovery has succeeded), the processor returns to action 1420 by re-entering tracking of the head posture. If head posture recovery has failed, the processor returns to action 1410 to establish a new session. As part of the new session, all cached data is invalidated and the head posture is re-established thereafter. Any suitable method of head tracking may be used with Fig.59 Head tracking is described in conjunction with the process described in U.S. Patent Application No. 16 / 221,065, which is incorporated herein by reference in its entirety.

[0438] Remote positioning

[0439] Various embodiments may utilize remote resources to facilitate a persistent and consistent cross-reality experience between individuals and / or groups of users. The inventors have recognized and appreciated that it is possible to Fig.30 The benefits of operating an XR device utilizing canonical maps as described herein may be realized in the context of a set of canonical maps as shown. For example, the benefits may be realized by sending feature and posture information to a remote service that maintains a set of canonical maps. A device seeking to use canonical maps to position virtual content at locations specified relative to the canonical maps may receive one or more transformations between the features and the canonical maps from the remote service. These transformations may be used on a device that maintains information about the locations of these features in the physical world to position virtual content at locations specified relative to the canonical maps, or otherwise identify locations in the physical world specified relative to the canonical maps.

[0440] In some embodiments, spatial information is captured by the XR device and transmitted to a remote service, such as a cloud-based service, which uses the spatial information to position the XR device to a canonical map used by applications or other components of the XR system to specify the location of virtual content relative to the physical world. Once positioned, a transform that links a tracking map maintained by the device to the canonical map can be transmitted to the device. The transform can be used in conjunction with the tracking map to determine the location of the virtual content to be rendered relative to the canonical map, or otherwise identify a location in the physical world relative to the canonical map.

[0441] The inventors have appreciated that the data that needs to be exchanged between a device and a remote positioning service may be very small compared to the transmission of map data, which may occur when a device transmits a tracking map to a remote service and receives a set of canonical maps from the service for device-based positioning. In some embodiments, performing positioning functions on cloud resources requires only a small amount of information to be transmitted from the device to the remote service. For example, a complete tracking map need not be transmitted to the remote service to perform positioning. In some embodiments, feature and posture information, such as may be stored in association with a persistent posture as described above, may be transmitted to a remote server. As described above, in embodiments where features are represented by descriptors, the information uploaded may be smaller.

[0442] The result returned to the device from the positioning service may be one or more transformations relating the uploaded features to portions of the matching canonical map. These transformations may be used in the XR system in conjunction with its tracking map to identify the location of virtual content or otherwise identify locations in the physical world. In embodiments that use persistent spatial information such as the PCFs described above to specify locations relative to a canonical map, the positioning service may download the transformations between the features and one or more PCFs to the device after successful positioning.

[0443] As a result, the network bandwidth consumed by communications between the XR device and the remote service used to perform positioning may be low. The system can therefore support frequent positioning, enabling each device interacting with the system to quickly obtain information for positioning virtual content or performing other location-based functions. As the device moves through the physical environment, it may repeat requests for updated positioning information. In addition, the device may frequently obtain updates to the positioning information, such as when the canonical map changes, such as by incorporating additional tracking maps to expand the map or improve its accuracy.

[0444] Additionally, uploading features and downloading transforms can enhance privacy in XR systems that share map information among multiple users by increasing the difficulty of obtaining maps through spoofing. For example, an unauthorized user can be prevented from obtaining a map from the system by sending a false request for a canonical map representing a portion of the physical world in which the unauthorized user is not located. An unauthorized user is unlikely to access features in the area of ​​the physical world for which they are requesting map information if the unauthorized user is not physically present in that area. In embodiments where the feature information is formatted as feature descriptions, the difficulty of spoofing feature information in a request for map information is further compounded. Additionally, when the system returns transforms intended to be applied to a tracking map of a device operating in the area for which location information is requested, the information returned by the system may be of little or no use to an impostor.

[0445] According to one embodiment, the positioning service is implemented as a cloud-based microservice. In some examples, implementing a cloud-based positioning service can help save device computing resources and can enable the calculations required for positioning to be performed with very low latency. These operations can be supported by virtually unlimited computing power or other computing resources available by providing additional cloud resources, thereby ensuring the scalability of the XR system to support numerous devices. In one example, many canonical maps can be maintained in memory for almost instant access or stored in high-availability devices to reduce system latency.

[0446] In addition, performing positioning on multiple devices in the cloud service can achieve improvements to the process. Positioning telemetry and statistics can provide information about which canonical maps are in active memory and / or high-availability storage. For example, statistics of multiple devices can be used to identify the most frequently accessed canonical maps.

[0447] Additional accuracy may also be achieved as a result of processing in a cloud environment or other remote environment having substantial processing resources relative to the remote device. For example, localization may be performed on a higher density canonical map in the cloud relative to processing performed on a local device. Maps may be stored in the cloud, for example, with a greater number of PCFs or a higher density of feature descriptors per PCF, thereby increasing the accuracy of the match between a set of features from the device and the canonical map.

[0448] Fig.61 6100. A user device that displays cross-reality content during a user session can take a variety of forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As described above, these devices can be configured with software, such as applications or other components, and / or hardwired to generate local location information (e.g., tracking maps) that can be used to render virtual content on their respective displays.

[0449] The virtual content location information may be specified relative to the global location information, for example, the global location information may be formatted as a canonical map containing one or more PCFs. According to some embodiments, for example Fig.61 In the embodiment shown in , system 6100 is configured with a cloud-based service that supports the running and display of virtual content on a user device.

[0450] In one example, positioning functionality is provided as a cloud-based service 6106, which may be a microservice. The cloud-based service 6106 may be implemented on any of a plurality of computing devices from which computing resources may be allocated to one or more services executed in the cloud. Those computing devices may be interconnected with each other and accessible to devices such as the wearable XR device 6102 and the handheld device 6104. Such connectivity may be provided over one or more networks.

[0451] In some embodiments, the cloud-based service 6106 is configured to accept descriptor information from various user devices and "locate" the device to a matching one or more canonical maps. For example, the cloud-based positioning service matches the received descriptor information with the descriptor information of the corresponding canonical map. Canonical maps can be created using the techniques described above, which create canonical maps by merging maps provided by one or more devices with image sensors or other sensors that obtain information about the physical world. However, it is not required that canonical maps be created by devices that access them, as such maps can be created by map developers, for example, map developers can publish maps by making them available to the positioning service 6106.

[0452] According to some embodiments, the cloud service handles canonical map identification and may include operations to filter the repository of canonical maps to a set of potential matches. The filtering may be as follows: Fig.29 as shown, or by using any subset of the filter criteria and replacing Fig.29 or except the filter criteria shown in Fig.29 In one embodiment, geographic data may be used to limit the search for matching canonical maps to maps that represent an area proximate to the device requesting location. For example, regional attributes, such as Wi-Fi signal data, Wi-Fi fingerprint information, GPS data, and / or other device location information, may be used as a coarse filter on stored canonical maps to limit the analysis of descriptors to canonical maps that are known or likely to be proximate to the user's device. Similarly, the location history of each device may be maintained by a cloud service to prioritize searches for canonical maps near the device's last location. In some examples, filtering may include the above with respect to Fig.31B , Fig.32 , Fig.33 and Fig.34 Functionality discussed.

[0453] Fig.62It is an example process that can be executed by a device to use a cloud-based service to locate the position of the device using a canonical map and receive transformation information specifying one or more transformations between the local coordinate system of the device and the coordinate system of the canonical map. Various embodiments and examples may describe one or more transformations as specifying a transformation from a first coordinate frame to a second coordinate frame. Other embodiments include a transformation from a second coordinate frame to a first coordinate frame. In other embodiments, the transformation implements a transition from one coordinate frame to another, and the resulting coordinate frame depends only on the desired coordinate frame output (including, for example, the coordinate frame in which the content is displayed). In yet another embodiment, the coordinate system transformation may enable determination from a second coordinate frame to a first coordinate frame and from a first coordinate frame to a second coordinate frame.

[0454] According to some embodiments, information reflecting the transformation for each persistent gesture defined by the canonical map may be transmitted to the device.

[0455] According to one embodiment, process 6200 may start with a new session at 6202. Starting a new session on a device may initiate the capture of image information to build a tracking map for the device. In addition, the device may send a message to register with a server of a location service, prompting the server to create a session for the device.

[0456] In some embodiments, starting a new session on a device may optionally include sending adjustment data from the device to a location service. The location service returns one or more transformations calculated based on a set of features and associated postures to the device. If the posture of the feature is adjusted based on device-specific information before calculating the transformation and / or the transformation is adjusted based on device-specific information after calculating the transformation, instead of performing those calculations on the device, the device-specific information may be sent to the location service so that the location service can apply these adjustments. As a specific example, sending device-specific adjustment information may include capturing calibration data of the sensor and / or display. The calibration data can be used, for example, to adjust the position of the feature point relative to the measured position. Alternatively or additionally, the calibration data can be used to adjust the position of the command display to render virtual content so that it appears to be accurately positioned for the specific device. The calibration data can be obtained, for example, from multiple images of the same scene taken using sensors on the device. The position of the features detected in those images can be expressed as a function of the sensor position, so that multiple images generate a set of equations that can solve the sensor position. The calculated sensor position can be compared with the nominal position, and calibration data can be derived from any difference. In some embodiments, intrinsic information about the device construction can also enable calibration data to be calculated for the display, in some embodiments.

[0457] In embodiments where calibration data is generated for a sensor and / or display, the calibration data may be applied at any point in the measurement or display process. In some embodiments, the calibration data may be sent to a positioning server, which may store the calibration data in a data structure established for each device that has registered with the positioning server and is therefore in a session with the server. The positioning server may apply the calibration data to any transformations calculated as part of the positioning process for the device providing the calibration data. Thus, the computational burden of using the calibration data to improve the accuracy of the sensed and / or displayed information is borne by the calibration service, thereby providing a further mechanism to reduce the processing burden on the device.

[0458] Once the new session is established, process 6200 may continue to capture new frames of the device's environment at 6204. At 6206, each frame may be processed to generate descriptors for the captured frame (including, for example, the DSF values ​​discussed above). These values ​​may be calculated using some or all of the techniques described above, including those described above with respect to Fig.14 , Fig. 22 and Fig.23 Techniques discussed. As discussed, descriptors can be computed as a mapping of feature points, or in some embodiments, a mapping of image patches surrounding feature points to descriptors. The descriptors can have values ​​that enable valid matching between newly acquired frames / images and stored maps. In addition, the number of features extracted from the images can be limited to a maximum number of feature points per image, e.g., 200 feature points per image. As described above, feature points can be selected to represent points of interest. Thus, actions 6204 and 6206 can be performed as part of a device process for forming a tracking map or otherwise periodically collecting images of the physical world around the device, or can but need not be performed separately for positioning.

[0459] The feature extraction at 6206 may include attaching gesture information to the features extracted at 6206. The gesture information may be a gesture in the local coordinate system of the device. In some embodiments, the gesture may be relative to a reference point in a tracking map, such as a persistent gesture as described above. Alternatively or additionally, the gesture may be relative to the origin of the tracking map of the device. Such an embodiment may enable a positioning service as described herein to provide positioning services for a wide range of devices, even if they do not use persistent gestures. In any event, gesture information may be attached to each feature or group of features so that the positioning service may use the gesture information to calculate a transformation that may be returned to the device when matching the feature with a feature in a stored map.

[0460] Process 6200 can continue to decision box 6207, in which a decision is made whether to request positioning. One or more criteria can be applied to determine whether to request positioning. The criteria can include the passage of time, so that the device can request positioning after a certain threshold amount of time. For example, if positioning is not attempted within the threshold amount of time, the process can continue from decision box 6207 to action 6208, where positioning is requested from the cloud. The threshold amount of time can be between 10 and 30 seconds, such as 25 seconds. Alternatively or additionally, positioning can be triggered by the movement of the device. The device performing process 6200 can use an IMU and / or its tracking map to track its movement and initiate positioning when a movement exceeding a threshold distance from the location where the device was last requested to be positioned is detected. For example, the threshold distance can be between 1 and 10 meters, such as between 3 and 5 meters. As another alternative, positioning can be triggered in response to an event, such as when the device creates a new persistent posture or the current persistent posture of the device changes, as described above.

[0461] In some embodiments, decision block 6207 may be implemented so that a threshold for triggering positioning may be established dynamically. For example, in an environment where features are largely consistent so that the confidence in matching a set of extracted features with features of a stored map may be low, positioning may be requested more frequently to increase the chance that at least one positioning attempt will be successful. In this case, the threshold applied at decision block 6207 may be lowered. Similarly, in an environment where features are relatively few, the threshold applied at decision block 6207 may be lowered to increase the frequency of positioning attempts.

[0462] Regardless of how positioning is triggered, when triggered, process 6200 can proceed to action 6208, where the device sends a request to the positioning service, including data used by the positioning service to perform positioning. In some embodiments, data from multiple image frames can be provided for positioning attempts. For example, unless the features in multiple image frames produce consistent positioning results, the positioning service may not consider the positioning successful. In some embodiments, process 6200 can include saving feature descriptors and additional posture information to a buffer. The buffer can be, for example, a circular buffer that stores feature sets extracted from the most recently captured frames. Therefore, the positioning request can be sent together with multiple feature sets accumulated in the buffer. In some settings, the buffer size is implemented to accumulate multiple data sets that are more likely to produce successful positioning. In some embodiments, the buffer size can be set to accumulate features from, for example, two, three, four, five, six, seven, eight, nine, or ten frames. Optionally, the buffer size can have a baseline setting that can be increased in response to positioning failure. In some examples, increasing the buffer size and the corresponding number of transmitted feature sets reduces the possibility that subsequent positioning functions cannot return results.

[0463] Regardless of how the buffer size is set, the device can transmit the contents of the buffer to the location service as part of a location request. Other information can be transmitted along with the feature points and additional posture information. For example, in some embodiments, geographic information can be transmitted. The geographic information can include, for example, GPS coordinates or a wireless signature associated with the device tracking a map or a current persistent posture.

[0464] In response to the request sent at 6208, the cloud positioning service may analyze the feature descriptors to locate the device to a canonical map or other persistent map maintained by the service. For example, the descriptors match a set of features in a map where the device is located. The cloud-based positioning service may perform positioning as described above relative to device-based positioning (e.g., may rely on any of the functions discussed above for positioning (including map ranking, map filtering, location estimation, filtered map selection, Figure 44 to Figure 4 6, and / or discussed with respect to positioning module, PCF and / or PP identification and matching, etc.). However, instead of transmitting the identified canonical map to the device (e.g., in device positioning), the cloud-based positioning service can continue to generate transformations based on the matching features of the canonical map and the relative orientation of the feature set sent from the device. The positioning service can return these transformations to the device, which the device can receive at box 6210.

[0465] In some embodiments, canonical maps maintained by a positioning service may employ PCFs, as described above. In such embodiments, feature points of a canonical map that match feature points sent from a device may have locations specified relative to one or more PCFs. Thus, the positioning service may identify one or more canonical maps and may calculate a transformation between the coordinate frame represented in the gesture sent with the positioning request and the one or more PCFs. In some embodiments, identification of one or more canonical maps is assisted by filtering potential maps based on geographic data of the corresponding device. For example, once filtered to a candidate set (e.g., by other options such as GPS coordinates), the candidate set of canonical maps may be analyzed in detail to determine matching feature points or PCFs as described above.

[0466] The data returned to the requesting device in action 6210 may be formatted as a persistent posture transformation table. The table may be accompanied by one or more canonical map identifiers indicating the canonical map to which the device is located by the positioning service. However, it should be understood that the positioning information may be formatted in other ways, including as a transformation list, with associated PCF and / or canonical map identifiers.

[0467] Regardless of how the transforms are formatted, the device can use these transforms to calculate the position of rendering virtual content that has been specified by an application or other component of the XR system relative to any PCF in action 6212. This information is used alternatively or additionally on the device to perform any position-based operations in which the position is specified based on the PCF.

[0468] In some scenarios, the location service may not be able to match the features sent from the device to any stored canonical map, or may not be able to match a sufficient number of feature sets transmitted with the request to the location service to consider the location successfully located. In such a scenario, the location service may indicate to the device that the location failed, rather than returning the transformation to the device as described above in conjunction with action 6210. In such a scenario, process 6200 may branch to action 6230 at decision box 6209, where the device may take one or more actions for failure handling. These actions may include increasing the size of the buffer that stores the feature sets sent for location positioning. For example, if the location service does not consider the location to be successful unless three feature sets match, the buffer size may be increased from 5 to 6, thereby increasing the chance that the three transmitted feature sets match the canonical map maintained by the location service.

[0469] Alternatively or additionally, failure handling can include adjusting operating parameters of the device to trigger more frequent positioning attempts. For example, the threshold time and / or threshold distance between positioning attempts can be reduced. As another example, the number of feature points in each feature set can be increased. When a sufficient number of features in the set sent from the device match the features of the map, a match between the feature set and the features stored in the canonical map can be considered to have occurred. Increasing the number of features sent can increase the chance of a match. As a specific example, the initial feature set size can be 50, which can be increased to 100, 150, and then 200 at each consecutive positioning failure. After a successful match, the set size can then return to its initial value.

[0470] Failure handling may also include obtaining positioning information in addition to from a positioning service. According to some embodiments, the user device may be configured to cache canonical maps. Cached maps allow the device to access and display content that is not available in the cloud. For example, cached canonical maps allow device-based positioning in the event of a communication failure or other unavailability.

[0471] According to various embodiments, Fig.62 A high-level process for device-initiated cloud-based positioning is described. In other embodiments, various one or more of the steps shown may be combined, omitted, or invoke other processes to complete the positioning and final visualization of virtual content in the corresponding device view.

[0472] In addition, it should be understood that although process 6200 shows that the device determines whether to initiate positioning at decision block 6207, the trigger for initiating positioning can come from outside the device, including from a positioning service. For example, the positioning service can maintain information about each device in session with it. For example, the information may include an identifier of the canonical map to which each device was most recently positioned. The positioning service or other components of the XR system can update the canonical map, including using the above combined Fig.26 When a canonical map is updated, a location service may send a notification to each device that was recently located to the map. The notification may serve as a trigger for the device to request a location and / or may include an updated transform that is recalculated using a feature set recently sent from the device.

[0473] Fig.63A , Fig.63B and Fig.63C 6350, 6352, 6354, and 6456 illustrate an example architecture and separation between components involved in a cloud-based positioning process. For example, a module, component, and / or software configured to process perception on a user device is shown at 6350 (e.g., 660, Fig. 6A ). Device functionality for persistent world operations is shown at 6352 (including, for example, as described above and with respect to the persistent world module (e.g., 662, Fig. 6A )). In other embodiments, separation between 6350 and 6352 is not required and the communications shown may be between processes executing on the device.

[0474] Similarly, shown at block 6354 is a cloud process (e.g., 802, 812, 813) configured to handle functionality associated with traversable worlds / traversable world modeling. Fig.26 ). Shown at box 6356 is a cloud process that is configured to handle functionality associated with locating the device to one or more maps in a repository of stored canonical maps based on information sent from the device.

[0475] In the illustrated embodiment, process 6300 begins at 6302 when a new session begins. Sensor calibration data is obtained at 6304. The calibration data obtained may depend on the device (e.g., multiple cameras, sensors, positioning devices, etc.) represented at 6350. Once sensor calibration is obtained for a device, the calibration may be cached at 6306. If device operation results in a change in frequency parameters (e.g., collection frequency, sampling frequency, matching frequency, and other options), the frequency parameters are reset to a baseline at 6308.

[0476] Once the new session functionality is complete (e.g., calibration, steps 6302-6306), process 6300 can continue with capturing new frames 6312. At 6314, features and their corresponding descriptors are extracted from the frames. In some examples, the descriptors can include DSFs, as described above. In accordance with some embodiments, the descriptors can have spatial information attached to them to enable subsequent processing (e.g., transform generation). At 6316, pose information generated on the device (e.g., information for locating features in the physical world relative to a tracking map specified by the device, as described above) can be attached to the extracted descriptors.

[0477] At 6318, the descriptor and pose information are added to the buffer. New frames are captured and added to the buffer as shown in steps 6312-6318 in a loop until the buffer size threshold is exceeded at 6319. At 6320, in response to determining that the buffer size is met, a positioning request is transmitted from the device to the cloud. According to some embodiments, the request may be processed by a navigable world service (e.g., 6354) instantiated in the cloud. In further embodiments, the functional operations for identifying candidate canonical maps may be separated from the operations for actual matching (e.g., shown as boxes 6354 and 6356). In one embodiment, a cloud service for map filtering and / or map ranking may be executed at 6354 and process the positioning request received from 6320. According to one embodiment, the map ranking operation is configured to determine a set of candidate maps that may include the location of the device at 6322.

[0478] In one example, the map ranking function includes operations for identifying candidate canonical maps based on geographic attributes or other location data (e.g., observed or inferred location information). For example, the other location data may include Wi-Fi signatures or GPS information.

[0479] According to other embodiments, location data may be captured during a cross-reality session with a device and user. Process 6300 may include additional operations to populate locations for a given device and / or session (not shown). For example, location data may be stored as a device region attribute value and an attribute value for selecting a candidate canonical map close to the device location.

[0480] Any one or more of the location options may be used to filter the canonical map sets to those that may represent an area that includes the location of the user device. In some embodiments, the canonical map may cover a relatively large area of ​​the physical world. The canonical map may be segmented into regions so that selection of a map may require selection of a map region. For example, a map region may be on the order of tens of square meters. Thus, the filtered canonical map set may be a region set of maps.

[0481] According to some embodiments, a positioning snapshot can be constructed from candidate canonical maps, posture features, and sensor calibration data. For example, an array of candidate canonical maps, posture features, and sensor calibration information can be sent with a request to determine a specific matching canonical map. Matching with a canonical map can be performed based on descriptors received from a device and stored PCF data associated with the canonical map.

[0482] In some embodiments, a feature set from the device is compared to a feature set stored as part of the canonical map. The comparison can be based on feature descriptors and / or gestures. For example, a candidate feature set for the canonical map can be selected based on the number of features in the candidate set whose descriptors are similar enough to the descriptors of the feature set from the device that they are likely to be the same features. For example, the candidate set can be features derived from an image frame used to form the canonical map.

[0483] In some embodiments, if the number of similar features exceeds a threshold, further processing can be performed on the candidate feature set. Further processing can determine the degree to which the gesture feature set from the device can be aligned with the candidate feature set. Posing can be performed for a feature set from the canonical map that is similar to the features from the device.

[0484] In some embodiments, features are formatted as high-dimensional embeddings (e.g., DSF, etc.) and can be compared using a nearest neighbor search. In one example, the system is configured (e.g., by executing process 6200 and / or 6300) to find the first two nearest neighbors using Euclidean distance, and a ratio test can be performed. If the nearest neighbor is closer than the second nearest neighbor, the system considers the nearest neighbor to be a match. For example, "closer" in this context can be determined by the ratio of the Euclidean distance relative to the second nearest neighbor exceeding a threshold multiple than the ratio of the Euclidean distance relative to the nearest neighbor. Once a feature from a device is considered to "match" a feature in a canonical map, the system can be configured to calculate a relative transformation using the posture of the matching feature. The transformation developed from the posture information can be used to indicate the transformation required to locate the device to the canonical map.

[0485] The number of inliers can be used as an indication of the quality of the match. For example, in the case of DSF matching, the number of inliers reflects the number of features that matched between the received descriptor information and the stored / canonical map. In further embodiments, the inliers determined in this embodiment can be determined by counting the number of "matched" features in each set.

[0486] Indications of the quality of the match may alternatively or additionally be determined in other ways. In some embodiments, for example, when a transformation is calculated to locate a map from a device that may contain multiple features to a canonical map based on the relative poses of the matching features, the transformation statistics calculated for each of the multiple matching features may serve as an indication of quality. For example, a larger difference may indicate a poor match quality. Alternatively or additionally, for the determined transformation, the system may calculate the average error between features with matching descriptors. The average error may be calculated for the transformation, reflecting the degree of position mismatch. Mean square error is a specific example of an error metric. Regardless of the specific error metric, if the error is below a threshold, it can be determined that the transformation can be used for the feature received from the device, and the calculated transformation is used to locate the device. Alternatively or additionally, the number of inner layers may also be used to determine whether there is a map that matches the descriptor received from the device and / or the location information of the device.

[0487] As described above, in some embodiments, the device may send multiple feature sets for positioning. Positioning can be considered successful when at least a threshold number of feature sets match the feature set from the specification map with an error below a threshold and / or a number of inner layers above a threshold. The threshold number can be, for example, three feature sets. However, it should be understood that the threshold for determining whether a sufficient number of feature sets have a suitable value can be determined empirically or in other suitable ways. Similarly, other thresholds or parameters of the matching process, such as the similarity between feature descriptors considered to be matched, the number of inner layers used to select candidate feature sets, and / or the size of the mismatch error, can be similarly determined empirically or in other suitable ways.

[0488] Once a match is determined, a set of persistent map features associated with the matched one or more canonical maps is identified. In embodiments where the match is based on map regions, the persistent map features may be map features in the matching regions. The persistent map features may be persistent gestures or PCFs as described above. In the example of FIG. 63 , the persistent map features are persistent gestures.

[0489] Regardless of the format of the persistent map features, each persistent map feature can have a predetermined orientation relative to the canonical map to which it belongs. This relative orientation can be applied to a calculated transformation to align the feature set from the device with the feature set from the canonical map, thereby determining the transformation between the feature set from the device and the persistent map feature. Any adjustments, such as those that may come from calibration data, can be applied to the calculated transformation. The resulting transformation can be a transformation between the local coordinate frame of the device and the persistent map feature. This calculation can be performed for each persistent map feature that matches the map area, and the results can be stored in a table, represented in 6326 as persistent_pose_table.

[0490] In one example, block 6326 returns a table of persistent pose transforms, canonical map identifiers, and inlier levels. According to some embodiments, a canonical map ID is an identifier used to uniquely identify a canonical map and canonical map version (or region of a map, in embodiments where positioning is based on map regions).

[0491] In various embodiments, at 6328, the calculated positioning data may be used to populate positioning statistics and telemetry maintained by the positioning service. This information may be stored for each device and may be updated for each positioning attempt and may be cleared when the device's session ends. For example, maps that have been matched by a device may be used to improve map ranking operations. For example, maps covering the same area that the device previously matched may be prioritized in the ranking. Similarly, maps covering adjacent areas may be given higher priority than more remote areas. In addition, adjacent maps may be prioritized based on the detected trajectory of the device over time, with map areas in the direction of motion being given higher priority than other map areas. The positioning service may use this information, for example, to limit the maps or map areas searched for a candidate feature set in a stored canonical map based on subsequent positioning requests from the device. If a match with a low error metric and / or a large number or percentage of inner layers is identified in the limited area, processing of maps outside the area may be avoided.

[0492] Process 6300 may continue with the transfer of information from the cloud (e.g., 6354) to the user device (e.g., 6352). According to one embodiment, at 6330, the persistent gesture table and the canonical map identifier are transferred to the user device. In one example, the persistent gesture table may be composed of elements including at least a string identifying the persistent gesture ID and a transformation that links the device's tracking map with the persistent gesture. In embodiments where the persistent map feature is a PCF, the table may instead indicate a transformation to a PCF that matches the map.

[0493] If positioning fails at 6336, process 6300 continues by adjusting parameters that can increase the amount of data sent from the device to the positioning service to increase the chance of successful positioning. For example, a failure can be indicated when a feature set with more than a threshold number of similar descriptors cannot be found in the canonical map, or when an error metric associated with all transformed candidate feature sets is above a threshold. As an example of a parameter that can be adjusted, a size constraint on the descriptor buffer can be increased (6319). For example, in the case of a descriptor buffer size of 5, a positioning failure can trigger an increase to at least six feature sets extracted from at least six image frames. In some embodiments, process 6300 may include a descriptor buffer increment value. In one example, the increment value can be used to control the rate at which the buffer size is increased, for example, in response to a positioning failure. Other parameters, such as parameters that control the rate of positioning requests, can be changed when a matching canonical map cannot be found.

[0494] In some embodiments, the execution of 6300 may generate an error condition at 6340, which includes the execution of a positioning request failing to work rather than returning an unmatched result. For example, an error may occur due to a network error causing the storage holding the canonical map database to be unavailable to the server performing the positioning service, or a request received for the positioning service containing incorrectly formatted information. In the case of an error condition, in this example, the process 6300 schedules a retry of the request at 6342.

[0495] When the positioning request succeeds, any parameters adjusted in response to the failure may be reset. At 6332, process 6300 may continue to operate to reset the frequency parameters to any default values ​​or baselines. In some embodiments, 6332 is performed regardless of any changes, thereby ensuring that a baseline frequency is always established.

[0496] At 6334, the device may use the received information to update the cached positioning snapshot. According to various embodiments, the corresponding transformations, canonical map identifiers, and other positioning data may be stored by the device and used to correlate positions specified relative to a canonical map, or their persistent map features such as persistent gestures or PCFs, with positions determined by the device relative to its local coordinate frame (such as may be determined from its tracking map).

[0497] Various embodiments of the process for positioning in the cloud can implement any one or more of the aforementioned steps and be based on the aforementioned architecture. Other embodiments can combine various one or more of the aforementioned steps, performing the steps simultaneously, in parallel, or in another order.

[0498] According to some embodiments, the location service in the cloud in the context of a cross-reality experience may include additional functionality. For example, canonical map caching may be performed to resolve connectivity issues. In some embodiments, the device may periodically download and cache canonical maps that it has located. If the location service in the cloud is not available, the device may perform location positioning itself (e.g., as described above - including information about Fig.26 ). In other embodiments, the transformations returned from a positioning request can be chained together and applied to subsequent sessions. For example, a device can cache a series of transformations and use the sequence of transformations to establish a position fix.

[0499] Various embodiments of the system may use the results of the positioning operation to update the transformation information. For example, the positioning service and / or the device may be configured to maintain the state information on the tracking map to the canonical map transformation. The received transformation may be averaged over time. According to one embodiment, the averaging operation may be limited to occur after a threshold number of positioning successes (e.g., three, four, five or more times). In further embodiments, other state information may be tracked in the cloud, such as through a traversable world module. In one example, the state information may include a device identifier, a tracking map ID, a canonical map reference (e.g., version and ID), and a transformation from a canonical map to a tracking map. In some examples, the system may use the state information to continuously update and obtain a more accurate canonical map to track map transformations with each execution of a cloud-based positioning function.

[0500] Additional enhancements to cloud-based positioning can include transmitting to the device outliers in a feature set that do not match features in a canonical map. The device can use this information, for example, to improve its tracking map, such as by removing outliers from the feature set used to construct its tracking map. Alternatively or additionally, information from the positioning service can enable the device to limit a bundle adjustment for its tracking map to a calculated adjustment based on inlier features or otherwise impose constraints on the bundle adjustment process.

[0501] According to another embodiment, various sub-processes or additional operations may be used in conjunction with and / or as an alternative to the processes and / or steps discussed for cloud-based positioning. For example, candidate map identification may include identifying a map based on the corresponding Figure 1 The canonical map is accessed using the region identifiers and / or region attributes stored therein.

[0502] Depth Correspondence

[0503] Methods and apparatus are described herein for efficiently and accurately finding matching feature point sets, such as may occur when localizing an XR device in real time in a large-scale environment. Accordingly, matching feature sets as part of localization is described herein to illustrate techniques that may result in fast and accurate matching. Some or all of these techniques may be applied when matching feature sets in other contexts, such as when searching for matches between a canonical map and a portion of a tracking map as part of a map merging process.

[0504] Localizing the XR device may require comparing to find a match between a set of 2D features from one or more images captured by the XR device and a set of feature points, which may be 3D map points in a stored canonical map. A map of a large-scale environment may include a large number of 3D map points.

[0505] Compared to 2D image features, some 3D map points may be captured at different times of the day or in different seasons. Different dimensions, different lighting conditions, and other conditions make it more difficult to accurately find a matching feature set. For example, in large-scale and ultra-large-scale environments, accurate positioning may require comparing a large number of 2D feature sets to provide accurate positioning results. Therefore, positioning XR devices in large-scale and ultra-large-scale environments takes more time and consumes more computing power, resulting in delays in displaying virtual content and affecting the realism of the XR experience.

[0506] The inventors have recognized and appreciated methods and apparatus for using a subset of features with matching descriptors to search for a matching feature set to reduce time and improve accuracy in localizing XR devices in large-scale and ultra-large-scale environments. The system may include a component that evaluates the likelihood that a pair of features having matching descriptions in a subset will result in finding a matching feature set.

[0507] In some embodiments, a positioning service directed by a component when a subset of featur...

Claims

1. A method for merging one or more environment maps stored in a database with a tracking map calculated based on sensor data collected by a device worn by a user, the method comprising: receiving the tracking map from the device, wherein the tracking map is aligned with respect to a direction of gravity; determining a transformation between the tracking map and an environment map; determining whether to merge the environment map with the tracking map, wherein determining whether to merge comprises: determining whether applying a transformation to the tracking map produces a transformed tracking map that is aligned with respect to the gravity direction; and Based on determining that the transformed tracking map is aligned with the gravity direction, the environment map is merged with the tracking map.

2. The method according to claim 1, wherein: Determining the transformation includes: for corresponding features in the tracking map and the environment map, selecting as the determined transformation a transformation that aligns the corresponding features with an error metric below a threshold.

3. The method according to claim 2, further comprising: The corresponding features are determined based on similarities of the identifiers assigned to the features.

4. The method according to claim 2, wherein: Selecting as the determined transformation further comprises selecting a transformation that, when applied, results in the transformed tracking map being aligned relative to the gravity direction.

5. The method according to claim 1, wherein: Determining the transformation includes applying a plurality of candidate transformations to the tracking map and selecting a candidate transformation from among the plurality of candidate transformations as the determined transformation.

6. The method according to claim 5, wherein: Selecting the determined transformation further comprises selecting the candidate transformation that, when applied, produces the transformed tracking map aligned with respect to the gravity direction as the determined transformation.

7. The method according to claim 1, wherein: Determining whether applying the transformation to the tracking map produces the transformed tracking map aligned with respect to the gravity direction comprises: determining whether applying the transformation to the tracking map produces the transformed tracking map rotated by more than a threshold amount relative to the direction of gravity; In response to determining that applying the transform to the tracking map does not produce the transformed tracking map rotated relative to the direction of gravity by more than the threshold amount, selecting the transform as the determined transform; and In response to determining that applying the transformation to the tracking map produces the transformed tracking map rotated relative to the gravity direction by more than the threshold amount, the transformation is discarded.

8. The method according to claim 1, further comprising: identifying from the database a set of environment maps to be merged with the tracking map; as well as For each environment map in the environment map set: determining the transformation between the tracking map and the environment map; determining whether to merge the environment map with the tracking map; as well as Based on determining that the transformed tracking map is aligned with the gravity direction, the environment map is merged with the tracking map.

9. The method according to claim 8, further comprising: For each environment map in the set of environment maps, Based on determining that the gravity direction of the environment map is not aligned with the gravity direction of the transformed tracking map, merging the environment map with the tracking map is avoided.

10. The method according to claim 8, wherein: Identifying the set of environment maps includes: determining a region identifier associated with the tracking map; and The set of environment maps is identified from the database based at least in part on the region identifier associated with the tracking map.

Citation Information

Patent Citations

  • Localization determination for mixed reality systems

    US10812936B2

  • Fully convolutional interest point detection and description via homographic adaptation

    US20190147341A1

  • Enhanced pose determination for display device

    US20190188474A1

  • Methods and apparatuses for determining and / or evaluating localizing maps of image display devices

    US20200034624A1