Cross-reality system with location services and location-based shared content

Through distributed computing environment and positioning services, the sensors of portable devices are used to generate local coordinate systems, which enables virtual content sharing with low computing and network resource consumption, solves the high resource consumption problem of existing cross-reality systems, and provides an efficient multi-user immersive experience.

CN120599183APending Publication Date: 2025-09-05MAGIC LEAP INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510672878.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-11-12
Filing Date
2020-11-11
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing cross-reality systems consume high computing resources and network bandwidth when constructing and updating representations of the physical world around users, making it difficult to provide efficient virtual content sharing and multi-user immersive experience.

Method used

Through a distributed computing environment and positioning services, location-based virtual content sharing is provided, sensors of multiple portable electronic devices are used to generate a local coordinate system, and storage maps and data structures are obtained through the network to achieve selective rendering and updating of virtual content, reducing computing and network burdens.

Benefits of technology

It achieves efficient virtual content sharing and immersive cross-reality experience among multiple users with low computing and network resource consumption, reduces device battery consumption and heat generation, and improves system latency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599183A_ABST
    Figure CN120599183A_ABST
Patent Text Reader

Abstract

A cross-reality system enables any of a plurality of devices to efficiently render shared location-based content. The cross-reality system may include a cloud-based service that responds to a request from a device to locate relative to a stored map. The service may return information to the device that locates the device relative to the stored map. In conjunction with the positioning information, the service may provide information related to a location in the physical world proximate to the device that has been provided with the virtual content. Based on information received from the service, the device may render or stop rendering the virtual content to each of the plurality of users based on the location of the user and the specified location of the virtual content.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application with an application date of November 11, 2020, application number 202080078518.0, and name "Cross-reality system with positioning services and location-based shared content".

[0002] Cross-reference to related applications

[0003] This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Patent Application No. 62 / 934,485, filed on November 12, 2019, entitled “CROSS REALITY SYSTEM WITH LOCALIZATION SERVICE AND SHARED LOCATION-BASED CONTENT,” the entire contents of which are incorporated herein by reference. Technical Field

[0004] The present application relates generally to cross-reality systems. Background Art

[0005] A computer can control a human user interface to create a cross-reality (XR) environment, wherein part or all of the XR environment perceived by the user is generated by a computer. These XR environments can be virtual reality (VR) environments, augmented reality (AR) environments, and mixed reality (MR) environments, wherein part or all of the XR environment can be generated by a computer in part using data describing the environment. The data can, for example, describe a virtual object that can be rendered in such a way that the user feels or perceives the virtual object as part of the physical world and can interact with the virtual object. Because the data is rendered and presented by a user interface device (such as a head-mounted display device), the user can experience these virtual objects. The data can be displayed for the user to view, or can control audio that is played for the user to hear, or can control a tactile (or haptic) interface, allowing the user to experience the tactile sensation that is felt or perceived when the user feels the virtual object.

[0006] XR systems can be used in a wide range of applications, including scientific visualization, medical training, engineering design and prototyping, remote operation and telepresence, and personal entertainment. Compared to VR, AR and MR involve one or more virtual objects that are associated with real objects in the physical world. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment of XR systems and opens the door to a wide range of applications that present realistic and easily understood information about how the physical world may be altered.

[0007] To realistically render virtual content, an XR system can construct a representation of the physical world surrounding the user of the system. For example, this representation can be constructed by processing images acquired using sensors on a wearable device that forms part of the XR system. In such a system, the user can perform an initialization routine by viewing the room or other physical environment in which the user intends to use the XR system, until the system has enough information to construct a representation of that environment. As the system runs and the user moves around the environment or to other environments, sensors on the wearable device may acquire additional information to expand or update the representation of the physical world. Summary of the Invention

[0008] Some aspects of the present application relate to methods and apparatus for providing cross-reality (XR) scenes. The techniques described herein can be used together, separately, or in any suitable combination.

[0009] According to some aspects, a network resource within a distributed computing environment is provided for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a 3D environment. The resource comprises one or more processors and at least one computer-readable medium, the at least one computer-readable medium comprising a plurality of stored maps of the 3D environment and a plurality of data structures, each of the plurality of data structures representing a respective area in the 3D environment where virtual content is to be displayed. Each of the plurality of data structures comprises information associating the data structure with a location in the plurality of stored maps and a link to virtual content to be rendered in the respective area in the 3D environment. The computer-readable medium further comprises computer-executable instructions. When executed by one of the one or more processors, the instructions implement a service for providing location information to portable electronic devices among the plurality of portable electronic devices, wherein the location information indicates the location of the plurality of portable electronic devices relative to one or more shared maps, and selectively providing a copy of at least one of the plurality of data structures to a portable electronic device among the plurality of portable electronic devices based on the location of the portable electronic device relative to the areas represented by the plurality of data structures.

[0010] According to some embodiments, the computer-executable instructions, when executed by the processor, further implement an authentication service for determining access rights for the portable electronic device. Furthermore, the computer-executable instructions for selectively providing the at least one data structure to the portable electronic device may determine whether to send the at least one data structure based in part on the access rights for the portable electronic device and access attributes associated with the at least one data structure.

[0011] According to some embodiments, each of the plurality of data structures in the network resource further includes a common attribute. In addition, the computer-executable instructions for selectively providing the at least one data structure to the portable electronic device may determine whether to send the at least one data structure based in part on the common attribute of the at least one data structure.

[0012] According to some embodiments, for a portion of the plurality of data structures, the link to the virtual content comprises a link to an application providing the virtual content.

[0013] According to some embodiments, each of the plurality of data structures further comprises display characteristics of a prism on the portable electronic device.The prism is a volume in which the virtual content linked to the data structure is displayed.

[0014] According to some embodiments, the display characteristic may include a size of the prism.

[0015] According to some embodiments, the display characteristics include behavior of virtual content rendered within the prism relative to a physical surface.

[0016] In accordance with some embodiments, the display characteristics include one or more of: an offset of the prism relative to a persistent position associated with a map, a spatial orientation of the prism, a behavior of virtual content rendered within the prism relative to the position of the portable electronic device, and a behavior of virtual content rendered within the prism relative to a direction facing the portable electronic device.

[0017] According to some aspects, a method for operating a portable electronic device to render virtual content in a 3D environment is provided. The method includes using one or more processors to: generate a local coordinate system on the portable electronic device based on outputs of one or more sensors on the portable electronic device; generate information indicating a location in the 3D environment on the portable electronic device based on the outputs of the one or more sensors and an indication of the location in the local coordinate system; send the information indicating the location and the indication of the location in the local coordinate system to a positioning service over a network; obtain a transformation between a coordinate system storing spatial information about the 3D environment and the local coordinate system from the positioning service; obtain one or more data structures from the positioning service, each data structure representing a corresponding area in the 3D environment and virtual content for display in the corresponding area; and render the virtual content represented in the one or more data structures in the corresponding area of ​​the one or more data structures.

[0018] According to some embodiments, rendering virtual content in the corresponding area includes creating a prism having a parameter set based on the data structure representing the corresponding area.

[0019] According to some embodiments, the virtual content is represented in at least one of the one or more data structures as an indicator of the location of the virtual content on the network.

[0020] According to some embodiments, rendering the virtual content includes executing an application that generates the virtual content on the portable electronic device.

[0021] According to some embodiments, rendering the virtual content further comprises determining whether the application is currently installed on the portable electronic device, and based on determining that the application is not currently installed, downloading the application to the portable electronic device.

[0022] According to some embodiments, the method further comprises: detecting that the portable electronic device has left an area represented by a data structure of the one or more data structures; and based on the detection, deleting the virtual content represented in the data structure.

[0023] According to some embodiments, the received one or more data structures include a first set of data structures, and the first set of data structures is received at a first time. The method may further include: storing rendering information associated with a first data structure in the first set of data structures; receiving a second set of data structures at a second time after the first time; and deleting the rendering information associated with the first data structure based on determining that the first data structure is not included in the second set.

[0024] According to some aspects, an electronic device configured to operate within a cross-reality system is provided. The electronic device includes: one or more sensors configured to capture information about a three-dimensional (3D) environment, the captured information including a plurality of images; at least one processor; and at least one computer-readable medium storing computer-executable instructions. The computer instructions, when executed on a processor in the at least one processor, include: maintaining a local coordinate system for representing a position in the 3D environment based on at least a first portion of the plurality of images; managing a prism associated with one or more applications so that virtual content generated by an application in the one or more applications is rendered within the prism; and sending information derived from the output of the one or more sensors to a service over a network. The instructions also receive from the service: positioning information; and a data structure representing corresponding virtual content and an area in the 3D environment for rendering the virtual content. The instructions also include creating a prism associated with the data structure so that the corresponding virtual content is rendered within the prism.

[0025] According to some embodiments, the computer-executable instructions further include computer-executable instructions for: retrieving the corresponding virtual content based on the information in the data structure; and rendering the retrieved virtual content within the prism.

[0026] According to some embodiments, obtaining the corresponding virtual content based on the information in the data structure includes accessing the corresponding virtual content through a network based on a virtual content location indicator in the data structure.

[0027] According to some embodiments, obtaining the corresponding virtual content based on the information in the data structure includes accessing the corresponding virtual content through the network based on the virtual content location indicator in the data structure.

[0028] According to some embodiments, the computer-executable instructions further include instructions for performing the following operations: detecting that the electronic device has left an area represented by the data structure; and based on the detection, deleting a prism Prism associated with the data structure.

[0029] According to some embodiments, rendering the acquired virtual content within the prism further comprises: using a coordinate system of the electronic device to determine a set of coordinates for rendering the virtual content in the 3D environment.

[0030] According to some aspects, a method for curating location-based virtual content for a cross-reality system operable with a plurality of portable electronic devices capable of rendering virtual content in a 3D environment is provided. The method includes: generating, using one or more processors, a data structure representing an area in the 3D environment in which virtual content is to be displayed; storing, using one or more processors, in the data structure information indicating virtual content to be rendered in the area in the 3D environment; and associating, using one or more processors, the data structure with locations in a map for locating the plurality of portable electronic devices in a shared coordinate system.

[0031] According to some embodiments, the method further comprises setting access rights to the data structure.

[0032] According to some embodiments, setting access rights regarding the data structure includes indicating that the data structure may be accessed by one or more particular categories of users of the plurality of portable electronic devices.

[0033] According to some embodiments, storing in the data structure information indicative of virtual content to be rendered in the area of ​​the 3D environment comprises specifying an application executable on the portable electronic device to generate the virtual content.

[0034] According to some embodiments, the method further comprises storing the data structure in association with a location service that locates the plurality of portable electronic devices using the map.

[0035] According to some embodiments, the method further comprises: receiving, through a user interface, a designation of the area of ​​the 3D environment and the virtual content.

[0036] According to some embodiments, the method further comprises: receiving, from an application via a programming interface, a designation of the area of ​​the 3D environment and the virtual content.

[0037] The foregoing summary is provided by way of illustration and is not intended to be limiting. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component illustrated in various figures is represented by a like reference numeral. For clarity, not every component is labeled in every figure. In the drawings:

[0039] Figure 1 is a sketch showing an example of a simplified augmented reality (AR) scene according to some embodiments;

[0040] Figure 2is a sketch illustrating an exemplary simplified AR scene illustrating an exemplary use case of an XR system in accordance with some embodiments;

[0041] Figure 3 is a schematic diagram illustrating data flow for a single user in an AR system configured to provide the user with an experience of AR content that interacts with the physical world, according to some embodiments;

[0042] Figure 4 is a schematic diagram illustrating an exemplary AR display system for displaying virtual content to a single user in accordance with some embodiments;

[0043] Figure 5A is a schematic diagram illustrating a user wearing an AR display system rendering AR content as the user moves in a physical world environment in accordance with some embodiments;

[0044] Figure 5B is a schematic diagram illustrating an observation optical assembly and accompanying components according to some embodiments;

[0045] Figure 6A is a schematic diagram illustrating an AR system using a world reconstruction system according to some embodiments;

[0046] Figure 6B is a schematic diagram illustrating components of an AR system that maintains a model of a connected world according to some embodiments;

[0047] Figure 7 A schematic diagram of a tracking map formed by devices traversing paths through the physical world.

[0048] Figure 8 is a schematic diagram illustrating a user of a cross-reality (XR) system perceiving virtual content in accordance with some embodiments;

[0049] Figure 9 is transformed between coordinate systems according to some embodiments Figure 8 A block diagram of components of a first XR device of an XR system;

[0050] Figure 10 is a diagram illustrating an exemplary transformation of an origin coordinate system to a destination coordinate system for correctly rendering local XR content according to some embodiments;

[0051] Figure 11 is a top view illustrating a pupil-based coordinate system according to some embodiments;

[0052] Figure 12 is a top view illustrating a camera coordinate system including all pupil positions according to some embodiments;

[0053] Figure 13 According to some embodiments Figure 9 A schematic diagram of a display system;

[0054] Figure 14 is a block diagram illustrating the creation of a persistent coordinate frame (PCF) and the attachment of XR content to the PCF in accordance with some embodiments;

[0055] Figure 15 is a flowchart illustrating a method of establishing and using a PCF according to some embodiments;

[0056] Figure 16 is according to some embodiments including a second XR device Figure 8 Block diagram of the XR system;

[0057] Figure 17 is a schematic diagram illustrating a room and key frames established for various areas in the room according to some embodiments;

[0058] Figure 18 is a schematic diagram illustrating the establishment of a keyframe-based persistent pose according to some embodiments;

[0059] Figure 19 is a schematic diagram illustrating the establishment of a persistent coordinate system (PCF) based on a persistent pose according to some embodiments;

[0060] Figures 20A to 20C is a schematic diagram illustrating an example of creating a PCF according to some embodiments;

[0061] Figure 21 is a block diagram illustrating a system for generating global descriptors for individual images and / or maps according to some embodiments;

[0062] Figure 22 is a flowchart illustrating a method of computing an image descriptor according to some embodiments;

[0063] Figure 23 is a flow chart illustrating a localization method using image descriptors according to some embodiments;

[0064] Figure 24 is a flowchart illustrating a method of training a neural network according to some embodiments;

[0065] Figure 25 is a block diagram illustrating a method of training a neural network according to some embodiments;

[0066] Figure 26 is a schematic diagram illustrating an AR system configured to rank and merge multiple environment maps according to some embodiments;

[0067] Figure 27is a simplified block diagram illustrating a plurality of canonical maps stored on a remote storage medium according to some embodiments;

[0068] Figure 28 is a schematic diagram illustrating a method of selecting a canonical map, for example, to locate a new tracking map in one or more canonical maps and / or to obtain a PCF from a canonical map, according to some embodiments;

[0069] Figure 29 is a flow chart illustrating a method of selecting a plurality of ranked environment maps according to some embodiments;

[0070] Figure 30 is a diagram showing a method according to some embodiments Figure 26 A schematic diagram of an exemplary map ranking portion of an AR system;

[0071] Figure 31A is a diagram illustrating examples of area attributes of a Tracking Map (TM) and an environment map in a database according to some embodiments;

[0072] Figure 31B is a diagram illustrating determining a method for Figure 29 A schematic diagram of an example of a geo-location filtered Tracking Map(TM);

[0073] Figure 32 is a diagram showing a method according to some embodiments Figure 29 A schematic diagram of an example of geographic location filtering;

[0074] Figure 33 is a diagram showing a method according to some embodiments Figure 29 A schematic diagram of an example of Wi-Fi BSSID filtering;

[0075] Figure 34 is a diagram illustrating the use of Figure 29 A schematic diagram of an example of positioning;

[0076] Figure 35 and Figure 36 is a block diagram of an XR system configured to rank and merge multiple environment maps according to some embodiments;

[0077] Figure 37 is a block diagram illustrating a method of creating an environment map of the physical world in a canonical form according to some embodiments;

[0078] Figure 38A and Figure 38B is shown in accordance with some embodiments by updating the tracking map with a new Figure 7 A tracking map, a schematic diagram of the environment map created in a canonical form.

[0079] Figures 39A to 39F is a schematic diagram illustrating an example of a merged map according to some embodiments;

[0080] Figure 40 According to some embodiments, Figure 9 The first XR device to generate the first 3D local tracking map (ground Figure 1 ) two-dimensional representation;

[0081] Figure 41 is a diagram illustrating how to convert the ground Figure 1 Upload from the first XR device to Figure 9 Block diagram of the server;

[0082] Figure 42 is a diagram showing a method according to some embodiments Figure 16 A schematic diagram of an XR system showing that after the first user terminates the first session, a second user initiates a second session using a second XR device of the XR system;

[0083] Figure 43A is a diagram showing a method according to some embodiments Figure 42 A block diagram of a new session for a second XR device;

[0084] Figure 43B is a diagram showing that according to some embodiments Figure 42 A block diagram of creation of a tracking map for a second XR device;

[0085] Figure 43C is a diagram illustrating downloading a specification map from a server to a Figure 42 a block diagram of a second XR device;

[0086] Figure 44 is to show that according to some embodiments, Figure 42 The second tracking map generated by the second XR device (ground Figure 2 ) a schematic diagram of a positioning attempt to locate the canonical map;

[0087] Figure 45 It is shown that according to some embodiments, it can be further developed and has the same Figure 2 PCF associated XR content Figure 44 Second tracking map (map Figure 2 ) a schematic diagram of a positioning attempt to locate the canonical map;

[0088] Figures 46A-46B is a diagram illustrating the Figure 45 land Figure 2 Schematic diagram of successful positioning on the canonical map;

[0089] Figure 47is a diagram illustrating that according to some embodiments, Figure 46A The canonical map of one or more PCFs includes Figure 45 land Figure 2 Schematic diagram of the canonical map generated by

[0090] Figure 48 is a diagram illustrating further expansion on a second XR device according to some embodiments. Figure 2 of Figure 47 Schematic diagram of the normative map;

[0091] Figure 49 is a diagram illustrating how to convert the ground Figure 2 Block diagram of uploading from the second XR device to the server;

[0092] Figure 50 is a diagram illustrating how to convert the ground Figure 2 a block diagram for merging with the canonical map;

[0093] Figure 51 is a block diagram illustrating sending a new canonical map from a server to a first XR device and a second XR device according to some embodiments;

[0094] Figure 52 is a diagram showing a location according to some embodiments Figure 2 Two-dimensional representation and reference ground Figure 2 a block diagram of a head coordinate system of a second XR device;

[0095] Figure 53 is a block diagram illustrating in two dimensions adjustments of a head coordinate system that may occur in six degrees of freedom, according to some embodiments;

[0096] Figure 54 is a diagram illustrating the relative position of sound relative to ground according to some embodiments. Figure 2 A block diagram of a canonical map of PCF positioning on a second XR device;

[0097] Figure 55 and Figure 56 is a perspective and block diagram illustrating a use case of an XR system when a first user terminates a first session and the first user initiates a second session using the XR system, according to some embodiments;

[0098] Figure 57 and Figure 58 is a perspective view and block diagram illustrating a use case of an XR system when three users are using the XR system simultaneously in the same session, according to some embodiments;

[0099] Figure 59 is a flow chart illustrating a method of recovering and resetting head posture according to some embodiments;

[0100] Figure 60 is a block diagram of a machine in the form of a computer that may find application in the system of the present invention according to some embodiments;

[0101] Figure 61 is a diagram of an example XR system in which any of a plurality of devices may access location services in accordance with some embodiments;

[0102] Figure 62 is an example process flow for operating a portable device as part of an XR system providing cloud-based positioning, according to some embodiments; and

[0103] Figure 63A 、 Figure 63B and Figure 63C is an example process flow for cloud-based positioning according to some embodiments.

[0104] Figure 64 is a schematic diagram of a system for managing and displaying location-based shared virtual content in a physical environment, and a sketch of how exemplary content may be presented to a user in the physical environment.

[0105] Figure 65 is a schematic diagram of an exemplary volume data structure and associated data.

[0106] Figure 66 is a block diagram of an exemplary software architecture for a cross-reality device configured to retrieve and display virtual content based on the position of the cross-reality device relative to a physical environment.

[0107] Figure 67 is a flow diagram illustrating the interaction between system components for retrieving and displaying location-based shared content.

[0108] Figure 68 is an exemplary software architecture that configures a device to work with a cross-reality system so that the device can retrieve and render content from the cross-reality system.

[0109] Figure 69 is a sketch of an exemplary physical environment with location content shared to be perceivable by any of a plurality of users in the physical environment. DETAILED DESCRIPTION

[0110] This document describes methods and apparatus for providing a cross-reality (XR) scene to any of multiple users who may be traversing the physical world. The system enables content managers to specify virtual content associated with a location in the physical world, so that when a user wearing an XR device passes near the location, the XR device can render the content for the user, causing the content to appear at the specified location in the physical world.

[0111] Such a system can be implemented to efficiently operate services that provide virtual content, as well as interact with these services to render virtual content to the user's XR device. In some embodiments, the virtual content can be provided by a positioning service that enables each of multiple XR devices to determine its location relative to a shared map. According to some embodiments, a volume can be defined in the shared map in which location-based virtual content will be rendered. Due to the positioning process, when the service determines that the XR device is located at a location within the map that has such a volume associated with it, the service can provide the XR device with an indication of the volume and the content to be displayed within the volume. The device can then use these indications to render the virtual content.

[0112] The information generated during positioning can be similarly used to remove content. When the XR device has moved away from the location where the volume will be rendered, the device can delete information indicating the location and nature of the virtual content. The movement of the XR device can be determined by the service that specifies the content or by a service on the device. Because positioning can be performed repeatedly to support other functions of the XR system, the computational burden and network bandwidth of identifying and / or removing location-based virtual content can be reduced.

[0113] In some embodiments, information about location-based virtual content can be represented in an efficient format when it is passed from a service to an XR device. For example, virtual content can be represented as a link to virtual content or a link to an application that generates virtual content. Therefore, less network bandwidth is consumed by the communication of content between the service and the XR device, allowing content information to be updated frequently. In addition, the volume in which the virtual content is displayed can correspond to a structure used by the XR device to manage the rendering of content specified by an application executed on the XR device. An example of an application executed on the XR device is prism. The XR device can use utilities that manage prism in other aspects to manage the rendering of location-based virtual content in combination with other virtual content. For example, when multiple applications specify virtual content for the same volume in the physical world, these utilities can determine the content to be rendered, associate user actions with the specific application that provides the virtual content to be rendered, and remove the virtual content and associated data from the XR device when the virtual content is no longer rendered.

[0114] The localization process (which can be used to identify location-based virtual content) is used for some functions of XR systems, such as providing realistic shared experiences for multiple users. In order to provide realistic XR experiences to multiple users, the XR system must understand the users' physical environment in order to correctly associate the positions of virtual objects with respect to real objects. The XR system can build an environmental map of the scene, which can be created by utilizing images and / or depth information collected by sensors that are part of the XR device worn by the user of the XR system.

[0115] In an XR system, each XR device can develop a local map of its physical environment by integrating information from one or more images collected at a point in time during a scan. In some embodiments, the coordinate system of this map is associated with the device's orientation when the scan begins. As a user interacts with the XR system, the orientation changes from session to session, whether different sessions are associated with different users (each with their own wearable device and sensors scanning the environment) or the same user using the same device at different times. The inventors have recognized and understood techniques for operating an XR system based on persistent spatial information that overcome the limitations of XR systems in which each user device relies solely on spatial information collected relative to an orientation that varies for different user instances of the system (e.g., a snapshot in time) or sessions (e.g., between being turned on and off). For example, these techniques can provide XR scenes with more efficient computation and a more immersive experience for single or multiple users by allowing any of multiple users of the XR system to create, store, and retrieve persistent spatial information.

[0116] Persistent spatial information can be represented by a persistent map, which can enable one or more functions that enhance the XR experience. The persistent map can be stored in a remote storage medium (e.g., the cloud). For example, after being turned on, a wearable device worn by a user can retrieve a previously created and stored appropriate stored map from a persistent storage such as cloud storage. The previously stored map may have been based on data about the environment collected using sensors on the user's wearable device during a previous session. Retrieving the stored map can allow the wearable device to be used without scanning the physical world through sensors on the wearable device. Alternatively or additionally, the system / device can similarly retrieve an appropriate stored map when entering a new area of ​​the physical world.

[0117] The stored map can be represented in a canonical form that is relative to the local reference frame on each XR device. In a multi-device XR system, the stored map accessed by one device may have been created and stored by another device, and / or may have been constructed by aggregating data about the physical world collected by sensors on multiple wearable devices, which previously existed in the portion of the physical world represented by the stored map.

[0118] The relationship between the canonical map and each device's local map may be determined through a positioning process. The positioning process may be performed on each XR device based on a set of canonical maps that are selected and sent to the device. However, the inventors have recognized and appreciated that network bandwidth and computing resources on the XR devices may be reduced by providing a positioning service that may be executed on a remote processor (such as may be implemented in the cloud). As a result, battery consumption and heat generation on the XR devices may be reduced, enabling the devices to devote resources such as computing time, network bandwidth, battery life, and thermal budgets to providing a more immersive user experience. However, through appropriate selection of the information transmitted between each XR device and the positioning service, positioning may be performed with the latency and accuracy required to support such an immersive experience.

[0119] Sharing data about the physical world across multiple devices enables a shared user experience for virtual content. For example, two XR devices with access to the same stored map can both be positioned relative to the stored map. Once positioned, the user device can render virtual content at that location by transforming the location specified by the reference stored map into a reference frame maintained by the user device. The user device can use this local reference frame to control the user device's display to render virtual content at the specified location.

[0120] To support these and other features, the XR system may include components that develop, maintain, and use persistent spatial information (including one or more stored maps) based on data about the physical world collected using sensors on the user device. These components may be distributed throughout the XR system, with some components operating, for example, on a head-mounted portion of the user device. Other components may operate on a computer associated with the user coupled to the head-mounted portion via a local area network or a personal area network. Still other components may operate at a remote location, such as one or more servers accessible via a wide area network.

[0121] For example, these components may include components that are capable of identifying information about the physical world collected by one or more user devices that is of sufficient quality to be stored as a persistent map or stored in a persistent map. An example of such a component, described in more detail below, is a map merging component. For example, such a component may receive input from a user device and determine the suitability of a portion of the input for updating a persistent map. For example, the map merging component may split a local map created by a user device into multiple parts, determine the mergeability of one or more parts with a persistent map, and merge the parts that meet the merge criteria into the persistent map. For example, the map merging component may also promote parts that were not merged with the persistent map to a separate persistent map.

[0122] As another example, these components can include components that help determine appropriate persistent maps that can be retrieved and used by a user device. An example of such a component, described in more detail below, is a map ranking component. For example, such a component can receive input from a user device and identify one or more persistent maps that may represent the area of ​​the physical world in which the device is operating. For example, the map ranking component can help select a persistent map to be used by a local device when rendering virtual content, collecting data about the environment, or performing other actions. Alternatively or additionally, the map ranking component can help identify persistent maps to be updated as additional information about the physical world is collected by one or more user devices.

[0123] Additional components may determine a transformation that transforms information captured or described with respect to one reference frame into another reference frame. For example, a sensor may be attached to a head-mounted display so that data read from the sensor indicates the location of objects in the physical world relative to the wearer's head pose. One or more transformations may be applied to associate the position information with the associated coordinate system of the persistent environment map. Similarly, data indicating the location where a virtual object will be rendered, expressed in the coordinate system of the persistent environment map, may undergo one or more transformations to be located in the reference frame of the display on the user's head. As described in more detail below, multiple such transformations may be required. These transformations may be partitioned across XR system components so that they can be efficiently updated and / or applied to a distributed system.

[0124] In some embodiments, a persistent map can be constructed based on information collected by multiple user devices. The XR device can capture local spatial information and construct a separate tracking map using information collected by sensors of each of the XR devices at different locations and times. Each tracking map can include points, each of which can be associated with a feature of a real-world object that can include multiple features. In addition to potentially providing input for creating and maintaining a persistent map, the tracking map can also be used to track the user's movement in the scene, thereby enabling the XR system to estimate the corresponding user's head pose based on the tracking map.

[0125] XR systems can operate using techniques that provide XR scenes to obtain a highly immersive user experience, such as estimating head pose at a frequency of 1kHz, with low usage of computing resources associated with the XR device, which can be configured with, for example, four video graphics array (VGA) cameras operating at a frequency of 30Hz, an inertial measurement unit (IMU) operating at a frequency of 1kHz, the computing power of a single Advanced RISC Machine (ARM) core, less than 1GB of memory, and less than 100Mbps of network bandwidth. These techniques involve reducing the processing required to generate and maintain maps and estimate head pose, and involve providing and using data with low computational overhead. The XR system can calculate its pose based on matched visual features. U.S. patent application No. 16 / 221,065 describes hybrid tracking, the entire contents of which are incorporated herein by reference.

[0126] These techniques may include reducing the amount of data processed when building a map, such as by constructing a sparse map with a set of plot points and keyframes and / or dividing the map into blocks to enable block-by-block updates. Plot points can be associated with points of interest in the environment. Keyframes include information selected from camera-captured data. U.S. Patent Application No. 16 / 520,582, the entire contents of which are incorporated herein by reference, describes determining and / or evaluating a localization map.

[0127] In some embodiments, persistent spatial information can be represented in a manner that is easily shared between users and between distributed components including applications. For example, information about the physical world can be represented as a persistent coordinate frame (PCF). The PCF can be defined based on one or more points that represent features identified in the physical world. Features can be selected so that they are the same between XR system user sessions. The PCF can exist sparsely, providing less than all available information about the physical world, thereby efficiently processing and sending this information. Techniques for processing persistent spatial information can include creating dynamic maps based on one or more coordinate frames in real space across one or more sessions, and also include generating a persistent coordinate frame (PCF) on the sparse map, which can be exposed to XR applications via an application programming interface (API), for example, supported by techniques for ranking and merging multiple maps created by one or more XR devices. Persistent spatial information can also quickly restore and reset head pose on each of one or more XR devices in a computationally efficient manner.

[0128] Furthermore, these techniques can enable efficient comparison of spatial information. In some embodiments, an image frame can be represented by a digital descriptor. The descriptor can be computed via a transformation that maps a set of features identified in the image to the descriptor. The transformation can be performed within a trained neural network. In some embodiments, the feature set provided as input to the neural network can be a filtered feature set extracted from the image using, for example, a technique that preferentially selects features that are likely to persist.

[0129] Representing image frames as descriptors allows, for example, efficient matching of new image information with stored image information. The XR system can store descriptors of one or more frames under the persistent map in conjunction with the persistent map. Local image frames acquired by the user device can similarly be converted to such descriptors. By selecting stored maps with descriptors similar to the descriptors of the local image frames, one or more persistent maps that are likely to represent the same physical space as the user device can be selected with a relatively small amount of processing. In some embodiments, descriptors of key frames in the local map and the persistent map can be calculated, thereby further reducing processing when comparing the maps. For example, such efficient comparison can be used to simplify looking up a persistent map to be loaded into the local device or for looking up a persistent map to be updated based on image information acquired with the local device.

[0130] The techniques described herein can be used with a variety of device types, together or individually, and in a variety of scenarios, including wearable or portable devices with limited computing resources that provide augmented or mixed reality scenarios. In some embodiments, these techniques can be implemented by one or more services that form part of an XR system.

[0131] AR System Overview

[0132] Figure 1 and Figure 2 A scene with virtual content displayed in conjunction with a portion of the physical world is shown. For illustration purposes, an AR system is used as an example of an XR system. Figure 3-6B An exemplary AR system is shown, including one or more processors, memory, sensors, and a user interface that can operate in accordance with the techniques described herein.

[0133] refer to Figure 1, shows an outdoor AR scene 354 in which a user of AR technology sees a physical-world park-like environment 356 that includes people, trees, buildings in the background, and a concrete platform 358 feature. In addition to these items, the user of AR technology also perceives that they "see" a robotic statue 357 standing on the physical-world concrete platform 358, as well as a flying cartoon avatar character 352 that appears to be a humanoid bumblebee, even though these elements (e.g., avatar character 352 and robotic statue 357) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is challenging to create AR technology that promotes a comfortable, natural, and rich presentation of virtual image elements within other virtual or physical-world image elements.

[0134] Such an AR scene can be implemented using a system that builds a map of the physical world based on tracking information, allowing users to place AR content in the physical world, determine the location of the placed AR content in the map of the physical world, maintain the AR scene so that the placed AR content can be reloaded to be displayed in the physical world during, for example, a different AR experience session, and allow multiple users to share the AR experience. The system can build and update a digital representation of the physical world surface around the user. This representation can be used to render virtual content so that it appears to be completely or partially obscured by physical objects between the user and the rendered location of the virtual content, for placing virtual objects in physics-based interactions, and for virtual character path planning and navigation, or for other operations that use information about the physical world.

[0135] Figure 2 Another example of an indoor AR scene 400, in accordance with some embodiments, illustrates an exemplary XR system use case. The exemplary scene 400 is a living room, which features a wall, a bookshelf located on one side of the wall, a floor lamp located in a corner of the room, a floor, a sofa located on the floor, and a coffee table. In addition to these physical objects, the user of AR technology can also perceive virtual objects, such as an image on the wall behind the sofa, a bird flying through a door, a deer peeking out from a bookshelf, and a decorative pinwheel placed on the coffee table.

[0136] For an image on a wall, AR technology needs information not only about the surface of the wall, but also about objects and surfaces in the room (such as the shape of a lamp) that obstruct the image to render the virtual object correctly. For a flying bird, AR technology needs information about all objects and surfaces in every corner of the room to render the bird avoiding objects and surfaces with realistic physics, or the bird rebounding when it hits these objects and surfaces. For a deer, AR technology needs information about surfaces such as the floor or coffee table to calculate where to place the deer. For a windmill, the system can identify it as a separate object from the table and determine that it is movable, while the corner of a bookshelf or the corner of a wall is determined to be stationary. This distinction can be used to determine which parts of the scene to use or update in each of the various operations.

[0137] Virtual objects can be placed in a previous AR experience session. When a new AR experience session starts in the living room, AR technology requires that the virtual objects be displayed accurately in the previously placed location and realistically visible from different angles. For example, a windmill should be displayed standing on a book, rather than floating somewhere else above the table without a book. This floating may occur if the user location of the new AR experience session is not accurately positioned in the living room. For another example, if the user's perspective of viewing the windmill is different from the perspective from which the windmill was placed, AR technology needs to display the corresponding side of the windmill.

[0138] The scene can be presented to the user via a system comprising multiple components, including a user interface that can stimulate one or more user senses (such as vision, sound and / or touch). In addition, the system can include one or more sensors that can measure parameters of the physical part of the scene, including the position and / or movement of the user within the physical part of the scene. In addition, the system can include one or more computing devices with associated computer hardware (such as, memory). These components can be integrated into a single device, or can be distributed among multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.

[0139] Figure 3An AR system 502 is shown configured to provide an experience of AR content interacting with a physical world 506 in accordance with some embodiments. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by a user as part of a headset, such that the user may wear the display over their eyes, such as a pair of goggles or glasses. At least a portion of the display may be transparent, such that the user may observe a see-through reality 510. The see-through reality 510 may correspond to a portion of the physical world 506 that is within a current viewpoint of the AR system 502, which corresponds to the user's viewpoint if the user is wearing a headset that includes both the AR system's display and sensors for acquiring information about the physical world.

[0140] AR content may also be presented on display 508, overlaying see-through reality 510. To provide accurate interaction between the AR content on display 508 and see-through reality 510, AR system 502 may include sensors 522 configured to capture information about the physical world 506.

[0141] Sensors 522 may include one or more depth sensors that output depth maps 512. Each depth map 512 may have multiple pixels, each of which may represent the distance to a surface in the physical world 506 in a particular direction relative to the depth sensor. Raw depth data is obtained from the depth sensors to create the depth map. Such depth maps can be updated as quickly as the depth sensors can form new images, hundreds or thousands of times per second. However, this data may be noisy and incomplete, and the depth map shown may have holes, which appear as black pixels.

[0142] The system may include other sensors, such as image sensors. Image sensors can acquire monocular or stereo information that can be processed to represent the physical world in other ways. For example, images can be processed in world reconstruction component 516 to create a mesh representing the connected parts of objects in the physical world. Metadata about these objects (including, for example, color and surface texture) can similarly be acquired using sensors and stored as part of the world reconstruction.

[0143] The system can also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, the head pose can be calculated in real time using a head pose tracking component of the system. The head pose tracking component can represent the user's head pose in a coordinate system with six degrees of freedom, including, for example, translation on three perpendicular axes (e.g., front / back, up / down, left / right), and rotation about three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, the sensor 522 can include an inertial measurement unit ("IMU") that can be used to calculate and / or determine the head pose 514. For example, the head pose 514 for a depth map can indicate the current viewpoint of the sensor capturing the depth map in six degrees of freedom, but the head pose 514 can also be used for other purposes, such as associating image information with a specific part of the physical world or associating the position of a display worn on the user's head with the physical world.

[0144] In some embodiments, head pose information may be derived in other ways besides from the IMU, such as by analyzing objects in an image. For example, the head pose tracking component may calculate the relative position and orientation of the AR device to the physical object based on the visual information captured by the camera and the inertial information captured by the IMU. The head pose tracking component may then calculate the head pose of the AR device by, for example, comparing the calculated relative position and orientation of the AR device to the physical object with features of the physical object. In some embodiments, this comparison may be performed by identifying features in images captured using one or more sensors 522 that remain stable over a period of time, such that changes in the positions of these features in images captured over a period of time can be associated with changes in the user's head pose.

[0145] In some embodiments, as a user moves around in the physical world with an AR device, the AR device can build a map based on feature points identified in consecutive images within a series of captured image frames. Although each image frame can be taken from a different posture as the user moves, the system can adjust the orientation of the features of each consecutive image frame to match the orientation of the initial image frame by matching the features of the consecutive image frame with the previously captured features. Each consecutive image frame can be aligned to match the orientation of the previously processed image frame using translation of the consecutive image frames (so that points representing the same features will match corresponding feature points from the previously collected image frame). The frames in the resulting map may have a common orientation established when the first image frame is added to the map. The map has a set of feature points located in a common reference frame that can be used to determine the user's posture in the physical world by matching features from the current image frame with the map. In some embodiments, the map can be referred to as a tracking map.

[0146] In addition to being able to track the user's posture in the environment, the map can also enable other components of the system, such as the world reconstruction component 516, to determine the position of physical objects relative to the user. The world reconstruction component 516 can receive the depth map 512 and the head pose 514, as well as any other data from the sensors, and integrate this data into the reconstruction 518. The reconstruction 518 can be more complete and less noisy than the sensor data. The world reconstruction component 516 can update the reconstruction 518 using spatial and temporal averaging of sensor data from multiple viewpoints over time.

[0147] Reconstruction 518 can include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, and the like. Different formats can represent alternative representations of the same portion of the physical world, or can represent different portions of the physical world. In the example shown, on the left side of reconstruction 518, a portion of the physical world is presented as a global surface; on the right side of reconstruction 518, a portion of the physical world is presented as a mesh.

[0148] In some embodiments, the map maintained by the head pose component 514 may be sparse relative to other maps of the physical world that may be maintained. A sparse map may indicate the locations of points of interest and / or structures, such as corners or edges, rather than providing information about the locations of surfaces and other possible features. In some embodiments, the map may include image frames captured by the sensor 522. These frames may be simplified to include features that may represent points of interest and / or structures. Information about the user's pose from which the frame was captured may also be stored as part of the map, along with each frame. In some embodiments, each image captured by the sensor may or may not be stored. In some embodiments, the system may process the images as they are collected by the sensor and select a subset of image frames for further computation. This selection may be based on one or more criteria that limit the amount of information added while ensuring that the map contains useful information. The system may add new image frames to the map based on, for example, overlap with previously added image frames or based on image frames containing a sufficient number of features determined to potentially represent stationary objects. In some embodiments, the selected image frames or groups of features from the selected image frames may serve as keyframes for the map, providing spatial information.

[0149] The AR system 502 can integrate sensor data from multiple viewpoints of the physical world over a period of time. As the device including the sensor moves, the pose (e.g., position and orientation) of the sensor can be tracked. Because the frame pose of the sensor and its relationship to other poses are known, each of these multiple viewpoints of the physical world can be fused together to form a single combined reconstruction of the physical world, which can serve as an abstraction layer for the map and provide spatial information. By using spatial and temporal averaging (i.e., averaging data from multiple viewpoints over a period of time) or any other suitable method, the reconstruction can be more complete and less noisy than the original sensor data.

[0150] exist Figure 3 In the illustrated embodiment, a map represents a portion of the physical world in which a user of a single wearable device is located. In this case, a head pose associated with a frame in the map can be represented as a local head pose, indicating an orientation relative to an initial orientation of the single wearable device at the start of the session. For example, when the device is powered on or otherwise operated to scan an environment to build a representation of that environment, the head pose can be tracked relative to the initial head pose.

[0151] In conjunction with the content representing the portion of the physical world, the map may include metadata. For example, the metadata may indicate the time at which sensor information used to form the map was captured. Alternatively or additionally, the metadata may indicate the location of the sensor at the time the information used to form the map was captured. The location may be expressed directly, such as using information from a GPS chip, or indirectly, such as using a wireless (e.g., Wi-Fi) signature (indicating the signal strength received from one or more wireless access points at the time the sensor data was collected), and / or using an identifier of the wireless access point to which the user device was connected at the time the sensor data was collected, such as a BSSID.

[0152] Reconstruction 518 can be used for AR functions, such as generating a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation can change as the user moves or objects in the physical world change. Some aspects of reconstruction 518 can be used, for example, by component 320, which generates a changing global surface representation in world coordinates that can be used by other components.

[0153] AR content is generated, such as by an AR application 504, based on the information. The AR application 504 may be a game program that, for example, performs one or more functions, such as visual occlusion, physics-based interaction, and environmental reasoning, based on information about the physical world. These functions may be performed by querying data in different formats from the reconstruction 518 generated by the world reconstruction component 516. In some embodiments, component 520 may be configured to output updates when a representation in a region of interest in the physical world changes. For example, the region of interest may be set to a portion of the physical world near a user of the system, such as within the user's field of view, or projected (predicted / determined) as being within the user's field of view.

[0154] The AR application 504 can use this information to generate and update AR content. The virtual portion of the AR content can be presented on the display 508 in combination with the see-through reality 510, creating a realistic user experience.

[0155] In some embodiments, an AR experience may be provided to a user via an XR device, which may be a wearable display device, which may be part of a system that may include remote processing and / or remote data storage and / or, in some embodiments, include other wearable display devices worn by other users.

[0156] To simplify the explanation, Figure 4 An example of a system 580 (hereinafter referred to as "system 580") is shown that includes a single wearable device. System 580 includes a head-mounted display device 562 (hereinafter referred to as "display device 562"), and various mechanical and electronic modules and systems that support the functionality of display device 562. Display device 562 can be coupled to a frame 564 that can be worn by a display system user or viewer 560 (hereinafter referred to as "user 560") and is configured to position display device 562 in front of the eyes of user 560. According to various embodiments, display device 562 can be a sequential display. Display device 562 can be a monocular or binocular display. In some embodiments, display device 562 can be Figure 3 Example of display 508 in .

[0157] In some embodiments, a speaker 566 is coupled to the frame 564 and positioned near the ear canal of the user 560. In some embodiments, another speaker (not shown) is positioned near the other ear canal of the user 560 to provide stereo / shapeable sound control. The display device 562 is operably coupled to a local data processing module 570, such as by a wired conductor or a wireless connection 568, which can be mounted in various configurations, such as being attached to the frame 564, being fixedly attached to a helmet or hat worn by the user 560, being embedded in headphones, or otherwise being removably attached to the user 560 (e.g., in a backpack configuration, in a belt-coupled configuration).

[0158] The local data processing module 570 may include a processor, as well as digital memory, such as non-volatile memory (e.g., flash memory), both of which may be used to assist in the processing, caching, and storage of data. The data may include: a) data captured from sensors that may be operably coupled to the frame 564 or otherwise attached to the user 560, such as an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a radio, and / or a gyroscope; and / or b) data acquired and / or processed using the remote processing module 572 and / or the remote data repository 574, which data may be transmitted to the display device 562 after such processing or retrieval.

[0159] In some embodiments, the wearable device can communicate with remote components. The local data processing module 570 can be operably coupled to the remote processing module 572 and the remote data repository 574 via communication links 576, 578, such as via wired or wireless communication links, respectively, so that these remote modules 572, 574 are operably coupled to each other and can be used as resources for the local data processing module 570. In further embodiments, in addition to or instead of the remote data repository 574, the wearable device can access a cloud-based remote data repository and / or service. In some embodiments, the head posture tracking component described above can be implemented at least in part in the local data processing module 570. In some embodiments, Figure 3 The world reconstruction component 516 in can be implemented at least in part in the local data processing module 570. For example, the local data processing module 570 can be configured to execute computer-executable instructions to generate a map and / or a physical world representation based at least in part on at least a portion of the data.

[0160] In some embodiments, processing can be distributed across local and remote processors. For example, local processing can be used to build a map (e.g., a tracking map) on the user device based on sensor data collected using sensors on the user device. Such maps can be used by applications on the user device. In addition, previously created maps (e.g., canonical maps) can be stored in a remote data repository 574. Where suitable stored maps or persistent maps are available, these maps can be used instead of or in addition to the tracking map created locally on the device. In some embodiments, the tracking map can be positioned to the stored map, thereby establishing a correspondence between the tracking map (which can be oriented relative to the position of the wearable device when the user turns on the system) and the canonical map (which can be oriented relative to one or more persistent features). In some embodiments, a persistent map can be loaded on the user device to allow the user device to render virtual content without the delay associated with the scan location, thereby building a tracking map of the user's complete environment based on the sensor data acquired during the scan. In some embodiments, the user device can access a remote persistent map (e.g., stored in the cloud) without having to download the persistent map on the user device.

[0161] In some embodiments, spatial information can be transmitted from the wearable device to a remote service, such as a cloud service configured to locate the device to a stored map maintained on the cloud service. According to some embodiments, positioning processing can be performed in the cloud that matches the device location to an existing map (such as a canonical map) and returns a transformation that links virtual content to the wearable device location. In such an embodiment, the system can avoid transmitting the map from the remote resource to the wearable device. Other embodiments can be configured for both device-based positioning and cloud-based positioning, for example to enable functionality where a network connection is unavailable or the user chooses not to enable cloud-based positioning.

[0162] Alternatively or additionally, the tracking map may be merged with previously stored maps to extend or improve the quality of those maps. The process of determining whether a previously created suitable environment map is available and / or merging the tracking map with one or more stored environment maps may be performed in the local data processing module 570 or the remote processing module 572.

[0163] In some embodiments, the local data processing module 570 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the computational budget of the local data processing module 570 but enable a smaller device. In some embodiments, the world reconstruction component 516 can use less than the computational budget of a single Advanced RISC Machine (ARM) core to generate a physical world representation in real time on a non-predefined space, leaving the remaining computational budget of the single ARM core available for other purposes, such as extracting a mesh.

[0164] In some embodiments, remote data repository 574 may comprise a digital data storage facility accessible via the Internet or other network configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, allowing for fully autonomous use from a remote module. In some embodiments, all data is stored and all computations are performed in the remote data repository 574, allowing for the use of smaller devices. For example, a world reconstruction may be stored in whole or in part in this repository 574.

[0165] In embodiments where data is stored remotely and accessible over a network, the data can be shared by multiple users of the augmented reality system. For example, user devices can upload their tracking maps to augment the database of environmental maps. In some embodiments, the tracking map upload occurs at the end of the user session with the wearable device. In some embodiments, the tracking map upload can occur continuously, semi-continuously, intermittently, at a predefined time, after a predefined time period from the previous upload, or when triggered by an event. A tracking map uploaded by any user device can be used to expand or enhance a previously stored map, whether based on data from that user device or any other user device. Similarly, a persistent map downloaded to a user device can be based on data from that user device or any other user device. In this way, users can easily obtain high-quality environmental maps to improve their experience with the AR system.

[0166] In further embodiments, persistent map downloads may be limited and / or avoided based on positioning performed on a remote resource (e.g., in the cloud). In such a configuration, a wearable device or other XR device transmits feature information coupled with pose information (e.g., positioning information of the device when sensing a feature represented by the feature information) to a cloud service. One or more components of the cloud service may match the feature information with a corresponding stored map (e.g., a canonical map) and generate a transformation between the coordinate systems of the tracking map and the canonical map maintained by the XR device. Each XR device that positions the tracking map relative to the canonical map may accurately render virtual content at a location specified relative to the canonical map based on its own tracking.

[0167] In some embodiments, the local data processing module 570 is operably coupled to a battery 582. In some embodiments, the battery 582 is a removable power source, such as a battery that can be purchased over the counter. In other embodiments, the battery 582 is a lithium-ion battery. In some embodiments, the battery 582 includes an internal lithium-ion battery that can be charged by the user 560 during non-operating hours of the system 580, as well as a removable battery, so that the user 560 can operate the system 580 for extended periods of time without being tethered to a power source to charge the lithium-ion battery or having to shut down the system 580 to replace the battery.

[0168] Figure 5A A user 530 is shown wearing an AR display system that renders AR content as the user 530 moves in a physical world environment 532 (hereinafter referred to as "environment 532"). Information captured by the AR system along the user's movement path can be processed into one or more tracking maps. The user 530 positions the AR display system at a location 534, and the AR display system records contextual information of the connected world relative to the location 534 (e.g., digital representations of real objects in the physical world that can be stored and updated as the real objects in the physical world change). This information can be stored as a gesture in combination with images, features, directional audio input, or other desired data. The location 534 is integrated into a data input 536, for example as part of a tracking map, and is processed by at least a connected world module 538, which can, for example, be Figure 4 In some embodiments, the connected world module 538 may include a head pose component 514 and a world reconstruction component 516 so that the processed information can be combined with information about physical objects used in rendering virtual content to indicate the location of objects in the physical world.

[0169] As determined based on data input 536, connected world module 538 determines, at least in part, where and how AR content 540 can be placed in the physical world. AR content is "placed" in the physical world by presenting representations of the physical world and AR content through a user interface, with the AR content rendered as if interacting with objects in the physical world, and the objects in the physical world presented as if the AR content obscures the user's view of those objects when appropriate. In some embodiments, AR content can be placed by appropriately selecting a portion of a fixed element 542 (e.g., a table) from a reconstruction (e.g., reconstruction 518) to determine the shape and position of AR content 540. As an example, the fixed element can be a table, and virtual content can be positioned so that it appears to be located on the table. In some embodiments, AR content can be placed within a structure in field of view 544, which can be the current field of view or an estimated future field of view. In some embodiments, AR content can be persisted relative to a model 546 (mesh) of the physical world.

[0170] As shown, fixed element 542 acts as a pointer (e.g., a digital copy) to any fixed element within the physical world, which can be stored in connected world module 538 so that user 530 can perceive content on fixed element 542 without the system having to map to fixed element 542 each time user 530 views content. Thus, fixed element 542 can be a mesh model from a previous modeling session, or determined by an individual user, but still stored by connected world module 538 for future reference by multiple users. Thus, connected world module 538 can recognize environment 532 from a previously drawn environment and display AR content without requiring user 530's device to first draw all or part of environment 532, thereby saving computational processes and cycles and avoiding any delay in rendering AR content.

[0171] A mesh model 546 of the physical world can be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying AR content 540 can be stored by the connected world module 538 for future retrieval by the user 530 or other users without having to rebuild the model in whole or in part. In some embodiments, data input 536 is input such as geographic location, user identification, and current activity that indicates to the connected world module 538 which of one or more fixed elements 542 is available, which AR content 540 was last placed on the fixed element 542, and whether to display the same content (such AR content is "persistent" regardless of whether the user is viewing a particular connected world model).

[0172] Even in embodiments where objects are considered stationary (e.g., a kitchen table), the connected world module 538 may update those objects in the physical world model from time to time to account for changes that may occur in the physical world. Models of stationary objects may be updated at a much lower frequency. Other objects in the physical world may be moving or not considered stationary (e.g., a kitchen chair). To render a realistic AR scene, the AR system may update the positions of these non-stationary objects at a much higher frequency than that used to update stationary objects. To be able to accurately track all objects in the physical world, the AR system may obtain information from multiple sensors, including one or more image sensors.

[0173] Figure 5B 548 and accompanying components. In some embodiments, two eye-tracking cameras 550 directed toward the user's eyes 549 detect indicators of the user's eyes 549, such as eye shape, eyelid closure, pupil direction, and bright spots on the user's eyes 549.

[0174] In some embodiments, one of the sensors may be a depth sensor 551, such as a time-of-flight sensor, which transmits signals to the world and detects reflections of these signals from nearby objects to determine the distance to a given object. For example, a depth sensor can quickly determine whether an object has entered the user's field of view due to the movement of those objects or a change in the user's posture. However, information about the location of objects in the user's field of view can alternatively or additionally be collected using other sensors. For example, depth information can be obtained from a stereoscopic image sensor or a plenoptic sensor.

[0175] In some embodiments, world camera 552 records a view larger than the periphery to draw and / or otherwise create a model of environment 532 and detect input that can affect AR content. In some embodiments, world camera 552 and / or camera 553 can be grayscale and / or color image sensors that can output grayscale and / or color image frames at fixed time intervals. Camera 553 can also capture images of the physical world within the user's field of view at a specific time. The pixels of a frame-based image sensor can be repeatedly sampled even if their values ​​are not changed. Each of world camera 552, camera 553 and depth sensor 551 has a respective field of view 554, 555 and 556 to collect data from and record the physical world scene, such as Figure 5A The physical world environment 532 is depicted in FIG.

[0176] Inertial measurement unit 557 can determine the movement and orientation of viewing optics assembly 548. In some embodiments, each component is operably coupled to at least one other component. For example, depth sensor 551 is operably coupled to eye tracking camera 550 as a measure of the actual distance at which the user's eye 549 is looking.

[0177] It should be understood that the viewing optics assembly 548 may include Figure 5B , and may include components that are alternatives to or in addition to those shown. In some embodiments, for example, the viewing optics assembly 548 may include two world cameras 552 instead of four. Alternatively or additionally, the cameras 552 and 553 need not capture visible light images of their full fields of view. The viewing optics assembly 548 may include other types of components. In some embodiments, the viewing optics assembly 548 may include one or more dynamic vision sensors (DVS) whose pixels may asynchronously respond to relative changes in light intensity that exceed a threshold.

[0178] In some embodiments, the observation optical component 548 may not include a depth sensor 551 based on time-of-flight information. In some embodiments, for example, the observation optical component 548 may include one or more plenoptic cameras whose pixels can capture light intensity and the angle of incident light, from which depth information can be determined. For example, the plenoptic camera may include an image sensor covered with a transmissive diffraction mask (TDM). Alternatively or in addition, the plenoptic camera may include an image sensor that includes angle-sensitive pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such a sensor can be used as a source of depth information in addition to or in lieu of the depth sensor 551.

[0179] It should also be understood that Figure 5B The configuration of components in FIG. 5 is provided as an example. Viewing optics assembly 548 may include components having any suitable configuration that can be arranged to provide the user with the maximum field of view practical for a particular set of components. For example, if viewing optics assembly 548 includes a world camera 552, the world camera may be placed in a central region of the viewing optics assembly rather than on the side.

[0180] The information from these sensors in the observation optical assembly 548 can be coupled to one or more processors in the system. The processor can generate data that can be used to render virtual content that allows the user to perceive interaction with objects in the physical world. This rendering can be implemented in any suitable manner, including generating image data that depicts physical and virtual objects. In other embodiments, physical and virtual content can be shown in one scene by modulating the opacity of a display device through which the user views the physical world. The opacity can be controlled to create the appearance of virtual objects and also prevent the user from seeing objects in the physical world that are obscured by virtual objects. In some embodiments, the image data may only include virtual content that can be modified so that when viewed through the user interface, the virtual content is perceived by the user as realistically interacting with the physical world (e.g., clipping content to solve occlusion problems).

[0181] The position of content displayed on the viewing optics 548 to create the impression of an object at a particular location can depend on the physical characteristics of the viewing optics. Furthermore, the user's head posture relative to the physical world and the direction of the user's eye gaze can affect where content displayed at a particular location on the viewing optics appears in the physical world. The aforementioned sensors can collect this information and / or provide information from which this information can be calculated, such that a processor receiving sensor input can calculate where an object should be rendered on the viewing optics 548 to create the appearance desired by the user.

[0182] Regardless of how the content is presented to the user, a physical world model may be used so that properties of virtual objects that may be affected by physical objects can be correctly calculated, including the shape, position, motion, and visibility of the virtual objects. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 518.

[0183] The model can be created based on data collected from sensors on a user's wearable device. However, in some embodiments, the model can be created based on data collected by multiple users, which can be aggregated in a computing device remote from all users (and can be "in the cloud").

[0184] The model may be created at least in part by a world reconstruction system, e.g. Figure 3 The world reconstruction component 516, which is in Figure 6A. The world reconstruction component 516 may include a perception module 660 that may generate, update, and store a representation of a portion of the physical world. In some embodiments, the perception module 660 may represent a portion of the physical world within the sensor reconstruction range as a plurality of voxels. Each voxel may correspond to a 3D cube of a predetermined volume in the physical world and include surface information indicating whether a surface exists in the volume represented by the voxel. A voxel may be assigned a value indicating whether its corresponding volume has been determined to include a surface of a physical object, determined to be empty, or has not yet been measured with a sensor, in which case its value is unknown. It should be understood that there is no need to explicitly store values ​​indicating voxels that are determined to be empty or unknown, as the values ​​of these voxels may be stored in computer memory in any suitable manner, including not storing any information for voxels that are determined to be empty or unknown.

[0185] In addition to generating information for the persistent world representation, the perception module 660 can identify and output indications of changes in the area surrounding the AR system user. These indications of changes can trigger updates to volumetric data stored as part of the persistent world, or trigger other functions, such as triggering the generation of AR content to update the AR content component 304.

[0186] In some embodiments, the perception module 660 can identify changes based on a signed distance function (SDF) model. The perception module 660 can be configured to receive sensor data, such as a depth map 660a and a head pose 660b, and then fuse the sensor data into an SDF model 660c. The depth map 660a can directly provide SDF information, and the image can be processed to obtain the SDF information. The SDF information represents the distance from the sensors used to capture the information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and therefore from the perspective of the user. The head pose 660b can relate the SDF information to voxels in the physical world.

[0187] In some embodiments, the perception module 660 can generate, update, and store a representation of a portion of the physical world within a perception range. The perception range can be determined at least in part based on a reconstruction range of the sensor, which can be determined at least in part based on the limits of the sensor's observation range. As a specific example, an active depth sensor operating using active IR pulses can reliably operate within a range of distances, thereby creating a sensor's observation range, which can range from a few centimeters or tens of centimeters to several meters.

[0188] The world reconstruction component 516 may include additional modules that can interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data acquired by the perception module 660. The persistent world module 662 may also include various representation formats of the physical world. For example, volume metadata 662b such as voxels, as well as meshes 662c and planes 662d may be stored. In some embodiments, other information such as depth maps may also be stored.

[0189] In some embodiments, such as Figure 6A The depicted representation of the physical world may provide relatively dense information about the physical world compared to sparse maps (eg, the feature point-based tracking maps described above).

[0190] In some embodiments, the perception module 660 may include a module that generates a representation of the physical world in various formats (e.g., including grids 660d, planes, and semantics 660e). The representation of the physical world can be stored on local and remote storage media. Depending on, for example, the location of the storage media, the representation of the physical world can be described in different coordinate systems. For example, a representation of the physical world stored in a device can be described in a coordinate system local to the device. The representation of the physical world can have a counterpart stored in the cloud. The counterpart in the cloud can be described in a coordinate system shared by all devices in the XR system.

[0191] In some embodiments, these modules can generate representations based on data within the perception range of one or more sensors at the time the representation is generated, as well as data captured at a previous time and information in the persistent world module 662. In some embodiments, these components can operate on depth information captured using a depth sensor. However, the AR system can include a visual sensor and can generate such representations by analyzing monocular or binocular visual information.

[0192] In some embodiments, these modules can operate on regions of the physical world. When the perception module 660 detects a change in the physical world in a sub-region, these modules can be triggered to update the sub-region of the physical world. For example, such a change can be detected by detecting a new surface in the SDF model 660c or other criteria (such as a change in the value of a sufficient number of voxels representing the sub-region).

[0193] World reconstruction component 516 may include component 664, which may receive a representation of the physical world from perception module 660. Information about the physical world may be pulled by these components based on usage requests, such as from applications. In some embodiments, information may be pushed to consuming components, such as via indications of changes in pre-identified areas or changes in the representation of the physical world within the perception range. Component 664 may include, for example, game programs and other components that handle visual occlusion, physics-based interactions, and environmental reasoning.

[0194] In response to a query from component 664, perception module 660 can send a representation of the physical world in one or more formats. For example, when component 664 indicates that the usage is for visual occlusion or physics-based interaction, perception module 660 can send a surface representation. When component 664 indicates that the usage is for environmental reasoning, perception module 660 can send meshes, planes, and semantics of the physical world.

[0195] In some embodiments, perception module 660 may include a component that formats the information to provide component 664. An example of such a component may be ray casting component 660f. Using a component (e.g., component 664), for example, information about the physical world from a particular viewpoint may be queried. Ray casting component 660f may select from one or more representations of the physical world data within the field of view from that viewpoint.

[0196] It should be understood from the foregoing that the perception module 660 or another component of the AR system can process data to create a 3D representation of a portion of the physical world. The data to be processed can be simplified by culling a portion of the 3D reconstruction volume based at least in part on the camera frustum and / or the depth image, extracting and persisting planar data, capturing, persisting, and updating 3D reconstruction data in blocks that allow local updates while maintaining neighborhood consistency, providing occlusion data to applications generating such scenes (where the occlusion data is derived from a combination of one or more depth data sources), and / or performing multi-stage mesh simplification. The reconstruction can contain data of varying complexity, including, for example, raw data such as real-time depth data, fused volumetric data such as voxels, and computed data such as meshes.

[0197] In some embodiments, the components of the connected world model can be distributed, with some portions executing locally on the XR device and some portions executing remotely, such as on a network-connected server or otherwise in the cloud. The distribution of information processing and storage between the local XR device and the cloud can impact the functionality and user experience of the XR system. For example, reducing processing on the local device by offloading it to the cloud can extend battery life and reduce heat generated on the local device. However, offloading excessive processing to the cloud can introduce undesirable latency, resulting in an unacceptable user experience.

[0198] Figure 6B 6 shows a distributed component architecture 600 configured for spatial computing according to some embodiments. The distributed component architecture 600 may include a connected world component 602 (e.g., Figure 5A 6), Lumin OS 604, API 606, SDK 608, and applications 610. Lumin OS 604 may include a Linux-based kernel with custom drivers compatible with XR devices. API 606 may include an application programming interface that allows XR applications (e.g., application 610) to access the spatial computing features of the XR device. SDK 608 may include a software development kit that allows the creation of XR applications.

[0199] One or more components in architecture 600 may create and maintain a model of the connected world. In this example, sensor data is collected on a local device. Processing of this sensor data may be performed partially locally on the XR device and partially in the cloud. PW 538 may include a map of the environment created at least in part based on data captured by AR devices worn by multiple users. During a session of an AR experience, a single AR device (such as the one described above in conjunction with Figure 4 The wearable device described) can create tracking maps, which is a type of map.

[0200] In some embodiments, the device may include components for building sparse maps and dense maps. A tracking map may serve as a sparse map and may include the head pose of the AR device scanning the environment and information about objects detected in the environment at each head pose. These head poses may be maintained locally for each device. For example, the head pose on each device is relative to the initial head pose when the device is turned on for a session. Thus, each tracking map is local to the device on which it was created. A dense map may include surface information that may be represented by a mesh or depth information. Alternatively or in addition, a dense map may include higher-level information derived from the surface or depth information, such as the location and / or properties of planes and / or other objects.

[0201] In some embodiments, the creation of a dense map can be independent of the creation of a sparse map. For example, the creation of a dense map and a sparse map can be performed in separate processing pipelines within the AR system. For example, the separate processing can enable the generation or processing of different types of maps to be performed at different rates. For example, a sparse map may be refreshed at a faster rate than a dense map. However, in some embodiments, the processing of dense maps and sparse maps may be related, even if performed in different pipelines. For example, a change in the physical world displayed in a sparse map may trigger an update to a dense map, and vice versa. Furthermore, even if maps are created independently, the maps can still be used together. For example, a coordinate system derived from a sparse map can be used to define the position and / or orientation of objects in a dense map.

[0202] Sparse maps and / or dense maps can be persisted for reuse by the same device and / or sharing with other devices. Such persistence can be achieved by storing the information in the cloud. The AR device can send the tracking map to the cloud to, for example, merge with an environment map selected from a persistent map previously stored in the cloud. In some embodiments, the selected persistent map can be sent from the cloud to the AR device for merging. In some embodiments, the persistent map can be oriented relative to one or more persistent coordinate systems. Such maps can be used as canonical maps because they can be used by any of multiple devices. In some embodiments, the model of the connected world can include or be created based on one or more canonical maps. Even if devices perform some operations based on their local coordinate system, they can still use the canonical map by determining the transformation between their device local coordinate system and the canonical map.

[0203] Normalize maps to Tracking Maps(TM) (e.g. Figure 31A The tracking map may be promoted to a canonical map, taking the TM 1102 in the canonical map as a starting point. The canonical map may be persisted so that a device accessing the canonical map, once it has determined the transformation between its local coordinate system and the coordinate system of the canonical map, can use the information in the canonical map to determine the position of objects represented in the canonical map in the physical world around the device. In some embodiments, the TM may be a sparse map of head pose created by the XR device. In some embodiments, the canonical map may be created when the XR device sends one or more TMs to a cloud server to be merged with additional TMs captured by the XR device at a different time or by other XR devices.

[0204] A canonical map or other map may provide information about the portion of the physical world represented by the data processed to create the corresponding map. Figure 7An exemplary tracking map 700 is shown in accordance with some embodiments. The tracking map 700 may provide a plan view 706 of physical objects in the corresponding physical world represented by points 702. In some embodiments, the map points 702 may represent features of a physical object which may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. These features may be derived from processed images, such as may be obtained using sensors of a wearable device in an augmented reality system. For example, features may be derived by processing image frames output by a sensor to identify features based on large gradients in the image or other suitable criteria. Further processing may limit the number of features in each frame. For example, the processing may select features that may represent persistent objects. One or more heuristics may be used to make this selection.

[0205] The tracking map 700 may include data about points 702 collected by the device. For each image frame having data points contained in the tracking map, a pose may be stored. The pose may represent the orientation at which the image frame was captured, such that feature points within each image frame may be spatially correlated. The pose may be determined by positioning information, such as may be derived from sensors on the wearable device (such as an IMU sensor). Alternatively or additionally, the pose may be determined by matching the image frame with other image frames that depict overlapping portions of the physical world. By finding such positional correlations, which may be achieved by matching subsets of feature points in two frames, a relative pose between the two frames may be calculated. Relative poses are sufficient for the tracking map because the map may be relative to a coordinate system local to the device that is established based on an initial pose of the device when construction of the tracking map begins.

[0206] Not all feature points and image frames collected by the device may be retained as part of the tracking map, as much of the information collected using the sensors may be redundant. Instead, only certain frames may be added to the map. These frames may be selected based on one or more criteria, such as the degree of overlap with image frames already in the map, the number of new features they contain, or a quality indicator of the features in the frame. Image frames that are not added to the tracking map may be discarded or may be used to modify the location of features. As a further alternative, all or most of the image frames represented as a feature set may be retained, but a subset of these frames may be designated as key frames for further processing.

[0207] The keyframes may be processed to generate a keyrig 704. The keyframes may be processed to generate a set of three-dimensional feature points and saved as a keyrig 704. Such processing may require, for example, comparing image frames simultaneously derived from two cameras to stereoscopically determine the 3D positions of the feature points. Metadata may be associated with these keyframes and / or keyrigs, such as pose.

[0208] The environment map can have any of a variety of formats, depending, for example, on where the environment map is stored, including, for example, local memory of the AR device and remote memory. For example, the resolution of the map in the remote memory may be higher than the map in the local memory on a wearable device with limited memory. To send the higher resolution map from the remote memory to the local memory, the map can be downsampled or otherwise converted to an appropriate format, such as by reducing the number of poses per unit area of ​​the physical world stored in the map and / or the number of feature points stored for each pose. In some embodiments, a slice or portion of the high-resolution map from the remote memory can be sent to the local memory without the slice or portion being downsampled.

[0209] A database of environment maps can be updated when a new tracking map is created. To determine which of a potentially large number of environment maps in the database to update, the update can include effectively selecting one or more environment maps stored in the database that are relevant to the new tracking map. The selected one or more environment maps can be ranked by relevance, and one or more highest-ranked maps can be selected for processing to merge the highly ranked selected environment maps with the new tracking map to create one or more updated environment maps. When the new tracking map represents a portion of the physical world for which no pre-existing environment map is to be updated, the tracking map can be stored in the database as a new environment map.

[0210] Watch standalone display

[0211] This article describes methods and apparatus for providing virtual content using an XR system that is independent of the position of the eyes viewing the virtual content. Traditionally, virtual content is re-rendered whenever the display system performs any action. For example, if a user wearing a display system views a virtual representation of a three-dimensional (3D) object on a display and walks around the area where the 3D object appears, the 3D object should be re-rendered for each viewpoint so that the user feels like they are walking around an object that occupies real space. However, re-rendering consumes a lot of the system's computing resources and causes artifacts due to latency.

[0212] The inventors have recognized and appreciated that head pose (e.g., the position and orientation of a user wearing an XR system) can be used to render virtual content independently of eye rotation within the user's head. In some embodiments, a dynamic map of a scene can be generated based on multiple coordinate systems in real space across one or more sessions, such that virtual content interacting with the dynamic map can be robustly rendered independent of eye rotation within the user's head and / or independent of sensor deformation caused by, for example, heat generated during high-speed, computationally intensive operations. In some embodiments, configuring multiple coordinate systems can enable a first XR device worn by a first user and a second XR device worn by a second user to identify a common location in the scene. In some embodiments, configuring multiple coordinate systems can enable users wearing XR devices to view virtual content in the same location of the scene.

[0213] In some embodiments, a tracking map may be constructed in a world coordinate system having a world origin. The world origin may be the first pose when the XR device is powered on. The world origin may be aligned with gravity so that developers of XR applications can obtain gravity alignment without additional work. Different tracking maps may be constructed in different world coordinate systems because tracking maps may be captured by the same XR device in different sessions and / or by different XR devices worn by different users. In some embodiments, a session of an XR device may extend from when the device is powered on to when the device is powered off. In some embodiments, the XR device may have a head coordinate system including a head origin. The head origin may be the pose of the XR device when the image is captured. The difference between the head pose of the world coordinate system and the head pose of the head coordinate system may be used to estimate the tracking route.

[0214] In some embodiments, an XR device may have a camera coordinate system including a camera origin. The camera origin may be the current pose of one or more sensors of the XR device. The inventors have recognized and appreciated that the configuration of the camera coordinate system enables robust display of virtual content independent of eye rotation within the user's head. This configuration also enables robust display of virtual content independent of sensor deformation caused by, for example, heat generated during operation.

[0215] In some embodiments, the XR device may have a head unit with a head-mounted frame (which the user can secure to their head) and may include two waveguides, one in front of each eye of the user. The waveguides may be transparent so that background light from real-world objects can be transmitted through the waveguides and the user can see the real-world objects. Each waveguide may transmit projected light from a projector to a corresponding eye of the user. The projected light may form an image on the retina of the eye. The retina of the eye thus receives both the background light and the projected light. The user may simultaneously see the real-world objects and one or more virtual objects created by the projected light. In some embodiments, the XR device may have sensors that detect real-world objects around the user. For example, these sensors may be cameras that capture images that can be processed to identify the location of the real-world objects.

[0216] In some embodiments, the XR system can assign a coordinate system to virtual content, as opposed to attaching the virtual content to the world coordinate system. Such a configuration enables the description of virtual content without regard to where the virtual content is rendered to the user, but the virtual content can be attached to a more persistent coordinate system location, such as with respect to, for example, Figure 14-20C The XR device can detect changes in the environment map and determine the movement of the head-mounted unit worn by the user relative to the real-world object when the position of the object changes.

[0217] Figure 8 1 shows a user in a physical environment experiencing virtual content rendered by an XR system 10. The XR system may include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is in a physical environment having real objects in the form of a table 16.

[0218] In the example shown, the first XR device 12.1 includes a head unit 22, a waist pack 24, and a cable connection 26. The first user 14.1 secures the head unit 22 to their head and secures the waist pack 24 to their waist, away from the head unit 22. The cable connection 26 connects the head unit 22 to the waist pack 24. The head unit 22 includes technology for displaying one or more virtual objects to the first user 14.1 while allowing the first user 14.1 to see real objects such as a table 16. The waist pack 24 primarily includes the processing and communication functions of the first XR device 12.1. In some embodiments, the processing and communication functions can reside entirely or partially in the head unit 22, so that the waist pack 24 can be removed or can be located in another device such as a backpack.

[0219] In the example shown, a belt pack 24 is connected to the network 18 via a wireless connection. A server 20 is connected to the network 18 and stores data representing local content. The belt pack 24 downloads the data representing the local content from the server 20 via the network 18. The belt pack 24 provides the data to the head unit 22 via a cable connection 26. The head unit 22 may include a display having a light source, such as a laser light source or a light emitting diode (LED), and a waveguide to guide the light.

[0220] In some embodiments, first user 14.1 may attach head unit 22 to their head and waist pack 24 to their waist. Waist pack 24 may download image data representing virtual content from server 20 via network 18. First user 14.1 may view table 16 through the display of head unit 22. A projector forming part of head unit 22 may receive the image data from waist pack 24 and generate light based on the image data. This light may travel through one or more waveguides forming part of the display of head unit 22. The light may then exit the waveguides and propagate onto the retina of first user 14.1's eye. The projector generates light in a pattern that is replicated on the retina of first user 14.1's eye. The light that strikes the retina of first user 14.1's eye may have a selected depth of field, causing first user 14.1 to perceive an image at a preselected depth behind the waveguide. Furthermore, first user 14.1's two eyes may receive slightly different images, causing first user 14.1's brain to perceive one or more three-dimensional images at a selected distance from head unit 22. In the example shown, the first user 14.1 perceives virtual content 28 above the desktop 16. The scale of the virtual content 28 and its position and distance from the first user 14.1 are determined by the data representing the virtual content 28 and the various coordinate systems used to display the virtual content 28 to the first user 14.1.

[0221] In the example shown, the virtual content 28 is invisible from the perspective of the drawing and is visible to the first user 14.1 using the first XR device 12.1. The virtual content 28 may initially exist as a data structure within the visual data and algorithms in the belt pack 24. The data structures may then manifest themselves as light when the projector of the head unit 22 generates light based on the data structures. It should be understood that although the virtual content 28 does not exist in the three-dimensional space in front of the first user 14.1, the virtual content 28 is still visible in the three-dimensional space. Figure 1 24 to illustrate what is perceived by a wearer of head unit 22. Visualization of computer data in three-dimensional space may be used in this description to illustrate how data structures that facilitate rendering are perceived by one or more users as being related to one another within data structures in waist pack 24.

[0222] Figure 9Components of a first XR device 12.1 according to some embodiments are shown. The first XR device 12.1 may include a head unit 22, and various components that form part of the visual data and algorithms, including, for example, a rendering engine 30, various coordinate systems 32, various origin and destination coordinate systems 34, and various origin-to-destination coordinate system transformers 36. The various coordinate systems may be based on intrinsic factors of the XR device, or may be determined by reference to other information (such as a persistent pose or a persistent coordinate system), as described herein.

[0223] Head unit 22 may include a head-mounted frame 40 , a display system 42 , a real object detection camera 44 , a motion tracking camera 46 , and an inertial measurement unit 48 .

[0224] The head-mounted frame 40 may have a Figure 8 The display system 42 , the real object detection camera 44 , the motion tracking camera 46 , and the inertial measurement unit 48 may be mounted to the head mounted frame 40 and thus move with the head mounted frame 40 .

[0225] Coordinate system 32 may include a local data system 52 , a world coordinate system 54 , a head coordinate system 56 , and a camera coordinate system 58 .

[0226] The local data system 52 may include a data channel 62, a local coordinate system determination routine 64, and local coordinate system storage instructions 66. The data channel 62 may be an internal software routine, a hardware component such as an external cable or radio frequency receiver, or a hybrid component such as an open port. The data channel 62 may be configured to receive image data 68 representing virtual content.

[0227] A local coordinate system determination routine 64 can be connected to the data channel 62. The local coordinate system determination routine 64 can be configured to determine a local coordinate system 70. In some embodiments, the local coordinate system determination routine can determine the local coordinate system based on a real-world object or real-world location. In some embodiments, the local coordinate system can be based on a top edge relative to a bottom edge of the browser window, a character's head or feet, a node on an outer surface of a prism or a bounding box surrounding virtual content, or any other suitable location for placing a coordinate system that defines the orientation of virtual content and the location of virtual content (e.g., a node such as a placement node or a PCF node).

[0228] The local coordinate system storage instructions 66 may be connected to the local coordinate system determination routine 64. Those skilled in the art will appreciate that software modules and routines are "connected" to each other through subroutines, calls, and the like. The local coordinate system storage instructions 66 may store the local coordinate system 70 as the local coordinate system 72 in the origin and destination coordinate system 34. In some embodiments, the origin and destination coordinate system 34 may be one or more coordinate systems that may be manipulated or transformed so that virtual content persists between sessions. In some embodiments, a session may be the period of time between startup and shutdown of an XR device. Two sessions may be two startup and shutdown cycles of a single XR device, or may be one startup and shutdown of two different XR devices.

[0229] In some embodiments, the origin and destination coordinate systems 34 may be coordinate systems that involve one or more transformations required to enable the first user's XR device and the second user's XR device to recognize a common location. In some embodiments, the destination coordinate system may be the output of a series of calculations and transformations applied to a target coordinate system so that the first user and the second user view virtual content at the same location.

[0230] The rendering engine 30 may be connected to the data channel 62. The rendering engine 30 may receive image data 68 from the data channel 62 so that the rendering engine 30 may render virtual content based at least in part on the image data 68.

[0231] The display system 42 may be connected to the rendering engine 30. The display system 42 may include components that convert the image data 68 into visible light. The visible light may be formed into two patterns, one for each eye. The visible light may enter Figure 8 The eye of the first user 14 . 1 can be detected on the retina of the eye of the first user 14 . 1 .

[0232] The real object detection camera 44 may include one or more cameras that can capture images from different sides of the head-mounted frame 40. The motion tracking camera 46 may include one or more cameras that capture image frames from the sides of the head-mounted frame 40. Instead of two sets of one or more cameras representing one or more real object detection cameras 44 and one or more motion tracking cameras 46, one set of one or more cameras may be used. In some embodiments, cameras 44 and 46 can capture images. As described above, these cameras can collect data used to build a tracking map.

[0233] Inertial measurement unit 48 may include multiple devices for detecting movement of head unit 22. Inertial measurement unit 48 may include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of inertial measurement unit 48, in combination, track the movement of head unit 22 in at least three orthogonal directions and about at least three orthogonal axes.

[0234] In the example shown, the world coordinate system 54 includes a world surface determination routine 78, a world coordinate system determination routine 80, and world coordinate system storage instructions 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 receives images and / or keyframes based on images captured by the real object detection camera 44 and processes the images to identify surfaces in the images. A depth sensor (not shown) can determine the distance to the surface. Thus, the surface is represented by three dimensions of data, including its size, shape, and distance from the real object detection camera.

[0235] In some embodiments, the world coordinate system 84 can be based on the origin at the time the head gesture session is initialized. In some embodiments, the world coordinate system can be at the location where the device was booted up, or at a new location if the head gesture was lost during the boot session. In some embodiments, the world coordinate system can be the origin at the start of the head gesture session.

[0236] In the example shown, a world coordinate system determination routine 80 is coupled to the world surface determination routine 78 and determines a world coordinate system 84 based on the position of the surface determined by the world surface determination routine 78. World coordinate system storage instructions 82 are coupled to the world coordinate system determination routine 80 to receive the world coordinate system 84 from the world coordinate system determination routine 80. The world coordinate system storage instructions 82 store the world coordinate system 84 as a world coordinate system 86 within the origin and destination coordinate system 34.

[0237] The head coordinate system 56 may include a head coordinate system determination routine 90 and head coordinate system storage instructions 92. The head coordinate system determination routine 90 may be connected to the motion tracking camera 46 and the inertial measurement unit 48. The head coordinate system determination routine 90 may use data from the motion tracking camera 46 and the inertial measurement unit 48 to calculate a head coordinate system 94. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The motion tracking camera 46 may continuously capture images that are used by the head coordinate system determination routine 90 to refine the head coordinate system 94. When Figure 8 As the first user 14 . 1 in FIG. 2 moves their head, the head unit 22 also moves. The motion tracking camera 46 and the inertial measurement unit 48 may continuously provide data to the head coordinate system determination routine 90 so that the head coordinate system determination routine 90 may update the head coordinate system 94 .

[0238] The head coordinate system storage instructions 92 may be connected to the head coordinate system determination routine 90 to receive a head coordinate system 94 from the head coordinate system determination routine 90. The head coordinate system storage instructions 92 may store the head coordinate system 94 in the origin and destination coordinate system 34 as a head coordinate system 96. When the head coordinate system determination routine 90 recalculates the head coordinate system 94, the head coordinate system storage instructions 92 may repeatedly store the updated head coordinate system 94 as the head coordinate system 96. In some embodiments, the head coordinate system may be the position of the wearable XR device 12.1 relative to the local coordinate system 72.

[0239] The camera coordinate system 58 may include camera intrinsics 98. The camera intrinsics 98 may include the dimensions of the head unit 22 as a feature of its design and manufacture. The camera intrinsics 98 may be used to calculate a camera coordinate system 100 that is stored within the origin and destination coordinate system 34.

[0240] In some embodiments, the camera coordinate system 100 may include Figure 8 1. As the left eye moves from left to right or from top to bottom, the pupil position of the left eye is located within the camera coordinate system 100. Additionally, the pupil position of the right eye is located within the camera coordinate system 100 for the right eye. In some embodiments, the camera coordinate system 100 may include the position of the camera relative to a local coordinate system when the image was captured.

[0241] The origin-to-destination coordinate system transformer 36 may include a local-to-world coordinate transformer 104, a world-to-head coordinate transformer 106, and a head-to-camera coordinate transformer 108. The local-to-world coordinate transformer 104 may receive the local coordinate system 72 and transform the local coordinate system 72 to the world coordinate system 86. The transformation of the local coordinate system 72 to the world coordinate system 86 may be represented as a local coordinate system 110 transformed within the world coordinate system 86 to the world coordinate system.

[0242] A world-to-head coordinate transformer 106 may transform from the world coordinate system 86 to the head coordinate system 96. The world-to-head coordinate transformer 106 may transform a local coordinate system 110 transformed to the world coordinate system to the head coordinate system 96. This transformation may be represented as a local coordinate system 112 transformed within the head coordinate system 96 to the head coordinate system.

[0243] The head-to-camera coordinate transformer 108 may transform from the head coordinate system 96 to the camera coordinate system 100. The head-to-camera coordinate transformer 108 may transform the local coordinate system 112 transformed to the head coordinate system to a local coordinate system 114 transformed to the camera coordinate system within the camera coordinate system 100. The local coordinate system 114 transformed to the camera coordinate system may be input to the rendering engine 30. The rendering engine 30 may render the image data 68 representing the local content 28 based on the local coordinate system 114 transformed to the camera coordinate system.

[0244] Figure 10 is a spatial representation of the various origin and destination coordinate systems 34. The figure shows a local coordinate system 72, a world coordinate system 86, a head coordinate system 96, and a camera coordinate system 100. In some embodiments, when virtual content is placed in the real world so that the user can view the virtual content, the local coordinate system associated with the XR content 28 can have a position and rotation relative to the local and / or world coordinate systems and / or PCF (e.g., a node and an orientation can be provided). Each camera can have its own camera coordinate system 100, which contains all the pupil positions of one eye. Reference numerals 104A and 106A respectively represent the camera coordinate system represented by Figure 9 The transformations are performed by the local to world coordinate transformer 104, the world to head coordinate transformer 106 and the head to camera coordinate transformer 108 in FIG.

[0245] Figure 11 A camera rendering protocol for transforming from a head coordinate system to a camera coordinate system is shown in accordance with some embodiments. In the example shown, the pupil of a single eye moves from position A to B. Virtual objects intended to appear stationary will be projected onto the depth plane at one of two positions, A or B, depending on the position of the pupil (assuming the camera is configured to use a pupil-based coordinate system). Therefore, using a pupil coordinate system transformed to a head coordinate system will result in jitter in stationary virtual objects as the eye moves from position A to position B. This situation is referred to as view-dependent display or projection.

[0246] like Figure 12 As shown, the camera coordinate system (e.g., CR) is positioned and contains all pupil positions, and regardless of pupil positions A and B, the object projection is now consistent. The head coordinate system is transformed to the CR coordinate system, which is called view-independent display or projection. Image reprojection can be applied to virtual content to cope with changes in eye position, however, jitter is minimized while rendering remains in the same location.

[0247] Figure 13 The display system 42 is shown in more detail. The display system 42 comprises a stereo analyser 144 which is connected to the rendering engine 30 and forms part of the visual data and algorithms.

[0248] Display system 42 also includes left projector 166A and right projector 166B and left waveguide 170A and right waveguide 170B. Left projector 166A and right projector 166B are connected to a power supply. Each projector 166A and 166B has a corresponding input so that image data is provided to the corresponding projector 166A or 166B. When powered, the corresponding projector 166A or 166B generates a two-dimensional pattern of light and emits light therefrom. Left waveguide 170A and right waveguide 170B are positioned to receive light from left projector 166A and right projector 166B, respectively. Left waveguide 170A and right waveguide 170B are transparent waveguides.

[0249] In use, the user mounts the head-mounted frame 40 on their head. Components of the head-mounted frame 40 may include, for example, a strap (not shown) that wraps around the back of the user's head. The left and right waveguides 170A, 170B are then positioned in front of the user's left and right eyes 220A, 220B.

[0250] The rendering engine 30 inputs the image data it receives into the stereo analyzer 144. The image data is Figure 8 3D image data of local content 28 in the video is projected onto multiple virtual planes. A stereo analyzer 144 analyzes the image data to determine left and right image datasets based on the image data projected onto each depth plane. The left and right image datasets represent 2D images that are projected onto three dimensions to give the user a sense of depth.

[0251] The stereo analyzer 144 inputs the left and right image data sets into the left projector 166A and the right projector 166B. The left projector 166A and the right projector 166B then create left and right light patterns. The components of the display system 42 are shown in plan view, but it should be understood that when shown in elevation, the left and right patterns are two-dimensional patterns. Each light pattern includes a plurality of pixels. For illustrative purposes, light rays 224A and 226A from two pixels are shown as leaving the left projector 166A and entering the left waveguide 170A. Light rays 224A and 226A reflect from the sides of the left waveguide 170A. Light rays 224A and 226A are shown propagating within the left waveguide 170A by internal reflection from left to right, although it should be understood that light rays 224A and 226A can also propagate into the paper in one direction using a refraction and reflection system.

[0252] Light rays 224A and 226A exit the left optical waveguide 170A through pupil 228A and then enter the left eye 220A through pupil 230A of the left eye 220A. Light rays 224A and 226A then fall on the retina 232A of the left eye 220A. In this manner, the left light pattern falls on the retina 232A of the left eye 220A. The user perceives the pixels formed on the retina 232A as pixels 234A and 236A, which the user perceives as being located at a certain distance on the side of the left waveguide 170A opposite the left eye 220A. The perception of depth is created by manipulating the focal length of light.

[0253] In a similar manner, stereo analyzer 144 inputs the right image data set into right projector 166B. Right projector 166B emits a right light pattern, represented by pixels in the form of light rays 224B and 226B. Light rays 224B and 226B reflect within right waveguide 170B and exit through pupil 228B. Light rays 224B and 226B then enter through pupil 230B of right eye 220B and fall on retina 232B of right eye 220B. The pixels of light rays 224B and 226B are perceived as pixels 134B and 236B behind right waveguide 170B.

[0254] The patterns created on retinas 232A and 232B are perceived as left and right images, respectively. The left and right images are slightly different from each other due to the function of stereo analyzer 144. The left and right images are perceived as a three-dimensional rendering in the user's mind.

[0255] As described above, the left waveguide 170A and the right waveguide 170B are transparent. Light from a real-life object, such as the table 16 on one side of the left waveguide 170A and the right waveguide 170B opposite to the eyes 220A and 220B, can be projected through the left waveguide 170A and the right waveguide 170B and fall on the retinas 232A and 232B.

[0256] Persistent Coordinate Framework (PCF)

[0257] This document describes methods and apparatus for providing spatial persistence across user instances within a shared space. Without spatial persistence, virtual content placed in the physical world by a user during one session may not exist or may be misplaced in the user's view in a different session. Without spatial persistence, virtual content placed in the physical world by one user may not exist or be misplaced in the view of a second user, even if the second user intends to share the experience of the same physical space with the first user.

[0258] The inventors have recognized and appreciated that spatial persistence can be provided through a persistent coordinate system (PCF). The PCF can be defined based on one or more points representing features (e.g., corners, edges) recognized in the physical world. Features can be selected so that they appear to be the same from one user instance of the XR system to another.

[0259] Additionally, tracking drift, which causes the calculated tracking path (e.g., camera trajectory) to deviate from the actual tracking path, can cause the position of virtual content to appear in inappropriate locations when rendered relative to a local map based solely on a tracking map. As the XR device collects more information about the scene over time, the tracking map of the space can be refined to correct for drift. However, if virtual content is placed on real objects before the map is refined and saved relative to the device's world coordinate system derived from the tracking map, the virtual content may appear displaced as if the real objects had moved during map refinement. PCFs can be updated based on map refinement because PCFs are defined based on features and are updated as features move during map refinement.

[0260] The PCF may include six degrees of freedom including translation and rotation relative to a map coordinate system. The PCF may be stored in a local storage medium and / or a remote storage medium. The translation and rotation of the PCF may be calculated relative to the map coordinate system, depending on, for example, the storage location. For example, a PCF used locally by a device may have translation and rotation relative to the device's world coordinate system. A PCF in the cloud may have translation and rotation relative to the canonical coordinate system of a canonical map.

[0261] PCF can provide a sparse representation of the physical world, providing less than all available information about the physical world so that it can be efficiently processed and transmitted. Techniques for processing persistent spatial information can include creating a dynamic map based on one or more coordinate systems in real space across one or more sessions, generating a persistent coordinate system (PCF) on top of the sparse map, which can be exposed to XR applications via an application programming interface (API), for example.

[0262] Figure 14 11 is a block diagram illustrating the creation of a persistent coordinate system (PCF) and the attachment of XR content to the PCF in accordance with some embodiments. Each block may represent digital information stored in computer memory. In the case of application 1180, the data may represent computer-executable instructions. For example, in the case of virtual content 1170, the digital information may define a virtual object, such as that specified by application 1180. In the case of other blocks, the digital information may represent some aspect of the physical world.

[0263] In the illustrated embodiment, one or more PCFs are created based on images captured using sensors on the wearable device. Figure 14 In an embodiment, the sensors are visual image cameras. These cameras can be the same cameras used to form the tracking map. Thus, Figure 14 Some of the processing suggested can be performed as part of updating the tracking map. However, Figure 14 It is shown that in addition to the tracking map, information providing persistence is also generated.

[0264] To derive a 3D PCF, two images 1110 from two cameras mounted on the wearable device are processed together in a configuration that allows stereo image analysis. Figure 14 Image 1 and Image 2 are shown, each originating from one of the cameras. For simplicity, a single image from each camera is shown. However, each camera can output a stream of image frames, and Figure 14 The illustrated processing may be performed for multiple image frames in a stream.

[0265] Therefore, Image 1 and Image 2 are each a frame in a sequence of image frames. Figure 14 The illustrated process may be repeated for successive image frames in the sequence until an image frame containing feature points that provide a suitable image for forming persistent spatial information is processed. Alternatively or additionally, Figure 14 The process can be repeated as the user moves so that the user is no longer close enough to a previously identified PCF to reliably use that PCF to determine position relative to the physical world. For example, the XR system maintains a current PCF for the user. When that distance exceeds a threshold, the system can switch to a new current PCF that is closer to the user, which can be based on Figure 14 The process is generated using image frames acquired at the user's current location.

[0266] Even when generating a single PCF, the stream of image frames can be processed to identify image frames depicting content in the physical world that is likely stable and can be easily recognized by devices near the area of ​​the physical world depicted in the image frames. Figure 14 In the embodiment shown, the process begins by identifying features in the image 1120. For example, features can be identified by finding gradient locations or other characteristics above a threshold in the image, which can correspond to corners of an object, for example. In the embodiment shown, the features are points, but other identifiable features, such as edges, can be used instead or in addition.

[0267] In the illustrated embodiment, a fixed number N of features 1120 are selected for further processing. These feature points can be selected based on one or more criteria, such as gradient magnitude or proximity to other feature points. Alternatively or additionally, feature points can be selected heuristically, such as based on properties that suggest the persistence of the feature point. For example, a heuristic can be defined based on the properties of a feature point that may correspond to a corner of a window or door or a large piece of furniture. This heuristic takes into account the feature point itself and its surrounding environment. As a specific example, the number of feature points per image can be between 100 and 500, or between 150 and 250, such as 200.

[0268] Regardless of the number of feature points selected, descriptors 1130 can be calculated for the feature points. In this example, a descriptor is calculated for each selected feature point, but descriptors can be calculated for multiple groups of feature points or subsets of feature points or all features within an image. The descriptors characterize the feature points so that feature points representing the same object in the physical world are assigned similar descriptors. The descriptors can facilitate alignment of two frames, such as may occur when one map is positioned relative to another. Initial alignment of the two frames can be performed by identifying feature points with similar descriptors, rather than searching for a relative frame orientation that minimizes the distance between feature points of the two images. Alignment of image frames can be based on the alignment of points with similar descriptors, which requires less processing than calculating the alignment of all feature points in the image.

[0269] Descriptors can be calculated as a mapping of feature points to descriptors, or in some embodiments, as a mapping of image blocks surrounding a feature point to descriptors. Descriptors can be numerical quantities. U.S. patent application Ser. No. 16 / 190,948 describes calculating descriptors for feature points, and its entirety is incorporated herein by reference.

[0270] exist Figure 14 In the example of , a descriptor is calculated for each feature point in each image frame 1130. Based on the descriptors and / or the feature points and / or the image itself, the image frame may be identified as a keyframe 1140. In the illustrated embodiment, a keyframe is an image frame that meets certain criteria, which is then selected for further processing. For example, when making a tracking map, image frames that add meaningful information to the map may be selected as keyframes to be integrated into the map. On the other hand, image frames that substantially overlap with areas for which image frames have already been integrated into the map may be discarded so that these image frames do not become keyframes. Alternatively or additionally, keyframes may be selected based on the number and / or type of feature points in the image frame. In Figure 14 In the embodiment of FIG. 1 , keyframes 1150 selected for inclusion in the tracking map may also be considered keyframes for determining the PCF, although different or additional criteria may be used to select keyframes for generating the PCF.

[0271] Although Figure 14 Keyframes are shown for further processing, but the information obtained from the images can also be processed in other forms. For example, feature points such as in Keyrig can be processed instead or in addition. Furthermore, while keyframes are described as being derived from a single image frame, there need not be a one-to-one relationship between keyframes and the image frames obtained. For example, keyframes can be obtained from multiple image frames, such as by stitching together or aggregating the image frames so that only features that appear in multiple images are retained in the keyframe.

[0272] The key frame may include image information and / or metadata associated with the image information. In some embodiments, the key frame may be captured by cameras 44, 46 ( Figure 9 ) captured images can be calculated into one or more keyframes (e.g., keyframes 1, 2). In some embodiments, a keyframe may include a camera pose. In some embodiments, a keyframe may include one or more camera images captured at the camera pose. In some embodiments, the XR system may determine that a portion of the camera image captured at the camera pose is useless and therefore does not include that portion in the keyframe. Therefore, using keyframes to align new images with early scene knowledge can reduce the use of XR system computing resources. In some embodiments, a keyframe may include an image and / or image data located at a position with a direction / angle. In some embodiments, a keyframe may include a position and direction from which one or more map points can be observed. In some embodiments, a keyframe may include a coordinate system with an ID. U.S. Patent Application No. 15 / 877,359 describes keyframes and is hereby incorporated by reference in its entirety.

[0273] Some or all of the keyframes 1140 may be selected for further processing, such as generating persistent gestures 1150 for the keyframes. This selection may be based on characteristics of all or a subset of feature points in the image frame. These characteristics may be determined by processing descriptors, features, and / or the image frame itself. As a specific example, this selection may be based on clusters of feature points identified as potentially associated with persistent objects.

[0274] Each keyframe is associated with the camera pose at which it was acquired. For keyframes selected for processing as persistent poses, this pose information can be saved along with other metadata about the keyframe (such as the acquisition time and / or WiFi fingerprint and / or GPS coordinates at the acquisition location). In some embodiments, metadata such as GPS coordinates can be used alone or in combination as part of the positioning process.

[0275] A persistent gesture is a source of information that the device uses to orient itself relative to previously acquired information about the physical world. For example, if the keyframe that created the persistent gesture is incorporated into a map of the physical world, the device can orient itself relative to the persistent gesture using a sufficient number of feature points in the keyframe that are associated with the persistent gesture. The device can align a current image of its surroundings that it is capturing with the persistent gesture. This alignment can be based on matching the current image with the image 1110, features 1120, and / or descriptors 1130 that gave rise to the persistent gesture, or any subset of the image or these features or descriptors. In some embodiments, the current image frame that matches the persistent gesture may be another keyframe that has been incorporated into the device's tracking map.

[0276] Information about persistent gestures can be stored in a format that facilitates sharing among multiple applications executing on the same or different devices. Figure 14 In an example, some or all persistent gestures can be reflected as a persistent coordinate system (PCF) 1160. Like persistent gestures, the PCF can be associated with a map and can include a feature set or other information that the device uses to determine its orientation relative to the PCF. The PCF can include a transform that defines its transformation relative to its map origin, such that the device can determine its position relative to any object in the physical world reflected in the map by relating its position to the PCF.

[0277] Because PCFs provide a mechanism for determining position relative to physical objects, applications such as application 1180 can define the position of virtual objects relative to one or more PCFs that serve as anchors for virtual content 1170. For example, Figure 14 App 1 is shown as associating its virtual content 2 with PCF 1.2. Similarly, App 2 associates its virtual content 3 with PCF 1.2. App 1 is also shown as associating its virtual content 1 with PCF 4.5, and App 2 is shown as associating its virtual content 4 with PCF 3. In some embodiments, PCF 3 can be based on image 3 (not shown), and PCF 4.5 can be based on image 4 and image 5 (not shown), similar to how PCF 1.2 is based on image 1 and image 2. When rendering this virtual content, the device can apply one or more transformations to calculate information such as the position of the virtual content relative to the device's display and / or the position of physical objects relative to the desired position of the virtual content. Using the PCF as a reference can simplify such calculations.

[0278] In some embodiments, a persistent gesture can be a coordinate location and / or orientation with one or more associated keyframes. In some embodiments, a persistent gesture can be automatically created after the user has traveled a certain distance (e.g., three meters). In some embodiments, a persistent gesture can serve as a reference point during positioning. In some embodiments, a persistent gesture can be stored in the connected world (e.g., connected world module 538).

[0279] In some embodiments, a new PCF can be determined based on a predefined distance allowed between adjacent PCFs. In some embodiments, one or more persistent gestures can be calculated as a PCF when the user travels a predetermined distance (e.g., five meters). In some embodiments, the PCF can be associated with one or more world coordinate systems and / or canonical coordinate systems, for example, in a connected world. In some embodiments, the PCF can be stored in a local and / or remote database, depending on, for example, security settings.

[0280] Figure 15 A method 4700 of establishing and using a persistent coordinate system according to some embodiments is shown. The method 4700 may capture (act 4702) an image of a scene (e.g., Figure 14 Multiple cameras can be used and one camera can generate multiple images, for example in one stream.

[0281] Method 4700 may include extracting (4704) points of interest (e.g., Figure 7 Map point 702, Figure 14 1120 in the features), generating (action 4706) a descriptor of the extracted interest point (e.g., Figure 14 1130 in the image) and generates (action 4708) a keyframe (e.g., keyframe 1140) based on the descriptor. In some embodiments, the method may compare the points of interest in the keyframes and form pairs of keyframes that share a predetermined number of points of interest. The method may use the separate keyframe pairs to reconstruct a portion of the physical world. The mapped portion of the physical world may be saved as a 3D feature (e.g., Figure 7Keyrig 704 in ). In some embodiments, selected portions of keyframe pairs can be used to construct 3D features. In some embodiments, the results of the drawing can be selectively saved. Keyframes that are not used to construct the 3D features can be associated with the 3D features by pose, for example, the distance between the keyframes is represented by a covariance matrix between the poses of the keyframes. In some embodiments, keyframe pairs can be selected to construct the 3D features so that the distance between each two constructed 3D features is within a predetermined distance, which can be determined to balance the amount of computation required and the accuracy of the resulting model. This approach can provide a suitable amount of data for the physical world model to facilitate efficient and accurate computation using the XR system. In some embodiments, the covariance matrix of two images can include a covariance matrix between the poses (e.g., six degrees of freedom) of the two images.

[0282] Method 4700 may include generating (act 4710) a persistent gesture based on the keyframes. In some embodiments, the method may include generating a persistent gesture based on a 3D feature reconstructed from the keyframe pair. In some embodiments, the persistent gesture may be attached to the 3D feature. In some embodiments, the persistent gesture may include a gesture of the keyframes used to construct the 3D feature. In some embodiments, the persistent gesture may include an average gesture of the keyframes used to construct the 3D feature. In some embodiments, the persistent gesture may be generated such that a distance between adjacent persistent gestures is within a predetermined value, such as within a range of one meter to five meters, any value therebetween, or any other suitable value. In some embodiments, the distance between adjacent persistent gestures may be represented by a covariance matrix of the adjacent persistent gestures.

[0283] Method 4700 may include generating (act 4712) a PCF based on the persistent gesture. In some embodiments, the PCF may be attached to the 3D feature. In some embodiments, the PCF may be associated with one or more persistent gestures. In some embodiments, the PCF may include the pose of one of the associated persistent gestures. In some embodiments, the PCF may include the average pose of the poses of the associated persistent gestures. In some embodiments, the PCF may be generated so that the distance between adjacent PCFs is within a predetermined value, such as within a range of three meters to ten meters, any value therebetween, or any other suitable value. In some embodiments, the distance between adjacent PCFs may be represented by a covariance matrix of adjacent PCFs. In some embodiments, the PCF may be exposed to the XR application, for example, via an application programming interface (API), so that the XR application can access a model of the physical world through the PCF without accessing the model itself.

[0284] Method 4700 may include associating image data of a virtual object to be displayed by the XR device with at least one PCF (act 4714). In some embodiments, the method may include calculating a translation and orientation of the virtual object relative to the associated PCF. It should be understood that it is not necessary to associate the virtual object with a PCF generated by the device on which the virtual object is placed. For example, the device may retrieve a PCF stored within a canonical map in the cloud and associate the virtual object with the retrieved PCF. It should be understood that as the PCF adjusts over time, the virtual object may move with the associated PCF.

[0285] Figure 16 Visual data and algorithms of the first XR device 12 . 1 and the second XR device 12 . 2 and the server 20 are shown in accordance with some embodiments. Figure 16 The components shown can be operated to perform some or all of the operations associated with generating, updating, and / or using spatial information (such as a persistent pose, a persistent coordinate system, a tracking map, or a canonical map as described herein). Although not shown, the first XR device 12.1 can be configured identically to the second XR device 12.2. The server 20 can have a map storage routine 118, a canonical map 120, a map sender 122, and a map merging algorithm 124.

[0286] The second XR device 12.2, which may be in the same scene as the first XR device 12.1, may include a persistent coordinate frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that can be used to render virtual content, and a frame embedding generator 308 (see Figure 21 In some embodiments, the map download system 126, the PCF identification system 128, the map Figure 2 , positioning module 130, canonical map merger 132, canonical map 133, and map publisher 136 may be grouped into a connected world unit 1304. The PCF integration unit 1300 may be connected to the connected world unit 1304 and other components of the second XR device 12.2 to enable retrieval, generation, use, upload, and download of PCF.

[0287] Maps including PCFs can achieve more persistence in a changing world. In some embodiments, locating a tracking map, such as one containing image matching features, can include selecting features representing persistent content from a map composed of PCFs, which enables fast matching and / or localization. For example, a world where people enter and exit a scene and objects such as doors move relative to the scene requires less storage space and lower transmission rates, and allows the scene to be mapped using individual PCFs and their relationships relative to each other (e.g., an integrated constellation of PCFs).

[0288] In some embodiments, the PCF integration unit 1300 may include a PCF 1306 that was previously stored in a data storage on a storage unit of the second XR device 12.2, a PCF tracker 1308, a persistent pose acquirer 1310, a PCF checker 1312, a PCF generation system 1314, a coordinate system calculator 1316, a persistent pose calculator 1318, and three transformers, including a tracking map and persistent pose transformer 1320, a persistent pose and PCF transformer 1322, and a PCF and image data transformer 1324.

[0289] In some embodiments, the PCF tracker 1308 may have an on prompt and a off prompt selectable by the application 1302. The application 1302 may be executed by the processor of the second XR device 12.2 to, for example, display virtual content. The application 1302 may have a call to turn on the PCF tracker 1308 via the on prompt. The PCF tracker 1308 may generate a PCF when the PCF tracker 1308 is turned on. The application 1302 may have a subsequent call to turn off the PCF tracker 1308 via the off prompt. When the PCF tracker 1308 is turned off, the PCF tracker 1308 terminates PCF generation.

[0290] In some embodiments, the server 20 may include a plurality of persistent gestures 1332 and a plurality of PCFs 1330 that have been previously saved in association with the canonical map 120. The map sender 122 may send the canonical map 120 along with the persistent gestures 1332 and / or PCFs 1330 to the second XR device 12.2. The persistent gestures 1332 and PCFs 1330 may be stored on the second XR device 12.2 in association with the canonical map 133. Figure 2 When locating to the standard map 133, you can Figure 2 The persistent gesture 1332 and the PCF 1330 are stored in association.

[0291] In some embodiments, the persistent gesture acquirer 1310 may acquire the Figure 2 A PCF checker 1312 may be connected to the persistent gesture retriever 1310. The PCF checker 1312 may retrieve a PCF from the PCF 1306 based on the persistent gesture retrieved by the persistent gesture retriever 1310. The PCFs retrieved by the PCF checker 1312 may form an initial PCF set for PCF-based image display.

[0292] In some embodiments, the application 1302 may need to generate additional PCFs. For example, if the user moves to an area that was not previously mapped, the application 1302 may turn on the PCF tracker 1308. The PCF generation system 1314 may be connected to the PCF tracker 1308 and generate additional PCFs as the user moves to an area that was not previously mapped. Figure 2 Start expanding and start based on the ground Figure 2 Generating PCFs. The PCFs generated by the PCF generation system 1314 may form a second set of PCFs that may be used for PCF-based image display.

[0293] The coordinate system calculator 1316 may be connected to the PCF checker 1312. After the PCF checker 1312 retrieves the PCF, the coordinate system calculator 1316 may call the head coordinate system 96 to determine the head pose of the second XR device 12.2. The coordinate system calculator 1316 may also call the persistent pose calculator 1318. The persistent pose calculator 1318 may be directly or indirectly connected to the frame embedding generator 308. In some embodiments, an image / frame may be designated as a keyframe after traveling a threshold distance (e.g., 3 meters) from the last keyframe. The persistent pose calculator 1318 may generate a persistent pose based on multiple (e.g., three) keyframes. In some embodiments, the persistent pose may be substantially an average of the coordinate systems of the multiple keyframes.

[0294] Tracking map and persistent pose changer 1320 can be connected to the ground Figure 2 and persistent pose calculator 1318. Tracking map and persistent pose converter 1320 can Figure 2 Transform to a persistent pose to determine relative to the ground Figure 2 The persistent posture at the origin of .

[0295] Persistent pose and PCF converter 1322 may be connected to tracking map and persistent pose converter 1320 and further connected to PCF checker 1312 and PCF generation system 1314. Persistent pose and PCF converter 1322 may convert a persistent pose (a persistent pose into which the tracking map has been converted) into a PCF from PCF checker 1312 and PCF generation system 1314 to determine a PCF relative to the persistent pose.

[0296] PCF and image data converter 1324 may be connected to persistent gesture and PCF converter 1322 and data channel 62. PCF and image data converter 1324 converts the PCF into image data 68. Rendering engine 30 may be connected to PCF and image data converter 1324 to display image data 68 relative to the PCF to a user.

[0297] The PCF integration unit 1300 may store additional PCFs generated using the PCF generation system 1314 within the PCF 1306. The PCF 1306 may be stored relative to the persistent gesture. Figure 2 When sending to the server 20, the map publisher 136 can retrieve the PCF 1306 and the persistent gesture associated with the PCF 1306. The map publisher 136 will also Figure 2 The associated PCF and persistent gesture are sent to the server 20. When the map storage routine 118 of the server 20 stores the map Figure 2 The map storage routine 118 may also store the persistent gestures and PCFs generated by the second viewing device 12.2. The map merging algorithm 124 may create a canonical map 120 where the map Figure 2 The persistent posture and PCF are associated with the canonical map 120 and stored in the persistent posture 1332 and PCF 1330, respectively.

[0298] The first XR device 12.1 may include a PCF integration unit similar to the PCF integration unit 1300 of the second XR device 12.2. When the map sender 122 sends the canonical map 120 to the first XR device 12.1, the map sender 122 may send a persistent gesture 1332 and a PCF 1330 associated with the canonical map 120 and originating from the second XR device 12.2. The first XR device 12.1 may store the PCF and the persistent gesture in a data store on a storage device of the first XR device 12.1. The first XR device 12.1 may then utilize the persistent gesture and PCF originating from the second XR device 12.2 for image display relative to the PCF. Additionally or alternatively, the first XR device 12.1 may retrieve, generate, use, upload, and download the PCF and the persistent gesture in a manner similar to that described above for the second XR device 12.2.

[0299] In the example shown, the first XR device 12.1 generates a local tracking map (hereinafter referred to as a “map”). Figure 1 ”) and the map storage routine 118 receives the map from the first XR device 12.1 Figure 1 The map storage routine 118 then stores the map Figure 1 Stored on the storage device of the server 20 as the specification map 120 .

[0300] The second XR device 12 . 2 includes a map download system 126 , an anchor point identification system 128 , a positioning module 130 , a canonical map merger 132 , a local content positioning system 134 , and a map publisher 136 .

[0301] In use, the map sender 122 sends the canonical map 120 to the second XR device 12 . 2 , and the map download system 126 downloads the canonical map 120 from the server 20 and stores it as a canonical map 133 .

[0302] The anchor point identification system 128 is connected to the world surface determination routine 78. The anchor point identification system 128 identifies anchor points based on objects detected by the world surface determination routine 78. The anchor point identification system 128 generates a second map (map) using the anchor points. Figure 2 As shown in loop 138, the anchor point identification system 128 continues to identify anchor points and continues to update the ground Figure 2 The anchor point positions are recorded as three-dimensional data based on data provided by the world surface determination routine 78. The world surface determination routine 78 receives images from the real object detection camera 44 and depth data from the depth sensor 135 to determine the positions of surfaces and their relative distances from the depth sensor 135.

[0303] The positioning module 130 is connected to the specification map 133 and the Figure 2 The positioning module 130 repeatedly attempts to Figure 2 The normative map 133 is located. The normative map merger 132 is connected to the normative map 133 and the map Figure 2 When the positioning module 130 sets the ground Figure 2 When the standard map 133 is located, the standard map merger 132 merges the standard map 133 into the map. Figure 2 The map is then updated with the missing data included in the canonical map. Figure 2 .

[0304] The local content location system 134 is connected to the ground Figure 2 For example, the local content location system 134 may be a system whereby a user locates local content at a specific location within a world coordinate system. The local content then attaches itself to the local Figure 2 The local to world coordinate converter 104 converts the local coordinate system to the world coordinate system based on the settings of the local content positioning system 134. The functions of the rendering engine 30, the display system 42 and the data channel 62 have been described with reference to Figure 2 Described.

[0305] Map publisher 136 will Figure 2 Upload to the server 20. The map storage routine 118 of the server 20 then Figure 2 Stored in the storage medium of the server 20.

[0306] Map Merge Algorithm 124 Figure 2Merge with the canonical map 120. When more than two maps (e.g., three or four maps relating to the same or adjacent areas of the physical world) are stored, a map merging algorithm 124 merges all of the maps into the canonical map 120 to render a new canonical map 120. A map sender 122 then sends the new canonical map 120 to any and all devices 12.1 and 12.2 in the area represented by the new canonical map 120. When devices 12.1 and 12.2 align their respective maps with the canonical map 120, the canonical map 120 becomes the improved map.

[0307] Figure 17 An example of keyframe generation for a scene map according to some embodiments is shown. In the example shown, a first keyframe, KF1, is generated for the door on the left wall of the room. A second keyframe, KF2, is generated for the corner region where the floor, left wall, and right wall intersect. A third keyframe, KF3, is generated for the window region on the right wall of the room. A fourth keyframe, KF4, is generated for the far end of the carpet on the floor of the room. A fifth keyframe, KF5, is generated for the carpet region closest to the user.

[0308] Figure 18 shows the generation of Figure 17 In some embodiments, a new persistent pose (PP) is created when the device measures a threshold distance traveled and / or when an application requests a new persistent pose (PP). In some embodiments, the threshold distance can be 3 meters, 5 meters, 20 meters, or any other suitable distance. Selecting a smaller threshold distance (e.g., 1m) can result in an increased computational load because a larger number of PPs will be created and managed compared to a larger threshold distance. Selecting a larger threshold distance (e.g., 40m) can result in increased virtual content placement errors because fewer PPs are created, which will result in fewer PCFs being created, meaning that virtual content attached to the PCF is at a larger distance from the PCF (e.g., 30m), and errors increase as the distance from the PCF to the virtual content increases.

[0309] In some embodiments, a PP may be created when a new session begins. This initial PP may be considered zero and may be visualized as the center of a circle with a radius equal to the threshold distance. When the device reaches the perimeter of the circle, and in some embodiments, an application requests a new PP, the new PP may be placed at the device's current location (at the threshold distance). In some embodiments, if the device can find an existing PP within the threshold distance from the device's new location, a new PP is not created at the threshold distance. In some embodiments, when a new PP is created (e.g., Figure 14When a PP 1150 is created in a PP, the device appends one or more recent keyframes to the PP. In some embodiments, the location of the PP relative to the keyframes can be based on the device's location when the PP is created. In some embodiments, a PP is not created when the device travels a threshold distance unless an application requests a PP.

[0310] In some embodiments, when an application has virtual content to display to the user, the application can request a PCF from the device. The PCF request from the application may trigger a PP request, and a new PP is created after the device travels a threshold distance. Figure 18 A first persistent pose PP1 is shown, which may have recent keyframes (eg, KF1, KF2, and KF3) attached thereto, for example, by computing relative poses between the keyframes and the persistent pose. Figure 18 Also shown is a second persistent pose PP2, which may be attached to the most recent keyframes (eg, KF4 and KF5).

[0311] Figure 19 shows the generation of Figure 17 1 and 2. In the example shown, PCF1 may include PP1 and PP2. As described above, PCFs may be used to display image data relative to the PCFs. In some embodiments, each PCF may have coordinates in another coordinate system (e.g., a world coordinate system) and a PCF descriptor that uniquely identifies the PCF, for example. In some embodiments, the PCF descriptor may be calculated based on feature descriptors of features in a frame associated with the PCF. In some embodiments, various PCF constellations may be combined to represent the real world in a persistent manner that requires less data and less data transmission.

[0312] 20A to 20C is a schematic diagram showing an example of establishing and using a persistent coordinate system. Figure 20A Two users 4802A, 4802B are shown, each with its own local tracking map 4804A, 4804B that has not yet been localized to a canonical map. Each user's origin 4806A, 4806B is depicted by a coordinate system (e.g., a world coordinate system) in their respective regions. These origins for each tracking map are local to each user, as they depend on the orientation of their respective devices when tracking is initiated.

[0313] As the user device's sensors scan the environment, the device can capture images, as shown above in conjunction with Figure 14 As described, images may contain features representing persistent objects, such that these images may be classified as keyframes for creating persistent gestures. In this example, track map 4802A includes persistent gestures (PPs) 4808A; track 4802B includes PPs 4808B.

[0314] In addition, as above combined Figure 14 As described above, some PPs may be classified as PCFs, which are used to determine the orientation of virtual content to render it to the user. Figure 20B It shows that the XR devices worn by each user 4802A, 4802B can create local PCFs 4810A, 4810B based on PP4808A, 4808B. Figure 20C It is shown that persistent content 4812A, 4812B (e.g., virtual content) can be attached to the PCF 4810A, 4810B through the corresponding XR device.

[0315] In this example, virtual content can have a virtual content coordinate system that can be used by the application generating the virtual content, regardless of how the virtual content will be displayed. For example, virtual content can be specified as surfaces, such as triangles of a mesh, at specific positions and angles relative to the virtual content coordinate system. To render the virtual content to the user, the positions of these surfaces can be determined relative to the user perceiving the virtual content.

[0316] Attaching virtual content to the PCF simplifies the computations involved in determining the position of the virtual content relative to the user. The position of the virtual content relative to the user can be determined by applying a series of transformations. Some of these transformations may vary and be updated frequently. Others of these transformations may be stable and updated frequently or not at all. Regardless, the transformations can be applied with relatively low computational overhead, allowing for frequent updates of the virtual content's position relative to the user, providing a realistic appearance to the rendered virtual content.

[0317] exist Figures 20A-20C In the example shown in FIG1 , user 1's device has a coordinate system related to a coordinate system defining the origin of the map by the transformation rig1_T_w1. User 2's device has a similar transformation rig2_T_w2. These transformations can be expressed as 6-DOF transformations, specifying a translation and a rotation to align the device coordinate system with the map coordinate system. In some embodiments, the transformations can be expressed as two separate transformations, one specifying the translation and the other specifying the rotation. Therefore, it should be understood that the transformations can be expressed in a form that simplifies computations or otherwise provides advantages.

[0318] The transformations between the origin of the tracking map and the PCF identified by each user device are denoted as pcf1_T_w1 and pcf2_T_w2. In this example, the PCF and PP are identical, so the same transformation also characterizes the PP.

[0319] The position of the user equipment relative to the PCF can therefore be calculated by a series of applications of these transformations, such as rig1_T_pcf1 = (rig1_T_w1) * (pcf1_T_w1).

[0320] like Figure 20C As shown, the virtual content is positioned relative to the PCF with a transformation obj1_T_pcf1. This transformation can be set by the application generating the virtual content, which can receive information from the world reconstruction system describing the physical objects about the PCF. In order to render the virtual content to the user, a transformation to the coordinate system of the user device is calculated, which can be calculated by associating the virtual content coordinate system with the origin of the tracking map through the transformation obj1_t_w1 = (obj1_T_pcf1) * (pcf1_T_w1). This transformation is then associated with the user device through a further transformation rig1_T_w1.

[0321] The position of virtual content can change based on the output from the application generating the virtual content. When this happens, the end-to-end transformation from the source coordinate system to the target coordinate system can be recalculated. Furthermore, the user's position and / or head pose can change as the user moves. Therefore, the transformation rig1_T_w1 may change, as may any end-to-end transformation that depends on the user's position or head pose.

[0322] The transformation rigl_T_wl can be updated by the user's movement based on tracking the user's position relative to stationary objects in the physical world. Such tracking can be performed by the headset tracking component (described above) or other components of the system that process the image sequence. Such updates can be performed by determining the user's posture relative to a stationary reference frame (such as PP).

[0323] In some embodiments, the position and orientation of the user device can be determined relative to the most recent persistent gesture, or in this example, the PCF (since the PP is used as the PCF). This determination can be made by identifying feature points that characterize the PP in the current image captured by the sensors on the device. Using image processing techniques such as stereo image analysis, the position of the device relative to these feature points can be determined. Based on this data, the system can calculate the change in transformation associated with the user's motion based on the relationship rig1_T_pcf1=(rig1_T_w1)*(pcf1_T_w1).

[0324] The system can determine and apply transformations in an order that saves computation. For example, the need to calculate rig1_T_w1 from measurements that generate rig1_T_pcf1 can be avoided by tracking user gestures and defining the position of virtual content relative to a PP or PCF based on persistent gestures. In this way, the transformation from the source coordinate system of the virtual content to the target coordinate system of the user device can be based on a transformation measured according to the expression (rig1_T_pcf1)*(obj1_t_pcf1), where the first transformation is measured by the system and subsequent transformations are provided by the application that specifies the virtual content to be rendered. In embodiments where the virtual content is positioned relative to the origin of a map, the end-to-end transformation can be based on a further transformation between map coordinates and PCF coordinates to relate the virtual content coordinate system to the PCF coordinate system. In embodiments where the virtual content is positioned relative to a different PP or PCF than the PP or PCF used to track the user's position, a transformation between the two can be applied. Such a transformation can be fixed and, for example, can be determined from a map where both appear.

[0325] For example, a transformation-based approach can be implemented in a device that has components that process sensor data to build a tracking map. As part of this process, these components can identify feature points that serve as persistent gestures, which in turn can be converted into PCFs. These components can limit the number of persistent gestures generated for the map to provide suitable spacing between persistent gestures while allowing the user (regardless of their location in the physical environment) to be close enough to the persistent gesture location to accurately calculate the user's gesture, as described above in conjunction with Figure 17-Figure 19 As described. When the persistent pose closest to the user is updated due to user movement, refinement of the tracking map, or other reasons, any transformations used to calculate the position of the virtual content relative to the user (depending on the position of the PP and, if a PCF is used, the position of the PCF) are updated and stored for use at least until the user leaves the persistent pose. Nevertheless, by calculating and storing the transformations, the computational burden of each update of the virtual content position can be relatively low, and thus can be performed with relatively low latency.

[0326] Figures 20A-20C Positioning is shown relative to the tracking map, and each device has its own tracking map. However, transformations can be generated relative to any map coordinate system. Content persistence across user sessions of the XR system can be achieved by using persistent maps. Shared experiences for users can also be facilitated by using maps that multiple user devices can be directed to.

[0327] In some embodiments, described in more detail below, the location of virtual content can be specified relative to coordinates in a canonical map, which can be formatted to allow any of multiple devices to use the map. Each device can maintain a tracking map and can determine changes in the user's posture relative to the tracking map. In this example, the transformation between the tracking map and the canonical map can be determined by a "localization" process, which can be performed by matching structures in the tracking map (such as one or more persistent postures) with one or more structures of the canonical map (such as one or more PCFs).

[0328] Techniques for creating and using canonical maps in this manner are described in more detail below.

[0329] Depth Keyframes

[0330] The techniques described herein rely on comparisons of image frames. For example, to establish the position of the device relative to a tracking map, a new image may be captured using sensors worn by the user, and the XR system may search the image set used to create the tracking map for an image that shares at least a predetermined number of points of interest with the new image. As an example of another scenario involving image frame comparisons, a tracking map can be localized to a canonical map by first finding image frames in the tracking map that are associated with persistent gestures (similar to image frames associated with PCFs in a canonical map). Alternatively, a transformation between two canonical maps can be computed by first finding similar image frames in the two maps.

[0331] Depth keyframes provide a method for reducing the amount of processing required to identify similar image frames. For example, in some embodiments, a comparison can be made between image features in the new 2D image (e.g., "2D features") and 3D features in the map. This comparison can be made in any suitable manner, such as by projecting the 3D image into a 2D plane. Conventional methods such as bag of words (BoW) search for the 2D features of the new image in a database that includes all 2D features in the map, which can require significant computational resources, especially when the map represents a large area. The conventional method then locates images that share at least one 2D feature with the new image, including images that are not useful for locating meaningful 3D features in the map. The conventional method then locates 3D features that are not meaningful relative to the 2D features in the new image.

[0332] The inventors have recognized and understood techniques for retrieving images from a map using fewer memory resources (e.g., one-quarter the memory resources used by BoW), higher efficiency (2.5 ms processing time per keyframe, 100 μs for 500 keyframes compared to), and higher accuracy (e.g., 20% higher retrieval recall than BoW for a 1024-dimensional model and 5% higher retrieval recall than BoW for a 256-dimensional model).

[0333] To reduce computation, a descriptor can be calculated for an image frame, which is used to compare the image frame with other image frames. The descriptor can be stored as an alternative to or in addition to the image frame and feature points. In a map that can generate persistent poses and / or PCFs from image frames, the descriptors of one or more image frames used to generate each persistent pose or PCF can be stored as part of the persistent pose and / or PCF.

[0334] In some embodiments, descriptors can be calculated based on feature points in an image frame. In some embodiments, a neural network is configured to calculate a unique frame descriptor representing an image. The image can have a resolution greater than 1 megabyte in order to capture sufficient detail of the 3D environment within the field of view of a device worn by the user. The frame descriptor can be much shorter, such as a string of numbers, for example, in the range of 128 bytes to 512 bytes, or any number in between.

[0335] In some embodiments, a neural network is trained so that the calculated frame descriptors indicate similarity between images. An image in a map can be located by identifying the nearest images in a database including images used to generate the map. These nearest images can have frame descriptors within a predetermined distance from the frame descriptor of the new image. In some embodiments, the distance between images can be represented by the difference between the frame descriptors of the two images.

[0336] Figure 21 is a block diagram illustrating a system for generating a descriptor for a single image, according to some embodiments. In the illustrated example, a frame embedding generator 308 is shown. In some embodiments, frame embedding generator 308 may be used within server 20, but, alternatively or additionally, may be executed in whole or in part within one of XR devices 12.1 and 12.2, or any other device that processes an image for comparison with other images.

[0337] In some embodiments, the frame embedding generator can be configured to generate a condensed image data representation from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes) that, despite the reduced size, still indicates the content in the image. In some embodiments, the frame embedding generator can be used to generate an image data representation that can be a keyframe or a frame used in other aspects. In some embodiments, the frame embedding generator 308 can be configured to convert an image at a particular position and orientation into a unique string of numbers (e.g., 256 bytes). In the example shown, an image 320 captured by an XR device can be processed by a feature extractor 324 to detect points of interest 322 in the image 320. The points of interest can be based on or can be derived based on the identified feature points, as described above for feature 1120 ( Figure 14 ) or as otherwise described herein. In some embodiments, a point of interest may be represented by a descriptor, such as that described above with reference to descriptor 1130 ( Figure 14 ), the descriptor can be generated using a deep sparse feature method. In some embodiments, each interest point 322 can be represented by a digital string (e.g., 32 bytes). For example, there may be n features (e.g., 100), each of which is represented by a 32-byte string.

[0338] In some embodiments, the frame embedding generator 308 may include a neural network 326. The neural network 326 may include a multilayer perceptron unit 312 and a max pooling unit 314. In some embodiments, the multilayer perceptron (MLP) unit 312 may include a multilayer perceptron that can be trained. In some embodiments, the interest points 322 (e.g., descriptors of the interest points) may be reduced by the multilayer perceptron 312 and may be output as a weighted combination 310 of the descriptors. For example, the MLP may reduce n features to m features, which is less than n features.

[0339] In some embodiments, the MLP unit 312 can be configured to perform matrix multiplication. The multilayer perceptron unit 312 receives a plurality of interest points 322 of the image 320 and converts each interest point into a corresponding string of numbers (e.g., 256). For example, there may be 100 features, each of which can be represented by a string of 256 numbers. In this example, a matrix with 100 horizontal rows and 256 vertical columns can be created. Each row has a series of 256 numbers, and these numbers are of different sizes, some smaller and some larger. In some embodiments, the output of the MLP can be an n×256 matrix, where n represents the number of interest points extracted from the image. In some embodiments, the output of the MLP can be an m×256 matrix, where m is the number of interest points reduced from n.

[0340] In some embodiments, the MLP 312 may have a training phase during which the model parameters of the MLP are determined and a usage phase. Figure 25 The training MLP is shown. The input training data can include a set of three data, each of which includes 1) a query image, 2) a positive sample, and 3) a negative sample. The query image can be considered as a reference image.

[0341] In some embodiments, positive samples may include images that are similar to the query image. For example, in some embodiments, similarity means having the same object in both the query image and the positive sample image, but viewing the object from different angles. In some embodiments, similarity means having the same object in both the query image and the positive sample image, but the object is moved relative to the other image (e.g., left, right, up, down).

[0342] In some embodiments, negative samples may include images that are dissimilar to the query image. For example, in some embodiments, a dissimilar image does not contain any objects that are prominent in the query image or contains only a small portion (e.g., <10%, 1%) of the objects that are prominent in the query image. In contrast, for example, a similar image has a large portion (e.g., >50% or >75%) of the objects in the query image.

[0343] In some embodiments, interest points can be extracted from images in the input training data and converted into feature descriptors. These descriptors can be used for both Figure 25 The training images shown, and Figure 21 The extracted features are calculated in the operation of the frame embedding generator 308. In some embodiments, a deep sparse feature (DSF) process can be used to generate descriptors (e.g., DSF descriptors), as described in U.S. patent application Ser. No. 16 / 190,948. In some embodiments, the DSF descriptors are n×32 dimensional. The descriptors can then be passed through the model / MLP to create a 256-byte output. In some embodiments, the model / MLP can have the same structure as the MLP 312, so that once the model parameters are set through training, the resulting trained MLP can be used as the MLP 312.

[0344] In some embodiments, the feature descriptors (e.g., the 256-byte output from the MLP model) can then be sent to a triplet edge loss module (which can be used only during the training phase of the MLP neural network, not during the use phase of the MLP neural network). In some embodiments, the triplet edge loss module can be configured to select model parameters to reduce the difference between the 256-byte output of the query image and the 256-byte output of the positive samples, and to increase the difference between the 256-byte output of the query image and the 256-byte output of the negative samples. In some embodiments, the training phase can include feeding multiple triplet input images into the learning process to determine the model parameters. For example, the training process can continue until the difference for positive images is minimized and the difference for negative images is maximized, or until other suitable exit criteria are met.

[0345] Return Reference Figure 21 , the frame embedding generator 308 may include a pooling layer, shown here as a max pooling unit 314. The max pooling unit 314 may analyze each column to determine the maximum number in the corresponding column. The max pooling unit 314 may combine the maximum values ​​of the numbers in each column of the output matrix of the MLP 312 into a global feature string 316 of, for example, 256 numbers. It should be understood that the images processed in the XR system preferably have high-resolution frames, potentially reaching millions of pixels. The global feature string 316 is a relatively small number that takes up relatively little memory and is easier to search than an image (e.g., having a resolution greater than 1 megabyte). Therefore, the image can be searched without analyzing every raw frame from the camera, and the cost of storing 256 bytes instead of a full frame is also lower.

[0346] Figure 22 22 is a flow chart illustrating a method 2200 for computing an image descriptor according to some embodiments. Method 2200 may begin by receiving (act 2202) a plurality of images captured by an XR device worn by a user. In some embodiments, method 2200 may include determining (act 2204) one or more keyframes from the plurality of images. In some embodiments, act 2204 may be skipped and / or may occur after step 2210.

[0347] Method 2200 may include identifying (act 2206) one or more points of interest in the plurality of images using an artificial neural network, and computing (act 2208) feature descriptors for the respective points of interest using the artificial neural network. The method may include computing (act 2210) a frame descriptor for each image to represent the image based at least in part on the feature descriptors of the identified points of interest in the image computed using the artificial neural network.

[0348] Figure 2323 is a flowchart illustrating a method 2300 for localization using image descriptors, according to some embodiments. In this example, a new image frame depicting the current location of the XR device can be compared to image frames stored in connection with points in a map (such as the persistent gestures or PCFs described above). Method 2300 can begin by receiving (act 2302) a new image captured by an XR device worn by a user. Method 2300 can include identifying (act 2304) one or more recent keyframes in a database comprising keyframes used to generate one or more maps. In some embodiments, the recent keyframes can be identified based on coarse spatial information and / or previously determined spatial information. For example, the coarse spatial information can indicate that the XR device is located in a geographic area represented by a 50mx50m area of ​​the map. Image matching can be performed only for points within this area. As another example, based on tracking, the XR system can know that the XR device was previously near a first persistent gesture in the map and is moving in the direction of a second persistent gesture in the map. The second persistent gesture can be considered the most recent persistent gesture, and the keyframes stored with it can be considered the most recent keyframes. Alternatively or additionally, other metadata such as GPS data or WiFi fingerprints may be used to select the most recent keyframe or most recent set of keyframes.

[0349] Regardless of how the nearest keyframe is selected, the frame descriptor can be used to determine whether the new image matches any of the frames selected to be associated with the nearby persistent gesture. This determination can be made by comparing the frame descriptor of the new image with the frame descriptors of the nearest keyframe, or a subset of keyframes in the database selected in any other suitable manner, and selecting a keyframe with a frame descriptor within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors can be calculated by taking the difference between two numeric strings representing the two frame descriptors. In embodiments where the strings are treated as strings of multiple quantities, the difference can be calculated as a vector difference.

[0350] Once a matching image frame is identified, the orientation of the XR device relative to the image frame can be determined. Method 2300 can include performing (act 2306) feature matching for 3D features in the map corresponding to the most recently identified keyframe, and calculating (act 2308) a pose of the device worn by the user based on the feature matching results. In this way, computationally intensive matching of feature points in two images can be performed for as few as one image that has been determined to be a possible match for the new image.

[0351] Figure 2424 is a flow chart illustrating a method 2400 for training a neural network according to some embodiments. Method 2400 may begin by generating (act 2402) a dataset comprising a plurality of image sets. Each of the plurality of image sets may include a query image, a positive sample image, and a negative sample image. In some embodiments, the plurality of image sets may include synthetic record pairs configured to, for example, impart basic information, such as shape, to the neural network. In some embodiments, the plurality of image sets may include real record pairs that may be recorded from the physical world.

[0352] In some embodiments, the correct data (inlier) can be calculated by fitting a fundamental matrix between the two images. In some embodiments, the sparse overlap can be calculated as the intersection over union (IoU) of the interest points seen in the two images. In some embodiments, the positive sample can include at least twenty interest points used as correct data that are the same as the interest points in the query image. The negative sample can include less than ten correct data points. The negative sample has less than half of its sparse points overlap with the sparse points of the query image.

[0353] The method 2400 may include, for each image set, calculating (act 2404) a loss by comparing the query image to the positive and negative images. The method 2400 may include modifying (act 2406) the artificial neural network based on the calculated loss so that a distance between a frame descriptor generated by the artificial neural network for the query image and a frame descriptor generated for the positive images is less than a distance between a frame descriptor generated for the query image and a frame descriptor generated for the negative images.

[0354] It should be understood that although the methods and apparatus described above are configured to generate global descriptors for individual images, these methods and apparatus can be configured to generate descriptors for individual maps. For example, a map can include multiple keyframes, each of which can have a frame descriptor as described above. The max pooling unit can analyze the frame descriptors of the map keyframes and combine the frame descriptors into a unique map descriptor for the map.

[0355] Furthermore, it should be understood that other architectures can be used for the processing described above. For example, separate neural networks are described for generating DSF descriptors and frame descriptors. This approach saves computation. However, in some embodiments, frame descriptors can be generated from selected feature points without first generating DSF descriptors.

[0356] Map Ranking and Merging

[0357] Described herein are methods and apparatus for ranking and merging multiple maps of an environment in a cross-reality (XR) system. Map merging can enable maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking maps can enable efficient performance of techniques as described herein, including map merging, which involves selecting maps from a set of atlases based on similarity. In some embodiments, for example, a set of canonical maps can be maintained by the system, the format of which can be set in a manner that can be accessed by any of a plurality of XR devices. These canonical maps can be formed by merging selected tracking maps from the devices with other tracking maps or previously stored canonical maps. For example, canonical maps can be ranked for selection of one or more canonical maps to merge with a new tracking map and / or selection of one or more canonical maps from a set for use in a device.

[0358] To provide a realistic XR experience to the user, the XR system must understand the user's physical environment so that the position of virtual objects can be correctly associated with real objects. Information about the user's physical environment can be obtained from an environmental map of the user's location.

[0359] The inventors have recognized and appreciated that an XR system can provide an enhanced XR experience including real and / or virtual content to multiple users sharing the same world by enabling efficient sharing of environmental maps of the real / physical world collected by multiple users, regardless of whether those users are present in the world at the same time or at different times. However, there are significant challenges in providing such a system. Such a system may store multiple maps generated by multiple users and / or the system may store multiple maps generated at different times. For operations that may be performed using previously generated maps, such as positioning as described above, a significant amount of processing may be required to identify relevant environmental maps of the same world (e.g., the same real-world location) from all environmental maps collected in the XR system. In some embodiments, a device may only have access to a small number of environmental maps, such as for positioning. In some embodiments, a device may have access to a large number of environmental maps. For example, the inventors have recognized and appreciated that it may be difficult to identify relevant environmental maps from all possible environmental maps (such as, Figure 28 A technique for quickly and accurately ranking the relevance of maps of an environment among the universe of all canonical maps 120 in the user's environment. Highly ranked maps can then be selected for further processing, such as rendering virtual objects on a user's display that realistically interact with the physical world around the user, or merging map data collected by the user with stored maps to create a larger or more accurate map.

[0360] In some embodiments, a stored map relevant to a user's task at a location in the physical world can be identified by filtering the stored maps based on multiple criteria. These criteria can dictate a comparison of a tracking map generated by a wearable device of the user at the location with candidate environment maps stored in a database. The comparison can be performed based on metadata associated with the map, such as a Wi-Fi fingerprint detected by the device generating the map and / or the set of BSSIDs to which the device was connected when forming the map. The comparison can also be performed based on the compressed or uncompressed content of the map. For example, a comparison based on a compressed representation can be performed by comparing vectors calculated from the map content. For example, a comparison based on an uncompressed map can be performed by locating a tracking map within a stored map, or vice versa. Multiple comparisons can be performed sequentially based on the computation time required to reduce the number of candidate maps considered, with comparisons involving less computation being performed earlier in the sequence than other comparisons requiring more computation.

[0361] Figure 26 FIG2 shows an AR system 800 configured to rank and merge one or more environment maps according to some embodiments. The AR system may include a connected world model 802 of an AR device. Information to populate the connected world model 802 may come from sensors on the AR device, which may include data stored in a processor 804 (e.g., Figure 4 The computer executable instructions in the local data processing module 570 in the processor can perform some or all of the processing to convert the sensor data into a map. Such a map can be a tracking map because it can be built as the AR device collects sensor data during operation in an area. Area attributes can be associated with the tracking map. Figure 1 Area attributes may be provided to indicate the area represented by the tracking map. These area attributes may be geographic location identifiers, such as coordinates expressed as longitude and latitude or an ID used by the AR system to represent a location. Alternatively or additionally, area attributes may be measurement characteristics that are most likely unique to the area. For example, area attributes may be derived from parameters of wireless networks detected in the area. In some embodiments, area attributes may be associated with unique addresses of access points that are near and / or connected to the AR system. For example, area attributes may be associated with a MAC address or basic service set identifier (BSSID) of a 5G base station / router, a Wi-Fi router, or the like.

[0362] exist Figure 26 In the example of FIG, the tracking map can be merged with other maps of the environment. Map ranking unit 806 receives the tracking map from device PW 802 and communicates with map database 808 to select and rank environment maps from map database 808. The selected maps with higher rankings are sent to map merging unit 810.

[0363] The map merging unit 810 can perform a merging process on the maps sent from the map ranking unit 806. This merging process entails merging the tracking map with some or all ranked maps and sending the new merged map to the connected world model 812. The map merging unit can merge the maps by identifying overlapping maps depicting the physical world. These overlapping portions can be aligned so that the information in the two maps can be aggregated into a final map. The canonical map can be merged with other canonical maps and / or tracking maps.

[0364] Aggregation entails extending one map with information from another map. Alternatively or additionally, aggregation entails adjusting the representation of the physical world in one map based on information in another map. For example, a subsequent map may reveal that an object that generated a feature point has moved, so that the map can be updated based on the subsequent information. Alternatively, two maps may represent the same area with different feature points, and aggregation entails selecting a set of feature points from both maps to better represent the area. Regardless of the specific processing that occurs during the merge, in some embodiments, PCFs from all maps being merged can be retained so that applications that locate content relative to these PCFs can continue to operate. In some embodiments, the merging of maps can result in redundant persistent gestures, and some persistent gestures may be deleted. When a PCF is associated with a persistent gesture that is to be deleted, merging the maps may require modifying the PCF to be associated with the persistent gesture that remains in the map after the merge.

[0365] In some embodiments, the map can be refined as it is expanded and / or updated. Refinement requires computation to reduce internal inconsistencies between feature points that may represent the same object in the physical world. The inconsistency may be caused by inaccuracies in the poses associated with keyframes that provide feature points that represent the same object in the physical world. For example, such inconsistencies may be caused by the XR device calculating poses relative to a tracking map that is in turn built based on estimated poses, such that errors in the pose estimates accumulate, resulting in "drift" in pose accuracy over time. The map can be refined by performing bundle adjustments or other operations to reduce inconsistencies in feature points from multiple keyframes.

[0366] During refinement, the position of a persistent point relative to the map origin can change. Consequently, the transform associated with that persistent point (e.g., a persistent pose or PCF) may change. In some embodiments, the XR system can, in conjunction with map refinement (whether performed as part of a merge operation or for other reasons), recalculate the transforms associated with any persistent point that has changed. These transforms can be pushed from the components that compute the transforms to the components that use them, so that any use of the transforms is based on the updated position of the persistent point.

[0367] The connected world model 812 can be a cloud model that can be shared by multiple AR devices. The connected world model 812 can store or otherwise access a map of the environment in the map database 808. In some embodiments, when a previously calculated map of the environment is updated, the previous version of the map can be deleted to remove outdated data from the database. In some embodiments, when a previously calculated map of the environment is updated, the previous version of the map can be archived, allowing the previous version of the environment to be retrieved / viewed. In some embodiments, permissions can be set so that only AR systems with specific read / write access permissions can trigger the deletion / archiving of previous versions of the map.

[0368] AR devices in the AR system can access these environment maps created based on tracking maps provided by one or more AR devices / systems. The map ranking unit 806 can also be used to provide environment maps to AR devices. The AR device can send a message requesting its current position in the environment map, and the map ranking unit 806 can be used to select and rank the environment maps related to the requesting device.

[0369] In some embodiments, the AR system 800 may include a downsampling unit 814 configured to receive a merged map from the cloud PW 812. The merged map received from the cloud PW 812 may be in a storage format suitable for the cloud, which may include high-resolution information, such as a large number of PCFs or multiple image frames per square meter or a large set of feature points associated with the PCFs. The downsampling unit 814 may be configured to downsample the cloud format map to a format suitable for storage on the AR device. The device format map has less data, such as fewer PCFs or less data stored for each PCF, to accommodate the limited local computing power and storage space of the AR device.

[0370] Figure 27 is a simplified block diagram illustrating a plurality of canonical maps 120 that may be stored in a remote storage medium (e.g., a cloud). Each canonical map 120 may include a plurality of canonical map identifiers that indicate the location of the canonical map in physical space, such as somewhere on the Earth. These canonical map identifiers may include one or more of the following identifiers: a region identifier represented by a series of longitudes and latitudes, a frame descriptor (e.g., Figure 21 ), Wi-Fi fingerprints, feature descriptors (e.g., Figure 21), and device identifiers indicating one or more devices contributing to the map. In the example shown, the canonical maps 120 are arranged geographically in a two-dimensional pattern, as they might exist on the surface of the Earth. Canonical maps 120 can be uniquely identified by corresponding longitudes and latitudes, as any canonical maps with overlapping longitudes and latitudes can be merged into a new canonical map.

[0371] Figure 28 is a schematic diagram illustrating a method of selecting canonical maps according to some embodiments, which method may be used to locate a new tracking map to one or more canonical maps. The method may begin by accessing (act 120) a universe of canonical maps 120, which may, for example, be stored in a database in a connected world (e.g., connected world module 538). The universe of canonical maps may include canonical maps from all previously visited locations. The XR system may filter the universe of all canonical maps to a small subset or to only a single map. It should be understood that in some embodiments, due to bandwidth limitations, it may not be possible to send all canonical maps to a viewing device. Selecting a subset of possible candidates that are selected as matching tracking maps to be sent to the device may reduce bandwidth and latency associated with accessing a database of remote maps.

[0372] The method may include filtering (act 300) the full domain of the canonical map based on regions having a predetermined size and shape. Figure 27 In the example shown, each square represents an area. Each square may cover 50m x 50m. Each square has six adjacent areas. In some embodiments, act 300 may select at least one matching canonical map 120 that covers the longitude and latitude of the location identifier received from the XR device, as long as at least one map exists at that longitude and latitude. In some embodiments, act 300 may select at least one adjacent canonical map that covers the longitude and latitude adjacent to the matching canonical map. In some embodiments, act 300 may select multiple matching canonical maps and multiple adjacent canonical maps. For example, act 300 may reduce the number of canonical maps by approximately tenfold, e.g., from thousands to hundreds, to form the first filtered selection. Alternatively or additionally, criteria other than latitude and longitude may be used to identify adjacent maps. For example, the XR device may have previously been localized using a canonical map in the collection as part of the same session. The cloud service may retain information about the XR device, including previously localized maps. In this example, the maps selected in act 300 may include those that cover areas adjacent to the map to which the XR device was localized.

[0373] The method may include filtering (act 302) a first filtered selection of canonical maps based on a Wi-Fi fingerprint. Act 302 may determine a latitude and longitude based on a Wi-Fi fingerprint received from the XR device as part of a location identifier. Act 302 may compare the latitude and longitude from the Wi-Fi fingerprint with the latitude and longitude of the canonical map 120 to determine one or more canonical maps that form a second filtered selection. Act 302 may reduce the number of canonical maps by approximately ten times, for example, from hundreds to tens (e.g., 50) of canonical maps that form the second selection. For example, the first filtered selection may include 130 canonical maps, the second filtered selection may include 50 of the 130 canonical maps, and may not include the other 80 of the 130 canonical maps.

[0374] The method may include filtering (act 304) a second filter selection of the canonical map based on the keyframes. Act 304 may compare data representing an image captured by the XR device with data representing the canonical map 120. In some embodiments, the data representing the image and / or map may include feature descriptors (e.g., Figure 25 DSF descriptor in ) and / or global feature string (e.g., Figure 21 316 in). Action 304 may provide a third filtered selection of canonical maps. In some embodiments, for example, the output of action 304 may be only five of the 50 canonical maps identified after the second filtered selection. The map transmitter 122 then sends one or more canonical maps based on the third filtered selection to the viewing device. Action 304 may reduce the number of canonical maps by approximately ten times, for example, from tens to a single digit (e.g., 5) of canonical maps for the third selection. In some embodiments, the XR device may receive the canonical map in the third filtered selection and attempt to locate within the received canonical map.

[0375] For example, action 304 can filter the canonical map 120 based on the global feature string 316 of the canonical map 120 and based on the global feature string 316 of an image captured by the viewing device (eg, an image that can be part of a user's local tracking map). Figure 27 Each canonical map 120 thus has one or more global feature strings 316 associated with it. In some embodiments, the global feature strings 316 can be obtained when the XR device submits an image or feature details to the cloud and the cloud processes the image or feature details to generate the global feature strings 316 for the canonical map 120.

[0376] In some embodiments, the cloud can receive feature details of a live / new / current image captured by the viewing device, and the cloud can generate a global feature string 316 for the live image. The cloud can then filter the canonical map 120 based on the live global feature string 316. In some embodiments, the global feature string can be generated locally on the viewing device. In some embodiments, the global feature string can be generated remotely, such as on the cloud. In some embodiments, the cloud can send the filtered canonical map along with the global feature string 316 associated with the filtered canonical map to the XR device. In some embodiments, when the viewing device positions its tracking map to the canonical map, it can do so by matching the global feature string 316 of the local tracking map with the global feature string of the canonical map.

[0377] It should be understood that the operation of the XR device may not perform all of the actions (300, 302, 304). For example, if the universe of the canonical map is relatively small (e.g., 500 maps), the XR device attempting localization may filter the universe of the canonical map based on Wi-Fi fingerprints (e.g., action 302) and keyframes (e.g., action 304), but omit the region-based filtering (e.g., action 300). Furthermore, it is not necessary to compare the entire map. In some embodiments, for example, the comparison of two maps may result in the identification of common persistent points, such as persistent gestures or PCFs that appear in both the new map and the selected map from the universe of maps. In this case, descriptors may be associated with the persistent points, and these descriptors may be compared.

[0378] Figure 29 is a flow chart illustrating a method 900 for selecting one or more ranked environment maps, according to some embodiments. In the illustrated embodiment, the user AR device that is creating the tracking map is ranked. Thus, the tracking map can be used to rank the environment maps. In embodiments where a tracking map is unavailable, some or all portions of the selection and ranking of the environment maps that do not explicitly rely on the tracking map can be used.

[0379] Method 900 may begin at act 902, where a set of maps from a database of environment maps (which may be formatted as canonical maps) near a location where a tracking map is being formed may be accessed and then filtered for ranking. Furthermore, at act 902, at least one area attribute of the area in which the user AR device is operating is determined. In a scenario where the user AR device is constructing a tracking map, the area attribute may correspond to the area in which the tracking map is being created. As a specific example, the area attribute may be calculated based on a received signal from an access point to a computer network when the AR device calculates the tracking map.

[0380] Figure 30An exemplary map ranking unit 806 of the AR system 800 is shown in accordance with some embodiments. The map ranking unit 806 can be executed in a cloud computing environment, as it can include portions executed on the AR device and portions executed on a remote computing system such as the cloud. The map ranking unit 806 can be configured to perform at least a portion of the method 900.

[0381] Figure 31A An example of area attributes AA1-AA8 of a tracking map (TM) 1102 and environment maps CM1-CM4 in a database according to some embodiments is shown. As shown, the environment map can be associated with multiple area attributes. The area attributes AA1-AA8 can include parameters of a wireless network detected by the AR device that calculates the tracking map 1102, such as the basic service set identifier (BSSID) of the network to which the AR device is connected and / or the strength of a received signal from an access point to the wireless network via, for example, a network tower 1104. The parameters of the wireless network can conform to protocols including Wi-Fi and 5G NR. Figure 32 In the example shown, the region attribute is a fingerprint of the area where the user's AR device collects sensor data to form a tracking map.

[0382] Figure 31B An example of a determined geographic location 1106 for a tracking map 1102 is shown in accordance with some embodiments. In the example shown, the determined geographic location 1106 includes a centroid point 1110 and an area 1108 formed as a circle around the centroid point. It should be understood that the determination of the geographic location in the present application is not limited to the format shown. The determined geographic location can have any suitable format, including, for example, different area shapes. In this example, the geographic location is determined from the area attributes using a database that associates area attributes with geographic locations. Commercially available databases are available, and for example, a database that associates Wi-Fi fingerprints with locations represented as latitude and longitude can be used to perform this operation.

[0383] exist Figure 29 In an embodiment, the map database containing environment maps may also include location data for these maps, including the latitude and longitude covered by the maps. The processing at action 902 entails selecting from the database an environment map set that covers the same latitude and longitude determined for the area attributes of the tracking map.

[0384] Action 904 is a first filtering of the set of environmental maps accessed in action 902. In action 902, environmental maps were retained in the set based on proximity to the geographic location of the tracking map. This filtering step can be performed by comparing the latitude and longitude associated with the tracking map with the latitude and longitude associated with the environmental maps in the set.

[0385] Figure 32 An example of action 904 according to some embodiments is shown. Each area attribute may have a corresponding geographic location 1202. The set of environment maps may include environment maps having at least one area attribute whose geographic location overlaps with the determined geographic location of the tracking map. In the example shown, the identified set of environment maps includes environment maps CM1, CM2, and CM4, each of which has at least one area attribute whose geographic location overlaps with the determined geographic location of the tracking map 1102. Environment map CM3, which is associated with area attribute AA6, is not included in the set because it is located outside the determined geographic location of the tracking map.

[0386] Other filtering steps may also be performed on the set of environment maps to reduce / rank the number of environment maps ultimately processed in the set (such as for map merging or providing connected world information to a user device). Method 900 may include filtering (act 906) the set of environment maps based on the similarity between one or more identifiers of network access points associated with tracking maps and environment maps in the set of environment maps. During map formation, devices collecting sensor data to generate maps may be connected to a network via a network access point, such as via Wi-Fi or a similar wireless communication protocol. Access points may be identified by a BSSID. As a user device moves through the area where data is collected to form a map, the user device may connect to multiple different access points. Similarly, when multiple devices provide information to form a map, these devices may have connected via different access points, and therefore multiple access points may be used to form the map. Therefore, there may be multiple access points associated with a map, and the set of access points may be an indication of the location of the map. The strength of the signals from the access points (which may be reflected as RSSI values) may provide further geographic information. In some embodiments, a list of BSSIDs and RSSI values ​​may form a regional attribute of the map.

[0387] In some embodiments, filtering the set of environment maps based on the similarity of one or more identifiers of the network access points may include retaining, in the set of environment maps, environment maps having the highest Jaccard similarity with respect to at least one area attribute of the tracking map based on the one or more identifiers of the network access points.

[0388] Figure 33An example of action 906 according to some embodiments is shown. In the example shown, the network identifier associated with area attribute AA7 can be determined as the identifier of tracking map 1102. Following action 906, the environment map set includes environment map CM2, whose area attribute has a high Jaccard similarity to AA7, and environment map CM4, which also includes area attribute AA7. Environment map CM1 is not included in the environment map set because it has the lowest Jaccard similarity to AA7.

[0389] The processing at actions 902-906 can be performed based on metadata associated with the map without actually accessing the content of the map stored in the map database. Other processing may involve accessing the content of the map. Action 908 indicates accessing the environment map retained in the subset after filtering based on the metadata. It should be understood that this action can be performed earlier or later in the process if subsequent operations can be performed on the accessed content.

[0390] Method 900 may include filtering (act 910) the environment atlas based on similarity of metrics representing the content of the tracking map and the environment maps in the environment atlas. The metrics representing the content of the tracking map and the environment maps may include vectors of values ​​calculated based on the content of the maps. For example, as described above, depth keyframe descriptors calculated for one or more keyframes used to form a map may provide metrics for comparing maps or portions of maps. The metrics may be calculated using the maps retrieved at act 908, or may be pre-calculated and stored as metadata associated with the maps. In some embodiments, filtering the environment atlas based on similarity of metrics representing the content of the tracking map and the environment maps in the environment atlas may include retaining the environment maps in the environment atlas having the smallest vector distance between a vector of a characteristic of the tracking map and a vector representing an environment map in the environment atlas.

[0391] Method 900 may include further filtering (act 912) the environment atlas based on a degree of match between a portion of the tracking map and multiple portions of the environment map in the environment atlas. The degree of match can be determined as part of the localization process. As a non-limiting example, localization can be performed by identifying key points in the tracking map and the environment map that are sufficiently similar (because they represent the same part of the physical world). In some embodiments, the key points can be features, feature descriptors, key frames / key rigs, persistent poses and / or PCFs. The set of key points in the tracking map can then be aligned to produce a best fit with the set of key points in the environment map. The mean squared distance between corresponding key points can be calculated and, if below a threshold for a particular area of ​​the tracking map, used as an indication that the tracking map and the environment map represent the same area of ​​the physical world.

[0392] In some embodiments, filtering the environment atlas based on a degree of match between a portion of the tracking map and portions of environment maps in the environment atlas may include calculating a volume of the physical world represented by the tracking map (which volume is also represented in environment maps in the environment atlas), and retaining in the environment atlas environment maps having a larger calculated volume than environment maps filtered from the set. Figure 34 An example of action 912 according to some embodiments is shown. In the example shown, the set of environment maps after action 912 includes environment map CM4, which has regions 1402 that match regions of tracking map 1102. Environment map CM1 is not included in the set because it has no regions that match regions of tracking map 1102.

[0393] In some embodiments, the environment atlas may be filtered in the order of act 906, act 910, and act 912. In some embodiments, the environment atlas may be filtered based on act 906, act 910, and act 912, which may be performed in order of lowest to highest processing required to perform the filtering. Method 900 may include loading (act 914) the environment atlas and data.

[0394] In the illustrated example, the user database stores a region identifier indicating the region in which the AR device is used. The region identifier may be a region attribute, which may include parameters of a wireless network detected by the AR device during use. The map database may store multiple environment maps constructed using data provided by the AR device and associated metadata. The associated metadata may include a region identifier derived from the region identifier of the AR device that provided the data required to construct the environment map. The AR device may send a message to the PW module indicating that a new tracking map has been created or is being created. The PW module may calculate a region identifier for the AR device and update the user database based on the received parameters and / or the calculated region identifier. The PW module may also determine the region identifier associated with the AR device requesting the environment map, identify multiple environment map sets from the map database based on the region identifier, filter the environment map sets, and send the filtered multiple environment map sets to the AR device. In some embodiments, the PW module may filter the multiple environment map sets based on one or more criteria, such as the geographic location of the tracking map, the similarity of one or more identifiers of network access points associated with the tracking map and the environment maps in the environment map set, the similarity of a metric representing the content of the tracking map and the environment maps in the environment map set, and the degree of match between a portion of the tracking map and a portion of the environment map in the environment map set.

[0395] Having thus described several aspects of some embodiments, it will be appreciated that various changes, modifications, and improvements will readily occur to those skilled in the art. As an example, the embodiments are described in conjunction with an augmented reality (AR) environment. It will be appreciated that some or all of the techniques described herein may be applied in an MR environment or more generally in other XR and VR environments.

[0396] As another example, embodiments are described in conjunction with devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), discrete applications, and / or any suitable combination of devices, networks, and discrete applications.

[0397] also, Figure 29 Examples of criteria that can be used to filter candidate maps to produce a highly ranked set of maps are provided. Other criteria can be used instead of or in addition to the criteria described. For example, if multiple candidate maps have similar values ​​for a metric used to filter out less desirable maps, the characteristics of the candidate maps can be used to determine which maps are retained as candidates or filtered out. For example, larger or more dense candidate maps can be prioritized over smaller candidate maps. In some embodiments, Figure 27-28 Can describe Figures 29-34 All or part of the systems and methods described in.

[0398] Figure 35 and Figure 36 is a diagram illustrating an XR system configured to rank and merge multiple maps of an environment in accordance with some embodiments. In some embodiments, a connected world (PW) may determine when to trigger ranking and / or merging of maps. In some embodiments, determining which map to use may be based at least in part on the above-described configuration of the map in accordance with some embodiments. Figure 21-25 Depth keyframes described.

[0399] Figure 37 is a block diagram illustrating a method 3700 for creating an environment map of the physical world according to some embodiments. The method 3700 may begin by positioning a tracking map captured by an XR device worn by a user (act 3702) to a canonical map set (e.g., a map of a physical world). Figure 28 methods and / or Figure 29 The method 900 may include selecting the canonical map. Action 3702 may include positioning the keyrig of the tracking map into the canonical map group. The positioning result of each keyrig may include the positioning pose of the keyrig and a set of 2D to 3D feature correspondences.

[0400] In some embodiments, method 3700 may include splitting the tracking map (act 3704) into connected components that can be robustly merged into maps by merging connected pieces. Each connected component may include a Keyrig within a predetermined distance. Method 3700 may include merging connected components greater than a predetermined threshold into one or more canonical maps (act 3706), and removing the merged connected components from the tracking map.

[0401] In some embodiments, method 3700 may include merging (act 3708) canonical maps in the group that are merged with the same connected components of the tracking map. In some embodiments, method 3700 may include promoting (act 3710) the remaining connected components of the tracking map that have not been merged with any canonical map to a canonical map. In some embodiments, method 3700 may include merging (act 3712) the persistent pose and / or PCF of the tracking map and the canonical map merged with at least one connected component of the tracking map. In some embodiments, method 3700 may include finalizing (act 3714) the canonical map, for example, by fusing map points and pruning redundant Keyrigs.

[0402] Figure 38A and Figure 38B 38. The environment map 3800 is shown as being created by updating the canonical map 700 according to some embodiments. The canonical map 700 may be a new tracking map created from the tracking map 700 ( Figure 7 ) is improved. Figure 7 As shown and described, the canonical map 700 may provide a floor plan 706 of a physical object reconstructed in the corresponding physical world, represented by a point 702. In some embodiments, the map point 702 may represent a feature of a physical object, which may include multiple features. A new tracking map of the physical world may be captured and uploaded to the cloud to be merged with the map 700. The new tracking map may include map points 3802 and keyrigs 3804, 3806. In the example shown, the keyrig 3804 represents a physical object that has been reconstructed, for example, by establishing a keyrig 704 (e.g., a physical object) with the map 700. Figure 38B ) is successfully located to the Keyrig of the canonical map. On the other hand, Keyrig 3806 indicates that it has not yet been located to the Keyrig of map 700. In some embodiments, Keyrig 3806 can be promoted to a separate canonical map.

[0403] Figures 39A to 39F is a diagram illustrating an example of a cloud-based persistent coordinate system that provides a shared experience for users in the same physical space. Figure 39A Shown is, for example, a canonical map 4814 from the cloud being Figures 20A-20CThe canonical map 4814 may have a canonical coordinate system 4806C. The canonical map 4814 may have a PCF 4810C with multiple associated PPs (e.g., 4818A, 4818B in 39C).

[0404] Figure 39B The XR devices are shown establishing a relationship between their respective world coordinate systems 4806A, 4806B and a canonical coordinate system 4806C. This can be accomplished, for example, by positioning the tracking map on the canonical map 4814 on the respective devices. For each device, positioning the tracking map on the canonical map can result in a transformation between its local world coordinate system and the coordinate system of the canonical map.

[0405] Figure 39C It is shown that due to positioning, a transformation (e.g., transformation 4816A, 4816B) between the local PCF (e.g., PCF 4810A, 4810B) on the respective device and the corresponding persistent gesture (e.g., PP 4818A, 4818B) on the canonical map can be calculated. With these transformations, each device can use its local PCF (which can be detected locally on the device by processing images detected by sensors on the device) to determine where on the local device to display virtual content attached to PP 4818A, 4818B or other persistent points on the canonical map. This approach can accurately position virtual content relative to each user and allow each user to have the same experience of virtual content in physical space.

[0406] Figure 39D A snapshot of the persistent pose from the canonical map to the local tracking map is shown. It can be seen that the local tracking maps are connected to each other via the persistent pose. Figure 39E It is shown that PCF 4810A on the device worn by user 4802A can be accessed in the device worn by user 4802B through PP 4818A. Figure 39F Tracking maps 4804A, 4804B and canonical map 4814 are shown as being merged. In some embodiments, some PCFs may be removed due to the merge. In the example shown, the merged map includes PCF 4810C from canonical map 4814, but does not include PCFs 4810A, 4810B from tracking maps 4804A, 4804B. PPs previously associated with PCFs 4810A, 4810B may be associated with PCF 4810C after the map merge.

[0407] Example

[0408] Figure 40 and Figure 41 Shown Figure 9An example of the first XR device 12.1 using tracking maps. Figure 40 According to some embodiments, Figure 9 The first XR device to generate the first 3D local tracking map (ground Figure 1 ) is a two-dimensional representation of . Figure 41 is a diagram showing that according to some embodiments Figure 1 Upload from the first XR device to Figure 9 Block diagram of the server.

[0409] Figure 40 The first XR device 12.1 is shown Figure 1 and virtual content (content 123 and content 456). Figure 1 Has an origin (origin 1). Figure 1 From the perspective of the first XR device 12.1, PCF a is located at the ground. Figure 1 , and has X, Y, and Z coordinates of (0, 0, 0). PCF b has X, Y, and Z coordinates of (-1, 0, 0). Content 123 is associated with PCF a. In this example, content 123 has an X, Y, and Z relationship with respect to PCF a of (1, 0, 0). Content 456 has a relationship with respect to PCF b. In this example, content 456 has an X, Y, and Z relationship with respect to PCF b of (1, 0, 0).

[0410] exist Figure 41 The first XR device 12.1 will be Figure 1 Uploaded to the server 20. In this example, the server does not store a canonical map of the same area of ​​the physical world represented by the tracking map, and the tracking map is stored as the initial canonical map. The server 20 now has a map based on the Figure 1 The first XR device 12.1 has a canonical map that is empty at this stage. For the purposes of discussion, and in some embodiments, the server 20 does not include a canonical map. Figure 1 The second XR device 12.2 does not store any maps.

[0411] The first XR device 12.1 also sends its Wi-Fi signature data to the server 20. The server 20 can use the Wi-Fi signature data to determine the approximate location of the first XR device 12.1 based on intelligence gathered from other devices that have connected to the server 20 or other servers in the past and the GPS locations that such other devices have recorded. The first XR device 12.1 can now end the first session (see 8) and can disconnect from the server 20.

[0412] Figure 42 is a diagram showing a method according to some embodiments Figure 16 Schematic diagram of an XR system showing that a second user 14 . 2 initiates a second session using a second XR device of the XR system after the first user 14 . 1 terminates the first session. Figure 43A is a block diagram illustrating the initiation of a second session by a second user 14.2. The first user 14.1 is shown with a dashed line because the first session of the first user 14.1 has ended. The second XR device 12.2 begins recording the object. The server 20 may use various systems with different granularities to determine that the second session of the second XR device 12.2 is in the same vicinity as the first session of the first XR device 12.1. For example, Wi-Fi signature data, Global Positioning System (GPS) location data, GPS data based on Wi-Fi signature data, or any other location-indicative data may be included in the first XR device 12.1 and the second XR device 12.2 to record their locations. Alternatively, the PCF identified by the second XR device 12.2 may be displayed in a manner consistent with the location. Figure 1 The similarity of PCF.

[0413] like Figure 43B As shown, the second XR device starts up and begins collecting data, such as images 1110 from one or more cameras 44, 46. Figure 14 As shown, in some embodiments, an XR device (e.g., a second XR device 12.2) may collect one or more images 1110 and perform image processing to extract one or more features / points of interest 1120. Each feature may be converted into a descriptor 1130. In some embodiments, the descriptors 1130 may be used to describe keyframes 1140, which may have a position and orientation with an associated image attached. One or more keyframes 1140 may correspond to a single persistent gesture 1150, which may be automatically generated after a threshold distance (e.g., 3 meters) from a previous persistent gesture 1150. One or more persistent gestures 1150 may correspond to a single PCF 1160, which may be automatically generated after a predetermined distance (e.g., every 5 meters). Over time, as the user continues to move around the user environment and the XR device continues to collect more data, such as while collecting images 1110, additional PCFs may be created (e.g., PCF 3 and PCFs 4 and 5). One or more applications 1180 may run on the XR device and provide virtual content 1170 to the XR device for presentation to the user. Virtual content can have an associated content coordinate system that can be placed relative to one or more PCFs. Figure 43B As shown, the second XR device 12 . 2 creates three PCFs. In some embodiments, the second XR device 12 . 2 may attempt to locate itself in one or more canonical maps stored on the server 20 .

[0414] In some embodiments, as Figure 43CAs shown, the second XR device 12.2 can download the canonical map 120 from the server 20. Figure 1 Includes PCFs a through d and origin 1. In some embodiments, the server 20 may have multiple canonical maps for different locations and may determine that the second XR device 12.2 is in the same vicinity as the first XR device 12.1 during the first session and send a canonical map of the vicinity to the second XR device 12.2.

[0415] Figure 44 The second XR device 12.2 begins to identify the PCF for generating the Figure 2 The second XR device 12.2 recognizes only a single PCF, PCF 1,2. The X, Y, and Z coordinates of PCF 1,2 of the second XR device 12.2 may be (1, 1, 1). Figure 2 has its own origin (origin 2), which may be based on the head pose of device 2 when the device was started for the current head pose session. In some embodiments, the second XR device 12.2 may immediately attempt to Figure 2 In some embodiments, the map Figure 2 It may not be possible to locate the standard map (map Figure 1 ) (i.e., localization may fail) because the system cannot identify any or sufficient overlap between the two maps. Localization can be performed by identifying parts of the physical world represented in the first map that are also represented in the second map, and computing the transformations between the first and second maps required to align these parts. In some embodiments, the system can localize based on a comparison of PCFs between the local map and the canonical map. In some embodiments, the system can localize based on a comparison of persistent poses between the local map and the canonical map. In some embodiments, the system can localize based on a comparison of keyframes between the local map and the canonical map.

[0416] Figure 45 The second XR device 12.2 identifies the Figure 2 The other PCFs (PCF 1,2, PCF 3, PCF 4,5) after Figure 2 The second XR device 12.2 tries again to Figure 2 Locate to the standard map. Figure 2 has been extended to overlap at least a portion of the canonical map, so the positioning attempt will succeed. Figure 2 The overlap between the CNN and canonical maps can be represented by PCFs, persistent poses, keyframes, or any other suitable intermediate or derived construct.

[0417] Furthermore, the second XR device 12.2 has associated content 123 and content 456 to the map. Figure 2 PCF 1,2 and PCF 3. Content 123 has X, Y and Z coordinates (1,0,0) relative to PCF 1,2. Similarly, content 456 is located at Figure 2 The X, Y, and Z coordinates relative to PCF 3 are (1,0,0).

[0418] Figure 46A and Figure 46B Shows the ground Figure 2 The successful localization to the canonical map. Localization can be based on matching features in one map to another. By appropriate transformations, here involving translation and rotation of one map relative to the other, the overlapping area / volume / portion of the map 1410 represents the ground Figure 1 and the common part of the normative map. Figure 2 PCF 3 and PCF4,5 were created before positioning, and the normative map was created before positioning. Figure 2 PCF a and PCF c were created previously, so different PCFs are created to represent the same volume in real space (e.g., in different maps).

[0419] like Figure 47 As shown, the second XR device 12.2 is extended Figure 2 To include PCF ad from the canonical map. Figure 2 In some embodiments, the XR system may perform an optimization step to remove duplicate PCFs from overlapping regions, such as PCF 1410, PCF 3, and PCF 4,5. Figure 2 After positioning, the placement of virtual content (such as content 456 and content 123) will be relative to the updated location. Figure 2 The most recent update of the PCF in the. Although the PCF attachment for the content has changed, and although the location Figure 2 The PCF is updated, but the virtual content appears at the same real-world location relative to the user.

[0420] like Figure 48 As shown, the second XR device 12.2 continues to expand Figure 2 , because, for example, when the user moves around in the real world, more PCFs (e.g., PCFe, f, g, and h) are recognized by the second XR device 12.2. It can also be noted that Figure 1 exist Figure 47 and Figure 48 Not expanded.

[0421] refer to Figure 49 , the second XR device 12.2 will be Figure 2Upload to server 20. Server 20 will Figure 2 With normative Figure 1 In some embodiments, the Figure 2 The data may be uploaded to the server 20 at the end of the session of the second XR device 12 . 2 .

[0422] The canonical map within the server 20 now includes PCFi, which was not included in the map on the first XR device 12.1. Figure 1 When a third XR device (not shown) uploads a map to the server 20 and the map includes PCFi, the canonical map on the server 20 has been extended to include PCFi.

[0423] exist Figure 50 In the process, the server 20 will Figure 2 Combined with the normative map to form a new normative map. The server 20 determines whether PCF a to d are normative maps and maps. Figure 2 The server extends the specification map to include Figure 2 The PCFe to h and PCF 1,2 to form a new canonical map. The canonical maps on the first XR device 12.1 and the second XR device 12.2 are based on the ground Figure 1 And outdated.

[0424] exist Figure 51 In some embodiments, this occurs when the first XR device 12.1 and the second XR device 12.2 attempt to locate themselves during a different session or a new session or a subsequent session. The first XR device 12.1 and the second device 12.2 continue to locate their respective local maps (respectively, the local maps) as described above. Figure 1 peacefully Figure 2 ) to locate the new canonical map.

[0425] like Figure 52 As shown, the head coordinate system 96 or "head pose" is Figure 2 In some embodiments, the origin of the map (origin 2) is based on the head pose of the second XR device 12.2 at the start of the session. Since the PCF is created during the session, the PCF is placed relative to the origin 2 of the world coordinate system. Figure 2 The PCF of is used as a persistent coordinate system relative to the canonical coordinate system, where the world coordinate system is the world coordinate system of the previous session (e.g., Figure 40 middle ground Figure 1 The origin of the earth1). These coordinate systems are used to Figure 2 Position the same transformation to the canonical map and associate it, as above combined Figure 46B discussed.

[0426] The transformation from the world coordinate system to the head coordinate system 96 has been previously referred to Figure 9 Discussed. Figure 52 The head coordinate system 96 shown has only two orthogonal axes, which are located relative to the ground. Figure 2 The specific coordinate position of the PCF is relative to the ground Figure 2 However, it should be understood that the head coordinate system 96 is relative to the ground. Figure 2 The PCF is in three-dimensional position and has three orthogonal axes in the three-dimensional space.

[0427] exist Figure 53 In the figure, the head coordinate system 96 is relative to the ground. Figure 2 The PCF of the head coordinate system 96 has moved because the second user 14.2 has moved his head. The user can move his head in six degrees of freedom (6dof). The head coordinate system 96 can therefore move in 6dof, i.e. from its original position in the Figure 52 Starting from the previous position in the Figure 2 The three orthogonal axes of the PCF move in three-dimensional space. Figure 9 The head coordinate system 96 is adjusted when the real object detection camera 44 and the inertial measurement unit 48 in the head unit 22 detect motion of the real object and the head unit 22, respectively. More information about head pose tracking is disclosed in U.S. patent application Ser. No. 16 / 221,065, entitled “Enhanced Pose Determination for Display Device,” which is incorporated herein by reference in its entirety.

[0428] Figure 54 It is shown that sounds can be associated with one or more PCFs. For example, a user can wear headphones or earphones with stereo sound. The location of the sound through the earphones can be simulated using conventional techniques. The location of the sound can be located at a fixed position so that when the user rotates their head to the left, the location of the sound rotates to the right so that the user perceives the sound as coming from the same location in the real world. In this example, the location of the sound is represented by sound 123 and sound 456. For the purposes of discussion, Figure 54 Analysis and Figure 48 Similarly, when the first user 14.1 and the second user 14.2 are in the same room at the same or different times, they perceive the sound 123 and the sound 456 as coming from the same location in the real world.

[0429] Figure 55 and Figure 56 A further implementation of the above technology is shown. Figure 8As mentioned above, the first user 14.1 has initiated a first session. Figure 55 As shown, the first user 14.1 has terminated the first session, as indicated by the dashed line. At the end of the first session, the first XR device 12.1 Figure 1 The first user 14.1 has now initiated a second session at a later time than the first session. The first XR device 12.1 does not download the address from the server 20. Figure 1 , because the ground Figure 1 Already stored on the first XR device 12.1. Figure 1 If lost, the first XR device 12.1 downloads the address from the server 20 Figure 1 The first XR device 12.1 then proceeds to construct Figure 2 PCF, positioned to the ground Figure 1 , and further develop the canonical map as described above. The first XR device 12.1 Figure 2 It is then used to associate the above-mentioned local content, head coordinate system, local sound, etc.

[0430] refer to Figure 57 and Figure 58 It is also possible that more than one user interacts with the server in the same session. In this example, a third user 14.3 with a third XR device 12.3 joins the first user 14.1 and the second user 14.2. Each XR device 12.1, 12.2 and 12.3 starts generating its own map, respectively. Figure 1 ,land Figure 2 peacefully Figure 3 As XR devices 12.1, 12.2, and 12.3 continue to develop Figure 1 ,land Figure 2 peacefully Figure 3 , and continuously upload these maps to the server 20. The server 20 merges Figure 1 ,land Figure 2 peacefully Figure 3 The canonical map is then sent from the server 20 to each of the XR devices 12 . 1 , 12 . 2 , and 12 . 3 .

[0431] Figure 59Aspects of a viewing method for restoring and / or resetting a head pose are shown in accordance with some embodiments. In the illustrated example, at act 1400, power is applied to the viewing device. At act 1410, in response to being powered, a new session is initiated. In some embodiments, the new session may include establishing a head pose. One or more capture devices on a head-mounted frame secured to the user's head capture a surface of an environment by first capturing an image of the environment and then determining the surface based on the image. In some embodiments, the surface data may be combined with data from a gravity sensor to establish a head pose. Other suitable methods of establishing a head pose may be used.

[0432] In action 1420, the processor of the viewing device enters a routine for tracking head pose. As the user moves their head to determine the orientation of the head mounted frame relative to the surface, the capture device continues to capture the surface of the environment.

[0433] In action 1430, the processor determines whether the head pose has been lost. The head pose may be lost due to "edge" conditions that result in insufficient feature acquisition (such as too many reflective surfaces, insufficient lighting, blank walls, being outdoors, etc.), or due to the dynamic situation of a group of people moving and forming part of the map. The routine at 1430 allows a certain amount of time (e.g., 10 seconds) so that there is enough time to determine whether the head pose has been lost. If the head pose has not been lost, the processor returns to 1420 and enters head pose tracking again.

[0434] If the head pose has been lost at action 1430, the processor enters a routine to recover the head pose at 1440. If the head pose is lost due to insufficient lighting, a message such as the following is displayed to the user via the display of the viewing device:

[0435] The system is detecting low light conditions. Please move to a brighter area.

[0436] The system will continue to monitor whether sufficient light is available and whether the head pose can be recovered. The system may alternatively determine that low texture on the surface is causing the head pose to be lost, in which case the following hint is provided to the user in the display as a suggestion to improve surface capture:

[0437] The system is unable to detect enough finely textured surfaces. Please move to an area with less rough, finer textures.

[0438] In action 1450, the processor enters a routine to determine whether head pose recovery has failed. If head pose recovery has not failed (i.e., head pose recovery has been successful), the processor returns to action 1420 by re-entering head pose tracking. If head pose recovery has failed, the processor returns to action 1410 to establish a new session. As part of the new session, all cached data is invalidated and the head pose is re-established. Any suitable method of head tracking may be used with Figure 59 The process described is used in conjunction with US Patent Application No. 16 / 221,065, which is incorporated herein by reference in its entirety, describing head tracking.

[0439] Remote positioning

[0440] Various embodiments may utilize remote resources to facilitate a persistent and consistent cross-reality experience between individuals and / or groups of users. The inventors have recognized and appreciated that the operational benefits of an XR device having a canonical map as described herein may be achieved without downloading a canonical atlas. Figure 30 An example implementation of downloading a canonical map to a device is shown. For example, the benefits of not downloading a map can be realized by sending feature and pose information to a remote service that maintains a canonical map set. According to some embodiments, a device seeking to use a canonical map to position virtual content at a location relative to the canonical map can receive one or more transformations between features and the canonical map from the remote service. These transformations can be used on a device that maintains information about the locations of these features in the physical world to position virtual content at a location relative to the canonical map, or otherwise identify a location in the physical world relative to the canonical map.

[0441] In some embodiments, spatial information is captured by the XR device and transmitted to a remote service, such as a cloud-based service, which uses the spatial information to position the XR device relative to a canonical map used by applications or other components of the XR system to specify the location of virtual content relative to the physical world. Once positioned, a transformation linking a tracking map maintained by the device to the canonical map can be transmitted to the device. The transformation can be used in conjunction with the tracking map to determine the location of rendered virtual content relative to the canonical map, or otherwise identify a location in the physical world relative to the canonical map.

[0442] The inventors have recognized that the data that needs to be exchanged between a device and a remote location service can be very small compared to the transmission of map data, which can occur when a device transmits a tracking map to a remote service and receives a canonical map set from the remote service to perform device-based positioning. In some embodiments, performing positioning functions on cloud resources requires only a small amount of information to be sent from the device to the remote service. For example, a complete tracking map need not be transmitted to the remote service to perform positioning. In some embodiments, feature and gesture information, such as described above in conjunction with persistent gesture storage, can be sent to a remote server. In embodiments where features are represented by descriptors, as described above, the uploaded information can be even smaller.

[0443] The result returned to the device from the positioning service can be one or more transformations that associate the uploaded features with multiple parts of the matching canonical map. These transformations can be used in conjunction with the XR system's tracking map to identify the location of virtual content or otherwise identify locations in the physical world. In embodiments where persistent spatial information such as the PCF described above is used to specify a location relative to the canonical map, the positioning service can download the transformation between the feature and one or more PCFs to the device after successful positioning.

[0444] Therefore, the network bandwidth consumed by communications between the XR device and the remote service to perform positioning can be low. The system can therefore support frequent positioning, allowing each device interacting with the system to quickly obtain information used to locate virtual content or perform other location-based functions. As the device moves through the physical environment, it may repeatedly request updated positioning information. In addition, the device may frequently obtain updates to positioning information, such as when the canonical map changes, such as by incorporating additional tracking maps to expand the map or improve its accuracy.

[0445] Additionally, uploading features and downloading transforms can enhance privacy in XR systems that share map information among multiple users by increasing the difficulty of obtaining maps through spoofing. For example, an unauthorized user can be prevented from obtaining a map from the system by sending a false request for a canonical map representing a part of the physical world that the user is not in. When an unauthorized user is not physically in the area of ​​the physical world for which they are requesting map information, it is unlikely that they will be able to access features in that area. In embodiments where the feature information is formatted as a feature description, the difficulty of spoofing feature information by requesting map information is further complicated. Additionally, when the system returns transforms that are intended to be applied to a tracking map of a device operating in the area for which location information is requested, the information returned by the system is of little or no use to an imposter.

[0446] According to some embodiments, the positioning service is implemented as a cloud-based microservice. In some examples, implementing a cloud-based positioning service can help conserve device computing resources and enable the calculations required for positioning to be performed with very low latency. These operations can be supported by virtually unlimited computing power or other computing resources available through the provision of additional cloud resources, thereby ensuring that the XR system can scale to support a large number of devices. In one example, many canonical maps can be maintained in memory for near-instant access, or alternatively stored in high-availability devices to reduce system latency.

[0447] Furthermore, performing location mapping on multiple devices in a cloud service can lead to process improvements. Location telemetry and statistics can provide information about which canonical maps are in active storage and / or high-availability storage. For example, statistics from multiple devices can be used to identify the most frequently accessed canonical maps.

[0448] Additional accuracy can also be achieved due to processing being performed in a cloud environment or other remote environment with significant processing resources relative to the remote device. For example, localization can be performed on a higher density canonical map in the cloud relative to processing performed on the local device. Maps can be stored in the cloud, for example, with more PCFs or a higher density of feature descriptors per PCF, thereby improving the accuracy of the match between the device's feature set and the canonical map.

[0449] Figure 61 6100 is a schematic diagram of an XR system. The user device that displays cross-reality content during a user session can take many forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As described above, these devices can be configured with software, such as applications or other components, and / or hardwired to generate local location information (e.g., a tracking map) that can be used to render virtual content on their respective displays.

[0450] The virtual content location information may be specified relative to the global location information, for example, the global location information may be formatted as a canonical map containing one or more PCFs. Figure 61 In the illustrated embodiment, the system 6100 is configured with cloud-based services that support the execution and display of virtual content on user devices.

[0451] In one example, positioning functionality is provided as a cloud-based service 6106, which can be a microservice. The cloud-based service 6106 can be implemented on any of a plurality of computing devices, from which computing resources can be allocated to one or more services executed in the cloud. These computing devices can be interconnected and accessible to devices such as the wearable XR device 6102 and the handheld device 6104. These connections can be provided via one or more networks.

[0452] In some embodiments, the cloud-based service 6106 is configured to accept descriptor information from respective user devices and "locate" the devices to one or more matching canonical maps. For example, the cloud-based positioning service matches the received descriptor information with the descriptor information of the corresponding canonical maps. Canonical maps can be created using the techniques described above, which create canonical maps by merging maps provided by one or more devices having image sensors or other sensors that obtain information about the physical world. However, canonical maps do not have to be created by the devices accessing them, as these maps can be created by map developers, for example, developers who publish maps by making them available to the positioning service 6106.

[0453] According to some embodiments, the cloud service handles canonical map identification and may include operations to filter the repository of canonical maps into a set of potential matches. Figure 29 Filtering is performed as shown, or as Figure 29 Filtering may be performed using any subset of the filtering criteria and other filtering criteria, in lieu of or in addition to the filtering criteria shown. In one embodiment, geographic data may be used to limit the search for matching canonical maps to maps representing an area proximate to the device requesting location. For example, regional attributes, such as Wi-Fi signal data, Wi-Fi fingerprint information, GPS data, and / or other device location information, may be used as a coarse filter on stored canonical maps, thereby limiting the analysis of descriptors to canonical maps that are known or likely to be proximate to the user's device. Similarly, a location history for each device may be maintained by a cloud service to prioritize searches of canonical maps that are proximate to the device's last location. In some examples, filtering may include the above with respect to Figure 31B 、 Figure 32 、 Figure 33 and Figure 34 Function of discussion.

[0454] Figure 62is an example process flow that may be executed by a device to locate the device using a cloud-based service utilizing one or more canonical maps, and receive transformation information specifying one or more transformations between a local coordinate system of the device and a canonical map coordinate system. Various embodiments and examples may describe one or more transformations as specifying a transformation from a first coordinate system to a second coordinate system. Other embodiments include a transformation from a second coordinate system to a first coordinate system. In yet another embodiment, the transformation implements a transformation from one coordinate system to another, with the resulting coordinate system depending only on the desired coordinate system output, such as the coordinate system of displayed content, output. In yet other embodiments, the coordinate system transformation allows a first coordinate system to be determined based on a second coordinate system, and a second coordinate system to be determined based on a first coordinate system.

[0455] According to some embodiments, information reflecting the transformation of each persistent gesture defined relative to the canonical map may be sent to the device.

[0456] According to some embodiments, process 6200 may begin with a new session at 6202. Initiating a new session on a device may initiate the capture of image information to build a tracking map for the device. Additionally, the device may send a message to register for location services with the server, prompting the server to create a session for the device.

[0457] In some embodiments, starting a new session on a device optionally includes sending adjustment data from the device to a location service. The location service returns one or more transformations calculated based on the features and associated gesture sets to the device. If feature gestures are adjusted based on device-specific information before the transformations are calculated and / or the transformations are adjusted based on device-specific information after the transformations are calculated, rather than performing these calculations on the device, the device-specific information may be sent to the location service so that the location service can apply the adjustments. As a specific example, sending device-specific adjustment information may include capturing calibration data for sensors and / or displays. For example, the calibration data may be used to adjust the position of feature points relative to a measured position. Alternatively or additionally, the calibration data may be used to adjust the position of virtual content rendered by the display so that it appears accurately positioned for the specific device. For example, this calibration data may be obtained from multiple images of the same scene captured using a sensor on the device. The positions of features detected in these images can be represented as functions of the sensor position, such that the multiple images generate a system of equations that can be solved for the sensor position. The calculated sensor position can be compared to the nominal position, and calibration data can be derived from any differences. In some embodiments, calibration data for the display can also be calculated using intrinsic information about the device's configuration.

[0458] In embodiments where calibration data is generated for a sensor and / or display, the calibration data may be applied at any point in the measurement or display process. In some embodiments, the calibration data may be sent to a positioning server, which may store the calibration data in a data structure established for each device that has registered with the positioning server and is therefore in session with the server. The positioning server may apply the calibration data to any transformations calculated as part of the positioning process for the device providing the calibration data. Thus, the computational burden of using the calibration data to improve the accuracy of the sensed and / or displayed information is borne by the calibration service, thereby providing a further mechanism to reduce the processing burden on the device.

[0459] Once the new session is established, process 6200 can continue capturing new frames of the device context at 6204. At 6206, each frame can be processed to generate a descriptor of the captured frame (e.g., including the DSF values ​​discussed above). These values ​​can be calculated using some or all of the techniques described above, including those described above with respect to Figure 14 、 Figure 22 and Figure 23 Techniques discussed. As discussed, descriptors can be computed as a mapping of feature points to descriptors, or in some embodiments, as a mapping of image patches surrounding feature points to descriptors. The descriptors can have values ​​that enable valid matching between newly acquired frames / images and stored maps. In addition, the number of features extracted from the images can be limited to a maximum number of feature points per image, such as 200 feature points per image. As described above, feature points can be selected to represent points of interest. Thus, actions 6204 and 6206 are performed as part of a device process for forming a tracking map or otherwise periodically collecting images of the physical world around the device, or can, but need not, be performed separately for localization.

[0460] The feature extraction at 6206 may include attaching gesture information to the extracted features at 6206. The gesture information may be a gesture in the device's local coordinate system. In some embodiments, the gesture may be relative to a reference point in a tracking map, such as a persistent gesture as described above. Alternatively or additionally, the gesture may be relative to the origin of the device's tracking map. Such an embodiment may enable the location service described herein to provide location services for a wide range of devices, even if they do not use persistent gestures. In any event, gesture information may be attached to each feature or each set of features such that the location service can use the gesture information to calculate a transformation that can be returned to the device when matching the feature with features in a stored map.

[0461] Process 6200 may continue to decision block 6207, where a decision is made as to whether to request a fix. One or more criteria may be applied to determine whether to request a fix. The criteria may include the passage of time, allowing the device to request a fix after a certain threshold amount of time. For example, if no fix is ​​attempted within the threshold amount of time, the process may continue from decision block 6207 to action 6208, which requests a fix from the cloud. The threshold amount of time may be between 10 and 30 seconds, such as 25 seconds. Alternatively or additionally, a fix may be triggered by the movement of the device. The device performing process 6200 may track its movement using an IMU and / or its tracking map, and initiate a fix upon detecting movement exceeding a threshold distance from the location where the device last requested a fix. For example, the threshold distance may be between 1 and 10 meters, such as between 3 and 5 meters. As another alternative, a fix may be triggered in response to an event, such as when the device creates a new persistent pose or when the device's current persistent pose changes, as described above.

[0462] In some embodiments, decision block 6207 can be implemented so that the threshold for triggering a position fix can be dynamically established. For example, in environments where the features are substantially consistent, resulting in a low confidence level in matching the extracted feature set with the features of the stored map, position fixes may be requested more frequently to increase the chances of at least one position fix attempt being successful. In this case, the threshold applied at decision block 6207 can be lowered. Similarly, in environments where there are relatively few features, the threshold applied at decision block 6207 can be lowered to increase the frequency of position fix attempts.

[0463] Regardless of how a location fix is ​​triggered, when triggered, process 6200 can proceed to act 6208, where the device sends a request to the location service, including data used by the location service to perform the location fix. In some embodiments, data from multiple image frames can be provided for the location fix attempt. For example, the location service will not consider a location fix successful unless features in multiple image frames produce consistent location fix results. In some embodiments, process 6200 can include saving feature descriptors and additional pose information to a buffer. For example, the buffer can be a circular buffer that stores feature sets extracted from the most recently captured frame. Thus, the location fix request can be sent with multiple feature sets accumulated in the buffer. In some configurations, the buffer size is implemented to accumulate multiple data sets, which are more likely to produce a successful location fix. In some embodiments, the buffer size can be set to accumulate features from, for example, two, three, four, five, six, seven, eight, nine, or ten frames. Optionally, the buffer size can have a baseline setting that increases in response to a location fix failure. In some examples, increasing the buffer size and the corresponding number of feature sets transmitted reduces the likelihood that subsequent location fixes will fail to return a result.

[0464] Regardless of how the buffer size is set, the device can transmit the contents of the buffer to the location service as part of a location request. Other information can be sent along with the feature points and additional gesture information. For example, in some embodiments, geographic information can be sent. Geographic information can include, for example, GPS coordinates or a wireless signature associated with a device tracking map or the current persistent gesture.

[0465] In response to the request sent at 6208, the cloud positioning service may analyze the feature descriptors to locate the device to a canonical map or other persistent map maintained by the service. For example, the descriptors are matched to a feature set in a map to which the device is located. The cloud-based positioning service may perform positioning as described above with respect to positioning relative to the device (e.g., positioning may rely on any of the functions discussed above, including map ranking, map filtering, position estimation, filtered map selection, Figure 44-Figure 4 6, and / or discussed with respect to positioning modules, PCF and / or PP identification and matching, etc.). However, rather than sending the identified canonical map to the device (e.g., in device positioning), the cloud-based positioning service can continue to generate transformations based on the relative orientation of the feature set sent from the device and the matching features of the canonical map. The positioning service can return these transformations to the device, which can receive the transformations at block 6210.

[0466] In some embodiments, the canonical maps maintained by the positioning service can employ PCFs, as described above. In such embodiments, feature points of the canonical map that match feature points sent from the device can have locations specified relative to one or more PCFs. Thus, the positioning service can identify one or more canonical maps and can calculate a transformation between the coordinate system represented in the gesture sent with the positioning request and the one or more PCFs. In some embodiments, identification of one or more canonical maps is assisted by filtering potential maps based on geographic data of the corresponding device. For example, once filtered into a candidate set (e.g., by GPS coordinates and other options), the candidate set of canonical maps can be analyzed in detail to determine the matching feature points or PCFs described above.

[0467] The data returned to the requesting device at action 6210 can be formatted as a persistent gesture transform table. The table can include one or more canonical map identifiers indicating the canonical map to which the device was located by the location service. However, it should be understood that the location information can be formatted in other ways, including as a transform list with associated PCF and / or canonical map identifiers.

[0468] Regardless of how the transforms are formatted, the device can use these transforms to calculate a position to render virtual content that has been specified by an application or other component of the XR system relative to any of the PCFs at action 6212. This information can alternatively or additionally be used on the device to perform any position-based operations where the position is specified based on the PCF.

[0469] In some scenarios, the location service may not be able to match the features sent from the device to any stored canonical map, or may not be able to match a sufficient number of feature sets communicated to the location service request to deem a successful location. In such a scenario, the location service may indicate to the device that the location failed, rather than returning the transformation to the device as described above in conjunction with action 6210. In such a scenario, process 6200 may branch to action 6230 at decision box 6209, at which time the device may take one or more actions to handle the failure. These actions include increasing the size of the buffer that holds the feature sets sent for location positioning. For example, if the location service does not consider the location to be successful unless three features are positively matched, the buffer size may be increased from 5 to 6, thereby increasing the chance that the three feature sets sent will match the canonical map maintained by the location service.

[0470] Alternatively or additionally, failure handling may include adjusting device operating parameters to trigger more frequent positioning attempts. For example, the threshold time and / or threshold distance between positioning attempts may be reduced. As another example, the number of feature points in each feature set may be increased. When a sufficient number of features in a feature set sent from the device match features of the map, it may be considered that there is a match between the feature set and the features stored within the canonical map. Increasing the number of features sent increases the chance of a match. As a specific example, the initial feature set size may be 50, which may be increased to 100, 150, and then 200 upon each successive positioning failure. Upon a successful match, the size of the feature set may be returned to its initial value.

[0471] Failure handling may also include obtaining positioning information from sources other than location services. According to some embodiments, the user device may be configured to cache canonical maps. Cached maps allow the device to access and display content in situations where the cloud is unavailable. For example, cached canonical maps allow for device-based positioning in the event of a communication failure or other unavailability.

[0472] According to various embodiments, Figure 62 A high-level process for a device to initiate cloud-based positioning is described. In other embodiments, one or more of the various steps shown can be combined, omitted, or invoke other processes to complete positioning and ultimately achieve visualization of virtual content in the view of the corresponding device.

[0473] Furthermore, it should be understood that while process 6200 illustrates the device determining whether to initiate positioning at decision block 6207, the trigger for initiating positioning may come from outside the device, including from a positioning service. For example, a positioning service may maintain information about each of the devices in its session. For example, this information may include an identifier of the canonical map to which each device was most recently positioned. The positioning service or other components of the XR system may update the canonical map, including using the above in conjunction with Figure 26 When a canonical map is updated, the location service can send a notification to each device that was recently located on that map. This notification can serve as a trigger for the device to request a location fix and / or can include updated transforms recalculated using the feature set most recently sent from the device.

[0474] Figure 63A 、 Figure 63B and Figure 63C is an example process flow illustrating operations and communications between a device and a cloud service. Boxes 6350, 6352, 6354, and 6456 illustrate an example architecture and separation between components involved in a cloud-based positioning process. For example, modules, components, and / or software configured to process perception on a user device are shown at 6350 (e.g., Figure 6A Device functionality for persistent world operations is shown at 6352 (e.g., including the above description of the persistent world module (e.g., Figure 6A In other embodiments, separation between 6350 and 6352 is not required and the communication shown can exist between processes executing on the device.

[0475] Similarly, block 6354 shows a block configured to process data related to connected world / connected world modeling (e.g., Figure 26 Block 6356 illustrates a cloud process configured to process functionality associated with locating a device to one or more maps of a repository of stored canonical maps based on information sent from the device.

[0476] In the illustrated embodiment, when a new session begins, process 6300 begins at 6302. Sensor calibration data is obtained at 6304. The calibration data obtained may depend on the device (e.g., multiple cameras, sensors, positioning devices, etc.), represented at 6350. Once sensor calibration for the device is obtained, the calibration may be cached at 6306. If device operation results in a change in frequency parameters (e.g., collection frequency, sampling frequency, matching frequency, etc.), the frequency parameters are reset to a baseline at 6308.

[0477] Once the new session functionality is complete (e.g., calibration, steps 6302-6306), process 6300 may continue with the capture of a new frame 6312. Features and their corresponding descriptors are extracted from the frame at 6314. In some examples, as described above, the descriptors may include DSFs. According to some embodiments, the descriptors may have spatial information attached to them to facilitate subsequent processing (e.g., transform generation). Gesture information generated on the device (e.g., information specified for locating features in the physical world relative to the device's tracking map as discussed above) may be attached to the extracted descriptors at 6316.

[0478] At 6318, the descriptor and pose information are added to the buffer. The capture of new frames and additions to the buffer shown in steps 6312-6318 are performed in a loop until the buffer size threshold is exceeded at 6319. In response to determining that the buffer size is met, a positioning request is sent from the device to the cloud at 6320. In accordance with some embodiments, the request may be processed by a connected world service instantiated in the cloud (e.g., 6354). In other embodiments, the functional operations for identifying candidate canonical maps may be separated from the operations for actual matching (e.g., shown as boxes 6354 and 6356). In one embodiment, a cloud service for map filtering and / or map ranking may be executed at 6354 and process the positioning request received from 6320. In accordance with some embodiments, the map ranking operation is configured to determine a set of candidate maps that may include the device location at 6322.

[0479] In one example, the map ranking functionality includes operations for identifying candidate canonical maps based on geographic attributes or other location data (e.g., observed or inferred location information). For example, the other location data may include Wi-Fi signatures or GPS information.

[0480] According to other embodiments, location data may be captured during a cross-reality session with a device and user. Process 6300 may include additional operations to populate the location for a given device and / or session (not shown). For example, the location data may be stored as a device region attribute value and an attribute value for selecting a candidate canonical map close to the device location.

[0481] Any one or more location options may be used to filter the plurality of canonical map sets to a canonical map set that may represent an area including the user device's location. In some embodiments, the canonical map may cover a relatively large area of ​​the physical world. The canonical map may be segmented into a plurality of regions such that map selection requires selection of a map region. For example, a map region may be on the order of tens of square meters. Thus, the filtered canonical map set may be a set of map regions.

[0482] According to some embodiments, a localization snapshot can be constructed based on candidate canonical maps, posture signatures, and sensor calibration data. For example, an array of candidate canonical maps, posture signatures, and sensor calibration information can be sent along with a request to determine a specific matching canonical map. Matching against the canonical map can be performed based on descriptors received from the device and stored PCF data associated with the canonical map.

[0483] In some embodiments, a feature set from the device is compared to a feature set stored as part of the canonical map. The comparison can be based on feature descriptors and pose. For example, a candidate feature set for the canonical map can be selected based on a plurality of features in the candidate set whose descriptors are sufficiently similar to the descriptors of the feature set from the device that the features can be the same features. For example, the candidate set can be features derived from an image frame used to form the canonical map.

[0484] In some embodiments, if the number of similar features exceeds a threshold, further processing can be performed on the candidate feature set. The further processing can determine the degree to which the pose feature set from the device aligns with features in the candidate feature set. The feature set from the canonical map can be posed, such as the features from the device.

[0485] In some embodiments, the features are formatted as a high-dimensional embedding (e.g., DSF, etc.) and can be compared using a nearest neighbor search. In one example, the system is configured (e.g., by executing process 6200 and / or 6300) to find the first two nearest neighbors using Euclidean distance, and a ratio test can be performed. If the nearest neighbor is closer than the next nearest neighbor, the system considers the nearest neighbor to be a match. For example, "closer" in this context can be determined based on the ratio of the Euclidean distance relative to the next nearest neighbor being greater than a threshold multiplied by the ratio of the Euclidean distance relative to the nearest neighbor. Once a feature from the device is deemed to be "matched" to a feature in the canonical map, the system can be configured to calculate a relative transformation using the posture of the matching feature. The transformation developed from the posture information can be used to indicate the transformation required to locate the device to the canonical map.

[0486] The number of correct data can be used as an indicator of the quality of the match. For example, in the case of DSF matching, the number of correct data reflects the number of features that matched between the received descriptor information and the stored map / canonical map. In other embodiments, the correct data can be determined in this embodiment by counting the number of "matched" features in each set.

[0487] Alternatively or additionally, indications of the quality of the match may be determined in other ways. In some embodiments, for example, when computing a transformation to position a map from a device containing multiple features to a canonical map based on the relative poses of the matching features, the transformation statistics computed for each of the multiple matching features may be used as an indication of quality. For example, a large variance may indicate a poor quality match. Alternatively or additionally, for the determined transformation, the system may compute the average error between features with matching descriptors. The average error of the transformation may be computed to reflect the degree of position mismatch. Mean squared error is a specific example of an error metric. Regardless of the specific error metric, if the error is below a threshold, it may be determined that the transformation is usable for the features received from the device, and the computed transformation is used to locate the device. Alternatively or additionally, the amount of correct data may also be used to determine whether there is a map that matches the device location information and / or descriptors received from the device.

[0488] As described above, in some embodiments, the device may send multiple feature sets for positioning. Positioning may be considered successful when at least a threshold number of feature sets match the feature sets from the canonical map, the error is below a threshold, and / or the number of correct data is above a threshold. For example, the threshold number may be three feature sets. However, it will be appreciated that the threshold for determining whether a sufficient number of feature sets have suitable values ​​may be determined empirically or in other suitable ways. Similarly, other thresholds or parameters of the matching process, such as the degree of similarity between feature descriptors considered to be matched, the number of correct data used to select a candidate feature set, and / or the size of the mismatch error, may be determined empirically or in other suitable ways.

[0489] Once a match is determined, a set of persistent map features associated with the matching canonical map or maps is identified. In embodiments where matching is based on map regions, the persistent map features may be map features within the matching region. The persistent map features may be persistent gestures or PCFs as described above. In the example of FIG63 , the persistent map features are persistent gestures.

[0490] Regardless of the format of the persistent map features, each persistent map feature can have a predetermined orientation relative to the canonical map to which it belongs. This relative orientation is applied to a transformation that is calculated to align the feature set from the device with the feature set from the canonical map to determine the transformation between the feature set from the device and the persistent map feature. Any adjustments, such as may be derived from calibration data, can then be applied to this calculated transformation. The resulting transformation can be a transformation between the device's local coordinate system and the persistent map feature. This calculation can be performed for each persistent map feature that matches the map area, and the results can be stored in a table, represented at 6326 as persistent_pose_table.

[0491] In one example, block 6326 returns the persistent gesture transformation table, the canonical map identifier, and the number of correct data. According to some embodiments, the canonical map ID is an identifier used to uniquely identify a canonical map and canonical map version (or map region, in embodiments where positioning is based on a map region).

[0492] In various embodiments, the calculated positioning data can be used to populate positioning statistics and telemetry maintained by the positioning service at 6328. This information can be stored for each device and updated for each positioning attempt; at the end of a device session, the information can be cleared. For example, which maps the device matched to can be used to improve map ranking operations. For example, maps covering the same area that the device previously matched to can be prioritized in the ranking. Similarly, maps covering adjacent areas can be given higher priority than more distant areas. Furthermore, adjacent maps can be prioritized based on the detected trajectory of the device over time, with map areas in the direction of motion being given higher priority than other map areas. The positioning service can use this information, for example, when a subsequent positioning request from the device restricts the search to maps or map areas within a stored canonical map to a candidate feature set. If a match with a low error metric and / or a large amount or percentage of correct data is identified within that limited area, processing of maps outside that area can be avoided.

[0493] Process 6300 continues by transmitting the information cloud (e.g., 6354) to the user device (e.g., 6352). According to some embodiments, at 6330, the persistent gesture table and the canonical map identifier are transmitted to the user device. In one example, the persistent gesture table can be composed of multiple elements, including at least one string identifying the persistent gesture ID and a transformation linking the device's tracking map to the persistent gesture. In embodiments where the persistent map feature is a PCF, the table can instead indicate a transformation to the PCF of the matching map.

[0494] If positioning fails at 6336, process 6300 continues by adjusting parameters that can increase the amount of data sent from the device to the positioning service, thereby increasing the chance of successful positioning. For example, failure can be indicated when no feature set with more than a threshold number of similar descriptors can be found in the canonical map, or when the error metric associated with all transformed candidate feature sets is above a threshold. As an example of an adjustable parameter, the size constraint of the descriptor buffer can be increased (6319). For example, with a descriptor buffer size of 5, a positioning failure would trigger an increase to at least six feature sets extracted from at least six image frames. In some embodiments, process 6300 can include a descriptor buffer increment value. In one example, the increment value can be used to control the rate at which the buffer size is increased in response to a positioning failure, for example. Other parameters, such as parameters that control the rate of positioning requests, can be changed when a matching canonical map cannot be found.

[0495] In some embodiments, execution 6300 may generate an error condition at 6340, which includes execution in the event that a positioning request fails to work rather than returning an unmatched result. For example, an error may occur when a network error causes the memory holding the canonical map database to be unavailable to the server performing the positioning service, or when a request for the positioning service is received that contains incorrectly formatted information. In this example, in the event of an error condition, process 6300 schedules a retry of the request at 6342.

[0496] When the positioning request succeeds, any parameters that were adjusted in response to the failure can be reset. At 6332, process 6300 can continue to operate to reset the frequency parameters to any default values ​​or baselines. In some embodiments, 6332 is performed regardless of any changes, thereby ensuring that a baseline frequency is always established.

[0497] At 6334, the device may use the received information to update the cached positioning snapshot. According to various embodiments, the corresponding transform, canonical map identifier, and other positioning data may be stored by the device and used to associate the position specified relative to the canonical map or its persistent map features (such as a persistent pose or PCF) with the position determined by the device relative to its local coordinate system (such as may be determined from its tracking map).

[0498] Various embodiments of the process for positioning in the cloud can implement any one or more of the aforementioned steps and be based on the aforementioned architecture. Other embodiments can combine one or more of the aforementioned steps, performing the steps simultaneously, in parallel, or in another order.

[0499] According to some embodiments, the location service in the cloud in the context of a cross-reality experience may include additional functionality. For example, canonical map caching may be performed to address connectivity issues. In some embodiments, the device may periodically download and cache the canonical map to which it is located. If the location service in the cloud is unavailable, the device may perform its own location (e.g., as described above—including information about the location of the device). Figure 26 In other embodiments, the transformations returned from a location request can be chained together and applied to subsequent sessions. For example, a device can cache a series of transformations and use that sequence of transformations to establish a location.

[0500] Various embodiments of the system may use the results of the positioning operation to update the transformation information. For example, the positioning service and / or the device may be configured to maintain state information about the transformation from the tracking map to the canonical map. The transformations received over a period of time may be averaged. According to some embodiments, the averaging operation may be limited to occur after a threshold number of positioning successes (e.g., three, four, five, or more times). In other embodiments, other state information may be tracked in the cloud, such as by a connected world module. In one example, the state information may include a device identifier, a tracking map ID, a canonical map reference (e.g., version and ID), and a transformation from the canonical map to the tracking map. In some examples, the system may use the state information to continuously update and obtain a more accurate canonical map to tracking map transformation each time it performs a cloud-based positioning function.

[0501] Additional enhancements to cloud-based positioning may include communicating erroneous data (outliers) in a feature set that do not match features in a canonical map to the device. For example, the device may use this information to improve its tracking map, such as by removing erroneous data from the feature set used to construct its tracking map. Alternatively or additionally, information from the positioning service may enable the device to limit bundle adjustments for its tracking map to adjustments calculated based on correct data features or otherwise impose constraints on the bundle adjustment process. According to another embodiment, various sub-processes or additional operations may be used in conjunction with and / or as an alternative to the processes and / or steps discussed for cloud-based positioning. For example, candidate map identification may include identifying a candidate map based on a corresponding Figure 1 The canonical map is accessed using the region identifiers and / or region attributes stored therein.

[0502] Sharing content based on location

[0503] In some embodiments, the XR system can enable virtual content to be associated with specific locations and shared with users who visit those locations. This location-based virtual content can be added to the system in any variety of ways. For example, the XR system can include an authentication service that allows or blocks specific users from accessing functionality that associates virtual content with specific locations in the physical world. This association can be achieved by editing a stored map, such as the canonical map described above.

[0504] Users authorized to add virtual content to canonical maps can ...

Claims

1. A network resource within a distributed computing environment for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a 3D environment, the network resource comprising: one or more processors; At least one computer-readable medium comprising: a plurality of stored maps of the 3D environment, the maps comprising a plurality of coordinate systems, each of the plurality of coordinate systems being determined based on one or more feature points in the 3D environment; a plurality of data structures, each data structure in the plurality of data structures being associated with a corresponding volumetric region in the 3D environment where virtual content is to be displayed, wherein each data structure in the plurality of data structures comprises: information associating the data structure with a coordinate system of the plurality of coordinate systems in the plurality of stored maps; and a link to virtual content for rendering in the corresponding area in the 3D environment; and Computer-executable instructions that, when executed by at least one of the one or more processors: implementing a service for providing positioning information to portable electronic devices of the plurality of portable electronic devices, wherein the positioning information indicates positions of the plurality of portable electronic devices relative to one or more shared maps of the plurality of stored maps of the 3D environment; and A copy of a data structure in the plurality of data structures is provided to the portable electronic device in the plurality of portable electronic devices when the location of the portable electronic device is within a threshold distance of a coordinate system associated with the data structure.

2. The network resource according to claim 1, wherein: The computer-executable instructions, when executed by the at least one processor, further implement an authentication service that determines access rights to the portable electronic device; as well as The computer-executable instructions for selectively providing the data structure to the portable electronic device determine whether to send the data structure based in part on the access rights of the portable electronic device and access attributes associated with the data structure.

3. The network resource according to claim 1, wherein: Each data structure of the plurality of data structures further includes common attributes; and The computer-executable instructions for selectively providing the data structure to the portable electronic device determine whether to transmit the at least one data structure based in part on a common attribute of the at least one data structure.

4. The network resource according to claim 1, wherein: For a portion of the plurality of data structures, the link to the virtual content includes a link to an application that provides the virtual content.

5. The network resource of claim 1 , wherein: Each data structure of the plurality of data structures further includes display characteristics for rendering the volumetric region of the virtual content linked to the data structure.

6. The network resource according to claim 5, wherein: The display characteristics include behavior of virtual content rendered within the volumetric region relative to a physical surface.

7. The network resource according to claim 5, wherein: The display characteristics include one or more of: a size of the volumetric area, an offset of the volumetric area relative to a permanent coordinate system associated with a map, a spatial orientation of the volumetric area, a behavior of virtual content rendered within the volumetric area relative to the position of the portable electronic device, and a behavior of virtual content rendered within the volumetric area relative to a direction facing the portable electronic device.

8. The network resource of claim 1 , wherein: The computer executable instructions, when executed by the at least one processor, further detecting whether the portable electronic device has moved outside the threshold distance of the coordinate system associated with the data structure; as well as Based on the detecting, the copy of the data structure is deleted.

9. A portable electronic device configured to render virtual content in a 3D environment, the portable electronic device comprising: one or more processors; at least one computer-readable medium containing computer-executable instructions that, when executed by at least one of the one or more processors: generating information indicating a position of the portable electronic device in a local coordinate system within the 3D environment; sending the information indicating the position in the local coordinate system to a positioning service over a network; Obtaining from the positioning service a transformation between a coordinate system of stored spatial information about the 3D environment and the local coordinate system; obtaining a data structure from the location service, the data structure representing a volumetric region within the 3D environment, virtual content for display in the volumetric region, and an application for displaying the virtual content, the data structure further comprising one or more dimensions of the volumetric region, an offset of the volumetric region from a persistent location associated with the stored spatial information, and behavior of the virtual content relative to a location of the portable electronic device; as well as The virtual content represented in the data structure is rendered within the volumetric region of the data structure by executing the application of the data structure.

10. The portable electronic device according to claim 9, wherein: The computer-executable instructions also include computer-executable instructions for: A determination is made as to whether the application of the data structure is installed on the portable electronic device. The portable electronic device according to claim 10 , wherein: The computer-executable instructions also include computer-executable instructions for: When it is determined that the application of the data structure is not installed on the portable electronic device, the application is obtained from a remote server.

12. The portable electronic device according to claim 9, wherein: The virtual content is represented in the data structure as an indicator of the location of the virtual content on the network.

13. The portable electronic device according to claim 9, wherein: The computer-executable instructions also include computer-executable instructions for: detecting that the portable electronic device has left an area represented by the data structure; and Based on the detecting, the virtual content represented in the data structure is deleted.

14. The portable electronic device according to claim 9, wherein: Rendering the virtual content in the volumetric region includes creating a prism, parameters of the prism being set based on the data structure.

15. A network resource within a distributed computing environment for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a 3D environment, the network resource comprising: one or more processors; At least one computer-readable medium comprising: a plurality of stored maps of the 3D environment; Multiple data structures, among which: Each of the plurality of data structures comprises: information associating the data structure with locations in the plurality of stored maps, a link to virtual content for rendering in a corresponding area in the 3D environment, and display characteristics of the corresponding areas on the plurality of portable electronic devices; and the display characteristics comprising one or more of: a size of the corresponding area, an offset of the corresponding area from the location in the plurality of stored maps, a spatial orientation of the corresponding area, and a behavior of virtual content rendered in the corresponding area; and Computer-executable instructions that, when executed by at least one of the one or more processors: A copy of a data structure in the plurality of data structures is provided to a portable electronic device in the plurality of portable electronic devices based on a coordinate system of the portable electronic device relative to the plurality of data structures.

16. The network resource of claim 15, wherein: The computer executable instructions, when executed by the at least one processor, authenticate a service, the authentication service determining access rights to the portable electronic device; as well as The computer-executable instructions that provide the copy of the data structure to the portable electronic device determine whether to send the copy of the data structure based in part on the access rights of the portable electronic device and access attributes associated with the copy of the data structure.

17. The network resource of claim 15, wherein: Each of the plurality of data structures further includes common attributes; and The computer-executable instructions that provide the copy of the data structure to the portable electronic device determine whether to send the copy of the data structure based in part on the public attribute of the copy of the data structure.

18. The network resource of claim 15, wherein: For a portion of the plurality of data structures, the link to the virtual content includes a link to an application that provides the virtual content.

19. The network resource of claim 15, wherein: The corresponding area is a volume in which the virtual content linked to the data structure is displayed.

20. The network resource of claim 15, wherein: The behavior of the virtual content includes the behavior of the virtual content relative to a physical surface.

21. The network resource of claim 15, wherein: The behavior of the virtual content rendered in the corresponding area includes a behavior of the virtual content relative to the coordinate system of the portable electronic device and a behavior of the virtual content relative to a direction facing the portable electronic device.

22. A method of operating a portable electronic device to render virtual content in a 3D environment, the method comprising, utilizing one or more processors: sending information indicating the location in the 3D environment to a service over a network; Retrieve from the service one or more data structures, each data structure representing a corresponding region in the 3D environment and virtual content for display in the corresponding region; rendering the virtual content represented in the one or more data structures in the corresponding areas of the one or more data structures; detecting that the portable electronic device has left an area represented by a data structure of the one or more data structures; as well as Based on the detecting, the virtual content represented in the data structure is deleted.

23. The method of claim 22, wherein: Rendering the virtual content in the corresponding area includes creating the corresponding area with a parameter set based on the data structure representing the corresponding area.

24. The method according to claim 22, wherein The virtual content is represented in at least one of the one or more data structures as an indicator of a location of the virtual content on a network.

25. The method of claim 22, wherein: Rendering the virtual content includes executing an application that generates the virtual content on the portable electronic device.

26. The method according to claim 25, wherein Rendering the virtual content further includes: determining whether the application is currently installed on the portable electronic device; and Based on determining that the application is not currently installed, the application is downloaded to the portable electronic device.

27. The method of claim 22, wherein: Each of the one or more data structures contains display characteristics of the corresponding area on the portable electronic device; and The display characteristics include one or more of: a size of the corresponding area, an offset of the corresponding area relative to the position in the 3D environment, a spatial orientation of the corresponding area, and a behavior of virtual content rendered in the corresponding area.

28. The method of claim 22, wherein: The one or more data structures include a first set of data structures; The first set of data structures is received at a first time; The method further comprises: storing rendering information associated with a first data structure in the first set of data structures; receiving a second set of data structures at a second time after the first time; and Based on determining that the first data structure is not included in the second set of data structures, the rendering information associated with the first data structure is deleted.

29. An electronic device configured to operate in a cross-reality system, the electronic device comprising: one or more sensors configured to capture information about a three-dimensional (3D) environment, the captured information comprising a plurality of images; at least one processor; as well as at least one computer-readable medium storing computer-executable instructions that, when executed on a processor of the at least one processor: Maintain local coordinate system; managing a corresponding area associated with one or more applications so that the virtual content generated by an application among the one or more applications is rendered within the corresponding area; On first reception from the service: a first data structure representing corresponding virtual content and a corresponding region of the 3D environment for rendering the virtual content; Associating the corresponding area with the first data structure so that the corresponding virtual content is rendered within the corresponding area; receiving from the service at a second time after the first time: a second set of data structures; and Based on determining that the first data structure is not included in the second set of data structures, the rendering information associated with the first data structure is deleted.

30. The electronic device according to claim 29, wherein The computer-executable instructions also include computer-executable instructions for: Based on the information in the first data structure, obtaining the corresponding virtual content; and The acquired virtual content is rendered in the corresponding area.

31. The electronic device according to claim 30, wherein Retrieving the corresponding virtual content based on the information in the first data structure includes accessing the corresponding virtual content through a network based on an indicator of a location of the virtual content in the first data structure.

32. The electronic device according to claim 31, wherein Based on the information in the first data structure, obtaining the corresponding virtual content includes: based on the indicator of the location of the virtual content in the first data structure, downloading an application that generates the corresponding virtual content through a network.

33. The electronic device according to claim 29, wherein The computer-executable instructions also include instructions for: detecting that the electronic device has left the area represented by the first data structure; and Based on the detecting, the corresponding region associated with the first data structure is deleted.

34. The electronic device according to claim 30, wherein Rendering the acquired virtual content within the corresponding area further includes: using a coordinate system of the electronic device to determine a set of coordinates for rendering the virtual content in the 3D environment.

35. A network resource within a distributed computing environment for providing location-based content to a plurality of portable electronic devices capable of rendering virtual content in a 3D environment, the network resource comprising: One or more processors: At least one computer-readable medium comprising: a stored map of the 3D environment; and The computer executable instructions, when executed by at least one of the one or more processors, perform the following operations: implementing an authentication service that determines whether a user has permission to access a location in the 3D environment; receiving, from an authorized user, an indication and data of an area in the 3D environment where the virtual content is to be displayed, wherein: The data includes data representing the virtual content and display properties for rendering the virtual content in the area of ​​the 3D environment by a device of the plurality of portable electronic devices; and The display properties include one or more of: a size of the area, an offset of the area from a location in the stored map, a spatial orientation of the area, and a behavior of the virtual content rendered in the area; and A data structure is stored in association with the location in the stored map, the data structure including an indication of the display attributes and the virtual content.

36. The network resource of claim 35, wherein: The computer-executable instructions include computer-executable instructions that, when executed by at least one of the one or more processors, A copy of the data structure is provided to a portable electronic device of the plurality of portable electronic devices based on a coordinate system of the portable electronic device relative to the area represented by the data structure.

37. The network resource of claim 35, wherein: The computer-executable instructions include computer-executable instructions that, when executed by at least one of the one or more processors, Users are assigned different access permissions to one or more portions of the stored map of the 3D environment.

38. The network resource of claim 35, wherein: The computer-executable instructions include computer-executable instructions that, when executed by at least one of the one or more processors, Based on authorized user input, data representing the virtual content and / or display properties of the area on the plurality of portable electronic devices are updated.

39. The network resource of claim 36, wherein: The authentication service further determines access rights to the portable electronic device; as well as The computer-executable instructions that provide the copy of the data structure to the portable electronic device determine whether to send the copy of the data structure based in part on the access rights of the portable electronic device and access attributes associated with the copy of the data structure.

40. The network resource of claim 36, wherein: The data structure further includes public attributes; and Computer-executable instructions that provide the data structure to the portable electronic device determine whether to send the data structure based in part on the common attributes of the data structure.

41. The network resource of claim 35, wherein: The data structure includes a link to an application that provides the virtual content.

42. The network resource of claim 35, wherein: The area is a volume in which the virtual content linked to the data structure is displayed.

43. The network resource of claim 35, wherein: The behavior of the virtual content rendered in the area includes behavior of the virtual content relative to a physical surface.

44. The network resource of claim 36, wherein: The behavior of the virtual content rendered within the area includes the behavior of the virtual content relative to a coordinate system of the portable electronic device and the behavior of the virtual content relative to a direction the portable electronic device is facing.

45. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: generating a data structure representing an area in the 3D environment where virtual content is to be displayed; storing information indicative of virtual content to be rendered in the area of ​​the 3D environment in the data structure; as well as The data structure is associated with a location in a map for positioning a plurality of portable electronic devices capable of rendering virtual content in the 3D environment in a shared coordinate system of a cross-reality system operable with the plurality of portable electronic devices.

46. ​​The non-transitory computer-readable medium of claim 45, wherein: Storing information in the data structure indicating virtual content to be rendered in the area in the 3D environment includes specifying a link to the virtual content.

47. The non-transitory computer-readable medium of claim 46, wherein: The link to the virtual content includes a link to an application that provides the virtual content.

Citation Information

Patent Citations

  • Localization determination for mixed reality systems

    US10812936B2

  • System and method for augmented and virtual reality

    US20140306866A1

  • Centralized rendering

    US20180286116A1

  • Matching content to a spatial 3D environment

    US20180315248A1

  • Fully convolutional interest point detection and description via homographic adaptation

    US20190147341A1