Cross reality system with positioning services

By introducing cloud-hosted positioning services into the cross-reality system, and using sensor data to construct a local coordinate framework and match it with a cloud map, the problem of high resource consumption in multi-user and multi-device scenarios is solved, and efficient virtual content sharing and immersive experience are achieved.

CN114600064BActive Publication Date: 2026-04-24MAGIC LEAP INC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MAGIC LEAP INC
Filing Date
2020-10-15
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing cross-reality systems face problems of high resource consumption and inconsistent user experience when building and maintaining representations of the user's physical environment, especially in multi-user and multi-device scenarios, where devices rely on their own sensor scanning, resulting in battery consumption and wasted computing resources.

Method used

By introducing cloud-hosted positioning services into the cross-reality system, a local coordinate framework is constructed using sensor data and matched with a map database stored in the cloud, reducing the computing and network burden on devices and enabling efficient positioning and virtual content sharing between devices.

Benefits of technology

It enhances the immersive experience of the user experience and improves device battery life, while reducing network bandwidth and computing resource consumption, enabling virtual content sharing among multiple users and devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114600064B_ABST
    Figure CN114600064B_ABST
Patent Text Reader

Abstract

A cross reality system enables any of a plurality of devices to efficiently and accurately access previously stored maps and render virtual content specified in association with those maps. The cross reality system can include a cloud-based localization service that responds to requests from devices to localize with respect to stored maps. The request can include one or more sets of feature descriptors extracted from images of the physical world surrounding the device. The features can be posed with respect to a coordinate frame used by the local device. The localization service can utilize the matching sets of features to identify one or more stored maps. Based on a transformation required to align the features from the device with the matching sets of features, the localization service can compute a transformation that relates its local coordinate frame to the coordinate frame of the stored map and return the transformation to the device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 915,599, filed October 15, 2019, entitled “CROSSREALITY SYSTEM WITH WIRELESS FINGERPRINTS”, pursuant to Section 119(e) of 35 USC, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application generally relates to cross-reality systems. Background Technology

[0004] Computers can control human user interfaces to create cross-reality (XR) environments, in which some or all of the XR environment is generated by the computer as perceived by the user. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments, some or all of which can be generated by the computer using data describing the environment. For example, this data can describe virtual objects that can be rendered in a way that the user feels or perceives as part of the physical world, and can be interacted with. Because the data is rendered and presented through user interface devices such as, for example, head-mounted displays, the user can experience these virtual objects. The data can be displayed to the user, or it can control audio played to the user, or it can control a haptic (or tactile) interface, allowing the user to experience the tactile sensations of virtual objects.

[0005] XR systems can be used in a wide range of applications across scientific visualization, medical training, engineering design and prototyping, remote manipulation and telepresence, and personal entertainment. Compared to VR, AR and MR involve one or more virtual objects that are related to real-world objects. The experience of interacting with real-world objects significantly enhances the user experience of XR systems and opens the door to a variety of applications that present realistic and easily understandable information about how the physical world can be altered.

[0006] To realistically render virtual content, an XR system can establish a representation of the physical world surrounding the user. For example, this representation can be constructed by processing images acquired using sensors on a wearable device that forms part of the XR system. In such a system, the user can perform an initialization routine by looking around the room or other physical environment in which they intend to use the XR system until the system acquires enough information to construct a representation of that environment. As the system operates and the user moves within or to other environments, the sensors on the wearable device can acquire additional information to expand or update the representation of the physical world. Summary of the Invention

[0007] This application relates to methods and apparatus for providing X-Reality (cross-reality or XR) scenes. The techniques described herein can be used together, individually, or in any suitable combination.

[0008] According to one aspect, an electronic device configured to operate within a cross-reality system is provided. The electronic device includes: one or more sensors configured to capture information about a three-dimensional 3D environment, the captured information including multiple images of the 3D environment; and at least one processor configured to execute computer-executable instructions, wherein the computer-executable instructions include instructions for: generating a local coordinate frame representing a position in the 3D environment; extracting multiple features from the multiple images of the 3D environment; transmitting information about the multiple features and positional information of the multiple features expressed in the local coordinate frame to a positioning service via a network; and receiving from the positioning service at least one transformation associating the local coordinate frame with a second coordinate frame.

[0009] According to one embodiment, the electronic device includes a display, and the computer-executable instructions further include instructions for rendering virtual content having a position specified in the second coordinate frame at a position calculated at least in part based on a transformation of the at least one of the transformations on the display. According to one embodiment, the computer-executable instructions further include instructions for generating descriptors for the plurality of features, and sending information about the plurality of features includes sending the descriptors for the plurality of features.

[0010] According to one embodiment, the computer-executable instructions further include: instructions for storing the information about the plurality of features and the location information of the plurality of features in a buffer, and instructions for transmitting over the network including: transmitting the contents of the buffer together, such that the information about the plurality of features and the location information of the plurality of features are transmitted together. According to one embodiment, the buffer includes an adjustable size, and the computer-executable instructions further include: instructions for increasing the size of the buffer in response to a failure indication received from the location service via the network.

[0011] According to one embodiment, extracting the plurality of features from the plurality of images includes: extracting up to a threshold number of features from each image; and the computer-executable instructions further include: instructions for increasing the threshold number in response to a failure indication received from the location service via the network.

[0012] According to one embodiment, the computer-executable instructions further include: instructions for sending a request to initiate a session with the location service via the network. According to one embodiment, the request to initiate a session with the location service includes: an identifier of the electronic device and calibration data for the electronic device. According to one embodiment, the information about the plurality of features and the location information of the plurality of features sent via the network includes: a request for the location service; and the computer-executable instructions further include: instructions for sending a location request based on one or more trigger conditions being met. According to one embodiment, the one or more trigger conditions include: the distance the electronic device has moved since the last successful location request.

[0013] According to one aspect, a method for operating an electronic device configured to operate within a cross-reality system, the method comprising: receiving information about a three-dimensional 3D environment, the information including multiple images of the 3D environment; generating a local coordinate frame for representing a position in the 3D environment; extracting multiple features from the multiple images of the 3D environment; transmitting information about the multiple features and positional information of the multiple features expressed in the local coordinate frame to a positioning service via a network; and receiving from the positioning service at least one transformation associating the local coordinate frame with a second coordinate frame.

[0014] According to one embodiment, the method further includes rendering virtual content having a position specified in the second coordinate frame on a display of the electronic device at a position calculated at least in part based on a transformation in the at least one transformation.

[0015] According to one embodiment, the method further includes: generating descriptors for the plurality of features, and sending information about the plurality of features includes: sending the descriptors for the plurality of features.

[0016] According to one embodiment, the method further includes: storing the information about the plurality of features and the location information of the plurality of features in a buffer, and transmitting over a network including: transmitting the contents of the buffer together, such that the information about the plurality of features and the location information of the plurality of features are transmitted together.

[0017] According to one embodiment, the buffer includes an adjustable size, and the method further includes: increasing the size of the buffer in response to a failure indication received from the location service via the network.

[0018] According to one embodiment, extracting the plurality of features from the plurality of images includes: extracting up to a threshold number of features from each image, and the method further includes: increasing the threshold number in response to a failure indication received from the location service via the network.

[0019] According to one embodiment, the method further includes: sending a request to the location service via the network to initiate a session with the location service, wherein the request to initiate a session with the location service includes an identifier of the electronic device and calibration data for the electronic device.

[0020] According to one embodiment, the information about the plurality of features and the location information of the plurality of features sent through the network includes a request for a location service, and the method further includes: sending a location request based on one or more triggering conditions being met.

[0021] According to one embodiment, the one or more triggering conditions include: the distance the electronic device has moved since the last successful location request.

[0022] According to one aspect, a computer-readable medium storing computer-executable instructions configured to perform a method for operating an electronic device when executed by at least one processor, the electronic device being configured to operate within a cross-reality system. The method includes: receiving information about a three-dimensional (3D) environment, the information including a plurality of images of the 3D environment; generating a local coordinate frame for representing a position in the 3D environment; extracting a plurality of features from the plurality of images of the 3D environment; transmitting information about the plurality of features and positional information of the plurality of features expressed in the local coordinate frame to a positioning service via a network; and receiving from the positioning service at least one transformation associating the local coordinate frame with a second coordinate frame.

[0023] According to one aspect, an XR system is provided that supports specifying the location of virtual content relative to a stored map in a stored map database. The system includes: one or more computing devices configured to network communicate with one or more portable electronic devices, comprising: a communication component configured to receive from the portable electronic device information about a feature set in a three-dimensional 3D environment of the portable electronic device and positional information of features in the received feature set expressed in a first coordinate frame; a positioning component connected to the communication component, the positioning component being configured to: select a stored map from the stored map database based on a selected map having a feature set matching the received feature set, wherein the selected map includes a second coordinate frame; generate a transformation between the first coordinate frame and the second coordinate frame based on a computed alignment between the received feature set in the 3D environment of the portable electronic device and the matching feature set in the selected map; and send the transformation to the portable electronic device.

[0024] According to one embodiment, the XR system is configured to receive from an application a designation of the location of virtual content relative to persistent map features. According to one embodiment, the stored map includes a plurality of persistent map features, and the positioning component is configured to represent a transformation between a generated first coordinate frame and a second coordinate frame as a plurality of transformations between the first coordinate frame and each of the plurality of persistent map features in the selected map. According to one embodiment, the positioning component is configured to select a stored map from the stored map database by filtering maps in the stored map database based at least in part on location information associated with the portable electronic device. According to one embodiment, the location information includes one or more of the following: wireless fingerprints and / or GPS coordinates received from the portable electronic device. According to one embodiment, the positioning component is further configured to maintain information about each portable electronic device, the maintained information including a location history, and the location information including the corresponding location history of the portable electronic device.

[0025] According to one embodiment, the stored map is divided into multiple regions; and the positioning component is configured to select the stored map by selecting regions of the stored map having a feature set that matches the received feature set. According to one embodiment, the information regarding the received feature set includes descriptors calculated for the received feature set, and the positioning component is configured to select the stored map having a feature set that matches the received feature set by: identifying a candidate map set of feature sets having a number of features exceeding a threshold, wherein the descriptors of the features match the descriptors of the features in the received feature set; and selecting the stored map from the candidate map set based on an error metric associated with a transformation calculated between the received feature set and the feature set of the selected candidate map.

[0026] According to one embodiment, the positioning component is configured to generate a positioning failure indication based on the fact that no map with an error metric below a threshold is returned when searching the stored map database. According to one embodiment, the positioning service is further configured to maintain status information for each of a plurality of portable electronic devices, wherein for each of the portable electronic devices, the status information includes at least one or any combination of the following: device ID, tracking map ID corresponding to the first coordinate frame, a previously generated map reference, and / or a transformation from a canonical map to a tracking map. According to one embodiment, the received information about the received feature sets includes information about multiple feature sets, and the positioning component is configured to select a stored map by selecting a stored map whose feature sets match more than a threshold number of received feature sets.

[0027] According to one aspect, a method is provided for operating a portable electronic device to render virtual content in a 3D environment. The method includes: generating a local coordinate frame on the portable electronic device based on the output of one or more sensors on the portable electronic device; generating multiple descriptors of multiple features sensed in the 3D environment on the portable electronic device; sending the multiple descriptors of the multiple features and position information of the multiple features expressed in the local coordinate frame to a location service via a network; obtaining a transformation between a stored coordinate frame and the local coordinate frame from the location service regarding stored spatial information of the 3D environment; receiving a designation of a virtual object having a virtual object coordinate frame and the position of the virtual object relative to the stored coordinate frame; and rendering the virtual object on a display of the portable electronic device at a position determined at least in part based on the calculated transformation and the received position of the virtual object.

[0028] According to one embodiment, sending via the network includes: sending via the network the plurality of descriptors of the plurality of features and the location information of the plurality of features expressed in the local coordinate frame to a cloud-hosted location service. According to one embodiment, the transformation between a stored coordinate frame and the local coordinate frame for obtaining stored spatial information about the 3D environment from the location service includes: obtaining the stored coordinate frame via an application programming interface (API).

[0029] According to one embodiment, the portable electronic device includes a first portable electronic device containing a first processor, and the system further includes a second portable electronic device containing a second processor, wherein each of the first processor and the second processor: obtains a transformation between their respective local coordinate frames and the same stored coordinate frames, receives a designation of the virtual object, and renders the virtual object on a respective display for a user of each of the first portable electronic device and the second portable electronic device.

[0030] According to one embodiment, an application is executed to generate a specification of the virtual object and the position of the virtual object relative to the stored coordinate frame for rendering in the local coordinate frame. According to one embodiment, maintaining a local coordinate frame on the portable electronic device includes: for each of the first and second portable electronic devices: capturing multiple images of the 3D environment from one or more sensors of the portable electronic device; generating spatial information about the 3D environment based at least in part on one or more calculated persistent poses; the method further includes: for each of the first and second portable electronic devices, transmitting the generated spatial information to a remote server; and obtaining the transformation includes: receiving the transformation from a cloud-hosted positioning service.

[0031] According to one embodiment, each of the first portable electronic device and the second portable electronic device includes: a download system configured to download the stored coordinate frame from a server, enabling a transformation from a local coordinate frame to the stored coordinate frame to be performed locally in a device-based positioning mode.

[0032] According to one aspect, a method is provided for operating a portable electronic device to render virtual content in a three-dimensional 3D environment, said 3D environment comprising a framework of an XR system that provides a shared experience to each of a plurality of users using a canonical coordinate frame. The method includes: generating a local coordinate frame on the portable electronic device based on outputs from one or more sensors on the portable electronic device and a tracking map constructed on the portable electronic device; generating multiple features of said multiple images of the 3D environment on the portable electronic device; transmitting indications of said multiple features and location information of said multiple features expressed in the local coordinate frame from the portable electronic device to a location service via a network; and receiving from the location service at least one transformation associating the local coordinate frame with a second coordinate frame.

[0033] According to one embodiment, information about the 3D environment is captured from one or more sensors, the captured information including a plurality of images; and a map of at least a portion of the 3D environment is generated based on the plurality of images. According to one embodiment, the portable electronic device further includes a display, and the method further includes rendering virtual content having a position specified in a second coordinate frame on the display at a position calculated at least partially based on a transformation in the at least one transformation. According to one embodiment, the method further includes generating descriptors for the plurality of features, and sending information about the plurality of features includes sending the descriptors for the plurality of features.

[0034] According to one embodiment, the method further includes: storing the information about the plurality of features and the location information of the plurality of features in a buffer, and transmitting over the network including: transmitting the contents of the buffer together such that the information about the plurality of features and the location information of the plurality of features are transmitted together. According to one embodiment, the buffer includes an adjustable size, and the method further includes: increasing the size of the buffer in response to a failure indication. According to one embodiment, extracting the plurality of features from the plurality of images includes: extracting up to a threshold number of features from each image, and wherein the method further includes: increasing the threshold number in response to a failure indication received from the location service via the network.

[0035] According to one embodiment, the method further includes: sending a request to initiate a session with the location service via the network. According to one embodiment, the method further includes: receiving a request to initiate a session with the location service, the request including an identifier of the portable electronic device and calibration data for the portable electronic device. According to one embodiment, sending indications of the plurality of features and location information of the plurality of features expressed in the local coordinate frame from the portable electronic device to the location service via the network includes: a request for the location service; and wherein the method further includes: sending a location-specific request in response to one or more triggering conditions being met.

[0036] According to one embodiment, the method further includes: determining that the one or more triggering conditions have been met in response to identifying the distance the portable electronic device has moved since the last successful location request.

[0037] According to one aspect, a method is provided for operating a portable electronic device to render virtual content in a three-dimensional 3D environment, the 3D environment including the framework of an XR system that provides a shared experience to each of a plurality of users using a stored coordinate frame. The method includes using one or more processors to: receive information about a feature set in the 3D environment of the portable electronic device and location information of features in the received feature set expressed in a first coordinate frame at a cloud-hosted location service; select a stored map from a stored map database based on a selected map having a feature set that matches the received feature set, wherein the selected map includes a second coordinate frame; generate a transformation between the first coordinate frame and the second coordinate frame based on a computed alignment between the received feature set in the 3D environment of the portable electronic device and the feature set in the matched selected map; and send the transformation to the portable electronic device.

[0038] According to one embodiment, the method further includes: receiving from an application a designation of the location of virtual content relative to persistent map features, wherein the stored map includes a plurality of persistent map features; and representing the transformation between the generated first coordinate frame and the second coordinate frame as a plurality of transformations between the first coordinate frame and each of the plurality of persistent map features in the selected map. According to one embodiment, the method further includes: selecting a stored map from the stored map database by filtering maps in the stored map database at least in part based on location information associated with the portable electronic device. According to one embodiment, the location information includes one or more of the following: wireless fingerprints and / or GPS coordinates received from the portable electronic device. According to one embodiment, the method further includes: maintaining information about each portable electronic device, the maintained information including a location history, and the location information including the corresponding location history of the portable electronic device.

[0039] According to one embodiment, the method further includes: dividing the stored map into multiple regions; and selecting the stored map by the location service by selecting regions of the stored map having a feature set that matches the received feature set. According to one embodiment, the information regarding the received feature set includes descriptors calculated for the received feature set, and the method further includes: selecting the stored map having a feature set that matches the received feature set by: identifying a candidate map set of feature sets having a number of features exceeding a threshold, wherein the descriptors of the features match the descriptors of the features in the received feature set; and selecting the stored map from the candidate map set based on an error metric associated with a transformation calculated between the received feature set and the feature set of the selected candidate map.

[0040] According to one embodiment, the method further includes: generating a positioning failure indication based on searching the stored map database for maps with an error metric below a threshold and finding no maps with an error metric below the threshold returned. According to one embodiment, the method further includes: maintaining state information for each of a plurality of portable electronic devices, wherein for each of the portable electronic devices, the state information includes at least one or any combination of the following: a device ID, a tracking map ID corresponding to the first coordinate frame, a previously generated map reference, and / or a transformation from a canonical map to a tracking map.

[0041] According to one embodiment, the received information about the received feature sets includes information about multiple feature sets, and the method further includes: selecting a stored map by selecting stored maps that match more than a threshold number of received feature sets. According to one embodiment, the method further includes: obtaining a second transformation from the cloud-hosted positioning service, the second transformation defining a transformation between a stored coordinate frame of stored spatial information about the 3D environment and the device's local coordinate frame. According to one embodiment, the method further includes: receiving a designation of a virtual object having a virtual object coordinate frame and the position of the virtual object relative to the stored coordinate frame.

[0042] According to one embodiment, the method further includes rendering the virtual object on a display of the portable electronic device at a location determined at least in part based on the calculated transformation and the position of the received virtual object. According to one embodiment, matching the descriptors of the received feature set with the descriptors of the feature set of the stored map includes at least matching a first persistent coordinate frame (PCF) of the local coordinate frame with a first PCF of the stored map (e.g., a previously generated stored map or canonical map). According to one embodiment, matching the descriptors of the received feature set with the descriptors of the feature set of the stored map includes at least matching a second PCF of the local coordinate frame with a second PCF of the stored map (e.g., a previously generated stored map or canonical map). According to one embodiment, the method further includes obtaining or deriving a corresponding feature descriptor from any one or more, or any combination thereof: a frame, a portion of a frame, a keyframe, a persistent pose, or a persistent coordinate frame (PCF).

[0043] The foregoing overview is provided illustratively and is not intended to be limiting. Attached Figure Description

[0044] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in the various figures is represented by similar numbers. For clarity, not every component is labeled in every figure. In the drawings:

[0045] Figure 1 This is a schematic diagram illustrating an example of a simplified augmented reality (AR) scene according to some embodiments;

[0046] Figure 2 This is an exemplary simplified AR scene based on some embodiments, illustrating an exemplary use case of an XR system;

[0047] Figure 3This is a schematic diagram illustrating a data flow for a single user in an AR system according to some embodiments, the AR system being configured to provide the user with an experience of AR content that interacts with the physical world;

[0048] Figure 4 This is a schematic diagram illustrating an exemplary AR display system according to some embodiments, which displays virtual content for a single user;

[0049] Figure 5A This is a schematic diagram illustrating, according to some embodiments, how an AR display system renders AR content as the user moves through the physical world environment while wearing the system.

[0050] Figure 5B This is a schematic diagram illustrating viewing optical components and accompanying parts according to some embodiments;

[0051] Figure 6A This is a schematic diagram illustrating an AR system using a world reconstruction system according to some embodiments;

[0052] Figure 6B This is a schematic diagram illustrating components of an AR system that maintains a model of a passable world according to some embodiments.

[0053] Figure 7 It is a schematic diagram of a tracing graph formed by the path that the device traverses through the physical world.

[0054] Figure 8 This is a schematic diagram illustrating a user of an XR system that perceives virtual content according to some embodiments;

[0055] Figure 9 It involves transformations between coordinate systems according to some embodiments. Figure 8 A block diagram of the components of the first XR device in an XR system;

[0056] Figure 10 This is a schematic diagram illustrating, according to some embodiments, the exemplary transformation of the origin coordinate frame to the destination coordinate frame in order to correctly render local XR content;

[0057] Figure 11 This is a top view plan of a pupil-based coordinate frame according to some embodiments;

[0058] Figure 12 This is a top view of a camera coordinate frame including all pupil positions according to some embodiments;

[0059] Figure 13 According to some embodiments Figure 9 A schematic diagram of the display system;

[0060] Figure 14 This is a block diagram illustrating the creation of a persistent coordinate system (PCF) according to some embodiments and the attachment of XR content to the PCF;

[0061] Figure 15 This is a flowchart illustrating a method for creating and using a PCF according to some embodiments;

[0062] Figure 16 According to some embodiments, it includes a second XR device. Figure 8 A block diagram of an XR system;

[0063] Figure 17 This is a schematic diagram illustrating a room according to some embodiments and keyframes created for the various areas within the room;

[0064] Figure 18 This is a schematic diagram illustrating the establishment of a persistent pose based on keyframes according to some embodiments;

[0065] Figure 19 This is a schematic diagram illustrating the establishment of a persistent coordinate frame (PCF) based on persistent pose according to some embodiments;

[0066] Figures 20A to 20C This is a schematic diagram illustrating an example of creating a PCF according to some embodiments;

[0067] Figure 21 This is a block diagram illustrating a system for generating global descriptors for a single image and / or map, according to some embodiments;

[0068] Figure 22 This is a flowchart illustrating a method for calculating an image descriptor according to some embodiments;

[0069] Figure 23 This is a flowchart illustrating a localization method using image descriptors according to some embodiments;

[0070] Figure 24 This is a flowchart illustrating a method for training a neural network according to some embodiments;

[0071] Figure 25 This is a block diagram illustrating a method for training a neural network according to some embodiments;

[0072] Figure 26 This is a schematic diagram illustrating an AR system configured to rank and merge multiple environment maps according to some embodiments;

[0073] Figure 27 This is a simplified block diagram illustrating multiple specification maps stored on a remote storage medium according to some embodiments;

[0074] Figure 28 This is a schematic diagram illustrating a method for selecting a canonical map according to some embodiments to, for example, locate a new tracking map in one or more canonical maps and / or obtain a PCF from a canonical map;

[0075] Figure 29 This is a flowchart illustrating a method for selecting multiple rankings of an environment map according to some embodiments;

[0076] Figure 30 This illustrates some embodiments. Figure 26 A schematic diagram of an exemplary map ranking section of an AR system;

[0077] Figure 31A This is a schematic diagram illustrating examples of regional attributes of tracking maps (TM) and environmental maps in a database according to some embodiments;

[0078] Figure 31B This illustrates the determination of the purpose of using according to some embodiments. Figure 29 A schematic diagram illustrating an example of a geographic location filtering tracking map (TM);

[0079] Figure 32 This illustrates some embodiments. Figure 29 A schematic diagram illustrating an example of geolocation filtering;

[0080] Figure 33 This illustrates some embodiments. Figure 29 A schematic diagram illustrating an example of Wi-Fi BSSID filtering;

[0081] Figure 34 This illustrates the use according to some embodiments. Figure 29 A schematic diagram illustrating an example of positioning;

[0082] Figure 35 and 36 This is a block diagram of an XR system configured to rank and merge multiple environmental maps according to some embodiments.

[0083] Figure 37 This is a block diagram illustrating a method for creating an environment map of the physical world in a specification form according to some embodiments;

[0084] Figure 38A and 38B This illustrates, according to some embodiments, updating with a new tracking map. Figure 7 The tracking map is a schematic diagram of an environment map created in a standardized format.

[0085] Figures 39A to 39F This is a schematic diagram illustrating an example of a merged map according to some embodiments;

[0086] Figure 40 According to some embodiments, it can be derived from Figure 9 The first 3D local tracking map generated by the first XR device (ground) Figure 1 Two-dimensional representation of )

[0087] Figure 41 This illustrates, according to some embodiments, the transmission from a first XR device to... Figure 9 Server upload location Figure 1 A block diagram;

[0088] Figure 42 This illustrates some embodiments. Figure 16 A schematic diagram of an XR system, showing that after a first user has terminated a first session, a second user has initiated a second session using a second XR device of the XR system;

[0089] Figure 43A This illustrates a method for using according to some embodiments. Figure 42 A block diagram of a new session for the second XR device;

[0090] Figure 43B This illustrates a method for using according to some embodiments. Figure 42 A block diagram illustrating the creation of the tracking map for the second XR device;

[0091] Figure 43C This illustrates a method for transmitting data from a server to a client according to some embodiments. Figure 42 A block diagram of the second XR device download specification map;

[0092] Figure 44 This illustrates that, according to some embodiments, it can be generated by Figure 42 The second tracking map (map) generated by the second XR device Figure 2 A diagram illustrating the positioning attempt on a standard map;

[0093] Figure 45 This illustrates, according to some embodiments, that... Figure 44 The second tracking map (ground) Figure 2 This is a schematic diagram of a positioning attempt to locate a standard map, which can be further developed and has features related to the local area. Figure 2 PCF-related XR content;

[0094] Figures 46A to 46B This illustrates, according to some embodiments, that... Figure 45 land Figure 2 A diagram illustrating successful location on a standard map;

[0095] Figure 47 This illustrates, according to some embodiments, the method of using data from... Figure 46A One or more PCFs of the specification map include to Figure 45 land Figure 2 A schematic diagram of the standardized map generated in the process;

[0096] Figure 48 This illustrates some embodiments. Figure 47 Standard maps and the location on the second XR device Figure 2 A further extended schematic diagram;

[0097] Figure 49 This illustrates the uploading of data from a second XR device to a server according to some embodiments. Figure 2 A block diagram;

[0098] Figure 50 This illustrates the use of ground according to some embodiments. Figure 2 A diagram merging the standard map;

[0099] Figure 51 This is a block diagram illustrating the transmission of a new specification map from a server to a first XR device and a second XR device according to some embodiments;

[0100] Figure 52 This illustrates the ground according to some embodiments. Figure 2 Two-dimensional representation and reference ground Figure 2 A block diagram of the head coordinate frame of the second XR device;

[0101] Figure 53 This is a block diagram illustrating, in two dimensions, the adjustment of the head coordinate frame that can occur in six degrees of freedom, according to some embodiments;

[0102] Figure 54 This is a block diagram illustrating a specification map on a second XR device according to some embodiments, wherein sound is relative to the ground. Figure 2 The PCF was located;

[0103] Figure 55 and Figure 56 This is a perspective view and block diagram illustrating the use of the XR system according to some embodiments when the first user has terminated the first session and the first user has initiated a second session using the XR system;

[0104] Figure 57 and Figure 58 This is a perspective view and block diagram illustrating the use of an XR system when three users simultaneously use the XR system in the same session, according to some embodiments.

[0105] Figure 59 This is a flowchart illustrating a method for restoring and resetting head posture according to some embodiments;

[0106] Figure 60 This is a block diagram of a computer-type machine that can be found in the system of the present invention according to some embodiments;

[0107] Figure 61 This is a schematic diagram of an example XR system according to some embodiments, wherein any of a plurality of devices can access location services;

[0108] Figure 62 This is an example processing flow for operating a portable device according to some embodiments, wherein the portable device is part of an XR system providing cloud-based positioning; and

[0109] Figure 63A , Figure 63B and Figure 63C This is an example processing flow for cloud-based positioning according to some embodiments. Detailed Implementation

[0110] This document describes methods and apparatus for providing cross-reality (XR) scenes. To provide a realistic XR experience to multiple users, the XR system must understand the users' physical environment in order to correctly associate the locations of virtual objects with real-world objects. The XR system can construct an environmental map of the scene, created from images and / or depth information collected by sensors that are part of the XR device worn by the user of the XR system.

[0111] The inventors have recognized and understood that it can be beneficial to have an XR system in which each XR device develops a local map of its physical environment by integrating information from one or more images collected at a point in time during scanning. In some embodiments, the coordinate system of this map is bound to the orientation of the device at the start of scanning. This orientation may vary from session to session as a user interacts with the XR system, whether different sessions are associated with different users, each user has their own wearable device with sensors for the scanning environment, or the same user uses the same device at different times. The inventors have recognized and understood techniques for operating an XR system based on persistent spatial information that overcome the limitations of XR systems where each user device relies solely on spatial information it collects that differs in orientation relative to different user instances (e.g., snapshots) or system sessions (e.g., time between on and off). For example, these techniques can provide computationally more efficient and immersive XR scenarios for single or multiple users by allowing any of the multiple users of the XR system to create, store, and retrieve persistent spatial information.

[0112] Persistent spatial information can be represented by persistent maps, which can enable one or more features that enhance the XR experience. Persistent maps can be stored in remote storage media (e.g., the cloud). For example, a wearable device worn by a user, once activated, can retrieve a previously created and stored map from persistent storage such as cloud storage. The previously stored map might be based on data about the environment collected by sensors on the user's wearable device during a previous session. Retrieving the stored map enables the use of the wearable device without requiring sensors on the wearable device to scan the physical world. Alternatively or additionally, the system / device can similarly retrieve a suitable stored map when entering a new area of ​​the physical world.

[0113] The stored map can be represented in a specification form associated with a local reference frame on each XR device. In a multi-device XR system, the stored map accessed by one device may have been created and stored by another device, and / or may have been constructed by aggregating data about the physical world collected by sensors on multiple wearable devices, which previously existed at least in a part of the physical world represented by the stored map.

[0114] The relationship between the local map and the canonical map for each device can be determined through a positioning process. This positioning process can be performed on each XR device based on a set of canonical maps selected and sent to the device. However, the inventors have recognized and realized that network bandwidth and computing resources on XR devices can be reduced by providing positioning services that can be executed on remote processors, such as in the cloud. Therefore, battery consumption and heat generation on XR devices can be reduced, allowing the device to allocate resources such as computing time, network bandwidth, battery life, and thermal budget to provide a more immersive user experience. Furthermore, by appropriately selecting the information transmitted between each XR device and the positioning service, positioning can be performed with the latency and accuracy required to support such an immersive experience.

[0115] Sharing data about the physical world across multiple devices enables a shared user experience for virtual content. For example, two XR devices accessing the same stored map can both be positioned relative to that map. Once positioned, the user device can render virtual content with the location specified by the referenced stored map by translating the location into a frame or reference maintained by the user device. The user device can then use this local reference frame to control its display to render the virtual content at the specified location.

[0116] To support these and other functions, an XR system may include components that develop, maintain, and use persistent spatial information (including one or more stored maps) based on data about the physical world collected by sensors on the user device. These components may be distributed across the XR system, for example, through some operations on the head-mounted portion of the user device. Other components may operate on a computer and be associated with a user coupled to the head-mounted portion via a local area network (LAN) or personal area network (PAN). Still others may operate at remote locations, such as at one or more servers accessible via a wide area network (WAN).

[0117] These components might include, for example, those that can identify information of sufficient quality from information about the physical world collected by one or more user devices that can be stored as a persistent map or within a persistent map. An example of such a component, described in more detail below, is a map merging component. Such a component, for example, can receive input from a user device and determine the suitability of different portions of the input used to update the persistent map. The map merging component, for example, can divide a local map created by a user device into multiple portions, determine the merging compatibility of one or more portions with the persistent map, and merge portions that meet the qualifying merging criteria into the persistent map. The map merging component, for example, can also promote portions not merged with the persistent map to separate persistent maps.

[0118] As another example, these components may include those that help determine the appropriate persistent map that can be obtained and used by a user device. An example of such a component, described in more detail below, is a map ranking component. For instance, such a component may receive input from a user device and identify one or more persistent maps that may represent areas of the physical world in which the device operates. For example, a map ranking component may help select the persistent map that the local device will use when rendering virtual content, collecting data about the environment, or performing other actions. Alternatively or additionally, a map ranking component may help identify persistent maps that are updated as one or more user devices collect additional information about the physical world.

[0119] Other components can determine transformations that convert information captured or described with respect to one reference frame into another. For example, a sensor can be attached to a head-mounted display such that data read from the sensor indicates the physical location of an object relative to the wearer's head pose. One or more transformations can be applied to correlate this location information with a coordinate frame associated with a persistent environment map. Similarly, data indicating where to render a virtual object when expressed in the coordinate frame of the persistent environment map can be transformed once or multiple times to fit within the reference frame of the display on the user's head. As described in more detail below, multiple such transformations may exist. These transformations can be partitioned across components of the XR system, allowing them to be updated efficiently or applied to a distributed system.

[0120] In some embodiments, persistent maps can be constructed based on information collected by multiple user devices. XR devices can capture local spatial information and use information collected by sensors of each XR device at various locations and times to construct separate tracking maps. Each tracking map may include points, and each point may be associated with features of a real-world object that may include multiple features. In addition to potentially providing input for creating and maintaining persistent maps, tracking maps can also be used to track user movement within a scene, enabling the XR system to estimate the head pose of the corresponding user based on the tracking map.

[0121] The interdependence between map creation and head pose estimation presents a significant challenge. A substantial amount of processing may be required to create a map and simultaneously estimate head pose. Processing must be rapid as objects move within the scene (e.g., moving a cup on a table) and as the user moves within the scene, as latency diminishes the realism of the XR experience. On the other hand, XR devices offer limited computing resources because they must be lightweight for comfortable wear. Adding more sensors cannot compensate for the lack of computing resources, as this would undesirably increase weight. Furthermore, more sensors or more computing resources would generate heat, potentially causing deformation of the XR device.

[0122] The inventors have recognized and understand XR scenarios for operating XR systems to provide a more immersive user experience, such as estimating head pose at a frequency of 1 kHz, low utilization of computing resources associated with XR devices, such as four video graphics array (VGA) cameras that can be configured to operate at 30 Hz, one inertial measurement unit (IMU) operating at 1 kHz, computing power of a single advanced RISC machine (ARM) core, less than 1 GB of memory, and network bandwidth of less than 100 Mbps. These techniques involve reducing the processing required to generate and maintain maps and estimate head pose, and involve providing and consuming data with low computational overhead. XR systems can calculate their pose based on matched visual features. Hybrid tracking is described in U.S. Patent Application Serial No. 16 / 221,065, and its entire contents are incorporated herein by reference.

[0123] These techniques may include reducing the amount of data processed during map construction, such as constructing sparse maps by employing a set of map build points and keyframes and / or dividing maps into tiles for tile-by-tile updates. Map build points may be associated with points of interest in the environment. Keyframes may include information selected from data captured by cameras. U.S. Patent Application Serial No. 16 / 520,582 describes the determination and / or evaluation of localization maps, and its entire contents are incorporated herein by reference.

[0124] In some embodiments, persistent spatial information can be represented in a way that is easily shared between users and among distributed components including the application. For example, information about the physical world can be represented as a persistent coordinate frame (PCF). A PCF can be defined based on one or more points representing features identified in the physical world. Features can be selected such that they may be the same across user sessions of the XR system. PCFs may exist sparsely, providing less information than all available information about the physical world, so that they can be processed and transferred efficiently. Techniques for processing persistent spatial information can include creating dynamic maps based on one or more coordinate systems in the real space spanning one or more sessions, and generating persistent coordinate frames (PCFs) on the sparse maps, which can be exposed to XR applications via, for example, an application programming interface (API). These capabilities can be supported by techniques for ranking and merging multiple maps created by one or more XR devices. Persistent spatial information can also enable the rapid recovery and reset of head pose on each of one or more XR devices in a computationally efficient manner.

[0125] Furthermore, these techniques enable efficient comparison of spatial information. In some embodiments, image frames can be represented by digital descriptors. These descriptors can be computed by a transformation that maps a set of features identified in the image to a descriptor. This transformation can be performed in a trained neural network. In some embodiments, the set of features provided as input to the neural network can be a filtered set of features extracted from the image using techniques such as prioritizing potentially persistent features.

[0126] Representing image frames as descriptors enables, for example, efficient matching of new image information with stored image information. The XR system can store persistent map descriptors along with one or more frames below the persistent map. Local image frames acquired by the user device can be similarly converted into such descriptors. By selecting a stored map with a descriptor similar to that of the local image frames, one or more persistent maps that may represent the same physical space as the user device can be selected with relatively little processing. In some embodiments, descriptors can be computed for keyframes in both the local and persistent maps, further reducing processing when comparing maps. For example, this efficient comparison can be used to simplify the search for a persistent map to load onto the local device, or to search for a persistent map to update based on image information acquired using the local device.

[0127] The technologies described herein can be used with or alone in many types of devices and for many types of scenarios, including wearable or portable devices that provide augmented or mixed reality scenarios with limited computing resources. In some embodiments, the technology can be implemented through one or more services that form part of an XR system.

[0128] AR System Overview

[0129] Figure 1 and Figure 2 Scenes with virtual content are shown, displayed alongside a portion of the physical world. For illustrative purposes, an AR system is used as an example of an XR system. Figure 3-6B An exemplary AR system is shown, which includes one or more processors, memory, sensors, and a user interface that can operate according to the techniques described herein.

[0130] refer to Figure 1The text describes an outdoor AR scene 354 in which the user of the AR technology sees a park-like setting 356 in the physical world, characterized by people, trees, buildings in the background, and a concrete platform 358. In addition to these items, the user of the AR technology also perceives that they "see" a robot statue 357 standing on the concrete platform 358 in the physical world, and a flying cartoonish avatar 352 that appears to be the head of a bumblebee, even though these elements (e.g., the avatar 352 and the robot statue 357) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, producing an AR technology that fosters a comfortable, natural feeling and rich presentation of virtual image elements within other virtual or physical world image elements is challenging.

[0131] Such AR scenarios can be achieved through a system that builds a map of the physical world based on tracking information. This allows users to place AR content in the physical world, determine the location of the AR content on the map, preserve the AR scene so that the placed AR content can be reloaded and displayed in the physical world during, for example, different AR experience sessions, and allow multiple users to share the AR experience. The system can build and update a digital representation of the physical world surface around the user. This representation can be used to render virtual content as if it were fully or partially occluded by physical objects between the user and the rendering location of the virtual content, for placing virtual objects in physical-based interactions, for virtual character path planning and navigation, or for other operations where information about the physical world is used.

[0132] Figure 2 Another example of an indoor AR scene 400 according to some embodiments is depicted, illustrating an exemplary use case of an XR system. Exemplary scene 400 is a living room with walls, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical objects, the user of the AR technology can also perceive virtual objects such as images on the wall behind the sofa, birds flying through the door, a deer peeking out from the bookshelf, and decorative items in the form of a windmill placed on the coffee table.

[0133] For images on walls, AR technology needs not only information about the wall surface but also about objects and surfaces within the room (such as the shape of lights) to occlude the image and correctly render virtual objects. For flying birds, AR technology needs information about all objects and surfaces around the room to render the birds with realistic physical effects, allowing them to avoid objects and surfaces or bounce off collisions. For deer, AR technology needs information about surfaces (such as the floor or a coffee table) to calculate the deer's placement. For windmills, the system can identify objects detached from the table and determine if they are movable, while corners of shelves or walls can be identified as stationary. This distinction can be used to determine which parts of the scene are used or updated in each of various operations.

[0134] Virtual objects can be placed within previous AR experience sessions. When a new AR experience session begins in the living room, the AR technology needs to accurately display the virtual objects in the locations where they were previously placed and are actually visible from different perspectives. For example, a windmill should appear to be standing on a book, not floating above a table in a different location where there are no books. This floating could occur if the user's location in the new AR experience session is not accurately positioned in the living room. As another example, if the user views the windmill from a different perspective than when it was placed, the AR technology needs to display the corresponding side of the windmill.

[0135] A scene can be presented to a user via a system comprising multiple components, including a user interface capable of stimulating one or more user senses, such as vision, sound, and / or touch. Additionally, the system may include one or more sensors capable of measuring parameters of physical portions of the scene, including the user's position and / or movement within those physical portions. Furthermore, the system may include one or more computing devices, and associated computer hardware, such as memory. These components may be integrated into a single device or distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated into a wearable device.

[0136] Figure 3An AR system 502 according to some embodiments is depicted, configured to provide an experience of AR content interacting with the physical world 506. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by a user as part of a headset, allowing the user to wear the display over their eyes like a pair of goggles or glasses. At least a portion of the display may be transparent, allowing the user to observe a perspective reality 510. Perspective reality 510 may correspond to a portion of the physical world 506 within the current viewpoint of the AR system 502, which may correspond to the user's viewpoint when the user wears a headset incorporating the AR system's display and sensors to obtain information about the physical world.

[0137] AR content can also be displayed on display 508, overlaid on perspective reality 510. To provide accurate interaction between AR content and perspective reality 510 on display 508, AR system 502 may include sensor 522 configured to capture information about the physical world 506.

[0138] Sensor 522 may include one or more depth sensors that output depth maps 512. Each depth map 512 may have multiple pixels, each pixel representing a distance from a surface in the physical world 506 relative to the depth sensor in a particular orientation. Raw depth data may be derived from the depth sensors to create the depth map. This depth map may be updated as quickly as the depth sensors can form new images, potentially hundreds or thousands of times per second. However, this data may be noisy and incomplete, and may have holes shown as black pixels on the illustrated depth map.

[0139] The system may include other sensors, such as image sensors. Image sensors can acquire monocular or stereoscopic information, which can be processed to represent the physical world in other ways. For example, images can be processed in world reconstruction component 516 to create a mesh that represents the connected parts of objects in the physical world. Metadata about such objects, including, for example, color and surface texture, can be similarly acquired using sensors and stored as part of the world reconstruction.

[0140] The system can also acquire information about the user's head pose (or "pose") relative to the physical world. In some embodiments, the system's head pose tracking component can be used to calculate the head pose in real time. The head pose tracking component can represent the user's head pose in a coordinate system with six degrees of freedom, including, for example, translations (e.g., forward / backward, up / down, left / right) and rotations (e.g., pitch, yaw, and roll) about the three vertical axes. In some embodiments, sensor 522 may include an inertial measurement unit ("IMU") that can be used to calculate and / or determine head pose 514. Head pose 514 for depth mapping may indicate, for example, the current viewpoint of the sensor that captures the depth map with six degrees of freedom, but head pose 514 can be used for other purposes, such as associating image information with a specific part of the physical world or associating the position of a display worn on the user's head with the physical world.

[0141] In some embodiments, head pose information can be derived in a manner different from that of an IMU (such as analyzing objects in an image). For example, a head pose tracking component can calculate the relative position and orientation of the AR device relative to a physical object based on visual information captured by a camera and inertial information captured by an IMU. The head pose tracking component can then calculate the head pose of the AR device, for example, by comparing the calculated relative position and orientation of the AR device relative to a physical object with features of the physical object. In some embodiments, this comparison can be performed by identifying features in images captured using one or more sensors 522 that are time-stable so that changes in the position of these features in the captured images over time can be correlated with changes in the user's head pose.

[0142] In some embodiments, an AR device can construct a map based on feature points identified in a series of consecutive images captured as a user moves through the physical world with the AR device. Although each image frame can be taken from different poses of the user's movement, the system can adjust the orientation of the features of each consecutive image frame to match the orientation of the initial image frame by matching the features of the consecutive image frames with those of previously captured image frames. Translation of consecutive image frames such that points representing the same feature will match corresponding feature points in previously collected image frames can be used to align each consecutive image frame to match the orientation of previously processed image frames. Frames in the generated map can have a common orientation established when the first image frame is added to the map. The map has multiple sets of feature points in a common reference frame, which can be used to determine the user's pose in the physical world by matching features in the current image frame with the map. In some embodiments, this map can be referred to as a tracking map.

[0143] In addition to tracking the user's posture in the environment, this map enables other components of the system, such as world reconstruction component 516, to determine the position of physical objects relative to the user. World reconstruction component 516 can receive depth map 512 and head pose 514 from sensors, along with any other data, and integrate this data into reconstruction 518. Reconstruction 518 can be more complete and less noisy than the sensor data. World reconstruction component 516 can update reconstruction 518 using spatial and temporal averaging of sensor data from multiple viewpoints over time.

[0144] Reconstruction 518 can include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats can represent alternative representations of the same part of the physical world or can represent different parts of the physical world. In the example shown, on the left side of Reconstruction 518, the part of the physical world is rendered as a global surface; on the right side of Reconstruction 518, the part of the physical world is rendered as a mesh.

[0145] In some embodiments, the map maintained by the head pose component may be sparse relative to other maps of the physical world that may be maintained. A sparse map may indicate the location of points of interest and / or structures (e.g., corners or edges) rather than providing information about the location of surfaces and other possible features. In some embodiments, the map may include image frames captured by sensor 522. These frames may be simplified to features that can represent points of interest and / or structures. Information about the pose of the user from which the frame was acquired may also be stored as part of the map, in conjunction with each frame. In some embodiments, each image acquired by the sensor may or may not be stored. In some embodiments, as images are collected by the sensor, the system may process the images and select a subset of image frames for further computation. This selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may add new image frames to the map, for example, based on overlap with previous image frames already added to the map or based on image frames containing a sufficient number of features determined to likely represent stationary objects. In some embodiments, the selected image frames or sets of features from the selected image frames may be used as keyframes of the map, which provide spatial information.

[0146] AR system 502 can integrate sensor data from multiple perspectives of the physical world over time. As devices including sensors move, the sensor poses (e.g., position and orientation) can be tracked. Since the sensor frame poses and their relationships to other poses are known, each of these multiple viewpoints of the physical world can be fused together to form a single composite reconstruction of the physical world. This can be used as an abstract layer of a map and provide spatial information. By using spatial and temporal averaging (i.e., averaging data from multiple viewpoints over time) or any other suitable method, the reconstruction can be more complete and less noisy than the original sensor data.

[0147] exist Figure 3 In the illustrated embodiment, the map represents a portion of the physical world in which a user with a single wearable device is present. In that case, the head pose associated with a frame in the map can be represented as a local head pose, indicating an orientation relative to the initial orientation of the single device at the start of the session. For example, the head pose can be tracked relative to the initial head pose when the device is turned on, or otherwise operated to scan the environment to establish a representation of that environment.

[0148] Incorporating the portion representing the physical world, a map can include metadata. Metadata can, for example, indicate the time when sensor information used to form the map was captured. Alternatively or additionally, metadata can indicate the sensor's location at the time the information used to form the map was captured. Location can be represented directly, such as using information from a GPS chip, or indirectly, such as using a wireless (e.g., Wi-Fi) signature indicating the strength of a signal received from one or more wireless access points while the sensor data was being collected, and / or an identifier such as a BSSID of the wireless access point to which the user equipment was connected while the sensor data was being collected.

[0149] Reconstruction 518 can be used for AR functions, such as generating a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation may change as the user moves or objects in the real world change. Aspects of Reconstruction 518 can be used, for example, by component 520, which generates a global surface representation that changes in world coordinates, and can be used by other components.

[0150] AR content can be generated based on this information, such as through AR application 504. AR application 504 can be, for example, a game program that performs one or more functions based on information about the physical world, such as visual occlusion, physics-based interaction, and environmental reasoning. It can perform these functions by querying data in different formats from the reconstruction 518 generated by world reconstruction component 516. In some embodiments, component 520 can be configured to update its output when the representation in the region of interest of the physical world changes. For example, the region of interest can be set to approximate a portion of the physical world near the system user, such as a portion within the user's field of vision, or projected (predicted / determined) to enter the user's field of vision.

[0151] AR application 504 can use this information to generate and update AR content. The virtual portion of the AR content can be combined with perspective reality 510 and displayed on display 508 to create a realistic user experience.

[0152] In some embodiments, an AR experience can be provided to a user via an XR device, which may be a wearable display device that is part of a system that may include remote processing and / or remote data storage and / or, in some embodiments, other wearable display devices worn by other users. For simplicity of illustration, Figure 4 An example of a system 580 (hereinafter referred to as "System 580") including a single wearable device is shown. System 580 includes a head-mounted display device 562 (hereinafter referred to as "Display Device 562"), and various mechanical and electronic modules and systems supporting the functionality of Display Device 562. Display Device 562 may be coupled to a frame 564 that may be worn by a user or viewer of the display system 560 (hereinafter referred to as "User 560") and configured to position Display Device 562 in front of User 560's eyes. According to various embodiments, Display Device 562 may be displayed sequentially. Display Device 562 may be monocular or binocular. In some embodiments, Display Device 562 may be... Figure 3 Example of display 508 in the image.

[0153] In some embodiments, speaker 566 is coupled to frame 564 and positioned near the ear canal of user 560. In some embodiments, another speaker (not shown) is positioned near another ear canal of user 560 to provide stereo / plastic sound control. Display device 562 is operatively coupled to local data processing module 570, such as via a wired cable or wireless connection 568, which can be mounted in various configurations, such as being fixedly attached to frame 564, fixedly attached to a helmet or hat worn by user 560, embedded in headphones, or otherwise removably attached to user 560 (e.g., in a backpack configuration, in a belt-coupled configuration).

[0154] The local data processing module 570 may include a processor and digital memory such as non-volatile memory (e.g., flash memory), both of which can be used to assist in the processing, caching, and storage of data. Data includes: a) data captured from sensors (e.g., operatively coupled to frame 564) or otherwise attached to user 560, such as image capture devices (e.g., cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, radios, and / or gyroscopes; and / or b) data acquired and / or processed using remote processing module 572 and / or remote data repository 574, which may then be passed to display device 562.

[0155] In some embodiments, the wearable device can communicate with remote components. The local data processing module 570 can be operatively coupled to the remote processing module 572 and the remote data repository 574 via communication links 576, 578 (such as via wired or wireless communication links), such that these remote processing modules 572 and the remote data repository 574 are operatively coupled to each other and can be used as resources of the local data processing module 570. In a further embodiment, as a supplement to or alternative to the remote data repository 574, the wearable device can access cloud-based remote data repositories and / or services. In some embodiments, the head pose tracking component described above can be implemented at least partially in the local data processing module 570. In some embodiments, Figure 3 The world reconstruction component 516 can be implemented at least partially in the local data processing module 570. For example, the local data processing module 570 can be configured to execute computer-executable instructions to generate a map and / or a physical world representation based at least partially on at least a portion of the data.

[0156] In some embodiments, processing may be distributed across a local processor and a remote processor. For example, local processing may be used to construct a map (e.g., a tracking map) on the user device based on sensor data collected using sensors on the user device. Such a map may be used by applications on the user device. Additionally, previously created maps (e.g., canonical maps) may be stored in a remote data repository 574. Where appropriate stored or persistent maps are available, they may be used in place of tracking maps created locally on the device or in addition to tracking maps created locally on the device. In some embodiments, a tracking map may be mapped to a stored map, such that a correspondence is established between the tracking map and the canonical map, wherein the tracking map may be oriented relative to the location of the wearable device when the user turns on the system, and the canonical map may be oriented relative to one or more persistent features. In some embodiments, a persistent map may be loaded onto the user device to allow the user device to render virtual content without the latency associated with scanning the location, thereby constructing a tracking map of the user's entire environment based on sensor data acquired during the scanning. In some embodiments, the user device may access a remote persistent map (e.g., stored in the cloud) without needing to download the persistent map to the user device.

[0157] In some embodiments, spatial information can be transmitted from the wearable device to a remote service, such as a cloud service configured to locate the device to a stored map maintained on a cloud service. According to one embodiment, the location process can be performed in the cloud, matching the device location against an existing map (e.g., a canonical map) and returning a transformation that links virtual content to the wearable device's location. In such embodiments, the system can avoid transmitting maps from a remote resource to the wearable device. Other embodiments can be configured for both device-based and cloud-based location, for example, to enable functionality when network connectivity is unavailable or the user chooses not to enable cloud-based location.

[0158] Alternatively or additionally, the tracking map can be merged with previously stored maps to expand or improve the quality of those maps. The process of determining whether a suitable previously created environment map is available and / or merging the tracking map with one or more stored environment maps can be performed in the local data processing module 570 or the remote processing module 572.

[0159] In some embodiments, the local data processing module 570 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which limits the computational budget of the local data processing module 570 but enables a smaller device. In some embodiments, the world reconstruction component 516 may use a computational budget smaller than that of a single advanced RISC machine (ARM) core to generate a physical world representation in real time in a non-predefined space, allowing the remaining computational budget of a single ARM core to be accessed for other purposes, such as, for example, mesh extraction.

[0160] In some embodiments, the remote data repository 574 may include a digital data storage facility that is available via the Internet or other networked configurations in a “cloud” resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, thereby allowing fully autonomous use from the remote module. In some embodiments, all data is stored and all or most computations are performed in the remote data repository 574, thereby allowing for smaller devices. For example, world reconstruction may be stored wholly or partially in this repository 574.

[0161] In embodiments where data is stored remotely and accessible via a network, the data can be shared by multiple users of the augmented reality system. For example, user devices can upload their tracking maps to enhance the database of environment maps. In some embodiments, tracking map uploads occur at the end of a user session with the wearable device. In some embodiments, tracking map uploads can occur continuously, semi-continuously, or intermittently at a predefined time, after a predefined period following a previous upload, or when triggered by an event. Tracking maps uploaded by any user device, whether based on data from that user device or any other user device, can be used to expand or improve previously stored maps. Similarly, persistent maps downloaded to a user device can be based on data from that user device or any other user device. In this way, users can easily obtain high-quality environment maps to improve their experience in the AR system.

[0162] In another embodiment, persistent map downloads can be limited and / or avoided based on positioning performed on a remote resource (e.g., in the cloud). In such a configuration, the wearable device or other XR device transmits feature information combined with gesture information (e.g., device positioning information when features represented in the feature information are sensed) to a cloud service. One or more components of the cloud service can match the feature information with a corresponding stored map (e.g., a canonical map) and generate a transformation between the coordinate systems of the tracking map maintained by the XR device and the canonical map. Each XR device, whose tracking map is positioned relative to the canonical map, can accurately render virtual content at a location specified relative to the canonical map based on its own tracking.

[0163] In some embodiments, the local data processing module 570 is operatively coupled to the battery 582. In some embodiments, the battery 582 is a removable power source, such as above a counter battery. In other embodiments, the battery 582 is a lithium-ion battery. In some embodiments, the battery 582 includes both an internal lithium-ion battery that can be charged by the user 560 during the non-operational period of the system 580 and a removable battery, allowing the user 560 to operate the system 580 for longer periods without having to connect to a power source to charge the lithium-ion battery or without having to shut down the system 580 to replace the battery.

[0164] Figure 5A The illustration depicts a user 530 wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as "environment 532"). Information captured by the AR system along the user's movement path can be processed into one or more tracking maps. The user 530 positions the AR display system at location 534, and the AR display system records environmental information of the accessible world relative to location 534 (e.g., digital representations of real objects in the physical world, which can be stored and updated as real objects change in the physical world). This information can be combined with images, features, directional audio input, or other desired data and stored as a gesture. Location 534 is aggregated into data input 536, for example, as part of the tracking map, and is processed at least by an accessible world module 538, which can, for example, through... Figure 4 The processing is implemented on the remote processing module 572. In some embodiments, the traversable world module 538 may include a head pose component and a world reconstruction component 516, such that the processed information can be combined with other information related to the physical objects used in rendering the virtual content to indicate the location of the objects in the physical world.

[0165] The walkable world module 538 at least partially determines the location and manner in which the AR content 540, as determined from data input 536, can be placed in the physical world. The AR content is "placed" in the physical world by presenting both the physical world rendering and the AR content itself via a user interface; the AR content is rendered as if interacting with objects in the physical world, and the objects in the physical world are rendered as if the AR content obscures the user's view of these objects when appropriate. In some embodiments, the shape and position of the AR content 540 can be determined by appropriately selecting portions of a fixed element 542 (e.g., a table) from a reconstruction (e.g., reconstruction 518). As an example, the fixed element could be a table, and the virtual content could be positioned such that it appears to appear to be on that table. In some embodiments, the AR content can be placed within a structure in a field of view 544, which could be the current field of view or an estimated future field of view. In some embodiments, the AR content can persist relative to a model 546 (e.g., a grid) in the physical world.

[0166] As depicted, fixed element 542 serves as a proxy (e.g., a digital copy) for any fixed element within the physical world that can be stored in the walkable world module 538, allowing user 530 to perceive content on fixed element 542 without the system having to map to fixed element 542 every time user 530 sees it. Therefore, fixed element 542 can be a mesh model from a previous modeling session, or it can be determined by a single user but still stored by the walkable world module 538 for future reference by multiple users. Thus, the walkable world module 538 can recognize environment 532 from a previously mapped environment and display AR content without requiring user 530's device to first map all or part of environment 532, saving computation time and cycles and avoiding latency in any rendered AR content.

[0167] A physical world mesh model 546 can be created using an AR display system, and appropriate surfaces and measurements for interacting with and displaying AR content 540 can be stored by the walkable world module 538 for future retrieval by user 530 or other users without requiring complete or partial model re-creation. In some embodiments, data input 536 is input such as geolocation, user ID, and current activity to indicate to the walkable world module 538 which of one or more fixed elements 542 is available, which AR content 540 was last placed on a fixed element 542, and whether that same content is displayed (this AR content is "persistent" regardless of how the user views a particular walkable world model).

[0168] Even in embodiments where objects are considered fixed (e.g., a kitchen table), the walkable world module 538 can periodically update those objects in the physical world model to account for the possibility of changes in the physical world. The model of fixed objects may be updated very infrequently. Other objects in the physical world may be moving or otherwise not considered fixed (e.g., a kitchen chair). To render a realistic AR scene, the AR system can update the positions of these non-fixed objects at a much higher frequency than it would be for updating fixed objects. To accurately track all objects in the physical world, the AR system can acquire information from multiple sensors, including one or more image sensors.

[0169] Figure 5B This is a schematic diagram of viewing the optical assembly 548 and accompanying components. In some embodiments, two eye-tracking cameras 550, pointed at the user's eye 549, detect measurements of the user's eye 549, such as eye shape, eyelid occlusion, pupil direction, and flicker.

[0170] In some embodiments, one of the sensors may be a depth sensor 551, such as a time-of-flight sensor, which emits signals into the world and detects reflections of those signals from nearby objects to determine the distance to a given object. The depth sensor can, for example, quickly determine whether an object has entered the user's field of vision due to the movement of those objects or changes in the user's posture. However, information about the object's position in the user's field of vision may alternatively or additionally be collected by other sensors. Depth information may, for example, be obtained from a stereoscopic image sensor or an all-light sensor.

[0171] In some embodiments, the world camera 552 records a view larger than its periphery to map and / or otherwise create a model of the environment 532 and detect input that can affect AR content. In some embodiments, the world camera 552 and / or camera 553 may be grayscale and / or color image sensors that can output grayscale and / or color image frames at fixed time intervals. Camera 553 may further capture images of the physical world within the user's field of view at specific times. Pixels can be repeatedly sampled even if the pixel values ​​of the frame-based image sensor remain unchanged. Each of the world camera 552, camera 553, and depth sensor 551 has a corresponding field of view 554, 555, and 556 to capture images from sources such as... Figure 34 Data is collected and recorded in the physical world scene of 532 depicted in A.

[0172] The inertial measurement unit 557 can determine the motion and orientation of the viewing optical assembly 548. In some embodiments, each component is operatively coupled to at least one other component. For example, the depth sensor 551 is operatively coupled to the eye-tracking camera 550 to confirm the measured adaptation relative to the actual distance at which the user's eye 549 is looking.

[0173] It should be understood that viewing optical component 548 may include Figure 34 Some of the components shown in B may be included, and components other than those shown may also be included. For example, in some embodiments, the viewing optics 548 may include two world cameras 552 instead of four. Alternatively or additionally, cameras 552 and 553 may not need to capture visible light images of their entire field of view. The viewing optics 548 may include other types of components. In some embodiments, the viewing optics 548 may include one or more dynamic vision sensors (DVS) whose pixels can asynchronously respond to relative changes in light intensity exceeding a threshold.

[0174] In some embodiments, the viewing optics 548 may not include the depth sensor 551, based on time-of-flight information. For example, in some embodiments, the viewing optics 548 may include one or more all-optical cameras whose pixels can capture light intensity and the angle of incident light, thereby determining depth information. For example, the all-optical camera may include an image sensor covered with a transmission diffraction mask (TDM). Alternatively or additionally, the all-optical camera may include an image sensor containing angle-sensitive pixels and / or phase-detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such a sensor can be used as a source of depth information, either in place of or in addition to the depth sensor 551.

[0175] It should also be understood that Figure 5B The configuration of the components is provided as an example. The viewing optics 548 may include components with any suitable configuration, which may be set to provide the user with the maximum field of view actually feasible for a particular set of components. For example, if the viewing optics 548 has a world camera 552, the world camera may be placed in the central area of ​​the viewing optics rather than on the side.

[0176] Information from sensors in the viewing optics 548 can be coupled to one or more processors in the system. The processors can generate data that can be rendered to allow a user to perceive interaction with objects in the physical world. This rendering can be implemented in any suitable manner, including generating image data depicting both physical and virtual objects. In other embodiments, physical and virtual content can be depicted in a scene by modulating the opacity of a display device as the user views the physical world. The opacity can be controlled to create the appearance of virtual objects and also prevent the user from seeing objects in the physical world that are occluded by the virtual objects. In some embodiments, when viewed through a user interface, the image data may consist only of virtual content that can be modified to allow the user to perceive the virtual content as realistically interacting with the physical world (e.g., clipping content to account for occlusion).

[0177] The position at which the displayed content on the viewing optics 548 creates the impression of an object at a specific location can depend on the physical properties of the viewing optics. Furthermore, the user's head posture relative to the physical world and the direction of the user's gaze can influence the specific position of the displayed content on the viewing optics that will appear in the physical world. The sensor, as described above, can collect this information and / or provide information from which it can be calculated, allowing the processor receiving the sensor input to calculate the position on the viewing optics 548 where the object should be rendered, thereby creating the desired appearance for the user.

[0178] Regardless of how content is presented to the user, a model of the physical world can be used to correctly calculate the characteristics of virtual objects that are affected by physical objects, including the shape, position, motion, and visibility of virtual objects. In some embodiments, the model may include a reconstruction of the physical world, such as Reconstruction 518.

[0179] The model can be created based on data collected from sensors on a user's wearable device. However, in some embodiments, the model can be created from data collected from multiple users, which can be aggregated on computing devices remotely from all users (and this data can be in the "cloud").

[0180] The model can be created at least in part by a world reconstruction system, such as, for example, Figure 6A A more detailed description Figure 3The world reconstruction component 516 may include a sensing module 660 that can generate, update, and store a representation of a portion of the physical world. In some embodiments, the sensing module 660 may represent a portion of the physical world within the reconstruction range of a sensor as a plurality of voxels. Each voxel may correspond to a 3D cube of a predetermined volume in the physical world and includes surface information indicating the presence of a surface within the volume represented by the voxel. Values ​​may be assigned to voxels indicating whether their corresponding volumes have been determined to include surfaces of physical objects, whether they are determined to be empty, or whether they have not yet been measured by a sensor and are therefore unknown. It should be understood that it is not necessary to explicitly store the values ​​of voxels that indicate they are determined to be empty or unknown, as the values ​​of voxels can be stored in computer memory in any suitable manner, including not storing information about voxels determined to be empty or unknown.

[0181] In addition to generating information for the representation of the persistent world, the perception module 660 can also identify and output indications of changes in the area surrounding the user of the AR system. Such indications of change can trigger updates to volumetric data stored as part of the persistent world, or trigger other functions, such as triggering the generation of AR content to update the AR content triggering component 604.

[0182] In some embodiments, the sensing module 660 can identify changes based on a signed distance function (SDF) model. The sensing module 660 can be configured to receive sensor data such as, for example, a depth map 660a and a head pose 660b, and then fuse the sensor data into an SDF model 660c. The depth map 660a can directly provide SDF information, and the image can be processed to obtain SDF information. SDF information represents the distance to the sensor used to capture this information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and therefore from the user's perspective. The head pose 660b enables the SDF information to be correlated with voxels in the physical world.

[0183] In some embodiments, the sensing module 660 may generate, update, and store representations of portions of the physical world within the sensing range. The sensing range may be determined at least in part based on the sensor's reconstructed range, which may be determined at least in part based on the limitations of the sensor's observation range. As a specific example, an active depth sensor operating using active IR pulses can reliably operate over a distance range, thereby creating an observation range for the sensor that can range from a few centimeters or tens of centimeters to several meters.

[0184] The world reconstruction component 516 may include additional modules that can interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data acquired by the perception module 660. The persistent world module 662 may also include representations of the physical world in various formats. For example, it may store volume metadata 662b such as voxels, as well as meshes 662c and planes 662d. In some embodiments, other information, such as depth maps, may be stored.

[0185] In some embodiments, the representation of the physical world (such as Figure 6A The representation shown can provide relatively dense information about the physical world compared to sparse maps (such as feature-point-based tracking maps as described above).

[0186] In some embodiments, the perception module 660 may include modules that generate representations of the physical world in various formats, including, for example, a mesh 660d, a plane, and semantics 660e. Representations of the physical world can be stored across local and remote storage media. Depending on, for example, the location of the storage media, representations of the physical world can be described in different coordinate frames. For example, a representation of the physical world stored on the device can be described in a coordinate frame relative to the device's locality. The representation of the physical world may have a corresponding representation stored in the cloud. The corresponding representation in the cloud can be described in a coordinate frame shared by all devices in the XR system.

[0187] In some embodiments, these modules may generate a representation based on data within the perception range of one or more sensors at the time of representation generation, as well as data captured in previous time and information in the persistent world module 662. In some embodiments, these components may operate with respect to depth information captured using a depth sensor. However, the AR system may include a vision sensor and may generate such a representation by analyzing monocular or binocular visual information.

[0188] In some embodiments, these modules can operate on regions of the physical world. When the sensing module 660 detects a change in the physical world within a sub-region of the physical world, those modules can be triggered to update the sub-region. For example, such a change can be detected by detecting new surfaces or other criteria (e.g., changing the values ​​of a sufficient number of voxels representing the sub-region) in the SDF model 660c.

[0189] The world reconstruction component 516 may include component 664 that can receive a representation of the physical world from the perception module 660. Information about the physical world may be extracted by these components based on, for example, a usage request from an application. In some embodiments, information may be pushed to the usage component, such as via indication of changes in a pre-identified area or changes in the representation of the physical world within the perception range. Component 664 may include, for example, a game program and other components that perform processing for visual occlusion, physics-based interaction, and environmental reasoning.

[0190] In response to a query from component 664, perception module 660 may send a representation of the physical world in one or more formats. For example, when component 664 indicates that the use is for visual occlusion or physical-based interaction, perception module 660 may send a representation of a surface. When component 664 indicates that the use is for environmental reasoning, perception module 660 may send a mesh, plane, and semantics of the physical world.

[0191] In some embodiments, the perception module 660 may include a component that formats information to provide component 664. An example of such a component may be a ray-projection component 660f. Using a component (e.g., component 664), information about the physical world can be queried from a specific viewpoint, for example. The ray-projection component 660f can select from one or more representations of physical world data within the field of view from that viewpoint.

[0192] As should be understood from the above description, the perception module 660, or another component of the AR system, can process data to create a 3D representation of parts of the physical world. The amount of data to be processed can be reduced by: culling portions of the 3D reconstructed volume based at least in part on camera frustum and / or depth images; extracting and preserving planar data; capturing, preserving, and updating 3D reconstructed data in blocks that allow for local updates while maintaining nearest-neighbor consistency; providing occlusion data to applications that generate such scenes, where the occlusion data is derived from a combination of one or more depth data sources; and / or performing multi-stage mesh simplification. The reconstruction can contain data of varying complexity, including, for example, raw data (e.g., real-time depth data), fused volumetric data (e.g., voxels), and computational data (e.g., meshes).

[0193] In some embodiments, the components of a passable world model can be distributed, with some parts executing locally on the XR device and others executing remotely, such as on a network-connected server or in the cloud. The allocation of information processing and storage between the local XR device and the cloud can impact the functionality and user experience of the XR system. For example, reducing processing on the local device by distributing processing to the cloud can extend battery life and reduce heat generated on the local device. However, distributing too much processing to the cloud can introduce undesirable latency, leading to an unacceptable user experience.

[0194] Figure 6B A distributed component architecture 600 configured for spatial computing according to some embodiments is depicted. The distributed component architecture 600 may include a world-accessible component 602 (e.g., Figure 5A The components include PW 538, Lumin OS 604, API 606, SDK 608, and Application 610. LuminOS 604 may include a Linux-based kernel with custom drivers compatible with XR devices. API 606 may include an application programming interface that allows XR applications (e.g., Application 610) to access the spatial computing features of the XR device. SDK 608 may include a software development kit that allows the creation of XR applications.

[0195] One or more components in architecture 600 can create and maintain a model of a walkable world. In this example, sensor data is collected on a local device. Processing of this sensor data can be performed partly locally on the XR device and partly in the cloud. PW 538 may include an environment map created at least in part based on data captured by AR devices worn by multiple users. During an AR experience session, the individual AR devices (such as those combined above) Figure 4 The wearable device described can create tracking maps, which are a type of map.

[0196] In some embodiments, the device may include components for constructing sparse and dense maps. The tracking map can serve as a sparse map and may include the head pose of the AR device scanning the environment and information about objects detected within that environment at each head pose. Those head poses may be maintained locally for each device. For example, the head pose on each device may be relative to the initial head pose when the device initiates its session. As a result, each tracking map may be local to the device that created it. The dense map may include surface information, which may be represented by a grid or depth information. Alternatively or additionally, the dense map may include higher-level information derived from the surface or depth information, such as the location and / or features of planes and / or other objects.

[0197] In some embodiments, the creation of dense maps can be independent of the creation of sparse maps. For example, the creation of dense and sparse maps can be performed in separate processing pipelines within the AR system. Separate processing allows for the generation or processing of different types of maps at different rates. For example, the refresh rate of the sparse map may be faster than that of the dense map. However, in some embodiments, even when performed in different pipelines, the processing of dense and sparse maps may be related. For example, changes in the physical world revealed in the sparse map can trigger updates to the dense map, and vice versa. Furthermore, even if created independently, these maps can be used together. For example, a coordinate system derived from the sparse map can be used to define the position and / or orientation of objects in the dense map.

[0198] Sparse and / or dense maps can be persistently stored for reuse by the same device and / or shared with other devices. Such persistence can be achieved by storing the information in the cloud. AR devices can send tracking maps to the cloud, thereby merging them, for example, with an environment map selected from previously stored persistent maps in the cloud. In some embodiments, selected persistent maps can be sent from the cloud to the AR device for merging. In some embodiments, persistent maps can be oriented relative to one or more persistent coordinate frames. Such maps can be used as canonical maps because they can be used by any of multiple devices. In some embodiments, a model of a walkable world can include one or more canonical maps or be created from one or more canonical maps. Even when performing some operations based on the device's local coordinate frame, the device can use the canonical map by determining the transformation between the device's local coordinate frame and the canonical map.

[0199] Standard maps can originate from tracking maps (TM) (e.g., Figure 31A The TM (1102) can be promoted to a canonical map. The canonical map can be persistently stored so that once a device accessing the canonical map determines the transformation between its local coordinate system and the coordinate system of the canonical map, it can use the information in the canonical map to determine the location of objects represented in the canonical map in the physical world around the device. In some embodiments, the TM can be a sparse head pose map created by the XR device. In some embodiments, a canonical map can be created when the XR device sends one or more TMs to a cloud server to be merged with additional TMs captured by the XR device at different times or by other XR devices.

[0200] Standard maps or other maps can provide information about the various parts of the physical world represented by the data processed to create the corresponding map. Figure 7An exemplary tracking map 700 according to some embodiments is depicted. The tracking map 700 can provide a plan view 706 of corresponding physical objects in the physical world, represented by points 702. In some embodiments, map points 702 can represent features of physical objects that may include multiple features. For example, each corner of a table can be a feature represented by points on the map. These features can be derived by processing images, such as images acquired using sensors from a wearable device in an augmented reality system. For example, features can be derived by processing image frames output from the sensors to identify features based on large gradients or other suitable criteria in the images. Further processing may limit the number of features in each frame. For example, the processing may select features that might represent persistent objects. One or more heuristics may be applied to this selection.

[0201] The tracking map 700 may include data about points 702 collected by the device. For each image frame containing data points included in the tracking map, a pose may be stored. The pose may represent the orientation from which the image frame was captured, such that feature points within each image frame can be spatially correlated. This pose can be determined using location information, such as that derived from sensors on the wearable device (e.g., IMU sensors). Alternatively or additionally, the pose can be determined by matching the image frame to other image frames depicting overlapping portions of the physical world. This location correlation can be found by matching subsets of feature points in two frames, allowing the calculation of a relative pose between the two frames. A relative pose may be sufficient for the tracking map, as it can be relative to a device-local coordinate system established based on the device's initial pose when the tracking map was first constructed.

[0202] Not all feature points and image frames collected by the device can be retained as part of the tracking map, as much of the information collected by the sensors may be redundant. Instead, only certain frames can be added to the map. These frames can be selected based on one or more criteria, such as the degree of overlap with existing image frames in the map, the number of new features they contain, or a quality measure of the features in the frame. Image frames not added to the tracking map can be discarded or used to modify the location of features. As another alternative, all or most image frames representing a set of features can be retained, but a subset of these frames can be designated as keyframes for further processing.

[0203] Keyframes can be processed to generate key rigs 704. Keyframes can be processed to generate a 3D set of feature points and saved as a key rig 704. For example, this processing might require comparing image frames simultaneously obtained from two cameras to stereo-determine the 3D positions of feature points. Metadata can be associated with these keyframes and / or key rigs (e.g., pose).

[0204] Environmental maps can be in any of a variety of formats, depending on, for example, the storage location of the environmental map, including, for example, local storage and remote storage of an AR device. For instance, on a memory-constrained wearable device, a map in remote storage may have a higher resolution than a map in local storage. To send a higher-resolution map from remote storage to local storage, the map can be downsampled or otherwise converted to a suitable format, for example, by reducing the number of poses and / or the number of feature points stored for each pose in the physical world stored in the map. In some embodiments, tiles or portions of a high-resolution map from remote storage can be sent to local storage, where the tiles or portions are not downsampled.

[0205] When a new tracking map is created, the database of environment maps can be updated. To determine which of the potentially very large number of environment maps in the database will be updated, the update can include effectively selecting one or more environment maps stored in the database that are relevant to the new tracking map. The selected one or more environment maps can be ranked by relevance, and one or more of the highest-ranking maps can be selected for processing to merge the higher-ranking selected environment maps with the new tracking map to create one or more updated environment maps. When the new tracking map represents a part of the physical world for which there is no pre-existing environment map to update, the tracking map can be stored as a new environment map in the database.

[0206] Watch standalone display

[0207] This document describes methods and apparatus for providing virtual content using an XR system that is independent of the eye's position when viewing the virtual content. Traditionally, virtual content is re-rendered whenever the display system moves. For example, if a user wearing the display system views a virtual representation of a three-dimensional (3D) object on the display and walks around the area where the 3D object appears, the 3D object should be re-rendered for each viewpoint so that the user has the feeling that he or she is walking around an object occupying real space. However, re-rendering consumes significant computational resources of the system and causes artifacts due to latency.

[0208] The inventors have recognized and understood that head posture (e.g., the position and orientation of a user wearing an XR system) can be used to render virtual content independent of eye rotation within the user's head. In some embodiments, a dynamic map of the scene can be generated based on multiple coordinate frames across the real space of one or more sessions, such that virtual content interacting with the dynamic map can be robustly rendered regardless of eye rotation within the user's head and / or sensor deformation caused by heat generated during, for example, during high-speed, computationally intensive operations. In some embodiments, the configuration of the multiple coordinate frames enables a first XR device worn by a first user and a second XR device worn by a second user to identify common locations within the scene. In some embodiments, the configuration of the multiple coordinate frames enables users wearing XR devices to view virtual content from the same location within the scene.

[0209] In some embodiments, a tracking map can be constructed in a world coordinate frame, which may have a world origin. When the XR device is powered on, the world origin can be the XR device's initial pose. The world origin can be aligned with gravity, allowing XR application developers to perform gravity alignment without additional work. Different tracking maps can be constructed in different world coordinate frames because tracking maps can be captured by the same XR device in different sessions and / or by different XR devices worn by different users. In some embodiments, an XR device session can start from device power-on and end when the device is powered off. In some embodiments, the XR device may have a head coordinate frame, which may have a head origin. The head origin can be the XR device's current pose at the time of image capture. The difference between the head pose in the world coordinate frame and the head pose in the head coordinate frame can be used to estimate the tracking path.

[0210] In some embodiments, the XR device may have a camera coordinate frame, which may have a camera origin. The camera origin may be the current pose of one or more sensors of the XR device. The inventors have recognized and understood that the configuration of the camera coordinate frame enables robust display of virtual content independent of eye rotation within the user's head. This configuration also enables robust display of virtual content independent of sensor deformation, for example, due to heat generated during operation.

[0211] In some embodiments, the XR device may have a head unit with a headband that a user can attach to their head, and may include two waveguides, one in front of each of the user's eyes. The waveguides may be transparent, allowing ambient light from real-world objects to pass through them, and the user to see the real-world objects. Each waveguide can send projected light from a projector to the user's corresponding eye. The projected light can form an image on the retina of the eye. Thus, the retina receives both ambient light and projected light. The user can simultaneously see real-world objects and one or more virtual objects created by the projected light. In some embodiments, the XR device may have sensors that detect real-world objects around the user. These sensors may, for example, be cameras that capture images that can be processed to identify the location of real-world objects.

[0212] In some embodiments, instead of attaching virtual content to a world coordinate frame, an XR system can assign a coordinate frame to the virtual content. This configuration allows virtual content to be described without considering where it will be rendered to the user; however, the virtual content can be attached to a more persistent frame location, such as a coordinate frame that will be rendered at a specified location. Figures 14 to 20C The description uses a persistent coordinate frame (PCF). When an object's position changes, the XR device can detect changes in the environment map and determine the motion of the user-worn head unit relative to the real-world object.

[0213] Figure 8 The illustration depicts a user experiencing virtual content rendered by an XR system 10 in a physical environment, according to some embodiments. The XR system may include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is in a physical environment with real objects in the form of a table 16.

[0214] In the illustrated example, the first XR device 12.1 includes a head unit 22, a waist pack 24, and a cable connection 26. A first user 14.1 attaches the head unit 22 to their head and the waist pack 24, located away from the head unit 22, to their waist. The cable connection 26 connects the head unit 22 to the waist pack 24. The head unit 22 includes technologies for displaying one or more virtual objects to the first user 14.1 while allowing the first user 14.1 to see a real object, such as a table 16. The waist pack 24 primarily includes the processing and communication capabilities of the first XR device 12.1. In some embodiments, the processing and communication capabilities may reside wholly or partially in the head unit 22, allowing the waist pack 24 to be removed or located in another device, such as a backpack.

[0215] In the example shown, the waist pack 24 is connected to network 18 via a wireless connection. Server 20 is connected to network 18 and maintains data representing local content. The waist pack 24 downloads data representing local content from server 20 via network 18. The waist pack 24 provides data to head unit 22 via cable connection 26. Head unit 22 may include a display with a light source (e.g., a laser light source or a light-emitting diode (LED) light source) and a waveguide for guiding the light.

[0216] In some embodiments, a first user 14.1 may attach a head unit 22 to their head and a waist pack 24 to their waist. The waist pack 24 may download image data from a server 20 via a network 18. The first user 14.1 may see a table 16 through a display on the head unit 22. A projector forming part of the head unit 22 may receive image data from the waist pack 24 and generate light based on that image data. The light may travel through one or more waveguides forming part of the display on the head unit 22. The light may then leave the waveguides and propagate onto the retina of the first user 14.1's eye. The projector may generate light in the form of a pattern replicated on the retina of the first user 14.1's eye. The light falling on the retina of the first user 14.1's eye may have a selected depth of field, allowing the first user 14.1 to perceive an image at a preselected depth beyond the waveguides. Additionally, the first user 14.1's two eyes may receive slightly different images, allowing the first user 14.1's brain to perceive one or more three-dimensional images at a selected distance from the head unit 22. In the example shown, the first user 14.1 perceives virtual content 28 above table 16. The scale of the virtual content 28 and its position and distance from the first user 14.1 are determined by data representing the virtual content 28 and various coordinate frames used to display the virtual content 28 to the first user 14.1.

[0217] In the example shown, the virtual content 28 is invisible from the perspective of the accompanying drawing, but is visible to the first user 14.1 using the first XR device 12.1. The virtual content 28 can initially reside as a data structure within the visual data and algorithms in the waist pack 24. Then, when the projector of the head unit 22 generates light based on the data structure, the data structure can represent itself as light. It should be understood that although the virtual content 28 does not exist in the three-dimensional space in front of the first user 14.1, the virtual content 28 is still represented in three-dimensional space. Figure 1 The description illustrates the wearer's perception of head unit 22. Visualizations of computer data in three-dimensional space can be used in this description to show how data structures perceived by one or more users, contributing to rendering, are related to each other within the data structures of the waist pack 24.

[0218] Figure 9Components of a first XR device 12.1 according to some embodiments are shown. The first XR device 12.1 may include a head unit 22, and various components forming part of visual data and algorithms, including, for example, a rendering engine 30, various coordinate systems 32, various origin and destination coordinate frames 34, and various origin-to-destination coordinate frame transformers 36. The various coordinate systems may be based on the intrinsic properties of the XR device, or may be determined by referring to other information, such as persistent poses or persistent coordinate systems described herein.

[0219] The head unit 22 may include a head-mounted frame 40, a display system 42, a real object detection camera 44, a motion tracking camera 46, and an inertial measurement unit 48.

[0220] The headband 40 can be fixed to Figure 8 The shape of the head of the first user 14.1. The display system 42, the real object detection camera 44, the motion tracking camera 46 and the inertial measurement unit 48 can be mounted to the head-mounted frame 40 and thus move with the head-mounted frame 40.

[0221] The coordinate system 32 may include a local data system 52, a world frame system 54, a head frame system 56, and a camera frame system 58.

[0222] The local data system 52 may include a data channel 62, a local frame determination routine 64, and a local frame storage instruction 66. The data channel 62 may be an internal software routine, a hardware component such as an external cable or radio frequency receiver, or a hybrid component such as an open port. The data channel 62 may be configured to receive image data 68 representing virtual content.

[0223] The local frame determination routine 64 can be connected to the data channel 62. The local frame determination routine 64 can be configured to determine the local coordinate frame 70. In some embodiments, the local frame determination routine can determine the local coordinate frame based on a real-world object or a real-world location. In some embodiments, the local coordinate frame can be based on the top edge relative to the bottom edge of the browser window, the head or foot of a character, a node on the outer surface of a prism or bounding box surrounding virtual content, or any other suitable location of the coordinate frame that defines the orientation of the virtual content and the location where the virtual content is placed (e.g., a node, such as a placement node or a PCF node).

[0224] The local frame storage instruction 66 can be connected to the local frame determination routine 64. Those skilled in the art will understand that software modules and routines are "connected" to each other through subroutines, calls, etc. The local frame storage instruction 66 can store the local coordinate frame 70 as a local coordinate frame 72 within the origin and destination coordinate frames 34. In some embodiments, the origin and destination coordinate frames 34 can be one or more coordinate frames that can be manipulated or transformed to persist virtual content between sessions. In some embodiments, a session can be a time period between the startup and shutdown of an XR device. Two sessions can be two startup and shutdown time periods of a single XR device, or startup and shutdown time periods of two different XR devices.

[0225] In some embodiments, the origin and destination coordinate frames 34 may be coordinate frames involved in one or more transformations required to enable the XR devices of the first user and the second user to recognize a common location. In some embodiments, the destination coordinate frame may be the output of a series of calculations and transformations applied to the target coordinate frame so that the first and second users view virtual content in the same location.

[0226] The rendering engine 30 can be connected to the data channel 62. The rendering engine 30 can receive image data 68 from the data channel 62, so that the rendering engine 30 can render virtual content at least in part based on the image data 68.

[0227] Display system 42 can be connected to rendering engine 30. Display system 42 may include components that transform image data 68 into visible light. Visible light can form two patterns, one for each eye. Visible light can enter... Figure 8 The first user 14.1's eye, and can be detected on the retina of the first user 14.1's eye.

[0228] The real object detection camera 44 may include one or more cameras that can capture images from different sides of the head-mounted frame 40. The motion tracking camera 46 may include one or more cameras that can capture images from the sides of the head-mounted frame 40. One or more cameras may be used instead of two sets of one or more cameras representing the real object detection camera 44 and the motion tracking camera 46. In some embodiments, cameras 44, 46 may capture images. As described above, these cameras may collect data for constructing a tracking map.

[0229] The inertial measurement unit 48 may include multiple devices for detecting the motion of the head unit 22. The inertial measurement unit 48 may include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of the inertial measurement unit 48 track the motion of the head unit 22 in at least three orthogonal directions and about at least three orthogonal axes in combination.

[0230] In the example shown, the world frame system 54 includes a world surface determination routine 78, a world frame determination routine 80, and a world frame storage instruction 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 accepts images and / or keyframes based on images captured by the real object detection camera 44 and processes the images to identify surfaces within them. A depth sensor (not shown) can determine the distance to the surfaces. Therefore, these surfaces are represented by data in three dimensions, including their size, shape, and distance from the real object detection camera.

[0231] In some embodiments, the world coordinate frame 84 may be based on the origin at the time the head pose session is initialized. In some embodiments, the world coordinate frame may be located at the location where the device is started, or it may be located at a new location if the head pose is lost during the start-up session. In some embodiments, the world coordinate frame may be the origin at the start of the head pose session.

[0232] In the example shown, world frame determination routine 80 is connected to world surface determination routine 78, and determines world coordinate frame 84 based on the position of the surface determined by world surface determination routine 78. World frame storage instruction 82 is connected to world frame determination routine 80 to receive world coordinate frame 84 from world frame determination routine 80. World frame storage instruction 82 stores world coordinate frame 84 as world coordinate frame 86 within origin and destination coordinate frame 34.

[0233] The headframe system 56 may include a headframe determination routine 90 and a headframe storage instruction 92. The headframe determination routine 90 may be connected to a motion tracking camera 46 and an inertial measurement unit 48. The headframe determination routine 90 may use data from the motion tracking camera 46 and the inertial measurement unit 48 to calculate a head coordinate frame 94. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The motion tracking camera 46 may continuously capture images used by the headframe determination routine 90 to refine the head coordinate frame 94. Figure 8 When the first user 14.1 moves their head, the head unit 22 moves. The motion tracking camera 46 and the inertial measurement unit 48 can continuously provide data to the head frame determination routine 90, so that the head frame determination routine 90 can update the head coordinate frame 94.

[0234] Headframe storage instruction 92 can be connected to headframe determination routine 90 to receive head coordinate frame 94 from headframe determination routine 90. Headframe storage instruction 92 can store head coordinate frame 94 as head coordinate frame 96 in origin and destination coordinate frames 34. Headframe storage instruction 92 can also repeatedly store the updated head coordinate frame 94 as head coordinate frame 96 when headframe determination routine 90 recalculates head coordinate frame 94. In some embodiments, head coordinate frame can be the position of wearable XR device 12.1 relative to local coordinate frame 72.

[0235] The camera frame system 58 may include camera intrinsic characteristics 98. Camera intrinsic characteristics 98 may include the dimensions of the head unit 22, which is a design and manufacturing feature. Camera intrinsic characteristics 98 can be used to calculate the camera coordinate frame 100 stored within the origin and destination coordinate frames 34.

[0236] In some embodiments, the camera coordinate frame 100 may include Figure 8 The positions of all pupils of the left eye of the first user 14.1 are shown. The pupil position of the left eye is located within the camera coordinate frame 100 as the left eye moves from left to right or up and down. Additionally, the pupil position of the right eye is located within the camera coordinate frame 100 for the right eye. In some embodiments, the camera coordinate frame 100 may include the position of the camera relative to a local coordinate frame when an image is captured.

[0237] The origin-to-destination coordinate frame transformer 36 may include a local-to-world coordinate transformer 104, a world-to-head coordinate transformer 106, and a head-to-camera coordinate transformer 108. The local-to-world coordinate transformer 104 can receive a local coordinate frame 72 and transform it into a world coordinate frame 86. The transformation from local coordinate frame 72 to world coordinate frame 86 can be represented as a transformation within world coordinate frame 86 into a local coordinate frame of world coordinate frame 110.

[0238] The world-to-head coordinate transformer 106 can transform from world coordinate frame 86 to head coordinate frame 96. The world-to-head coordinate transformer 106 can also transform a local coordinate frame transformed to world coordinate frame 110 back to head coordinate frame 96. This transformation can be represented as a transformation within head coordinate frame 96 to a local coordinate frame transformed to head coordinate frame 112.

[0239] The head-to-camera coordinate transformer 108 can transform from head coordinate frame 96 to camera coordinate frame 100. The head-to-camera coordinate transformer 108 can also transform the local coordinate frame, transformed to head coordinate frame 112, into a local coordinate frame, transformed to camera coordinate frame 114, within camera coordinate frame 100. This local coordinate frame, transformed to camera coordinate frame 114, can be input into the rendering engine 30. The rendering engine 30 can then render image data 68 representing local content 28 based on this local coordinate frame, transformed to camera coordinate frame 114.

[0240] Figure 10 This is a spatial representation of various origin and destination coordinate frames 34. The figure shows the local coordinate frame 72, world coordinate frame 86, head coordinate frame 96, and camera coordinate frame 100. In some embodiments, when virtual content is placed in the real world so that a user can view it, the local coordinate frame associated with the XR content 28 may have a position and rotation relative to the local and / or world coordinate frames and / or PCF (e.g., node and facing direction may be provided). Each camera may have its own camera coordinate frame 100 containing all pupil positions of one eye. Reference numerals 104A and 106A respectively indicate... Figure 9 The transformations performed by the local-to-world coordinate transformer 104, the world-to-head coordinate transformer 106, and the head-to-camera coordinate transformer 108 are described.

[0241] Figure 11 A camera rendering protocol for transforming from a head coordinate frame to a camera coordinate frame, according to some embodiments, is described. In the example shown, the pupil of a single eye moves from position A to position B. The virtual object to be displayed as stationary will depend on the position of the pupil projected onto the depth plane of either position A or B (assuming the camera is configured to use a pupil-based coordinate frame). As a result, using a pupil coordinate frame transformed to a head coordinate frame when the eye moves from position A to position B will cause jitter in the stationary virtual object. This situation is called view-dependent display or projection.

[0242] like Figure 12 As shown, the camera coordinate frame (e.g., CR) is positioned and includes all pupil positions, and the object projection will now be consistent regardless of pupil positions A and B. The head coordinate frame is transformed into a CR frame, which is referred to as view-independent display or projection. The image can be reprojected onto the virtual content to accommodate changes in eye position; however, since the rendering remains in the same location, jitter can be minimized.

[0243] Figure 13The display system 42 is shown in more detail. The display system 42 includes a stereo analyzer 144, which is connected to the rendering engine 30 and forms part of the visual data and algorithms.

[0244] The display system 42 further includes a left projector 166A and a right projector 166B, as well as a left waveguide 170A and a right waveguide 170B. The left projector 166A and the right projector 166B are connected to a power source. Each projector 166A and 166B has a corresponding input for image data to be provided to the respective projector 166A or 166B. When powered on, the respective projector 166A or 166B generates and emits light from a two-dimensional pattern. The left waveguide 170A and the right waveguide 170B are positioned to receive light from the left projector 166A and the right projector 166B, respectively. The left waveguide 170A and the right waveguide 170B are transparent waveguides.

[0245] In use, the user attaches the headband 40 to their head. Components of the headband 40 may include, for example, a strap (not shown) wrapped around the back of the user's head. The left waveguide 170A and the right waveguide 170B are then positioned in front of the user's left eye 220A and right eye 220B.

[0246] The rendering engine 30 inputs the received image data into the stereo analyzer 144. This image data is... Figure 8 The three-dimensional image data of the local content 28 is projected onto multiple virtual planes. A stereo analyzer 144 analyzes the image data to determine a left image dataset and a right image dataset based on the image data used for projection onto each depth plane. The left and right image datasets are datasets representing two-dimensional images projected in three dimensions to give the user a sense of depth.

[0247] The stereo analyzer 144 inputs the left and right image datasets to the left projector 166A and the right projector 166B. The left and right projectors 166A and 166B then create left and right illumination patterns. The components of the display system 42 are shown in plan view; however, it should be understood that when shown in a front view, the left and right patterns are two-dimensional patterns. Each light pattern comprises multiple pixels. For illustrative purposes, light rays 224A and 226A from two pixels are shown leaving the left projector 166A and entering the left waveguide 170A. Light rays 224A and 226A are reflected from the sides of the left waveguide 170A. Light rays 224A and 226A are shown propagating from left to right within the left waveguide 170A via internal reflection; however, it should be understood that light rays 224A and 226A also propagate in a certain direction into the paper using a refraction and reflection system.

[0248] Light rays 224A and 226A exit the left waveguide 170A through the pupil 228A and then enter the left eye 220A through the pupil 230A. Light rays 224A and 226A then fall onto the retina 232A of the left eye 220A. In this way, the left light pattern falls onto the retina 232A of the left eye 220A. The user perceives the pixels formed on the retina 232A as pixels 234A and 236A at a certain distance on the side of the left waveguide 170A opposite the left eye 220A. Depth perception is created by manipulating the focal length of the light.

[0249] In a similar manner, stereo analyzer 144 inputs the right image dataset into right projector 166B. Right projector 166B transmits a right light pattern, represented by pixels in the form of rays 224B and 226B. Rays 224B and 226B are reflected within right waveguide 170B and exit through pupil 228B. Rays 224B and 226B then enter through pupil 230B of right eye 220B and fall on retina 232B of right eye 220B. The pixels of rays 224B and 226B are perceived as pixels 134B and 236B behind right waveguide 170B.

[0250] The patterns created on retinas 232A and 232B are perceived as left and right images, respectively. Due to the function of stereo analyzer 144, the left and right images are slightly different from each other. The left and right images are perceived as a three-dimensional rendering in the user's mind.

[0251] As mentioned, the left waveguide 170A and the right waveguide 170B are transparent. Light from a real object such as the left waveguide 170A and the right waveguide 170B on the side of the table 16 opposite to the eyes 220A and 220B can be projected through the left waveguide 170A and the right waveguide 170B and fall on the retinas 232A and 232B.

[0252] Persistent Coordinate Frame (PCF)

[0253] This document describes methods and apparatus for providing spatial persistence between user instances within a shared space. Without spatial persistence, virtual content placed by a user in the physical world during a session may not exist or may be misplaced in the user's view in different sessions. Without spatial persistence, virtual content placed by one user in the physical world may not exist or may be misplaced in the view of a second user, even if the second user intends to share the same physical space experience as the first user.

[0254] The inventors have recognized and understood that spatial persistence can be provided through a persistent coordinate frame (PCF). A PCF can be defined based on one or more points that represent features identified in the physical world (e.g., corners, edges). Features can be selected such that they appear identical from one user instance to another in an XR system.

[0255] Furthermore, when rendered relative to a local map based solely on the tracking map, drift during tracking that causes the calculated tracking path (e.g., camera trajectory) to deviate from the actual tracking path can lead to misalignment of the virtual content. As the XR device gathers more information about the scene over time, the spatial tracking map can be refined to correct this drift. However, if the virtual content is placed on top of a real object and saved relative to the device's world coordinate frame derived from the tracking map before map refinement, the virtual content may be displaced, just as if the real object had moved during map refinement. The PCF can be updated based on map refinement because the PCF is feature-based and updated as features move during map refinement.

[0256] PCFs can include six degrees of freedom for translation and rotation relative to the map coordinate system. PCFs can be stored on local and / or remote storage media. Depending on, for example, the storage location, the translation and rotation of the PCF can be calculated relative to the map coordinate system. For example, a PCF used locally on a device might have translation and rotation relative to the device's world coordinate frame. A PCF in the cloud might have translation and rotation relative to the canonical coordinate frame of a canonical map.

[0257] PCFs can provide a sparse representation of the physical world, offering less information about it than all available information, making them efficient to process and transfer. Techniques for processing persistent spatial information may include creating dynamic maps based on one or more coordinate systems in real space spanning one or more sessions, generating persistent coordinate frames (PCFs) on the sparse maps, which can be exposed to XR applications, for example, through application programming interfaces (APIs).

[0258] Figure 14 This is an additional block diagram illustrating the creation of a Persistent Coordinate Frame (PCF) according to some embodiments and the mapping of XR content to the PCF. Each block may represent digital information stored in computer memory. In the case of application 1180, the data may represent computer-executable instructions. In the case of virtual content 1170, the digital information may define, for example, a virtual object specified by application 1180. In the cases of other blocks, the digital information may characterize certain aspects of the physical world.

[0259] In the illustrated embodiment, one or more PCFs are created based on images captured by sensors on the wearable device. Figure 14 In this embodiment, the sensor is a visual image camera. These cameras can be the same cameras used to form the tracking map. Therefore, by Figure 14 Some of the suggested processes can be performed as part of updating the tracking map. However, Figure 14 This demonstrates that in addition to tracking maps, information that provides persistence is also generated.

[0260] To export the 3D PCF, two images 1110 from two cameras mounted on a wearable device in a configuration capable of stereo image analysis are processed together. Figure 14 Images 1 and 2 are shown, each originating from one of the cameras. For simplicity, only a single image from each camera is shown. However, each camera can output a stream of image frames, and multiple image frames in the stream can be processed. Figure 14 The processing.

[0261] Therefore, image 1 and image 2 can each be a frame in an image frame sequence. Repeating consecutive image frames in the sequence is possible. Figure 14 The process shown continues until an image frame containing feature points provides a suitable image from which persistent spatial information is formed. Alternatively or additionally, this process can be repeated when user movement causes the user to no longer be sufficiently close to the previously identified PCF to reliably determine their position relative to the physical world. Figure 14 This involves processing. For example, an XR system can maintain the current PCF (Position of Facing) for the user. When that distance exceeds a threshold, the system can switch to a new current PCF that is closer to the user, which can be determined based on... Figure 14 The process uses image frames acquired at the user's current location to generate the image.

[0262] Even when generating a single PCF, it is possible to process a stream of image frames to identify image frames that describe content in the physical world, which may be stable and easily identifiable by devices near the physical world region depicted in the image frame. Figure 14 In this embodiment, the process begins with the identification of feature 1120 in the image. For example, a feature can be identified by finding locations in the image where the gradient exceeds a threshold or other characteristic, such as a corner of an object. In the illustrated embodiment, the feature is a point, but other identifiable features, such as edges, can be used alternatively or additionally.

[0263] In the illustrated embodiment, a fixed number N of features 1120 are selected for further processing. These feature points can be selected based on one or more criteria, such as the magnitude of the gradient or proximity to other feature points. Alternatively or additionally, feature points can be selected tentatively, for example, based on properties that suggest the feature point is persistent. For example, a heuristic can be defined based on the properties of feature points that might correspond to a window, door, or corner of a large piece of furniture. Such a heuristic may take into account the feature point itself and its surroundings. As a specific example, the number of feature points per image can be between 100 and 500 or between 150 and 250, for example, 200.

[0264] Regardless of the number of feature points selected, descriptors 1130 can be computed for each feature point. In this example, a descriptor is computed for each selected feature point, but descriptors can be computed for a group of feature points, a subset of feature points, or all features within an image. Descriptors characterize feature points so that feature points representing the same object in the physical world are assigned similar descriptors. Descriptors can enable the alignment of two frames, which may occur, for example, when locating one map relative to another. Instead of searching for the relative orientation of frames that minimizes the distance between feature points in two images, initial alignment of two frames can be performed by identifying feature points with similar descriptors. Image frame alignment can be based on alignment points with similar descriptors, which may require less processing compared to computed alignment of all feature points in an image.

[0265] A descriptor can be computed as a mapping from feature points to descriptors, or in some embodiments, as a mapping from patches of the image surrounding the feature points to descriptors. The descriptor can be a numerical quantity. U.S. Patent Application 16 / 190,948 describes a computed descriptor for a feature point, and its entirety is incorporated herein by reference.

[0266] exist Figure 14 In the example, a descriptor 1130 is calculated for each feature point in each image frame. Based on the descriptor and / or feature points and / or the image itself, the image frame can be identified as a keyframe 1140. In the illustrated embodiment, a keyframe is an image frame that meets a certain criterion, and that image frame is then selected for further processing. For example, when creating a tracking map, image frames for which meaningful information can be added to the map can be selected as keyframes to be integrated into the map. On the other hand, image frames that substantially overlap with areas that have already been integrated into the map can be discarded, so that they do not become keyframes. Alternatively or additionally, keyframes can be selected based on the number and / or type of feature points in the image frame. Figure 14In some embodiments, the keyframes selected to be included in the tracking map can also be considered as keyframes for determining the PCF, but different or additional criteria can be used to select the keyframes for generating the PCF.

[0267] although Figure 14 The diagram shows that keyframes are used for further processing, but information obtained from the images can be processed in other ways. For example, feature points in key assemblies can be processed alternatively or additionally. Furthermore, although keyframes are described as being derived from a single image frame, there does not necessarily have to be a one-to-one relationship between the keyframe and the acquired image frame. For example, keyframes can be obtained from multiple image frames, such as by stitching or aggregating the image frames together, so that only features appearing in multiple images are retained in the keyframe.

[0268] Keyframes may include image information and / or metadata associated with the image information. In some embodiments, keyframes may be generated by cameras 44, 46 ( Figure 9 The captured image is computed into one or more keyframes (e.g., keyframe 1, 2). In some embodiments, a keyframe may include a camera pose. In some embodiments, a keyframe may include one or more camera images captured in a camera pose. In some embodiments, the XR system may determine that a portion of the camera images captured in a camera pose is useless and therefore not include that portion in the keyframe. Therefore, using keyframes to align new images with early knowledge of the scene can reduce the use of XR system computational resources. In some embodiments, a keyframe may include images and / or image data at locations with orientation / angle. In some embodiments, a keyframe may include the location and orientation of one or more map points that can be observed. In some embodiments, a keyframe may include a coordinate frame with an ID. Keyframes are described in U.S. Patent Application No. 15 / 877,359, the entire contents of which are incorporated herein by reference.

[0269] Some or all of the keyframes 1140 can be selected for further processing, such as generating a persistent pose 1150 for the keyframes. This selection can be based on the characteristics of all feature points or a subset thereof in the image frame. These characteristics can be determined based on processing of descriptors, features, and / or the image frame itself. As a specific example, the selection can be based on clustering of feature points identified as potentially related to a persistent object.

[0270] Each keyframe is associated with the pose of the camera that acquired that keyframe. For keyframes selected for processing into a persistent pose, this pose information may be stored along with other metadata about the keyframe, such as a WiFi fingerprint and / or GPS coordinates at the time of acquisition and / or at the acquisition location. In some embodiments, metadata such as GPS coordinates may be used, alone or in combination, as part of the localization process.

[0271] A persistent pose is a source of information from which a device can orient itself relative to previously acquired information about the physical world. For example, if keyframes from which a persistent pose is created are incorporated into a map of the physical world, the device can orient itself relative to that persistent pose using a sufficient number of feature points in the keyframes associated with it. The device can align its current image of the surrounding environment with the persistent pose. This alignment can be based on matching the current image with the image 1110, feature 1120, and / or descriptor 1130 that caused the persistent pose, or any subset of that image or those features or descriptors. In some embodiments, the current image frame that matches the persistent pose can be another keyframe that has been incorporated into the device's tracking map.

[0272] Information about persistent poses can be stored in a format that facilitates sharing among multiple applications that can run on the same or different devices. Figure 14 In the example, some or all of the persistent poses can be represented as a persistent coordinate frame (PCF) 1160. Like the persistent pose, the PCF can be associated with a map and can include a set of features or other information that a device can use to determine its orientation relative to that PCF. The PCF may include a transformation that defines a transformation relative to the origin of its map, such that by associating its position with the PCF, the device can determine its position relative to any object in the physical world reflected in the map.

[0273] Because PCFs provide a mechanism for determining position relative to physical objects, applications (e.g., application 1180) can define the positions of virtual objects relative to one or more PCFs, which serve as anchor points for virtual content 1170. For example, Figure 14 It is shown that App 1 has associated its virtual content 2 with PCF 1.2. Similarly, App 2 has associated its virtual content 3 with PCF 1.2. It is also shown that App 1 has associated its virtual content 1 with PCF 4.5, and App 2 has associated its virtual content 4 with PCF 3. In some embodiments, PCF 3 may be based on image 3 (not shown), and PCF 4.5 may be based on images 4 and 5 (not shown), similar to how PCF 1.2 is based on images 1 and 2. When rendering this virtual content, the device may apply one or more transformations to calculate information such as the position of the virtual content relative to the device's display and / or the position of physical objects relative to the desired position of the virtual content. Using PCF as a reference simplifies such calculations.

[0274] In some embodiments, a persistent pose can be a coordinate position and / or orientation with one or more associated keyframes. In some embodiments, a persistent pose can be automatically created after the user has traveled a certain distance (e.g., three meters). In some embodiments, a persistent pose can be used as a reference point during positioning. In some embodiments, a persistent pose can be stored in a traversable world (e.g., traversable world module 538).

[0275] In some embodiments, a new PCF can be determined based on a predetermined distance allowed between adjacent PCFs. In some embodiments, one or more persistent poses can be calculated into the PCF as the user travels a predetermined distance (e.g., five meters). In some embodiments, the PCF can be associated with one or more world coordinate frames and / or canonical coordinate frames, for example, in a traversable world. In some embodiments, depending on, for example, security settings, the PCF can be stored in a local database and / or a remote database.

[0276] Figure 15 A method 4700 for establishing and using a persistent coordinate frame is illustrated according to some embodiments. Method 4700 may begin by capturing (action 4702) images of a scene using one or more sensors of an XR device (e.g., Figure 14 (Image 1 and Image 2 in the image). Multiple cameras can be used, and one camera can generate multiple images, for example, in the form of a stream.

[0277] Method 4700 may include extracting (4704) points of interest from the captured image (e.g., Figure 7 Map point 702, Figure 14 Feature 1120 in the middle), generate descriptors of the interest points extracted by (action 4706) (e.g., Figure 14 The method uses descriptor 1130 in the model and generates keyframes (e.g., keyframe 1140) based on the descriptor (action 4708). In some embodiments, the method can compare points of interest in the keyframes and form keyframe pairs that share a predetermined number of points of interest. The method can use the individual keyframe pairs to reconstruct a portion of the physical world. The mapped portion of the physical world can be saved as 3D features (e.g., ...). Figure 7(Key assembly 704 in the text). In some embodiments, selected portions of a keyframe pair can be used to construct 3D features. In some embodiments, the results of the mapping can be selectively saved. Keyframes not used to construct 3D features can be associated with 3D features through poses, for example, by using the covariance matrix between the poses of the keyframes to represent the distance between the keyframes. In some embodiments, keyframe pairs can be selected to construct 3D features such that the distance between any two constructed 3D features is within a predetermined distance, which can be determined to balance the required computational cost and the accuracy level of the resulting model. Such a method enables XR systems to provide models of the physical world with a data volume suitable for efficient and accurate computation. In some embodiments, the covariance matrix of two images can include the covariance between the poses (e.g., six degrees of freedom) of the two images.

[0278] Method 4700 may include generating a persistent pose based on keyframes (action 4710). In some embodiments, the method may include generating a persistent pose based on 3D features reconstructed from keyframe pairs. In some embodiments, the persistent pose may be attached to the 3D features. In some embodiments, the persistent pose may include the pose of the keyframes used to construct the 3D features. In some embodiments, the persistent pose may include the average pose of the keyframes used to construct the 3D features. In some embodiments, persistent poses may be generated such that the distance between adjacent persistent poses is within a predetermined value, such as the range of one to five meters, any value in between, or any other suitable value. In some embodiments, the distance between adjacent persistent poses may be represented by the covariance matrix of adjacent persistent poses.

[0279] Method 4700 may include generating a PCF (action 4712) based on a persistent pose. In some embodiments, the PCF may be attached to a 3D feature. In some embodiments, the PCF may be associated with one or more persistent poses. In some embodiments, the PCF may include the pose of one of the associated persistent poses. In some embodiments, the PCF may include the average pose of the poses of the associated persistent poses. In some embodiments, the PCF may be generated such that the distance between adjacent PCFs is within a predetermined value, such as in the range of three to ten meters, any value between these two values, or any other suitable value. In some embodiments, the distance between adjacent PCFs may be represented by the covariance matrix of adjacent PCFs. In some embodiments, the PCF may be exposed to an XR application via, for example, an application programming interface (API), allowing the XR application to access a model of the physical world through the PCF without accessing the model itself.

[0280] Method 4700 may include associating image data of a virtual object to be displayed by an XR device with at least one of the PCFs (action 4714). In some embodiments, the method may include calculating the translation and orientation of the virtual object relative to the associated PCF. It should be understood that it is not necessary to associate the virtual object with a PCF generated by the device placing the virtual object. For example, the device may acquire a saved PCF in a canonical map in the cloud and associate the virtual object with the acquired PCF. It should be understood that the virtual object may move along with the associated PCF as the PCF is adjusted over time.

[0281] Figure 16 Visual data and algorithms of a first XR device 12.1, a second XR device 12.2, and a server 20 according to some embodiments are shown. Figure 16 The components shown can be operated to perform some or all of the operations associated with generating, updating, and / or using spatial information (such as persistent pose, persistent coordinate frame, tracking map, or canonical map) as described herein. Although not shown, the first XR device 12.1 can be configured to be identical to the second XR device 12.2. Server 20 may have map storage routine 118, canonical map 120, map transmitter 122, and map merging algorithm 124.

[0282] A second XR device 12.2, which can be in the same scene as the first XR device 12.1, may include a Permanent Coordinate Frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that can be used to render virtual objects, and a frame embedding generator 308 (see FIG. 21). In some embodiments, a map download system 126, a PCF recognition system 128, and a map embedding generator may be included. Figure 2 The positioning module 130, the standard map merger 132, the standard map 133, and the map publisher 136 are combined into a navigable world unit 1304. The PCF integration unit 1300 can be connected to the navigable world unit 1304 and other components of the second XR device 12.2 to allow the acquisition, generation, use, uploading, and downloading of PCF.

[0283] Maps incorporating PCFs can achieve greater persistence in a changing world. In some embodiments, the localization of a tracking map, which includes matching features such as images, can include features representing persistent content selected from a map composed of PCFs, enabling fast matching and / or localization. For example, in a world where people enter and exit scenes and objects such as doors move relative to the scene, less storage space and transfer rates are required, and scenes can be mapped using individual PCFs and their relationships with each other (e.g., an integrated constellation of PCFs).

[0284] In some embodiments, the PCF integration unit 1300 may include PCF 1306, PCF tracker 1308, persistent pose acquirer 1310, PCF checker 1312, PCF generation system 1314, coordinate frame calculator 1316, persistent pose calculator 1318, and three converters including a tracking map and persistent pose converter 1320, a persistent pose and PCF converter 1322, and a PCF and image data converter 1324, all stored in data previously stored in the storage unit of the second XR device 12.2.

[0285] In some embodiments, PCF tracker 1308 may have open and close prompts selectable by application 1302. Application 1302 may be executable by the processor of a second XR device 12.2 for example, to display virtual content. Application 1302 may have a call to open PCF tracker 1308 via the open prompt. When PCF tracker 1308 is open, PCF tracker 1308 may generate PCFs. Application 1302 may have subsequent calls that can close PCF tracker 1308 via the close prompt. When PCF tracker 1308 is closed, PCF tracker 1308 terminates PCF generation.

[0286] In some embodiments, server 20 may include a plurality of persistent poses 1332 and a plurality of PCFs 1330 previously saved in association with specification map 120. Map transmitter 122 may send specification map 120 together with persistent poses 1332 and / or PCFs 1330 to second XR device 12.2. Persistent poses 1332 and PCFs 1330 may be stored on second XR device 12.2 in association with specification map 133. Figure 2 When located on standard map 133, it can be compared with the ground Figure 2 Persistent pose 1332 and PCF 1330 are stored in association.

[0287] In some embodiments, the persistent pose acquirer 1310 can acquire the ground Figure 2 The persistent pose is obtained. The PCF inspector 1312 can be connected to the persistent pose acquirer 1310. The PCF inspector 1312 can obtain PCFs from PCFs 1306 based on the persistent poses obtained by the persistent pose acquirer 1310. The PCFs obtained by the PCF inspector 1312 can form an initial group of PCFs for PCF-based image display.

[0288] In some embodiments, application 1302 may need to generate additional PCFs. For example, if a user moves to an area that was not previously mapped, application 1302 may activate PCF tracker 1308. PCF generation system 1314 may connect to PCF tracker 1308 and, as the map is generated... Figure 2 Start expanding and start being based on land Figure 2 PCF generation. The PCF generated by PCF generation system 1314 can form a second set of PCFs, which can be used for PCF-based image display.

[0289] The coordinate frame calculator 1316 can be connected to the PCF checker 1312. After the PCF checker 1312 obtains the PCF, the coordinate frame calculator 1316 can invoke the head coordinate frame 96 to determine the head pose of the second XR device 12.2. The coordinate frame calculator 1316 can also invoke the persistent pose calculator 1318. The persistent pose calculator 1318 can be directly or indirectly connected to the frame embedding generator 308. In some embodiments, an image / frame can be specified as a keyframe after traveling at a threshold distance (e.g., 3 meters) from a previous keyframe. The persistent pose calculator 1318 can generate a persistent pose based on multiple (e.g., three) keyframes. In some embodiments, the persistent pose can be essentially the average of the coordinate frames of multiple keyframes.

[0290] The tracking map and persistent pose converter 1320 can be connected to the ground. Figure 2 And the persistent pose calculator 1318. The tracking map and persistent pose converter 1320 can track the ground... Figure 2 Transform into a persistent posture to determine relative to the ground Figure 2 The persistent posture at the origin.

[0291] The persistent pose and PCF transformer 1322 can be connected to the tracking map and persistent pose transformer 1320, and further connected to the PCF inspector 1312 and the PCF generation system 1314. The persistent pose and PCF transformer 1322 can transform the persistent pose (to which the tracking map has been transformed) from the PCF inspector 1312 and the PCF generation system 1314 into a PCF to determine the PCF relative to the persistent pose.

[0292] PCF and image data transformer 1324 can be connected to persistent pose and PCF transformer 1322 and data channel 62. PCF and image data transformer 1324 transforms PCF into image data 68. Rendering engine 30 can be connected to PCF and image data transformer 1324 to display image data 68 to the user relative to PCF.

[0293] PCF integration unit 1300 can store additional PCFs generated using PCF generation system 1314 within PCF 1306. PCF 1306 can be stored relative to persistent pose storage. When map publisher 136 sends a map to server 20... Figure 2At that time, map publisher 136 can obtain PCF 1306 and the persistent gesture associated with PCF 1306, and map publisher 136 also sends data related to the map to server 20. Figure 2 Associated PCF and persistent pose. When server 20's map storage routine 118 stores the location... Figure 2 At the same time, map storage routine 118 can also store persistent poses and PCFs generated by the second viewing device 12.2. Map merging algorithm 124 can use maps associated with the canonical map 120 and stored in persistent poses 1332 and PCFs 1330 respectively. Figure 2 Persistent poses and PCF are used to create a canonical map 120.

[0294] The first XR device 12.1 may include a PCF integration unit similar to the PCF integration unit 1300 of the second XR device 12.2. When the map transmitter 122 sends a specification map 120 to the first XR device 12.1, the map transmitter 122 may send a persistent pose 1332 and a PCF 1330 associated with the specification map 120 and originating from the second XR device 12.1. The first XR device 12.1 may store the PCF and persistent pose in data storage on its storage device. The first XR device 12.1 may then utilize the persistent pose and PCF originating from the second XR device 12.2 for image display relative to the PCF. Alternatively or additionally, the first XR device 12.1 may acquire, generate, use, upload, and download the PCF and persistent pose in a manner similar to that of the second XR device 12.2 as described above.

[0295] In the example shown, the first XR device 12.1 generates a local tracking map (hereinafter referred to as "the map"). Figure 1 ", and the map storage routine 118 receives the map from the first XR device 12.1 Figure 1 Then, map storage routine 118 will store the map... Figure 1 The standard map 120 is stored on the storage device of server 20.

[0296] The second XR device 12.2 includes a map download system 126, an anchor point recognition system 128, a positioning module 130, a standard map merger 132, a local content positioning system 134, and a map publisher 136.

[0297] In use, map transmitter 122 sends standard map 120 to second XR device 12.2, and map download system 126 downloads and stores standard map 120 from server 20 as standard map 133.

[0298] Anchor point identification system 128 is connected to world surface determination routine 78. Anchor point identification system 128 identifies anchor points based on objects detected by world surface determination routine 78. Anchor point identification system 128 uses the anchor points to generate a second map (map). Figure 2 As shown in loop 138, the anchor point identification system 128 continues to identify anchor points and continue to update the location. Figure 2 Based on the data provided by the world surface determination routine 78, the positions of the anchor points are recorded as three-dimensional data. The world surface determination routine 78 receives images from the real object detection camera 44 and depth data from the depth sensor 135 to determine the positions of the surfaces and their relative distances to the depth sensor 135.

[0299] Positioning module 130 connects to standard map 133 and ground Figure 2 The positioning module 130 repeatedly attempts to locate the ground. Figure 2 Locate the standard map 133. The standard map merger 132 connects to the standard map 133 and the map. Figure 2 When the positioning module 130 will... Figure 2 When the standardized map 133 is located, the standardized map merger 132 merges the standardized map 133 into the map. Figure 2 The anchor points are then used to update the map using the missing data included in the canonical map. Figure 2 .

[0300] Local content location system 134 connects to the ground Figure 2 The local content positioning system 134 could be, for example, a system where a user can locate local content at a specific location within a world coordinate frame. The local content then attaches itself to the local coordinate system. Figure 2 An anchor point. The local-to-world coordinate transformer 104 transforms the local coordinate frame to the world coordinate frame based on the settings of the local content positioning system 134. (Already referenced...) Figure 2 The functions of the rendering engine 30, the display system 42, and the data channel 62 are described.

[0301] Map publisher 136 will land Figure 2 Uploaded to server 20. Server 20's map storage routine 118 then... Figure 2 It is stored in the storage medium of server 20.

[0302] Map merging algorithm 124 will merge the maps Figure 2Merging with canonical map 120. When more than two maps have been stored (e.g., three or four maps related to the same or adjacent areas of the physical world), map merging algorithm 124 merges all maps into canonical map 120 to render the new canonical map 120. Map transmitter 122 then sends the new canonical map 120 to any and all devices 12.1 and 12.2 located in the area represented by the new canonical map 120. When devices 12.1 and 12.2 locate their respective maps to canonical map 120, canonical map 120 becomes the upgraded map.

[0303] Figure 17 The illustration shows an example of generating keyframes for a scene map according to some embodiments. In the example shown, a first keyframe KF1 is generated for the door on the left wall of the room. A second keyframe KF2 is generated for the corner area where the floor, left wall, and right wall of the room intersect. A third keyframe KF3 is generated for the window area on the right wall of the room. A fourth keyframe KF4 is generated on the floor of the wall, at the far end of the carpet. A fifth keyframe KF5 is generated for the area of ​​the carpet closest to the user.

[0304] Figure 18 The following are examples of some embodiments. Figure 17 An example of map generation of persistent poses. In some embodiments, a new persistent pose is created when the device measures a threshold distance traveled, and / or when the application requests a new persistent pose (PP). In some embodiments, the threshold distance can be 3 meters, 5 meters, 20 meters, or any other suitable distance. Choosing a smaller threshold distance (e.g., 1 m) can lead to an increased computational load because a larger number of PPs can be created and managed compared to a larger threshold distance. Choosing a larger threshold distance (e.g., 40 m) can lead to increased virtual content placement errors because fewer PPs will be created, which will result in fewer PCFs being created, meaning that virtual content attached to PCFs may be at a relatively large distance from the PCF (e.g., 30 m), and the error increases with the increasing distance from the PCF to the virtual content.

[0305] In some embodiments, a Point of Purchase (PP) can be created at the start of a new session. This initial PP can be considered zero and can be visualized as the center of a circle with a radius equal to a threshold distance. When the device reaches the circumference of the circle, and in some embodiments, when the application requests a new PP, the new PP can be placed at the device's current location (at the threshold distance). In some embodiments, if the device can find an existing PP within a threshold distance from the device's new location, a new PP will not be created at the threshold distance. In some embodiments, when a new PP is created (e.g., ...), Figure 14In PP 1150, the device attaches one or more of the nearest keyframes to the PP. In some embodiments, the position of the PP relative to the keyframes may be based on the device's position at the time the PP is created. In some embodiments, a PP will not be created when the device has traveled a threshold distance unless the application requests a PP.

[0306] In some embodiments, when an application has virtual content to display to the user, the application can request a PCF from the device. A PCF request from the application can trigger a PP request, and a new PP will be created after the device has traveled a threshold distance. Figure 18 The first permanent pose PP1 is shown, which may have the closest keyframes (e.g., KF1, KF2, and KF3) attached by, for example, calculating the relative pose between the keyframes and the permanent pose. Figure 18 A second permanent pose PP2 is also shown, which may have additional nearest keyframes (e.g., KF4 and KF5).

[0307] Figure 19 The following are examples of some embodiments. Figure 17 An example of a map generating PCF is provided. In the illustrated example, PCF1 may include PP1 and PP2. As described above, PCFs can be used to display image data associated with a PCF. In some embodiments, each PCF may have coordinates in another coordinate frame (e.g., a world coordinate frame) and a PCF descriptor, for example, uniquely identifying the PCF. In some embodiments, the PCF descriptor may be computed based on feature descriptors of features in the frame associated with the PCF. In some embodiments, various constellations of PCFs may be combined to represent the real world in a persistent manner requiring less data and less data transfer.

[0308] Figures 20A to 20C This is a schematic diagram illustrating an example of creating and using a persistent coordinate frame. Figure 20A The diagram shows two users 4802A and 4802B with corresponding local tracking maps 4804A and 4804B that have not yet been mapped to a standard map. The origins 4806A and 4806B for each user are depicted by a coordinate system (e.g., a world coordinate system) within their respective regions. These origins for each tracking map may be local to each user, as they depend on the orientation of their respective devices at the time tracking is initiated.

[0309] When the user device's sensors scan the environment, the device can capture the above-mentioned combinations. Figure 14 The description may include images representing features of persistent objects, such that those images can be classified as keyframes from which persistent poses can be created. In this example, tracking map 4802A includes persistent pose (PP) 4808A; tracking 4802B includes PP 4808B.

[0310] Similarly, as described above... Figure 14 Some PPs can be classified as PCFs, which are used to determine the orientation of virtual content in order to render it to the user. Figure 20B The diagram shows that the XR devices worn by the corresponding users 4802A and 4802B can create local PCF 4810A and 4810B based on PP 4808A and 4808B. Figure 20C It is shown that persistent content 4812A, 4812B (e.g., virtual content) can be attached to PCF 4810A, 4810B via a corresponding XR device.

[0311] In this example, the virtual content can have a virtual content coordinate frame, which can be used by the application that generates the virtual content, regardless of how the virtual content should be displayed. For example, the virtual content can be specified as a surface at a specific position and angle relative to the virtual content coordinate frame, such as triangles in a grid. In order to render the virtual content to the user, the positions of those surfaces can be determined relative to the user who wants to perceive the virtual content.

[0312] Attaching virtual content to the PCF simplifies the computations involved in determining the virtual content's position relative to the user. The position of the virtual content relative to the user can be determined by applying a series of transformations. Some of these transformations may change and may be updated frequently. Others may be stable, may be updated frequently, or may not be updated at all. Regardless, transformations can be applied with a relatively low computational burden, allowing the position of the virtual content to be updated frequently relative to the user, thus providing a realistic appearance to the rendered virtual content.

[0313] exist Figures 20A to 20C In the example, User 1's device has a coordinate system that relates to a coordinate system whose map origin is defined by the transformation rig1_T_w1. User 2's device has a similar transformation rig2_T_w2. These transformations can be represented as six transformation degrees, specifying translation and rotation to align the device coordinate system with the map coordinate system. In some embodiments, the transformations can be represented as two separate transformations, one specifying translation and the other specifying rotation. Therefore, it should be understood that transformations can be expressed in a form that simplifies calculations or otherwise provides advantages.

[0314] The transformation between the origin of the tracking map and the PCF identified by the corresponding user equipment is represented as pcf1_T_w1 and pcf2_T_w2. In this example, the PCF and PP are the same, so the same transformation also characterizes PP.

[0315] Therefore, the position of the user equipment relative to the PCF can be calculated through the serial application of these transformations, for example, rig1_T_pcf1 = (rig1_T_w1) * (pcf1_T_w1).

[0316] like Figure 20C As shown, the virtual content is positioned relative to the PCF through the transformation obj1_T_pcf1. This transformation can be set by the application generating the virtual content, which receives information describing physical objects relative to the PCF from the world reconstruction system. To render the virtual content to the user, the transformation to the user's device coordinate system is calculated. This can be done by associating the virtual content coordinate frame with the origin of the tracking map through the transformation obj1_t_w1 = (obj1_T_pcf1) * (pcf1_T_w1). This transformation can then be further correlated with the user's device through the transformation rig1_T_w1.

[0317] The position of the virtual content can change based on the output from the application that generates it. When this changes, the end-to-end transformation from the source coordinate system to the destination coordinate system can be recalculated. Additionally, the user's position and / or head pose can change as the user moves. Consequently, the transformation rig1_T_w1 can change, and any end-to-end transformation depending on the user's position or head pose can also be altered.

[0318] The transformation rig1_T_w1 can be updated as the user moves, based on tracking the user's position relative to a stationary object in the physical world. This tracking can be performed by the headphone positioning component or other components of the system that process image sequences as described above. Such updates can be made by determining the user's pose relative to a fixed reference frame (e.g., PP).

[0319] In some embodiments, since the PP is used as the PCF, the position and orientation of the user equipment can be determined relative to the most recent persistent pose, or in this example, the PCF. This determination can be made by identifying feature points characterizing the PP in a current image captured using sensors on the device. Using image processing techniques such as stereo image analysis, the position of the device relative to those feature points can be determined. Based on this data, the system can calculate the changes in transformation associated with the user's motion based on the relationship rig1_T_pcf1 = (rig1_T_w1) * (pcf1_T_w1).

[0320] The system can determine and apply transformations in a computationally efficient order. For example, the need to calculate rig1_T_w1 from the measurements that generate rig1_T_pcf1 can be avoided by tracking the user's pose and defining the position of the virtual content relative to a PP or PCF constructed based on the persistent pose. Thus, the transformation from the source coordinate system of the virtual content to the destination coordinate system of the user device can be based on the measured transformation according to the expression (rig1_T_pcf1)*(obj1_t_pcf1), where the first transformation is measured by the system and the latter is provided by the application that specifies the virtual content to be rendered. In embodiments where the virtual content is positioned relative to the origin of the map, the end-to-end transformation can correlate the virtual object coordinate system with the PCF coordinate system based on a further transformation between map coordinates and PCF coordinates. In embodiments where the virtual content is positioned relative to a different PP or PCF than the one used to track the user's position, a transformation can be performed between the two. Such a transformation can be fixed and can be determined, for example, from a map where both exist.

[0321] For example, a transformation-based method can be implemented in a device with components that process sensor data to construct a tracking map. As part of this process, these components can identify feature points that can be used as persistent poses, which can then be transformed into PCFs. These components can limit the number of persistent poses generated for the map to provide appropriate intervals between them, while allowing the user to be sufficiently close to the persistent pose location regardless of their physical location to accurately calculate the user's pose, as described above. Figures 17 to 19 As shown, with updates to the most recent persistent pose, the refinement of the tracking map or other elements due to user movement allows any transformations of the virtual content relative to the user's position—which are used to calculate the position depending on the PP (or PCF, if in use)—to be updated and stored for use, at least until the user leaves that persistent pose. Nevertheless, by calculating and storing transformations, the computational burden of updating the position of the virtual content each time can be relatively low, thus allowing it to be performed with relatively low latency.

[0322] Figures 20A to 20C The diagram illustrates positioning relative to a tracking map, with each device having its own tracking map. However, transformations can be generated relative to any map coordinate system. Content persistence between user sessions in an XR system can be achieved using persistent maps. Shared user experiences can also be achieved by using maps that can be directed to by multiple user devices.

[0323] In some embodiments described in more detail below, the location of the virtual content can be specified relative to coordinates in a canonical map, the canonical map being formatted so that any of a plurality of devices can use it. Each device may maintain a tracking map and can determine changes in the user's pose relative to the tracking map. In this example, the transformation between the tracking map and the canonical map can be determined by a "localization" process, which can be performed by matching structures in the tracking map (such as one or more persistent poses) to one or more structures in the canonical map (such as one or more PCFs).

[0324] The techniques for creating and using canonical maps in this way are described in more detail below.

[0325] Depth keyframe

[0326] The techniques described herein rely on the comparison of image frames. For example, to establish the device's position relative to a tracking map, new images can be captured using sensors worn by the user, and the XR system can search the image set used to create the tracking map for images that share at least a predetermined number of points of interest with the new images. As an example of another scenario involving image frame comparison, the tracking map can be localized to the canonical map by first finding image frames in the tracking map associated with a persistent pose that are similar to image frames in the canonical map associated with a PCF (Positive Contour Field). Alternatively, the transformation between the two canonical maps can be calculated by first finding similar image frames in the two maps.

[0327] Depth keyframes provide a method to reduce the amount of processing required to identify similar image frames. For example, in some embodiments, the comparison can be between image features (e.g., “2D features”) in a new 2D image and 3D features in a map. This comparison can be performed in any suitable manner, such as by projecting the 3D image onto a 2D plane. Conventional methods such as Bag of Words (BoW) search for 2D features of the new image in a database that includes all 2D features in the map, which can require significant computational resources, especially when the map represents a large area. Conventional methods then locate images that share at least one 2D feature with the new image, which may include images that are not useful for locating 3D features that are meaningful in the map. Conventional methods then locate 3D features that are not meaningful relative to the 2D features in the new image.

[0328] The inventors have recognized and understood techniques for retrieving images from maps using fewer memory resources (e.g., a quarter of the memory resources used by BoW), with higher efficiency (e.g., 2.5 ms processing time per keyframe and 100 µs for comparisons of 500 keyframes), and higher accuracy (e.g., 20% better retrieval recall than BoW for a 1024-dimensional model and 5% better retrieval recall than BoW for a 256-dimensional model).

[0329] To reduce computation, descriptors can be computed for each image frame, which can be used to compare the image frame with other image frames. Descriptors can be stored in place of image frames and feature points, or in addition to image frames and feature points. In a map from which persistent poses and / or PCFs can be generated based on image frames, descriptors of one or more image frames upon which each persistent pose or PCF is based can be stored as part of the persistent pose and / or PCF.

[0330] In some embodiments, a descriptor can be computed based on feature points in an image frame. In some embodiments, a neural network is configured to compute a unique frame descriptor representing the image. The image can have a resolution greater than 1 megabyte, thereby capturing sufficient detail in the image of the 3D environment within the field of view of the device worn by the user. The frame descriptor can be much shorter, such as a numeric string, for example, a numeric string in the range of 128 bytes to 512 bytes or any number of numeric strings in between.

[0331] In some embodiments, the neural network is trained to compute frame descriptors that indicate the similarity between images. Images in a map can be located by identifying the nearest images that have frame descriptors within a predetermined distance from the frame descriptors of a new image in a database comprising images used to generate the map. In some embodiments, the distance between images can be represented by the difference between the frame descriptors of the two images.

[0332] Figure 21 This is a block diagram illustrating a system for generating descriptors for individual images according to some embodiments. In the example shown, a frame embedding generator 308 is illustrated. In some embodiments, the frame embedding generator 308 may be used within server 20, but may alternatively or additionally be implemented wholly or partially in one of XR devices 12.1 and 12.2 or any other device that processes images for comparison with other images.

[0333] In some embodiments, the frame embedding generator can be configured to generate a reduced data representation of an image from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes), which, despite the reduced size, still indicates the content in the image. In some embodiments, the frame embedding generator can be used to generate a data representation of an image, which may be a keyframe or a frame used in other ways. In some embodiments, the frame embedding generator 308 can be configured to convert an image located at a specific position and orientation into a unique numeric string (e.g., 256 bytes). In the example shown, an image 320 captured by an XR device can be processed by a feature extractor 324 to detect points of interest 322 in the image 320. Points of interest may or may not be derived from the feature points identified as described above for feature 1120. Figure 14 ) or as otherwise described herein. In some embodiments, concerns may be directed to descriptor 1130 as described above ( Figure 14 The descriptors described in the diagram are used to represent these concerns, which can be generated using a deep sparse feature approach. In some embodiments, each concern 322 can be represented by a numeric string (e.g., 32 bytes). For example, there can be n features (e.g., 100), and each feature can be represented by a 32-byte string.

[0334] In some embodiments, the frame embedding generator 308 may include a neural network 326. The neural network 326 may include a multilayer perceptron unit (MLP) 312 and a max pooling unit 314. In some embodiments, the multilayer perceptron unit 312 may include a multilayer perceptron that can be trained. In some embodiments, the multilayer perceptron unit 312 may be used to reduce concerns 322 (e.g., descriptors for concerns) and may output them as a weighted combination of descriptors. For example, the MLP may reduce n features to m features (less than n features).

[0335] In some embodiments, the MLP unit 312 can be configured to perform matrix multiplication. The multilayer perceptron unit 312 receives multiple points of interest 322 from the image 320 and converts each point of interest into a corresponding numeric string (e.g., 256). For example, there may be 100 features, and each feature can be represented by a string of 256 numbers. In this example, a matrix with 100 horizontal rows and 256 vertical columns can be created. Each row may have a series of 256 numbers, which vary in size, with some smaller and others larger. In some embodiments, the output of the MLP can be an n×256 matrix, where n represents the number of points of interest extracted from the image. In some embodiments, the output of the MLP can be an m×256 matrix, where m is the number of points of interest reduced from n.

[0336] In some embodiments, the MLP unit 312 may have a training phase and a usage phase, during which model parameters for the MLP are determined. In some embodiments, it may be as follows: Figure 25 The MLP is trained as shown in the diagram. Input training data can consist of sets of three data points: 1) a query image, 2) positive samples, and 3) negative samples. The query image can be considered a reference image.

[0337] In some embodiments, a positive sample may include an image similar to the query image. For example, in some embodiments, similarity may mean having the same object in both the query image and the positive sample image, but viewed from different angles. In some embodiments, similarity may mean having the same object in both the query image and the positive sample image, but that object is shifted relative to the other image (e.g., to the left, right, up, or down).

[0338] In some embodiments, negative samples may include images that are dissimilar to the query image. For example, in some embodiments, dissimilar images may not contain any objects that are prominent in the query image, or may contain only a small fraction (e.g., <10%, 1%) of the prominent objects in the query image. In contrast, similar images may contain a large majority (e.g., >50% or >75%) of the objects in the query image.

[0339] In some embodiments, attention points can be extracted from images in the input training data, and these attention points can be converted into feature descriptors. This can be applied to, for example... Figure 25 The training images shown and the targets Figure 21The features extracted from the operation of the frame embedding generator 308 are used to compute these descriptors. In some embodiments, as described in U.S. Patent Application 16 / 190,948, deep sparse feature (DSF) processing can be used to generate descriptors (e.g., DSF descriptors). In some embodiments, the DSF descriptor is n×32 dimensional. The descriptors can then be passed via a model / MLP to create a 256-byte output. In some embodiments, the model / MLP can have the same structure as MLP unit 312, such that once the model parameters are set via training, the resulting trained MLP can be used as MLP unit 312.

[0340] In some embodiments, feature descriptors (e.g., 256 bytes from the MLP model output) can then be sent to a triplet boundary loss module (which may be used only during the training phase and not during the usage phase of the MLP neural network). In some embodiments, the triplet boundary loss module can be configured to select model parameters to reduce the difference between the 256-byte output from the query image and the 256-byte output from the positive sample, and to increase the difference between the 256-byte output from the query image and the 256-byte output from the negative sample. In some embodiments, the training phase may include feeding multiple triplet input images into the learning process to determine model parameters. This training process may continue, for example, until the difference between the positive images is minimized and the difference between the negative images is maximized, or until other appropriate exit criteria are met.

[0341] Refer again Figure 21 The frame embedding generator 308 may include a pooling layer, shown herein as a max pooling unit 314. The max pooling unit 314 can analyze each column to determine the maximum number in the corresponding column. The max pooling unit 314 can combine the maximum values ​​of each column of numbers in the output matrix of the MLP unit 312 into a global feature string 316 of, for example, 256 numbers. It should be understood that images processed in an XR system may be expected to have high-resolution frames, potentially with millions of pixels. The global feature string 316 is a relatively small number that occupies relatively little memory and is easier to search compared to images (e.g., with a resolution greater than 1 megabyte). Therefore, images can be searched without analyzing every raw frame from the camera, and storing 256 bytes instead of the entire frame is also cheaper.

[0342] Figure 22This is a flowchart illustrating a method 2200 for calculating an image descriptor according to some embodiments. Method 2200 may begin by receiving (action 2202) multiple images captured by an XR device worn by a user. In some embodiments, method 2200 may include determining (action 2204) one or more keyframes from the multiple images. In some embodiments, action 2204 may be skipped and / or may occur instead after step 2210.

[0343] Method 2200 may include: identifying (action 2206) one or more points of interest in multiple images using an artificial neural network; and calculating (action 2208) feature descriptors for each point of interest using an artificial neural network. The method may include calculating (action 2210) frame descriptors for each image, thereby representing the image at least in part based on feature descriptors calculated using an artificial neural network for the identified points of interest in the image.

[0344] Figure 23 This is a flowchart illustrating a method 2300 for localization using image descriptors according to some embodiments. In this example, a new image frame describing the current location of an XR device can be compared with an image frame stored in conjunction with points in a map (e.g., persistent poses or PCFs as described above). Method 2300 can begin by receiving (action 2302) a new image captured by an XR device worn by a user. Method 2300 may include identifying (action 2304) one or more recent keyframes in a database that includes keyframes used to generate one or more maps. In some embodiments, recent keyframes can be identified based on coarse spatial information and / or previously determined spatial information. For example, coarse spatial information may indicate that the XR device is located in a geographic area represented by a 50m × 50m area of ​​a map. Image matching can be performed only for points within that area. As another example, based on tracking, the XR system can know that the XR device previously approached a first persistent pose in the map and is currently moving in the direction of a second persistent pose in the map. This second persistent pose can be considered the most recent persistent pose, and the keyframes stored with it can be considered the most recent keyframes. Alternatively or additionally, other metadata, such as GPS data or WiFi fingerprints, can be used to select the most recent keyframe or a set of most recent keyframes.

[0345] Regardless of how the nearest keyframe is selected, frame descriptors can be used to determine whether a new image matches any frame selected as being associated with a nearby persistent pose. This determination can be performed by comparing the frame descriptor of the new image with the frame descriptors of the nearest keyframe or a subset of keyframes in a database selected in any other suitable manner, and selecting a keyframe whose frame descriptor is within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors can be calculated by obtaining the difference between two numeric strings that can represent the two frame descriptors. In embodiments where the strings are processed as multiple strings, the difference can be calculated as a vector difference.

[0346] Once a matching image frame is identified, the orientation of the XR device relative to that image frame can be determined. Method 2300 may include: performing feature matching (action 2306) on 3D features in a map corresponding to the identified nearest keyframe, and calculating (action 2308) the pose of the device worn by the user based on the feature matching result. In this way, computationally dense matching of feature points in two images can be performed for as few as one image that has been determined to be a possible match with a new image.

[0347] Figure 24 This is a flowchart illustrating a method 2400 for training a neural network according to some embodiments. Method 2400 may begin by generating (action 2402) a dataset comprising multiple image sets. Each of the multiple image sets may include a query image, positive sample images, and negative sample images. In some embodiments, the multiple image sets may include synthetic record pairs configured to, for example, teach basic information (such as shape) to the neural network. In some embodiments, the multiple image sets may include real record pairs that can be based on physical world records.

[0348] In some embodiments, inliers can be computed by fitting a fundamental matrix between the two images. In some embodiments, sparse overlap can be computed as the intersection-over-union (IoU) ratio of the points of interest seen in the two images. In some embodiments, positive samples may include at least twenty identical points of interest from the query image as inliers. Negative samples may include fewer than ten inliers. Negative samples may have less than half of their sparse points overlapping with the sparse points of the query image.

[0349] Method 2400 may include calculating (action 2404) a loss for each image set by comparing the query image with positive and negative sample images. Method 2400 may include modifying (action 2406) the artificial neural network based on the calculated loss such that the distance between the frame descriptor generated by the artificial neural network for the query image and the frame descriptor for the positive sample images is smaller than the distance between the frame descriptor for the query image and the frame descriptor for the negative sample images.

[0350] It should be understood that although the methods and apparatus described above are configured to generate global descriptors for individual images, the methods and apparatus can be configured to generate descriptors for individual maps. For example, a map may include multiple keyframes, each of which may have a frame descriptor as described above. The max-pooling unit can analyze the frame descriptors of the keyframes of the map and combine the frame descriptors into a unique map descriptor for that map.

[0351] Furthermore, it should be understood that other architectures can be used for the processing described above. For example, a separate neural network for generating DSF descriptors and frame descriptors is described. This approach is computationally efficient. However, in some embodiments, frame descriptors can be generated based on selected feature points without first generating DSF descriptors.

[0352] Rank and merge maps

[0353] This document describes methods and apparatus for ranking and merging multiple environmental maps in an X-Reality (XR) system. Map merging allows maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking the maps enables the efficient execution of the techniques described herein, including map merging, which involves selecting maps from a set of maps based on similarity. In some embodiments, for example, the system may maintain a set of canonical maps formatted in a manner accessible to any of a number of XR devices. These canonical maps can be formed by merging selected tracking maps from those devices with other tracking maps or previously stored canonical maps. Canonical maps can be ranked, for example, to select one or more canonical maps to merge with a new tracking map and / or to select one or more canonical maps from a set for use in a device.

[0354] To provide users with a realistic XR experience, an XR system must understand the user's physical environment in order to correctly correlate the positions of virtual objects with real objects. Information about the user's actual environment can be obtained from an environment map of the user's location.

[0355] The inventors have recognized and understood that XR systems can provide an enhanced XR experience to multiple users sharing the same world, including real and / or virtual content, by enabling efficient sharing of environmental maps of the real / physical world collected by multiple users, regardless of whether these users are in the world at the same time or at different times. However, significant challenges exist in providing such a system. Such a system may store multiple maps generated by multiple users and / or the system may store multiple maps generated at different times. For operations that may be performed using previously generated maps (such as, for example, positioning as described above), a significant amount of processing may be required to identify relevant environmental maps of the same world (e.g., the same real-world location) from all environmental maps collected by the XR system. In some embodiments, only a small number of environmental maps may exist that are accessible to the device for, for example, positioning. In some embodiments, a large number of environmental maps may exist that are accessible to the device. The inventors have recognized and understood the need for rapid and accurate processing of environmental maps from all possible environmental maps (such as, for example, positioning as described above). Figure 28 The technique ranks the relevance of environmental maps in all 120 canonical maps of a universe. High-ranking maps can then be selected for further processing, such as rendering virtual objects on a user's display to realistically interact with the physical world around the user, or merging the user-collected map data with stored maps to create a larger or more accurate map.

[0356] In some embodiments, stored maps relevant to a user's task at a physical location can be identified by filtering stored maps based on multiple criteria. These criteria can indicate a comparison between a tracking map generated by the user's wearable device at that location and candidate environment maps stored in a database. Comparisons can be performed based on metadata associated with the maps, such as a Wi-Fi fingerprint detected by the device that generated the map and / or a set of BSSIDs to which the device was connected at the time the map was formed. Comparisons can also be performed based on compressed or uncompressed content of the maps. Comparisons based on compressed representations can be performed by comparing vectors computed from the map content. For example, a comparison based on an uncompressed map can be performed by locating the tracking map within a stored map, and vice versa. Multiple comparisons can be performed sequentially based on the computation time required to reduce the number of candidate maps to be considered, where comparisons involving less computation are performed earlier in sequence compared to other comparisons requiring more computation.

[0357] Figure 26An AR system 800, configured to rank and merge one or more environmental maps according to some embodiments, is depicted. The AR system may include a navigable world model 802 of an AR device. Information populating the navigable world model 802 may come from sensors on the AR device, which may include data stored in a processor 804 (e.g., ...). Figure 4 The local data processing module 570 contains computer-executable instructions that allow the processor to perform some or all of the processing to convert sensor data into a map. This map can be a tracking map, as a tracking map can be built while the AR device is operating in the area, collecting sensor data along the way. Figure 1 The system can provide regional attributes to indicate the area represented by the tracking map. These regional attributes can be geographic location identifiers, such as coordinates expressed as latitude and longitude, or IDs used by the AR system to represent a location. Alternatively or additionally, regional attributes can be measurable characteristics that have a high probability of being unique to that area. Regional attributes can, for example, be derived from parameters of wireless networks detected in that area. In some embodiments, regional attributes can be associated with the unique addresses of access points nearby and / or connected to by the AR system. For example, regional attributes can be associated with the MAC address or Basic Service Set Identifier (BSSID) of a 5G base station / router, Wi-Fi router, etc.

[0358] exist Figure 26 In the example, the tracking map can be merged with other environmental maps. The map ranking section 806 receives the tracking map from device PW 802 and communicates with the map database 808 to select and rank environmental maps from the map database 808. The selected maps with higher rankings are sent to the map merging section 810.

[0359] The map merging section 810 can perform merging processing on the maps sent from the map ranking section 806. The merging process may require merging the tracking map with some or all of the ranking maps and sending the new merged map to the walkable world model 812. The map merging section can merge maps by identifying overlapping portions depicting the physical world. Those overlapping portions can be aligned so that information from the two maps can be aggregated into the final map. The canonical map can be merged with other canonical maps and / or tracking maps.

[0360] Aggregation may require expanding a map with information from another map. Alternatively or additionally, aggregation may require adjusting the representation of the physical world in one map based on information from another map. For example, the later map may reveal objects that caused feature points to have moved, allowing the map to be updated based on the later information. Alternatively, two maps may characterize the same area with different feature points, and aggregation may require selecting a set of feature points from the two maps to better represent the area. Regardless of the specific processing that occurs during the merging process, in some embodiments, PCFs from all the merged maps may be retained, allowing applications that locate content relative to them to continue doing so. In some embodiments, map merging may result in redundant persistent poses, and some persistent poses may be deleted. When a PCF is associated with a persistent pose to be deleted, the merged map may need to modify the PCF to be associated with the persistent poses retained in the merged map.

[0361] In some embodiments, maps can be refined as they are expanded and / or updated. Refinement may require computation to reduce internal inconsistencies between feature points that may represent the same objects in the physical world. Such inconsistencies may arise from inaccurate poses associated with keyframes that provide feature points representing the same objects in the physical world. For example, such inconsistencies may arise when an XR device calculates pose relative to a tracking map, which in turn is built based on estimated poses, and thus errors in pose estimation accumulate, resulting in a “drift” in pose accuracy over time. Maps can be refined by performing bundle adjustment or other operations to reduce inconsistencies from feature points across multiple keyframes.

[0362] During refinement, the position of a persistent point relative to the map origin can change. Therefore, the transformations associated with that persistent point, such as persistent pose or PCF, may change. In some embodiments, an XR system incorporating map refinement (whether performed as part of a merging operation or for other reasons) can recalculate the transformations associated with any changed persistent points. These transformations may be pushed from the component that computes the transformations to the component that uses them, so that any use of the transformations can be based on the updated position of the persistent point.

[0363] The walkable world model 812 can be a cloud model that can be shared by multiple AR devices. The walkable world model 812 can store or otherwise access environmental maps in the map database 808. In some embodiments, when a previously calculated environmental map is updated, a previous version of the map can be deleted to remove outdated maps from the database. In some embodiments, when a previously calculated environmental map is updated, a previous version of the map can be archived, making it possible to access / view previous versions of the environment. In some embodiments, permissions can be set such that only AR systems with certain read / write access permissions can trigger the deletion / archiving of previous versions of the map.

[0364] These environmental maps, created from tracking maps provided by one or more AR devices / systems, can be accessed by AR devices within the AR system. The map ranking section 806 can also be used to provide environmental maps to AR devices. An AR device can send a message requesting an environmental map of its current location, and the map ranking section 806 can be used to select and rank the environmental maps relevant to the requesting device.

[0365] In some embodiments, the AR system 800 may include a downsampling unit 814 configured to receive a merged map from a cloud PW 812. The merged map received from the cloud PW 812 may be a cloud-based storage format that may include high-resolution information, such as a large number of PCFs per square meter or multiple image frames or a large set of feature points associated with a PCF. The downsampling unit 814 may be configured to downsample the cloud-formatted map to a format suitable for storage on an AR device. The device-formatted map may contain less data, such as fewer PCFs or less data stored per PCF, to accommodate the limited local computing power and storage space of the AR device.

[0366] Figure 27 This is a simplified block diagram illustrating multiple specification maps 120 that can be stored in a remote storage medium, such as the cloud. Each specification map 120 may include multiple specification map identifiers that indicate the location of the specification map in physical space, such as a location on Earth. These specification map identifiers may include one or more of the following identifiers: a region identifier represented by a range of longitude and latitude, a frame descriptor (e.g., ... Figure 21 Global feature string 316), Wi-Fi fingerprint, feature descriptor (e.g., Figure 21 The feature descriptor 310 in the map, and the device identifier indicating one or more devices that contribute to the map.

[0367] In the example shown, the canonical maps 120 are geographically arranged in a two-dimensional pattern because they can exist on the Earth's surface. The canonical maps 120 can be uniquely identified by their respective longitudes and latitudes because any canonical maps with overlapping longitudes and latitudes can be merged into a new canonical map.

[0368] Figure 28 This is a schematic diagram illustrating a method for selecting a canonical map according to some embodiments, which can be used to locate a new tracking map to one or more canonical maps. The method can begin by accessing (action 120) the world of canonical map 120, which, as an example, canonical map 120's world can be stored in a database of traversable worlds (e.g., traversable world module 538). The world of the canonical map can include canonical maps from all previously visited locations. The XR system can filter all the worlds of the canonical maps into a small subset or just one map. It should be understood that in some embodiments, it is not possible to send all canonical maps to the viewing device due to bandwidth limitations. Selecting a subset of candidates that can be used to match the tracking map and sending it to the device can reduce the bandwidth and latency associated with accessing a remote database of maps.

[0369] This method may include filtering the world of a canonical map based on regions of predetermined size and shape (Action 300). Figure 27 In the example, each square can represent an area. Each square can cover 50 m × 50 m. Each square can have six adjacent areas. In some embodiments, action 300 can select at least one matching canonical map 120 covering longitude and latitude, wherein the longitude and latitude include the longitude and latitude of the location identifier received from the XR device, provided that at least one map exists at that longitude and latitude. In some embodiments, action 300 can select at least one adjacent canonical map covering the longitude and latitude adjacent to the matching canonical map. In some embodiments, action 300 can select multiple matching canonical maps and multiple adjacent canonical maps. Action 300 can, for example, reduce the number of canonical maps by about ten times, for example, from thousands to hundreds, to form a first filtering selection. Alternatively or additionally, criteria other than latitude and longitude can be used to identify adjacent maps. For example, the XR device may have previously used canonical maps in the collection for positioning as part of the same session. The cloud service can retain information about the XR device, including previously located maps. In this example, the map selected at action 300 may include those maps that cover the area adjacent to the map located by the XR device.

[0370] The method may include a first filtering selection based on Wi-Fi fingerprints (action 302) of canonical maps. Action 302 may determine latitude and longitude based on the Wi-Fi fingerprint received from the XR device as part of a location identifier. Action 302 may compare the latitude and longitude from the Wi-Fi fingerprint with the latitude and longitude of canonical map 120 to determine one or more canonical maps forming a second filtering selection. Action 302 may reduce the number of canonical maps by approximately tenfold, for example, from hundreds of canonical maps to dozens (e.g., 50) of canonical maps forming the second selection. For example, the first filtering selection may include 130 canonical maps, the second filtering selection may include 50 of the 130 canonical maps, and may exclude the other 80 of the 130 canonical maps.

[0371] The method may include a second filtering selection based on keyframes (action 304) of the canonical map. Action 304 may compare data representing an image captured by the XR device with data representing the canonical map 120. In some embodiments, the data representing the image and / or map may include feature descriptors (e.g., Figure 25 DSF descriptors in the DSF descriptor) and / or global feature strings (e.g., Figure 21 (316 in the original text). Action 304 can provide a third filtering selection of the canonical maps. In some embodiments, for example, the output of action 304 may be only five canonical maps out of 50 canonical maps identified after the second filtering selection. Map transmitter 122 then sends one or more canonical maps based on the third filtering selection to the viewing device. Action 304 can reduce the number of canonical maps by approximately tenfold, for example, from dozens of canonical maps to a single-digit number (e.g., 5) forming the third selection. In some embodiments, the XR device may receive canonical maps in the third filtering selection and attempt to locate within the received canonical maps.

[0372] For example, action 304 can filter the canonical map 120 based on the global feature string 316 of the canonical map 120 and the global feature string 316 of an image captured by the viewing device (e.g., an image that may be part of the user's local tracking map). Therefore, Figure 27 Each canonical map 120 in the system has one or more global feature strings 316 associated with it. In some embodiments, the global feature strings 316 can be obtained when the XR device submits images or feature details to the cloud and processes these images or feature details in the cloud to generate global feature strings 316 for the canonical map 120.

[0373] In some embodiments, the cloud may receive feature details of a real-time / new / current image captured by a viewing device, and the cloud may generate a global feature string 316 for the real-time image. The cloud may then filter the canonical map 120 based on the real-time global feature string 316. In some embodiments, the global feature string may be generated on a local viewing device. In some embodiments, the global feature string may be generated remotely in the cloud, for example. In some embodiments, the cloud may send the filtered canonical map along with the global feature string 316 associated with the filtered canonical map to the XR device. In some embodiments, when the viewing device positions its tracking map to the canonical map, it may do so by matching the global feature string 316 of the local tracking map with the global feature string of the canonical map.

[0374] It should be understood that the operation of an XR device may not perform all actions (300, 302, 304). For example, if the canonical map world is relatively small (e.g., 500 maps), an XR device attempting localization may filter the canonical map world based on Wi-Fi fingerprints (e.g., action 302) and keyframes (e.g., action 304), but omit region-based filtering (e.g., action 300). Furthermore, it is not necessary to compare the entire map. For example, in some embodiments, comparing two maps may lead to the identification of common persistent points, such as persistent poses or PCFs that appear in both the new map and the map selected from the map world. In that case, descriptors can be associated with persistent points, and those descriptors can be compared.

[0375] Figure 29 This is a flowchart illustrating a method 900 for selecting one or more rankings of environmental maps according to some embodiments. In the illustrated embodiment, ranking is performed on the AR device of a user creating a tracking map. Therefore, the tracking map can be used to rank the environmental maps. In embodiments where the tracking map is unavailable, some or all portions of the selection and ranking of environmental maps that do not explicitly depend on the tracking map can be used.

[0376] Method 900 may begin with action 902, where a set of maps in a database of environmental maps (which may be formatted as canonical maps) located near the location where the tracking map is formed can be accessed and then filtered for ranking. Additionally, at action 902, at least one region attribute of the area in which the user's AR device is operating is determined. In a scenario where the user's AR device is constructing a tracking map, the region attribute may correspond to the area on which the tracking map is created. As a specific example, the region attribute may be calculated based on received signals from an access point to a computer network while the AR device is calculating the tracking map.

[0377] Figure 30An exemplary map ranking portion 806 of an AR system 800 according to some embodiments is depicted. The map ranking portion 806 can execute in a cloud computing environment because it can include a portion executing on an AR device and a portion executing on a remote computing system such as the cloud. The map ranking portion 806 can be configured to execute at least a portion of method 900.

[0378] Figure 31A Examples of area attributes AA1-AA8 of a tracking map (TM) 1102 and environment maps CM1-CM4 in a database according to some embodiments are depicted. As shown, the environment map can be associated with multiple area attributes. Area attributes AA1-AA8 may include parameters of the wireless network detected by the tracking map 1102, calculated by the AR device, such as the Basic Service Set Identifier (BSSID) of the network to which the AR device is connected and / or the strength of the received signal to the access point of the wireless network via, for example, network tower 1104. The parameters of the wireless network may conform to protocols including Wi-Fi and 5G NR. Figure 32 In the example shown, the area attribute is a fingerprint of the area in which the user's AR device collects sensor data to form a tracking map.

[0379] Figure 31B An example of a determined geographic location 1106 for a tracking map 1102 according to some embodiments is depicted. In the illustrated example, the determined geographic location 1106 includes a centroid point and a region 1108 surrounding the centroid point. It should be understood that the determination of geographic location in this application is not limited to the format shown. The determined geographic location can have any suitable format, including, for example, different region shapes. In this example, the geographic location is determined from region attributes using a database that associates region attributes with geographic locations. The database is commercially available, for example, a database that associates Wi-Fi fingerprints with locations expressed as latitude and longitude and can be used for this operation.

[0380] exist Figure 29 In this embodiment, the map database containing environmental maps may also include location data for those maps, including the latitude and longitude covered by the maps. The processing at action 902 may require selecting a set of environmental maps from this database that cover the same latitude and longitude determined for the area attributes of the tracking map.

[0381] Action 904 is the first filtering of the set of environmental maps accessed in Action 902. In Action 902, environmental maps are retained in the set based on their geographical proximity to the tracking maps. This filtering step can be performed by comparing the latitude and longitude associated with the tracking maps and environmental maps in the set.

[0382] Figure 32An example of action 904 according to some embodiments is depicted. Each region attribute may have a corresponding geographic location 1202. The set of environmental maps may include environmental maps having at least one region attribute having a geographic location overlapping with a determined geographic location of the tracking map. In the example shown, a set of identified environmental maps includes environmental maps CM1, CM2, and CM4, each environmental map having at least one region attribute having a geographic location overlapping with a determined geographic location of the tracking map 1102. CM3, associated with region attribute AA6, is not included in this set because it is outside the determined geographic location of the tracking map.

[0383] Further filtering steps can be performed on the set of environmental maps to reduce / rank the number of environmental maps ultimately processed in the set (e.g., for map merging or to provide navigable world information to user devices). Method 900 may include filtering (action 906) the set of environmental maps based on the similarity of one or more identifiers of network access points associated with the tracking map and the environmental maps of the set of environmental maps. During map formation, devices that collect sensor data to generate maps may connect to a network via network access points (e.g., via Wi-Fi or similar wireless communication protocols). Access points can be identified by BSSID. When a user device moves through an area where data is collected to form a map, the user device may connect to multiple different access points. Similarly, when multiple devices provide information to form a map, the device may have been connected via different access points, and therefore, multiple access points may also be used when forming a map for this reason. Thus, multiple access points may exist associated with a map, and the set of access points may be an indication of map location. Signal strength from access points may be reflected as RSSI values, which can provide further geographic information. In some embodiments, a list of BSSIDs and RSSI values ​​may form regional attributes for the map.

[0384] In some embodiments, filtering the set of environment maps based on the similarity of one or more identifiers of network access points may include: retaining the environment map with the highest Jaccard similarity to at least one area attribute of the tracking map in the set of environment maps based on one or more identifiers of the network access points. Figure 33 An example of action 906 according to some embodiments is depicted. In the example shown, a network identifier associated with region attribute AA7 can be identified as the identifier of tracking map 1102. The set of environment maps following action 906 includes: environment map CM2, which may have a region attribute with a higher Jaccard similarity than AA7; and environment map CM4, which also includes region attribute AA7. Environment map CM1 is not included in this set because it has the lowest Jaccard similarity to AA7.

[0385] Actions 902-906 can be performed based on metadata associated with the map without actually accessing the map content stored in the map database. Other processes may involve accessing map content. Action 908 instructs access to the environment map retained in a subset after filtering based on metadata. It should be understood that if subsequent operations can be performed on the accessed content, this action can be performed earlier or later in the process.

[0386] Method 900 may include filtering (action 910) a set of environment maps based on the similarity of metrics representing the content of the tracking map and a set of environment maps. Metrics representing the content of the tracking map and environment maps may include vectors of values ​​computed from the content of the maps. For example, as described above, depth keyframe descriptors computed for one or more keyframes used to form a map can provide metrics for comparing maps or portions of maps. Metrics may be computed from maps obtained at action 908, or may be pre-computed and stored as metadata associated with those maps. In some embodiments, filtering a set of environment maps based on the similarity of metrics representing the content of the tracking map and a set of environment maps may include retaining the environment map with the minimum vector distance between the feature vector of the tracking map and the vector representing the environment map in the set of environment maps within the set of environment maps.

[0387] Method 900 may include further filtering (action 912) a set of environment maps based on the degree of matching between a portion of the tracking map and a portion of the environment map. The degree of matching may be determined as part of the localization process. As a non-limiting example, localization may be performed by identifying critical points in the tracking map and environment map that are sufficiently similar to the same parts of the physical world they can represent. In some embodiments, critical points may be features, feature descriptors, keyframes, key assemblies, persistent poses, and / or PCFs. A set of critical points in the tracking map may then be aligned to produce an optimal fit with that set of critical points in the environment map. The mean square distance between corresponding critical points may be calculated and used as an indication that the tracking map and environment map represent the same area of ​​the physical world if it is below a threshold for a specific region in the tracking map.

[0388] In some embodiments, filtering a set of environment maps based on the degree of matching between a portion of a tracking map and a portion of an environment map in a set of environment maps may include: calculating the volume of the physical world represented by the tracking map, which is also represented in the environment map of the set of environment maps; and retaining the environment map with a larger calculated volume than the environment map filtered from the set in the set of environment maps. Figure 34An example of action 912 according to some embodiments is depicted. In the example shown, a set of environment maps following action 912 includes environment map CM4, which has a region 1402 that matches the region of tracking map 1102. Environment map CM1 is not included in this set because it does not have a region that matches the region of tracking map 1102.

[0389] In some embodiments, the set of environment maps can be filtered in the order of actions 906, 910, and 912. In some embodiments, the set of environment maps can be filtered based on actions 906, 910, and 912, which can be performed in ascending order of the processing required to perform the filtering. Method 900 may include loading (action 914) the set of environment maps and data.

[0390] In the example shown, the user database stores area identifiers indicating the region where the AR device is used. The area identifier can be an area attribute, which may include parameters of the wireless network detected by the AR device during use. The map database can store multiple environmental maps constructed from data provided by the AR device and associated metadata. The associated metadata may include area identifiers derived from the area identifiers of the AR device providing the data, from which the environmental maps are constructed. The AR device can send a message to the PW module indicating that a new tracking map is being created or is being created. The PW module can calculate an area identifier for the AR device and update the user database based on the received parameters and / or the calculated area identifier. The PW module can also determine the area identifier associated with the AR device requesting the environmental map, identify the set of environmental maps from the map database based on the area identifier, filter the set of environmental maps, and send the filtered set of environmental maps to the AR device. In some embodiments, the PW module may filter the set of environment maps based on one or more criteria, including, for example, the geographic location of the tracking map, the similarity of one or more identifiers of network access points associated with the tracking map and the environment maps of the set of environment maps, the similarity of a measure representing the content of the tracking map and the environment maps of the set of environment maps, and the degree of matching between a portion of the tracking map and a portion of the environment maps of the set of environment maps.

[0391] Several aspects of some embodiments have been described so far, and it should be understood that various changes, modifications, and improvements will readily occur to those skilled in the art. As an example, the embodiments are described in conjunction with an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein can be applied in MR environments or more generally in other XR and VR environments.

[0392] As another example, embodiments are described in conjunction with devices such as wearable devices. It should be understood that some or all of the technologies described herein can be implemented via networks (such as the cloud), discrete applications and / or devices, or any suitable combination of networks and discrete applications.

[0393] also, Figure 29 Examples of criteria that can be used to filter candidate maps to produce a set of high-ranking maps are provided. Other criteria may be used instead of or in addition to the described criteria. For example, if multiple candidate maps have similar values ​​for metrics used to filter out less desirable maps, the characteristics of the candidate maps can be used to determine which maps are retained as candidate maps or filtered out. For example, larger or denser candidate maps may be preferred over smaller candidate maps. In some embodiments, Figure 27-28 Can describe Figures 29-34 The systems and methods described in whole or in part.

[0394] Figure 35 and 36 This is a schematic diagram illustrating an XR system configured to rank and merge multiple environment maps according to some embodiments. In some embodiments, the traversable world (PW) can determine when to trigger map ranking and / or merging. In some embodiments, determining the map to be used can be based at least in part on the above regarding Figures 21 to 25 Describes the depth keyframe.

[0395] Figure 37 This is a block diagram illustrating a method 3700 for creating an environment map of the physical world according to some embodiments. Method 3700 can localize (action 3702) a tracking map captured by a user-worn XR device to a canonical map (e.g., via...). Figure 28 The method and / or the method of Figure 900 (selected from the canonical map) group. Action 3702 may include locating key assemblies from the tracking map into the group of canonical maps. The localization result for each key assembly may include a set of localized poses and 2D-to-3D feature correspondences for the key assembly.

[0396] In some embodiments, method 3700 may include splitting the tracking map (action 3704) into connected portions, which can robustly merge the map by merging the connected fragments. Each connected portion may include a key assembly within a predetermined distance. Method 3700 may include: merging connected portions larger than a predetermined threshold (action 3706) into one or more canonical maps; and removing the merged connected portions from the tracking map.

[0397] In some embodiments, method 3700 may include merging (action 3708) canonical maps in a group that are merged with the same connected portions of the tracking map. In some embodiments, method 3700 may include lifting (action 3710) the remaining connected portions of the tracking map that have not yet been merged with any canonical map into a canonical map. In some embodiments, method 3700 may include merging (action 3712) the persistent pose and / or PCF of the tracking map and the canonical map, wherein the canonical map is merged with at least one connected portion of the tracking map. In some embodiments, method 3700 may include finalizing (action 3714) the canonical map, for example, by fusing map points and pruning redundant critical assemblies.

[0398] Figure 38A and 38B An environment map 3800, created by updating a specification map 700 according to some embodiments, is shown. This specification map 700 can be used from a tracking map 700 with a new tracking map. Figure 7 Upgrade accordingly. For example, compared to... Figure 7 As illustrated and described, the specification map 700 can provide a planar view 706 of reconstructed physical objects in the corresponding physical world, represented by points 702. In some embodiments, map points 702 can represent features of physical objects, which may include multiple features. A new tracking map of the physical world can be captured and uploaded to the cloud for merging with map 700. The new tracking map may include map points 3802 and key assemblies 3804, 3806. In the example shown, key assembly 3804 represents a mapping established, for example, with key assembly 704 in map 700 (e.g., ...). Figure 38B (As shown) a key assembly that has been successfully located on the specification map. On the other hand, key assembly 3806 indicates a key assembly that has not yet been located on map 700. In some embodiments, key assembly 3806 can be promoted to a separate specification map.

[0399] Figures 39A to 39F This is a schematic diagram illustrating an example of a cloud-based persistent coordinate system that provides a shared experience for users in the same physical space. Figure 39A This shows, for example, a specification map 4814 from the cloud. Figures 20A to 20C The XR devices worn by users 4802A and 4802B are received. The specification map 4814 may have a specification coordinate frame 4806C. The specification map 4814 may have a PCF 4810C with multiple associated PPs (e.g., Figure 39C (4818A and 4818B).

[0400] Figure 39BThe relationship established between an XR device and its corresponding world coordinate system and a canonical coordinate frame 4806C is illustrated. This can be accomplished, for example, by locating the tracking map to a canonical map 4814 on the corresponding device. For each device, locating the tracking map to the canonical map results in a transformation between the local world coordinate system and the coordinate system of the canonical map for each device.

[0401] Figure 39C The diagram illustrates transformations (e.g., transformation 4816A, transformation 4816B) between a local PCF (e.g., PCF 4810A, PCF 4810B) on the respective device and a corresponding persistent pose on the canonical map (e.g., PP 4818A, PP 4818B) as a result of localization. Using these transformations, each device can determine, relative to its local PCF, where to display virtual content attached to PP 4818A, PP 4818B, or other persistent points on the canonical map, where the local PCF can be detected locally on the device by processing images detected using sensors on the device. This method allows for accurate localization of virtual content relative to each user and enables each user to have the same experience with virtual content in physical space.

[0402] Figure 39D This shows a persistent pose snapshot from the canonical map to the local tracking map. It can be seen that the local tracking maps are interconnected through persistent poses. Figure 39E It is shown that the PCF 4810A on the device worn by user 4802A can be accessed in the device worn by user 4802B via PP 4818A. Figure 39F The diagram illustrates that tracking maps 4804A, 4804B, and specification map 4814 can be merged. In some embodiments, some PCFs can be removed due to the merging. In the example shown, the merged map includes PCF 4810C of specification map 4814, but excludes PCFs 4810A and 4810B of tracking maps 4804A and 4804B. After the map merging, PPs previously associated with PCFs 4810A and 4810B can be associated with PCF 4810C.

[0403] Example

[0404] Figure 40 and Figure 41 It shows the result of Figure 9 Example of using a tracking map in the first XR device 12.1. Figure 40 It is a three-dimensional first local tracking map (map) according to some embodiments. Figure 1 The two-dimensional representation of ) can be derived from Figure 9 The first XR device was generated. Figure 41 This illustrates a method based on some embodiments. Figure 9 The first XR device uploads the location to the server Figure 1 A block diagram.

[0405] Figure 40 The ground plane on the first XR device 12.1 is shown. Figure 1 And virtual content (content 123 and content 456). Location Figure 1 It has an origin (origin 1). (The rest of the text appears to be a fragment and doesn't translate directly.) Figure 1 This includes many PCFs (PCF a to PCF d). From the perspective of the first XR device 12.1, PCF a is, for example, located on the ground. Figure 1 The origin is located at PCF a with X, Y, and Z coordinates of (0, 0, 0), and PCF b has X, Y, and Z coordinates of (-1, 0, 0). Content 123 is associated with PCF a. In this example, content 123 has X, Y, and Z relationships relative to PCF a (1, 0, 0). Content 456 has a relationship relative to PCF b. In this example, content 456 has X, Y, and Z relationships relative to PCF b (1, 0, 0).

[0406] exist Figure 41 In the middle, the first XR device 12.1 will be on the ground. Figure 1 Uploaded to server 20. In this example, since the server does not store a canonical map for the same region of the physical world represented by the tracking map, the tracking map is stored as the initial canonical map. Server 20 now has a location-based... Figure 1 The specification map. The first XR device 12.1 has a specification map that is empty at this stage. For the purposes of discussion, and in some embodiments, the server 20, in addition to the specification map, has a specification map that is empty at this stage. Figure 1 This excludes other maps. No maps are stored on the second XR device 12.2.

[0407] The first XR device 12.1 also sends its Wi-Fi signature data to server 20. Server 20 can use the Wi-Fi signature data to determine the approximate location of the first XR device 12.1 based on intelligence gathered from other devices that have previously connected to server 20 or other servers along with their recorded GPS locations. The first XR device 12.1 can now terminate the first session (see...). Figure 8 And it can disconnect from server 20.

[0408] Figure 42 It is shown that according to some embodiments Figure 16 The diagram illustrates an XR system where, after the first user 14.1 terminates the first session, the second user 14.2 initiates a second session using the second XR device of the XR system. Figure 43A A block diagram showing the second user 14.2 initiating a second session is shown. Since the first user 14.1's first session has ended, the first user 14.1 is shown as a dashed line. The second XR device 12.2 begins recording objects. The server 20 can use various systems with different granularities to determine that the second session of the second XR device 12.2 is in the same vicinity as the first XR device 12.1's first session. For example, the first XR device 12.1 and the second XR device 12.2 may include Wi-Fi signature data, Global Positioning System (GPS) positioning data, GPS data based on Wi-Fi signature data, or any other data indicating location to record their locations. Alternatively, the PCF identified by the second XR device 12.2 can be displayed in relation to the location. Figure 1 The similarity of PCF.

[0409] like Figure 43B As shown, the second XR device is activated and begins collecting data, such as images 1110 from one or more cameras 44, 46. Figure 14 As shown, in some embodiments, an XR device (e.g., a second XR device 12.2) may collect one or more images 1110 and perform image processing to extract one or more features / points of interest 1120. Each feature may be converted into a descriptor 1130. In some embodiments, the descriptor 1130 may be used to describe a keyframe 1140, which may have the position and orientation of an additional associated image. One or more keyframes 1140 may correspond to a single persistent pose 1150, which may be automatically generated after a threshold distance (e.g., 3 meters) from a previous persistent pose 1150. One or more persistent poses 1150 may correspond to a single PCF 1160, which may be automatically generated after a predetermined distance (e.g., every 5 meters). Over time, as the user continues to move around the user's environment and the XR device continues to collect more data (such as images 1110), additional PCFs (e.g., PCF 3 and PCF 4, 5) may be created. One or more applications 1180 can run on the XR device and provide virtual content 1170 to the XR device for presentation to the user. The virtual content may have an associated content coordinate frame, which can be positioned relative to one or more PCFs. For example... Figure 43B As shown, the second XR device 12.2 creates three PCFs. In some embodiments, the second XR device 12.2 may attempt to locate one or more specification maps stored on server 20.

[0410] In some embodiments, such as Figure 43CAs shown, the second XR device 12.2 can download the specification map 120 from the server 20. The map on the second XR device 12.2... Figure 1 This includes PCF a to d and origin 1. In some embodiments, server 20 may have multiple canonical maps for each location and may determine that the second XR device 12.2 is located in the same vicinity as the first XR device 12.1 during the first session, and send the canonical map of that vicinity to the second XR device 12.2.

[0411] Figure 44 The second XR device 12.2 is shown to begin identifying the PCF for generating the ground. Figure 2 The second XR device 12.2 only recognized a single PCF, namely PCF 1,2. The X, Y, and Z coordinates of PCF 1,2 of the second XR device 12.2 can be (1, 1, 1). Figure 2 It has its own origin (origin 2), which can be based on the head pose of device 2 at the start of the current head pose session. In some embodiments, the second XR device 12.2 can immediately attempt to... Figure 2 Locate the specified map. In some embodiments, because the system cannot recognize any or sufficient overlap between the two maps, the location... Figure 2 It may not be possible to locate on a standard map (land). Figure 1 (i.e., localization may fail). Localization can be performed by identifying a portion of the physical world represented in both a first and a second map, and calculating the transformation between the first and second maps required to align these portions. In some embodiments, the system may perform localization based on a PCF comparison between a local map and a canonical map. In some embodiments, the system may perform localization based on a persistent pose comparison between a local map and a canonical map. In some embodiments, the system may perform localization based on a keyframe comparison between a local map and a canonical map.

[0412] Figure 45 The second XR device 12.2 is shown to identify the location. Figure 2 The other PCFs (PCF 1, 2, PCF 3, PCF 4, 5) followed by the ground Figure 2 The second XR device, 12.2, attempted to [transfer the data] again. Figure 2 Locate on a standard map. Due to the terrain... Figure 2 The local tracking map has been expanded to overlap with at least a portion of the canonical map, therefore the localization attempt will succeed. In some embodiments, the local tracking map ... Figure 2 The overlap between the map and the canonical map can be represented by PCF, persistent pose, keyframes, or any other suitable intermediate or derived construct.

[0413] In addition, the second XR device 12.2 has connected content 123 and content 456 with the ground. Figure 2 PCF 1, 2, and PCF 3 are associated. Content 123 has X, Y, and Z coordinates (1, 0, 0) relative to PCF 1 and 2. Similarly, relative to the ground Figure 2 In PCF 3, the X, Y, and Z coordinates of content 456 are (1, 0, 0).

[0414] Figure 46A and Figure 46B Showing the ground Figure 2 Successful localization to a standard map. Localization can be based on matching features in one map with those in another. This involves appropriate transformations, including translation and rotation of one map relative to another, with the overlapping area / volume / section of map 1410 representing the location. Figure 1 And the common parts of standard maps. Due to the land Figure 2 PCF 3, 4, and 5 were created before localization, while the standard map was created at the location. Figure 2 PCFs a and c were created previously, so different PCFs were created to represent the same volume in real space (e.g., in different maps).

[0415] like Figure 47 As shown, the second XR device 12.2 extends the ground Figure 2 This includes PCF ad from the specification map. PCF ad representation is included. Figure 2 Location to a standardized map. In some embodiments, the XR system may perform optimization steps to remove duplicate PCFs from overlapping areas, such as PCFs 3 and 4, 5 in 1410. Figure 2 Once located, the placement of virtual content (such as content 456 and content 123) will be updated. Figure 2 The closest updated PCF is associated with the virtual content. The virtual content appears in the same real-world location relative to the user, despite changes to the content's PCF attachment, and despite updates to the location. Figure 2 PCF.

[0416] like Figure 48 As shown, the second XR device 12.2 continues to extend... Figure 2 For example, when a user walks around the real world, the second XR device 12.2 will identify other PCFs (PCFs e, f, g, and h). It should also be noted that... Figure 1 exist Figure 47 and Figure 48 There is no extension in it.

[0417] refer to Figure 49 The second XR device 12.2 will be located on the ground. Figure 2Uploaded to server 20. Server 20 will... Figure 2 With standard Figure 1 Storage. In some embodiments, when a session for the second XR device 12.2 ends, the storage... Figure 2 It can be uploaded to server 20.

[0418] The specification map within server 20 now includes PCF i, which is not included on the map of the first XR device 12.1. Figure 1 In the middle. When a third XR device (not shown) uploads a map to server 20 and that map includes PCF i, the canonical map on server 20 may have been expanded to include PCF i.

[0419] exist Figure 50 In the middle, server 20 will be located Figure 2 Merging with the canonical map to form a new canonical map. Server 20 determines PCF a to d for the canonical map and the map. Figure 2 It is common. The server extension specification map includes PCF e to h and from the local area. Figure 2 PCF 1 and 2 are used to form a new specification map. The specification maps on the first XR device 12.1 and the second XR device 12.2 are based on the ground. Figure 1 And it's outdated.

[0420] exist Figure 51 In this process, server 20 sends a new canonical map to the first XR device 12.1 and the second XR device 12.2. In some embodiments, this may occur when the first XR device 12.1 and the second device 12.2 attempt to locate during a different, new, or subsequent session. The first XR device 12.1 and the second XR device 12.2 perform this as described above to retrieve their respective local maps (respectively...) Figure 1 peacefully Figure 2 Locate the new standard map.

[0421] like Figure 52 As shown, the head coordinate frame 96, or "head pose," is related to the ground. Figure 2 The PCF (Potentially Frame Component) is related to this. In some embodiments, the map origin, origin 2, is based on the head pose of the second XR device 12.2 at the start of the session. When the PCF is created during the session, it is positioned relative to the world coordinate frame origin 2. Figure 2 The PCF is used as a persistent coordinate frame relative to the canonical coordinate frame, where the world coordinate frame can be the world coordinate frame of the previous session (e.g., ...). Figure 40 In the land Figure 1 The origin 1). These coordinate frames are used to represent the origin 1). Figure 2 The same transformations are related to the location on the standardized map, as described above. Figure 46B The subject of discussion.

[0422] Previously referenced Figure 9 The transformation from the world coordinate frame to the head coordinate frame 96 was discussed. Figure 52 The head coordinate frame 96 shown has only two orthogonal axes, which are relative to the ground. Figure 2 The PCF is located at a specific coordinate position, and relative to the ground. Figure 2 At a specific angle. However, it should be understood that the head coordinate frame 96 is relative to the ground. Figure 2 The PCF is located in a three-dimensional position and has three orthogonal axes in three-dimensional space.

[0423] exist Figure 53 In the middle, the head coordinate frame 96 is already relative to the ground. Figure 2 The PCF has moved. Since the second user 14.2 has moved its head, the head coordinate frame 96 has moved. The user can move its head in six degrees of freedom (6DOF). The head coordinate frame 96 can therefore move in 6DOF (i.e., from its position in...). Figure 52 The previous position in three dimensions, and relative to the ground Figure 2 The PCF moves around three orthogonal axes. Figure 9 When the real object detection camera 44 and the inertial measurement unit 48 detect the real object and motion of the head unit 22, respectively, the head coordinate frame 96 is adjusted. More information regarding head pose tracking is disclosed in U.S. Patent Application Serial No. 16 / 221,065 entitled “Enhanced Pose Determination for Display Device,” which is incorporated herein by reference in its entirety.

[0424] Figure 54 This illustrates that a sound can be associated with one or more PCFs. A user can, for example, wear a stereo headset or headphones. The sound position through the headphones can be simulated using conventional techniques. The sound position can be fixed such that when the user rotates their head to the left, the sound position rotates to the right, allowing the user to perceive a sound from the same location in the real world. In this example, the sound position is represented by sounds 1, 2, 3 and 4, 5, 6. For ease of discussion, Figure 54 In terms of analysis and Figure 48 Similarity. When the first user 14.1 and the second user 14.2 are in the same room at the same or different times, they perceive sounds 123 and 456 as originating from the same location in the real world.

[0425] Figure 55 and Figure 56 Another implementation of the above technology is shown. See reference... Figure 8 As stated, the first user 14.1 has initiated the first session. Figure 55 As shown in the diagram, the first user 14.1 has terminated the first session, as indicated by the dashed line. At the end of the first session, the first XR device 12.1 will... Figure 1 Uploaded to server 20. First user 14.1 has now initiated a second session at a later time than the first session. Due to... Figure 1 The data is already stored on the first XR device 12.1, therefore the first XR device 12.1 will not download the data from the server 20. Figure 1 If the land is lost Figure 1 Then the first XR device 12.1 downloads the location from server 20. Figure 1 Then, the first XR device 12.1 continued to be built. Figure 2 PCF, located to the ground Figure 1 And further develop the specification map as described above. Then, as described above, the location of the first XR device 12.1 Figure 2 Used to associate local content, header coordinate frames, local audio, etc.

[0426] refer to Figure 57 and Figure 58 It's also possible that more than one user interacts with the server in the same session. In this example, the first user 14.1 and the second user 14.2 are combined with the third user 14.3 and the third XR device 12.3. Each XR device 12.1, 12.2, and 12.3 starts generating its own map, which is respectively the map... Figure 1 ,land Figure 2 peacefully Figure 3 As XR devices 12.1, 12.2, and 12.3 continue to be developed... Figure 1 , 2 At time 3, the map was incrementally uploaded to server 20. Server 20 merged the maps. Figure 1 , 2 The standard map is then generated from server 20 and sent to each of the XR devices 12.1, 12.2, and 12.3.

[0427] Figure 59Aspects of a viewing method for restoring and / or resetting head posture, according to some embodiments, are illustrated. In the example shown, at action 1400, the viewing device is powered on. At action 1410, in response to power-on, a new session is initiated. In some embodiments, the new session may include establishing a head posture. The surface of the environment is captured by one or more capture devices fixed to a head-mounted frame attached to the user's head by first capturing an image of the environment and then determining the surface from the image. In some embodiments, the surface data may be combined with data from a gravity sensor to establish the head posture. Other suitable methods for establishing head posture may be used.

[0428] At action 1420, the viewing device's processor inputs a routine for tracking head posture. As the user moves their head to determine the orientation of the head-mounted frame relative to the surface, the capturing device continues to capture the surface of the environment.

[0429] At action 1430, the processor determines whether head pose has been lost. Head pose may be lost due to "edge" conditions, such as excessive reflective surfaces that can lead to low feature acquisition, low light, blank walls, being outdoors, etc.; or due to dynamic conditions such as movement or crowds forming part of the map. The routine at 1430 allows a certain amount of time, such as 10 seconds, to allow sufficient time to determine whether head pose has been lost. If head pose has not been lost, the processor returns to 1420 and resumes head pose tracking.

[0430] If head pose has been lost at action 1430, the processor enters a routine at 1440 to restore the head pose. If head pose loss is due to low light, a message such as the following will be displayed to the user via the viewing device's monitor:

[0431] The system is detecting low-light conditions. Please move to a better-lit area.

[0432] The system will continue to monitor whether sufficient light is available and whether head posture can be recovered. Alternatively, the system can determine that low surface texture is causing head posture loss, in which case the following prompts will be displayed to the user as suggestions for improving surface capture:

[0433] The system cannot detect enough surfaces with fine textures. Please move to an area with a less coarse texture and a finer texture.

[0434] At action 1450, the processor enters a routine to determine if head pose recovery has failed. If head pose recovery has not failed (i.e., head pose recovery has succeeded), the processor returns to action 1420 by re-entering head pose tracking. If head pose recovery has failed, the processor returns to action 1410 to establish a new session. As part of the new session, all cached data is invalidated, and the head pose is then re-established. Any suitable head tracking method can be used in conjunction with... Figure 59 The process described herein is used in conjunction with that described in U.S. Patent Application No. 16 / 221,065, which describes head tracking and is therefore incorporated herein by reference in its entirety.

[0435] Remote positioning

[0436] Various embodiments can leverage remote resources to facilitate persistent and consistent cross-reality experiences among individuals and / or groups of users. The inventors have recognized and understand that the benefits of operating XR devices using canonical maps as described herein can be achieved without downloading a set of canonical maps. The above-discussed... Figure 30 An example implementation of downloading specifications to a device is shown. For example, the benefit of not downloading maps can be achieved by sending feature and pose information to a remote service that maintains a set of specification maps. According to one embodiment, a device seeking to use the specification maps to position virtual content at a location specified relative to the specification maps can receive one or more transformations between features and the specification maps from the remote service. These transformations can be used on the device, which maintains information about the locations of these features in the physical world, to position the virtual content at a location specified relative to the specification maps, or otherwise identify a location in the physical world relative to the specification maps.

[0437] In some embodiments, spatial information is captured by the XR device and transmitted to a remote service, such as a cloud-based service, which uses the spatial information to locate the XR device relative to a canonical map used by applications or other components of the XR system, thereby specifying the location of the virtual content relative to the physical world. Once located, a transformation linking a tracking map maintained by the device to the canonical map can be transmitted to the device. The transformation can be used in conjunction with the tracking map to determine the location of the rendered virtual content relative to the canonical map, or otherwise to identify the location in the physical world relative to the canonical map.

[0438] The inventors have recognized that the amount of data exchanged between a device and a remote positioning service may be very small compared to the amount of map data being transmitted. This transmission of map data may occur when a device transmits a tracking map to a remote service and receives a set of canonical maps from that service for device-based positioning. In some embodiments, performing positioning functionality on cloud resources requires only a small amount of information to be transmitted from the device to the remote service. For example, it is not necessary to transmit the complete tracking map to the remote service to perform positioning. In some embodiments, feature and pose information, such as that which may be stored in relation to persistent poses as described above, may be transmitted to a remote server. As mentioned above, in embodiments where features are represented by descriptors, the amount of information uploaded may be even smaller.

[0439] The result returned from the location service to the device may be one or more transformations relating the uploaded features to a portion of a matching canonical map. These transformations can be combined with the tracking map in an XR system to identify the location of virtual content or otherwise identify locations in the physical world. In embodiments that use persistent spatial information such as the PCF described above to specify a location relative to a canonical map, the location service may download the transformations between the features and one or more PCFs to the device after successful location.

[0440] As a result, the network bandwidth consumed by communication between the XR device and the remote service used to perform location tracking can be very low. The system can therefore support frequent location tracking, enabling each device interacting with the system to quickly obtain information for locating virtual content or performing other location-based functions. As the device moves in the physical environment, it may repeatedly request updated location information. Furthermore, the device may frequently obtain updates to its location information, such as when the canonical map changes, for example, by incorporating additional tracking maps to expand the map or improve its accuracy.

[0441] Furthermore, uploading features and downloading transformations can enhance privacy in XR systems that share map information among multiple users by increasing the difficulty of obtaining maps through deception. For example, it can prevent unauthorized users from obtaining maps from the system by sending spoofed requests for canonical maps representing portions of the physical world where the unauthorized user is not located. An unauthorized user is less likely to access features in a physical world area of ​​the map information they are requesting if they are not actually present in that area. In embodiments where feature information is formatted as feature descriptions, the difficulty of deceiving feature information in a request for map information will be even more complex. Moreover, when the system returns transformations intended for use on tracking maps of devices operating in the requested location information area, the information returned may be of little or no use to imposters.

[0442] According to one embodiment, the location service is implemented as a cloud-based microservice. In some examples, implementing a cloud-based location service can help save device computing resources and enable the computations required for location to be performed with very low latency. These operations can be supported by virtually unlimited computing power or additional computing resources available by providing additional cloud resources, thereby ensuring the scalability of the XR system to support numerous devices. In one example, many canonical maps can be maintained in memory for near-instantaneous access or stored on a highly available device to reduce system latency.

[0443] Furthermore, performing location tracking on multiple devices within a cloud service can improve the process. Location telemetry and statistics can provide information about which canonical maps reside in active storage and / or high-availability storage. For example, statistics from multiple devices can be used to identify the most frequently accessed canonical maps.

[0444] Additional accuracy can also be achieved as a result of processing in a cloud environment or other remote environments with significantly more processing resources than the remote device. For example, localization can be performed on a higher-density canonical map in the cloud compared to processing performed on a local device. The map can be stored in the cloud, for example, with more PCFs or higher-density feature descriptors per PCF, thereby improving the accuracy of matching a set of features from the device with the canonical map.

[0445] Figure 61 This is a schematic diagram of the XR system 6100. The user device displaying cross-reality content during a user session can take many forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As described above, these devices can be configured with software, such as applications or other components, and / or hardwired to generate local location information (e.g., tracking maps) that can be used to render virtual content on their respective displays.

[0446] Virtual content location information can be specified relative to global location information; for example, the global location information can be formatted as a canonical map containing one or more PCFs. According to some embodiments, system 6100 is configured with cloud-based services that support the operation and display of virtual content on user devices.

[0447] In one example, location functionality is provided as a cloud-based service 6106, which can be a microservice. The cloud-based service 6106 can be implemented on any of multiple computing devices, and computing resources can be allocated from these devices to one or more services executing in the cloud. Those computing devices can interconnect with each other and are accessible to devices such as wearable XR devices 6102 and handheld devices 6104. Such connectivity can be provided through one or more networks.

[0448] In some embodiments, the cloud-based service 6106 is configured to receive descriptor information from various user devices and "locate" the devices to one or more matching canonical maps. For example, the cloud-based location service matches the received descriptor information with the descriptor information of the corresponding canonical map. Canonical maps can be created using techniques described above, which create canonical maps by merging maps provided by one or more devices with image sensors or other sensors acquiring information about the physical world. However, it is not required that the canonical maps be created by the devices accessing them, as such maps can be created by map developers, for example, by making the maps available to the location service 6106.

[0449] According to some embodiments, the cloud service processes canonical map recognition and may include an operation of filtering a repository of canonical maps into a set of potential matches. Filtering may be as follows: Figure 29 To perform as shown, or by using any subset and alternative of the filtering criteria. Figure 29 The filtering criteria shown or in addition to Figure 29 Other filtering criteria besides those shown can be used to perform filtering. In one embodiment, geographic data can be used to limit the search for matching canonical maps to maps representing areas near the device requesting location. For example, regional attributes, such as Wi-Fi signal data, Wi-Fi fingerprint information, GPS data, and / or other device location information, can be used as coarse filters on stored canonical maps, thereby limiting the analysis of descriptors to canonical maps that are known or likely to be near the user's device. Similarly, the location history of each device can be maintained by a cloud service to prioritize searching canonical maps near the device's last location. In some examples, filtering may include the criteria mentioned above regarding... Figure 31B , Figure 32 , Figure 33 and Figure 34 The discussion function.

[0450] Figure 62This is an example process that can be executed by a device to use a cloud-based service to locate the device's position using a canonical map and receive transformation information between one or more transformations between a specified device local coordinate system and the coordinate system of the canonical map. Various embodiments and examples describe the transformation as a transformation specifying a transformation from a first coordinate frame to a second coordinate frame. Other embodiments include a transformation from a second coordinate frame to a first coordinate frame. In any other embodiment, the transformation implements a transition from one coordinate frame to another, the resulting coordinate frame depending only on the desired coordinate frame output (including, for example, the coordinate frame in which content is displayed). In yet another embodiment, the coordinate system transformation enables the determination of transformations from a second coordinate frame to a first coordinate frame and from a first coordinate frame to a second coordinate frame.

[0451] According to some embodiments, information reflecting changes in each persistent pose defined by the canonical map can be transmitted to the device.

[0452] According to one embodiment, process 6200 can begin with a new session at 6202. Starting a new session on the device can initiate the capture of image information to build a tracking map of the device. Furthermore, the device can send a message to register with the location service server, prompting the server to create a session for the device.

[0453] In some embodiments, initiating a new session on the device may optionally include sending adjustment data from the device to a positioning service. The positioning service returns to the device one or more transformations calculated based on a set of features and associated poses. Device-specific information may be sent to the positioning service so that it can apply these adjustments, if the poses of features are adjusted based on device-specific information before and / or the transformations are adjusted based on device-specific information after the transformations are calculated, rather than performing those calculations on the device. As a particular example, sending device-specific adjustment information may include capturing calibration data for sensors and / or the display. The calibration data can be used, for example, to adjust the position of feature points relative to a measurement location. Alternatively or additionally, the calibration data can be used to adjust the position in which the display renders virtual content so that it appears accurately positioned for that particular device. This calibration data may be obtained, for example, from multiple images of the same scene taken using sensors on the device. The positions of features detected in those images may be represented as a function of sensor positions, such that the multiple images produce a set of equations from which the sensor positions can be solved. The calculated sensor positions can be compared to nominal positions, and calibration data can be derived from any differences. In some embodiments, intrinsic information about the device construction may also enable the calculation of calibration data for the display, in some implementations.

[0454] In embodiments that generate calibration data for sensors and / or displays, the calibration data can be applied to any point in the measurement or display process. In some embodiments, the calibration data can be sent to a positioning server, which can store the calibration data in a data structure established for each device, each of which has registered with the positioning server and is therefore in a session with the server. The positioning server can apply the calibration data to any transformations calculated as part of the positioning process for the device providing the calibration data. Thus, the computational burden of using calibration data to improve the accuracy of sensing and / or display information is borne by the calibration service, thereby providing a further mechanism to reduce the processing burden on the device.

[0455] Once a new session is established, process 6200 can continue capturing new frames of the device's environment at 6204. At 6206, each frame can be processed to generate a descriptor for the captured frame (including, for example, the DSF value discussed above). These values ​​can be calculated using some or all of the techniques described above, including those mentioned above regarding... Figure 14 , Figure 22 and Figure 23 The techniques discussed. As discussed, a descriptor can be computed as a mapping of feature points, or in some embodiments, a mapping of image patches around feature points to a descriptor. The descriptor can have values ​​capable of efficient matching between newly acquired frames / images and stored maps. Furthermore, the number of features extracted from an image can be limited to a maximum number of feature points per image, such as 200 feature points per image. As mentioned above, feature points can be selected to represent points of interest. Therefore, actions 6204 and 6206 can be performed as part of a device process for forming a tracking map or otherwise periodically collecting images of the physical world around the device, or can but not necessarily be performed separately for localization.

[0456] Feature extraction at 6206 may include appending pose information to the features extracted at 6206. The pose information may be a pose in the device's local coordinate system. In some embodiments, the pose may be relative to a reference point in the tracking map, such as a persistent pose as described above. Alternatively or additionally, the pose may be relative to the origin of the device's tracking map. Such embodiments enable the positioning service, as described herein, to provide positioning services to a wide range of devices, even those that do not use a persistent pose. In any case, pose information may be appended to each feature or each set of features, allowing the positioning service to use the pose information to compute transformations that can be returned to the device when matching features with features in a stored map.

[0457] Process 6200 can continue to decision box 6207, where a decision is made on whether to request location. One or more criteria can be applied to determine whether to request location. These criteria may include the elapsed time, allowing the device to request location after a certain threshold time period. For example, if no location attempt is made within the threshold time period, the process can continue from decision box 6207 to action 6208, where location is requested from the cloud. This threshold time period can be between 10 and 30 seconds, for example, 25 seconds. Alternatively or additionally, location can be triggered by the device's movement. The device performing process 6200 can use its IMU and its tracking map to track its movement and initiate location when movement beyond a threshold distance is detected from the location where the device was last requested to be located. For example, the threshold distance can be between 1 and 10 meters, for example, between 3 and 5 meters. As yet another alternative, location can be triggered in response to an event, such as when the device creates a new persistent pose or when the device's current persistent pose changes, as described above.

[0458] In some embodiments, decision box 6207 can be implemented to dynamically establish a threshold for triggering localization. For example, in environments where features are largely consistent, the confidence in matching a set of extracted features with features from a stored map may be low, leading to more frequent localization requests to increase the chances that at least one localization attempt will succeed. In this case, the threshold applied to decision box 6207 can be reduced. Similarly, in environments with relatively few features, the threshold applied to decision box 6207 can be reduced to increase the frequency of localization attempts.

[0459] Regardless of how positioning is triggered, when triggered, process 6200 can proceed to action 6208, whereby the device sends a request to a positioning service, including data used by the positioning service to perform positioning. In some embodiments, data from multiple image frames may be provided for the positioning attempt. For example, the positioning service may not consider positioning successful unless features from multiple image frames produce consistent positioning results. In some embodiments, process 6200 may include saving feature descriptors and additional pose information to a buffer. The buffer may be, for example, a circular buffer storing feature sets extracted from the most recently captured frames. Thus, the positioning request may be sent along with multiple feature sets accumulated in the buffer. In some setups, the buffer size is implemented to accumulate multiple datasets that are more likely to produce successful positioning. In some embodiments, the buffer size may be set to accumulate features from, for example, two, three, four, five, six, seven, eight, nine, or ten frames. Optionally, the buffer size may have a baseline setting that can be increased in response to positioning failures. In some examples, increasing the buffer size and the corresponding number of feature sets transmitted reduces the likelihood that subsequent positioning functions will fail to return results.

[0460] Regardless of the buffer size, the device can transmit the buffer's contents to the location service as part of a location request. Other information can be transmitted along with feature points and additional pose information. For example, in some embodiments, geographic information can be transmitted. Geographic information may include, for example, GPS coordinates or a wireless signature or current persistent pose associated with the device tracking the map.

[0461] In response to the request sent at 6208, the cloud-based location service can analyze feature descriptors to locate the device in a canonical map or other persistent map maintained by the service. For example, the descriptor matches a set of features in the map to which the device is located. The cloud-based location service can perform positioning as described above relative to device-based positioning (e.g., it can rely on any of the positioning functions discussed above, including map ranking, map filtering, location estimation, filtered map selection, etc.). Figures 44 to 4 (The example in 6, and / or discussed relative to the positioning module, PCF, and / or PP identification and matching, etc.). However, instead of transmitting the identified canonical map to the device (e.g., in device positioning), the cloud-based positioning service can continue to generate transformations based on the matching features of the canonical map and the relative orientation of the feature set sent from the device. The positioning service can then return these transformations to the device, which can receive them at box 6210.

[0462] In some embodiments, the canonical map maintained by the location service may employ a PCF, as described above. In such embodiments, feature points in the canonical map that match feature points sent from the device may have positions specified relative to one or more PCFs. Therefore, the location service can identify one or more canonical maps and can calculate the transformation between the coordinate frame represented in the pose sent with the location request and one or more PCFs. In some embodiments, the identification of one or more canonical maps is aided by filtering potential maps based on the geographic data of the respective device. For example, once filtered to a candidate set (e.g., by other options such as GPS coordinates), the candidate set of canonical maps can be analyzed in detail to determine matching feature points or PCFs as described above.

[0463] The data returned to the requesting device in action 6210 can be formatted as a persistent pose transformation table. This table may be accompanied by one or more canonical map identifiers indicating the canonical map to which the device has been located by the positioning service. However, it should be understood that the positioning information can be formatted in other ways, including as a transformation list with associated PCFs and / or canonical map identifiers.

[0464] Regardless of how the transformations are formatted, in action 6212, the device can use these transformations to calculate the position of the rendered virtual content, the position of which has been specified by the XR system's application or other components relative to any PCF. This information is alternatively or additionally used on the device to perform any position-based operations where the position is specified based on the PCF.

[0465] In some scenarios, the location service may fail to match the features sent from the device to any stored canonical map, or it may fail to match a sufficient number of feature sets transmitted with the request to the location service to consider the location successful. In such scenarios, the location service may indicate a location failure to the device instead of returning a transformation to the device as described above in conjunction with action 6210. In such scenarios, process 6200 may branch to action 6230 at decision box 6209, where the device may take one or more actions for failure handling. These actions may include increasing the size of the buffer that stores the feature sets sent for location. For example, if the location service does not consider the location successful unless three feature sets match, the buffer size may be increased from 5 to 6, thereby increasing the chance that the three transmitted feature sets match the canonical map maintained by the location service.

[0466] Alternatively or additionally, failure handling may include adjusting the device's operating parameters to trigger more frequent location attempts. For example, the threshold time and / or threshold distance between location attempts could be reduced. As another example, the number of feature points in each feature set could be increased. A match between the feature set and features stored in the canonical map is considered to have occurred when a sufficient number of features from the set sent by the device match features in the map. Increasing the number of features sent increases the chance of a match. As a specific example, the initial feature set size could be 50, which could be increased to 100, 150, and then 200 with each consecutive location failure. After a successful match, the set size could then be returned to its initial value.

[0467] Failure handling may also include obtaining location information from sources other than location services. According to some embodiments, the user device can be configured to cache canonical maps. Cached maps allow the device to access and display content unavailable in the cloud. For example, cached canonical maps allow for device-based positioning in the event of communication failures or other unavailability.

[0468] According to various embodiments, Figure 62 A high-level process for device-initiated cloud-based location is described. In other embodiments, various steps, one or more, may be combined, omitted, or other processes may be invoked to complete the location and final visualization of virtual content in the corresponding device view.

[0469] Furthermore, it should be understood that although process 6200 shows the device determining whether to initiate location at decision box 6207, the trigger for initiating location can originate from outside the device, including from a location service. For example, a location service may maintain information about each device in its session. This information may include, for example, an identifier of the canonical map where each device was recently located. The location service or other components of the XR system may update the canonical map, including using the above-mentioned combination... Figure 26 The described technology. When the canonical map is updated, the location service can send a notification to each device recently located on the map. This notification can serve as a trigger for a device to request location and / or may include an updated transformation recalculated using a set of features recently sent from the device.

[0470] Figure 63A , Figure 63B and Figure 63C This illustrates an example process flow of operation and communication between a device and a cloud service. Boxes 6350, 6352, 6354, and 6456 show the example architecture and the separation between components involved in the cloud-based positioning process. For example, modules, components, and / or software configured to handle perception on a user device are shown at 6350 (e.g., 660, ...). Figure 6A The device functions for persistent world operations are shown at 6352 (including, for example, as described above and regarding the persistent world module (e.g., 662, ...). Figure 6A In other embodiments, separation between 6350 and 6352 is not required, and the communication shown can occur between processes performed on the device.

[0471] Similarly, shown at box 6354 is a cloud process (e.g., 802, 812) configured to handle functions associated with walkable world / walkable world modeling. Figure 26 Box 6356 shows a cloud process configured to handle functions associated with locating a device to one or more maps in a repository of stored canonical maps based on information sent from the device.

[0472] In the illustrated embodiment, process 6300 begins at 6302, at which point a new session starts. Sensor calibration data is obtained at 6304. The obtained calibration data may depend on the device represented at 6350 (e.g., multiple cameras, sensors, positioning devices, etc.). Once sensor calibration is obtained for the device, the calibration can be cached at 6306. If device operation causes changes in frequency parameters (e.g., collection frequency, sampling frequency, matching frequency, and other options), the frequency parameters are reset to the baseline at 6308.

[0473] Once the new session functionality is complete (e.g., calibration, steps 6302-6306), process 6300 can proceed to capture new frames 6312. At 6314, features and their corresponding descriptors are extracted from the frames. In some examples, the descriptors may include DSFs, as described above. According to some embodiments, the descriptors may have spatial information attached to them for subsequent processing (e.g., transform generation). At 6316, pose information generated on the device (e.g., as described above, information for locating features in the physical world relative to the device's tracking map) can be attached to the extracted descriptors.

[0474] At 6318, descriptor and pose information are added to a buffer. New frame captures and additions to the buffer, as shown in steps 6312-6318, are performed cyclically until a buffer size threshold is exceeded at 6319. At 6320, in response to determining that the buffer size is met, a location request is transmitted from the device to the cloud. According to some embodiments, this request may be handled by a navigable world service instantiated in the cloud (e.g., 6354). In a further embodiment, the functional operations for identifying candidate canonical maps may be separated from the operations for actual matching (e.g., shown as boxes 6354 and 6356). In one embodiment, a cloud service for map filtering and / or map ranking may be performed at 6354 and process the location request received from 6320. According to one embodiment, the map ranking operation is configured at 6322 to determine a candidate map set that may include the device's location.

[0475] In one example, the map ranking function includes operations for identifying candidate canonical maps based on geographic attributes or other location data, such as observed or inferred location information. Other location data could include Wi-Fi signatures or GPS information.

[0476] According to other embodiments, location data can be captured during cross-reality sessions with the device and the user. Process 6300 may include additional operations such as populating the location for a given device and / or session (not shown). For example, the location data may be stored as device region attribute values ​​and attribute values ​​for selecting candidate canonical maps of locations near the device.

[0477] Any one or more location options can be used to filter canonical map atlases to those that may represent areas including the location of the user's device. In some embodiments, the canonical map may cover a relatively large area of ​​the physical world. The canonical map may be segmented into regions, such that map selection may require the selection of map regions. For example, map regions may be on the order of tens of square meters. Therefore, the filtered canonical map atlas may be a set of map regions.

[0478] According to some embodiments, a positioning snapshot can be constructed from candidate specification maps, pose features, and sensor calibration data. For example, an array of candidate specification maps, pose features, and sensor calibration information can be sent along with a request to determine a specific matching specification map. Matching with the specification map can be performed based on descriptors received from the device and stored PCF data associated with the specification map.

[0479] In some embodiments, a feature set from the device is compared with a feature set stored as part of a canonical map. This comparison may be based on feature descriptors and / or poses. For example, a candidate feature set for the canonical map may be selected based on the number of features in a candidate set whose descriptors are sufficiently similar to those in the feature set from the device to suggest that they are likely the same features. For example, the candidate set may be features derived from image frames used to form the canonical map.

[0480] In some embodiments, if the number of similar features exceeds a threshold, further processing can be performed on the candidate feature set. This further processing can determine the extent to which the pose feature set from the device can be aligned with the candidate feature set. Poseming can be performed on the feature set from the canonical map (similar to features from the device).

[0481] In some embodiments, features are formatted as high-dimensional embeddings (e.g., DSF, etc.) and can be compared using nearest neighbor search. In one example, the system is configured (e.g., by performing procedures 6200 and / or 6300) to find the first two nearest neighbors using Euclidean distance, and a ratio test can be performed. If a nearest neighbor is closer than the second nearest neighbor, the system considers the nearest neighbor a match. For example, "closer" in this context can be determined by the ratio of the Euclidean distance to the second nearest neighbor exceeding a threshold multiple. Once features from the device are considered to "match" features in the canonical map, the system can be configured to compute a relative transformation using the pose of the matching features. The transformation derived from the pose information can be used to indicate the transformation required to position the device onto the canonical map.

[0482] The number of inliers can be used as an indicator of matching quality. For example, in the case of DSF matching, the number of inliers reflects the number of features that match between the received descriptor information and the stored / canonical map. In another embodiment, the inlier determined in this embodiment can be determined by counting the number of “matched” features in each set.

[0483] An indication of matching quality can be determined alternatively or additionally in other ways. In some embodiments, for example, when a transformation is calculated to localize a map from the device, which may contain multiple features, to a canonical map based on the relative pose of the matching features, the transformation statistics calculated for each of the multiple matching features can serve as a quality indicator. For example, a large difference can indicate poor matching quality. Alternatively or additionally, for a given transformation, the system can calculate an average error between features with matching descriptors. An average error can be calculated for the transformation, reflecting the degree of location mismatch. Mean squared error is a specific example of an error metric. Regardless of the specific error metric, if the error is below a threshold, it can be determined that the transformation is available for features received from the device, and the calculated transformation is used to locate the device. Alternatively or additionally, the number o...

Claims

1. An electronic device configured to operate within a cross-reality system, the electronic device comprising: One or more sensors are configured to capture information about a three-dimensional 3D environment, the captured information including multiple images of the 3D environment; At least one processor is configured to execute computer-executable instructions, wherein the computer-executable instructions include instructions for the following operations: Generate a local coordinate frame to represent the position in the 3D environment; Extract multiple features from the multiple images of the 3D environment; Sending information about a subset of the plurality of features and the location information of the subset of the plurality of features expressed in the local coordinate frame to a remote positioning service via a network; and In response to sending the information about the subset of the plurality of features and the location information, at least one transformation is received from the remote positioning service, the at least one transformation associating the local coordinate frame used to represent the location in the 3D environment with a second coordinate frame used to render virtual content having a location specified in the second coordinate frame.

2. The electronic device according to claim 1, wherein: The computer-executable instructions further include instructions for: applying a transformation of the at least one transformation to the position of virtual content specified in the second coordinate frame to calculate the position in the 3D environment specified in the local coordinate frame for rendering the virtual content.

3. The electronic device according to claim 1, wherein: The electronic device includes a display; and The computer-executable instructions also include instructions for rendering virtual content having the position specified in the second coordinate frame on the display at a position calculated at least in part based on a transformation in the at least one transformation.

4. The electronic device according to claim 1, wherein: The computer-executable instructions further include: instructions for generating descriptors for the plurality of features; and Sending the information about the subset of the plurality of features includes sending a corresponding descriptor for the subset of the plurality of features.

5. The electronic device according to claim 4, wherein: The computer-executable instructions further include: instructions for storing the information about the subset of the plurality of features and the location information of the subset of the plurality of features in a buffer; and Sending over the network includes sending the contents of the buffer together, such that the information about the subset of the plurality of features and the location information about the subset of the plurality of features are sent together.

6. The electronic device according to claim 5, wherein: The buffer includes an adjustable size; The computer-executable instructions also include instructions for increasing the size of the buffer in response to a failure indication received from the remote location service via the network.

7. The electronic device according to claim 6, wherein: Extracting the multiple features from the multiple images includes: extracting up to a threshold number of features from each image; and The computer-executable instructions further include: instructions for increasing the number of thresholds in response to a failure indication received from the remote location service via the network.

8. The electronic device according to claim 1, wherein, The computer-executable instructions further include: instructions for sending a request to initiate a session with the remote location service via the network.

9. The electronic device according to claim 8, wherein, The request to initiate a session with the remote location service includes: the identifier of the electronic device, and calibration data for the electronic device.

10. The electronic device according to claim 1, wherein, The information sent regarding the subset of the plurality of features and the location information for the subset of the plurality of features are not in the map.

11. The electronic device according to claim 1, wherein: The information about the subset of the plurality of features and the location information about the subset of the plurality of features transmitted through the network include: a request for a location service; and The computer-executable instructions also include instructions for sending a location request based on one or more triggering conditions being met.

12. The electronic device according to claim 11, wherein, The one or more triggering conditions include: the distance the electronic device has moved since the last successful location request.

13. A method for operating an electronic device, the electronic device being configured to operate within a cross-reality system, the method comprising: Receive information about a three-dimensional (3D) environment, the information including multiple images of the 3D environment; Generate a local coordinate frame to represent the position in the 3D environment; Extract multiple features from the multiple images of the 3D environment; The system sends information about a subset of the plurality of features and the location information of the subset of the plurality of features expressed in the local coordinate frame to a remote positioning service via a network. as well as In response to sending the information about the subset of the plurality of features and the location information, at least one transformation is received from the remote positioning service, the at least one transformation associating the local coordinate frame used to represent the location in the 3D environment with a second coordinate frame used to render virtual content having a location specified in the second coordinate frame.

14. The method of claim 13, wherein: The local coordinate frame is used to represent the position in the 3D environment. The second coordinate frame is used to represent the position of the virtual content, and The method includes: applying a transformation of the at least one transformation to the position of virtual content specified in the second coordinate frame to calculate the position in the 3D environment specified in the local coordinate frame for rendering the virtual content.

15. The method of claim 13, further comprising: Virtual content having a position specified in the second coordinate frame is rendered on the display of the electronic device at a position calculated at least in part based on a transformation in at least one of the at least one transformation.

16. The method according to any one of claims 13 to 15, wherein: The method further includes: generating descriptors for the plurality of features; and Sending information about a subset of the plurality of features includes sending the descriptors for the plurality of features.

17. The method of claim 16, wherein: The method further includes: storing the information about a subset of the plurality of features and the location information of the subset of the plurality of features in a buffer; and Sending over the network includes sending the contents of the buffer together, such that the information about a subset of the plurality of features and the location information about the subset of the plurality of features are sent together.

18. The method of claim 17, wherein: The buffer includes an adjustable size; The method further includes increasing the size of the buffer in response to a failure indication received from the remote location service via the network.

19. The method of claim 18, wherein: Extracting the multiple features from the multiple images includes: extracting up to a threshold number of features from each image; and The method further includes: increasing the number of thresholds in response to a failure indication received from the remote location service via the network.

20. The method according to claim 13, wherein, The method further includes: sending a request to initiate a session with the remote positioning service via the network, wherein the request to initiate a session with the remote positioning service includes an identifier of the electronic device and calibration data for the electronic device.

21. The method according to claim 13, wherein: The information about the subset of the plurality of features and the location information about the subset of the plurality of features sent through the network include requests for location services; as well as The method further includes sending a location request based on one or more triggering conditions being met.

22. The method according to claim 21, wherein, The one or more triggering conditions include: the distance the electronic device has moved since the last successful location request.

23. The method according to any one of claims 20 to 22, wherein: The method further includes: generating descriptors for the plurality of features; and Sending information about the subset of the plurality of features includes sending the descriptor for the plurality of features.

24. A computer-readable medium storing computer-executable instructions configured to perform a method for operating an electronic device when executed by at least one processor, the electronic device being configured to operate within a cross-reality system, the method comprising: Receive information about a three-dimensional (3D) environment, the information including multiple images of the 3D environment; Generate a local coordinate frame to represent the position in the 3D environment; Extract multiple features from the multiple images of the 3D environment; The system sends information about a subset of the plurality of features and the location information of the subset of the plurality of features expressed in the local coordinate frame to a remote positioning service via a network. as well as In response to sending the information about the subset of the plurality of features and the location information, at least one transformation is received from the remote positioning service, the at least one transformation associating the local coordinate frame used to represent the location in the 3D environment with a second coordinate frame used to render virtual content having a location specified in the second coordinate frame.

25. The computer-readable medium of claim 24, wherein: The local coordinate frame is used to represent the position in the 3D environment. The second coordinate frame is used to represent the position of the virtual content, and The method includes: applying a transformation of the at least one transformation to the position of virtual content specified in the second coordinate frame to calculate the position in the 3D environment specified in the local coordinate frame for rendering the virtual content.

Citation Information

Patent Citations

  • Localization determination for mixed reality systems

    US10812936B2

  • Fully convolutional interest point detection and description via homographic adaptation

    US20190147341A1

  • Enhanced pose determination for display device

    US20190188474A1

  • Methods and apparatuses for determining and / or evaluating localizing maps of image display devices

    US20200034624A1

  • Apparatuses, methods and systems for sharing virtual elements

    US20170237789A1