Cross-Reality System for Map Processing Using Multi-Resolution Frame Descriptors
Through network resources and multi-resolution frame descriptor technology in distributed computing environments, the problem of existing XR systems positioning and sharing virtual content in large-scale environments is solved, and efficient position positioning and computing efficiency improvement is achieved.
Patent Information
- Application Number
- CN202180027922.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-13
- Filing Date
- 2021-02-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-02-11
AI Technical Summary
Existing Cross-Reality (XR) systems have difficulty efficiently positioning and sharing location-based virtual content when creating or using maps, especially in large-scale environments.
Through network resources in a distributed computing environment, shared location-based content is provided, and using multi-resolution frame descriptor technology, the image frame descriptor from the portable electronic device is compared with the frame descriptor in the storage map, and then the appropriate map is selected to locate the device.
It realizes efficient positioning and sharing of virtual content in an inter-realistic system, and improves the system's position accuracy and computing efficiency in large-scale environments.
Smart Images

Figure CN115398314B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 976,240, filed on February 13, 2020, entitled "CROSSREALITY SYSTEM WITH MAP PROCESSING USING MULTI - RESOLUTION FRAME DESCRIPTORS", which is hereby incorporated by reference in its entirety. Technical field
[0003] This application generally relates to cross - reality systems. Background art
[0004] Computers can control human - user interfaces to create cross - reality (XR) environments, in which some or all of the XR environment perceived by the user is generated by the computer. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments, and some or all of these XR environments can be generated by the computer using, in part, data that describes the environment. For example, the data can describe virtual objects, which can be rendered in a way that the user senses or perceives them as part of the physical world and can interact with the virtual objects. Since the data is rendered and presented through a user - interface device (such as, for example, a head - mounted display device), the user can experience these virtual objects. The data can be shown to the user, or it can control the audio played to the user, or it can control a haptic (or tactile) interface, enabling the user to experience the feeling of touching a virtual object that the user senses or perceives as being felt.
[0005] XR systems can be used in many applications across the fields of scientific visualization, medical training, engineering design and prototyping, tele - operation and tele - presence, and personal entertainment. Compared to VR, AR and MR include one or more virtual objects related to real objects in the physical world. The experience of virtual objects interacting with real objects significantly enhances the user's enjoyment of using XR systems and also opens the door to various applications for presenting information about how to change the physical world in a realistic and easily understandable way.
[0006] To realistically render virtual content, an XR system can build a representation of the physical world of the user surrounding the system. For example, this representation can be constructed by processing images obtained using sensors on wearable devices, where the wearable devices form part of the XR system. In such a system, a user can perform an initialization routine by looking around the room or other physical environment in which the user intends to use the XR system until the system has obtained sufficient information to build a representation of that environment. As the system operates and the user moves within the environment or moves to other environments, sensors on the wearable devices can obtain additional information to extend or update the representation of the physical world. SUMMARY OF THE INVENTION
[0007] Aspects of the present application relate to methods and apparatus for creating or using maps in a cross-reality (XR) system. The techniques described herein can be used together, separately, or in any suitable combination.
[0008] According to one aspect, a network resource in a distributed computing environment is provided for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a three-dimensional (3D) environment. The network resource includes: one or more processors, and at least one computer-readable medium including a plurality of stored maps of the 3D environment. The medium further includes computer-executable instructions. When executed by the one or more processors, these instructions cause the network resource to: receive information about a plurality of features detected in an image captured by the portable electronic device, and compute a frame descriptor for the image, wherein the computed frame descriptor has a resolution greater than 512 bits.
[0009] According to some embodiments, at least one of the plurality of stored maps of the 3D environment is associated with at least one frame descriptor having a resolution greater than 512 bits.
[0010] According to some embodiments, the computer-executable instructions, when executed by one or more processors, further cause the network resource to: compare the computed image descriptor with the at least one frame descriptor associated with at least one of the plurality of stored maps of the 3D environment.
[0011] According to some embodiments, the computer-executable instructions, when executed by one or more processors, further cause the network resource to: select one or more maps from the plurality of stored maps to localize the portable electronic device to a shared coordinate system based on a comparison of the computed image frame descriptor with the at least one frame descriptor associated with at least one of the plurality of stored maps.
[0012] According to some embodiments, when the computer-executable instructions are executed by a processor of the one or more processors, the computer-executable instructions further cause the network resource to: send one or more selected maps to the portable electronic device.
[0013] According to some embodiments, when the computer-executable instructions are executed by a processor of the one or more processors, the computer-executable instructions further cause the network resource to: determine whether the location of the portable electronic device corresponds to a stored map of the plurality of stored maps from the 3D environment based on a comparison of the calculated image frame descriptor and the at least one frame descriptor associated with at least one of the plurality of stored maps.
[0014] According to some embodiments, when the computer-executable instructions are executed by a processor of the one or more processors, the computer-executable instructions further cause the network resource to: receive a tracking map from the portable electronic device; and merge the tracking map with the stored map to generate a merged map including location information based on the location information of the stored map and the tracking map. The instructions then cause the network resource to: store the merged map in the computer-readable medium.
[0015] According to some embodiments, the portable electronic device is selected from the group consisting of: a wearable device including a head-mounted display having a plurality of cameras mounted thereon; and a portable computing device (e.g., a smart phone or a tablet) including a camera and a display and configured with computer-executable instructions for rendering virtual content related to an image acquired by the camera on the display.
[0016] According to some embodiments, the computer-executable instructions include instructions for implementing a neural network to calculate a frame descriptor.
[0017] According to one aspect, a method is provided for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a 3D environment. The method includes, on the portable electronic device: obtaining one or more images of the 3D environment; identifying one or more features from the one or more images; transmitting information about the one or more features identified in the one or more images of the 3D environment to a network resource; and calculating, based on the one or more features, one or more first frame descriptors of the one or more images, the first frame descriptors having a first resolution. The method further includes, on the network resource: storing a plurality of maps of the 3D environment; and calculating, based on the information about the one or more features, one or more second frame descriptors representing the one or more images, wherein the one or more second frame descriptors have a second resolution greater than the first resolution.
[0018] According to some embodiments, the method further includes, on the portable electronic device, selecting at least a portion of a local map based on the one or more first frame descriptors; and, on the network resource, selecting at least a portion of a shared map based on the one or more second frame descriptors.
[0019] According to some embodiments, the step of selecting at least a portion of the shared map includes comparing the one or more second frame descriptors with one or more frame descriptors associated with the plurality of maps of the 3D environment.
[0020] According to some embodiments, the method further includes, on the network resource, determining, based on a comparison of the one or more second frame descriptors with one or more frame descriptors associated with the plurality of maps of the 3D environment, one or more maps for positioning the portable electronic device in a shared coordinate system.
[0021] According to some embodiments, the method further includes, for the one or more maps determined for positioning the portable electronic device, calculating one or more third frame descriptors, wherein the one or more third frame descriptors have the first resolution; and transmitting, from the network resource to the portable electronic device, the one or more maps determined for positioning the portable electronic device and the one or more third frame descriptors.
[0022] According to some embodiments, the method further includes: on the network resource, determining whether the location of the portable electronic device corresponds to a stored map from the plurality of stored maps of the 3D environment using a comparison of one or more frame descriptors calculated based on information about a pixel group received from the portable electronic device and one or more frame descriptors associated with the plurality of maps of the 3D environment.
[0023] According to some embodiments, the information about the one or more features includes a tracking map, and the method further includes: at the network resource, merging the tracking map with the stored map to generate a merged map including location information based on the location information of the stored map and the tracking map; and at the network resource, storing the merged map together with the one or more second frame descriptors in the computer-readable medium.
[0024] According to some embodiments, the portable electronic device is selected from the group consisting of: a wearable device including a head-mounted display having a plurality of cameras mounted thereon; and a portable computing device including a camera and a display and configured with computer-executable instructions for rendering virtual content related to an image acquired by the camera on the display.
[0025] According to some embodiments, the action of calculating one or more first frame descriptors and / or the action of calculating one or more second frame descriptors is performed using a neural network.
[0026] According to one aspect, a system for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a 3D environment is provided. The system includes: at least one portable electronic device configured to render virtual content, and at least one network resource. Each of the at least one portable electronic devices includes at least one processor, at least one camera, and at least one computer-readable medium including instructions that, when executed, cause the at least one processor to perform: capturing at least one image of the 3D environment with the at least one camera; for the at least one image, identifying multiple sets of pixels representing features; for the multiple sets of pixels, calculating descriptors representing the multiple sets of pixels; creating a data structure including the descriptors calculated for the multiple sets of pixels; sending the data structure to the network resource; calculating at least one first frame descriptor based on the descriptors calculated for the multiple sets of pixels, the at least one first frame descriptor having a first resolution; and comparing image frames local to the portable electronic device based on the at least one first frame descriptor having the first resolution. The at least one network resource includes one or more processors and at least one computer-readable medium, the at least one computer-readable medium including: multiple stored maps of the 3D environment, wherein at least one of the multiple stored maps is associated with at least one frame descriptor. The at least one network resource further includes: computer-executable instructions that, when executed by the one or more processors, cause the network resource to: calculate at least one second frame descriptor using a neural network based on the descriptors calculated for the multiple sets of pixels, wherein the at least one second frame descriptor has a second resolution higher than the first resolution; compare the at least one second frame descriptor with the at least one frame descriptor associated with the at least one stored map of the 3D environment; and position the portable electronic device in a shared coordinate system based on the comparison of the at least one second frame descriptor and the at least one frame descriptor associated with the at least one stored map.
[0027] According to some embodiments, the at least one portable electronic device is selected from the group including: a wearable device including a head-mounted display having a plurality of cameras mounted thereon; and a portable computing device (e.g., a smart phone or a tablet) including a camera and a display and configured with computer-executable instructions for rendering virtual content related to an image acquired by the camera on the display.
[0028] According to some embodiments, the computer-executable instructions of the at least one network resource further include computer-executable instructions that, when executed by the one or more processors, cause the network resource to: receive a tracking map from a portable electronic device among the at least one portable electronic device; and merge the tracking map with the at least one stored map to generate a merged map including location information based on the location information of the at least one stored map and the tracking map. The instructions further cause the resource to: store the merged map in the computer-readable medium.
[0029] The foregoing summary is provided by way of illustration and not by way of limitation. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings are not necessarily to scale. In the drawings, each identical or nearly identical component that is illustrated in various drawings is represented by a like numeral. For clarity, not every component may be labeled in every drawing. In the drawings:
[0031] Figure 1 is a schematic diagram showing an example of a simplified augmented reality (AR) scenario according to some embodiments;
[0032] Figure 2 is a schematic diagram of an exemplary simplified AR scenario according to some embodiments, showing an exemplary usage scenario of an XR system;
[0033] Figure 3 is a schematic diagram showing a data flow for a single user in an AR system according to some embodiments, the AR system being configured to provide a user with an experience of interacting with the physical world with AR content;
[0034] Figure 4 is a schematic diagram showing an exemplary AR display system according to some embodiments, the exemplary AR display system displaying virtual content for a single user;
[0035] Figure 5A is a schematic diagram according to some embodiments showing that when a user wears an AR display system, the AR display system renders AR content when the user moves through a physical world environment;
[0036] Figure 5B is a schematic diagram showing a viewing optical component and accompanying components according to some embodiments;
[0037] Fig. 6A is a schematic diagram of an AR system using a world reconstruction system according to some embodiments;
[0038] Figure 6BIt is a schematic diagram showing components of an AR system that maintains a traversable world model according to some embodiments.
[0039] Figure 7 It is a schematic diagram of a tracking map formed by a path traversed by a device through the physical world.
[0040] Figure 8 It is a schematic diagram showing a user of a cross-reality (XR) system that perceives virtual content according to some embodiments;
[0041] Fig. 9 It is for transforming between coordinate systems according to some embodiments Figure 8 a block diagram of components of a first XR device of an XR system;
[0042] Fig.10 It is a schematic diagram showing an exemplary transformation of an origin coordinate frame to a destination coordinate frame for correctly rendering local XR content according to some embodiments;
[0043] Fig.11 It is a top view plan showing a pupil-based coordinate frame according to some embodiments;
[0044] Fig.12 It is a top view plan showing a camera coordinate frame including all pupil positions according to some embodiments;
[0045] Fig.13 It is according to some embodiments Fig. 9 a schematic diagram of a display system;
[0046] Fig.14 It is a block diagram showing the creation of a persistent coordinate frame (PCF) and the attachment of XR content to the PCF according to some embodiments;
[0047] Fig.15 It is a flowchart showing a method of establishing and using a PCF according to some embodiments;
[0048] Fig.16 It is an XR system according to some embodiments including a second XR device Figure 8 a block diagram;
[0049] Fig.17 It is a schematic diagram showing a room and key frames established for respective regions in the room according to some embodiments;
[0050] Fig.18 It is a schematic diagram showing the establishment of a key-frame based persistent pose according to some embodiments;
[0051] Fig.19is a schematic diagram showing the establishment of a persistent coordinate frame (PCF) based on a persistent pose according to some embodiments;
[0052] FIG. 20A to FIG. 20C is a schematic diagram showing an example of creating a PCF according to some embodiments;
[0053] Fig.21 is a block diagram showing a system for generating a global descriptor for a single image and / or map according to some embodiments;
[0054] Fig. 22 is a flowchart showing a method for calculating an image descriptor according to some embodiments;
[0055] Fig.23 is a flowchart showing a localization method using an image descriptor according to some embodiments;
[0056] Fig.24 is a flowchart showing a method for training a neural network according to some embodiments;
[0057] Fig.25 is a block diagram showing a method for training a neural network according to some embodiments;
[0058] Fig.26 is a schematic diagram showing an AR system configured to rank and merge multiple environmental maps according to some embodiments;
[0059] Fig. 27 is a simplified block diagram showing multiple canonical maps stored on a remote storage medium according to some embodiments;
[0060] Fig.28 is a schematic diagram showing a method for selecting a canonical map to, for example, localize a new tracking map in one or more canonical maps and / or obtain a PCF from the canonical map according to some embodiments;
[0061] Fig.29 is a flowchart showing a method for selecting multiple ranked environmental maps according to some embodiments;
[0062] Fig.30 is a schematic diagram showing an AR system according to some embodiments Fig.26 of an exemplary map ranking section;
[0063] Fig.31A is a schematic diagram showing an example of the regional attributes of a tracking map (TM) and an environmental map in a database according to some embodiments;
[0064] Fig.31B is a schematic diagram showing determining for Fig.29Schematic diagram of an example of the geographical location of a tracking map (TM) filtered by geographical location;
[0065] Fig.32 Shows according to some embodiments Fig.29 Schematic diagram of an example of geographical location filtering;
[0066] Fig.33 Shows according to some embodiments Fig.29 Schematic diagram of an example of Wi-Fi BSSID filtering;
[0067] Fig.34 Shows according to some embodiments the use of Fig.29 Schematic diagram of an example of positioning;
[0068] Fig.35 And 36 Is a block diagram of an XR system configured to rank and merge multiple environmental maps according to some embodiments.
[0069] Fig.37 Is a block diagram showing a method of creating an environmental map of the physical world in canonical form according to some embodiments;
[0070] Fig.38A And 38B Is a schematic diagram showing an environmental map created in canonical form by updating the tracking map of Figure 7 with a new tracking map according to some embodiments.
[0071] Figures 39A to 39F Is a schematic diagram showing an example of merging maps according to some embodiments;
[0072] Fig.40 Is according to some embodiments that can be generated by Fig. 9 The two-dimensional representation of the three-dimensional first local tracking map (ground Figure 1 ) generated by the first XR device;
[0073] Fig.41 Shows according to some embodiments uploading from the first XR device to Fig. 9 The server of the ground Figure 1 Block diagram;
[0074] Fig.42 Shows according to some embodiments Fig.16 Schematic diagram of an XR system, showing that after the first user has terminated the first session, the second user has initiated a second session using the second XR device of the XR system;
[0075] Fig.43A Shows according to some embodiments for Fig.42 Block diagram of a new session of the second XR device;
[0076] Fig.43B is a block diagram showing the creation of a tracking map for a second XR device according to some embodiments for Fig.42 ;
[0077] Fig.43C is a block diagram showing the downloading of a canonical map from a server to a second XR device according to some embodiments for Fig.42 ;
[0078] Fig.44 is a schematic diagram showing a positioning attempt to align a second tracking map (ground Fig.42 ) that can be generated by a second XR device to a canonical map according to some embodiments; Figure 2 )
[0079] Fig.45 is a schematic diagram showing a positioning attempt to align a second tracking map (ground Fig.44 ) that can be further developed and has XR content associated with the PCF of ground Figure 2 to a canonical map according to some embodiments; Figure 2 ;
[0080] FIG. 46A to FIG. 46B is a schematic diagram showing the successful alignment of ground Fig.45 to a canonical map according to some embodiments; Figure 2 ;
[0081] Fig.47 is a schematic diagram showing a canonical map generated by including one or more PCFs from a canonical map of Fig.46A into ground Fig.45 according to some embodiments; Figure 2 ;
[0082] Fig.48 is a schematic diagram showing a canonical map of Fig.47 and a further expansion of ground Figure 2 on a second XR device according to some embodiments;
[0083] Fig.49 is a block diagram showing the uploading of ground Figure 2 from a second XR device to a server according to some embodiments;
[0084] Fig.50 is a block diagram showing the merging of ground Figure 2 with a canonical map according to some embodiments;
[0085] Fig.51 is a block diagram showing the transmission of a new canonical map from a server to a first XR device and a second XR device according to some embodiments;
[0086] Fig.52 is a block diagram showing a two - dimensional representation of the ground according to some embodiments and the head coordinate frame of a second XR device with reference to the ground; Figure 2 and the second XR device; Figure 2 is a block diagram showing the adjustment of the head coordinate frame that can occur in six degrees of freedom in a two - dimensional manner according to some embodiments;
[0087] Fig.53 is a block diagram showing the adjustment of the head coordinate frame that can occur in six degrees of freedom in a two - dimensional manner according to some embodiments;
[0088] Fig.54 is a block diagram showing a canonical map on a second XR device according to some embodiments, where sound is positioned relative to the ground Figure 2 of the PCF;
[0089] Fig.55 and Fig.56 is a perspective view and a block diagram showing the use of the XR system when a first user has terminated a first session and the first user has initiated a second session using the XR system according to some embodiments;
[0090] Fig.57 and Fig.58 is a perspective view and a block diagram showing the use of the XR system when three users are using the XR system simultaneously in the same session according to some embodiments;
[0091] Fig.59 is a flowchart showing a method for restoring and resetting the head pose according to some embodiments;
[0092] Fig.60 is a block diagram of a machine in computer form that can find application in the system of the present invention according to some embodiments;
[0093] Fig.61 is a schematic diagram of an exemplary XR system according to some embodiments, where any of the multiple devices can access the positioning service;
[0094] Fig.62 is an exemplary processing flow for operating a portable device according to some embodiments, where the portable device is part of an XR system providing cloud - based positioning; and
[0095] Fig.63A 、 Fig.63B and Fig.63C are exemplary processing flows for cloud - based positioning according to some embodiments.
[0096] Fig.64 、 Fig.65 、 Fig.66 、 Fig.67 and Fig.68A series of schematic diagrams showing a portable XR device constructing a tracking map using multiple cells with wireless fingerprints when a user wearing the XR device traverses a 3D environment.
[0097] Fig.69 It is to use wireless fingerprints to select a set of cells in a stored map as candidate cells for positioning construction Fig.64 …68. A schematic diagram of a portable XR device with a tracking map.
[0098] Fig.70 A flowchart showing a method of operating a portable XR device to generate wireless fingerprints according to some embodiments.
[0099] Fig.71A A block diagram of components for calculating a low-resolution frame descriptor according to some embodiments.
[0100] Fig.71B A block diagram of components for calculating a high-resolution frame descriptor according to some embodiments.
[0101] Fig.72 A flowchart showing a method of using a multi-resolution frame descriptor in combination with an image captured by a user device according to some embodiments.
[0102] Fig.73 An illustration of a posed feature rig (PFR) according to some embodiments.
[0103] Fig.74 A flowchart showing a method of positioning using a high-resolution descriptor according to some embodiments. Detailed Description
[0104] Methods and devices for providing XR scenarios are described herein. To provide a realistic XR experience to multiple users, an XR system must know the location of the user within the physical world in order to correctly associate the positions of virtual objects with real objects. The inventors have recognized and realized methods and devices for positioning XR devices in large-scale and ultra-large-scale environments (e.g., neighborhoods, cities, countries, globally) with reduced time and increased accuracy.
[0105] An XR system can build an environmental map of a scene, which can be created based on images and / or depth information collected by sensors that are part of an XR device worn by a user of the XR system. Each XR device can develop a local map of its physical environment by integrating information from one or more images collected during operation of the device. In some embodiments, when the device initially begins scanning the physical world (e.g., starts a new session), the coordinate system of the map is associated with the position and / or orientation of the device. As the user interacts with the XR system, the position and / or orientation of the device may change as the session progresses, whether different sessions are associated with different users, each having their own wearable device with sensors for scanning the environment, or the same user uses the same device at different times.
[0106] The XR system can implement one or more techniques to enable operation based on persistent spatial information. For example, these techniques can provide a more computationally efficient and immersive XR experience for single or multiple users by allowing any one of multiple users of the XR system to create, store, and retrieve persistent spatial information. The persistent spatial information can also be used to quickly restore and reset the head pose on each of one or more XR devices in a computationally efficient manner.
[0107] The persistent spatial information can be represented by a persistent map. The persistent map can be stored in a remote storage medium (e.g., the cloud). For example, after a wearable device worn by a user is turned on, it can retrieve a previously created and stored appropriate map from persistent storage such as cloud storage. The previously stored map may be based on data about the environment collected by sensors on the user's wearable device during a previous session. Retrieving the stored map can enable use of the wearable device without having to complete a scan of the physical world with the sensors on the wearable device. Alternatively or additionally, when the system / device enters a new area of the physical world, it can similarly retrieve an appropriate stored map.
[0108] The stored map can be represented in a canonical form, and the local reference frame on each XR device can be related to this canonical form. In a multi-device XR system, a stored map accessed by one device may have been created and stored by another device and / or may have been constructed by aggregating data about the physical world collected by sensors on multiple wearable devices, where at least a portion of the physical world represented by the stored map was previously occupied by the multiple wearable devices.
[0109] In some embodiments, persistent space information can be represented in a way that can be easily shared among users and among distributed components including applications. A canonical map can provide information about the physical world, e.g., as a Persistent Coordinate Frame (PCF). The PCF can be defined based on a set of features identified in the physical world. The features can be selected such that they may be the same across user sessions of the XR system. The PCF may exist sparsely, providing less than all available information about the physical world so that they can be efficiently processed and transmitted. Techniques for processing persistent space information can include creating a dynamic map across one or more sessions based on the local coordinate frames of one or more devices. These maps can be sparse maps, representing the physical world based on a subset of feature points detected in the images used to form the map. The Persistent Coordinate Frame (PCF) can be generated from the sparse map and can be exposed to XR applications, e.g., through an Application Programming Interface (API). These capabilities can be supported by techniques for forming a canonical map by merging multiple maps created by one or more XR devices.
[0110] The relationship between the local map of each device and the canonical map can be determined through a localization process. The localization process can be performed on each XR device based on a set of canonical maps selected and sent to the device. Alternatively or additionally, a localization service can be provided on a remote processor, e.g., the localization service can be implemented in the cloud.
[0111] Sharing data about the physical world among multiple devices can enable a shared user experience of virtual content. For example, two XR devices accessing the same stored map can both be localized relative to the stored map. Once localized, the user device can render virtual content at that location by converting the location specified by reference to the stored map into the reference frame maintained by the user device. The user device can use this local reference frame to control the display of the user device to render the virtual content at the specified location.
[0112] To support these and other functions, an XR system can include components that develop, maintain, and use persistent space information (including one or more stored maps) based on data about the physical world collected by sensors on the user device. These components can be distributed across the XR system, e.g., through some operations on the head-mounted portion of the user device. Other components can operate on a computer, associated with a user coupled to the head-mounted portion through a local area network or a personal area network. Still others can operate at a remote location, such as at one or more servers accessible through a wide area network.
[0113] For example, these components can include components that build a persistent map from information about the physical world collected by one or more user devices. An example of such a component, described in more detail below, is a map merging component. The merging process may need to find a portion of a stored map that represents the same area of the physical world as other information to be merged into a set of stored maps.
[0114] In some embodiments, the information to be merged into a set of persistent maps can be a tracking map collected by multiple user devices. An XR device can each build its own tracking map with information about the physical world collected by the sensors of the XR device at different locations and times. In addition to possibly providing input for creating and maintaining a persistent map, the tracking map can also be used to track the movement of a user in a scene, enabling the XR system to estimate the head pose of its corresponding user relative to the reference frame established by the tracking map on that user device. The processes of matching a portion of the tracking map with a stored map and determining the head pose relative to the tracking map can both involve searching for matching image frames, which can be simplified by using frame descriptors.
[0115] To support these and other functions, the XR system can include components that help select an appropriate set of one or more persistent maps, where the appropriate set may represent the same area of the physical world as indicated by the location information provided by the user device. An example of such a component, described in more detail below, is a map ranking and map selection component. For example, such a component can receive input from a user device and identify one or more persistent maps that may represent the area in the physical world where the device is operating. For example, the map ranking component can help select a persistent map to be used by the local device when rendering virtual content, collecting data about the environment, or performing other actions. Alternatively or additionally, the map ranking component can help identify a persistent map to be updated as additional information about the physical world is collected by one or more user devices. The processes in these components may also need to find matching frames.
[0116] The XR system can be configured to create, share, and use persistent spatial information with low computational resource usage and / or low latency to provide an immersive user experience. Some such techniques can enable efficient comparison of spatial information, including finding sets of matching features or matching image frames.
[0117] In some embodiments, the comparison of sets of feature points can be simplified by using feature descriptors. The descriptors can have numerical values assigned by a trained neural network, enabling the comparison of features. Features that may represent the same feature point in the physical world are assigned feature descriptors with similar values, so that feature points representing the same location in the physical world can be quickly identified based on descriptors with similar values.
[0118] Finding similar image frames can also be simplified by representing the image frames with numerical descriptors. The descriptors can be calculated by a transformation that maps a set of features identified in the image to a frame descriptor. The transformation can be performed in a trained neural network. In some embodiments, the set of features provided as input to the neural network can be a filtered set of features, which are extracted from the image using techniques that preferentially select features that are likely to be persistent.
[0119] Representing feature points and image frames within an image as descriptors enables efficient matching of new image information with stored image information. The XR system can store in combination with the persistent map descriptors of one or more frames under the persistent map. The local image frames acquired by the user device can be similarly transformed into such descriptors. By searching the stored map for descriptors similar to the descriptors of the local image frames, one or more persistent maps that may represent the same physical space as the user device can be selected with relatively little processing. In some embodiments, descriptors can be calculated only for key frames in the persistent map and the local map, thereby further reducing processing when comparing maps. For example, such efficient comparison can be used to simplify finding the persistent map to be loaded into the local device or finding the persistent map to be updated based on the image information acquired with the local device.
[0120] Even when using simple techniques to comprehensively compare image frames based on frame descriptors, intensive computation may be required. Each frame descriptor can have a limited number of bits, which may be much less than the number of bits of information in the image frame. Therefore, ambiguity can be created by using frame descriptors because some images of different scenes in the physical world may have the same or very similar descriptor values. Comparison of sets of feature points can be additionally used for certain operations and may be computationally intensive. For example, two frames with matching frame descriptors can be determined to match only after a correspondence between the sets of feature points in these frames is found with a sufficiently low error.
[0121] To further compound the computational requirements, the number of frames in the persistent map increases as the scale of the environment grows, which in turn increases the risk that similar descriptors may be assigned to images at different locations in the physical world. For example, a map of a single room may have many frames. A building may have many rooms. In addition to outdoor areas such as streets and parks, a community may include many buildings. A city may include many communities, etc. Even when using techniques to limit the search space in a large stored map, the large map may have a large number of frames with similar descriptors (e.g., it may represent a large office with many similar desks and chairs). If there is a large amount of ambiguity, a large number of matching frames may be identified, and a large number of comparisons between feature sets may be required to find an exact match between the image frames. Such processing may result in computational latency, such as for localization, in identifying the matching image frames.
[0122] The inventors have recognized and realized that using frame descriptors of different resolutions in different parts of an XR system can enable balancing computational requirements based on available computing resources and ambiguity, thereby providing an overall improvement in system performance. For example, using a frame descriptor with a large number of bits can provide a higher resolution to reduce ambiguity and reduce the latency associated with subsequent processing to resolve that ambiguity. Such a higher-resolution frame descriptor can be used where there is greater ambiguity and / or more available computing resources to process the larger descriptor. Conversely, a smaller, lower-resolution descriptor can be used where there is less ambiguity and / or fewer available computing resources to process the descriptor.
[0123] Since the cloud-based components of an XR system can process larger maps that result in more ambiguity and can also access more computing resources than a user device, such as processor cycles and memory, the frame descriptors used in the cloud may be longer than the map descriptors stored on a local device. For example, a locally generated descriptor may have a length of 256 bytes, while a cloud-based descriptor may have a length of 1024 bytes.
[0124] In some embodiments, a user device can send a set of features identified in an image to the cloud. The networked computers forming the cloud can then calculate a high-resolution frame descriptor by a transformation that associates the set of features identified in the image with the descriptor. The transformation can be performed in a trained neural network. In some embodiments, the set of features provided as input to the neural network can be a filtered set of features, which are extracted from the image using techniques that preferentially select features that are likely to be persistent, for example.
[0125] The cloud can then select a cloud-stored map with a frame descriptor that is similar to the descriptor calculated from the feature data sent from the local device. The cloud can then perform processing, whether for localization, map merging, or other functions, and send the results back to the local device. For localization, the result may be a transformation between the coordinate system of the map in the cloud and the tracking map maintained by the user device. Alternatively or additionally, the result of the processing in the cloud can be one or more maps sent to the local device such that the device can localize to the selected map or perform other processing on those maps.
[0126] In some embodiments, when sending a map from the cloud to the device, only a portion of the map can be sent. Sending a portion of the map representing the device's current vicinity can enable the local device to operate on a map with fewer resources than those used in cloud-based processing. Reducing the computational requirements can result in a more immersive user experience since the local device may be lighter, generate less heat, have a longer battery life, be more responsive, and exhibit other desirable performance characteristics.
[0127] Thus, when performing the comparison of image frames with a map on a local device, there may be fewer image frames to compare. This may be the case where the map being compared is only a part of a set of maps that exist in the cloud or a locally generated tracking map. As a result, the ambiguity of the comparison result of the image frames based on the frame descriptors may be less. Using a lower-resolution frame descriptor on the device than in the cloud may result in a net gain, as the computational burden and the associated latency in deriving and comparing the frame descriptors may be reduced more than the increased computational burden of resolving the ambiguity of multiple frames with similar descriptors.
[0128] The techniques described herein can be used together or separately with many types of devices and for many types of scenarios, including wearable or portable devices with limited computational resources that provide augmented or mixed reality scenarios. In some embodiments, the techniques can be implemented by one or more services that form part of an XR system.
[0129] AR System Overview
[0130] Figure 1 and Figure 2 Scenes with virtual content are shown, which are displayed together with a part of the physical world. For illustrative purposes, an AR system is used as an example of an XR system. Figure 3-6B An exemplary AR system is shown, which includes one or more processors, a memory, sensors, and a user interface that can operate according to the techniques described herein.
[0131] Reference Figure 1 , depicts an outdoor AR scene 354, where a user of AR technology sees a park-like setting 356 of the physical world, characterized by people, trees, buildings in the background, and a concrete platform 358. In addition to these items, the user of AR technology also perceives that they "see" a robotic statue 357 standing on the concrete platform 358 of the physical world, and a flying cartoon-like avatar character 352 that appears to be the head of a bumblebee, even though these elements (e.g., avatar character 352 and robotic statue 357) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is challenging to produce AR technology that promotes a comfortable, natural feeling and rich presentation of virtual image elements among other virtual or physical world image elements.
[0132] Such an AR scene can be implemented by a system that builds a map of the physical world based on tracking information, enabling a user to place AR content in the physical world, determine the location to place the AR content in the map of the physical world, preserve the AR scene so that the placed AR content can be reloaded and displayed in the physical world during, for example, different AR experience sessions, and enable multiple users to share the AR experience. The system can build and update a digital representation of the surfaces of the physical world around the user. This representation can be used to render virtual content to appear fully or partially occluded by physical objects between the user and the rendering location of the virtual content for placing virtual objects in physics-based interactions, as well as for virtual character path planning and navigation, or for other operations that use information about the physical world.
[0133] Figure 2 Depicts another example of an indoor AR scene 400 according to some embodiments, which shows an exemplary usage scenario of an XR system. The exemplary scene 400 is a living room with walls, a bookshelf on one side of the walls, a floor lamp at the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of AR technology can also perceive virtual objects, such as an image on the wall behind the sofa, a bird flying across the door, a deer peeking out from the bookshelf, and an ornament in the form of a windmill placed on the coffee table.
[0134] For the image on the wall, AR technology requires not only information about the wall surface but also information about objects and surfaces in the room (such as the shape of the lamp), which occludes the image to correctly render the virtual object. For the flying bird, AR technology requires information about all objects and surfaces around the room to render the bird with realistic physical effects, to avoid objects and surfaces or to bounce when the bird collides. For the deer, AR technology requires information about the surface (such as the floor or the coffee table) to calculate the placement position of the deer. For the windmill, the system can identify that it is an object separate from the table and can determine that it is movable, while the corner of the shelf or the corner of the wall can be determined to be stationary. This distinction can be used to determine which parts of the scene to use or update in each of various operations.
[0135] Virtual objects can be placed during a previous AR experience session. When a new AR experience session starts in the living room, AR technology needs to accurately display the virtual objects at the positions where they were previously placed and are actually visible from different viewpoints. For example, the windmill should be shown standing on a book, rather than floating above the table at a different position without the book. Such floating may occur if the position of the user in the new AR experience session is not accurately located in the living room. As another example, if the user views the windmill from a different viewpoint than when it was placed, AR technology needs to display the corresponding side of the windmill.
[0136] A scene can be presented to a user via a system including a plurality of components, the plurality of components including a user interface that can stimulate one or more of the user's senses such as vision, sound, and / or touch. Additionally, the system can include one or more sensors that can measure parameters of a physical portion of the scene, including the position and / or movement of the user within the physical portion of the scene. Further, the system can include one or more computing devices, as well as associated computer hardware such as memory. These components can be integrated into a single device or can be distributed across multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.
[0137] Figure 3 Depicted is an AR system 502 according to some embodiments, which is configured to provide an experience of interacting with an AR content with the physical world 506. The AR system 502 can include a display 508. In the illustrated embodiment, the display 508 can be worn by the user as part of a head-mounted headset such that the user can wear the display over their eyes like a pair of goggles or glasses. At least a portion of the display can be transparent such that the user can observe a see-through reality 510. The see-through reality 510 can correspond to the portion of the physical world 506 that is within the current viewpoint of the AR system 502, which can correspond to the user's viewpoint in the case where the user wears a head-mounted headset incorporating the display and sensors of the AR system to obtain information about the physical world.
[0138] The AR content can also be presented on the display 508, overlaid on the see-through reality 510. To provide accurate interaction between the AR content and the see-through reality 510 on the display 508, the AR system 502 can include sensors 522 configured to capture information about the physical world 506.
[0139] The sensors 522 can include one or more depth sensors that output depth maps 512. Each depth map 512 can have a plurality of pixels, each pixel representing the distance from a surface in the physical world 506 relative to the depth sensor in a particular direction. Raw depth data can be received from the depth sensors to create the depth maps. The depth maps can be updated as fast as the depth sensors can form new images, which can be hundreds or thousands of times per second. However, the data can be noisy and incomplete and have holes shown as black pixels on the illustrated depth maps.
[0140] The system may include other sensors, such as an image sensor. The image sensor may acquire monocular or stereo information, which may be processed to represent the physical world in other ways. For example, the image may be processed in the world reconstruction component 516 to create a mesh that represents the connected parts of objects in the physical world. Metadata about such objects (including, for example, color and surface texture) may similarly be acquired using sensors and stored as part of the world reconstruction.
[0141] The system may also acquire information about the user's head pose (or "pose") relative to the physical world. In some embodiments, the head pose tracking component of the system may be used to compute the head pose in real time. The head pose tracking component may represent the user's head pose in a coordinate system with six degrees of freedom, which includes, for example, translations along three perpendicular axes (e.g., forward / backward, up / down, left / right) and rotations about those three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, the sensor 522 may include an inertial measurement unit ("IMU") that may be used to compute and / or determine the head pose 514. The head pose 514 for the depth map may indicate, for example, the current viewing point of the sensor that captures the depth map in six degrees of freedom, but the head-mounted headset 514 may be used for other purposes, such as associating image information with a particular part of the physical world or associating the position of a display worn on the user's head with the physical world.
[0142] In some embodiments, the head pose information may be derived in other ways different from the IMU (such as analyzing objects in images). For example, the head pose tracking component may compute the relative position and orientation of the AR device with respect to physical objects based on visual information captured by a camera and inertial information captured by the IMU. The head pose tracking component may then compute the head pose of the AR device, for example, by comparing the computed relative position and orientation of the AR device with respect to the physical objects with the features of the physical objects. In some embodiments, the comparison may be performed by identifying features in images captured using one or more sensors 522 that are stable over time so that the change in the position of these features in the images captured over time may be correlated with the change in the user's head pose.
[0143] The inventors have recognized and understood techniques for operating an XR system to provide XR scenes for a more immersive user experience, such as estimating head pose at a frequency of 1 kHz, with a low usage of computing resources associated with the XR device, which can be configured with, for example, four video graphics array (VGA) cameras operating at 30 Hz, one inertial measurement unit (IMU) operating at 1 kHz, the computing power of a single advanced RISC machine (ARM) core, memory less than 1 GB, and network bandwidth less than 100 Mbps. These techniques involve reducing the processing required to generate and maintain maps and estimate head pose, and providing and using data with low computational overhead. The XR system can calculate its pose based on matching visual features. U.S. Patent Application No. 16 / 221,065 describes hybrid tracking, and the entire content thereof is hereby incorporated by reference.
[0144] In some embodiments, the AR device can build a map based on feature points identified in successive images of a series of image frames captured as the user moves the AR device throughout the physical world. Although each image frame can be taken from a different pose as the user moves, the system can adjust the orientation of the features of each successive image frame to match the orientation of the initial image frame by matching the features of the successive image frames to the previously captured image frames. The translation of successive image frames such that points representing the same feature will match corresponding feature points in the previously collected image frames can be used to align each successive image frame to match the orientation of the previously processed image frame. The frames in the generated map can have a common orientation established when the first image frame was added to the map. The map has multiple sets of feature points in a common reference frame, and the map can be used to determine the user's pose in the physical world by matching the features in the current image frame to the map. In some embodiments, the map can be referred to as a tracking map.
[0145] In addition to being able to track the user's pose in the environment, the map can also enable other components of the system (such as the world reconstruction component 516) to determine the position of physical objects relative to the user. The world reconstruction component 516 can receive the depth map 512, the head pose 514, and any other data from the sensors and integrate this data into the reconstruction 518. The reconstruction 518 can be more complete and less noisy than the sensor data. The world reconstruction component 516 can update the reconstruction 518 using the spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0146] The reconstruction 518 may include a representation of the physical world in one or more data formats (including, for example, voxels, meshes, planes, etc.). Different formats may represent alternative representations of the same part of the physical world or may represent different parts of the physical world. In the example shown, on the left side of the reconstruction 518, a portion of the physical world is presented as a global surface; on the right side of the reconstruction 518, a portion of the physical world is presented as a mesh.
[0147] In some embodiments, the map maintained by the head pose component 514 may be sparse relative to other maps of the physical world that may be maintained. A sparse map may indicate the locations of areas of interest and / or structures (such as corners or edges) rather than providing information about the location of surfaces and possibly other features. In some embodiments, the map may include image frames captured by the sensor 522. These frames may be reduced to features that can represent areas of interest and / or structures. In conjunction with each frame, information about the pose of the user from whom the frame was obtained may also be stored as part of the map. In some embodiments, not every image acquired by the sensor may be stored or may be stored. In some embodiments, when images are collected by the sensor, the system may process the images and select a subset of the image frames for further computation. The selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may add new image frames to the map, for example, based on the overlap with previous image frames that have already been added to the map or based on image frames that contain a sufficient number of features that are determined to likely represent stationary objects. In some embodiments, the selected image frames or groups of features from the selected image frames may be used as key frames of the map, which are used to provide spatial information.
[0148] In some embodiments, the amount of data processed when constructing the map may be reduced, such as by constructing a sparse map with a set of mapped points and key frames and / or dividing the map into chunks to enable per-chunk updates. The mapped points may be associated with areas of interest in the environment. The key frames may include information selected from the data captured by the camera. U.S. Patent Application No. 16 / 520,582 describes determining and / or evaluating a localization map and is hereby incorporated by reference in its entirety.
[0149] The AR system 502 can integrate sensor data over time from multiple perspectives of the physical world. When a device including sensors moves, the pose of the sensors (e.g., position and orientation) can be tracked. Since the frame pose of the sensors and their relationships to other poses are known, each of these multiple viewpoints of the physical world can be fused together to form a single combined reconstruction of the physical world, which can be used as an abstract layer of a map and provide spatial information. By using spatial and temporal averaging (i.e., averaging data from multiple viewpoints over time) or any other suitable method, the reconstruction can be more complete and less noisy than the original sensor data.
[0150] In Figure 3 the illustrated embodiment, the map represents the part of the user's physical world in which there is a single wearable device. In that case, the head pose associated with the frame in the map can be represented as a local head pose, indicating the orientation relative to the initial orientation of the single device at the start of the session. For example, the head pose can be tracked relative to the initial head pose when the device is turned on, or otherwise operated to scan the environment to establish a representation of that environment.
[0151] In combination with the content characterizing that part of the physical world, the map can include metadata. The metadata can, for example, indicate the time when the sensor information captured to form the map was captured. Alternatively or additionally, the metadata can indicate the location of the sensors when the information captured to form the map was captured. The location can be represented directly, such as by using information from a GPS chip, or indirectly, such as by using a wireless (e.g., Wi-Fi) signature, which indicates the strength of the signals received from one or more wireless access points while the sensor data is being collected, and / or an identifier such as a BSSID of the wireless access point to which the user device is connected while the sensor data is being collected.
[0152] The reconstruction 518 can be used for AR functions, such as generating a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation may change as the user moves or objects in the real world change. Aspects of the reconstruction 518 can, for example, be used by components 520 that generate a changing global surface representation in world coordinates, which can be used by other components.
[0153] Based on this information, AR content can be generated, such as through the AR application 504. The AR application 504 can be, for example, a game program that performs one or more functions based on information about the physical world, such as visual occlusion, physics-based interaction, and environmental reasoning. It can perform these functions by querying different formats of data from the reconstruction 518 generated by the world reconstruction component 516. In some embodiments, the component 520 can be configured to output an update when the representation in the region of interest of the physical world changes. For example, the region of interest can be set to approximate a part of the physical world near the system user, such as the part within the user's field of view, or projected (predicted / determined) to enter the user's field of view.
[0154] The AR application 504 can use this information to generate and update AR content. The virtual part of the AR content can be presented on the display 508 in combination with the see-through reality 510, thus creating a realistic user experience.
[0155] In some embodiments, an AR experience can be provided to the user through an XR device, which can be a wearable display device and can be part of a system that can include remote processing and / or remote data storage and / or, in some embodiments, other wearable display devices worn by other users. For the sake of simplicity in illustration, Figure 4 an example of a system 580 (hereinafter referred to as "system 580") including a single wearable device is shown. The system 580 includes a head-mounted display device 562 (hereinafter referred to as "display device 562"), and various mechanical and electronic modules and systems that support the functions of the display device 562. The display device 562 can be coupled to a frame 564, which can be worn by the user or viewer 560 (hereinafter referred to as "user 560") of the display system and is configured to position the display device 562 in front of the eyes of the user 560. According to various embodiments, the display device 562 can display sequentially. The display device 562 can be monocular or binocular. In some embodiments, the display device 562 can be Figure 3 an example of the display 508 in
[0156] In some embodiments, a speaker 566 is coupled to the frame 564 and positioned near the ear canal of the user 560. In some embodiments, another speaker (not shown) is positioned near the other ear canal of the user 560 to provide stereo / plastic sound control. The display device 562 is operatively coupled to a local data processing module 570, such as through a wired wire or a wireless connection 568, and the local data processing module 570 can be installed in various configurations, such as fixedly attached to the frame 564, fixedly attached to a helmet or hat worn by the user 560, embedded in the earphone, or otherwise removably attached to the user 560 (e.g., in a backpack configuration, in a belt-coupled configuration).
[0157] The local data processing module 570 may include a processor and a digital memory such as a non-volatile memory (e.g., flash memory), both of which can be used to assist in the processing, caching, and storage of data. The data includes: a) data captured from sensors (e.g., that may be operatively coupled to the frame 564) or otherwise attached to the user 560, such as an image capture device (such as a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a radio device, and / or a gyroscope; and / or b) data obtained and / or processed using the remote processing module 572 and / or the remote data repository 574, and possibly passed to the display device 562 after such processing or acquisition.
[0158] In some embodiments, the wearable device may communicate with remote components. The local data processing module 570 may be operatively coupled to the remote processing module 572 and the remote data repository 574 via communication links 576, 578 (such as via a wired or wireless communication link), respectively, such that these remote modules 572, 574 are operatively coupled to each other and can be used as resources for the local data processing module 570. In a further embodiment, as a supplement or alternative to the remote data repository 574, the wearable device may access a cloud-based remote data repository and / or service. In some embodiments, the above-mentioned head pose tracking component may be implemented at least partially in the local data processing module 570. In some embodiments, Figure 3 the world reconstruction component 516 in can be implemented at least partially in the local data processing module 570. For example, the local data processing module 570 may be configured to execute computer-executable instructions to generate a map and / or a physical world representation based at least in part on at least a portion of the data.
[0159] In some embodiments, the processing can be distributed across local and remote processors. For example, local processing can be used to construct a map (e.g., a tracking map) on the user device based on sensor data collected using sensors on the user device. Such a map can be used by an application on the user device. Additionally, previously created maps (e.g., canonical maps) can be stored in a remote data repository 574. Where a suitable stored or persistent map is available, it can be used in place of or in addition to the tracking map created locally on the device. In some embodiments, the tracking map can be aligned to the stored map such that a correspondence is established between the tracking map and the canonical map, where the tracking map may be oriented relative to the position of the wearable device when the user turns on the system, and the canonical map can be oriented relative to one or more persistent features. In some embodiments, the persistent map can be loaded on the user device to allow the user device to render virtual content without the latency associated with scanning for a location, thereby constructing a tracking map of the user's entire environment based on sensor data acquired during the scan. In some embodiments, the user device can access a remote persistent map (e.g., stored in the cloud) without downloading the persistent map on the user device.
[0160] In some embodiments, spatial information can be transmitted from the wearable device to a remote service, such as a cloud service configured to localize the device to a stored map maintained on the cloud service. According to one embodiment, the localization process can be performed in the cloud, matching the device location to an existing map (e.g., a canonical map) and returning a transformation that links virtual content to the wearable device location. In such an embodiment, the system can avoid transmitting the map from the remote resource to the wearable device. Other embodiments can be configured for device-based and cloud-based localization, e.g., to enable functionality where network connectivity is unavailable or the user chooses not to enable cloud-based localization.
[0161] Alternatively or additionally, the tracking map can be merged with previously stored maps to extend or improve the quality of those maps. The process of determining whether a suitable previously created environmental map is available and / or merging the tracking map with one or more stored environmental maps can be done in the local data processing module 570 or the remote processing module 572.
[0162] In some embodiments, the local data processing module 570 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which would limit the computational budget of the local data processing module 570 but enable a smaller device. In some embodiments, the world reconstruction component 516 may use less than the computational budget of a single Advanced RISC Machine (ARM) core to generate a physical world representation in real time on an undefined space, such that the remaining computational budget of a single ARM core can be accessed for other purposes, such as, for example, extracting a mesh.
[0163] In some embodiments, the remote data repository 574 may include a digital data storage facility that may be made available via the Internet or other networking configurations in a “cloud” resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, allowing for fully autonomous use from a remote module. In some embodiments, all data is stored and all or most computations are performed in the remote data repository 574, allowing for a smaller device. For example, world reconstruction may be stored in whole or in part in the repository 574.
[0164] In embodiments where data is remotely stored and accessible via a network, the data may be shared by multiple users of the augmented reality system. For example, user devices may upload their tracking maps to augment the environmental map database. In some embodiments, tracking map upload occurs at the end of a user session with the wearable device. In some embodiments, tracking map upload may occur continuously, semi-continuously, intermittently at a predefined time, after a predefined period from a previous upload, or when triggered by an event. Regardless of whether the data is based on that from the user device or any other user device, the tracking maps uploaded by any user device can be used to extend or improve a previously stored map. Similarly, the persistent maps downloaded to a user device may be based on data from that user device or any other user device. In this way, users can easily obtain high-quality environmental maps to improve their experience in the AR system.
[0165] In additional embodiments, restricting and / or avoiding persistent map downloads may be based on positioning performed on remote resources (e.g., in the cloud). In such a configuration, a wearable device or other XR device transmits feature information (e.g., the device's positioning information when a feature represented in the feature information is sensed) combined with pose information to a cloud service. One or more components of the cloud service may match the feature information to a corresponding stored map (e.g., a canonical map) and generate a transformation between the coordinate systems of the tracking map maintained by the XR device and the canonical map. Each XR device whose tracking map is positioned relative to the canonical map may accurately render virtual content at a position specified relative to the canonical map based on its own tracking.
[0166] In some embodiments, the local data processing module 570 is operatively coupled to the battery 582. In some embodiments, the battery 582 is a removable power source, such as over a coin cell battery. In other embodiments, the battery 582 is a lithium-ion battery. In some embodiments, the battery 582 includes both an internal lithium-ion battery that can be charged by the user 560 during non-operating times of the system 580 and a removable battery, such that the user 560 can operate the system 580 for longer periods of time without having to connect to a power source to charge the lithium-ion battery or without having to power off the system 580 to replace the battery.
[0167] Figure 5A Shown is a user 530 wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as "environment 532"). Information captured by the AR system along the user's movement path can be processed into one or more tracking maps. The user 530 positions the AR display system at position 534, and the AR display system records environmental information of the traversable world relative to position 534 (e.g., a digital representation of real objects in the physical world, which can be stored and updated as the real objects change in the physical world). This information can be stored as a pose in combination with images, features, directional audio inputs, or other desired data. Position 534 is aggregated into the data input 536, for example, as part of a tracking map, and is processed by at least the traversable world module 538, which can be implemented, for example, by processing on Figure 4 the remote processing module 572. In some embodiments, the traversable world module 538 may include a head pose component 514 and a world reconstruction component 516 such that the processed information can be combined with other information related to physical objects used in rendering virtual content to indicate the position of the objects in the physical world.
[0168] The traversable world module 538 at least partially determines the location and manner in which AR content 540, as determined from data input 536, can be placed in the physical world. The AR content is "placed" in the physical world by presenting both the physical world rendering and the AR content via the user interface, with the AR content rendered as if interacting with objects in the physical world and the objects in the physical world presented as if the AR content occludes the user's view of those objects when appropriate. In some embodiments, the shape and location of the AR content 540 can be determined to place the AR content by appropriately selecting portions of a fixed element 542 (such as a table) from the reconstruction (e.g., reconstruction 518). As an example, the fixed element can be a table, and the virtual content can be positioned such that it appears to be on the table. In some embodiments, the AR content can be placed within a structure in the field of view 544, which can be the current field of view or an estimated future field of view. In some embodiments, the AR content can persist relative to a model 546 (such as a mesh) of the physical world.
[0169] As depicted, the fixed element 542 serves as a proxy (e.g., digital copy) for any fixed element within the physical world that can be stored in the traversable world module 538, such that the user 530 can perceive the content on the fixed element 542 without the system having to map build to the fixed element 542 every time the user 530 views the fixed element 542. Thus, the fixed element 542 can be a mesh model from a previous modeling session or can be determined by a separate user but still stored by the traversable world module 538 for future reference by multiple users. Accordingly, the traversable world module 538 can identify the environment 532 from a previously map built environment and display the AR content without the user 530's device first having to map build all or a portion of the environment 532, thereby saving computational processes and cycles and avoiding latency in any rendered AR content.
[0170] A mesh model 546 of the physical world can be created by the AR display system, and the appropriate surfaces and metrics for interacting with and displaying the AR content 540 can be stored by the traversable world module 538 for future access by the user 530 or other users without having to recreate the model in whole or in part. In some embodiments, the data input 536 is an input such as a geographical location, user identification, and current activity to indicate to the traversable world module 538 which fixed element 542 among one or more fixed elements is available, which AR content 540 was last placed on the fixed element 542, and whether to display that same content (such AR content is "persistent" content regardless of how the user views a particular traversable world model).
[0171] Even in embodiments where an object is considered stationary (e.g., a kitchen table), the traversable world module 538 may also update the model of those objects in the physical world model from time to time to account for the possibility of changes in the physical world. The model of a stationary object may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not considered stationary (e.g., a kitchen chair). To render an AR scene with a sense of realism, the AR system may update the positions of these non-stationary objects at a much higher frequency than that used to update stationary objects. To be able to accurately track all objects in the physical world, the AR system may obtain information from multiple sensors, including one or more image sensors.
[0172] Figure 5B is a schematic diagram of a viewing optical component 548 and accompanying components. In some embodiments, two eye tracking cameras 550 pointing at the user's eyes 549 detect metrics of the user's eyes 549, such as the eye shape, eyelid occlusion, pupil direction, and blink on the user's eyes 549.
[0173] In some embodiments, one of the sensors may be a depth sensor 551, such as a time-of-flight sensor, which emits signals into the world and detects the reflections of those signals from nearby objects to determine the distance to a given object. The depth sensor can, for example, quickly determine whether an object has entered the user's field of view due to the movement of those objects or a change in the user's pose. However, information about the position of an object in the user's field of view may alternatively or additionally be collected by other sensors. Depth information can be obtained, for example, from a stereo vision image sensor or a plenoptic sensor.
[0174] In some embodiments, the world camera 552 records a view larger than the periphery to map the environment 532 and / or otherwise create a model of the environment 532 and detect inputs that may affect the AR content. In some embodiments, the world camera 552 and / or the camera 553 may be grayscale and / or color image sensors that can output grayscale and / or color image frames at fixed time intervals. The camera 553 may further capture an image of the physical world within the user's field of view at a specific time. Even if the values of the pixels of a frame-based image sensor do not change, the pixels may be sampled repeatedly. Each of the world camera 552, the camera 553, and the depth sensor 551 has a corresponding field of view 554, 555, and 556 to collect data from and record a physical world scene such as the physical world environment 532 depicted in Fig.34 A.
[0175] The inertial measurement unit 557 can determine the motion and orientation of the viewing optical assembly 548. In some embodiments, each component is operatively coupled to at least one other component. For example, the depth sensor 551 is operatively coupled to the eye tracking camera 550 to confirm the measured accommodation relative to the actual distance at which the user's eye 549 is gazing.
[0176] It should be understood that the viewing optical assembly 548 may include Fig.34 some of the components shown in B and may include components in place of or in addition to the shown components. For example, in some embodiments, the viewing optical assembly 548 may include two world cameras 552 instead of four. Alternatively or additionally, the cameras 552 and 553 do not need to capture visible light images of their entire fields of view. The viewing optical assembly 548 may include other types of components. In some embodiments, the viewing optical assembly 548 may include one or more dynamic vision sensors (DVS) whose pixels may respond asynchronously to relative changes in light intensity above a threshold.
[0177] In some embodiments, based on time-of-flight information, the viewing optical assembly 548 may not include the depth sensor 551. For example, in some embodiments, the viewing optical assembly 548 may include one or more plenoptic cameras whose pixels may capture light intensity and the angle of incident light, from which depth information may be determined. For example, a plenoptic camera may include an image sensor covered with a transmissive diffraction mask (TDM). Alternatively or additionally, a plenoptic camera may include an image sensor that includes angle-sensitive pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Instead of or in addition to the depth sensor 551, such sensors may be used as a source of depth information.
[0178] It should also be understood that Figure 5B the configuration of the components in is provided as an example. The viewing optical assembly 548 may include components having any suitable configuration that may be set to provide the user with the maximum field of view that is practically feasible for a particular set of components. For example, if the viewing optical assembly 548 has one world camera 552, that world camera may be placed in the central region of the viewing optical assembly rather than on the side.
[0179] Information from sensors in the viewing optical assembly 548 can be coupled to one or more processors in the system. The processor can generate data that can be rendered to enable a user to perceive interacting with an object in the physical world. The rendering can be implemented in any suitable manner, including generating image data depicting both physical and virtual objects. In other embodiments, physical and virtual content can be depicted in a scene by modulating the opacity of a display device that the user views through in the physical world. The opacity can be controlled to create the appearance of a virtual object and also to prevent the user from seeing objects in the physical world that are occluded by the virtual object. In some embodiments, when viewed through a user interface, the image data can include only virtual content, which can be modified so that the virtual content is perceived by the user as realistically interacting with the physical world (e.g., clipping the content to account for occlusion).
[0180] The location at which display content on the viewing optical assembly 548 is viewed to create the impression that an object is located at a particular position can depend on the physical properties of the viewing optical assembly. Additionally, the pose of the user's head relative to the physical world and the direction of the user's eye gaze can affect the location at which the display content in the physical world will appear at a particular location on the viewing optical assembly. The sensors as described above can collect this information, and / or provide information from which this information can be calculated, such that a processor receiving sensor input can calculate the location at which an object should be rendered on the viewing optical assembly 548 to create the desired appearance for the user.
[0181] Regardless of how content is presented to the user, a model of the physical world can be used such that the characteristics of virtual objects that can be affected by physical objects can be correctly calculated, including the shape, location, motion, and visibility of the virtual objects. In some embodiments, the model can include a reconstruction of the physical world, such as reconstruction 518.
[0182] The model can be created based on data collected from sensors on the user's wearable device. However, in some embodiments, the model can be created from data collected from multiple users, which can be aggregated in a computing device remote from all users (and the data can be "in the cloud").
[0183] The model can be created at least in part by a world reconstruction system, such as, for example, Fig. 6A more particularly depicted in Figure 3The world reconstruction component 516. The world reconstruction component 516 can include a perception module 660 that can generate, update, and store a representation of a portion of the physical world. In some embodiments, the perception module 660 can represent a portion of the physical world within the reconstruction range of the sensor as a plurality of voxels. Each voxel can correspond to a 3D cube of a predetermined volume in the physical world and include surface information that indicates whether there is a surface within the volume represented by the voxel. A value can be assigned to the voxel that indicates whether its corresponding volume has been determined to include a surface of a physical object, determined to be empty, or not yet measured by the sensor and thus its value is unknown. It should be understood that it is not necessary to explicitly store the values of voxels determined to be empty or unknown, as the values of the voxels can be stored in computer memory in any suitable manner, including not storing information for voxels determined to be empty or unknown.
[0184] In addition to generating information for the persistent world representation, the perception module 660 can also identify and output an indication of a change in the area around the user of the AR system. This indication of change can trigger an update to the volumetric data stored as part of the persistent world or trigger other functions, such as triggering the trigger component 604 that generates AR content to update the AR content.
[0185] In some embodiments, the perception module 660 can identify changes based on a signed distance function (SDF) model. The perception module 660 can be configured to receive sensor data such as, for example, a depth map 660a and a head pose 660b, and then fuse the sensor data into the SDF model 660c. The depth map 660a can directly provide SDF information, and the image can be processed to obtain SDF information. The SDF information represents the distance from the sensor used to capture the information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and thus from the perspective of the user. The head pose 660b can enable the SDF information to be related to the voxels in the physical world.
[0186] In some embodiments, the perception module 660 can generate, update, and store a representation of a portion of the physical world within the perception range. The perception range can be determined at least in part based on the reconstruction range of the sensor, which can be determined at least in part based on the limitations of the observation range of the sensor. As a specific example, an active depth sensor that operates using active IR pulses can operate reliably within a certain distance range, creating an observation range of the sensor that can range from a few centimeters or tens of centimeters to several meters.
[0187] The world reconstruction component 516 may include additional modules that can interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data acquired by the perception module 660. The persistent world module 662 may also include representations of the physical world in various formats. For example, volumetric metadata 662b such as voxels, as well as meshes 662c and planes 662d, may be stored. In some embodiments, other information such as depth maps may be saved.
[0188] In some embodiments, the representation of the physical world (such as Fig. 6A the representation shown in) may provide relatively dense information about the physical world compared to a sparse map (such as the feature point-based tracking map described above).
[0189] In some embodiments, the perception module 660 may include modules that generate representations of the physical world in various formats, including for example meshes 660d, planes, and semantics 660e. Representations of the physical world may be stored across local storage media and remote storage media. Depending on, for example, the location of the storage media, the representation of the physical world may be described in different coordinate frames. For example, the representation of the physical world stored in the device may be described in a coordinate frame that is relatively local to the device. The representation of the physical world may have a counterpart stored in the cloud. The counterpart in the cloud may be described in a coordinate frame that can be shared by all devices in the XR system.
[0190] In some embodiments, these modules may generate the representation based on data within the perception range of one or more sensors at the time of generating the representation, as well as data captured at a previous time and information in the persistent world module 662. In some embodiments, these components may operate with respect to depth information captured by a depth sensor. However, the AR system may include visual sensors and may generate such a representation by analyzing monocular or binocular visual information.
[0191] In some embodiments, these modules may operate on regions of the physical world. When the perception module 660 detects a change in the physical world in a sub-region of the physical world, those modules may be triggered to update the sub-region of the physical world. For example, such a change may be detected by detecting a new surface or other criteria (such as changing the values of a sufficient number of voxels representing the sub-region) in the SDF model 660c.
[0192] The world reconstruction component 516 may include a component 664 that can receive a representation of the physical world from the perception module 660. Information about the physical world can be extracted by these components based on, for example, usage requests from applications. In some embodiments, information can be pushed to the usage component, such as via an indication of a change in a pre-identified region or a change in the representation of the physical world within the perception range. The component 664 may include, for example, a game program and other components that perform processing for visual occlusion, physics-based interaction, and environmental reasoning.
[0193] In response to a query from the component 664, the perception module 660 may send a representation of the physical world in one or more formats. For example, when the component 664 indicates that the usage is for visual occlusion or physics-based interaction, the perception module 660 may send a representation of the surface. When the component 664 indicates that the usage is for environmental reasoning, the perception module 660 may send the mesh, plane, and semantics of the physical world.
[0194] In some embodiments, the perception module 660 may include a component that formats information to provide to the component 664. An example of such a component may be the ray casting component 660f. A usage component (e.g., the component 664) may query information about the physical world from a particular viewpoint, for example. The ray casting component 660f may be selected from one or more representations of the physical world data within the field of view from that viewpoint.
[0195] It should be understood from the above description that the perception module 660 or another component of the AR system may process data to create a 3D representation of a portion of the physical world. The data to be processed can be reduced by: culling portions of the 3D reconstruction volume at least partially based on the camera frustum and / or depth image; extracting and retaining plane data; capturing, retaining, and updating 3D reconstruction data in blocks that allow for local updates while maintaining neighborhood consistency; providing occlusion data to applications that generate such scenes, where the occlusion data is derived from a combination of one or more depth data sources; and / or performing multi-stage mesh simplification. The reconstruction may include data of different levels of complexity, including, for example, raw data (e.g., real-time depth data), fused volume data (e.g., voxels), and computed data (e.g., meshes).
[0196] In some embodiments, the components of the traversable world model can be distributed, with some parts executed locally on the XR device and some parts executed remotely, such as on a network-connected server or in the cloud. The distribution of information processing and storage between the local XR device and the cloud can affect the functionality and user experience of the XR system. For example, reducing the processing on the local device by offloading it to the cloud can extend battery life and reduce the heat generated on the local device. However, allocating too much processing to the cloud may introduce undesirable latency, resulting in an unacceptable user experience.
[0197] Figure 6B FIG. 600 depicts a distributed component architecture configured for spatial computing according to some embodiments. The distributed component architecture 600 can include a traversable world component 602 (e.g., Figure 5A PW 538 in FIG. 538), Lumin OS 604, API 606, SDK 608, and application 610. Lumin OS 604 can include a Linux-based kernel with custom drivers compatible with the XR device. API 606 can include an application programming interface that permits XR applications (e.g., application 610) to access the spatial computing features of the XR device. SDK 608 can include a software development kit that allows for the creation of XR applications.
[0198] One or more components in architecture 600 can create and maintain a model of the traversable world. In this example, sensor data is collected on the local device. The processing of this sensor data can be executed partially locally on the XR device and partially in the cloud. PW 538 can include an environmental map created at least in part based on data captured by AR devices worn by multiple users. During a session of an AR experience, individual AR devices (such as the wearable devices described above in connection with Figure 4 can create a tracking map, which is a type of map.
[0199] In some embodiments, the device can include components for building sparse maps and dense maps. The tracking map can be used as a sparse map and can include the head pose of the AR device scanning the environment and information about objects detected within the environment at each head pose. Those head poses can be maintained locally for each device. For example, the head pose on each device can be relative to the initial head pose when the device initiated its session. As a result, each tracking map can be local to the device that created it. The dense map can include surface information, which can be represented by a mesh or depth information. Alternatively or additionally, the dense map can include higher-level information derived from the surface or depth information, such as the location and / or characteristics of planes and / or other objects.
[0200] In some embodiments, the creation of the dense map can be independent of the creation of the sparse map. For example, the creation of the dense map and the sparse map can be performed in separate processing pipelines within the AR system. For example, separate processing can enable different types of map generation or processing to be performed at different rates. For example, the refresh rate of the sparse map may be faster than that of the dense map. However, in some embodiments, even if performed in different pipelines, the processing of the dense map and the sparse map may be related. For example, changes in the physical world revealed in the sparse map can trigger an update of the dense map, and vice versa. Additionally, even if created independently, these maps can be used together. For example, the coordinate system derived from the sparse map can be used to define the position and / or orientation of objects in the dense map.
[0201] The sparse map and / or the dense map can be persisted for reuse by the same device and / or shared with other devices. Such persistence can be achieved by storing the information in the cloud. The AR device can send the tracking map to the cloud, for example, to be merged with an environmental map selected from a persistent map previously stored in the cloud. In some embodiments, the selected persistent map can be sent from the cloud to the AR device for merging. In some embodiments, the persistent map can be oriented with respect to one or more persistent coordinate systems. Such maps can be used as canonical maps because they can be used by any of a plurality of devices. In some embodiments, the model of the traversable world can include or be created from or based on one or more canonical maps. Even if some operations are performed based on the device-local coordinate frame, the device can use the canonical map by determining the transformation between the device-local coordinate frame and the canonical map.
[0202] The canonical map can originate from a tracking map (TM) (e.g., Fig.31A TM 1102 in
[0203] ), which can be promoted to a canonical map. The canonical map can be persisted so that a device accessing the canonical map can use the information in the canonical map to determine the position of the objects represented in the canonical map in the physical world around the device once it determines the transformation between its local coordinate system and the coordinate system of the canonical map. In some embodiments, the TM can be a head pose sparse map created by the XR device. In some embodiments, the canonical map can be created when the XR device sends one or more TMs to the cloud server for merging with additional TMs captured by the XR device at different times or by other XR devices. Figure 7Depicts an exemplary tracking map 700 according to some embodiments. The tracking map 700 may provide a floor plan 706 of a physical object in the corresponding physical world represented by points 702. In some embodiments, the map points 702 may represent features of a physical object that may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. These features may be derived by processing an image, for example, the image may be acquired by a sensor of a wearable device in an augmented reality system. For example, features may be derived by processing an image frame output by the sensor to identify features based on large gradients or other suitable criteria in the image. Further processing may limit the number of features in each frame. For example, the processing may select features that may represent persistent objects. One or more heuristics may be applied to this selection.
[0204] The tracking map 700 may include data about the points 702 collected by the device. For each image frame having data points included in the tracking map, a pose may be stored. The pose may represent the orientation from which the image frame was captured such that the feature points within each image frame may be spatially related. The pose may be determined by positioning information such as may be derived by sensors (such as IMU sensors) on the wearable device. Alternatively or additionally, the pose may be determined by matching the image frame to other image frames depicting overlapping portions of the physical world. By finding such position correlations, which may be achieved by matching a subset of the feature points in two frames, the relative pose between the two frames may be calculated. The relative pose may be sufficient for the tracking map since the map may be relative to a local coordinate system of the device based on the initial pose of the device when starting to build the tracking map.
[0205] Not all of the feature points and image frames collected by the device may be retained as part of the tracking map because much of the information collected by the sensors is likely to be redundant. Instead, only certain frames may be added to the map. Those frames may be selected based on one or more criteria, such as the degree of overlap with image frames already present in the map, the number of new features they contain, or a quality metric of the features in the frame. Image frames not added to the tracking map may be discarded or may be used to modify the positions of the features. As another alternative, all or most of the image frames represented as a set of features may be retained, but a subset of those frames may be designated as key frames for further processing.
[0206] The key frames may be processed to produce a key rig 704. The key frames may be processed to produce a three-dimensional set of feature points and saved as the key rig 704. For example, such processing may require comparing image frames simultaneously obtained from two cameras to stereoscopically determine the 3D positions of the feature points. Metadata may be associated with these key frames and / or the key rig (e.g., the pose).
[0207] The environmental map can have any one of a variety of formats depending on, for example, the storage location of the environmental map, which includes, for example, local storage and remote storage of the AR device. For example, on a wearable device with limited memory, the map in remote storage can have a higher resolution than the map in local storage. To send a higher resolution map from remote storage to local storage, the map can be downsampled or otherwise converted to an appropriate format, such as by reducing the number of poses for each region of the physical world stored in the map and / or the number of feature points stored for each pose. In some embodiments, slices or portions of the high resolution map from remote storage can be sent to local storage, where the slices or portions are not downsampled.
[0208] When a new tracking map is created, the environmental map database can be updated. To determine which one of the potentially very large number of environmental maps in the database will be updated, the update can include efficiently selecting one or more environmental maps stored in the database that are relevant to the new tracking map. The selected one or more environmental maps can be ranked by relevance, and one or more of the highest ranked maps can be selected for processing to merge the higher ranked selected environmental maps with the new tracking map to create one or more updated environmental maps. When the new tracking map represents a portion of the physical world for which there is no pre-existing environmental map to update, the tracking map can be stored in the database as a new environmental map.
[0209] Watch standalone display
[0210] Methods and apparatus for providing virtual content using an XR system independent of the position of the eyes viewing the virtual content are described herein. Conventionally, virtual content is re-rendered upon any movement of the display system. For example, if a user wearing a display system views a virtual representation of a three-dimensional (3D) object on a display and walks around the area where the 3D object appears, the 3D object should be re-rendered for each viewpoint so that the user has the feeling that he or she is walking around an object that occupies real space. However, re-rendering consumes a large amount of computational resources of the system and results in artifacts due to latency.
[0211] The inventors have recognized and understood that head pose (e.g., the position and orientation of a user wearing an XR system) can be used to render virtual content independent of eye rotation within the user's head. In some embodiments, a dynamic map of a scene can be generated based on multiple coordinate frames in the real space across one or more sessions, such that virtual content interacting with the dynamic map can be rendered robustly, independent of eye rotation within the user's head and / or independent of sensor distortion caused by heat generated, for example, during computationally intensive operations at high speed. In some embodiments, the configuration of the multiple coordinate frames can enable a first XR device worn by a first user and a second XR device worn by a second user to identify a common location in the scene. In some embodiments, the configuration of the multiple coordinate frames can enable a user wearing an XR device to view virtual content at the same location in the scene.
[0212] In some embodiments, a tracking map can be constructed in a world coordinate frame that can have a world origin. When the XR device is powered on, the world origin can be the first pose of the XR device. The world origin can be aligned with gravity, such that developers of XR applications can perform gravity alignment without additional work. Different tracking maps can be constructed in different world coordinate frames, as the tracking maps can be captured by the same XR device in different sessions and / or by different XR devices worn by different users. In some embodiments, a session of the XR device can start from when the device is powered on to when the device is powered off. In some embodiments, the XR device can have a head coordinate frame that can have a head origin. The head origin can be the current pose of the XR device when an image is captured. The difference between the head pose of the world coordinate frame and the head pose of the head coordinate frame can be used to estimate the tracking route.
[0213] In some embodiments, the XR device can have a camera coordinate frame that can have a camera origin. The camera origin can be the current pose of one or more sensors of the XR device. The inventors have recognized and understood that the configuration of the camera coordinate frame enables robust display of virtual content independent of eye rotation within the user's head. The configuration also enables robust display of virtual content independent of sensor distortion caused, for example, by heat generated during operation.
[0214] In some embodiments, an XR device may have a head unit with a head-mounted frame that a user can secure to their head and may include two waveguides, one in front of each of the user's eyes. The waveguides may be transparent such that ambient light from objects in the real world can pass through the waveguides and the user can see the real-world objects. Each waveguide may send projection light from a projector to the corresponding eye of the user. The projection light may form an image on the retina of the eye. Thus, the retina of the eye receives ambient light and projection light. The user may simultaneously see real-world objects as well as one or more virtual objects created by the projection light. In some embodiments, the XR device may have sensors that detect real-world objects around the user. These sensors may be, for example, cameras that capture images that can be processed to identify the positions of real-world objects.
[0215] In some embodiments, instead of attaching virtual content to a world coordinate frame, the XR system may assign a coordinate frame to the virtual content. Such a configuration enables the virtual content to be described without regard to where the virtual content is rendered to the user, but the virtual content may be attached to a more persistent frame location, such as a persistent coordinate frame (PCF) that will be rendered at a specified location, e.g., Figures 14 to 20C as described. When the position of an object changes, the XR device may detect the change in the environmental map and determine the movement of the head unit worn by the user relative to real-world objects.
[0216] Figure 8 Shown is a user experiencing virtual content rendered by an XR system 10 in a physical environment. The XR system may include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is in a physical environment with a real object in the form of a table 16.
[0217] In the example shown, the first XR device 12.1 includes a head unit 22, a hip pack 24, and a cable connection 26. The first user 14.1 secures the head unit 22 to their head and secures the hip pack 24, which is remote from the head unit 22, to their waist. The cable connection 26 connects the head unit 22 to the hip pack 24. The head unit 22 includes technology for displaying one or more virtual objects to the first user 14.1 while allowing the first user 14.1 to see real objects such as the table 16. The hip pack 24 primarily includes the processing and communication capabilities of the first XR device 12.1. In some embodiments, the processing and communication capabilities may reside wholly or partially in the head unit 22 such that the hip pack 24 may be removed or may be located in another device such as a backpack.
[0218] In the example shown, the waist pack 24 is connected to the network 18 via a wireless connection. The server 20 is connected to the network 18 and maintains data representing local content. The waist pack 24 downloads the data representing the local content from the server 20 via the network 18. The waist pack 24 provides the data to the head unit 22 via a cable connection 26. The head unit 22 may include a display having a light source, such as a laser light source or a light emitting diode (LED) light source, and a waveguide to guide the light.
[0219] In some embodiments, the first user 14.1 may mount the head unit 22 to his head and the waist pack 24 to his waist. The waist pack 24 may download image data from the server 20 via the network 18. The first user 14.1 may see the table 16 through the display of the head unit 22. A projector forming part of the head unit 22 may receive image data from the waist pack 24 and generate light based on the image data. The light may travel through one or more waveguides forming part of the display of the head unit 22. The light may then leave the waveguide and propagate onto the retina of the eye of the first user 14.1. The projector may generate light in a pattern that is replicated on the retina of the eye of the first user 14.1. The light that falls on the retina of the eye of the first user 14.1 may have a selected depth of field so that the first user 14.1 perceives an image at a preselected depth behind the waveguide. In addition, the two eyes of the first user 14.1 may receive slightly different images so that the brain of the first user 14.1 perceives one or more three-dimensional images at a selected distance from the head unit 22. In the example shown, the first user 14.1 perceives virtual content 28 above the table 16. The scale of the virtual content 28 and its position and distance from the first user 14.1 are determined by the data representing the virtual content 28 and the various coordinate frames used to display the virtual content 28 to the first user 14.1.
[0220] In the example shown, the virtual content 28 is invisible from the perspective of the drawing and is visible to the first user 14.1 using the first XR device 12.1. The virtual content 28 may initially reside as a data structure within the visual data and algorithms in the waist pack 24. The data structure may then manifest itself as light when the projector of the head unit 22 generates light based on the data structure. It should be understood that although the virtual content 28 does not exist in the three-dimensional space in front of the first user 14.1, the virtual content 28 still represents the three-dimensional space. Figure 1 , to illustrate the wearer perception of the head unit 22. Visualization of computer data in three-dimensional space may be used in this description to show how data structures that contribute to the rendering by one or more users relate to each other within the data structures in the waist pack 24.
[0221] Fig. 9Illustrates components of a first XR device 12.1 according to some embodiments. The first XR device 12.1 may include a head unit 22, and various components that form part of the visual data and algorithms, including, for example, a rendering engine 30, various coordinate frames 32, various origin and destination coordinate frames 34, and various origin-to-destination coordinate frame transformers 36. The various coordinate systems may be based on the intrinsic properties of the XR device, or may be determined by reference to other information, such as the persistent pose or persistent coordinate system described herein.
[0222] The head unit 22 may include a head-mounted frame 40, a display system 42, a real object detection camera 44, a motion tracking camera 46, and an inertial measurement unit 48.
[0223] The head-mounted frame 40 may have a shape that can be fixed to Figure 8 the head of a first user 14.1. The display system 42, the real object detection camera 44, the motion tracking camera 46, and the inertial measurement unit 48 may be mounted to the head-mounted frame 40 and thus move with the head-mounted frame 40.
[0224] The coordinate system 32 may include a local data system 52, a world frame system 54, a head frame system 56, and a camera frame system 58.
[0225] The local data system 52 may include a data channel 62, a local frame determination routine 64, and local frame storage instructions 66. The data channel 62 may be an internal software routine, a hardware component such as an external cable or radio frequency receiver, or a hybrid component such as an open port. The data channel 62 may be configured to receive image data 68 representing virtual content.
[0226] The local frame determination routine 64 may be connected to the data channel 62. The local frame determination routine 64 may be configured to determine a local coordinate frame 70. In some embodiments, the local frame determination routine may determine the local coordinate frame based on a real-world object or real-world location. In some embodiments, the local coordinate frame may be based on the top edge relative to the bottom edge of a browser window, the head or feet of a character, a node on the outer surface of a prism or bounding box enclosing virtual content, or any other suitable location of a coordinate frame that defines the orientation of the virtual content and the location where the virtual content is placed (e.g., a node, such as a placement node or a PCF node).
[0227] The local frame storage instruction 66 can be connected to the local frame determination routine 64. Those skilled in the art will understand that software modules and routines are "connected" to each other through subroutines, calls, etc. The local frame storage instruction 66 can store the local coordinate frame 70 as the local coordinate frame 72 within the origin and destination coordinate frames 34. In some embodiments, the origin and destination coordinate frames 34 can be one or more coordinate frames that can be manipulated or transformed to make virtual content persistent between sessions. In some embodiments, a session can be the time period between the startup and shutdown of an XR device. Two sessions can be two startup and shutdown time periods of a single XR device, or the startup and shutdown time periods of two different XR devices.
[0228] In some embodiments, the origin and destination coordinate frames 34 can be the coordinate frames involved in one or more transformations required for the XR devices of the first user and the second user to identify a common location. In some embodiments, the destination coordinate frame can be the output of a series of calculations and transformations applied to the target coordinate frame so that the first and second users can view virtual content in the same location.
[0229] The rendering engine 30 can be connected to the data channel 62. The rendering engine 30 can receive the image data 68 from the data channel 62, such that the rendering engine 30 can render virtual content at least partially based on the image data 68.
[0230] The display system 42 can be connected to the rendering engine 30. The display system 42 can include components that transform the image data 68 into visible light. The visible light can form two patterns, one for each eye. The visible light can enter Figure 8 the eyes of the first user 14.1 and can be detected on the retinas of the eyes of the first user 14.1.
[0231] The real object detection camera 44 can include one or more cameras that can capture images from different sides of the head-mounted frame 40. The motion tracking camera 46 can include one or more cameras that can capture images on the sides of the head-mounted frame 40. A single set of one or more cameras can be used instead of the two sets of one or more cameras representing the real object detection camera 44 and the motion tracking camera 46. In some embodiments, the cameras 44, 46 can capture images. As described above, these cameras can collect data for constructing a tracking map.
[0232] The inertial measurement unit 48 can include multiple devices for detecting the motion of the head unit 22. The inertial measurement unit 48 can include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of the inertial measurement unit 48 collectively track the motion of the head unit 22 in at least three orthogonal directions and around at least three orthogonal axes.
[0233] In the example shown, the world frame system 54 includes a world surface determination routine 78, a world frame determination routine 80, and world frame storage instructions 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 receives an image and / or keyframes based on the image captured by the real object detection camera 44, and processes the image to identify the surfaces in the image. A depth sensor (not shown) may determine the distance to the surfaces. Thus, the surfaces are represented by data in three dimensions including their size, shape, and distance from the real object detection camera.
[0234] In some embodiments, the world coordinate frame 84 may be based on the origin at the time of initializing the head pose session. In some embodiments, the world coordinate frame may be located at the position where the device is powered on, or if the head pose is lost during the startup session, the world coordinate frame may be located at a new location. In some embodiments, the world coordinate frame may be the origin at the start of the head pose session.
[0235] In the example shown, the world frame determination routine 80 is connected to the world surface determination routine 78, and determines the world coordinate frame 84 based on the positions of the surfaces determined by the world surface determination routine 78. The world frame storage instructions 82 are connected to the world frame determination routine 80 to receive the world coordinate frame 84 from the world frame determination routine 80. The world frame storage instructions 82 store the world coordinate frame 84 as the world coordinate frame 86 within the origin and the destination coordinate frame 34.
[0236] The head frame system 56 may include a head frame determination routine 90 and head frame storage instructions 92. The head frame determination routine 90 may be connected to the motion tracking camera 46 and the inertial measurement unit 48. The head frame determination routine 90 may use data from the motion tracking camera 46 and the inertial measurement unit 48 to calculate the head coordinate frame 94. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The motion tracking camera 46 may continuously capture images that the head frame determination routine 90 uses to refine the head coordinate frame 94. When Figure 8 the first user 14.1 in moves their head, the head unit 22 moves. The motion tracking camera 46 and the inertial measurement unit 48 may continuously provide data to the head frame determination routine 90 such that the head frame determination routine 90 can update the head coordinate frame 94.
[0237] The head frame storage instruction 92 can be connected to the head frame determination routine 90 to receive the head coordinate frame 94 from the head frame determination routine 90. The head frame storage instruction 92 can store the head coordinate frame 94 as the head coordinate frame 96 in the origin and destination coordinate frame 34. The head frame storage instruction 92 can repeatedly store the updated head coordinate frame 94 as the head coordinate frame 96 when the head frame determination routine 90 recalculates the head coordinate frame 94. In some embodiments, the head coordinate frame can be the position of the wearable XR device 12.1 relative to the local coordinate frame 72.
[0238] The camera frame system 58 can include camera intrinsic characteristics 98. The camera intrinsic characteristics 98 can include the dimensions of the head unit 22 as its design and manufacturing features. The camera intrinsic characteristics 98 can be used to calculate the camera coordinate frame 100 stored within the origin and destination coordinate frame 34.
[0239] In some embodiments, the camera coordinate frame 100 can include Figure 8 all the pupil positions of the left eye of the first user 14.1. When the left eye moves from left to right or up and down, the pupil position of the left eye is located within the camera coordinate frame 100. Additionally, the pupil position of the right eye is located within the camera coordinate frame 100 of the right eye. In some embodiments, the camera coordinate frame 100 can include the position of the camera relative to the local coordinate frame when an image is captured.
[0240] The origin to destination coordinate frame transformer 36 can include a local to world coordinate transformer 104, a world to head coordinate transformer 106, and a head to camera coordinate transformer 108. The local to world coordinate transformer 104 can receive the local coordinate frame 72 and transform the local coordinate frame 72 into the world coordinate frame 86. The transformation of the local coordinate frame 72 to the world coordinate frame 86 can be represented as the local coordinate frame transformed into the world coordinate frame 110 within the world coordinate frame 86.
[0241] The world to head coordinate transformer 106 can transform from the world coordinate frame 86 into the head coordinate frame 96. The world to head coordinate transformer 106 can transform the local coordinate frame transformed into the world coordinate frame 110 into the head coordinate frame 96. This transformation can be represented as the local coordinate frame transformed into the head coordinate frame 112 within the head coordinate frame 96.
[0242] The head-to-camera coordinate transformer 108 can transform from the head coordinate frame 96 to the camera coordinate frame 100. The head-to-camera coordinate transformer 108 can transform the local coordinate frame transformed to the head coordinate frame 112 into the local coordinate frame transformed to the camera coordinate frame 114 within the camera coordinate frame 100. The local coordinate frame transformed to the camera coordinate frame 114 can be input into the rendering engine 30. The rendering engine 30 can render the image data 68 representing the local content 28 based on the local coordinate frame transformed to the camera coordinate frame 114.
[0243] Fig.10 is a spatial representation of various origin and destination coordinate frames 34. The local coordinate frame 72, the world coordinate frame 86, the head coordinate frame 96, and the camera coordinate frame 100 are represented in this figure. In some embodiments, when virtual content is placed in the real world so that the user can view the virtual content, the local coordinate frame associated with the XR content 28 can have a position and rotation relative to the local and / or world coordinate frames and / or the PCF (e.g., nodes and orientation directions can be provided). Each camera can have its own camera coordinate frame 100 that includes all the pupil positions of one eye. Reference numerals 104A and 106A respectively represent the transformations performed by Fig. 9 the local-to-world coordinate transformer 104, the world-to-head coordinate transformer 106, and the head-to-camera coordinate transformer 108 in
[0244] Fig.11 depicts a camera rendering protocol for transforming from the head coordinate frame to the camera coordinate frame according to some embodiments. In the example shown, the pupil of a single eye moves from position A to position B. The virtual object to appear stationary will depend on the pupil position projected onto one of the two positions A or B on the depth plane (assuming the camera is configured to use a pupil-based coordinate frame). As a result, when the eye moves from position A to position B, using the pupil coordinate frame transformed to the head coordinate frame will cause jitter in the stationary virtual object. This situation is called view-dependent display or projection.
[0245] As Fig.12 shown, the camera coordinate frame (e.g., CR) is placed and includes all pupil positions, and now the object projection will be consistent regardless of the pupil positions A and B. The head coordinate frame is transformed into the CR frame, which is called view-independent display or projection. Image reprojection can be applied to the virtual content to address changes in the eye position. However, since the rendering remains in the same position, jitter can be minimized.
[0246] Fig.13The display system 42 is shown in more detail. The display system 42 includes a stereoscopic analyzer 144 that is connected to the rendering engine 30 and forms part of the visual data and algorithms.
[0247] The display system 42 further includes a left projector 166A and a right projector 166B, as well as a left waveguide 170A and a right waveguide 170B. The left projector 166A and the right projector 166B are connected to a power source. Each projector 166A and 166B has a corresponding input for the image data to be provided to the respective projector 166A or 166B. The respective projector 166A or 166B generates and emits light in a two-dimensional pattern when powered on. The left waveguide 170A and the right waveguide 170B are positioned to receive light from the left projector 166A and the right projector 166B, respectively. The left waveguide 170A and the right waveguide 170B are transparent waveguides.
[0248] In use, the user mounts the head-mounted frame 40 onto their head. The components of the head-mounted frame 40 can include, for example, a strap (not shown) that wraps around the back of the user's head. The left waveguide 170A and the right waveguide 170B are then positioned in front of the user's left eye 220A and right eye 220B.
[0249] The rendering engine 30 inputs the image data it receives into the stereoscopic analyzer 144. The image data is Figure 8 the three-dimensional image data of the local content 28. The image data is projected onto a plurality of virtual planes. The stereoscopic analyzer 144 analyzes the image data to determine a left image data set and a right image data set based on the image data projected onto each depth plane. The left image data set and the right image data set are data sets representing two-dimensional images that are projected in three dimensions to give the user a sense of depth.
[0250] The stereoscopic analyzer 144 inputs the left image data set and the right image data set into the left projector 166A and the right projector 166B. Then, the left projector 166A and the right projector 166B create a left illumination pattern and a right illumination pattern. The components of the display system 42 are shown in a plan view, but it should be understood that when shown in a front view, the left pattern and the right pattern are two-dimensional patterns. Each light pattern includes a plurality of pixels. For illustrative purposes, light rays 224A and 226A from two pixels are shown leaving the left projector 166A and entering the left waveguide 170A. The light rays 224A and 226A are reflected from the side of the left waveguide 170A. The light rays 224A and 226A are shown propagating from left to right within the left waveguide 170A by total internal reflection, but it should be understood that the light rays 224A and 226A also propagate in a certain direction into the paper using a refraction and reflection system.
[0251] Light rays 224A and 226A leave the left optical waveguide 170A through the pupil 228A, and then enter the left eye 220A through the pupil 230A of the left eye 220A. Then, the light rays 224A and 226A fall on the retina 232A of the left eye 220A. In this way, the left light pattern falls on the retina 232A of the left eye 220A. To the user, the pixels formed on the retina 232A are perceived as pixels 234A and 236A at a certain distance on the side of the left optical waveguide 170A opposite to the left eye 220A. Depth perception is created by manipulating the focal length of the light.
[0252] In a similar manner, the stereo analyzer 144 inputs the right image dataset into the right projector 166B. The right projector 166B transmits a right light pattern, which is represented by pixels in the form of light rays 224B and 226B. The light rays 224B and 226B are reflected within the right optical waveguide 170B and exit through the pupil 228B. The light rays 224B and 226B then enter through the pupil 230B of the right eye 220B and fall on the retina 232B of the right eye 220B. The pixels of the light rays 224B and 226B are perceived as pixels 134B and 236B behind the right optical waveguide 170B.
[0253] The patterns created on the retinas 232A and 232B are perceived as the left image and the right image, respectively. Due to the function of the stereo analyzer 144, the left image and the right image are slightly different from each other. The left image and the right image are perceived as a three-dimensional rendering in the user's mind.
[0254] As mentioned, the left optical waveguide 170A and the right optical waveguide 170B are transparent. Light from real objects such as a table 16 on the side of the left optical waveguide 170A and the right optical waveguide 170B opposite to the eyes 220A and 220B can be projected through the left optical waveguide 170A and the right optical waveguide 170B and fall on the retinas 232A and 232B.
[0255] Persistent Coordinate Framework (PCF)
[0256] Methods and apparatuses for providing spatial persistence between user instances in a shared space are described herein. Without spatial persistence, virtual content placed by a user in the physical world during a session may not be present or may be misplaced in the user's view in different sessions. Without spatial persistence, virtual content placed by one user in the physical world may not be present or may be misaligned in the view of a second user, even if the second user intends to share the same physical space experience as the first user.
[0257] The inventors have recognized and understood that spatial persistence can be provided through a Persistent Coordinate Frame (PCF). The PCF can be defined based on one or more points that represent features identified in the physical world (e.g., corners, edges). The features can be selected such that they appear the same from one user instance of the XR system to another.
[0258] In addition, when rendering relative to a local map based only on a tracking map, drift during tracking that causes the computed tracking path (e.g., camera trajectory) to deviate from the actual tracking path can result in misalignment of the position of virtual content. As the XR device collects more information about the scene over time, the tracking map of the space can be refined to correct the drift. However, if virtual content is placed on a real object and saved relative to the device's world coordinate frame derived from the tracking map before the map is refined, the virtual content may appear displaced as if the real object has moved during the map refinement. The PCF can be updated based on the map refinement because the PCF is defined based on features and is updated as the features move during the map refinement.
[0259] In some embodiments, persistent spatial information can be represented in a manner that can be easily shared among users and among distributed components including applications. For example, information about the physical world can be represented as a Persistent Coordinate Frame (PCF). The PCF can be defined based on one or more points that represent features identified in the physical world. The features can be selected such that they may be the same from user session to user session in the XR system. The PCF may be sparsely present, providing less than all available information about the physical world, such that they can be efficiently processed and transmitted. Techniques for processing persistent spatial information can include creating a dynamic map based on one or more coordinate systems in the real space across one or more sessions, and generating a Persistent Coordinate Frame (PCF) on the sparse map, which can be exposed to XR applications via, for example, an Application Programming Interface (API). These capabilities can be supported by techniques for ranking and merging multiple maps created by one or more XR devices. Persistent spatial information can also enable fast recovery and reset of the head pose on each of one or more XR devices in a computationally efficient manner.
[0260] The PCF can include six degrees of freedom for translation and rotation with respect to the map coordinate system. The PCF can be stored in a local storage medium and / or a remote storage medium. Depending on, for example, the storage location, the translation and rotation of the PCF can be computed with respect to the map coordinate system. For example, a PCF used locally by a device may have translation and rotation with respect to the device's world coordinate frame. A PCF in the cloud may have translation and rotation with respect to a canonical coordinate frame of a canonical map.
[0261] PCFs can provide a sparse representation of the physical world, providing less information about the physical world than all available information, enabling them to be processed and transferred efficiently. Techniques for processing persistent spatial information can include creating a dynamic map based on one or more coordinate systems in the real space spanning one or more sessions, generating persistent coordinate frames (PCFs) on the sparse map, which can be exposed to XR applications through, for example, an application programming interface (API).
[0262] Fig.14 is a block diagram showing the creation of a persistent coordinate frame (PCF) and the addition of XR content to the PCF according to some embodiments. Each block can represent digital information stored in a computer memory. In the case of application 1180, the data can represent computer-executable instructions. In the case of virtual content 1170, the digital information can define, for example, a virtual object specified by application 1180. In the case of other blocks, the digital information can characterize certain aspects of the physical world.
[0263] In the illustrated embodiment, one or more PCFs are created based on images captured by sensors on the wearable device. In Fig.14 the embodiment, the sensors are visual image cameras. These cameras can be the same cameras as those used to form the tracking map. Thus, some of the processing proposed by Fig.14 can be performed as part of updating the tracking map. However, Fig.14 shows that information providing persistence is generated in addition to the tracking map.
[0264] To derive a 3D PCF, two images 1110 from two cameras installed on the wearable device in a configuration capable of performing stereoscopic image analysis are processed together. Fig.14 Images 1 and 2 are shown, each of images 1 and 2 being from one of the cameras. For simplicity, a single image from each camera is shown. However, each camera can output a stream of image frames, and the processing of Fig.14 can be performed for multiple image frames in the stream.
[0265] Thus, images 1 and 2 can be respectively one frame in a sequence of image frames. The processing shown in Fig.14 can be repeated for consecutive image frames in the sequence until an image frame containing feature points provides a suitable image to form persistent spatial information based on that image. Alternatively or additionally, when the user moves such that the user is no longer close enough to a previously identified PCF to reliably use that PCF to determine the position relative to the physical world, the processing of Fig.14Processing. For example, the XR system can maintain the current PCF for the user. When the distance exceeds a threshold, the system can switch to a new current PCF that is closer to the user, which can be generated using the image frames obtained at the user's current location according to Fig.14 the process.
[0266] Even when generating a single PCF, a stream of image frames can be processed to identify image frames that depict content in the physical world, which may be stable and can be easily recognized by a device near the physical world region depicted in the image frame. In Fig.14 an embodiment, the processing begins with the identification of features 1120 in the image. For example, features can be identified by looking for the locations in the image where the gradient exceeds a threshold or other features, which may correspond to, for example, the corners of an object. In the illustrated embodiment, the features are points, but other recognizable features, such as edges, can be used alternatively or additionally.
[0267] In the illustrated embodiment, a fixed number N of features 1120 are selected for further processing. Those feature points can be selected based on one or more criteria, such as the magnitude of the gradient or proximity to other feature points. Alternatively or additionally, feature points can be tentatively selected, for example, based on characteristics that imply the feature points are persistent. For example, heuristics can be defined based on the characteristics of feature points that may correspond to the corners of windows or doors or large pieces of furniture. Such heuristics may take into account the feature points themselves and the things around them. As a specific example, the number of feature points per image can be between 100 and 500 or between 150 and 250, such as 200.
[0268] Regardless of the number of feature points selected, descriptors 1130 can be calculated for the feature points. In this example, descriptors are calculated for each selected feature point, but descriptors can be calculated for groups of feature points or subsets of feature points or all features within the image. The descriptors characterize the feature points so that feature points representing the same object in the physical world are assigned similar descriptors. The descriptors can enable the alignment of two frames, such as may occur when localizing one map relative to another. Instead of searching for the relative orientation of the frames that minimizes the distance between the feature points of two images, the initial alignment of the two frames can be performed by identifying feature points with similar descriptors. The alignment of the image frames can be based on the aligned points with similar descriptors, which may require less processing compared to calculating the alignment of all feature points in the image.
[0269] A descriptor can be calculated as a mapping from feature points to descriptors, or in some embodiments, as a mapping from patches of the image around the feature points to descriptors. The descriptor can be a numerical quantity. U.S. Patent Application Publication No. 2019 / 0147341, entitled "Fully Convolutional Interest Point Detection and Description via Homographic Adaptation," describes a computational descriptor for feature points and is hereby incorporated by reference in its entirety.
[0270] In Fig.14 the example of, a descriptor 1130 is calculated for each feature point in each image frame. Based on the descriptor and / or the feature points and / or the image itself, an image frame can be identified as a key frame 1140. In the illustrated embodiment, a key frame is an image frame that meets a certain criterion and is then selected for further processing. For example, when creating a tracking map, an image frame that adds meaningful information to the map can be selected as a key frame to be integrated into the map. On the other hand, an image frame that substantially overlaps with an area where an image frame has already been integrated into the map can be discarded so that it does not become a key frame. Alternatively or additionally, a key frame can be selected based on the number and / or type of feature points in the image frame. In Fig.14 the embodiment of, a key frame 1150 that is selected to be included in the tracking map can also be considered a key frame for determining the PCF, but different or additional criteria for selecting the key frame for generating the PCF can be used.
[0271] Although Fig.14 it is shown that key frames are used for further processing, the information obtained from the images can be processed in other forms. For example, feature points in a key assembly can be processed alternatively or additionally. Also, although key frames are described as being derived from a single image frame, there does not have to be a one-to-one relationship between a key frame and the acquired image frames. For example, a key frame can be obtained from multiple image frames, such as by stitching or aggregating the image frames together so that only features that appear in multiple images are retained in the key frame.
[0272] A key frame can include image information and / or metadata associated with the image information. In some embodiments, it can be the cameras 44, 46 ( Fig. 9)The captured images are computed as one or more key frames (e.g., key frames 1, 2). In some embodiments, the key frames may include camera poses. In some embodiments, the key frames may include one or more camera images captured at the camera pose. In some embodiments, the XR system may determine that a portion of the camera image captured at the camera pose is useless and thus does not include that portion in the key frame. Thus, aligning new images with an earlier perception of the scene using key frames can reduce the use of computing resources of the XR system. In some embodiments, the key frames may include images and / or image data at positions with orientations / angles. In some embodiments, the key frames may include positions and orientations from which one or more map points can be observed. In some embodiments, the key frames may include coordinate frames with IDs. U.S. Patent Application No. 15 / 877,359 describes key frames and incorporates its entire content herein by reference.
[0273] Some or all of the key frames 1140 may be selected for further processing, such as generating a persistent pose 1150 for the key frame. This selection may be based on the characteristics of all or a subset of the feature points in the image frame. These characteristics may be determined by processing the descriptors, features, and / or the image frame itself. As a specific example, the selection may be based on clusters of feature points identified as likely related to persistent objects.
[0274] Each key frame is associated with the pose of the camera that acquired the key frame. For the key frames selected for processing into persistent poses, this pose information may be saved along with other metadata about the key frame, such as WiFi fingerprints and / or GPS coordinates at the time of acquisition and / or at the acquisition location. In some embodiments, metadata such as GPS coordinates may be used, either alone or in combination, as part of the localization process.
[0275] A persistent pose is an information source that the device can use to orient itself relative to previously acquired information about the physical world. For example, if the key frame from which the persistent pose was created is incorporated into a map of the physical world, the device can use a sufficient number of feature points in the key frame associated with the persistent pose to orient itself relative to that persistent pose. The device can align its current image of the surrounding environment with the persistent pose. This alignment may be based on matching the current image with the image 1110, features 1120, and / or descriptors 1130 that gave rise to the persistent pose, or any subset of that image or those features or descriptors. In some embodiments, the current image frame that matches the persistent pose may be another key frame that has been incorporated into the device's tracking map.
[0276] The information about the persistent pose may be stored in a format that facilitates sharing among multiple applications that may execute on the same or different devices. In Fig.14In an example, some or all of the persistent poses may be reflected as a Persistent Coordinate Frame (PCF) 1160. Similar to the persistent poses, the PCF may be associated with a map and may include a set of features or other information that the device can use to determine its orientation relative to the PCF. The PCF may include a transformation that defines a transformation relative to the origin of the map such that by associating its position with the PCF, the device can determine its position relative to any object in the physical world reflected in the map.
[0277] Since the PCF provides a mechanism for determining positions relative to physical objects, an application (such as application 1180) can define the positions of virtual objects relative to one or more PCFs, and these positions are used as anchors for virtual content 1170. For example, Fig.14 shows that App 1 has associated its virtual content 2 with PCF 1.2. Similarly, App 2 has associated its virtual content 3 with PCF 1.2. It is also shown that App 1 associates its virtual content 1 with PCF 4.5, and it is shown that App 2 associates its virtual content 4 with PCF 3. In some embodiments, PCF 3 may be based on Image 3 (not shown), and PCF 4.5 may be based on Image 4 and Image 5 (not shown), similar to how PCF 1.2 is based on Image 1 and Image 2. When rendering this virtual content, the device can apply one or more transformations to calculate information such as the position of the virtual content relative to the device's display and / or the position of the physical object relative to the desired position of the virtual content. Using the PCF as a reference can simplify such calculations.
[0278] In some embodiments, the persistent pose may be a coordinate position and / or orientation with one or more associated key frames. In some embodiments, a persistent pose may be automatically created after the user has traveled a certain distance (such as three meters). In some embodiments, the persistent pose may be used as a reference point during positioning. In some embodiments, the persistent pose may be stored in the traversable world (e.g., traversable world module 538).
[0279] In some embodiments, a new PCF may be determined based on a predetermined distance allowed between adjacent PCFs. In some embodiments, when the user travels a predetermined distance (such as five meters), one or more persistent poses may be calculated into the PCF. In some embodiments, the PCF may be associated with one or more world coordinate frames and / or canonical coordinate frames in the traversable world, for example. In some embodiments, depending on, for example, security settings, the PCF may be stored in a local database and / or a remote database.
[0280] Fig.15A method 4700 of establishing and using a persistent coordinate system according to some embodiments is shown. Method 4700 may begin with capturing (act 4702) an image of a scene (e.g., images 1 and 2 in Fig.14 ) using one or more sensors of an XR device. Multiple cameras may be used, and one camera may generate multiple images, e.g., in the form of a stream.
[0281] Method 4700 may include extracting (4704) points of interest (e.g., map point 702 in Figure 7 , feature 1120 in Fig.14 ) from the captured image, generating (act 4706) a descriptor of the extracted points of interest (e.g., descriptor 1130 in Fig.14 ), and generating (act 4708) a key frame (e.g., key frame 1140) based on the descriptor. In some embodiments, the method may compare the points of interest in the key frames and form pairs of key frames that share a predetermined amount of points of interest. The method may use each pair of key frames to reconstruct a portion of the physical world. The mapped portion of the physical world may be saved as 3D features (e.g., key assembly 704 in Figure 7 ). In some embodiments, a selected portion of the pair of key frames may be used to construct 3D features. In some embodiments, the result of the mapping may be selectively saved. Key frames not used to construct 3D features may be associated with the 3D features by pose, e.g., representing the distance between key frames using the covariance matrix between the poses of the key frames. In some embodiments, pairs of key frames may be selected to construct 3D features such that the distance between each two of the constructed 3D features is within a predetermined distance, which may be determined to balance the required computational amount and the accuracy level of the resulting model. Such a method can provide an XR system with a model of the physical world having an amount of data suitable for performing efficient and accurate calculations. In some embodiments, the covariance matrix of two images may include the covariance between the poses (e.g., six degrees of freedom) of the two images.
[0282] Method 4700 may include generating (act 4710) a persistent pose based on the key frames. In some embodiments, the method may include generating a persistent pose based on the 3D features reconstructed from pairs of key frames. In some embodiments, the persistent pose may be attached to the 3D features. In some embodiments, the persistent pose may include the poses of the key frames used to construct the 3D features. In some embodiments, the persistent pose may include the average pose of the key frames used to construct the 3D features. In some embodiments, a persistent pose may be generated such that the distance between adjacent persistent poses is within a predetermined value, e.g., in the range of one meter to five meters, any value in between, or any other suitable value. In some embodiments, the distance between adjacent persistent poses may be represented by the covariance matrix of the adjacent persistent poses.
[0283] Method 4700 may include generating (act 4712) a PCF based on a persistent pose. In some embodiments, the PCF may be attached to a 3D feature. In some embodiments, the PCF may be associated with one or more persistent poses. In some embodiments, the PCF may include the pose of one of the associated persistent poses. In some embodiments, the PCF may include the average pose of the poses of the associated persistent poses. In some embodiments, the PCF may be generated such that the distance between adjacent PCFs is within a predetermined value, such as in the range of three meters to ten meters, any value therebetween, or any other suitable value. In some embodiments, the distance between adjacent PCFs may be represented by the covariance matrix of the adjacent PCFs. In some embodiments, the PCF may be exposed to an XR application via, for example, an application programming interface (API) such that the XR application can access the model of the physical world through the PCF without accessing the model itself.
[0284] Method 4700 may include associating (act 4714) image data of a virtual object to be displayed by an XR device with at least one of the PCFs. In some embodiments, the method may include calculating the translation and orientation of the virtual object relative to the associated PCF. It should be understood that it is not necessary to associate the virtual object with the PCF generated by the device that places the virtual object. For example, the device may obtain a saved PCF in a canonical map in the cloud and associate the virtual object with the obtained PCF. It should be understood that when the PCF is adjusted over time, the virtual object may move together with the associated PCF.
[0285] Fig.16 Visual data and algorithms of a first XR device 12.1, a second XR device 12.2, and a server 20 according to some embodiments are shown. Fig.16 The components shown therein may operate to perform some or all of the operations associated with generating, updating, and / or using spatial information (such as persistent poses, persistent coordinate systems, tracking maps, or canonical maps) as described herein. Although not shown, the first XR device 12.1 may be configured to be the same as the second XR device 12.2. The server 20 may have a map storage routine 118, a canonical map 120, a map transmitter 122, and a map merging algorithm 124.
[0286] A second XR device 12.2 that may be in the same scene as the first XR device 12.1 may include a permanent coordinate frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that may be used to render a virtual object, and a frame embedding generator 308 (see Fig.21 ). In some embodiments, a map download system 126, a PCF recognition system 128, the ground Figure 2 The positioning module 130, the canonical map merger 132, the canonical map 133, and the map publisher 136 are assembled into the traversable world unit 1304. The PCF integration unit 1300 can be connected to the traversable world unit 1304 and other components of the second XR device 12.2 to allow for the acquisition, generation, use, upload, and download of PCFs.
[0287] Maps that include PCFs can achieve more persistence in a changing world. In some embodiments, the positioning, such as tracking a map of matching features, e.g., an image, can include selecting features representing persistent content from a map composed of PCFs, which enables fast matching and / or positioning. For example, in a world where people enter and exit a scene and objects such as doors move relative to the scene, less storage space and transmission rate are required, and it is possible to map the scene using individual PCFs and their relationships to each other (e.g., the integrated constellation of PCFs).
[0288] In some embodiments, the PCF integration unit 1300 can include a PCF 1306, a PCF tracker 1308, a persistent pose acquirer 1310, a PCF inspector 1312, a PCF generation system 1314, a coordinate frame calculator 1316, a persistent pose calculator 1318, and three transformers including a tracking map and persistent pose transformer 1320, a persistent pose and PCF transformer 1322, and a PCF and image data transformer 1324, previously stored in the data storage of the storage unit of the second XR device 12.2.
[0289] In some embodiments, the PCF tracker 1308 can have an open prompt and a close prompt selectable by the application 1302. The application 1302 can be executable by the processor of the second XR device 12.2 to, for example, display virtual content. The application 1302 can have a call to open the PCF tracker 1308 via the open prompt. When the PCF tracker 1308 is open, the PCF tracker 1308 can generate a PCF. The application 1302 can have a subsequent call to close the PCF tracker 1308 via the close prompt. When the PCF tracker 1308 is closed, the PCF tracker 1308 terminates PCF generation.
[0290] In some embodiments, the server 20 can include a plurality of persistent poses 1332 and a plurality of PCFs 1330 that have been previously saved in association with the canonical map 120. The map transmitter 122 can send the canonical map 120 together with the persistent poses 1332 and / or PCF 1330 to the second XR device 12.2. The persistent poses 1332 and PCFs 1330 can be stored on the second XR device 12.2 in association with the canonical map 133. When loc Figure 2When located to the canonical map 133, it can store the persistent pose 1332 and the PCF 1330 in association with the ground. Figure 2 The persistent pose 1332 and the PCF 1330 can be stored in association with the ground.
[0291] In some embodiments, the persistent pose acquirer 1310 can acquire the persistent pose of the ground. The PCF checker 1312 can be connected to the persistent pose acquirer 1310. The PCF checker 1312 can obtain the PCF from the PCF 1306 based on the persistent pose obtained by the persistent pose acquirer 1310. The PCF obtained by the PCF checker 1312 can form an initial group of PCFs for image display based on the PCF. Figure 2
[0292] Figure 2 In some embodiments, the application 1302 may need to generate additional PCFs. For example, if the user moves to an area that has not been mapped before, the application 1302 can open the PCF tracker 1308. The PCF generation system 1314 can be connected to the PCF tracker 1308 and start generating PCFs based on the ground as it starts to expand. The PCFs generated by the PCF generation system 1314 can form a second group of PCFs that can be used for image display based on the PCF. Figure 2
[0293] The coordinate frame calculator 1316 can be connected to the PCF checker 1312. After the PCF checker 1312 obtains the PCF, the coordinate frame calculator 1316 can call the head coordinate frame 96 to determine the head pose of the second XR device 12.2. The coordinate frame calculator 1316 can also call the persistent pose calculator 1318. The persistent pose calculator 1318 can be directly or indirectly connected to the frame embedding generator 308. In some embodiments, an image / frame can be specified as a key frame after traveling a threshold distance (e.g., 3 meters) from a previous key frame. The persistent pose calculator 1318 can generate a persistent pose based on multiple (e.g., three) key frames. In some embodiments, the persistent pose can be substantially the average of the coordinate frames of multiple key frames.
[0294] Figure 2 Figure 2 Figure 2 The tracking map and persistent pose transformer 1320 can be connected to the ground and the persistent pose calculator 1318. The tracking map and persistent pose transformer 1320 can transform the ground into a persistent pose to determine the persistent pose at the origin with respect to the ground.
[0295] The persistent pose and PCF transformer 1322 can be connected to the tracking map and persistent pose transformer 1320, and further connected to the PCF checker 1312 and the PCF generation system 1314. The persistent pose and PCF transformer 1322 can transform the persistent pose (to which the tracking map has been transformed) from the PCF checker 1312 and the PCF generation system 1314 into a PCF to determine the PCF relative to the persistent pose.
[0296] The PCF and image data transformer 1324 can be connected to the persistent pose and PCF transformer 1322 and the data channel 62. The PCF and image data transformer 1324 transforms the PCF into image data 68. The rendering engine 30 can be connected to the PCF and image data transformer 1324 to display the image data 68 to the user relative to the PCF.
[0297] The PCF integration unit 1300 can store additional PCFs generated by the PCF generation system 1314 in the PCF 1306. The PCF 1306 can be stored relative to the persistent pose. When the map publisher 136 sends the map to the server 20, Figure 2 the map publisher 136 can obtain the PCF 1306 and the persistent pose associated with the PCF 1306, and the map publisher 136 also sends the PCF and the persistent pose associated with the map to the server 20. Figure 2 When the map storage routine 118 of the server 20 stores the map, Figure 2 the map storage routine 118 can also store the persistent pose and PCF generated by the second viewing device 12.2. The map merging algorithm 124 can use the persistent pose and PCF of the map associated with the canonical map 120 and stored in the persistent pose 1332 and the PCF 1330 respectively to create the canonical map 120. Figure 2
[0298] The first XR device 12.1 can include a PCF integration unit similar to the PCF integration unit 1300 of the second XR device 12.2. When the map sender 122 sends the canonical map 120 to the first XR device 12.1, the map sender 122 can send the persistent pose 1332 and the PCF 1330 associated with the canonical map 120 and originating from the second XR device 12.2. The first XR device 12.1 can store the PCF and the persistent pose in the data storage on the storage device of the first XR device 12.1. Then, the first XR device 12.1 can utilize the persistent pose and PCF originating from the second XR device 12.2 for image display relative to the PCF. Additionally or alternatively, the first XR device 12.1 can obtain, generate, use, upload, and download PCFs and persistent poses in a manner similar to that of the second XR device 12.2 described above.
[0299] In the example shown, the first XR device 12.1 generates a local tracking map (hereinafter referred to as "the Figure 1 "), and the map storage routine 118 receives the Figure 1 from the first XR device 12.1. Then, the map storage routine 118 stores the Figure 1 as the canonical map 120 on the storage device of the server 20.
[0300] The second XR device 12.2 includes a map download system 126, an anchor recognition system 128, a positioning module 130, a canonical map merger 132, a local content positioning system 134, and a map publisher 136.
[0301] In use, the map transmitter 122 sends the canonical map 120 to the second XR device 12.2, and the map download system 126 downloads and stores the canonical map 120 as the canonical map 133 from the server 20.
[0302] The anchor recognition system 128 is connected to the world surface determination routine 78. The anchor recognition system 128 identifies anchors based on objects detected by the world surface determination routine 78. The anchor recognition system 128 uses the anchors to generate a second map (the Figure 2 ). As shown in loop 138, the anchor recognition system 128 continues to identify anchors and continues to update the Figure 2 . Based on the data provided by the world surface determination routine 78, the positions of the anchors are recorded as three-dimensional data. The world surface determination routine 78 receives images from the real object detection camera 44 and depth data from the depth sensor 135 to determine the positions of the surfaces and their relative distances from the depth sensor 135.
[0303] The positioning module 130 is connected to the canonical map 133 and the Figure 2 . The positioning module 130 repeatedly attempts to position the Figure 2 to the canonical map 133. The canonical map merger 132 is connected to the canonical map 133 and the Figure 2 . When the positioning module 130 positions the Figure 2 to the canonical map 133, the canonical map merger 132 merges the canonical map 133 into the anchors of the Figure 2 . Then, the missing data included in the canonical map is used to update the Figure 2 .
[0304] The local content positioning system 134 is connected to the Figure 2 . The local content positioning system 134 can be, for example, a system that allows a user to position local content at a specific location within the world coordinate framework. Then, the local content attaches itself to the Figure 2An anchor point. The local-to-world coordinate transformer 104 transforms the local coordinate frame into the world coordinate frame based on the settings of the local content positioning system 134. It has been referenced Figure 2 The functions of the rendering engine 30, the display system 42, and the data channel 62 have been described.
[0305] The map publisher 136 uploads the map Figure 2 to the server 20. The map storage routine 118 of the server 20 then stores the map Figure 2 in the storage medium of the server 20.
[0306] The map merging algorithm 124 merges the map Figure 2 with the canonical map 120. When more than two maps have been stored (e.g., three or four maps related to the same or adjacent regions of the physical world), the map merging algorithm 124 merges all the maps into the canonical map 120 to render a new canonical map 120. Then, the map transmitter 122 sends the new canonical map 120 to any and all devices 12.1 and 12.2 located in the region represented by the new canonical map 120. When the devices 12.1 and 12.2 align their respective maps to the canonical map 120, the canonical map 120 becomes the upgraded map.
[0307] Fig.17 Illustrates examples of generating key frames for a map of a scene according to some embodiments. In the example shown, a first key frame KF1 is generated for the door on the left wall of the room. A second key frame KF2 is generated for the corner area where the floor, left wall, and right wall of the room intersect. A third key frame KF3 is generated for the window area on the right wall of the room. On the floor of the wall, a fourth key frame KF4 is generated for the distal area of the carpet. A fifth key frame KF5 is generated for the area of the carpet closest to the user.
[0308] Fig.18 Shows examples of generating persistent poses for a map according to some embodiments. Fig.17 In some embodiments, a new persistent pose is created when the device measures a threshold distance of travel, and / or when the application requests a new persistent pose (PP). In some embodiments, the threshold distance can be 3 meters, 5 meters, 20 meters, or any other suitable distance. Selecting a smaller threshold distance (e.g., 1m) may result in an increase in computational load because a larger number of PPs can be created and managed compared to a larger threshold distance. Selecting a larger threshold distance (e.g., 40m) may result in an increase in virtual content placement errors because a smaller number of PPs will be created, which will result in a smaller number of PCFs being created, meaning that the virtual content attached to the PCF may be a relatively large distance (e.g., 30m) away from the PCF, and the error increases with the increasing distance from the PCF to the virtual content.
[0309] In some embodiments, a PP may be created at the start of a new session. This initial PP may be considered as zero and may be visualized as the center of a circle having a radius equal to the threshold distance. When the device reaches the circumference of the circle and, in some embodiments, the application requests a new PP, the new PP may be placed at the current location of the device (at the threshold distance). In some embodiments, if the device is able to find an existing PP within the threshold distance from the new location of the device, a new PP will not be created at the threshold distance. In some embodiments, when a new PP is created (e.g., Fig.14 PP 1150 in
[0310] ), the device attaches one or more of the closest keyframes to the PP. In some embodiments, the position of the PP relative to the keyframes may be based on the position of the device when the PP was created. In some embodiments, a PP will not be created when the device travels the threshold distance unless the application requests a PP. Fig.18 Shows a first permanent pose PP1, which may have the closest keyframes (e.g., KF1, KF2, and KF3) attached by calculating the relative pose between the keyframes and the persistent pose. Fig.18 Also shows a second permanent pose PP2, which may have the closest keyframes (e.g., KF4 and KF5) attached.
[0311] Fig.19 Shows an example of generating a PCF for a Fig.17 map according to some embodiments. In the example shown, PCF1 may include PP1 and PP2. As described above, the PCF may be used to display image data related to the PCF. In some embodiments, each PCF may have coordinates and a PCF descriptor in another coordinate frame (e.g., the world coordinate frame), e.g., to uniquely identify the PCF. In some embodiments, the PCF descriptor may be calculated based on the feature descriptors of the features in the frames associated with the PCF. In some embodiments, the various constellations of the PCF may be combined to represent the real world in a persistent manner that requires less data and less data transmission.
[0312] Figures 20A to 20C is a schematic diagram showing an example of establishing and using a persistent coordinate system. Fig. 20AShows two users 4802A, 4802B with respective local tracking maps 4804A, 4804B that have not been located to a canonical map. The origins 4806A, 4806B of each user are depicted by a coordinate system (e.g., a world coordinate system) in their respective regions. These origins of each tracking map may be local to each user as they depend on the orientation of their respective devices when tracking is initiated.
[0313] When the sensors of the user device scan the environment, the device can capture images as described above in connection with Fig.14 that may contain features representing persistent objects such that those images can be classified as key frames and persistent poses can be created based on those key frames. In this example, the tracking map 4802A includes a persistent pose (PP) 4808A; the tracking 4802B includes a PP 4808B.
[0314] Similarly as described above in connection with Fig.14 Some PPs can be classified as PCFs, which are used to determine the orientation of virtual content for rendering it to the user. Fig. 20B Shows that the XR devices worn by the respective users 4802A, 4802B can create local PCFs 4810A, 4810B based on the PPs 4808A, 4808B. Fig. 20C Shows that persistent content 4812A, 4812B (e.g., virtual content) can be attached to the PCFs 4810A, 4810B by the respective XR devices.
[0315] In this example, the virtual content can have a virtual content coordinate frame that can be used by the application generating the virtual content regardless of how the virtual content is to be displayed. For example, the virtual content can be specified as a surface at a particular position and angle relative to the virtual content coordinate frame, such as a triangle of a mesh. To render this virtual content to the user, the positions of those surfaces can be determined relative to the user who is to perceive the virtual content.
[0316] Attaching the virtual content to the PCF can simplify the calculations involved in determining the position of the virtual content relative to the user. The position of the virtual content relative to the user can be determined by applying a series of transformations. Some of these transformations may change and may be updated frequently. Other of these transformations may be stable and may be updated frequently or not at all. In any case, the transformations can be applied with a relatively low computational burden such that the position of the virtual content can be updated frequently relative to the user, thereby providing a realistic appearance for the rendered virtual content.
[0317] In FIG. 20A to FIG. 20CIn the example, the device of User 1 has a coordinate system that is related to the coordinate system defining the map origin through the transformation rig1_T_w1. The device of User 2 has a similar transformation rig2_T_w2. These transformations can be represented as 6 degrees of transformation, specifying translation and rotation to align the device coordinate system with the map coordinate system. In some embodiments, the transformation can be represented as two separate transformations, one specifying translation and the other specifying rotation. Thus, it should be understood that the transformation can be expressed in a form that simplifies calculations or provides other advantages.
[0318] The transformation tracking the origin of the map and the PCF identified by the corresponding user device is represented as pcf1_T_w1 and pcf2_T_w2. In this example, the PCF and the PP are the same, so the same transformation also characterizes the PP.
[0319] Thus, the position of the user device relative to the PCF can be calculated by the serial application of these transformations, e.g., rig1_T_pcf1 = (rig1_T_w1)*(pcf1_T_w1).
[0320] As Fig. 20C shown, the virtual content is positioned relative to the PCF through the transformation of obj1_T_pcf1. This transformation can be set by the application generating the virtual content, which can receive information describing the physical object relative to the PCF from the world reconstruction system. To render the virtual content to the user, the transformation to the coordinate system of the user device is calculated, which can be calculated by associating the virtual content coordinate frame to the origin of the tracking map through the transformation obj1_t_w1 = (obj1_T_pcf1)*(pcf1_T_w1). Then, this transformation can be related to the user's device through the further transformation rig1_T_w1.
[0321] Based on the output from the application generating the virtual content, the position of the virtual content can change. When it changes, the end-to-end transformation from the source coordinate system to the destination coordinate system can be recalculated. Additionally, the position and / or head pose of the user can change as the user moves. As a result, the transformation rig1_T_w1 can change, and any end-to-end transformation depending on the user's position or head pose can also change.
[0322] The transformation rig1_T_w1 can be updated as the user moves based on tracking the user's position relative to stationary objects in the physical world. Such tracking can be performed by the head pose tracking component or other components of the system that processes the image sequence as described above. Such an update can be made by determining the user's pose relative to a fixed reference frame (e.g., the PP).
[0323] In some embodiments, since the PP is used as the PCF, the position and orientation of the user device can be determined relative to the most recent persistent pose or, in this example, the PCF. Such determination can be made by identifying feature points characterizing the PP in the current image captured by sensors on the device. Using image processing techniques such as stereo image analysis, the position of the device relative to those feature points can be determined. Based on this data, the system can calculate the change in the transformation associated with the user movement based on the relationship rig1_T_pcf1 = (rig1_T_w1)*(pcf1_T_w1).
[0324] The system can determine and apply the transformation in a computationally efficient order. For example, by tracking the user's pose and defining the position of the virtual content relative to the PP or PCF built based on the persistent pose, the need to calculate rig1_T_w1 from the measurements that produce rig1_T_pcf1 can be avoided. Thus, the transformation from the source coordinate system of the virtual content to the destination coordinate system of the user device can be based on the measured transformation according to the expression (rig1_T_pcf1)*(obj1_t_pcf1), where the first transformation is measured by the system and the latter transformation is provided by the application that specifies the rendering of the virtual content. In embodiments where the virtual content is positioned relative to the origin of the map, the end-to-end transformation can relate the virtual object coordinate system to the PCF coordinate system based on a further transformation between the map coordinates and the PCF coordinates. In embodiments where the virtual content is positioned relative to a PP or PCF different from the PP or PCF for which the user's position is tracked, a transformation can be made between the two. Such a transformation can be fixed and can be determined, for example, from a map where both appear.
[0325] For example, a transformation-based method can be implemented in a device having components that process sensor data to build a tracking map. As part of this process, these components can identify feature points that can be used as persistent poses, which in turn can become the PCF. These components can limit the number of persistent poses generated for the map to provide an appropriate spacing between the persistent poses while allowing the user to be close enough to the persistent pose locations regardless of their position in the physical environment to accurately calculate the user's pose, as described above in conjunction with Figures 17 to 19 shown. As the closest persistent pose to the user is updated due to user movement, the refinement of the tracking map or otherwise allows any transformation for calculating the position of the virtual content relative to the user, which depends on the PP (or PCF, if being used), to be updated and stored for use, at least until the user leaves that persistent pose. Nevertheless, by calculating and storing the transformation, the computational burden each time the position of the virtual content is updated can be relatively low, such that it can be performed with a relatively low latency.
[0326] FIG. 20A to FIG. 20C Shows positioning relative to a tracking map, and each device has its own tracking map. However, a transformation can be generated relative to any map coordinate system. Content persistence between user sessions of an XR system can be achieved by using a persistent map. A shared experience for users can also be achieved by using a map to which multiple user devices can be oriented.
[0327] In some embodiments described in more detail below, the position of virtual content can be specified relative to coordinates in a canonical map, the format of the canonical map being set such that any one of a plurality of devices can use the map. Each device may maintain a tracking map and can determine changes in the user's pose relative to the tracking map. In this example, the transformation between the tracking map and the canonical map can be determined by a "localization" process, which can be performed by matching structures in the tracking map (such as one or more persistent poses) to one or more structures in the canonical map (such as one or more PCFs).
[0328] Techniques for creating and using canonical maps in this manner are described in more detail below.
[0329] Depth Keyframes
[0330] The techniques described herein rely on the comparison of image frames. For example, to establish the position of a device relative to a tracking map, new images can be captured using sensors worn by the user, and the XR system can search an image set used to create the tracking map for images that share at least a predetermined number of points of interest with the new images. As an example of another scenario involving image frame comparison, the tracking map can be localized to the canonical map by first looking for an image frame in the tracking map associated with a persistent pose that is similar to an image frame in the canonical map associated with a PCF. Alternatively, the transformation between two canonical maps can be calculated by first looking for similar image frames in the two maps.
[0331] The techniques described herein can enable an efficient comparison of spatial information. In some embodiments, an image frame can be represented by a digital descriptor. The descriptor can be calculated by mapping a set of features identified in the image to a transformation of the descriptor. The transformation can be performed in a trained neural network. In some embodiments, the set of features provided as input to the neural network can be a filtered set of features extracted from the image using techniques that preferentially select features that are likely to be persistent.
[0332] Representing an image frame as a descriptor enables, for example, efficient matching of new image information with stored image information. The XR system can store with the persistent map descriptors of one or more frames underlying the persistent map. A local image frame acquired by a user device can similarly be converted to such a descriptor. By selecting a stored map having a descriptor similar to that of the local image frame, one or more persistent maps that may represent the same physical space as the user device can be selected with relatively little processing. In some embodiments, descriptors can be computed for key frames in the local map and the persistent map, further reducing the processing when comparing maps. For example, this efficient comparison can be used to simplify finding a persistent map to load into a local device or finding a persistent map to update based on image information acquired using the local device.
[0333] Depth key frames provide a way to reduce the processing amount required to identify similar image frames. For example, in some embodiments, the comparison can be between image features (e.g., "2D features") in a new 2D image and 3D features in the map. This comparison can be performed in any suitable way, such as by projecting the 3D image onto a 2D plane. Conventional methods such as Bag of Words (BoW) search for the 2D features of a new image in a database that includes all 2D features in the map, which may require a large amount of computing resources, especially when the map represents a large area. The conventional method then locates images that share at least one 2D feature with the new image, and these images may include images that are not useful for locating meaningful 3D features in the map. The conventional method then locates 3D features that are not meaningful relative to the 2D features in the new image.
[0334] The inventors have recognized and understood techniques for retrieving images in a map using fewer memory resources (e.g., one quarter of the memory resources used by BoW), higher efficiency (e.g., 2.5 ms of processing time per key frame, 100 μs for comparing 500 key frames), and higher accuracy (e.g., for a 1024 - dimensional model, the retrieval recall rate is 20% better than BoW, and for a 256 - dimensional model, the retrieval recall rate is 5% better than BoW).
[0335] To reduce computation, a descriptor can be computed for an image frame, which can be used to compare the image frame with other image frames. The descriptor can be stored instead of, or in addition to, the image frame and feature points. In a map in which persistent poses and / or PCFs can be generated based on image frames, the descriptors of one or more image frames according to which each persistent pose or PCF is generated can be stored as part of the persistent pose and / or PCF.
[0336] In some embodiments, descriptors can be calculated based on feature points in an image frame. In some embodiments, a neural network is configured to calculate a unique frame descriptor representing an image. The image can have a resolution higher than 1 megabyte, thereby capturing sufficient details of the 3D environment within the field of view of the device worn by the user in the image. The frame descriptor can be much shorter, such as a string of numbers, for example, in the range of 128 bytes to 512 bytes or any number of strings of numbers in between.
[0337] In some embodiments, the neural network is trained such that the calculated frame descriptor indicates the similarity between images. The image in the map can be located by identifying, in a database including images for generating a map, the nearest image that can have a frame descriptor within a predetermined distance from the frame descriptor of the new image. In some embodiments, the distance between images can be represented by the difference between the frame descriptors of the two images.
[0338] Fig.21 is a block diagram showing a system for generating a descriptor for a single image according to some embodiments. In the example shown, a frame embedding generator 308 is shown. In some embodiments, the frame embedding generator 308 can be used within the server 20, but alternatively or additionally can be executed in whole or in part in one of the XR devices 12.1 and 12.2 or any other device that processes images for comparison with other images.
[0339] In some embodiments, the frame embedding generator can be configured to generate a reduced data representation of an image from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes), which reduced data representation still indicates the content in the image despite the reduced size. In some embodiments, the frame embedding generator can be used to generate a data representation of an image, which image can be a key frame or otherwise a frame used. In some embodiments, the frame embedding generator 308 can be configured to convert an image located at a specific position and orientation into a unique string of numbers (e.g., 256 bytes). In the example shown, the image 320 captured by the XR device can be processed by a feature extractor 324 to detect the point of interest 322 in the image 320. The point of interest can or can not be derived from the feature points identified as described above for the feature 1120 ( Fig.14 ) or as otherwise described herein. In some embodiments, the point of interest can be represented by a descriptor as described above for the descriptor 1130 ( Fig.14 ), and these points of interest can be generated using a depth sparse feature method. In some embodiments, each point of interest 322 can be represented by a string of numbers (e.g., 32 bytes). For example, there can be n features (e.g., 100), and each feature is represented by a string of 32 bytes.
[0340] In some embodiments, the frame embedding generator 308 may include a neural network 326. The neural network 326 may include a multi-layer perceptron unit 312 and a max pooling unit 314. In some embodiments, the multi-layer perceptron (MLP) unit 312 may include a multi-layer perceptron, and the multi-layer perceptron may be trained. In some embodiments, the focus points 322 (e.g., descriptors for the focus points) may be reduced by the multi-layer perceptron 312 and output as a weighted combination 310 of descriptors. For example, the MLP may reduce n features to m features less than n features.
[0341] In some embodiments, the MLP unit 312 may be configured to perform matrix multiplication. The multi-layer perceptron unit 312 receives multiple focus points 322 of the image 320 and converts each focus point into a corresponding string of numbers (e.g., 256). For example, there may be 100 features, and each feature may be represented by a string of 256 numbers. In this example, a matrix with 100 horizontal rows and 256 vertical columns may be created. Each row may have a series of 256 numbers, and this series of 256 numbers varies in size, some being smaller and some being larger. In some embodiments, the output of the MLP may be an n×256 matrix, where n represents the number of focus points extracted from the image. In some embodiments, the output of the MLP may be an m×256 matrix, where m is the number of focus points reduced from n.
[0342] In some embodiments, the MLP 312 may have a training phase and a usage phase, during which the model parameters for the MLP are determined. In some embodiments, the MLP may be trained as shown in Fig.25 The input training data may include triplets of data, and a triplet includes 1) a query image, 2) a positive sample, and 3) a negative sample. The query image may be considered a reference image.
[0343] In some embodiments, the positive sample may include an image similar to the query image. For example, in some embodiments, being similar may mean having the same object in the query image and the positive sample image, but viewed from different angles. In some embodiments, being similar may mean having the same object in the query image and the positive sample image, but the object is shifted relative to the other image (e.g., left, right, up, down).
[0344] In some embodiments, negative samples may include images that are not similar to the query image. For example, in some embodiments, a dissimilar image may not contain any object that stands out in the query image, or may contain only a small portion (e.g., <10%, 1%) of the prominent objects in the query image. In contrast, for example, a similar image may have a majority (e.g., >50% or >75%) of the objects in the query image.
[0345] In some embodiments, regions of interest can be extracted from the images in the input training data and the regions of interest can be converted into feature descriptors. These descriptors can be calculated for both the training images as shown in Fig.25 and the features extracted from the operation of the frame embedding generator 308 for Fig.21 . In some embodiments, as described in U.S. Patent Application Publication No. 2019 / 0147341 entitled "Fully Convolutional Interest Point Detection and Description via Homographic Adaptation", deep sparse feature (DSF) processing can be used to generate descriptors (e.g., DSF descriptors). In some embodiments, the DSF descriptor is n×32 dimensional. The descriptor can then be passed through a model / MLP to create a 256-byte output. In some embodiments, the model / MLP can have the same structure as the MLP 312 such that once the model parameters are set through training, the resulting trained MLP can be used as the MLP 312.
[0346] In some embodiments, the feature descriptor (e.g., the 256 bytes output from the MLP model) can then be sent to a triplet margin loss module (which can be used only during the training phase and not during the usage phase of the MLP neural network). In some embodiments, the triplet margin loss module can be configured to select the parameters of the model to reduce the difference between the 256-byte output from the query image and the 256-byte output from the positive sample, and increase the difference between the 256-byte output from the query image and the 256-byte output from the negative sample. In some embodiments, the training phase can include feeding multiple triplet input images into the learning process to determine the model parameters. This training process can continue, for example, until the difference for the positive images is minimized and the difference for the negative images is maximized, or until other suitable exit criteria are met.
[0347] Referring again to Fig.21, the frame embedding generator 308 may include a pooling layer, shown here as a max pooling unit 314. The max pooling unit 314 may analyze each column to determine the maximum number in the corresponding column. The max pooling unit 314 may combine the maximum values of the numbers in each column of the output matrix of the MLP 312 into a global feature string 316 of, for example, 256 numbers. It should be understood that images processed in an XR system may be expected to have high-resolution frames, potentially having millions of pixels. The global feature string 316 is a relatively small number, which occupies relatively little memory and is easy to search compared to an image (e.g., having a resolution higher than 1 megabyte). Thus, the image can be searched without analyzing each original frame from the camera, and it is also cheaper to store 256 bytes instead of the full frame.
[0348] Fig. 22 is a flowchart showing a method 2200 for calculating an image descriptor according to some embodiments. The method 2200 may start with receiving (act 2202) a plurality of images captured by an XR device worn by a user. In some embodiments, the method 2200 may include determining (act 2204) one or more key frames from the plurality of images. In some embodiments, act 2204 may be skipped and / or may instead occur after step 2210.
[0349] The method 2200 may include: identifying (act 2206) one or more points of interest in the plurality of images using an artificial neural network; and calculating (act 2208) feature descriptors for the respective points of interest. The method may include calculating (act 2210) a frame descriptor for each image, thereby representing the image at least in part based on the feature descriptors calculated for the points of interest identified in the image using an artificial neural network.
[0350] Fig.23FIG. 2300 is a flow chart illustrating a method 2300 for localization using image descriptors according to some embodiments. In this example, a new image frame depicting the current position of an XR device can be compared to image frames stored in association with points in a map (e.g., persistent poses or PCFs as described above). Method 2300 can begin with receiving (act 2302) a new image captured by an XR device worn by a user. Method 2300 can include identifying (act 2304) one or more recent key frames in a database that includes key frames for generating one or more maps. In some embodiments, the recent key frames can be identified based on rough spatial information and / or previously determined spatial information. For example, the rough spatial information can indicate that the XR device is located in a geographic region represented by a 50m × 50m area of the map. Image matching can be performed only on points within this region. As another example, based on tracking, the XR system can know that the XR device was previously near a first persistent pose in the map and was moving in the direction of a second persistent pose in the map at that time. This second persistent pose can be considered the most recent persistent pose, and the key frame stored with it can be considered the most recent key frame. Alternatively or additionally, other metadata such as GPS data or WiFi fingerprints can be used to select the most recent key frame or a set of the most recent key frames.
[0351] Regardless of how the most recent key frames are selected, frame descriptors can be used to determine whether the new image matches any of the frames selected as being associated with nearby persistent poses. This determination can be made by comparing the frame descriptor of the new image to the frame descriptors of the most recent key frames or a subset of key frames in the database selected in any other suitable manner, and selecting the key frames that have frame descriptors within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors can be calculated by obtaining the difference between two numeric strings that can represent the two frame descriptors. In embodiments where the strings are processed as a plurality of strings, the difference can be calculated as a vector difference.
[0352] Once a matching image frame is identified, the orientation of the XR device relative to that image frame can be determined. Method 2300 can include performing (act 2306) feature matching on 3D features in the map corresponding to the identified most recent key frame, and calculating (act 2308) the pose of the device worn by the user based on the feature matching results. In this way, computationally intensive matching of feature points in two images can be performed for as few as one image that has been determined to be a possible match for the new image.
[0353] Fig.24is a flowchart showing a method 2400 for training a neural network according to some embodiments. The method 2400 may start by generating (act 2402) a data set including a plurality of image sets. Each of the plurality of image sets may include a query image, a positive sample image, and a negative sample image. In some embodiments, the plurality of image sets may include synthetic record pairs configured to teach the neural network basic information (such as shape), for example. In some embodiments, the plurality of image sets may include real record pairs, which may be recorded according to the physical world.
[0354] In some embodiments, inliers may be calculated by fitting an essential matrix between two images. In some embodiments, the sparse overlap may be calculated as the intersection over union (IoU) of the points of interest seen in two images. In some embodiments, the positive sample may include at least twenty points of interest that are the same as those in the query image as inliers. The negative sample may include fewer than ten inliers. The negative sample may have fewer than half of the sparse points that overlap with the sparse points of the query image.
[0355] The method 2400 may include calculating (act 2404) a loss for each image set by comparing the query image with the positive sample image and the negative sample image. The method 2400 may include modifying (act 2406) the artificial neural network based on the calculated loss such that the distance between the frame descriptor generated by the artificial neural network for the query image and the frame descriptor for the positive sample image is less than the distance between the frame descriptor for the query image and the frame descriptor for the negative sample image.
[0356] It should be understood that although the methods and apparatuses for generating global descriptors for individual images are described above, the methods and apparatuses may be configured to generate descriptors for individual maps. For example, a map may include a plurality of key frames, and each key frame may have a frame descriptor as described above. A max pooling unit may analyze the frame descriptors of the key frames of the map and combine the frame descriptors into a unique map descriptor for the map.
[0357] In addition, it should be understood that other architectures may be used for the processing as described above. For example, a separate neural network for generating DSF descriptors and frame descriptors is described. This method is computationally efficient. However, in some embodiments, frame descriptors may be generated based on the selected feature points without first generating DSF descriptors.
[0358] Ranking and merging maps
[0359] This document describes methods and apparatus for ranking multiple environmental maps in a cross-reality (XR) system, as a precursor to further processing, such as merging maps or positioning a device relative to a map. Map merging can enable maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking the maps can enable efficient execution of the techniques described herein, including map merging or positioning, both of which can involve selecting maps from a set based on similarity. In some embodiments, for example, the system can maintain a set of canonical maps that are formatted in a way that any of a number of XR devices can access them. These canonical maps can be formed by merging selected tracking maps from those devices with other tracking maps or previously stored canonical maps.
[0360] For example, the canonical maps can be ranked for selecting one or more canonical maps to merge with a new tracking map and / or for selecting one or more canonical maps from the set for use within a device. The ranking can indicate canonical maps having regions similar to regions of the tracking map, such that when attempting to merge the tracking map into a canonical map, there is likely to be a correspondence between a portion of the tracking map and a portion of the canonical map. Such a correspondence can be a precursor to further processing in which the tracking map is aligned with the canonical map so that it can be merged into the canonical map.
[0361] In some embodiments, the tracking map can be merged with tiles or other defined regions of the canonical map in order to limit the amount of processing required to merge the tracking map with the canonical map. Operating on regions of the canonical map, particularly for canonical maps representing relatively large areas, such as multiple rooms in a building, may require less processing than operating on the entire map. In those embodiments, map ranking or other selection preprocessing can identify one or more regions of a previously stored canonical map to attempt to merge.
[0362] For positioning, information about the location of a portable device is compared with information in the canonical map. If a set of candidate regions in a set of canonical maps is identified by ranking the canonical maps in the set, then the corresponding regions between the information from the local device and the canonical map can be identified more quickly.
[0363] The inventors have recognized and understood that XR systems can provide an enhanced XR experience to multiple users sharing the same world that includes real and / or virtual content by enabling the effective sharing of environmental maps of the real / physical world collected by multiple users, regardless of whether these users are present in the world at the same time or at different times. However, significant challenges exist in providing such a system. Such a system can store multiple maps generated by multiple users and / or the system can store multiple maps generated at different times. For operations that may be performed using previously generated maps, such as, for example, localization or map merging, a large amount of processing may be required to identify relevant environmental maps of the same world (e.g., the same real-world location) from all the environmental maps collected by the XR system.
[0364] In some systems, there may be a large number of environmental maps from which a selection of maps is made. The inventors have recognized and understood techniques for quickly and accurately ranking the relevance of environmental maps from all possible environmental maps, such as, for example Fig.28 the universe of all canonical maps 120 in. A high-ranked map or a small group of high-ranked maps can then be selected for further processing. This further processing can include determining the position of the user device relative to the selected map, or rendering virtual objects on the user display, interacting realistically with the physical world around the user, or merging the map data collected by the user device with the stored map to create a larger or more accurate map.
[0365] In some embodiments, stored maps relevant to the task of a user at a location in the physical world can be identified by filtering the stored maps based on multiple criteria. These criteria can indicate a comparison of a tracking map or other location information generated by the user's wearable device at that location with candidate environmental maps stored in a database. The comparison can be performed based on metadata associated with the map, such as Wi-Fi fingerprints detected by the device that generated the map and / or a set of BSSIDs to which the device was connected while the map was being formed. Geolocation information, such as GPS data, can alternatively or additionally be used to select one or more candidate environmental maps. Prior information can also be used for the selection of candidate maps. For example, if the position of the device has recently been determined relative to a canonical map, the canonical maps representing the previously determined position and adjacent positions can be selected.
[0366] Comparisons can be performed based on compressed or uncompressed content of a map. Comparisons based on a compressed representation can be performed by comparing vectors calculated from the map content. For example, comparisons based on an uncompressed representation can be performed by locating a tracking map within a stored map, and vice versa. For example, locating can be performed by matching one or more sets of features representing a portion of one map to a set of features in another map. Multiple comparisons can be performed in sequence based on the computational time required to reduce the number of candidate maps to be considered, where comparisons involving less computation will be performed earlier in sequence compared to other comparisons that require more computation.
[0367] In some embodiments, the comparison can be based on a portion of a tracking map, and the portion tracked can vary depending on the process for which an environmental map is selected. For example, for locating, candidate maps can be selected based on persistent poses in the tracking map that are closest to the user when the locating is to be performed. Conversely, when selecting candidate maps for a map merging operation, candidate maps can be selected based on the similarity between a canonical map and any one or more persistent poses in the tracking map. Regardless of what portion of the two maps is being compared, finding a matching portion can require a substantial process associated with finding sets of features in the maps that may correspond to the same location in the physical world.
[0368] Fig.26 An AR system 800 configured to rank one or more environmental maps according to some embodiments is depicted. In this example, the ranking is performed as a precursor to a merging operation. The AR system can include a traversable world model 802 of the AR device. Information populating the traversable world model 802 can come from sensors on the AR device, which can include computer-executable instructions stored in a processor 804 (e.g., Figure 4 the local data processing module 570 therein), which can perform some or all of the processing to convert sensor data into a map. Such a map can be a tracking map, as a tracking map can be built while collecting sensor data as the AR device operates in an area. Along with the tracking map, Figure 1 area attributes can be provided that can be formatted as metadata of the map to indicate the area represented by the tracking map. These area attributes can be geographical location identifiers, such as coordinates represented as latitude and longitude, or an ID used by the AR system to represent a location. Alternatively or additionally, the area attributes can be measured characteristics that are highly likely to be unique to the area. Area attributes can be derived, for example, from parameters of a wireless network detected in the area. In some embodiments, the area attributes can be associated with the unique address of an access point that the AR system is near and / or connected to. For example, the area attributes can be associated with the MAC address or basic service set identifier (BSSID) of a 5G base station / router, Wi-Fi router, etc.
[0369] Attributes can be attached to the map itself or to the locations represented in the map. For example, attributes for tracking the map can be associated with persistent poses in the map. For a canonical map, attributes can be associated with the persistent coordinate system or with tiles or other regions of the map.
[0370] In Fig.26 the example of, a tracking map can be merged with other maps of the environment. The map ranking section 806 receives the tracking map from the device PW 802 and communicates with the map database 808 to select and rank the environmental maps from the map database 808. The selected maps with higher rankings are sent to the map merging section 810.
[0371] The map merging section 810 can perform a merging process on the maps sent from the map ranking section 806. The merging process may involve merging the tracking map with some or all of the ranked maps and sending the new merged map to the traversable world model 812. The map merging section can merge maps by identifying maps that depict overlapping portions of the physical world. Those overlapping portions can be aligned so that the information in the two maps can be aggregated into a final map.
[0372] In some embodiments, the overlapping portions can be identified by first identifying a set of features in the canonical map that may correspond to a set of features associated with a persistent pose in the tracking map. The set of features in the canonical map can be a set of features that define the persistent coordinate system. However, it is not required that the set of features be associated with a single persistent location in the canonical map or the tracking map. In some embodiments, the process can be simplified by identifying correspondences between sets of features to which attributes have been assigned such that the values of those attributes can be compared to select sets of features for further processing.
[0373] Once a set of candidate maps, or in some embodiments, candidate regions of the map, have been identified, further steps can reduce the candidate set and can ultimately result in the identification of a matching map. One such step can be to compute a transformation of a set of feature points in the tracking map to align with a candidate set of features in the canonical map. The transformation can be applied to the entire tracking map such that the overlapping portions of the tracking map and the canonical map can be identified. A correspondence above a threshold between the features in the overlapping portions can indicate a strong likelihood that the tracking map has been registered with the canonical map. If the correspondence is below the threshold, other candidate sets of features in the tracking map and / or the canonical map can be processed to attempt to find a matching region of the map. When a large enough correspondence is found, the tracking map can be considered to be positioned relative to the canonical map, where the transformation that results in the large correspondence indicates the position of the tracking map relative to the canonical map.
[0374] Similar processing can be performed to merge a canonical map with other canonical maps and a tracking map. The merging of canonical maps can be performed from time to time to identify scenarios where a canonical map has been extended through repeated merging of tracking maps, such that multiple maps represent overlapping regions of the physical world. Regardless of the types of maps being merged, when there is an overlap between the maps being merged, the merging may require aggregating data from both maps. Aggregation may require extending one map with information from the other map. Alternatively or additionally, aggregation may require adjusting the representation of the physical world in one map based on information in the other map. For example, the later map may reveal an object that has caused a feature point to move, and thus the map can be updated based on the later information. Alternatively, two maps may characterize the same region with different feature points, and aggregation may require selecting a set of feature points from both maps to better represent the region.
[0375] Regardless of the specific processing that occurs during the merging process, in some embodiments, the persistent positions (e.g., persistent poses or PCFs) from all maps being merged can be retained such that applications that position content relative to them can continue to do so. In some embodiments, the merging of maps can result in redundant persistent poses, and some persistent poses can be deleted. When a PCF is associated with a persistent pose to be deleted, merging the maps may require modifying the PCF to be associated with the persistent poses that remain in the map after merging.
[0376] In some embodiments, as maps are extended and / or updated, they can be refined. Refinement may require performing calculations to reduce internal inconsistencies between feature points that may represent the same object in the physical world. Such inconsistencies may arise from inaccurate poses associated with keyframes that provide feature points representing the same object in the physical world. For example, such inconsistencies may arise when an XR device calculates a pose relative to a tracking map that is in turn based on an estimated pose, and thus errors in the pose estimation accumulate, resulting in a "drift" in pose accuracy over time. The map can be refined by performing bundle adjustment or other operations to reduce the inconsistencies of feature points from multiple keyframes.
[0377] During refinement, the position of a persistent point relative to the map origin can change. Accordingly, the transforms associated with that persistent point, such as a persistent pose or PCF, may change. In some embodiments, an XR system that incorporates map refinement (either as part of a merging operation or performed for other reasons) can recalculate the transforms associated with any persistent points that have changed. These transforms may be pushed from the component that calculates the transform to the component that uses the transform so that any use of the transform can be based on the updated position of the persistent point.
[0378] Merging may also involve assigning location metadata to locations in the merged map based on attributes associated with the locations in the maps being merged. For example, the attributes of a canonical map can be derived from the attributes of one or more tracking maps of the regions used to form the canonical map. In scenarios where the persistent coordinate system of the canonical map is defined based on the persistent poses in the tracking maps merged into the canonical map, the same attributes as the persistent poses can be assigned to the persistent coordinate system. In cases where the persistent coordinate system is derived from one or more persistent poses from one or more tracking maps, the attributes assigned to the persistent coordinate system can be an aggregation of the attributes of the persistent poses. The way of aggregating attributes may vary based on the type of the attributes. Geographical location attributes can be aggregated by interpolating the attribute values of the location of the persistent coordinate system according to the geographical location information associated with the persistent pose. For other attributes, the values can be aggregated by considering all of them applied to the persistent coordinate system. Similar logic can be applied to derive the attributes of tiles or other regions of the map.
[0379] In scenarios where the tracking map submitted for merging has insufficient relevance to any previously stored canonical map, the system can create a new canonical map based on the tracking map. In some embodiments, the tracking map can be promoted to a canonical map. In this way, the area of the physical world represented by the canonical map set can grow as the user device submits tracking maps for merging.
[0380] The traversable world model 812 can be a cloud model that can be shared by multiple AR devices. The traversable world model 812 can store or otherwise access the environmental map in the map database 808. In some embodiments, when updating a previously computed environmental map, the previous version of the map can be deleted to remove the outdated map from the database. In some embodiments, when updating a previously computed environmental map, the previous version of the map can be archived, enabling access / viewing of the previous version of the environment. In some embodiments, permissions can be set such that only AR systems with certain read / write access rights can trigger the deletion / archiving of the previous version of the map.
[0381] These environmental maps created from the tracking maps provided by one or more AR devices / systems can be accessed by the AR devices in the AR system. The map ranking section 806 can also be used to provide environmental maps to the AR devices. The AR device can send a message requesting the environmental map of its current location, and the map ranking section 806 can be used to select and rank the environmental maps relevant to the requesting device.
[0382] In some embodiments, the AR system 800 may include a downsampling unit 814 configured to receive a merged map from the cloud PW 812. The merged map received from the cloud PW 812 may be in a storage format for the cloud, which may include high-resolution information, such as a large number of PCFs per square meter, or multiple image frames, or a large set of feature points associated with the PCFs. The downsampling unit 814 may be configured to downsample the cloud-format map into a format suitable for storage on the AR device. The device-format map may contain less data, such as fewer PCFs or less data stored for each PCF, to accommodate the limited local computing power and storage space of the AR device.
[0383] Fig. 27 is a simplified block diagram showing a plurality of canonical maps 120 that may be stored in a remote storage medium such as the cloud. Each canonical map 120 may include a plurality of attributes that serve as canonical map identifiers, which indicate the location of the canonical map within the physical space, such as somewhere on the Earth. These canonical map identifiers may include one or more of the following identifiers: a region identifier represented by a longitude and latitude range, a frame descriptor (e.g., Fig.21 the global feature string 316 in), a Wi-Fi fingerprint, a feature descriptor (e.g., Fig.21 the feature descriptor 310 in), and a device identifier indicating one or more devices that contributed to the map. These attributes can be used to define which maps and / or which parts or maps to compare when searching for correspondences between maps.
[0384] Although Fig. 27 each of the elements in may be considered a separate map, in some embodiments, processing may be performed on tiles or other regions of the canonical map. For example, the canonical maps 120 may each represent a region of a larger map. In those embodiments, attributes may be assigned to regions instead of or in addition to the entire map. Thus, the discussion of map selection herein may alternatively or additionally apply to selecting segments of a map, such as one or more tiles.
[0385] In the example shown, the canonical maps 120 are geographically arranged in a two-dimensional pattern because they may exist on the Earth's surface. The canonical maps 120 may be uniquely identifiable by their respective longitudes and latitudes because any canonical maps having overlapping longitudes and latitudes may be merged into a new canonical map.
[0386] Fig.28FIG. is a schematic diagram showing a method of selecting a canonical map according to some embodiments, which can be used to position a new tracking map to one or more canonical maps. Such a process can be used to identify candidate maps for merging or can be used to position a user device relative to a canonical map. The method can start by accessing (action 120) the world of the canonical map 120, which, as an example, can be stored in a database of traversable worlds (e.g., traversable world module 538). The world of the canonical map can include canonical maps from all previously visited locations. The XR system can filter the world of all canonical maps into a small subset or just one map. The selected set can be sent to the user device for further processing on the device. Alternatively or additionally, some or all of the further processing can be performed in the cloud, and the results of the processing can be sent to the user device and / or stored in the cloud, as in the case of merged maps.
[0387] The method can include filtering (action 300) the world of the canonical map. In some embodiments, action 300 can select at least one matching canonical map 120 that covers a longitude and a latitude, where the longitude and latitude include the longitude and latitude of a location identifier received from the XR device, provided that there is at least one map at that longitude and latitude. In some embodiments, action 300 can select at least one adjacent canonical map that covers a longitude and a latitude adjacent to the matching canonical map. In some embodiments, action 300 can select multiple matching canonical maps and multiple adjacent canonical maps. Action 300 can, for example, reduce the number of canonical maps by approximately a factor of ten, e.g., from thousands to hundreds, to form a first filtered selection. Alternatively or additionally, criteria other than latitude and longitude can be used to identify adjacent maps. For example, the XR device may have previously been positioned using a canonical map in the set as part of the same session. The cloud service can retain information about the XR device, including the previously positioned map. In this example, the maps selected at action 300 can include those that cover areas adjacent to the map to which the XR device was positioned.
[0388] The method may include a first filtering selection of the canonical map based on Wi-Fi fingerprints (action 302). Action 302 may determine latitude and longitude based on the Wi-Fi fingerprints received from the XR device as part of a location identifier. Action 302 may compare the latitude and longitude from the Wi-Fi fingerprints with the latitude and longitude of the canonical map 120 to determine one or more canonical maps that form a second filtering selection. Action 302 may reduce the number of canonical maps by approximately ten times, e.g., from hundreds of canonical maps to dozens (e.g., 50) of canonical maps that form the second selection. For example, the first filtering selection may include 130 canonical maps, the second filtering selection may include 50 out of the 130 canonical maps, and may exclude the other 80 out of the 130 canonical maps.
[0389] The method may include a second filtering selection of the canonical map based on key frames (action 304). Action 304 may compare data representing an image captured by the XR device with data representing the canonical map 120. In some embodiments, the data representing the image and / or the map may include feature descriptors (e.g., Fig.25 the DSF descriptor in Fig.21 and / or global feature strings (e.g.,
[0390] the 316 in Fig. 27 ). Action 304 may provide a third filtering selection of the canonical map. In some embodiments, for example, the output of action 304 may be only five canonical maps out of the 50 canonical maps identified after the second filtering selection. Then, the map transmitter 122 sends one or more canonical maps based on the third filtering selection to the viewing device. Action 304 may reduce the number of canonical maps by approximately ten times, e.g., from dozens of canonical maps to a single-digit number of canonical maps (e.g., 5) that form the third selection. In some embodiments, the XR device may receive the canonical maps in the third filtering selection and attempt to locate to the received canonical maps.
[0391] In some embodiments, the cloud can receive the feature details of the live / new / current image captured by the viewing device, and the cloud can generate the global feature string 316 of the live image. Then, the cloud can filter the canonical map 120 based on the live global feature string 316. In some embodiments, the global feature string can be generated on the local viewing device. In some embodiments, the global feature string can be remotely generated in the cloud, for example. In some embodiments, the cloud can send the filtered canonical map together with the global feature string 316 associated with the filtered canonical map to the XR device. In some embodiments, when the viewing device aligns its tracking map to the canonical map, it can do so by matching the global feature string 316 of the local tracking map with the global feature string of the canonical map.
[0392] It should be understood that the operations of the XR device may not perform all of the actions (300, 302, 304). For example, if the world of the canonical map is relatively small (e.g., 500 maps), the XR device attempting to align can filter the world of the canonical map based on Wi-Fi fingerprints (e.g., action 302) and key frames (e.g., action 304), but omits region-based filtering (e.g., action 300). Also, it is not necessary to compare the entire maps. For example, in some embodiments, the comparison of two maps can result in the identification of common persistent points, such as persistent poses or PCFs that appear in both the new map and the map selected from the map world. In that case, descriptors can be associated with the persistent points, and those descriptors can be compared.
[0393] Fig.29 is a flowchart showing a method 900 for selecting one or more ranked environmental maps according to some embodiments. In the illustrated embodiment, ranking is performed on the AR device of the user who is creating the tracking map. Thus, the tracking map can be used to rank the environmental maps. In embodiments where the tracking map is not available, some or all of the selection and ranking of environmental maps that do not explicitly rely on the tracking map can be used.
[0394] Method 900 can start with action 902, where a set of maps in a database of environmental maps (which can be formatted as canonical maps) located near the location where the tracking map is being formed can be accessed and then filtered for ranking. Additionally, at action 902, at least one regional attribute of the region in which the user's AR device is operating is determined. In the scenario where the user's AR device is constructing a tracking map, the regional attribute can correspond to the region on which the tracking map is being created. As a specific example, the regional attribute can be calculated based on the received signal from the access point to the computer network while the AR device is computing the tracking map.
[0395] Fig.30Depicts an exemplary map ranking section 806 of an AR system 800 according to some embodiments. The map ranking section 806 can be executed in a cloud computing environment as it can include a part executed on an AR device and a part executed on a remote computing system such as the cloud. The map ranking section 806 can be configured to execute at least a part of method 900.
[0396] Fig.31A Depicts an example of area attributes AA1 - AA8 of a tracking map (TM) 1102 and environmental maps CM1 - CM4 in a database according to some embodiments. As shown, the environmental maps can be associated with multiple area attributes. The area attributes AA1 - AA8 can include parameters of a wireless network detected by an AR device calculating the tracking map 1102, e.g., the basic service set identifier (BSSID) of the network to which the AR device is connected and / or the strength of the received signal from the access point of the wireless network through, e.g., a network tower 1104. The parameters of the wireless network can conform to protocols including Wi-Fi and 5G NR. In Fig.32 the example shown, the area attribute is a fingerprint of the area where the user's AR device collects sensor data to form the tracking map.
[0397] Fig.31B Depicts an example of a determined geographical location 1106 of a tracking map 1102 according to some embodiments. In the example shown, the determined geographical location 1106 includes a centroid point 1110 and an area 1108 surrounding the centroid point. It should be understood that the determination of the geographical location in this application is not limited to the format shown. The determined geographical location can have any suitable format, including, for example, different area shapes. In this example, using a database that associates area attributes with geographical locations, the geographical location is determined from the area attributes. The database is commercially available, e.g., a database that associates Wi-Fi fingerprints with locations expressed as latitude and longitude and can be used for this operation.
[0398] In Fig.29 embodiments, the map database containing the environmental maps can also include location data of those maps, including the latitude and longitude covered by the maps. The processing at action 902 may require selecting a set of environmental maps from the database that cover the same latitude and longitude determined for the area attributes of the tracking map.
[0399] Action 904 is a first filtering of the set of environmental maps accessed in action 902. In action 902, the environmental maps are retained in the set based on their proximity to the geographical location of the tracking map. This filtering step can be performed by comparing the latitudes and longitudes associated with the tracking map and the environmental maps in the set.
[0400] Fig.32Depicts an example of operation 904 according to some embodiments. Each region attribute may have a corresponding geographical location 1202. The set of environmental maps may include an environmental map having at least one region attribute having a geographical location that overlaps with the determined geographical location of the tracking map. In the example shown, the set of identified environmental maps includes environmental maps CM1, CM2, and CM4, each having at least one region attribute having a geographical location that overlaps with the determined geographical location of the tracking map 1102. CM3 associated with region attribute AA6 is not included in the set because it is outside the determined geographical location of the tracking map.
[0401] Other filtering steps may also be performed on the set of environmental maps to reduce / rank the number of environmental maps in the set that are ultimately processed (such as for map merging or providing traversable world information to a user device). Method 900 may include filtering (operation 906) the set of environmental maps based on the similarity of one or more identifiers of network access points associated with the tracking map and the environmental maps of the set. During the formation of a map, a device that collects sensor data to generate the map may be connected to a network through a network access point (such as via Wi-Fi or a similar wireless communication protocol). An access point may be identified by a BSSID. As a user device moves through the area where data is collected to form a map, the user device may connect to multiple different access points. Similarly, when multiple devices provide information to form a map, the devices may have connected through different access points, and for this reason, multiple access points may also be used when forming the map. Thus, there may be multiple access points associated with a map, and the set of access points may be an indication of the map location. The signal strength from an access point may be reflected as an RSSI value, which may provide further geographical information. In some embodiments, a list of BSSID and RSSI values may form a region attribute for the map.
[0402] In some embodiments, filtering the set of environmental maps based on the similarity of one or more identifiers of network access points may include: retaining in the set of environmental maps the environmental map having the highest Jaccard similarity with at least one region attribute of the tracking map based on the one or more identifiers of the network access point. Fig.33 Depicts an example of operation 906 according to some embodiments. In the example shown, the network identifier associated with region attribute AA7 may be determined as the identifier of the tracking map 1102. The set of environmental maps after operation 906 includes: environmental map CM2, which may have a region attribute within a higher Jaccard similarity with AA7; and environmental map CM4, which also includes region attribute AA7. Environmental map CM1 is not included in the set because it has the lowest Jaccard similarity with AA7.
[0403] The processing of actions 902 - 906 can be performed based on metadata associated with the map without actually accessing the content of the map stored in the map database. Other processing may involve accessing the content of the map. Action 908 indicates accessing the environmental map remaining in the subset after filtering based on the metadata. It should be understood that if subsequent operations can be performed on the accessed content, this action can be performed earlier or later in the process.
[0404] Method 900 may include filtering (action 910) a set of environmental maps based on the similarity of metrics representing the content of the tracking map and the environmental maps of a set of environmental maps. The metrics representing the content of the tracking map and the environmental maps may include a vector of values calculated from the content of the map. For example, as described above, the depth key - frame descriptors calculated for one or more key frames used to form the map can provide metrics for comparing the map or parts of the map. The metrics can be calculated from the maps obtained at action 908 or can be pre - calculated and stored as metadata associated with those maps. In some embodiments, filtering a set of environmental maps based on the similarity of metrics representing the content of the tracking map and the environmental maps of a set of environmental maps may include: retaining in the set of environmental maps the environmental map having the minimum vector distance between the feature vector of the tracking map and the vectors representing the environmental maps in the set of environmental maps.
[0405] Method 900 may include further filtering (action 912) a set of environmental maps based on the degree of match between a portion of the tracking map and portions of the environmental maps of a set of environmental maps. The degree of match can be determined as part of the localization process. As a non - limiting example, localization can be performed by identifying critical points in the tracking map and the environmental maps that are similar enough to represent the same part of the physical world. In some embodiments, the critical points can be features, feature descriptors, key frames, key assemblies, persistent poses, and / or PCFs. Then, a set of critical points in the tracking map may be aligned to produce an optimal fit with the set of critical points in the environmental map. The mean - square distance between the corresponding critical points may be calculated and, if below a threshold for a particular region of the tracking map, used as an indication that the tracking map and the environmental map represent the same region of the physical world.
[0406] In some embodiments, filtering a set of environmental maps based on the degree of match between a portion of the tracking map and portions of the environmental maps of a set of environmental maps may include: calculating the volume of the physical world represented by the tracking map, which is also represented in the environmental maps of the set of environmental maps; and retaining in the set of environmental maps the environmental maps having a larger calculated volume than the environmental maps filtered out from the set. Fig.34Depicts an example of action 912 according to some embodiments. In the example shown, a set of environmental maps after action 912 includes environmental map CM4, which has an area 1402 that matches an area of tracking map 1102. Environmental map CM1 is not included in the set because it does not have an area that matches an area of tracking map 1102.
[0407] In some embodiments, the set of environmental maps can be filtered in the order of action 906, action 910, and action 912. In some embodiments, the set of environmental maps can be filtered based on action 906, action 910, and action 912, and action 906, action 910, and action 912 can be performed according to the order from lowest to highest based on the processing required for filtering. Method 900 can include loading (action 914) the set of environmental maps and data.
[0408] In the example shown, the user database stores a region identifier indicating the region where the AR device is used. The region identifier can be a region attribute, which can include parameters of a wireless network detected by the AR device during use. The map database can store multiple environmental maps constructed from data provided by the AR device and associated metadata. The associated metadata can include a region identifier derived from the region identifier of the AR device that provided the data, and the environmental map is constructed from this data. The AR device can send a message to the PW module indicating that a new tracking map is being created or has been created. The PW module can calculate a region identifier for the AR device and update the user database based on the received parameters and / or the calculated region identifier. The PW module can also determine the region identifier associated with the AR device that requested the environmental map, identify the set of environmental maps from the map database based on the region identifier, filter the set of environmental maps, and send the filtered set of environmental maps to the AR device. In some embodiments, the PW module can filter the set of environmental maps based on one or more criteria, which can include, for example, the geographical location of the tracking map, the similarity of one or more identifiers of network access points associated with the tracking map and the environmental maps of the set, the similarity of a measure representing the content of the tracking map and the environmental maps of the set, and the degree of match between a part of the tracking map and a part of the environmental maps of the set.
[0409] Several aspects of some embodiments have been described so far. It should be understood that various changes, modifications, and improvements will readily occur to those skilled in the art. As an example, the embodiments are described in the context of an augmented (AR) environment. It should be understood that some or all of the techniques described herein can be applied in an MR environment or more generally in other XR environments and VR environments.
[0410] As another example, embodiments are described in the context of a device such as a wearable device. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), discrete applications, and / or devices, any suitable combination of a network and discrete applications.
[0411] In addition, Fig.29 Examples of criteria that can be used to filter candidate maps to produce a set of highly ranked maps are provided. Instead of or in addition to the described criteria, other criteria may be used. For example, if multiple candidate maps have similar values of a metric for filtering out less desirable maps, the characteristics of the candidate maps can be used to determine which maps are retained as candidates or filtered out. For example, larger or denser candidate maps may be prioritized over smaller candidate maps. In some embodiments, Figure 27-28 may describe Figure 29-34 all or part of the systems and methods described in
[0412] Fig.35 and 36 is a schematic diagram showing an XR system configured to rank and merge multiple environmental maps according to some embodiments. In some embodiments, the Passable World (PW) may determine when to trigger ranking and / or merging of maps. In some embodiments, determining which maps to use may be at least partially based on the depth keyframes described above with respect to Figure 21 to Figure 25 described.
[0413] Fig.37 is a block diagram showing a method 3700 for creating an environmental map of a physical world according to some embodiments. Method 3700 may localize a tracking map (act 3702) captured by an XR device worn by a user to a group of canonical maps (e.g., canonical maps selected by the method of Fig.28 and / or the method 900 of FIG. 900). Act 3702 may include localizing the key assemblies of the tracking map into the group of canonical maps. The localization result of each key assembly may include the localized pose of the key assembly and a set of 2D to 3D feature correspondences.
[0414] In some embodiments, method 3700 may include splitting (act 3704) the tracking map into connected parts, which may robustly merge the maps by merging the connected segments. Each connected part may include key assemblies within a predetermined distance. Method 3700 may include: merging (act 3706) connected parts larger than a predetermined threshold into one or more canonical maps; and removing the merged connected parts from the tracking map.
[0415] In some embodiments, method 3700 may include merging (action 3708) canonical maps in a group that are merged with the same connected portion of the tracking map. In some embodiments, method 3700 may include promoting (action 3710) the remaining connected portions of the tracking map that have not been merged with any canonical maps to canonical maps. In some embodiments, method 3700 may include merging (action 3712) the persistent poses and / or PCFs of the tracking map and the canonical maps, where the canonical maps are merged with at least one connected portion of the tracking map. In some embodiments, method 3700 may include finalizing (action 3714) the canonical map, for example, by fusing map points and pruning redundant key assemblies.
[0416] Fig.38A and 38B illustrates an environment map 3800 created by updating a canonical map 700 according to some embodiments. The canonical map 700 can be upgraded from a tracking map 700 ( Figure 7 ) with a new tracking map. As illustrated and described with respect to Figure 7 , the canonical map 700 can provide a floor plan 706 of reconstructed physical objects represented by points 702 in the corresponding physical world. In some embodiments, the map points 702 can represent features of a physical object, and the physical object can include multiple features. A new tracking map of the physical world can be captured and uploaded to the cloud for merging with the map 700. The new tracking map can include map points 3802 and key assemblies 3804, 3806. In the example shown, the key assembly 3804 represents a key assembly that has been successfully located in the canonical map by, for example, establishing a correspondence with the key assembly 704 of the map 700 (as Fig.38B shown). On the other hand, the key assembly 3806 represents a key assembly that has not been located in the map 700. In some embodiments, the key assembly 3806 can be promoted to a separate canonical map.
[0417] Figures 39A to 39F is a schematic diagram showing an example of a cloud-based persistent coordinate system that provides a shared experience for users in the same physical space. Fig.39A illustrates that a canonical map 4814 from the cloud, for example, is received by XR devices worn by users 4802A and 4802B of FIG. 20A to FIG. 20C . The canonical map 4814 can have a canonical coordinate frame 4806C. The canonical map 4814 can have a PCF 4810C with multiple associated PPs (e.g., 4818A, 4818B in Fig.39C ).
[0418] Fig.39BShows the relationship established by the XR device between its respective world coordinate systems 4806A, 4806B and the canonical coordinate frame 4806C. For example, this can be done by locating the canonical map 4814 on the respective device. For each device, aligning the tracking map to the canonical map can result in a transformation between its local world coordinate system and the coordinate system of the canonical map for each device.
[0419] Fig.39C Shows that the transformation (e.g., transformation 4816A, transformation 4816B) between the local PCF (e.g., PCF 4810A, PCF 4810B) on the respective device and the corresponding persistent pose (e.g., PP 4818A, PP 4818B) on the canonical map can be calculated as a result of the alignment. Using these transformations, each device can use its local PCF to determine where to display virtual content attached to PP 4818A, PP 4818B or other persistent points on the canonical map relative to the local device, where the local PCF can be detected locally on the device by processing the images detected by the sensors on the device. Such a method can accurately position virtual content relative to each user and can enable each user to have the same experience of virtual content in the physical space.
[0420] Fig.39D Shows a snapshot of the persistent pose from the canonical map to the local tracking map. It can be seen that the local tracking maps are interconnected by the persistent poses. Fig.39E Shows that the PCF 4810A on the device worn by user 4802A can be accessed through PP 4818A in the device worn by user 4802B. Fig.39F Shows that the tracking maps 4804A, 4804B and the canonical map 4814 can be merged. In some embodiments, some PCFs can be removed due to the merger. In the example shown, the merged map includes the PCF 4810C of the canonical map 4814, but does not include the PCF 4810A, PCF 4810B of the tracking maps 4804A, 4804B. After the map merger, the PPs previously associated with PCF 4810A, PCF 4810B can be associated with PCF 4810C.
[0421] Example
[0422] Fig.40 and Fig.41 Shows an example of using the tracking map by Fig. 9 the first XR device 12.1. Fig.40 is a two-dimensional representation of a three-dimensional first local tracking map (map Figure 1 ) according to some embodiments, which can be obtained by Fig. 9Generated by the first XR device. Fig.41 Is a block diagram showing, according to some embodiments, the upload from Fig. 9 The first XR device to the server Figure 1 Of the map.
[0423] Fig.40 Shows the map on the first XR device 12.1 Figure 1 And virtual content (Content 123 and Content 456). The map Figure 1 Has an origin (Origin 1). The map Figure 1 Includes many PCFs (PCF a to PCF d). From the perspective of the first XR device 12.1, PCF a is located, for example, at Figure 1 The origin of the map and has X, Y, and Z coordinates of (0, 0, 0), and PCF b has X, Y, and Z coordinates of (-1, 0, 0). Content 123 is associated with PCF a. In this example, Content 123 has X, Y, and Z relationships relative to PCF a of (1, 0, 0). Content 456 has a relationship relative to PCF b. In this example, Content 456 has X, Y, and Z relationships relative to PCF b of (1, 0, 0).
[0424] In Fig.41 The first XR device 12.1 uploads the map Figure 1 To the server 20. In this example, since the server does not store a canonical map for the same area of the physical world represented by the tracking map, and the tracking map is stored as an initial canonical map. The server 20 now has a canonical map based on the map Figure 1 The first XR device 12.1 has an empty canonical map at this stage. For the purpose of discussion, and in some embodiments, the server 20 does not include other maps except the map Figure 1 The second XR device 12.2 does not have a map stored on it.
[0425] The first XR device 12.1 also sends its Wi-Fi signature data to the server 20. The server 20 can use the Wi-Fi signature data to determine the approximate location of the first XR device 12.1 based on intelligence collected from other devices that have been connected to the server 20 or other servers in the past along with the GPS locations of such other recorded devices. The first XR device 12.1 can now end the first session (see Figure 8 ) and can disconnect from the server 20.
[0426] Fig.42 Is a block diagram showing, according to some embodiments, Fig.16Schematic diagram of an XR system, showing that after the first user 14.1 terminated the first session, the second user 14.2 has initiated a second session using the second XR device of the XR system. Fig.43A Block diagram showing the second user 14.2 initiating a second session. Since the first session of the first user 14.1 has ended, the first user 14.1 is shown as a dashed line. The second XR device 12.2 starts recording objects. The server 20 can use various systems with different granularities to determine that the second session of the second XR device 12.2 is in the same vicinity as the first session of the first XR device 12.1. For example, the first XR device 12.1 and the second XR device 12.2 can include Wi-Fi signature data, Global Positioning System (GPS) positioning data, GPS data based on Wi-Fi signature data, or any other data indicating location to record their positions. Alternatively, the PCF identified by the second XR device 12.2 can show the similarity with the Figure 1 PCF of the ground.
[0427] As Fig.43B shown, the second XR device starts and begins to collect data, such as images 1110 from one or more cameras 44, 46. As Fig.14 shown, in some embodiments, an XR device (e.g., the second XR device 12.2) can collect one or more images 1110 and perform image processing to extract one or more features / areas of interest 1120. Each feature can be converted into a descriptor 1130. In some embodiments, the descriptor 1130 can be used to describe a key frame 1140, which can have the position and orientation of additional associated images. One or more key frames 1140 can correspond to a single persistent pose 1150, which can be automatically generated after a threshold distance (e.g., 3 meters) from the previous persistent pose 1150. One or more persistent poses 1150 can correspond to a single PCF 1160, which can be automatically generated after a predetermined distance (e.g., every 5 meters). Over time, as the user continues to move around the user's environment and the XR device continues to collect more data (such as images 1110), additional PCFs (e.g., PCF 3 and PCF 4, 5) may be created. One or more applications 1180 can run on the XR device and provide virtual content 1170 to the XR device for presentation to the user. The virtual content can have an associated content coordinate frame, which can be placed relative to one or more PCFs. As Fig.43B shown, the second XR device 12.2 creates three PCFs. In some embodiments, the second XR device 12.2 can attempt to locate one or more canonical maps stored on the server 20.
[0428] In some embodiments, as Fig.43C shown in Figure 1 , the second XR device 12.2 may download the canonical map 120 from the server 20. The map on the second XR device 12.2
[0429] Fig.44 includes PCFs a to d and the origin 1. In some embodiments, the server 20 may have multiple canonical maps for respective locations and may determine that the second XR device 12.2 is located in the same vicinity as the first XR device 12.1 during a first session and send the canonical map of that vicinity to the second XR device 12.2. Figure 2 shows the second XR device 12.2 starting to identify PCFs for generating the map Figure 2 . The second XR device 12.2 has only identified a single PCF, namely PCF 1,2. The X, Y, and Z coordinates of PCF 1,2 of the second XR device 12.2 may be (1, 1, 1). The map Figure 2 has its own origin (origin 2), which may be based on the head pose of device 2 during the current head pose session at the start of the device. In some embodiments, the second XR device 12.2 may immediately attempt to align the map Figure 2 to the canonical map. In some embodiments, the map Figure 1 may not be aligned to the canonical map (i.e., the alignment may fail) because the system cannot identify any or sufficient overlap between the two maps. Alignment can be performed by identifying a portion of the physical world represented in the first map that is also represented in the second map and calculating the transformation between the first map and the second map required to align these portions. In some embodiments, the system may perform alignment based on a PCF comparison between the local map and the canonical map. In some embodiments, the system may perform alignment based on a persistent pose comparison between the local map and the canonical map. In some embodiments, the system may perform alignment based on a key frame comparison between the local map and the canonical map.
[0430] Fig.45 shows the map Figure 2 after the second XR device 12.2 has identified other PCFs of the map Figure 2 (PCF 1,2, PCF 3, PCF 4,5). The second XR device 12.2 attempts to align the map Figure 2 to the canonical map again. Since the map Figure 2 has been extended to overlap at least a portion of the canonical map, the alignment attempt will succeed. In some embodiments, the overlap between the local tracking map, the map Figure 2 and the canonical map may be represented by PCFs, persistent poses, key frames, or any other suitable intermediate or derived constructs.
[0431] In addition, the second XR device 12.2 has associated content 123 and content 456 with the PCFs 1, 2, and 3 of the ground Figure 2 The content 123 has X, Y, and Z coordinates (1, 0, 0) with respect to the PCFs 1 and 2 of the ground. Similarly, with respect to the PCF 3 in the ground Figure 2 the X, Y, and Z coordinates of the content 456 are (1, 0, 0).
[0432] Fig.46A and Fig.46B shows a successful positioning to the canonical map. The positioning can be based on matching features in one map with those in another map. Through appropriate transformations, involving translation and rotation of one map with respect to another here, the overlapping region / volume / cross-section of the map 1410 represents the common part of the ground Figure 2 and the canonical map. Since the PCFs 3 and 4, 5 were created before the positioning, while the canonical map had the PCFs a and c created before the ground Figure 1 was created, different PCFs were created to represent the same volume in the actual space (e.g., in different maps). Figure 2 In Figure 2 As shown, the second XR device 12.2 extends the ground
[0433] such as Figure 47 to include the PCFs a - d from the canonical map. The inclusion of the PCFs a - d represents a positioning to the canonical map. In some embodiments, the XR system can perform an optimization step to remove duplicate PCFs from the overlapping region, such as the PCFs in 1410, PCF 3, and PCFs 4, 5. After the ground Figure 2 is positioned, the placement of virtual content (such as content 456 and content 123) will be associated with the closest updated PCF in the updated ground Figure 2 . The virtual content appears at the same real-world position relative to the user, despite the change in the PCF attachment of the content and despite the update of the PCFs in the ground Figure 2 . Figure 2 As shown in Figure 2 the second XR device 12.2 continues to extend the ground
[0434] such as Figure 48 when the user walks around the real world, the second XR device 12.2 will identify other PCFs (PCFs e, f, g, and h). It should also be noted that the ground Figure 2 does not extend in Figure 1 in Figure 47 and Figure 48 .
[0435] Referring to Figure 49 the second XR device 12.2 will ground Figure 2Uploaded to server 20. Server 20 will ground Figure 2 Stored in accordance with the specification Figure 1 When the session for the second XR device 12.2 ends, it can be uploaded to server 20 Figure 2
[0436] The canonical map within server 20 now includes PCF i, which was not included in the ground on the first XR device 12.1 Figure 1 When a third XR device (not shown) uploads a map to server 20 and the map includes PCF i, the canonical map on server 20 may have been extended to include PCF i
[0437] In Figure 50 Server 20 combines the ground Figure 2 With the canonical map to form a new canonical map. Server 20 determines that PCF a to d are common to the canonical map and the ground Figure 2 The server extends the canonical map to include PCF e to h and PCF 1, 2 from the ground Figure 2 To form a new canonical map. The canonical maps on the first XR device 12.1 and the second XR device 12.2 are based on the ground Figure 1 And are outdated
[0438] In Figure 51 Server 20 sends the new canonical map to the first XR device 12.1 and the second XR device 12.2. In some embodiments, this may occur when the first XR device 12.1 and the second device 12.2 attempt to localize during a different or new or subsequent session. The first XR device 12.1 and the second XR device 12.2 proceed as described above to align their respective local maps (ground Figure 1 And ground Figure 2 Respectively) to the new canonical map
[0439] As Figure 52 Shown in, the head coordinate frame 96 or "head pose" is related to the PCF in the ground Figure 2 In some embodiments, the origin of the map, origin 2, is based on the head pose of the second XR device 12.2 at the start of the session. When creating the PCF during the session, the PCF is placed relative to the world coordinate frame origin 2. The PCF of the ground Figure 2 Is used as a persistent coordinate system relative to the canonical coordinate frame, where the world coordinate frame can be the world coordinate frame of the previous session (e.g., Figure 40 The origin 1 of the ground Figure 1 In). These coordinate frames are related by the same transformation used to align the ground Figure 2 To the canonical map, as described above in connection with Figure 46B as discussed
[0440] has been previously referenced Figure 9 the transformation from the world coordinate frame to the head coordinate frame 96 has been discussed. Figure 52 The head coordinate frame 96 shown in has only two orthogonal axes, which are at a specific coordinate position relative to the Figure 2 PCF of the ground, and at a specific angle relative to the Figure 2 ground. However, it should be understood that the head coordinate frame 96 is located at a three-dimensional position relative to the Figure 2 PCF of the ground and has three orthogonal axes in three-dimensional space.
[0441] In Figure 53 the head coordinate frame 96 has been moved relative to the Figure 2 PCF of the ground. Since the second user 14.2 has moved their head, the head coordinate frame 96 has moved. The user can move their head with six degrees of freedom (6dof). The head coordinate frame 96 can thus move in 6dof (i.e., in three dimensions from its previous position in Figure 52 and relative to the Figure 2 PCF of the ground about three orthogonal axes). When Figure 9 the real object detection camera 44 and the inertial measurement unit 48 in respectively detect the real object and the movement of the head unit 22, the head coordinate frame 96 is adjusted. More information about head pose tracking is disclosed in U.S. Patent Application No. 16 / 221,065, titled "Enhanced Pose Determination for DisplayDevice", and is hereby incorporated by reference in its entirety.
[0442] Figure 54 shows that sounds can be associated with one or more PCFs. The user can, for example, wear a headset or headphones with stereo. The position of the sound through the headphones can be simulated using conventional techniques. The position of the sound can be located at a fixed position such that when the user rotates their head to the left, the position of the sound rotates to the right, so that the user perceives the sound from the same position in the real world. In this example, the positions of the sounds are represented by sound 123 and sound 456. For the sake of discussion, Figure 54 is similar in terms of analysis to Figure 48 When the first user 14.1 and the second user 14.2 are in the same room at the same or different times, they perceive the sounds 123 and 456 as coming from the same position in the real world.
[0443] Figure 55 and Figure 56 show another implementation of the above technique. As referenced in Figure 8As described, the first user 14.1 has initiated a first session. As Figure 55 shown, the first user 14.1 has terminated the first session, as indicated by the dashed line. At the end of the first session, the first XR device 12.1 will upload Figure 1 to the server 20. The first user 14.1 has now initiated a second session at a time later than the first session. Since Figure 1 has been stored on the first XR device 12.1, the first XR device 12.1 will not download Figure 1 from the server 20. If Figure 1 is lost, then the first XR device 12.1 downloads Figure 1 from the server 20. Then, the first XR device 12.1 continues to build the PCF of Figure 2 , locates to Figure 1 , and further develops the canonical map as described above. Then, as described above, the Figure 2 of the first XR device 12.1 is used to associate local content, head coordinate frame, local sound, etc.
[0444] Reference Figure 57 and Figure 58 , it is also possible that more than one user interacts with the server in the same session. In this example, the first user 14.1 and the second user 14.2 are combined by the third user 14.3 with the third XR device 12.3. Each XR device 12.1, 12.2, and 12.3 starts to generate its own map, namely Figure 1 , Figure 2 and Figure 3 . When the XR devices 12.1, 12.2, and 12.3 continue to develop Figure 1 , 2 and Figure 1 , 2 and
[0445] Figure 59 , the maps are incrementally uploaded to the server 20. The server 20 merges Figure 1 , 2 and Figure 1 to form the canonical map. Then the canonical map is sent from the server 20 to each of the XR devices 12.1, 12.2, and 12.3.Illustrates aspects of a viewing method for restoring and / or resetting a head pose according to some embodiments. In the example shown, at action 1400, the viewing device is powered on. At action 1410, in response to being powered on, a new session is initiated. In some embodiments, the new session may include establishing a head pose. One or more capture devices on a head-mounted frame fixed to the user's head capture the surface of the environment by first capturing an image of the environment and then determining the surface from the image. In some embodiments, the surface data may be combined with data from a gravity sensor to establish the head pose. Other suitable methods for establishing the head pose may be used.
[0446] At action 1420, the processor of the viewing device inputs a routine for tracking the head pose. As the user moves their head to determine the orientation of the head-mounted frame relative to the surface, the capture devices continue to capture the surface of the environment.
[0447] At action 1430, the processor determines whether the head pose has been lost. The head pose may be lost due to "edge" cases such as excessive reflective surfaces, low light, blank walls, outdoors, etc. that can result in low feature acquisition; or due to dynamic situations such as moving crowds and parts of the map. The routine at 1430 allows a certain amount of time, such as 10 seconds, to allow sufficient time to determine whether the head pose has been lost. If the head pose has not been lost, the processor returns to 1420 and enters the tracking of the head pose again.
[0448] If the head pose has been lost at action 1430, the processor enters a routine at 1440 to restore the head pose. If the head pose is lost due to low light, a message such as the following will be displayed to the user via the display of the viewing device:
[0449] The system is detecting low light conditions. Please move to a more well-lit area.
[0450] The system will continue to monitor whether sufficient light is available and whether the head pose can be restored. The system may alternatively determine that the low texture of the surface is causing the head pose to be lost, in which case the following prompt is given to the user in the display as a suggestion to improve surface capture:
[0451] The system cannot detect sufficient surfaces with fine texture. Please move to an area with less rough and more finely textured surfaces.
[0452] At action 1450, the processor enters a routine to determine whether head pose recovery has failed. If head pose recovery has not failed (i.e., head pose recovery has been successful), the processor returns to action 1420 by again entering tracking of the head pose. If head pose recovery has failed, the processor returns to action 1410 to establish a new session. As part of the new session, all cached data is invalidated and head pose is re-established thereafter. Any suitable method of head tracking can be used in conjunction with Figure 59 the process described in. U.S. Patent Application No. 16 / 221,065 describes head tracking and is hereby incorporated by reference in its entirety.
[0453] Remote positioning
[0454] Various embodiments can utilize remote resources to facilitate persistent and consistent cross-reality experiences between individuals and / or groups of users. The inventors have recognized and understood that the benefits of operating an XR device using canonical maps as described herein can be achieved without downloading a set of canonical maps. The Figure 30 above discussion shows an example implementation of downloading a canonical to a device. For example, the benefit of not downloading a map can be achieved by sending feature and pose information to a remote service that maintains a set of canonical maps. According to one embodiment, a device seeking to use a canonical map to position virtual content at a location specified relative to the canonical map can receive one or more transforms between the features and the canonical map from the remote service. These transforms can be used on the device, which maintains information about the location of these features in the physical world, to position virtual content at a location specified relative to the canonical map or otherwise identify a location in the physical world specified relative to the canonical map.
[0455] In some embodiments, spatial information is captured by the XR device and transmitted to a remote service, such as a cloud-based service, that uses the spatial information to position the XR device to a canonical map used by an application or other component of the XR system, thereby specifying the position of virtual content relative to the physical world. Once positioned, the transform that links the tracking map maintained by the device to the canonical map can be transmitted to the device. The transform can be used in conjunction with the tracking map to determine the position to render virtual content specified relative to the canonical map or otherwise identify a location in the physical world specified relative to the canonical map.
[0456] The inventors have realized that the data that needs to be exchanged between the device and the remote location service may be very small compared to transmitting map data, which may occur when the device transmits a tracking map to the remote service and receives a set of canonical maps from the service for device-based positioning. In some embodiments, performing the positioning function on cloud resources only requires transmitting a small amount of information from the device to the remote service. For example, it is not necessary to transmit the complete tracking map to the remote service to perform positioning. In some embodiments, feature and pose information, such as may be stored in relation to the persistent pose as described above, can be transmitted to the remote server. As described above, in embodiments where features are represented by descriptors, the information uploaded may be even smaller.
[0457] The result returned from the location service to the device can be one or more transforms that relate the uploaded features to portions of the matching canonical map. These transforms can be used in the XR system in conjunction with its tracking map to identify the location of virtual content or otherwise identify locations in the physical world. In embodiments that use persistent spatial information such as the PCF described above to specify a location relative to the canonical map, the location service can download the transform between the features and one or more PCFs to the device after successful positioning.
[0458] As a result, the network bandwidth consumed by the communication between the XR device and the remote service for performing positioning can be very low. The system can thus support frequent positioning, enabling each device interacting with the system to quickly obtain information for positioning virtual content or performing other location-based functions. When the device moves in the physical environment, it may repeat requests for updated location information. Additionally, the device may frequently obtain updates to the location information, such as when the canonical map changes, for example by incorporating additional tracking maps to expand the map or improve its accuracy.
[0459] Furthermore, uploading features and downloading transforms can enhance privacy in XR systems that share map information among multiple users by increasing the difficulty of obtaining maps by spoofing. For example, unauthorized users can be prevented from obtaining maps from the system by sending false requests for canonical maps representing portions of the physical world where the unauthorized users are not located. An unauthorized user is unlikely to have access to features in the region of the physical world for which they are requesting map information if the unauthorized user is not actually present in that region. In embodiments where the feature information is formatted as a feature description, it will be even more complex to spoof the feature information in a request for map information. Additionally, when the system returns transforms of the tracking map intended for a device operating in the region where the location information was requested, the information returned by the system may be of little or no use to an impostor.
[0460] According to one embodiment, the location service is implemented as a cloud-based microservice. In some examples, implementing a cloud-based location service can help save device computing resources and enable the computations required for location to be performed with very low latency. These operations can be supported by almost unlimited computing power or other computing resources made available by provisioning additional cloud resources, thus ensuring the scalability of the XR system to support numerous devices. In one example, a number of canonical maps can be maintained in memory for near-instant access or stored in highly available devices to reduce system latency.
[0461] In addition, performing location on multiple devices in the cloud service can enable improvements to the process. Location telemetry and statistics can provide information on which canonical maps are in active memory and / or highly available storage. For example, statistics from multiple devices can be used to identify the most frequently accessed canonical maps.
[0462] As a result of processing in a cloud environment or other remote environment with a large amount of processing resources relative to the remote devices, additional accuracy can also be achieved. For example, location can be performed on a higher density of canonical maps in the cloud as compared to processing performed on a local device. The maps can be stored in the cloud, for example, with more PCFs or a higher density of feature descriptors per PCF, thus improving the accuracy of the match between a set of features from the device and the canonical maps.
[0463] Figure 61 is a schematic diagram of an XR system 6100. The user device that displays cross-reality content during a user session can take various forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As described above, these devices can be configured with software, such as an application or other components, and / or hardwired to generate local location information (e.g., a tracking map) that can be used to render virtual content on their respective displays.
[0464] The virtual content location information can be specified relative to the global location information. For example, the global location information can be formatted as a canonical map that includes one or more PCFs. According to some embodiments, the system 6100 is configured with a cloud-based service that supports the running and display of virtual content on the user device.
[0465] In one example, the positioning function is provided as a cloud-based service 6106, which can be a microservice. The cloud-based service 6106 can be implemented on any of a plurality of computing devices, and computing resources can be allocated from these computing devices to one or more services executing in the cloud. Those computing devices can be interconnected with each other and accessible to devices such as the wearable XR device 6102 and the handheld device 6104. Such connections can be provided through one or more networks.
[0466] In some embodiments, the cloud-based service 6106 is configured to receive descriptor information from various user devices and "locate" the devices to one or more matching canonical maps. For example, the cloud-based positioning service matches the received descriptor information with the descriptor information of the corresponding canonical map. The techniques described above can be used to create canonical maps, which are created by merging maps provided by one or more devices having image sensors or other sensors that acquire information about the physical world. However, it is not required that the canonical maps be created by the devices accessing them, as such maps can be created by map developers. For example, a map developer can publish a map by making it available for use by the positioning service 6106.
[0467] According to some embodiments, the cloud service processes canonical map recognition and can include an operation of filtering a repository of canonical maps into a set of potential matches. The filtering can be performed as Figure 29 shown, or by using any subset of the filtering criteria and instead of Figure 29 the filtering criteria shown in Figure 29 or in addition to Figure 31B , Figure 32 , Figure 33 and Figure 34 the filtering criteria shown in
[0468] Figure 62An example process that can be executed by a device to utilize a canonical map to locate the device's position using cloud-based services and receive transformation information specifying one or more transformations between the device's local coordinate system and the coordinate system of the canonical map. Various embodiments and examples describe the transformation as specifying a transformation from a first coordinate frame to a second coordinate frame. Other embodiments include a transformation from the second coordinate frame to the first coordinate frame. In any other embodiment, the transformation effects a transition from one coordinate frame to another, and the resulting coordinate frame depends only on the desired coordinate frame output (including, for example, the coordinate frame in which content is to be displayed). In yet another embodiment, the coordinate system transformation enables determination from the second coordinate frame to the first coordinate frame and from the first coordinate frame to the second coordinate frame.
[0469] According to some embodiments, information reflecting the transformation for each persistent pose defined by the canonical map can be transmitted to the device.
[0470] According to one embodiment, process 6200 can begin at 6202 with a new session. Starting a new session on the device can initiate the capture of image information to construct a tracking map of the device. Additionally, the device can send a message to register with the server of the location service, prompting the server to create a session for the device.
[0471] In some embodiments, starting a new session on the device can optionally include sending adjustment data from the device to the location service. The location service returns one or more transformations calculated based on a set of features and associated poses. If the poses of the features are adjusted based on device-specific information before calculating the transformation and / or the transformation is adjusted based on device-specific information after calculating the transformation, rather than performing those calculations on the device, the device-specific information may be sent to the location service so that the location service can apply these adjustments. As a specific example, sending device-specific adjustment information can include capturing calibration data for sensors and / or displays. The calibration data can be used, for example, to adjust the position of feature points relative to the measured positions. Alternatively or additionally, the calibration data can be used to adjust the position at which the command display renders virtual content so that it appears accurately positioned for that particular device. The calibration data can be obtained, for example, from multiple images of the same scene taken using sensors on the device. The positions of the features detected in those images can be expressed as a function of the sensor positions, such that the multiple images yield a set of equations that can be solved for the sensor positions. The calculated sensor positions can be compared to the nominal positions, and the calibration data can be derived from any differences. In some embodiments, intrinsic information about the device's construction can also enable the calculation of calibration data for the display, in some embodiments.
[0472] In embodiments for generating calibration data for a sensor and / or a display, the calibration data may be applied at any point during a measurement or display process. In some embodiments, the calibration data may be sent to a location server, which may store the calibration data in a data structure established for each device that has registered with the location server and is thus in a session with the server. The location server may apply the calibration data to any transformation computed as part of the location process for the device for which the calibration data is provided. Thus, the computational burden of using the calibration data to improve the accuracy of sensed and / or displayed information is borne by the calibration service, providing a further mechanism for reducing the processing burden on the device.
[0473] Once a new session is established, process 6200 may continue at 6204 to capture new frames of the device's environment. At 6206, each frame may be processed to generate a descriptor for the captured frame (including, for example, the DSF values discussed above). These values may be computed using some or all of the techniques described above, including those discussed above with respect to Figure 14 , Figure 22 and Figure 23 As discussed, the descriptor may be computed as a mapping of feature points, or in some embodiments, a mapping of image patches around the feature points to the descriptor. The descriptor may have values that enable effective matching between newly acquired frames / images and the stored map. Additionally, the number of features extracted from an image may be limited to a maximum number of feature points per image, such as 200 feature points per image. As described above, the feature points may be selected to represent areas of interest. Thus, actions 6204 and 6206 may be performed as part of a device process for forming a tracking map or otherwise periodically collecting images of the physical world around the device, or may be performed separately, but not necessarily, for localization.
[0474] Feature extraction at 6206 may include attaching pose information to the features extracted at 6206. The pose information may be the pose in the device's local coordinate system. In some embodiments, the pose may be relative to a reference point in the tracking map, such as the persistent pose described above. Alternatively or additionally, the pose may be relative to the origin of the device's tracking map. Such embodiments may enable the location service described herein to provide location services for a wide range of devices, even if they do not use a persistent pose. In any case, the pose information may be attached to each feature or group of features such that the location service may use the pose information to compute a transformation that may be returned to the device when matching the features to features in the stored map.
[0475] Process 6200 may continue to decision block 6207, where a decision is made whether to request a location. One or more criteria may be applied to determine whether to request a location. The criteria may include the passage of time such that the device may request a location after a certain threshold amount of time. For example, if no location attempt has been made within the threshold amount of time, the process may continue from decision block 6207 to action 6208, where a location is requested from the cloud. The threshold amount of time may be between 10 and 30 seconds, such as 25 seconds. Alternatively or additionally, the location may be triggered by the movement of the device. The device executing process 6200 may use the IMU and its tracking map to track its movement and initiate a location when it detects movement beyond a threshold distance from the location where the device was last requested to be located. For example, the threshold distance may be between 1 and 10 meters, such as between 3 and 5 meters. As yet another alternative, the location may be triggered in response to an event, such as when the device creates a new persistent pose or the current persistent pose of the device changes, as described above.
[0476] In some embodiments, decision block 6207 may be implemented such that the thresholds for triggering a location may be established dynamically. For example, in an environment where the features are largely consistent such that the confidence of matching a set of extracted features to the features of a stored map may be low, locations may be requested more frequently to increase the chance that at least one location attempt will succeed. In such a case, the threshold applied at decision block 6207 may be decreased. Similarly, in an environment where there are relatively few features, the threshold applied at decision block 6207 may be decreased to increase the frequency of location attempts.
[0477] Regardless of how the location is triggered, when triggered, process 6200 may proceed to action 6208, where the device sends a request to the location service, including data used by the location service to perform the location. In some embodiments, data from multiple image frames may be provided for a location attempt. For example, the location service may not consider the location successful unless the features in multiple image frames produce consistent location results. In some embodiments, process 6200 may include saving feature descriptors and additional pose information to a buffer. The buffer may be, for example, a circular buffer that stores a set of features extracted from the most recently captured frames. Thus, the location request may be sent with multiple sets of features accumulated in the buffer. In some settings, the buffer size is implemented to accumulate multiple data sets that are more likely to result in a successful location. In some embodiments, the buffer size may be set to accumulate features from, for example, two, three, four, five, six, seven, eight, nine, or ten frames. Optionally, the buffer size may have a baseline setting that may increase in response to a location failure. In some examples, increasing the buffer size and the corresponding number of sets of features transmitted reduces the likelihood that subsequent location functions will not return a result.
[0478] Regardless of how the buffer size is set, the device can transfer the contents of the buffer to the location service as part of a location request. Other information can be transmitted along with the feature points and additional pose information. For example, in some embodiments, geographic information can be transmitted. The geographic information can include, for example, GPS coordinates or a wireless signature associated with the device tracking the map or the current persistent pose.
[0479] In response to the request sent at 6208, the cloud location service can analyze the feature descriptors to locate the device to a canonical map or other persistent map maintained by the service. For example, the descriptors match a set of features in the map where the device is located. The cloud-based location service can perform the location as described above relative to the device-based location (e.g., can rely on any of the functions for location discussed above, including map ranking, map filtering, position estimation, filtered map selection, Figures 44 to 4 the example in 6, and / or as discussed with respect to location module, PCF, and / or PP identification and matching, etc.). However, instead of transmitting the identified canonical map to the device (e.g., in device location), the cloud-based location service can continue to generate a transformation based on the matching features of the canonical map and the relative orientation of the set of features sent from the device. The location service can return these transformations to the device, which can be received at block 6210.
[0480] In some embodiments, the canonical map maintained by the location service can employ a PCF, as described above. In such embodiments, the feature points of the canonical map that match the feature points sent from the device can have positions specified relative to one or more PCFs. Thus, the location service can identify one or more canonical maps and can calculate the transformation between the coordinate frame represented in the pose sent with the location request and the one or more PCFs. In some embodiments, identifying one or more canonical maps is aided by filtering potential maps based on the geographic data of the corresponding device. For example, once filtered to a candidate set (e.g., by other options such as GPS coordinates), the candidate set of canonical maps can be analyzed in detail to determine the matching feature points or PCFs as described above.
[0481] The data returned to the requesting device at action 6210 can be formatted as a persistent pose transformation table. This table can be accompanied by one or more canonical map identifiers indicating the canonical map to which the device has been located by the location service. However, it should be understood that the location information can be formatted in other ways, including as a list of transformations, with associated PCFs and / or canonical map identifiers.
[0482] Regardless of how the transforms are formatted, at action 6212, the device can use these transforms to calculate the position for rendering virtual content, the position of which has been specified by an application or other component of the XR system relative to any PCF. This information is alternatively or additionally used on the device to perform any location-based operations where the location is specified based on the PCF.
[0483] In some scenarios, the positioning service may not be able to match the features sent from the device to any stored canonical map, or may not be able to match a sufficient number of feature sets transmitted along with the request for the positioning service to consider the positioning successful. In such scenarios, the positioning service can indicate a positioning failure to the device, rather than returning the transforms to the device as described above in connection with action 6210. In such scenarios, process 6200 can branch to action 6230 at decision box 6209, where the device can take one or more actions for failure handling. These actions can include increasing the size of a buffer that stores the feature sets sent for positioning. For example, if the positioning service does not consider the positioning successful unless three feature sets match, the buffer size can be increased from 5 to 6, thereby increasing the chance that the three transmitted feature sets match the canonical map maintained by the positioning service.
[0484] Alternatively or additionally, the failure handling can include adjusting the operating parameters of the device to trigger more frequent positioning attempts. For example, the threshold time and / or threshold distance between positioning attempts can be reduced. As another example, the number of feature points in each feature set can be increased. A match between a feature set and the features stored in the canonical map can be considered to occur when a sufficient number of features in the set sent from the device match the features of the map. Increasing the number of features sent can increase the chance of a match. As a specific example, the initial feature set size can be 50, and at each successive positioning failure, it can be increased to 100, 150, and then 200. After a successful match, the size of the set can then return to its initial value.
[0485] The failure handling can also include obtaining positioning information from sources other than the positioning service. According to some embodiments, the user device can be configured to cache the canonical map. Caching the map allows the device to access and display content when the cloud is unavailable. For example, the cached canonical map allows for device-based positioning in the event of communication failure or other unavailability.
[0486] According to various embodiments, Figure 62 a high-level process for device-initiated cloud-based positioning is described. In other embodiments, various ones of the steps shown can be combined, omitted, or other processes can be invoked to complete the positioning and eventual visualization of virtual content in the corresponding device view.
[0487] In addition, it should be understood that although process 6200 shows the device determining whether to initiate localization at decision block 6207, the trigger for initiating localization can come from outside the device, including from a localization service. For example, a localization service can maintain information about each device in its session. For example, this information can include the identifier of the canonical map to which each device was most recently localized. The localization service or other components of the XR system can update the canonical map, including using the techniques described above in connection with Figure 26 When the canonical map is updated, the localization service can send a notification to each device that was most recently localized to that map. This notification can be used as a trigger for the device to request localization and / or can include an updated transformation recomputed using the set of features most recently sent from the device.
[0488] Figure 63A , Figure 63B and Figure 63C are example process flows showing the operations and communications between a device and a cloud service. What is shown in blocks 6350, 6352, 6354, and 6456 is an example separation of the architecture and components involved in a cloud-based localization process. For example, the modules, components, and / or software configured to handle perception on a user device are shown at 6350 (e.g., 660, Figure 6A ). The device functions for persistent world operations are shown at 6352 (including, for example, as described above and with respect to the persistent world module (e.g., 662, Figure 6A ). In other embodiments, the separation between 6350 and 6352 is not required and the communications shown can be between processes executed on the device.
[0489] Similarly, shown at block 6354 are cloud processes configured to handle functions associated with the traversable world / traversable world modeling (e.g., 802, 812, Figure 26 ). Shown at block 6356 are cloud processes configured to handle functions associated with localizing a device to one or more maps in a repository of stored canonical maps based on information sent from the device.
[0490] In the illustrated embodiment, process 6300 begins at 6302 when a new session starts. Sensor calibration data is obtained at 6304. The calibration data obtained can depend on the device represented at 6350 (e.g., multiple cameras, sensors, localization devices, etc.). Once sensor calibration is obtained for the device, the calibration can be cached at 6306. If a device operation causes a change in frequency parameters (e.g., collection frequency, sampling frequency, matching frequency, and other options), the frequency parameters are reset to a baseline at 6308.
[0491] Once the new session functionality is complete (e.g., calibration, steps 6302 - 6306), process 6300 can continue to capture new frame 6312. At 6314, features and their corresponding descriptors are extracted from the frame. In some examples, the descriptors can include DSF as described above. According to some embodiments, the descriptors can have spatial information attached to them for subsequent processing (e.g., transform generation). At 6316, pose information generated on the device (e.g., information for locating features in the physical world relative to the tracking map specified for the device as described above) can be attached to the extracted descriptors.
[0492] At 6318, the descriptors and pose information are added to the buffer. The new frame capture and addition to the buffer shown in steps 6312 - 6318 are executed in a loop until the buffer size threshold is exceeded at 6319. At 6320, in response to determining that the buffer size is met, a localization request is transmitted from the device to the cloud. According to some embodiments, the request can be processed by a traversable world service (e.g., 6354) instantiated in the cloud. In further embodiments, the functional operations for identifying candidate canonical maps can be separated from the operations for actual matching (e.g., shown as boxes 6354 and 6356). In one embodiment, a cloud service for map filtering and / or map ranking can be executed at 6354 and process the localization request received from 6320. According to one embodiment, the map ranking operation is configured to determine a set of candidate maps that may include the location of the device at 6322.
[0493] In one example, the map ranking function includes operations for identifying candidate canonical maps based on geographical attributes or other location data (e.g., observed or inferred location information). For example, other location data can include Wi-Fi signatures or GPS information.
[0494] According to other embodiments, location data can be captured during a cross - reality session with the device and the user. Process 6300 can include additional operations for populating locations for a given device and / or session (not shown). For example, the location data can be stored as device area attribute values and attribute values for selecting candidate canonical maps close to the device location.
[0495] Any one or more location options can be used to filter the set of canonical maps to those sets of canonical maps that may represent the area including the location of the user device. In some embodiments, the canonical maps can cover a relatively large area of the physical world. The canonical maps can be segmented into regions such that the selection of a map may require the selection of a map region. For example, the map region can be on the order of dozens of square meters. Thus, the filtered set of canonical maps can be a set of regions of the map.
[0496] According to some embodiments, a localization snapshot can be constructed from a candidate canonical map, pose features, and sensor calibration data. For example, an array of candidate canonical maps, pose features, and sensor calibration information can be sent along with a request to determine a specific matching canonical map. Matching with the canonical map can be performed based on descriptors received from the device and stored PCF data associated with the canonical map. Additionally, matching with the canonical map can be performed based on high-resolution ...
Claims
1. A network resource device in a distributed computing environment, the network resource device being configured to provide shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a three-dimensional (3D) environment, the network resource device comprising: one or more processors; and at least one computer-readable medium comprising: a plurality of stored maps of the 3D environment; and computer-executable instructions that, when executed by the one or more processors, cause the network resource device to: receive information from a portable electronic device regarding a plurality of features detected in an image captured by the portable electronic device; and compute a frame descriptor for the image, wherein the computed frame descriptor has a resolution greater than 512 bits.
2. The network resource device according to claim 1, wherein: at least one of the plurality of stored maps of the 3D environment is associated with at least one frame descriptor having a resolution greater than 512 bits.
3. The network resource device according to claim 2, wherein: the computer-executable instructions, when executed by one or more processors, further cause the network resource device to: compare the computed frame descriptor for the image with the at least one frame descriptor associated with at least one of the plurality of stored maps of the 3D environment.
4. The network resource device according to claim 3, wherein: the computer-executable instructions, when executed by one or more processors, further cause the network resource device to: select one or more maps from the plurality of stored maps to locate the portable electronic device in a shared coordinate system based on a comparison of the computed frame descriptor for the image with the at least one frame descriptor associated with at least one of the plurality of stored maps.
5. The network resource device according to claim 4, wherein: the computer-executable instructions, when executed by a processor of the one or more processors, further cause the network resource device to: send the selected one or more maps to the portable electronic device.
6. The network resource device according to claim 3, wherein: the computer-executable instructions, when executed by a processor of the one or more processors, further cause the network resource device to: determine whether the location of the portable electronic device corresponds to a stored map from the plurality of stored maps of the 3D environment based on a comparison of the computed frame descriptor for the image with the at least one frame descriptor associated with at least one of the plurality of stored maps.
7. The network resource device according to claim 6, wherein: the computer-executable instructions, when executed by a processor of the one or more processors, further cause the network resource device to: receive a tracking map from the portable electronic device; merge the tracking map with the stored map to generate a merged map including location information based on the location information of the stored map and the tracking map; and and store the merged map in the computer-readable medium.
8. The network resource device according to any one of claims 1 to 7, wherein, the portable electronic device is selected from the group consisting of: a wearable device including a head-mounted display having a plurality of cameras mounted thereon; and a portable computing device including a camera and a display and configured with computer-executable instructions for rendering virtual content related to an image acquired by the camera on the display.
9. The network resource device according to any one of claims 1 to 7, wherein, the computer-executable instructions include instructions for implementing a neural network to compute a frame descriptor.
10. The network resource device according to claim 1, wherein, the frame descriptor of the image represents the plurality of features detected in the image.
11. A method for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a three-dimensional (3D) environment, the method comprising: on the portable electronic device: acquiring one or more images of the 3D environment; identifying one or more features from the one or more images; transmitting information about the one or more features identified in the one or more images of the 3D environment to a network resource; and computing one or more first frame descriptors of the one or more images based on the one or more features, the first frame descriptors having a first resolution; on the network resource: storing a plurality of maps of the 3D environment; and computing one or more second frame descriptors representing the one or more images based on the information about the one or more features, wherein the one or more second frame descriptors have a second resolution greater than the first resolution.
12. The method according to claim 11, further comprising: on the portable electronic device, selecting at least a portion of a local map based on the one or more first frame descriptors; and on the network resource, selecting at least a portion of a shared map based on the one or more second frame descriptors.
13. The method according to claim 12, wherein, selecting at least a portion of the shared map includes: comparing the one or more second frame descriptors with one or more frame descriptors associated with the plurality of maps of the 3D environment.
14. The method according to claim 13, further comprising: on the network resource, determining one or more maps from the plurality of maps for positioning the portable electronic device in a shared coordinate system based on a comparison of the one or more second frame descriptors and one or more frame descriptors associated with the plurality of maps of the 3D environment.
15. The method according to claim 14, further comprising: for the one or more maps determined for positioning the portable electronic device, computing one or more third frame descriptors, wherein the one or more third frame descriptors have the first resolution; Send the determined one or more maps for positioning the portable electronic device and the one or more third frame descriptors from the network resource to the portable electronic device.
16. The method according to claim 12, further comprising: On the network resource, based on a comparison of one or more frame descriptors calculated according to information about a pixel group received from the portable electronic device and one or more frame descriptors associated with the plurality of maps of the 3D environment, determine whether the position of the portable electronic device corresponds to a stored map among the plurality of stored maps from the 3D environment.
17. The method according to claim 16, wherein: The information about the one or more features includes a tracking map; and the method further comprises: At the network resource, merge the tracking map with the stored map to generate a merged map including position information based on the position information of the stored map and the tracking map; and At the network resource, store the merged map together with the one or more second frame descriptors in the computer-readable medium.
18. The method according to any one of claims 11 to 17, wherein, The portable electronic device is selected from the group consisting of: A wearable device, which includes a head-mounted display, and the head-mounted display includes a plurality of cameras mounted thereon; and A portable computing device, which includes a camera and a display, and is configured with computer-executable instructions for rendering virtual content related to an image acquired by the camera on the display.
19. The method according to any one of claims 11 to 17, wherein, The action of calculating one or more first frame descriptors and / or the action of calculating one or more second frame descriptors is performed using a neural network.
20. The method according to claim 11, wherein, The first frame descriptor represents the one or more features.
21. A system for providing shared location-based content to a plurality of portable electronic devices capable of rendering virtual content in a three-dimensional (3D) environment, the system comprising: At least one portable electronic device configured to render virtual content; and At least one network resource; wherein each of the at least one portable electronic device includes at least one processor, at least one camera, and at least one computer-readable medium including instructions that, when executed, cause the at least one processor to perform: Capture at least one image of the 3D environment with the at least one camera; For the at least one image, identify multiple groups of pixels representing features; For the multiple groups of pixels, calculate descriptors representing the multiple groups of pixels; Create a data structure including the descriptors calculated for the multiple groups of pixels; Send the data structure to the network resource; Calculate at least one first frame descriptor according to the descriptors calculated for the multiple groups of pixels, the at least one first frame descriptor having a first resolution; and Compare the image frames local to the portable electronic device based on the at least one first frame descriptor having the first resolution; and wherein the at least one network resource includes one or more processors and at least one computer-readable medium, the at least one computer-readable medium comprising: multiple stored maps of the 3D environment, wherein at least one of the multiple stored maps is associated with at least one frame descriptor; and computer-executable instructions that, when executed by the one or more processors, cause the network resource to: calculate at least one second frame descriptor using a neural network based on the descriptors calculated for the multiple sets of pixels, wherein the at least one second frame descriptor has a second resolution higher than the first resolution; compare the at least one second frame descriptor with the at least one frame descriptor associated with the at least one stored map of the 3D environment; and locate the portable electronic device in a shared coordinate system based on the comparison of the at least one second frame descriptor and the at least one frame descriptor associated with the at least one stored map.
22. The system according to claim 21, wherein, the at least one portable electronic device is selected from the group consisting of: a wearable device that includes a head-mounted display, the head-mounted display including a plurality of cameras mounted thereon; and a portable computing device that includes a camera and a display and is configured with computer-executable instructions for rendering virtual content related to an image acquired by the camera on the display.
23. The system according to claim 21, wherein, the computer-executable instructions of the at least one network resource further include computer-executable instructions that, when executed by the one or more processors, cause the network resource to: receive a tracking map from the portable electronic device among the at least one portable electronic device; and merge the tracking map with the at least one stored map to generate a merged map including location information based on the location information of the at least one stored map and the tracking map, and store the merged map in the computer-readable medium.
24. The system according to claim 21, wherein, the at least one first frame descriptor represents the descriptor calculated for the multiple sets of pixels, and wherein the at least one second frame descriptor represents the descriptor calculated for the multiple sets of pixels.
Citation Information
Patent Citations
Localization determination for mixed reality systems
US10812936B2
Fully convolutional interest point detection and description via homographic adaptation
US20190147341A1
Enhanced pose determination for display device
US20190188474A1
Methods and apparatuses for determining and / or evaluating localizing maps of image display devices
US20200034624A1
Using a map of the world for augmented or virtual reality systems
US20150302656A1