Cross-reality system with wireless fingerprint
By forming and maintaining wireless fingerprints in the XR system, and using network access point information and persistent coordinate frameworks, the problem of limited computing resources in the XR system is solved, faster and more accurate positioning and virtual content display are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202080072241.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-15
- Filing Date
- 2020-10-15
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2040-10-15
AI Technical Summary
When building and maintaining physical environment maps, the computing resources are limited, resulting in positioning delays and excessive computing burdens, affecting the user experience.
By forming and maintaining wireless fingerprints on portable devices, using wireless signal network access point information, updating location information in the map, combining persistent coordinate framework and map merging technology, reducing computing needs and improving positioning efficiency.
It realizes faster and more accurate positioning and virtual content display, reduces latency, improves user immersion, and enhances data sharing and synchronization of virtual content between multiple devices.
Smart Images

Figure CN114616534B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit under 35 U.S.C. Section 119(e) of U.S. Provisional Patent Application Serial No. 62 / 915,559, filed on October 15, 2019, and entitled “CROSSREALITY SYSTEM WITH WIRELESS FINGERPRINTS,” the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application relates generally to cross-reality systems. Background Art
[0004] A computer can control a human user interface to create a cross-reality (XR) environment in which some or all of the XR environment perceived by the user is generated by the computer. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments, some or all of which can be generated in part by a computer using data describing the environment. For example, the data can describe virtual objects that can be rendered in a way that the user feels or perceives as part of the physical world, and the virtual objects can be interacted with. Because the data is rendered and presented through a user interface device (such as, for example, a head-mounted display device), the user can experience these virtual objects. The data can be displayed to the user, or can control audio that is played to the user, or can control a tactile (or haptic) interface, allowing the user to experience the touch sensation of the virtual objects that the user feels or perceives as being felt.
[0005] XR systems can be used for a wide range of applications across scientific visualization, medical training, engineering design and prototyping, telemanipulation and telepresence, and personal entertainment. Compared to VR, AR and MR involve one or more virtual objects that are associated with real objects in the physical world. The experience of virtual objects interacting with real objects significantly enhances the user's enjoyment of using XR systems and also opens the door to a variety of applications that present realistic and easily understood information about how to modify the physical world.
[0006] To realistically render virtual content, an XR system can build a representation of the physical world surrounding the user of the system. For example, this representation can be constructed by processing images acquired using sensors on a wearable device, where the wearable device forms part of the XR system. In such a system, the user can perform an initialization routine by looking around the room or other physical environment in which the user intends to use the XR system until the system has enough information to build a representation of that environment. As the system operates and the user moves within the environment or to other environments, sensors on the wearable device can acquire additional information to expand or update the representation of the physical world. Summary of the Invention
[0007] Aspects of the present application relate to methods and apparatus for providing an X-reality (cross reality or XR) scene. The techniques described herein may be used together, individually, or in any suitable combination.
[0008] According to one embodiment, a portable device is configured to operate and display virtual content of a cross-reality (XR) system in a three-dimensional (3D) environment. The portable device may include: at least one processor, and computer-executable instructions that can be executed by the at least one processor. The computer-executable instructions can be configured to perform a method when executed by the at least one processor, the method comprising: forming a map of the 3D environment as the portable device moves in the 3D environment, and maintaining a wireless fingerprint associated with the portable device. Maintaining the wireless fingerprint associated with the portable device may include repeatedly performing the following steps: obtaining network access point information from a network access point that transmits a wireless signal at a location within the 3D environment, the wireless signal being received by the portable device; selecting a location within the map that corresponds to the location of the portable device within the 3D environment; and updating information stored in association with the selected location within the map based on the obtained network access point information.
[0009] According to one embodiment, updating the information stored in association with the selected location includes combining the obtained network access point information with previously obtained network access point information associated with the selected location.
[0010] According to one embodiment, obtaining network access point information includes triggering a scan for network access points. According to one embodiment, the scan for network access points is triggered when an amount of time that has elapsed since a previous scan for network access points exceeds a threshold. According to one embodiment, the scan for network access points is triggered when a distance that the portable device has moved in the 3D environment exceeds a threshold. According to one embodiment, obtaining network access point information includes receiving the network access point information pushed from a wireless hardware component after the scan for network access points.
[0011] According to one embodiment, obtaining network access point information includes obtaining an access point identifier for the network access point. According to one embodiment, obtaining network access point information further includes obtaining one or more signal strength indicator values for the access point identified by the access point identifier. According to one embodiment, the access point identifier is a basic service set identifier (BSSID). According to one embodiment, the signal strength indicator value is a received signal strength indicator (RSSI) value.
[0012] According to one embodiment, updating the information stored associated with the selected location within the map includes storing the one or more signal strength indication values in association with the access point identifier. According to one embodiment, updating the information stored associated with the selected location within the map also includes identifying a subset of the access point identifiers to be excluded. According to one embodiment, the subset of access point identifiers to be excluded is based at least on the one or more signal strength indication values. According to one embodiment, storing the one or more signal strength indication values includes storing an average of a plurality of signal strength indication values in association with the access point identifier.
[0013] According to one embodiment, the selected location within the map comprises a persistent pose or persistent coordinate frame of the map.
[0014] According to one embodiment, obtaining the network access point information further includes filtering, clustering, and / or normalizing the network access point information.
[0015] According to one embodiment, the portable device further comprises a computer-readable medium connected to the at least one processor, and the method further comprises storing the map on the computer-readable medium.
[0016] According to one embodiment, updating information stored in association with the selected location within the map based on the obtained network access point information includes updating information stored on the computer-readable medium in association with the selected location within the map based on the obtained network access point information.
[0017] According to one embodiment, a method for operating a portable device configured to operate and display virtual content of a cross-reality system in a three-dimensional (3D) environment includes: forming a map of the 3D environment as the portable device moves within the 3D environment, and maintaining a wireless fingerprint associated with the portable device. Maintaining the wireless fingerprint associated with the portable device may include repeatedly performing the following steps: obtaining network access point information from a network access point transmitting a wireless signal at a location within the 3D environment, the wireless signal being received by the portable device; selecting a location within the map that corresponds to the location of the portable device within the 3D environment; and updating information stored in association with the selected location within the map based on the obtained network access point information.
[0018] According to one embodiment, a computer-readable medium storing computer-executable instructions is configured to, when executed by at least one processor, perform a method for operating a portable device configured to operate and display virtual content of a cross-reality system in a three-dimensional (3D) environment. The method includes forming a map of the 3D environment as the portable device moves within the 3D environment, and maintaining a wireless fingerprint associated with the portable device. Maintaining the wireless fingerprint associated with the portable device may include repeatedly performing the following steps: obtaining network access point information from a network access point transmitting a wireless signal at a location within the 3D environment, the wireless signal being received by the portable device; selecting a location within the map that corresponds to the location of the portable device within the 3D environment; and updating information stored associated with the selected location within the map based on the obtained network access point information.
[0019] According to one embodiment, a computing device configured for use in a cross-reality system, wherein a portable device operating in a three-dimensional (3D) environment renders virtual content, may include: at least one processor; a computer-readable medium connected to the at least one processor, wherein a plurality of maps are stored on the computer-readable medium, wherein the plurality of maps include information identifying locations within the 3D environment, the information including stored location information and stored network access point information associated with corresponding locations in the 3D environment; and computer-executable instructions configured to perform a method when executed by the at least one processor. The method performed by the computer-executable instructions may include: receiving information of the portable device indicating a location within the 3D environment, the received information of the portable device including location information and network access point information; selecting one or more locations within at least one of the multiple maps as candidate locations based at least on a comparison of the received network access point information of the portable device with stored network access point information of the multiple maps; determining whether the received location information of the portable device is associated with a location in the 3D environment that is the same as location information of a candidate location contained in the stored location information of the multiple maps; and, for a candidate location determined to be associated with the same location indicated by the received location information of the portable device, calculating a transformation between the received location information of the portable device and the location information of the candidate location contained in the stored location information of the multiple maps.
[0020] According to one embodiment, the received network access point information of the portable device and / or the stored network access point information of the plurality of maps comprises an access point identifier and a signal strength indicator value. According to one embodiment, the access point identifier is a BSSID and the signal strength indicator value is an RSSI value. According to one embodiment, the comparison of the received network access point information of the portable device with the stored network access point information of the plurality of maps comprises: determining a Jaccard similarity between the received network access point information of the portable device and the stored network access point information of the plurality of maps. According to one embodiment, the comparison of the received network access point information of the portable device with the stored network access point information of the plurality of maps further comprises: determining whether the Jaccard similarity is above a threshold value.
[0021] According to one embodiment, the received position information of the portable device includes a first coordinate frame, and the stored position information of the plurality of maps includes a plurality of coordinate frames. According to one embodiment, calculating a transformation between the received position information of the portable device and the position information of the candidate position contained in the stored position information of the plurality of maps includes calculating a transformation between the first coordinate frame and a second coordinate frame of the plurality of coordinate frames. According to one embodiment, calculating the transformation is based on an alignment calculated between a feature set in the 3D environment of the portable device and a matching feature set in the plurality of maps containing the candidate position.
[0022] According to one embodiment, the plurality of maps from which the candidate location is selected comprises a filtered subset of stored maps.
[0023] According to one embodiment, the received network access point information of the portable device is stored in a first map stored on the computer-readable medium, and the candidate location determined to have location information associated with the same location as the received location information of the portable device is included in a second stored map. The method may also include: merging the first stored map with the second stored map based on the first stored map and the second stored map to generate a merged map including location information and network access point information; and storing the merged map on the computer-readable medium.
[0024] According to one embodiment, the received network access point information of the portable device is stored in a tracking map received from the portable device, and the candidate location determined to have location information associated with the same location as the received location information of the portable device is included in a first stored map. The method may further include: based on the first stored map and the tracking map, merging the first stored map with the tracking map to generate a merged map including location information and network access point information; and storing the merged map on the computer-readable medium.
[0025] According to one embodiment, a method of operating a cross-reality system is provided. The cross-reality system may include a portable device and a remote computing device, wherein the portable device operates within a 3D environment and wherein the portable device and the remote computing device are configured to interact with each other. The method of operating the cross-reality system may include: accumulating, on the portable device, network access point information associated with each of a plurality of locations in the 3D environment over time, wherein the network access point information includes a plurality of wireless network access point identifiers and associated average signal strength values; sending a request from the portable device to the remote computing device, the request including at least the network access point information at the location of the portable device within the 3D environment and information representing the location within the 3D environment, the information being expressed in a coordinate frame associated with the portable device; and computing, on the remote computing device, a transformation between the coordinate frame associated with the portable device and a coordinate frame associated with at least one stored map, wherein the map includes a location matching the network access point information at the location of the portable device within the 3D environment and the information representing the location within the 3D environment.
[0026] According to one embodiment, the method of operating a cross-reality system may further include sending the transformation from the remote computing device to the portable device in response to the request. According to one embodiment, the network access point information includes a BSSID and an RSSI value. According to one embodiment, the remote computing device includes a distributed server network in a cloud computing configuration. According to one embodiment, the method is repeatedly performed as the portable device moves in the 3D environment. According to one embodiment, the portable device and the remote computing device interact with each other via a wireless communication channel.
[0027] According to one embodiment, the request from the portable device comprises a tracking map of the 3D environment of the portable device.According to one embodiment, the coordinate frame associated with the portable device is the coordinate frame of the tracking map.
[0028] According to one embodiment, a portable device is configured to operate and display virtual content of a cross-reality system in a three-dimensional (3D) environment. The portable device may include: at least one processor; a computer-readable medium connected to the at least one processor; and computer-executable instructions stored on the computer-readable medium, the computer-executable instructions configured to perform a method when executed by the at least one processor. The method may include: forming a map of the 3D environment as the portable device moves within the 3D environment; storing the map on the computer-readable medium; and maintaining a wireless fingerprint associated with the portable device. Maintaining the wireless fingerprint associated with the portable device may include repeatedly performing the following steps: obtaining network access point information from a network access point transmitting a wireless signal at a location within the 3D environment, the wireless signal being received by the portable device; selecting a location within the map that corresponds to the location of the portable device within the 3D environment; and updating information stored on the computer-readable medium associated with the selected location within the map based on the obtained network access point information.
[0029] The foregoing summary has been provided by way of illustration and is not intended to be limiting. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by a like numeral. For clarity, not every component may be labeled in every figure. In the drawings:
[0031] Figure 1 is a schematic diagram illustrating an example of a simplified augmented reality (AR) scene according to some embodiments;
[0032] Figure 2 is a schematic diagram of an exemplary simplified AR scene, illustrating an exemplary use case of an XR system, according to some embodiments;
[0033] Figure 3 is a diagram illustrating data flow for a single user in an AR system configured to provide the user with an experience of AR content that interacts with the physical world, according to some embodiments;
[0034] Figure 4 is a schematic diagram illustrating an exemplary AR display system that displays virtual content for a single user according to some embodiments;
[0035] Figure 5A is a diagram illustrating an AR display system rendering AR content as a user moves through a physical world environment while the user is wearing the AR display system, according to some embodiments;
[0036] Figure 5B is a schematic diagram illustrating a viewing optics assembly and accompanying components according to some embodiments;
[0037] Figure 6A is a schematic diagram illustrating an AR system using a world reconstruction system according to some embodiments;
[0038] Figure 6B is a diagram illustrating components of an AR system that maintains a model of a navigable world, according to some embodiments.
[0039] Figure 7 A diagram of a tracking graph formed by the path traversed by a device through the physical world.
[0040] Figure 8 is a schematic diagram illustrating a user of a cross-reality (XR) system perceiving virtual content in accordance with some embodiments;
[0041] Figure 9 is a method for transforming between coordinate systems according to some embodiments Figure 8 A block diagram of components of a first XR device of an XR system;
[0042] Figure 10 is a diagram illustrating an exemplary transformation of an origin coordinate frame into a destination coordinate frame for correctly rendering local XR content, according to some embodiments;
[0043] Figure 11 is a top plan view illustrating a pupil-based coordinate framework according to some embodiments;
[0044] Figure 12 is a top plan view illustrating a camera coordinate frame including all pupil positions according to some embodiments;
[0045] Figure 13 According to some embodiments Figure 9 A schematic diagram of a display system;
[0046] Figure 14 is a block diagram illustrating the creation of a persistent coordinate frame (PCF) and the attachment of XR content to the PCF in accordance with some embodiments;
[0047] Figure 15 is a flowchart illustrating a method of establishing and using a PCF according to some embodiments;
[0048] Figure 16 is according to some embodiments including a second XR device Figure 8 Block diagram of the XR system;
[0049] Figure 17is a schematic diagram illustrating a room and keyframes established for various areas in the room according to some embodiments;
[0050] Figure 18 is a schematic diagram illustrating the establishment of a keyframe-based persistent pose according to some embodiments;
[0051] Figure 19 is a schematic diagram illustrating the establishment of a persistent coordinate frame (PCF) based on persistent gestures according to some embodiments;
[0052] 20A to 20C is a schematic diagram illustrating an example of creating a PCF according to some embodiments;
[0053] Figure 21 is a block diagram illustrating a system for generating a global descriptor for a single image and / or map according to some embodiments;
[0054] Figure 22 is a flowchart illustrating a method of computing an image descriptor according to some embodiments;
[0055] Figure 23 is a flow chart illustrating a localization method using image descriptors according to some embodiments;
[0056] Figure 24 is a flowchart illustrating a method of training a neural network according to some embodiments;
[0057] Figure 25 is a block diagram illustrating a method of training a neural network according to some embodiments;
[0058] Figure 26 is a schematic diagram illustrating an AR system configured to rank and merge multiple environment maps according to some embodiments;
[0059] Figure 27 is a simplified block diagram illustrating a plurality of canonical maps stored on a remote storage medium according to some embodiments;
[0060] Figure 28 is a schematic diagram illustrating a method of selecting a canonical map, for example, to locate a new tracking map in one or more canonical maps and / or to obtain a PCF from the canonical map, according to some embodiments;
[0061] Figure 29 is a flow chart illustrating a method of selecting a plurality of ranked environment maps according to some embodiments;
[0062] Figure 30 is a diagram showing a method according to some embodiments Figure 26 A schematic diagram of an exemplary map ranking portion of an AR system;
[0063] Figure 31A is a diagram illustrating examples of area attributes of a Tracking Map (TM) and an environment map in a database according to some embodiments;
[0064] Figure 31B is a diagram illustrating determining a method for Figure 29 A schematic diagram of an example of a geo-location filtered Tracking Map(TM);
[0065] Figure 32 is a diagram showing a method according to some embodiments Figure 29 A schematic diagram of an example of geographic location filtering;
[0066] Figure 33 is a diagram showing a method according to some embodiments Figure 29 A schematic diagram of an example of Wi-Fi BSSID filtering;
[0067] Figure 34 is a diagram illustrating the use of Figure 29 A schematic diagram of an example of positioning;
[0068] Figure 35 and 36 is a block diagram of an XR system configured to rank and merge multiple environment maps, according to some embodiments.
[0069] Figure 37 is a block diagram illustrating a method of creating an environment map of the physical world in a canonical form according to some embodiments;
[0070] Figure 38A and 38B is shown in accordance with some embodiments by updating the tracking map with a new Figure 7 A schematic diagram of an environment map created in a canonical form.
[0071] Figures 39A to 39F is a schematic diagram illustrating an example of a merged map according to some embodiments;
[0072] Figure 40 According to some embodiments, Figure 9 The first XR device to generate the first 3D local tracking map (ground Figure 1 ) two-dimensional representation;
[0073] Figure 41 is an illustration of a flow from a first XR device to a Figure 9 Server upload location Figure 1 Block diagram;
[0074] Figure 42 is a diagram showing a method according to some embodiments Figure 16a schematic diagram of an XR system showing that a second user has initiated a second session using a second XR device of the XR system after the first user has terminated the first session;
[0075] Figure 43A is a diagram illustrating a method for Figure 42 A block diagram of a new session for a second XR device;
[0076] Figure 43B is a diagram illustrating a method for Figure 42 A block diagram of creation of a tracking map for a second XR device;
[0077] Figure 43C is a diagram illustrating a method for transmitting data from a server to a server according to some embodiments. Figure 42 A block diagram of a second XR device downloading a canonical map;
[0078] Figure 44 is to show that according to some embodiments, Figure 42 The second tracking map generated by the second XR device (ground Figure 2 ) a schematic diagram of a positioning attempt to locate the canonical map;
[0079] Figure 45 is a diagram illustrating the Figure 44 Second tracking map (map Figure 2 ) is a schematic diagram of a positioning attempt to locate the canonical map, the second tracking map can be further developed and has the same Figure 2 The XR content associated with the PCF;
[0080] Figures 46A to 46B is a diagram illustrating the Figure 45 land Figure 2 Schematic diagram of successful positioning on the canonical map;
[0081] Figure 47 is a diagram illustrating that according to some embodiments, Figure 46A The canonical map of one or more PCFs includes Figure 45 land Figure 2 Schematic diagram of the canonical map generated by
[0082] Figure 48 is a diagram showing a method according to some embodiments Figure 47 The standard map and the map on the second XR device Figure 2 A schematic diagram of further expansion of
[0083] Figure 49 FIG. 1 is a diagram illustrating uploading a map from a second XR device to a server according to some embodiments. Figure 2 Block diagram;
[0084] Figure 50 is a diagram illustrating how to convert the ground Figure 2 box plots merged with canonical maps;
[0085] Figure 51 is a block diagram illustrating transmitting a new canonical map from a server to a first XR device and a second XR device according to some embodiments;
[0086] Figure 52 is a diagram showing a location according to some embodiments Figure 2 Two-dimensional representation and reference ground Figure 2 a block diagram of a head coordinate frame of a second XR device;
[0087] Figure 53 is a block diagram illustrating in two dimensions adjustments of a head coordinate frame that may occur in six degrees of freedom, according to some embodiments;
[0088] Figure 54 is a block diagram illustrating a canonical map on a second XR device according to some embodiments, wherein sound is relative to ground. Figure 2 The PCF is located;
[0089] Figure 55 and Figure 56 are perspective and block diagrams illustrating use of an XR system when a first user has terminated a first session and the first user has initiated a second session using the XR system, according to some embodiments;
[0090] Figure 57 and Figure 58 is a perspective view and block diagram illustrating use of an XR system when three users are using the XR system simultaneously in the same session, according to some embodiments;
[0091] Figure 59 is a flow chart illustrating a method of recovering and resetting head posture according to some embodiments;
[0092] Figure 60 is a block diagram of a machine in the form of a computer that may find application in the system of the present invention according to some embodiments;
[0093] Figure 61 is a schematic diagram of an example XR system in which any of a plurality of devices can access location services, according to some embodiments;
[0094] Figure 62 is an example process flow for operating a portable device as part of an XR system providing cloud-based positioning, according to some embodiments; and
[0095] Figure 63A 、 Figure 63B and Figure 63Cis an example process flow for cloud-based positioning according to some embodiments.
[0096] Figure 64 、 Figure 65 、 Figure 66 、 Figure 67 and Figure 68 It is a series of schematic diagrams of a portable XR device using multiple units with wireless fingerprints to build a tracking map as a user wearing the XR device traverses a 3D environment.
[0097] Figure 69 Wireless fingerprints are used to select cells in a set of stored maps as candidate cells for positioning construction. Figures 64 to 68 Schematic diagram of a portable XR device that tracks maps.
[0098] Figure 70 is a flow chart illustrating a method of operating a portable XR device to generate a wireless fingerprint, according to some embodiments. DETAILED DESCRIPTION
[0099] Described herein are methods and apparatus for providing an X-reality (XR or cross-reality) scene. To provide a realistic XR experience to multiple users, the XR system must understand the users' physical environment in order to correctly associate the positions of virtual objects with real objects. The XR system can construct an environment map of the scene, which can be created from images and / or depth information collected by sensors that are part of an XR device worn by a user of the XR system.
[0100] In an XR system, each XR device can develop a local map of its physical environment by integrating information from one or more images collected at a point in time during a scan. In some embodiments, the coordinate system of this map is bound to the device's orientation at the time the scan begins. As a user interacts with the XR system, this orientation may vary from session to session, whether different sessions are associated with different users, each user having their own wearable device with sensors that scan the environment, or the same user using the same device at different times.
[0101] However, applications executing on the XR system can specify the location of virtual content based on persistent spatial information, such as can be obtained from a canonical map that can be accessed in a shared manner by multiple users interacting with the XR system. The persistent spatial information can be represented by a persistent map. The persistent map can be stored in a remote storage medium (e.g., the cloud). For example, a wearable device worn by a user, after being turned on, can retrieve a previously created and stored appropriate stored map from persistent storage such as cloud storage. Retrieving the stored map can enable the use of the wearable device without the need to scan the physical world using sensors on the wearable device. Alternatively or additionally, the system / device can similarly retrieve an appropriate stored map when entering a new area of the physical world.
[0102] The stored maps can be represented in a canonical form that can be relative to a local reference frame on each XR device. The relationship between the local map and the canonical map for each device can be determined by a localization process. This localization process can be performed on each XR device based on a set of canonical maps that are selected and sent to the device. Alternatively or additionally, localization services can be performed by a localization service that can be implemented on a remote processor, such as one deployed in the cloud.
[0103] Regardless of where localization is performed, the inventors have recognized and appreciated that effectively selecting one or a small number of canonical maps, or more specifically, one or more segments of a canonical map to utilize for localization attempts, can reduce the computational requirements and latency of the localization process. As a result, localization can be performed faster, more frequently, and / or with lower power consumption, allowing virtual content to be displayed with less latency or more accurately, creating a more immersive user experience.
[0104] In some embodiments, positioning can be made more efficient by storing the wireless fingerprint in association with a persistent location of a canonical map that can be attempted for positioning. The portable device can maintain the wireless fingerprint in association with location information that may be used during positioning. For example, the persistent location of the canonical map can be a persistent coordinate frame that can be used within the XR system to specify the location of virtual content. The persistent location on the portable device can be a persistent posture that can be used to position the portable device relative to the canonical map, for example. Based on the similarity of the wireless fingerprint to the persistent posture in the local map that is closest to the device when positioning is to be performed, the persistent coordinate frame in the canonical map can be selected as a candidate for positioning.
[0105] As a device moves around an area of the physical world, maintaining the wireless fingerprint may require repeated updates of the wireless fingerprint associated with the tracking map. Updating in this manner allows for filtering and averaging of wireless characteristics near the location represented by the persistent gesture within the physical world area. This process can produce a stable and accurate wireless fingerprint, which in turn can lead to more accurate matches with other similarly created wireless fingerprints. Furthermore, this process increases the usability of the wireless fingerprint and reduces the latency associated with the positioning process. With more efficient positioning, the system has a greater ability to share data about the physical world between multiple devices.
[0106] Sharing data about the physical world between multiple devices enables a shared user experience for virtual content. For example, two XR devices accessing the same stored map can both be localized relative to the stored map. Once localized, the user device can render virtual content with a location specified by reference to the stored map by translating the location into a frame of reference maintained by the user device. The user device can use this local reference frame to control the user device's display to render virtual content in the specified location.
[0107] To support these and other functions, the XR system may include components that develop, maintain, and use persistent spatial information (including one or more stored maps) based on data about the physical world collected by sensors on the user device. These components may be distributed across the XR system, for example, with some operating on a head-mounted portion of the user device. Other components may operate on a computer associated with the user coupled to the head-mounted portion via a local area network or personal area network. Still others may operate at a remote location, such as at one or more servers accessible via a wide area network.
[0108] These components may, for example, include components that can identify information about the physical world collected by one or more user devices that is of sufficient quality to be stored as a persistent map or stored in a persistent map. An example of such a component, described in more detail below, is a map merging component. Such a component can, for example, receive input from a user device and determine the suitability of each portion of the input to be used to update the persistent map. The map merging component can, for example, divide a local map created by a user device into multiple parts, determine the mergeability of one or more parts with the persistent map, and merge the parts that meet the qualified mergeability criteria into the persistent map. The map merging component can also, for example, promote parts that were not merged with the persistent map to a separate persistent map.
[0109] As another example, these components can include components that can help determine appropriate persistent maps that can be retrieved and used by the user device. An example of such a component, described in more detail below, is a map ranking component. For example, such a component can receive input from a user device and identify one or more persistent maps that may represent the area of the physical world in which the device operates. For example, the map ranking component can help select a persistent map for the local device to use when rendering virtual content, collecting data about the environment, or performing other actions. Alternatively or additionally, the map ranking component can help identify persistent maps that are updated as one or more user devices collect additional information about the physical world.
[0110] Other components may determine a transformation that converts information captured or described with respect to one reference frame into another reference frame. For example, a sensor may be attached to a head-mounted display so that data read from the sensor indicates the location of an object in the physical world relative to the wearer's head pose. One or more transformations may be applied to relate this position information to a coordinate frame associated with a persistent environment map. Similarly, data indicating where a virtual object is to be rendered when expressed in the coordinate frame of the persistent environment map may be transformed one or more times to be placed in the reference frame of the display on the user's head. As described in more detail below, there may be multiple such transformations. These transformations may be partitioned across components of the XR system so that they can be updated efficiently or applied in a distributed system.
[0111] In some embodiments, a persistent map can be constructed based on information collected by multiple user devices. All or some of the XR devices can capture local spatial information and construct separate tracking maps using information collected by sensors of each XR device at various locations and times. Each tracking map can include points, each of which can be associated with a feature of a real-world object, which can include multiple features. In addition to potentially providing input to create and maintain a persistent map, the tracking map can also be used to track the movement of users in the scene, thereby enabling the XR system to estimate the head pose of the corresponding user based on the tracking map.
[0112] This interdependence between the creation of the map and the estimation of the head pose poses a significant challenge. A lot of processing may be required to create the map and estimate the head pose at the same time. As objects move in the scene (such as moving a cup on a table) and as the user moves in the scene, the processing must be completed quickly because the latency makes the XR experience less realistic for the user. On the other hand, XR devices may provide limited computing resources because the weight of the XR device should be light to be worn comfortably by the user. It is not possible to use more sensors to compensate for the lack of computing resources because adding sensors would also undesirably increase the weight. In addition, more sensors or more computing resources would generate heat, which could cause deformation of the XR device.
[0113] The inventors have recognized and understood XR scenarios for operating an XR system to provide techniques for a more immersive user experience, such as estimating head pose at a frequency of 1 kHz, using lower rates of computing resources associated with the XR device, such as four video graphics array (VGA) cameras that can be configured to operate at 30 Hz, one inertial measurement unit (IMU) operating at 1 kHz, the computing power of a single Advanced RISC Machine (ARM) core, less than 1 GB of memory, and less than 100 Mbps of network bandwidth. These techniques relate to reducing the processing required to generate and maintain maps and estimate head pose, and to providing and consuming data with low computational overhead. The XR system can compute its pose based on matched visual features. U.S. patent application serial number 16 / 221,065 describes hybrid tracking and is incorporated herein by reference in its entirety.
[0114] These techniques may include reducing the amount of data processed when constructing a map, such as by constructing a sparse map using a collection of map building points and keyframes and / or dividing the map into blocks to enable block-by-block updates. Map building points may be associated with points of interest in the environment. Keyframes may include information selected from camera-captured data. U.S. patent application Ser. No. 16 / 520,582 describes determining and / or evaluating a localization map and is incorporated herein by reference in its entirety.
[0115] In some embodiments, persistent spatial information can be represented in a way that can be easily shared between users and between distributed components including applications. For example, information about the physical world can be represented as a persistent coordinate frame (PCF). The PCF can be defined based on one or more points that represent features identified in the physical world. Features can be selected so that they may be the same between user sessions of the XR system. The PCF may be sparse, providing less information than all available information about the physical world so that they can be efficiently processed and transferred. Techniques for processing persistent spatial information may include: creating a dynamic map based on one or more coordinate systems in real space across one or more sessions, and generating a persistent coordinate frame (PCF) on the sparse map, which can be exposed to the XR application via, for example, an application programming interface (API). These functions can be supported by techniques for ranking and merging multiple maps created by one or more XR devices. Persistent spatial information can also enable rapid recovery and resetting of head poses on each of one or more XR devices in a computationally efficient manner.
[0116] Furthermore, these techniques can enable efficient comparison of spatial information. In some embodiments, an image frame can be represented by a digital descriptor. The descriptor can be computed by a transformation that maps a set of features identified in the image to the descriptors. The transformation can be performed in a trained neural network. In some embodiments, the set of features provided as input to the neural network can be a filtered set of features extracted from the image using, for example, a technique that preferentially selects features that are likely to persist.
[0117] Representing image frames as descriptors enables, for example, efficient matching of new image information with stored image information. The XR system may store along with persistent map descriptors for one or more frames underlying a persistent map. Local image frames acquired by a user device may similarly be converted to such descriptors. By selecting stored maps having descriptors similar to the descriptors of the local image frames, one or more persistent maps that are likely to represent the same physical space as the user device may be selected with relatively little processing. In some embodiments, descriptors may be calculated for key frames in both the local map and the persistent map, further reducing processing when comparing maps. For example, such efficient comparisons may be used to simplify finding persistent maps to load into a local device, or finding persistent maps to update based on image information acquired using a local device.
[0118] The technology described herein can be used with or alone on many types of devices and for many types of scenarios, including wearable or portable devices with limited computing resources that provide augmented or mixed reality scenarios. In some embodiments, the technology can be implemented by one or more services that form part of an XR system.
[0119] AR System Overview
[0120] Figure 1 and Figure 2 A scene with virtual content is shown, which is displayed together with a portion of the physical world. For illustration purposes, an AR system is used as an example of an XR system. Figure 3-6B An exemplary AR system is shown that includes one or more processors, memory, sensors, and a user interface that can operate according to the techniques described herein.
[0121] refer to Figure 1 , depicts an outdoor AR scene 354 in which the user of the AR technology sees a park-like setting 356 of the physical world, which features people, trees, buildings in the background, and a concrete platform 358. In addition to these items, the user of the AR technology also perceives that they "see" a robotic statue 357 standing on the physical concrete platform 358, and a flying cartoon-like avatar character 352 that appears to be the head of a bumblebee, even though these elements (e.g., avatar character 352 and robotic statue 357) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is challenging to create AR technology that promotes a comfortable, natural-feeling, rich presentation of virtual image elements among other virtual or physical world image elements.
[0122] Such an AR scene can be implemented by a system that builds a map of the physical world based on tracking information, enables users to place AR content in the physical world, determines the location of the AR content in the map of the physical world, preserves the AR scene so that the placed AR content can be reloaded for display in the physical world during, for example, different AR experience sessions, and enables multiple users to share the AR experience. The system can build and update a digital representation of the physical world surface around the user. This representation can be used to render virtual content to appear to be fully or partially occluded by physical objects between the user and the rendered location of the virtual content, to place virtual objects in physics-based interactions, and for virtual character path planning and navigation, or for other operations in which information about the physical world is used.
[0123] Figure 2Another example of an indoor AR scene 400 is depicted, illustrating an exemplary use case for an XR system, in accordance with some embodiments. The exemplary scene 400 is a living room with a wall, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, the user of the AR technology can also perceive virtual objects, such as an image on the wall behind the sofa, a bird flying through the door, a deer peeking from the bookshelf, and a decorative ornament in the form of a pinwheel placed on the coffee table.
[0124] For an image on a wall, AR technology needs information not only about the surface of the wall, but also about objects and surfaces in the room (such as the shape of a lamp) that would obstruct the image to render the virtual object correctly. For a flying bird, AR technology needs information about all the objects and surfaces around the room in order to render the bird with realistic physics to avoid objects and surfaces or avoid bouncing off the bird when it collides with them. For a deer, AR technology needs information about surfaces (such as the floor or coffee table) to calculate where to place the deer. In the case of a windmill, the system can recognize that it is a separate object from the table and can determine that it is movable, while the corner of a shelf or the corner of a wall can be determined to be stationary. This distinction can be used to determine which parts of the scene to use or update in each of the various operations.
[0125] Virtual objects can be placed in a previous AR experience session. When a new AR experience session starts in the living room, AR technology needs to display the virtual objects accurately in the location where they were previously placed and are actually visible from different perspectives. For example, a windmill should be displayed as standing on a book, rather than floating above the table in a different location without a book. This floating may occur if the user of the new AR experience session is not positioned accurately in the living room. As another example, if the user views the windmill from a different perspective than when the windmill was placed, AR technology needs to display the corresponding side of the windmill.
[0126] The scene can be presented to the user via a system including multiple components, including a user interface that can stimulate one or more user senses (such as vision, sound and / or touch). In addition, the system can include one or more sensors that can measure parameters of the physical part of the scene, including the position and / or movement of the user within the physical part of the scene. In addition, the system can include one or more computing devices, and associated computer hardware, such as memory. These components can be integrated into a single device, or can be distributed across multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.
[0127] Figure 3An AR system 502 is depicted that is configured to provide an experience of AR content that interacts with a physical world 506, in accordance with some embodiments. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by a user as part of a headset, such that the user may wear the display over their eyes like a pair of goggles or glasses. At least a portion of the display may be transparent, such that the user may observe a see-through reality 510. The see-through reality 510 may correspond to a portion of the physical world 506 that is within the current viewpoint of the AR system 502, which may correspond to the user's viewpoint if the user is wearing a headset that incorporates the AR system's display and sensors to obtain information about the physical world.
[0128] AR content may also be presented on display 508, overlaid on see-through reality 510. To provide accurate interaction between AR content and see-through reality 510 on display 508, AR system 502 may include sensors 522 configured to capture information about physical world 506.
[0129] Sensors 522 may include one or more depth sensors that output depth maps 512. Each depth map 512 may have multiple pixels, each of which may represent the distance to a surface in the physical world 506 in a particular direction relative to the depth sensor. Raw depth data may come from the depth sensors to create the depth map. The depth map may be updated as quickly as the depth sensor can form a new image, which may be hundreds or thousands of times per second. However, this data may be noisy and incomplete, with holes in the depth map shown, represented as black pixels.
[0130] The system may include other sensors, such as image sensors. Image sensors may acquire monocular or stereo information, which may be processed to represent the physical world in other ways. For example, images may be processed in world reconstruction component 516 to create a mesh representing connected parts of objects in the physical world. Metadata about such objects (including, for example, color and surface texture) may similarly be acquired using sensors and stored as part of the world reconstruction.
[0131] The system may also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, a head pose tracking component of the system may be used to calculate the head pose in real time. The head pose tracking component may represent the user's head pose in a coordinate system having six degrees of freedom, including, for example, translation about three perpendicular axes (e.g., forward / backward, up / down, left / right) and rotation about the three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, the sensor 522 may include an inertial measurement unit ("IMU") that may be used to calculate and / or determine the head pose 514. The head pose 514 for a depth map may indicate, for example, the current viewpoint of a sensor capturing a depth map in six degrees of freedom, but the headset 514 may be used for other purposes, such as associating image information with a specific part of the physical world or associating the position of a display worn on the user's head with the physical world.
[0132] In some embodiments, head pose information can be derived in other ways than an IMU (such as analyzing objects in an image). For example, the head pose tracking component can calculate the relative position and orientation of the AR device with respect to the physical object based on the visual information captured by the camera and the inertial information captured by the IMU. The head pose tracking component can then calculate the head pose of the AR device, for example, by comparing the calculated relative position and orientation of the AR device with respect to the physical object with features of the physical object. In some embodiments, the comparison can be performed by identifying features in images captured using one or more sensors 522, which are stable over time so that changes in the positions of these features in images captured over time can be associated with changes in the user's head pose.
[0133] In some embodiments, an AR device can build a map based on feature points identified in consecutive images in a series of image frames captured as the user moves through the physical world with the AR device. Although each image frame can be taken from a different posture of the user as they move, the system can adjust the orientation of the features of each consecutive image frame to match the orientation of the initial image frame by matching the features of the consecutive image frame with the previously captured image frame. The translation of consecutive image frames so that points representing the same feature will match corresponding feature points in the previously collected image frame can be used to align each consecutive image frame to match the orientation of the previously processed image frame. The frames in the generated map can have a common orientation established when the first image frame is added to the map. The map has multiple sets of feature points in a common reference frame, and the map can be used to determine the user's posture in the physical world by matching the features in the current image frame with the map. In some embodiments, the map can be referred to as a tracking map.
[0134] In addition to being able to track the user's posture in the environment, the map can also enable other components of the system (such as the world reconstruction component 516) to determine the position of physical objects relative to the user. The world reconstruction component 516 can receive the depth map 512 and head posture 514 and any other data from the sensor and integrate this data into the reconstruction 518. The reconstruction 518 can be more complete and less noisy than the sensor data. The world reconstruction component 516 can update the reconstruction 518 using spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0135] Reconstruction 518 can include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats can represent alternative representations of the same portion of the physical world or can represent different portions of the physical world. In the example shown, on the left side of reconstruction 518, the portion of the physical world is presented as a global surface; on the right side of reconstruction 518, the portion of the physical world is presented as a mesh.
[0136] In some embodiments, the map maintained by the head posture component 514 may be sparse relative to other maps of the physical world that may be maintained. A sparse map may indicate the location of points of interest and / or structures (e.g., corners or edges), rather than providing information about the location of surfaces and possibly other features. In some embodiments, the map may include image frames captured by the sensor 522. These frames may be simplified to include features that may represent points of interest and / or structures. Information about the user's posture from which the frame was captured may also be stored as part of the map in conjunction with each frame. In some embodiments, every image captured by the sensor may or may not be stored. In some embodiments, as images are collected by the sensor, the system may process the images and select a subset of image frames for further computation. This selection may be based on one or more criteria that limit the amount of information added while ensuring that the map contains useful information. The system may, for example, add new image frames to the map based on overlap with previously added image frames or based on image frames containing a sufficient number of features determined to potentially represent stationary objects. In some embodiments, the selected image frames or groups of features from the selected image frames may serve as keyframes for the map, providing spatial information.
[0137] The AR system 502 can integrate sensor data from multiple perspectives of the physical world over time. As the device including the sensor moves, the pose (e.g., position and orientation) of the sensor can be tracked. Since the frame pose of the sensor and its relationship to other poses are known, each of these multiple viewpoints of the physical world can be fused together to form a single combined reconstruction of the physical world, which can be used as an abstract layer for the map and provide spatial information. By using spatial and temporal averaging (i.e., averaging data from multiple viewpoints over time) or any other appropriate method, the reconstruction can be more complete and less noisy than the original sensor data.
[0138] exist Figure 3 In the illustrated embodiment, the map represents a portion of the user's physical world in which a single wearable device is present. In that case, the head pose associated with a frame in the map can be represented as a local head pose, indicating an orientation relative to the initial orientation of the single device at the start of the session. For example, the head pose can be tracked relative to the initial head pose when the device is turned on, or otherwise operated to scan the environment to build a representation of the environment.
[0139] In conjunction with the content representing that portion of the physical world, a map may include metadata. The metadata may, for example, indicate the time at which sensor information used to form the map was captured. Alternatively or in addition, the metadata may indicate the location of the sensor at the time the information used to form the map was captured. The location may be represented directly, such as using information from a GPS chip, or indirectly, such as using a wireless (e.g., Wi-Fi) signature that indicates the strength of the signal received from one or more wireless access points while the sensor data was being collected, and / or using an identifier, such as a BSSID, of the wireless access point to which the user device was connected while the sensor data was being collected.
[0140] Reconstruction 518 can be used for AR functions, such as generating a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation may change as the user moves or objects in the real world change. Aspects of reconstruction 518 can be used, for example, by a component 520 that generates a changing global surface representation in world coordinates, which can be used by other components.
[0141] AR content can be generated based on this information, such as by an AR application 504. The AR application 504 can be, for example, a game program that performs one or more functions, such as visual occlusion, physics-based interaction, and environmental reasoning, based on information about the physical world. It can perform these functions by querying data in different formats from the reconstruction 518 produced by the world reconstruction component 516. In some embodiments, component 520 can be configured to output updates when the representation in the area of interest of the physical world changes. For example, the area of interest can be set to approximate a portion of the physical world near the system user, such as within the user's field of view, or projected (predicted / determined) as entering the user's field of view.
[0142] The AR application 504 can use this information to generate and update AR content. The virtual portion of the AR content can be presented on the display 508 in combination with the see-through reality 510 to create a realistic user experience.
[0143] In some embodiments, an AR experience may be provided to a user via an XR device, which may be a wearable display device, which may be part of a system that may include remote processing and / or remote data storage and / or, in some embodiments, other wearable display devices worn by other users. Figure 4 An example of a system 580 (hereinafter referred to as "system 580") including a single wearable device is shown. System 580 includes a head-mounted display device 562 (hereinafter referred to as "display device 562"), and various mechanical and electronic modules and systems that support the functions of display device 562. Display device 562 can be coupled to a frame 564 that can be worn by a display system user or viewer 560 (hereinafter referred to as "user 560") and is configured to position display device 562 in front of the eyes of user 560. According to various embodiments, display device 562 can display sequentially. Display device 562 can be monocular or binocular. In some embodiments, display device 562 can be Figure 3 Example of display 508 in .
[0144] In some embodiments, a speaker 566 is coupled to the frame 564 and positioned near the ear canal of the user 560. In some embodiments, another speaker, not shown, is positioned near the other ear canal of the user 560 to provide stereo / modifiable sound control. The display device 562 is operably coupled, such as via a wired conductor or wireless connection 568, to a local data processing module 570, which can be mounted in various configurations, such as fixedly attached to the frame 564, fixedly attached to a helmet or hat worn by the user 560, embedded in headphones, or otherwise removably attached to the user 560 (e.g., in a backpack configuration, in a belt-coupled configuration).
[0145] The local data processing module 570 may include a processor and digital storage such as non-volatile memory (e.g., flash memory), both of which may be used to facilitate processing, caching, and storage of data, including: a) data captured from sensors (e.g., which may be operably coupled to the frame 564) or otherwise attached to the user 560, such as an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a radio, and / or a gyroscope; and / or b) data acquired and / or processed using the remote processing module 572 and / or remote data repository 574, and possibly transferred to the display device 562 after such processing or acquisition.
[0146] In some embodiments, the wearable device can communicate with remote components. The local data processing module 570 can be operably coupled to the remote processing module 572 and the remote data repository 574 respectively through communication links 576, 578 (such as via wired or wireless communication links), so that these remote modules 572, 574 are operably coupled to each other and can be used as resources for the local data processing module 570. In further embodiments, in addition to or instead of the remote data repository 574, the wearable device can access a cloud-based remote data repository and / or service. In some embodiments, the above-mentioned head posture tracking component can be implemented at least in part in the local data processing module 570. In some embodiments, Figure 3 The world reconstruction component 516 in can be implemented at least in part in the local data processing module 570. For example, the local data processing module 570 can be configured to execute computer-executable instructions to generate a map and / or physical world representation based at least in part on at least a portion of the data.
[0147] In some embodiments, processing can be distributed across a local processor and a remote processor. For example, local processing can be used to construct a map on the user device (e.g., a tracking map) based on sensor data collected using sensors on the user device. Such a map can be used by applications on the user device. In addition, a previously created map (e.g., a canonical map) can be stored in a remote data repository 574. If an appropriate stored or persistent map is available, it can be used instead of or in addition to a tracking map created locally on the device. In some embodiments, the tracking map can be positioned to a stored map so that a correspondence is established between the tracking map and the canonical map, where the tracking map may be oriented relative to the position of the wearable device when the user turns on the system, and the canonical map may be oriented relative to one or more persistent features. In some embodiments, the persistent map can be loaded on the user device to allow the user device to render virtual content without the delay associated with scanning a location, thereby constructing a tracking map of the user's entire environment based on the sensor data acquired during the scan. In some embodiments, the user device can access a remote persistent map (e.g., stored in the cloud) without having to download the persistent map on the user device.
[0148] In some embodiments, spatial information can be transmitted from the wearable device to a remote service, such as a cloud service configured to locate the device to a stored map maintained on the cloud service. According to one embodiment, the positioning process can be performed in the cloud, matching the device location to an existing map (e.g., a canonical map) and returning a transformation that links virtual content to the wearable device location. In such an embodiment, the system can avoid transmitting the map from the remote resource to the wearable device. Other embodiments can be configured for device-based and cloud-based positioning, for example, to enable functionality where a network connection is unavailable or the user chooses not to enable cloud-based positioning.
[0149] Alternatively or additionally, the tracking map can be merged with previously stored maps to extend or improve the quality of those maps. The process of determining whether a suitable previously created environment map is available and / or merging the tracking map with one or more stored environment maps can be performed in the local data processing module 570 or the remote processing module 572.
[0150] In some embodiments, the local data processing module 570 may include one or more processors (e.g., graphics processing units (GPUs)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the computational budget of the local data processing module 570 but enable a smaller device. In some embodiments, the world reconstruction component 516 may use less than the computational budget of a single Advanced RISC Machine (ARM) core to generate a representation of the physical world in real time over a non-predefined space, leaving the remaining computational budget of the single ARM core accessible for other uses, such as, for example, extracting a mesh.
[0151] In some embodiments, remote data repository 574 may comprise a digital data storage facility accessible via the internet or other networked configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, allowing for fully autonomous use from the remote module. In some embodiments, all data is stored and all or most computations are performed in remote data repository 574, allowing for smaller devices. For example, a world reconstruction may be stored in whole or in part in this repository 574.
[0152] In embodiments where data is stored remotely and accessible over a network, the data can be shared by multiple users of the augmented reality system. For example, user devices can upload their tracking maps to enhance the database of environmental maps. In some embodiments, the tracking map upload occurs at the end of the user session with the wearable device. In some embodiments, the tracking map upload can occur continuously, semi-continuously, or intermittently at a predefined time, after a predefined time period from a previous upload, or when triggered by an event. Whether based on data from that user device or any other user device, the tracking map uploaded by any user device can be used to expand or improve the previously stored map. Similarly, the persistent map downloaded to the user device can be based on data from that user device or any other user device. In this way, users can easily obtain high-quality environmental maps to improve their experience in the AR system.
[0153] In further embodiments, persistent map downloads may be limited and / or avoided based on positioning performed on a remote resource (e.g., in the cloud). In such a configuration, a wearable device or other XR device transmits feature information (e.g., positioning information of the device when the feature represented in the feature information is sensed) in combination with pose information to a cloud service. One or more components of the cloud service may match the feature information with a corresponding stored map (e.g., a canonical map) and generate a transformation between the coordinate systems of a tracking map maintained by the XR device and the canonical map. Each XR device whose tracking map is positioned relative to the canonical map can accurately render virtual content at a location specified relative to the canonical map based on its own tracking.
[0154] In some embodiments, the local data processing module 570 is operably coupled to a battery 582. In some embodiments, the battery 582 is a removable power source, such as a counter battery. In other embodiments, the battery 582 is a lithium-ion battery. In some embodiments, the battery 582 includes both an internal lithium-ion battery that can be charged by the user 560 during non-operating times of the system 580, and a removable battery, allowing the user 560 to operate the system 580 for extended periods of time without having to connect to a power source to charge the lithium-ion battery or shut down the system 580 to replace the battery.
[0155] Figure 5A A user 530 is shown wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as "environment 532"). Information captured by the AR system along the user's movement path can be processed into one or more tracking maps. The user 530 positions the AR display system at a location 534, and the AR display system records environmental information of the traversable world relative to the location 534 (e.g., digital representations of real objects in the physical world, which can be stored and updated as changes are made to the real objects in the physical world). This information can be combined with images, features, directional audio input, or other desired data and stored as a gesture. The location 534 is aggregated to a data input 536, for example, as part of a tracking map, and processed by at least a traversable world module 538, which can, for example, be processed by Figure 4 In some embodiments, the navigable world module 538 may include a head pose component 514 and a world reconstruction component 516 so that the processed information can be combined with other information related to physical objects used in rendering virtual content to indicate the location of the objects in the physical world.
[0156] The navigable world module 538 determines, at least in part, where and how AR content 540, as determined from the data input 536, can be placed in the physical world. AR content is "placed" in the physical world by presenting both the physical world presentation and the AR content via a user interface, with the AR content rendered as if interacting with objects in the physical world, and the objects in the physical world presented as if the AR content obscures the user's view of those objects, when appropriate. In some embodiments, the AR content can be placed by appropriately selecting portions of a fixed element 542 (e.g., a table) from a reconstruction (e.g., reconstruction 518) to determine the shape and position of the AR content 540. As an example, the fixed element can be a table, and the virtual content can be positioned so that it appears to be on the table. In some embodiments, the AR content can be placed within a structure in a field of view 544, which can be a current field of view or an estimated future field of view. In some embodiments, the AR content can be persistently relative to a model 546 (e.g., a grid) of the physical world.
[0157] As depicted, fixed element 542 serves as a proxy (e.g., a digital copy) for any fixed element within the physical world that may be stored in traversable world module 538, such that user 530 can perceive content on fixed element 542 without the system having to map to fixed element 542 each time user 530 sees fixed element 542. Thus, fixed element 542 may be a mesh model from a previous modeling session, or may be determined by a separate user but still stored by traversable world module 538 for future reference by multiple users. Thus, traversable world module 538 can identify environment 532 from a previously mapped environment and display AR content without requiring user 530's device to first map all or a portion of environment 532, thereby saving computational processes and cycles and avoiding any latency in rendering AR content.
[0158] A mesh model 546 of the physical world can be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying AR content 540 can be stored by the traversable world module 538 for future retrieval by the user 530 or other users without having to completely or partially recreate the model. In some embodiments, data input 536 is input such as geographic location, user identification, and current activity to indicate to the traversable world module 538 which of one or more fixed elements 542 is available, which AR content 540 was last placed on the fixed element 542, and whether to display the same content (such AR content is "persistent" regardless of how the user views the particular traversable world model).
[0159] Even in embodiments where objects are considered fixed (e.g., a kitchen table), the traversable world module 538 may update those objects in the physical world model from time to time to account for the possibility of changes in the physical world. The models of fixed objects may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not considered fixed (e.g., a kitchen chair). In order to render a realistic AR scene, the AR system may update the positions of these non-fixed objects at a much higher frequency than the frequency used to update fixed objects. In order to accurately track all objects in the physical world, the AR system may obtain information from multiple sensors (including one or more image sensors).
[0160] Figure 5B is a schematic diagram of viewing optics assembly 548 and accompanying components.In some embodiments, two eye tracking cameras 550 directed at the user's eyes 549 detect metrics of the user's eyes 549, such as eye shape, eyelid occlusion, pupil direction, and glint on the user's eyes 549.
[0161] In some embodiments, one of the sensors may be a depth sensor 551, such as a time-of-flight sensor, which transmits signals to the world and detects reflections of those signals from nearby objects to determine the distance to a given object. The depth sensor can quickly determine whether an object has entered the user's field of view, for example, due to the motion of those objects or a change in the user's posture. However, information about the position of an object in the user's field of view may alternatively or additionally be collected by other sensors. Depth information can be obtained, for example, from a stereoscopic image sensor or a plenoptic sensor.
[0162] In some embodiments, the world camera 552 records a view larger than the periphery to map the environment 532 and / or otherwise create a model of the environment 532 and detect inputs that can affect the AR content. In some embodiments, the world camera 552 and / or the camera 553 can be a grayscale and / or color image sensor that can output grayscale and / or color image frames at fixed time intervals. The camera 553 can further capture an image of the physical world within the user's field of view at a particular time. Even if the value of the pixels of the frame-based image sensor does not change, its pixels can be repeatedly sampled. Each of the world camera 552, camera 553 and depth sensor 551 has a corresponding field of view 554, 555 and 556 to obtain a view from, for example, Figure 34 Data is collected in the physical world scene of the physical world environment 532 depicted in A and the physical world scene is recorded.
[0163] Inertial measurement unit 557 can determine the motion and orientation of viewing optical assembly 548. In some embodiments, each component is operably coupled to at least one other component. For example, depth sensor 551 can be operably coupled to eye tracking camera 550 to confirm the measured accommodation relative to the actual distance at which the user's eye 549 is looking.
[0164] It should be understood that the viewing optics assembly 548 may include Figure 34 B, and may include components instead of or in addition to the components shown. For example, in some embodiments, viewing optics assembly 548 may include two world cameras 552 instead of four. Alternatively or in addition, cameras 552 and 553 need not capture visible light images of their entire fields of view. Viewing optics assembly 548 may include other types of components. In some embodiments, viewing optics assembly 548 may include one or more dynamic vision sensors (DVS) whose pixels may asynchronously respond to relative changes in light intensity exceeding a threshold.
[0165] In some embodiments, based on the time-of-flight information, the viewing optical assembly 548 may not include the depth sensor 551. For example, in some embodiments, the viewing optical assembly 548 may include one or more plenoptic cameras whose pixels can capture light intensity and the angle of incident light, from which depth information can be determined. For example, the plenoptic camera may include an image sensor covered with a transmissive diffraction mask (TDM). Alternatively or in addition, the plenoptic camera may include an image sensor that includes angle-sensitive pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors may be used as a source of depth information instead of or in addition to the depth sensor 551.
[0166] It should also be understood that Figure 5B The configuration of components in FIG. 5 is provided as an example. Viewing optics assembly 548 may include components having any suitable configuration that can be configured to provide the user with the maximum field of view practicable for a particular set of components. For example, if viewing optics assembly 548 includes a world camera 552, the world camera may be placed in a central region of the viewing optics assembly rather than on the side.
[0167] Information from the sensors in the viewing optical assembly 548 can be coupled to one or more processors in the system. The processor can generate data that can be rendered so that the user perceives virtual content interacting with objects in the physical world. This rendering can be achieved in any suitable manner, including generating image data that depicts both physical and virtual objects. In other embodiments, physical and virtual content can be depicted in one scene by modulating the opacity of a display device that the user is browsing in the physical world. The opacity can be controlled to create the appearance of virtual objects and also prevent the user from seeing objects in the physical world that are obscured by virtual objects. In some embodiments, when viewed through a user interface, the image data can include only virtual content, which can be modified so that the virtual content is perceived by the user as realistically interacting with the physical world (e.g., clipping the content to account for occlusion).
[0168] The position of displayed content on viewing optical assembly 548 to create the impression that an object is located at a particular location can depend on the physical properties of the viewing optical assembly. Furthermore, the position of the user's head relative to the physical world and the direction of the user's eye gaze can influence how the location of displayed content in the physical world will appear at a particular location on the viewing optical assembly. Sensors, such as those described above, can collect this information and / or provide information from which this information can be calculated, such that a processor receiving sensor input can calculate the location at which an object should be rendered on viewing optical assembly 548 to create the desired appearance for the user.
[0169] Regardless of how content is presented to the user, a model of the physical world can be used so that properties of virtual objects that can be affected by physical objects can be correctly calculated, including the shape, position, motion, and visibility of the virtual objects. In some embodiments, the model can include a reconstruction of the physical world, such as reconstruction 518.
[0170] The model may be created from data collected from sensors on a user's wearable device. However, in some embodiments, the model may be created from data collected from multiple users, which may be aggregated in a computing device remote from all users (and this data may be in the "cloud").
[0171] The model may be created at least in part by a world reconstruction system, such as, for example, Figure 6A Described in more detail in Figure 3The world reconstruction component 516 may include a perception module 660 that can generate, update, and store a representation of a portion of the physical world. In some embodiments, the perception module 660 can represent the portion of the physical world within the reconstruction range of the sensor as a plurality of voxels. Each voxel can correspond to a 3D cube of a predetermined volume in the physical world and include surface information that indicates whether a surface exists in the volume represented by the voxel. Voxels can be assigned a value that indicates whether their corresponding volume has been determined to include a surface of a physical object, determined to be empty, or has not yet been measured with the sensor and therefore its value is unknown. It should be understood that it is not necessary to explicitly store values indicating voxels that are determined to be empty or unknown, as the values of voxels can be stored in computer memory in any suitable manner, including not storing information for voxels that are determined to be empty or unknown.
[0172] In addition to generating information for the persistent world representation, the perception module 660 can also identify and output indications of changes in the area surrounding the user of the AR system. Such indications of changes can trigger updates to volumetric data stored as part of the persistent world, or trigger other functions, such as triggering the generation of AR content to update the trigger component 604 of the AR content.
[0173] In some embodiments, the perception module 660 can identify changes based on a signed distance function (SDF) model. The perception module 660 can be configured to receive sensor data such as, for example, a depth map 660a and a head pose 660b, and then fuse the sensor data into an SDF model 660c. The depth map 660a can directly provide SDF information, and the image can be processed to obtain the SDF information. The SDF information represents the distance from the sensors used to capture the information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and therefore from the perspective of the user. The head pose 660b can enable the SDF information to be correlated with voxels in the physical world.
[0174] In some embodiments, the perception module 660 can generate, update, and store a representation of the portion of the physical world within a perception range. The perception range can be determined at least in part based on a reconstruction range of the sensor, which can be determined at least in part based on a limit of the sensor's observation range. As a specific example, an active depth sensor operating using active IR pulses can reliably operate within a range of distances, thereby creating a sensor's observation range, which can range from a few centimeters or tens of centimeters to several meters.
[0175] The world reconstruction component 516 may include additional modules that can interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data acquired by the perception module 660. The persistent world module 662 may also include representations of the physical world in various formats. For example, volume metadata such as voxels 662b, as well as meshes 662c and planes 662d, may be stored. In some embodiments, other information such as a depth map may also be stored.
[0176] In some embodiments, a representation of the physical world (such as Figure 6A The representation shown in ) can provide relatively dense information about the physical world compared to sparse maps (such as the feature point-based tracking map described above).
[0177] In some embodiments, the perception module 660 may include modules that generate representations of the physical world in various formats, including, for example, grids 660d, planes, and semantics 660e. Representations of the physical world can be stored across local and remote storage media. Depending on, for example, the location of the storage media, representations of the physical world can be described in different coordinate frames. For example, a representation of the physical world stored in a device can be described in a coordinate frame local to the device. A representation of the physical world can have a counterpart stored in the cloud. The corresponding representation in the cloud can be described in a coordinate frame shared by all devices in the XR system.
[0178] In some embodiments, these modules can generate representations based on data within the perception range of one or more sensors at the time the representation is generated, as well as data captured at previous times and information in the persistent world module 662. In some embodiments, these components can operate based on depth information captured using a depth sensor. However, the AR system can include a visual sensor and can generate such representations by analyzing monocular or binocular visual information.
[0179] In some embodiments, these modules can operate on regions of the physical world. When perception module 660 detects a change in the physical world in a subregion of the physical world, these modules can be triggered to update the subregion of the physical world. For example, such a change can be detected by detecting a new surface or other criteria in SDF model 660c (e.g., a change in the value of a sufficient number of voxels representing the subregion).
[0180] World reconstruction component 516 may include component 664 that can receive a representation of the physical world from perception module 660. Information about the physical world can be extracted by these components based on, for example, usage requests from applications. In some embodiments, information can be pushed to the usage components, such as by indicating changes in pre-identified areas or changes in the representation of the physical world within the perception range. Component 664 may include, for example, game programs and other components that perform processing for visual occlusion, physics-based interaction, and environmental reasoning.
[0181] In response to a query from component 664, perception module 660 can send a representation of the physical world in one or more formats. For example, when component 664 indicates that the usage is for visual occlusion or physics-based interaction, perception module 660 can send a representation of a surface. When component 664 indicates that the usage is for environmental reasoning, perception module 660 can send meshes, planes, and semantics of the physical world.
[0182] In some embodiments, perception module 660 may include a component that formats information to provide component 664. An example of such a component may be ray casting component 660f. Using a component (e.g., component 664), for example, information about the physical world can be queried from a particular viewpoint. Ray casting component 660f can select from one or more representations of physical world data within the field of view from that viewpoint.
[0183] As will be appreciated from the above description, the perception module 660 or another component of the AR system can process data to create a 3D representation of a portion of the physical world. The data to be processed can be reduced by: culling portions of the 3D reconstruction volume based at least in part on the camera frustum and / or depth image; extracting and retaining planar data; capturing, retaining, and updating 3D reconstruction data in blocks that allow local updates while maintaining neighbor consistency; providing occlusion data to an application generating such a scene, wherein the occlusion data is derived from a combination of one or more depth data sources; and / or performing multi-stage mesh simplification. The reconstruction can contain data of varying complexity, including, for example, raw data (e.g., real-time depth data), fused volumetric data (e.g., voxels), and computed data (e.g., meshes).
[0184] In some embodiments, components of the traversable world model can be distributed, with some portions executing locally on the XR device and some portions executing remotely, such as on a network-connected server or in the cloud. The distribution of information processing and storage between the local XR device and the cloud can impact the functionality and user experience of the XR system. For example, reducing processing on the local device by offloading it to the cloud can extend battery life and reduce heat generated on the local device. However, offloading excessive processing to the cloud can introduce undesirable latency, resulting in an unacceptable user experience.
[0185] Figure 6B A distributed component architecture 600 configured for spatial computing according to some embodiments is depicted. The distributed component architecture 600 may include a traversable world component 602 (e.g., Figure 5A 6), LuminOS 604, API 606, SDK 608, and applications 610. LuminOS 604 may include a Linux-based kernel with custom drivers compatible with XR devices. API 606 may include an application programming interface that allows XR applications (e.g., application 610) to access the spatial computing features of the XR device. SDK 608 may include a software development kit that allows the creation of XR applications.
[0186] One or more components of the architecture 600 may create and maintain a model of the navigable world. In this example, sensor data is collected on a local device. Processing of this sensor data may be performed partially locally on the XR device and partially in the cloud. PW 538 may include a map of the environment created based at least in part on data captured by AR devices worn by multiple users. During a session of an AR experience, each AR device (such as the one described above in combination) may be used to create a map of the environment. Figure 4 The wearable device described herein can create a tracking map, which is a type of map.
[0187] In some embodiments, the device may include components for building sparse maps and dense maps. A tracking map may serve as a sparse map and may include the head poses of the AR device scanning the environment and information about objects detected within the environment at each head pose. Those head poses may be maintained locally for each device. For example, the head pose on each device may be an initial head pose relative to when the device started its session. As a result, each tracking map may be local to the device that created it. A dense map may include surface information, which may be represented by a mesh or depth information. Alternatively or in addition, a dense map may include higher-level information derived from the surface or depth information, such as the location and / or features of planes and / or other objects.
[0188] In some embodiments, the creation of dense maps can be independent of the creation of sparse maps. For example, the creation of dense maps and sparse maps can be performed in separate processing pipelines within the AR system. For example, the separate processing can enable the generation or processing of different types of maps to be performed at different rates. For example, the refresh rate of a sparse map may be faster than the refresh rate of a dense map. However, in some embodiments, the processing of dense maps and sparse maps may be related even if performed in different pipelines. For example, a change in the physical world revealed in a sparse map can trigger an update to a dense map, and vice versa. In addition, even if created independently, these maps can be used together. For example, a coordinate system derived from a sparse map can be used to define the position and / or orientation of objects in a dense map.
[0189] Sparse maps and / or dense maps may be persisted for reuse by the same device and / or shared with other devices. Such persistence may be achieved by storing the information in the cloud. The AR device may send a tracking map to the cloud to be merged, for example, with an environment map selected from a persistent map previously stored in the cloud. In some embodiments, the selected persistent map may be sent from the cloud to the AR device for merging. In some embodiments, the persistent maps may be oriented relative to one or more persistent coordinate frames. Such maps may be used as canonical maps in that they may be used by any of a plurality of devices. In some embodiments, a model of the navigable world may include or be created from one or more canonical maps. Even if some operations are performed based on a coordinate frame local to the device, the device may use a canonical map by determining a transformation between the coordinate frame local to the device and the canonical map.
[0190] A canonical map may be derived from a tracking map (TM) (e.g. Figure 31A The canonical map may be stored persistently so that a device accessing the canonical map, once it has determined the transformation between its local coordinate system and the coordinate system of the canonical map, can use the information in the canonical map to determine the location of objects represented in the canonical map in the physical world around the device. In some embodiments, the TM may be a sparse map of head pose created by the XR device. In some embodiments, the canonical map may be created when the XR device sends one or more TMs to a cloud server to be merged with additional TMs captured by the XR device at a different time or by other XR devices.
[0191] A canonical map or other map may provide information about the portions of the physical world represented by the data processed to create the corresponding map. Figure 7An exemplary tracking map 700 is depicted in accordance with some embodiments. The tracking map 700 may provide a plan view 706 of physical objects in the physical world represented by points 702. In some embodiments, the map points 702 may represent features of a physical object which may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. These features may be derived by processing an image, such as an image that may be acquired by a sensor of a wearable device in an augmented reality system. For example, features may be derived by processing image frames output by a sensor to identify features based on large gradients in the image or other appropriate criteria. Further processing may limit the number of features in each frame. For example, the processing may select features that may represent persistent objects. One or more heuristics may be applied to the selection.
[0192] The tracking map 700 may include data about points 702 collected by the device. For each image frame having data points included in the tracking map, a pose may be stored. The pose may represent the orientation from which the image frame was captured, so that feature points within each image frame may be spatially correlated. The pose may be determined by positioning information, such as may be derived by sensors on the wearable device (such as an IMU sensor). Alternatively or additionally, the pose may be determined by matching the image frame to other image frames depicting overlapping portions of the physical world. By looking for such positional correlations, which may be achieved by matching subsets of feature points in two frames, a relative pose between the two frames may be calculated. Relative poses may be sufficient for a tracking map because the map may be relative to a coordinate system local to the device that is established based on the initial pose of the device when construction of the tracking map begins.
[0193] Not all feature points and image frames collected by the device can be retained as part of the tracking map, as much of the information collected with the sensors is likely to be redundant. Instead, only certain frames can be added to the map. Those frames can be selected based on one or more criteria, such as the degree of overlap with image frames already in the map, the number of new features they contain, or a quality measure of the features in the frame. Image frames that are not added to the tracking map can be discarded or used to modify the location of the features. As another alternative, all or most of the image frames represented as a set of features can be retained, but a subset of these frames can be designated as key frames for further processing.
[0194] The keyframes may be processed to generate a keyrig 704. The keyframes may be processed to generate a three-dimensional set of feature points and saved as the keyrig 704. For example, such processing may require comparing image frames simultaneously acquired from two cameras to stereoscopically determine the 3D positions of the feature points. Metadata may be associated with these keyframes and / or keyrigs (e.g., poses).
[0195] The environment map can have any of a variety of formats depending on, for example, where the environment map is stored, including, for example, local storage of the AR device and remote storage. For example, on a wearable device with limited memory, the map in the remote storage may have a higher resolution than the map in the local storage. In order to send the higher resolution map from the remote storage to the local storage, the map can be downsampled or otherwise converted to an appropriate format, such as by reducing the number of poses for each area of the physical world stored in the map and / or the number of feature points stored for each pose. In some embodiments, slices or portions of the high-resolution map from the remote storage can be sent to the local storage, where the slices or portions are not downsampled.
[0196] When a new tracking map is created, a database of environment maps may be updated. To determine which of a potentially large number of environment maps in the database to update, the update may include effectively selecting one or more environment maps stored in the database that are related to the new tracking map. The selected one or more environment maps may be ranked by relevance, and one or more of the highest-ranked maps may be selected for processing to merge the higher-ranked selected environment maps with the new tracking map to create one or more updated environment maps. When the new tracking map represents a portion of the physical world for which no pre-existing environment map is to be updated, the tracking map may be stored in the database as a new environment map.
[0197] Watch standalone display
[0198] Methods and apparatus are described herein for providing virtual content using an XR system that is independent of the position of the eyes viewing the virtual content. Traditionally, virtual content is re-rendered upon any movement of the display system. For example, if a user wearing a display system views a virtual representation of a three-dimensional (3D) object on a display and walks around the area where the 3D object appears, the 3D object should be re-rendered for each viewpoint so that the user has the perception that he or she is walking around the object occupying real space. However, re-rendering consumes a significant amount of the system's computational resources and can result in artifacts due to latency.
[0199] The inventors have recognized and appreciated that head pose (e.g., the position and orientation of a user wearing an XR system) can be used to render virtual content that is independent of eye rotation within the user's head. In some embodiments, a dynamic map of a scene can be generated based on multiple coordinate frames in real space across one or more sessions, such that virtual content that interacts with the dynamic map can be robustly rendered, independent of eye rotation within the user's head and / or independent of sensor deformation caused by, for example, heat generated during high-speed, computationally intensive operations. In some embodiments, the configuration of multiple coordinate frames can enable a first XR device worn by a first user and a second XR device worn by a second user to identify a common location in a scene. In some embodiments, the configuration of multiple coordinate frames can enable users wearing XR devices to view virtual content at the same location in a scene.
[0200] In some embodiments, a tracking map may be constructed in a world coordinate frame, which may have a world origin. When the XR device is powered on, the world origin may be the first pose of the XR device. The world origin may be aligned with gravity, allowing developers of XR applications to perform gravity alignment without additional work. Different tracking maps may be constructed in different world coordinate frames because tracking maps may be captured by the same XR device in different sessions and / or by different XR devices worn by different users. In some embodiments, a session of an XR device may start from when the device is powered on to when the device is powered off. In some embodiments, the XR device may have a head coordinate frame, which may have a head origin. The head origin may be the current pose of the XR device when the image is captured. The difference between the head pose of the world coordinate frame and the head pose of the head coordinate frame may be used to estimate the tracking route.
[0201] In some embodiments, an XR device may have a camera coordinate frame that may have a camera origin. The camera origin may be the current pose of one or more sensors of the XR device. The inventors have recognized and appreciated that the configuration of the camera coordinate frame enables robust display of virtual content independent of eye rotation within the user's head. This configuration also enables robust display of virtual content independent of sensor deformation, such as due to heat generated during operation.
[0202] In some embodiments, an XR device may have a head unit having a head-mounted frame that a user can secure to their head and may include two waveguides, one in front of each eye of the user. The waveguides may be transparent so that ambient light from real-world objects can be transmitted through the waveguides and the user can see the real-world objects. Each waveguide may send projected light from a projector to a corresponding eye of the user. The projected light may form an image on the retina of the eye. Thus, the retina of the eye receives both the ambient light and the projected light. The user may simultaneously see the real-world objects and one or more virtual objects created by the projected light. In some embodiments, the XR device may have sensors that detect real-world objects around the user. These sensors may, for example, be cameras that capture images that can be processed to identify the location of the real-world objects.
[0203] In some embodiments, rather than attaching virtual content to a world coordinate frame, the XR system can assign a coordinate frame to the virtual content. Such a configuration enables description of virtual content without regard to where the virtual content is rendered to the user, but the virtual content can be attached to a more persistent frame location, such as a coordinate frame that will be rendered at a specified location. Figures 14 to 20C As the position of an object changes, the XR device can detect the change in the environment map and determine the motion of the head unit worn by the user relative to the real-world object.
[0204] Figure 8 1 is a diagram illustrating a user in a physical environment experiencing virtual content rendered by an XR system 10, according to some embodiments. The XR system may include a first XR device 12.1 worn by a first user 14.1, a network 18, and a server 20. The user 14.1 is in a physical environment with real objects in the form of a table 16.
[0205] In the example shown, the first XR device 12.1 includes a head unit 22, a waist pack 24, and a cable connection 26. The first user 14.1 secures the head unit 22 to their head and secures the waist pack 24, which is remote from the head unit 22, to their waist. The cable connection 26 connects the head unit 22 to the waist pack 24. The head unit 22 includes technology for displaying one or more virtual objects to the first user 14.1 while allowing the first user 14.1 to see real objects such as a table 16. The waist pack 24 primarily includes the processing and communication capabilities of the first XR device 12.1. In some embodiments, the processing and communication capabilities can reside in whole or in part in the head unit 22, so that the waist pack 24 can be removed or can be located in another device such as a backpack.
[0206] In the example shown, a belt pack 24 is connected to a network 18 via a wireless connection. A server 20 is connected to the network 18 and maintains data representing local content. The belt pack 24 downloads data representing local content from the server 20 via the network 18. The belt pack 24 provides the data to the head unit 22 via a cable connection 26. The head unit 22 may include a display having a light source (e.g., a laser light source or a light emitting diode (LED) light source) and a waveguide to guide the light.
[0207] In some embodiments, first user 14.1 may attach head unit 22 to their head and waist pack 24 to their waist. Waist pack 24 may download image data from server 20 via network 18. First user 14.1 may view table 16 through the display of head unit 22. A projector forming part of head unit 22 may receive image data from waist pack 24 and generate light based on the image data. The light may travel through one or more waveguides forming part of the display of head unit 22. The light may then exit the waveguides and propagate onto the retinas of first user 14.1's eyes. The projector may generate light in a pattern that is replicated on the retinas of first user 14.1's eyes. The light that strikes the retinas of first user 14.1's eyes may have a selected depth of field, allowing first user 14.1 to perceive an image at a preselected depth behind the waveguides. Furthermore, first user 14.1's two eyes may receive slightly different images, allowing first user 14.1's brain to perceive one or more three-dimensional images at a selected distance from head unit 22. In the example shown, the first user 14.1 perceives virtual content 28 above the table 16. The scale of the virtual content 28 and its position and distance from the first user 14.1 are determined by the data representing the virtual content 28 and the various coordinate frames used to display the virtual content 28 to the first user 14.1.
[0208] In the example shown, the virtual content 28 is invisible from the perspective of the drawing and is visible to the first user 14.1 using the first XR device 12.1. The virtual content 28 may initially reside as a data structure within the visual data and algorithms in the waist pack 24. Then, when the projector of the head unit 22 generates light based on the data structure, the data structure may manifest itself as light. It should be understood that although the virtual content 28 does not exist in the three-dimensional space in front of the first user 14.1, the virtual content 28 is still represented in the three-dimensional space. Figure 1 , to illustrate wearer perception of head unit 22. Visualizations of computer data in three-dimensional space may be used in this description to illustrate how data structures contributing to the rendering as perceived by one or more users relate to one another within the data structures in waist pack 24.
[0209] Figure 9Components of a first XR device 12.1 are shown in accordance with some embodiments. The first XR device 12.1 may include a head unit 22, and various components that form part of the visual data and algorithms, including, for example, a rendering engine 30, various coordinate frames 32, various origin and destination coordinate frames 34, and various origin-to-destination coordinate frame transformers 36. The various coordinate frames may be based on intrinsic properties of the XR device or may be determined by reference to other information, such as a persistent pose or persistent coordinate system as described herein.
[0210] Head unit 22 may include a head-mounted frame 40 , a display system 42 , a real object detection camera 44 , a motion tracking camera 46 , and an inertial measurement unit 48 .
[0211] The head-mounted frame 40 may have a Figure 8 The display system 42 , real object detection camera 44 , motion tracking camera 46 , and inertial measurement unit 48 may be mounted to the head mounted frame 40 and thus move with the head mounted frame 40 .
[0212] Coordinate system 32 may include a local data system 52 , a world frame system 54 , a head frame system 56 , and a camera frame system 58 .
[0213] The local data system 52 may include a data channel 62, a local frame determination routine 64, and local frame storage instructions 66. The data channel 62 may be an internal software routine, a hardware component such as an external cable or radio frequency receiver, or a hybrid component such as an open port. The data channel 62 may be configured to receive image data 68 representing virtual content.
[0214] A local frame determination routine 64 can be connected to the data channel 62. The local frame determination routine 64 can be configured to determine a local coordinate frame 70. In some embodiments, the local frame determination routine can determine the local coordinate frame based on a real-world object or real-world location. In some embodiments, the local coordinate frame can be based on a top edge relative to a bottom edge of the browser window, a head or foot of a character, a node on the outer surface of a prism or bounding box surrounding the virtual content, or any other suitable location for placing a coordinate frame that defines the facing direction of the virtual content and the location where the virtual content is placed (e.g., a node, such as a placement node or a PCF node).
[0215] The local frame storage instructions 66 may be connected to the local frame determination routine 64. Those skilled in the art will appreciate that software modules and routines are "connected" to each other through subroutines, calls, and the like. The local frame storage instructions 66 may store the local coordinate frame 70 as a local coordinate frame 72 within the origin and destination coordinate frame 34. In some embodiments, the origin and destination coordinate frames 34 may be one or more coordinate frames that may be manipulated or transformed to allow virtual content to persist between sessions. In some embodiments, a session may be a period of time between startup and shutdown of an XR device. Two sessions may be two startup and shutdown periods of a single XR device, or startup and shutdown periods of two different XR devices.
[0216] In some embodiments, the origin and destination coordinate frames 34 may be coordinate frames involved in one or more transformations required to enable the first user's XR device and the second user's XR device to recognize a common location. In some embodiments, the destination coordinate frame may be the output of a series of calculations and transformations applied to the target coordinate frame so that the first and second users view virtual content in the same location.
[0217] The rendering engine 30 may be connected to the data channel 62. The rendering engine 30 may receive image data 68 from the data channel 62 so that the rendering engine 30 may render virtual content based at least in part on the image data 68.
[0218] The display system 42 may be connected to the rendering engine 30. The display system 42 may include components that convert the image data 68 into visible light. The visible light may be formed into two patterns, one for each eye. The visible light may enter Figure 8 in the eye of the first user 14 . 1 and can be detected on the retina of the eye of the first user 14 . 1 .
[0219] The real object detection camera 44 may include one or more cameras that can capture images from different sides of the head-mounted frame 40. The motion tracking camera 46 may include one or more cameras that can capture images on the sides of the head-mounted frame 40. Instead of two sets of one or more cameras representing the real object detection camera 44 and the motion tracking camera 46, a single set of one or more cameras may be used. In some embodiments, the cameras 44 and 46 can capture images. As described above, these cameras can collect data used to construct a tracking map.
[0220] Inertial measurement unit 48 may include multiple devices for detecting the motion of head unit 22. Inertial measurement unit 48 may include a gravity sensor, one or more accelerometers, and one or more gyroscopes. The sensors of inertial measurement unit 48, in combination, track the motion of head unit 22 in at least three orthogonal directions and about at least three orthogonal axes.
[0221] In the example shown, the world frame system 54 includes a world surface determination routine 78, a world frame determination routine 80, and a world frame storage instruction 82. The world surface determination routine 78 is connected to the real object detection camera 44. The world surface determination routine 78 accepts images and / or keyframes based on images captured by the real object detection camera 44 and processes the images to identify surfaces in the images. A depth sensor (not shown) can determine the distance to the surfaces. Therefore, these surfaces are represented by data in three dimensions including their size, shape, and distance from the real object detection camera.
[0222] In some embodiments, the world coordinate frame 84 can be based on the origin when the head gesture session is initialized. In some embodiments, the world coordinate frame can be located at the location where the device was started, or if the head gesture is lost during the startup session, the world coordinate frame can be located in a new place. In some embodiments, the world coordinate frame can be the origin when the head gesture session begins.
[0223] In the example shown, a world frame determination routine 80 is coupled to the world surface determination routine 78 and determines a world coordinate frame 84 based on the position of the surface determined by the world surface determination routine 78. World frame storage instructions 82 are coupled to the world frame determination routine 80 to receive the world coordinate frame 84 from the world frame determination routine 80. The world frame storage instructions 82 store the world coordinate frame 84 as a world coordinate frame 86 within the origin and destination coordinate frame 34.
[0224] The head frame system 56 may include a head frame determination routine 90 and head frame storage instructions 92. The head frame determination routine 90 may be connected to the motion tracking camera 46 and the inertial measurement unit 48. The head frame determination routine 90 may use data from the motion tracking camera 46 and the inertial measurement unit 48 to calculate a head coordinate frame 94. For example, the inertial measurement unit 48 may have a gravity sensor that determines the direction of gravity relative to the head unit 22. The motion tracking camera 46 may continuously capture images that are used by the head frame determination routine 90 to refine the head coordinate frame 94. Figure 8 When the first user 14 . 1 in FIG. 1 moves their head, the head unit 22 moves. The motion tracking camera 46 and the inertial measurement unit 48 may continuously provide data to the head frame determination routine 90 so that the head frame determination routine 90 may update the head coordinate frame 94 .
[0225] Head frame storage instructions 92 may be coupled to head frame determination routine 90 to receive head coordinate frame 94 from head frame determination routine 90. Head frame storage instructions 92 may store head coordinate frame 94 as head coordinate frame 96 in origin and destination coordinate frames 34. Head frame storage instructions 92 may repeatedly store updated head coordinate frame 94 as head coordinate frame 96 whenever head frame determination routine 90 recalculates head coordinate frame 94. In some embodiments, the head coordinate frame may be the position of wearable XR device 12.1 relative to local coordinate frame 72.
[0226] The camera frame system 58 may include camera intrinsics 98. The camera intrinsics 98 may include the dimensions of the head unit 22 as a feature of its design and manufacture. The camera intrinsics 98 may be used to calculate a camera coordinate frame 100 that is stored within the origin and destination coordinate frame 34.
[0227] In some embodiments, the camera coordinate frame 100 may include Figure 8 1. As the left eye moves from left to right or up and down, the pupil position of the left eye is located within the camera coordinate frame 100. Additionally, the pupil position of the right eye is located within the camera coordinate frame 100 for the right eye. In some embodiments, the camera coordinate frame 100 may include the position of the camera relative to the local coordinate frame when the image was captured.
[0228] The origin-to-destination coordinate frame transformer 36 may include a local-to-world coordinate transformer 104, a world-to-head coordinate transformer 106, and a head-to-camera coordinate transformer 108. The local-to-world coordinate transformer 104 may receive the local coordinate frame 72 and transform the local coordinate frame 72 into the world coordinate frame 86. The transformation of the local coordinate frame 72 to the world coordinate frame 86 may be represented as a transformation of the local coordinate frame within the world coordinate frame 86 into the world coordinate frame 110.
[0229] The world-to-head coordinate transformer 106 may transform from the world coordinate frame 86 to the head coordinate frame 96. The world-to-head coordinate transformer 106 may transform the local coordinate frame transformed to the world coordinate frame 110 to the head coordinate frame 96. This transformation may be represented as a local coordinate frame transformed within the head coordinate frame 96 to the head coordinate frame 112.
[0230] The head-to-camera coordinate transformer 108 may transform from the head coordinate frame 96 to the camera coordinate frame 100. The head-to-camera coordinate transformer 108 may transform the local coordinate frame transformed to the head coordinate frame 112 to a local coordinate frame transformed to the camera coordinate frame 114 within the camera coordinate frame 100. The local coordinate frame transformed to the camera coordinate frame 114 may be input to the rendering engine 30. The rendering engine 30 may render the image data 68 representing the local content 28 based on the local coordinate frame transformed to the camera coordinate frame 114.
[0231] Figure 10 is a spatial representation of the various origin and destination coordinate frames 34. In this figure, a local coordinate frame 72, a world coordinate frame 86, a head coordinate frame 96, and a camera coordinate frame 100 are shown. In some embodiments, when virtual content is placed in the real world so that a user can view the virtual content, the local coordinate frame associated with the XR content 28 can have a position and rotation relative to the local and / or world coordinate frames and / or PCF (e.g., a node and a facing direction can be provided). Each camera can have its own camera coordinate frame 100 that contains all pupil positions of one eye. Reference numerals 104A and 106A respectively represent the camera coordinate frames 104 and 106A represented by Figure 9 The transformations performed by the local to world coordinate transformer 104, the world to head coordinate transformer 106 and the head to camera coordinate transformer 108 in FIG.
[0232] Figure 11 A camera rendering protocol for transforming from a head coordinate frame to a camera coordinate frame, according to some embodiments, is depicted. In the example shown, the pupil of a single eye moves from position A to position B. Virtual objects that appear stationary will be projected onto a depth plane at one of two locations, A or B, depending on the location of the pupil (assuming the camera is configured to use a pupil-based coordinate frame). As a result, using a pupil coordinate frame transformed to a head coordinate frame will cause stationary virtual objects to jitter as the eye moves from position A to position B. This situation is called view-dependent display or projection.
[0233] like Figure 12 As shown, a camera coordinate frame (e.g., CR) is placed to encompass all pupil positions, and object projection will now be consistent regardless of pupil positions A and B. The head coordinate frame is transformed into a CR frame, which is referred to as a view-independent display or projection. Image reprojection can be applied to virtual content to account for changes in eye position; however, since the rendering remains in the same location, jitter can be minimized.
[0234] Figure 13The display system 42 is shown in more detail. The display system 42 comprises a stereo analyser 144 which is connected to the rendering engine 30 and forms part of the visual data and algorithms.
[0235] Display system 42 further includes left and right projectors 166A and 166B, as well as left and right waveguides 170A and 170B. Left and right projectors 166A and 166B are connected to a power source. Each projector 166A and 166B has a corresponding input for providing image data to the corresponding projector 166A or 166B. When powered on, the corresponding projector 166A or 166B generates a two-dimensional pattern of light and emits light therefrom. Left and right waveguides 170A and 170B are positioned to receive light from left and right projectors 166A and 166B, respectively. Left and right waveguides 170A and 170B are transparent waveguides.
[0236] In use, the user mounts the head-mounted frame 40 on their head. Components of the head-mounted frame 40 may, for example, include a strap (not shown) that wraps around the back of the user's head. The left and right waveguides 170A, 170B are then positioned in front of the user's left and right eyes 220A, 220B.
[0237] The rendering engine 30 inputs the image data it receives into the stereo analyzer 144. The image data is Figure 8 3D image data of local content 28 in the video is projected onto multiple virtual planes. A stereo analyzer 144 analyzes the image data to determine a left image dataset and a right image dataset based on the image data for projection onto each depth plane. The left and right image datasets represent 2D images that are projected in 3D to give the user a sense of depth.
[0238] The stereo analyzer 144 inputs the left and right image data sets to the left and right projectors 166A and 166B. The left and right projectors 166A and 166B then create left and right illumination patterns. The components of the display system 42 are shown in plan view, but it should be understood that when shown in elevation, the left and right patterns are two-dimensional patterns. Each light pattern includes a plurality of pixels. For illustrative purposes, light rays 224A and 226A from two pixels are shown exiting the left projector 166A and entering the left waveguide 170A. Light rays 224A and 226A are reflected from the sides of the left waveguide 170A. Light rays 224A and 226A are shown propagating from left to right within the left waveguide 170A via internal reflection, but it should be understood that light rays 224A and 226A could also propagate into the paper in a certain direction using a refraction and reflection system.
[0239] Light rays 224A and 226A exit the left optical waveguide 170A through pupil 228A and then enter the left eye 220A through pupil 230A of the left eye 220A. Light rays 224A and 226A then fall on the retina 232A of the left eye 220A. In this manner, the left light pattern falls on the retina 232A of the left eye 220A. The user perceives the pixels formed on the retina 232A as pixels 234A and 236A at a certain distance on the side of the left waveguide 170A opposite the left eye 220A. Depth perception is created by manipulating the focal length of the light.
[0240] In a similar manner, stereo analyzer 144 inputs the right image dataset into right projector 166B. Right projector 166B transmits a right light pattern, represented by pixels in the form of light rays 224B and 226B. Light rays 224B and 226B reflect within right waveguide 170B and exit through pupil 228B. Light rays 224B and 226B then enter through pupil 230B of right eye 220B and fall on retina 232B of right eye 220B. The pixels of light rays 224B and 226B are perceived as pixels 134B and 236B behind right waveguide 170B.
[0241] The patterns created on retinas 232A and 232B are perceived as left and right images, respectively. The left and right images are slightly different from each other due to the function of stereo analyzer 144. The left and right images are perceived as a three-dimensional rendering in the user's mind.
[0242] As mentioned, the left waveguide 170A and the right waveguide 170B are transparent. Light from a real object on the side of the table 16 opposite the eyes 220A and 220B, such as the left waveguide 170A and the right waveguide 170B, can be projected through the left waveguide 170A and the right waveguide 170B and fall on the retinas 232A and 232B.
[0243] Persistent Coordinate Frame (PCF)
[0244] This document describes methods and apparatus for providing spatial persistence between user instances within a shared space. Without spatial persistence, virtual content placed in the physical world by a user in one session may not exist or may be misplaced in the view of a user in a different session. Without spatial persistence, virtual content placed in the physical world by one user may not exist or may be misplaced in the view of a second user, even if the second user intended to share the same physical space experience as the first user.
[0245] The inventors have recognized and appreciated that spatial persistence can be provided through a persistent coordinate frame (PCF). The PCF can be defined based on one or more points that represent features recognized in the physical world (e.g., corners, edges). Features can be selected so that they appear the same from one user instance of the XR system to another.
[0246] Additionally, when rendering relative to a local map based solely on a tracking map, drift during tracking that causes the calculated tracking path (e.g., camera trajectory) to deviate from the actual tracking path can cause the position of virtual content to appear misaligned. As the XR device collects more information about the scene over time, the tracking map of the space can be refined to correct for drift. However, if virtual content is placed on real objects before map refinement and saved relative to the device's world coordinate frame derived from the tracking map, the virtual content may appear displaced as if the real objects had moved during the map refinement process. The PCF can be updated based on map refinement because it is defined based on features and is updated as features move during map refinement.
[0247] The PCF may include six degrees of freedom, including translation and rotation relative to the map coordinate system. The PCF may be stored in a local storage medium and / or a remote storage medium. Depending on, for example, the storage location, the translation and rotation of the PCF may be calculated relative to the map coordinate system. For example, a PCF used locally on a device may have a translation and rotation relative to the device's world coordinate frame. A PCF stored in the cloud may have a translation and rotation relative to the canonical coordinate frame of the canonical map.
[0248] PCF can provide a sparse representation of the physical world, providing less information about the physical world than all available information, so that they can be efficiently processed and transferred. Techniques for processing persistent spatial information may include creating a dynamic map based on one or more coordinate systems in real space across one or more sessions, generating a persistent coordinate frame (PCF) on the sparse map, which can be exposed to XR applications through, for example, an application programming interface (API).
[0249] Figure 14 11 is a block diagram illustrating the creation of a persistent coordinate frame (PCF) and the attachment of XR content to the PCF in accordance with some embodiments. Each block may represent digital information stored in computer memory. In the case of application 1180, the data may represent computer-executable instructions. In the case of virtual content 1170, the digital information may define, for example, a virtual object specified by application 1180. In the case of other blocks, the digital information may represent certain aspects of the physical world.
[0250] In the illustrated embodiment, one or more PCFs are created based on images captured by sensors on the wearable device. Figure 14 In an embodiment of the invention, the sensors are visual image cameras. These cameras can be the same cameras used to form the tracking map. Figure 14 Some of the suggested processing can be performed as part of updating the tracking map. However, Figure 14 It is shown that in addition to the tracking map, information providing persistence is also generated.
[0251] To derive a 3D PCF, two images 1110 from two cameras mounted to the wearable device in a configuration enabling stereo image analysis are processed together. Figure 14 Images 1 and 2 are shown, each of which is from one of the cameras. For simplicity, a single image from each camera is shown. However, each camera may output a stream of image frames, and the image processing may be performed for multiple image frames in the stream. Figure 14 processing.
[0252] Thus, image 1 and image 2 can each be a frame in a sequence of image frames. Figure 14 The process shown is repeated until the image frame containing the feature point provides a suitable image to form persistent spatial information from the image. Alternatively or additionally, the process may be repeated when the user moves such that the user is no longer close enough to a previously identified PCF to reliably use the PCF to determine the position relative to the physical world. Figure 14 For example, the XR system can maintain the current PCF for the user. When the distance exceeds a threshold, the system can switch to a new current PCF that is closer to the user, which can be based on Figure 14 The process of generating is performed using image frames acquired at the user's current location.
[0253] Even when a single PCF is generated, the stream of image frames can be processed to identify image frames that depict content in the physical world that is likely stable and can be easily identified by devices near the area of the physical world depicted in the image frames. Figure 14 In the embodiment shown, the process begins with the identification of features 1120 in the image. For example, features can be identified by finding locations in the image where the gradient exceeds a threshold or other feature, which may correspond to corners of an object, for example. In the embodiment shown, the features are points, but other identifiable features, such as edges, may be used instead or in addition.
[0254] In the illustrated embodiment, a fixed number N of features 1120 are selected for further processing. Those feature points may be selected based on one or more criteria, such as the magnitude of a gradient or proximity to other feature points. Alternatively or additionally, feature points may be selected heuristically, for example, based on properties that suggest the feature points are persistent. For example, a heuristic may be defined based on properties of feature points that may correspond to the corners of windows or doors or large pieces of furniture. Such a heuristic may take into account the feature point itself and its surroundings. As specific examples, the number of feature points per image may be between 100 and 500 or between 150 and 250, such as 200.
[0255] Regardless of the number of feature points selected, descriptors may be calculated 1130 for the feature points. In this example, a descriptor is calculated for each selected feature point, but descriptors may be calculated for groups of feature points, subsets of feature points, or all features within an image. Descriptors characterize feature points so that feature points that represent the same object in the physical world are assigned similar descriptors. Descriptors may enable alignment of two frames, such as may occur when one map is positioned relative to another. Instead of searching for a relative orientation of the frames that minimizes the distance between feature points of the two images, an initial alignment of the two frames may be performed by identifying feature points with similar descriptors. Alignment of image frames may be based on alignment points with similar descriptors, which may require less processing than calculating an alignment of all feature points in the images.
[0256] Descriptors can be calculated as a mapping of feature points to descriptors, or in some embodiments, as a mapping of patches of the image surrounding the feature points to descriptors. Descriptors can be numerical quantities. U.S. patent application Ser. No. 16 / 190,948 describes computing descriptors for feature points and is incorporated herein by reference in its entirety.
[0257] exist Figure 14 In the example of , a descriptor is calculated for each feature point in each image frame 1130. Based on the descriptors and / or the feature points and / or the image itself, an image frame may be identified as a keyframe 1140. In the embodiment shown, a keyframe is an image frame that meets a certain criterion, which is then selected for further processing. For example, when making a tracking map, image frames that add meaningful information to the map may be selected as keyframes to be integrated into the map. On the other hand, image frames that substantially overlap with areas where image frames have already been integrated into the map may be discarded so that they do not become keyframes. Alternatively or additionally, keyframes may be selected based on the number and / or type of feature points in the image frame. In Figure 14In embodiments of the present invention, keyframes 1150 selected for inclusion in the tracking map may also be considered keyframes for determining the PCF, although different or additional criteria for selecting keyframes for generating the PCF may be used.
[0258] although Figure 14 While keyframes are shown as being used for further processing, the information obtained from the images can be processed in other ways. For example, feature points, such as in key assemblies, can be processed alternatively or additionally. Furthermore, while keyframes are described as being derived from a single image frame, there need not be a one-to-one relationship between keyframes and the image frames from which they are obtained. For example, a keyframe can be obtained from multiple image frames, such as by concatenating or aggregating the image frames together, such that only features that appear in multiple images are retained in the keyframe.
[0259] The key frame may include image information and / or metadata associated with the image information. In some embodiments, the key frame may be captured by cameras 44, 46 ( Figure 9 ) is calculated as one or more keyframes (e.g., keyframes 1, 2). In some embodiments, a keyframe may include a camera pose. In some embodiments, a keyframe may include one or more camera images captured at a camera pose. In some embodiments, the XR system may determine that a portion of a camera image captured at a camera pose is useless and therefore not include that portion in the keyframe. Therefore, using keyframes to align new images with early knowledge of the scene can reduce the use of XR system computing resources. In some embodiments, a keyframe may include an image and / or image data at a location with a direction / angle. In some embodiments, a keyframe may include a location and direction in which one or more map points can be observed. In some embodiments, a keyframe may include a coordinate frame with an ID. U.S. Patent Application No. 15 / 877,359 describes keyframes, which is incorporated herein by reference in its entirety.
[0260] Some or all of the keyframes 1140 may be selected for further processing, such as generating a persistent gesture 1150 for the keyframes. This selection may be based on characteristics of all or a subset of the feature points in the image frame. These characteristics may be determined by processing the descriptors, features, and / or the image frame itself. As a specific example, the selection may be based on clustering of feature points identified as potentially associated with a persistent object.
[0261] Each keyframe is associated with the pose of the camera that captured it. For keyframes selected for processing as persistent poses, this pose information can be saved along with other metadata about the keyframe, such as a WiFi fingerprint and / or GPS coordinates at the time and / or location of capture. In some embodiments, metadata such as GPS coordinates can be used alone or in combination as part of the positioning process.
[0262] A persistent gesture is a source of information that the device can use to orient itself relative to previously acquired information about the physical world. For example, if the keyframe from which the persistent gesture was created is incorporated into a map of the physical world, the device can orient itself relative to the persistent gesture using a sufficient number of feature points in the keyframe associated with the persistent gesture. The device can align a current image it takes of its surroundings with the persistent gesture. The alignment can be based on matching the current image with the image 1110, features 1120 and / or descriptors 1130 that gave rise to the persistent gesture, or any subset of the image or those features or descriptors. In some embodiments, the current image frame that is matched to the persistent gesture can be another keyframe that has been incorporated into the device's tracking map.
[0263] Information about persistent gestures can be stored in a format that facilitates sharing among multiple applications that may be executing on the same or different devices. Figure 14 In an example of FIG11 , some or all persistent gestures may be reflected as a persistent coordinate frame (PCF) 1160. Like persistent gestures, a PCF may be associated with a map and may include a set of features or other information that a device may use to determine its orientation relative to the PCF. A PCF may include a transform that defines a transformation relative to the origin of its map, such that by associating its position with the PCF, a device may determine its position relative to any object in the physical world reflected in the map.
[0264] Because PCFs provide a mechanism for determining positions relative to physical objects, applications (e.g., application 1180) can define the positions of virtual objects relative to one or more PCFs, which serve as anchor points for virtual content 1170. For example, Figure 14 App 1 is shown as having associated its virtual content 2 with PCF 1.2. Similarly, App 2 has associated its virtual content 3 with PCF 1.2. App 1 is also shown as associating its virtual content 1 with PCF 4.5, and App 2 is shown as associating its virtual content 4 with PCF 3. In some embodiments, PCF 3 can be based on image 3 (not shown), and PCF 4.5 can be based on image 4 and image 5 (not shown), similar to how PCF 1.2 is based on image 1 and image 2. When rendering this virtual content, the device can apply one or more transformations to calculate information such as the position of the virtual content relative to the device's display and / or the position of physical objects relative to the desired position of the virtual content. Using the PCF as a reference can simplify such calculations.
[0265] In some embodiments, a persistent gesture can be a coordinate position and / or orientation with one or more associated keyframes. In some embodiments, a persistent gesture can be automatically created after the user has traveled a certain distance (e.g., three meters). In some embodiments, a persistent gesture can be used as a reference point during positioning. In some embodiments, a persistent gesture can be stored in a traversable world (e.g., traversable world module 538).
[0266] In some embodiments, a new PCF can be determined based on a predetermined distance allowed between adjacent PCFs. In some embodiments, when the user travels a predetermined distance (e.g., five meters), one or more persistent gestures can be calculated into the PCF. In some embodiments, the PCF can be associated with one or more world coordinate frames and / or canonical coordinate frames, such as in a navigable world. In some embodiments, the PCF can be stored in a local database and / or a remote database, depending on, for example, security settings.
[0267] Figure 15 A method 4700 of establishing and using a persistent coordinate frame according to some embodiments is shown. The method 4700 may begin by capturing (act 4702) an image of a scene (e.g., Figure 14 1 and 2 in ). Multiple cameras can be used, and one camera can generate multiple images, for example in the form of a stream.
[0268] Method 4700 may include extracting (4704) points of interest (e.g., Figure 7 Map point 702, Figure 14 1120 in the features), generating (action 4706) a descriptor of the extracted point of interest (e.g., Figure 14 1130 in the image) and generates (act 4708) a keyframe (e.g., keyframe 1140) based on the descriptor. In some embodiments, the method may compare points of interest in the keyframes and form pairs of keyframes that share a predetermined amount of points of interest. The method may use the respective keyframe pairs to reconstruct a portion of the physical world. The mapped portion of the physical world may be saved as a 3D feature (e.g., Figure 7704 in the key assembly). In some embodiments, selected portions of the key frame pairs can be used to construct 3D features. In some embodiments, the results of the mapping can be selectively saved. Key frames that are not used to construct 3D features can be associated with 3D features by pose, for example, by representing the distance between key frames using a covariance matrix between the poses of the key frames. In some embodiments, key frame pairs can be selected to construct 3D features so that the distance between each two of the constructed 3D features is within a predetermined distance, which can be determined to balance the amount of computation required and the level of accuracy of the resulting model. Such an approach can provide the XR system with a model of the physical world with an amount of data suitable for efficient and accurate computation. In some embodiments, the covariance matrix of two images may include the covariance between the poses (e.g., six degrees of freedom) of the two images.
[0269] Method 4700 may include generating (action 4710) a persistent gesture based on the keyframes. In some embodiments, the method may include generating a persistent gesture based on a 3D feature reconstructed from a pair of keyframes. In some embodiments, the persistent gesture may be attached to the 3D feature. In some embodiments, the persistent gesture may include a gesture of the keyframes used to construct the 3D feature. In some embodiments, the persistent gesture may include an average pose of the keyframes used to construct the 3D feature. In some embodiments, the persistent gesture may be generated such that the distance between adjacent persistent gestures is within a predetermined value, such as within a range of one meter to five meters, any value therebetween, or any other appropriate value. In some embodiments, the distance between adjacent persistent gestures may be represented by a covariance matrix of adjacent persistent gestures.
[0270] Method 4700 may include generating (act 4712) a PCF based on the persistent pose. In some embodiments, the PCF may be attached to the 3D feature. In some embodiments, the PCF may be associated with one or more persistent poses. In some embodiments, the PCF may include the pose of one of the associated persistent poses. In some embodiments, the PCF may include the average pose of the poses of the associated persistent poses. In some embodiments, the PCF may be generated so that the distance between adjacent PCFs is within a predetermined value, such as within a range of three meters to ten meters, any value therebetween, or any other appropriate value. In some embodiments, the distance between adjacent PCFs may be represented by a covariance matrix of adjacent PCFs. In some embodiments, the PCF may be exposed to the XR application via, for example, an application programming interface (API) so that the XR application can access a model of the physical world through the PCF without accessing the model itself.
[0271] Method 4700 may include associating image data of a virtual object to be displayed by the XR device with at least one of the PCFs (action 4714). In some embodiments, the method may include calculating a translation and orientation of the virtual object relative to the associated PCF. It should be understood that it is not necessary to associate the virtual object with a PCF generated by the device where the virtual object is placed. For example, the device may retrieve a saved PCF from a canonical map in the cloud and associate the virtual object with the retrieved PCF. It should be understood that as the PCF is adjusted over time, the virtual object may move with the associated PCF.
[0272] Figure 16 Visual data and algorithms of a first XR device 12 . 1 and a second XR device 12 . 2 and a server 20 are shown, according to some embodiments. Figure 16 The components shown in FIG. 1 and FIG. 2 may be operable to perform some or all of the operations associated with generating, updating, and / or using spatial information (such as a persistent pose, a persistent coordinate frame, a tracking map, or a canonical map) as described herein. Although not shown, the first XR device 12.1 may be configured identically to the second XR device 12.2. The server 20 may have a map storage routine 118, a canonical map 120, a map sender 122, and a map merging algorithm 124.
[0273] The second XR device 12.2, which may be in the same scene as the first XR device 12.1, may include a persistent coordinate frame (PCF) integration unit 1300, an application 1302 that generates image data 68 that may be used to render a virtual object, and a frame embedding generator 308 (see Figure 21 In some embodiments, the map download system 126, the PCF identification system 128, the map Figure 2 , positioning module 130, canonical map merger 132, canonical map 133, and map publisher 136 are collectively referred to as a traversable world unit 1304. The PCF integration unit 1300 can be connected to the traversable world unit 1304 and other components of the second XR device 12.2 to allow acquisition, generation, use, upload, and download of PCF.
[0274] Maps including PCFs can achieve more persistence in a changing world. In some embodiments, locating a tracking map including matching features of, for example, an image can include selecting features representing persistent content from a map composed of PCFs, which enables fast matching and / or localization. For example, in a world where people enter and exit a scene and objects such as doors move relative to the scene, less storage space and transmission rates are required, and the scene can be mapped using separate PCFs and their relationships to each other (e.g., an integrated constellation of PCFs).
[0275] In some embodiments, the PCF integration unit 1300 may include a PCF 1306 previously stored in a data storage on a storage unit of the second XR device 12.2, a PCF tracker 1308, a persistent pose acquirer 1310, a PCF checker 1312, a PCF generation system 1314, a coordinate frame calculator 1316, a persistent pose calculator 1318, and three transformers including a tracking map and persistent pose transformer 1320, a persistent pose and PCF transformer 1322, and a PCF and image data transformer 1324.
[0276] In some embodiments, the PCF tracker 1308 may have an on prompt and a off prompt selectable by the application 1302. The application 1302 may be executable by the processor of the second XR device 12.2 to, for example, display virtual content. The application 1302 may have a call to turn on the PCF tracker 1308 via the on prompt. When the PCF tracker 1308 is on, the PCF tracker 1308 may generate PCFs. The application 1302 may have a subsequent call to turn off the PCF tracker 1308 via the off prompt. When the PCF tracker 1308 is off, the PCF tracker 1308 terminates PCF generation.
[0277] In some embodiments, the server 20 may include a plurality of persistent gestures 1332 and a plurality of PCFs 1330 that have been previously saved in association with the canonical map 120. The map sender 122 may send the canonical map 120 along with the persistent gestures 1332 and / or PCFs 1330 to the second XR device 12.2. The persistent gestures 1332 and PCFs 1330 may be stored on the second XR device 12.2 in association with the canonical map 133. Figure 2 When locating to the standard map 133, you can Figure 2 The persistent gesture 1332 and the PCF 1330 are stored in association.
[0278] In some embodiments, the persistent gesture acquirer 1310 may acquire the Figure 2 The PCF checker 1312 may be connected to the persistent gesture acquirer 1310. The PCF checker 1312 may acquire a PCF from the PCF 1306 based on the persistent gesture acquired by the persistent gesture acquirer 1310. The PCFs acquired by the PCF checker 1312 may form an initial set of PCFs for use in PCF-based image display.
[0279] In some embodiments, the application 1302 may need to generate additional PCFs. For example, if the user moves to an area that has not been mapped before, the application 1302 may open the PCF tracker 1308. The PCF generation system 1314 may connect to the PCF tracker 1308 and generate additional PCFs as the map is created. Figure 2 Start expanding and start based on the ground Figure 2 Generate PCFs. The PCFs generated by the PCF generation system 1314 may form a second set of PCFs, which may be used for PCF-based image display.
[0280] The coordinate frame calculator 1316 can be connected to the PCF checker 1312. After the PCF checker 1312 obtains the PCF, the coordinate frame calculator 1316 can call the head coordinate frame 96 to determine the head pose of the second XR device 12.2. The coordinate frame calculator 1316 can also call the persistent pose calculator 1318. The persistent pose calculator 1318 can be directly or indirectly connected to the frame embedding generator 308. In some embodiments, the image / frame can be designated as a keyframe that travels a threshold distance (e.g., 3 meters) from the previous keyframe. The persistent pose calculator 1318 can generate a persistent pose based on multiple (e.g., three) keyframes. In some embodiments, the persistent pose can be substantially an average of the coordinate frames of the multiple keyframes.
[0281] Tracking map and persistent pose changer 1320 can be connected to the ground Figure 2 and persistent pose calculator 1318. Tracking map and persistent pose converter 1320 can Figure 2 Transform to a persistent pose to determine relative to the ground Figure 2 The persistent posture at the origin of .
[0282] Persistent pose and PCF converter 1322 may be coupled to tracking map and persistent pose converter 1320 and further coupled to PCF checker 1312 and PCF generation system 1314. Persistent pose and PCF converter 1322 may convert the persistent pose (to which the tracking map has been converted) from PCF checker 1312 and PCF generation system 1314 into a PCF to determine a PCF relative to the persistent pose.
[0283] PCF and image data transformer 1324 may be connected to persistent gesture and PCF transformer 1322 and data channel 62. PCF and image data transformer 1324 transforms the PCF into image data 68. Rendering engine 30 may be connected to PCF and image data transformer 1324 to display image data 68 to a user relative to the PCF.
[0284] The PCF integration unit 1300 may store additional PCFs generated using the PCF generation system 1314 within the PCF 1306. The PCF 1306 may be stored relative to the persistent posture. Figure 2When the map publisher 136 obtains the PCF 1306 and the persistent gesture associated with the PCF 1306, the map publisher 136 also sends the map to the server 20. Figure 2 The associated PCF and persistent posture. When the map storage routine 118 of the server 20 stores the map Figure 2 The map storage routine 118 may also store the persistent gesture and PCF generated by the second viewing device 12.2. The map merging algorithm 124 may use the map associated with the canonical map 120 and stored in the persistent gesture 1332 and PCF 1330, respectively. Figure 2 The persistent pose and PCF are used to create a canonical map 120 .
[0285] The first XR device 12.1 may include a PCF integration unit similar to the PCF integration unit 1300 of the second XR device 12.2. When the map sender 122 sends the canonical map 120 to the first XR device 12.1, the map sender 122 may send a persistent gesture 1332 and a PCF 1330 associated with the canonical map 120 and originating from the second XR device 12.2. The first XR device 12.1 may store the PCF and the persistent gesture in a data store on a storage device of the first XR device 12.1. The first XR device 12.1 may then utilize the persistent gesture and PCF originating from the second XR device 12.2 for image display relative to the PCF. Additionally or alternatively, the first XR device 12.1 may obtain, generate, use, upload, and download the PCF and the persistent gesture in a manner similar to that of the second XR device 12.2 as described above.
[0286] In the example shown, the first XR device 12.1 generates a local tracking map (hereinafter referred to as “map”). Figure 1 ”), and the map storage routine 118 receives the map from the first XR device 12.1 Figure 1 The map storage routine 118 then stores the map Figure 1 The specification map 120 is stored on the storage device of the server 20 .
[0287] The second XR device 12 . 2 includes a map download system 126 , an anchor point identification system 128 , a positioning module 130 , a canonical map merger 132 , a local content positioning system 134 , and a map publisher 136 .
[0288] In use, the map sender 122 sends the canonical map 120 to the second XR device 12 . 2 , and the map download system 126 downloads and stores the canonical map 120 from the server 20 as the canonical map 133 .
[0289] The anchor point identification system 128 is connected to the world surface determination routine 78. The anchor point identification system 128 identifies anchor points based on the objects detected by the world surface determination routine 78. The anchor point identification system 128 generates a second map (map) using the anchor points. Figure 2 As shown in loop 138, the anchor point identification system 128 continues to identify anchor points and continues to update the ground Figure 2 The positions of the anchor points are recorded as three-dimensional data based on data provided by the world surface determination routine 78. The world surface determination routine 78 receives images from the real object detection camera 44 and depth data from the depth sensor 135 to determine the positions of surfaces and their relative distances from the depth sensor 135.
[0290] Positioning module 130 is connected to the specification map 133 and the Figure 2 The positioning module 130 repeatedly attempts to Figure 2 Locate the normative map 133. The normative map merger 132 is connected to the normative map 133 and the map Figure 2 When the positioning module 130 sets the ground Figure 2 When the standard map 133 is located, the standard map merger 132 merges the standard map 133 into the map. Figure 2 Then, the map is updated with the missing data included in the canonical map. Figure 2 .
[0291] The local content location system 134 is connected to the ground Figure 2 The local content location system 134 may be, for example, a system in which a user can locate local content at a specific location within a world coordinate frame. The local content then attaches itself to the local Figure 2 The local to world coordinate converter 104 converts the local coordinate frame into the world coordinate frame based on the settings of the local content positioning system 134. Figure 2 The functionality of the rendering engine 30, display system 42, and data pipeline 62 is described.
[0292] Map publisher 136 will be the map Figure 2 Upload to the server 20. The map storage routine 118 of the server 20 then Figure 2 Stored in the storage medium of the server 20.
[0293] Map Merge Algorithm 124 Figure 2Merge with the canonical map 120. When more than two maps are already stored (e.g., three or four maps relating to the same or adjacent areas of the physical world), the map merging algorithm 124 merges all of the maps into the canonical map 120 to render a new canonical map 120. The map sender 122 then sends the new canonical map 120 to any and all devices 12.1 and 12.2 located in the area represented by the new canonical map 120. When devices 12.1 and 12.2 position their respective maps to the canonical map 120, the canonical map 120 becomes the upgraded map.
[0294] Figure 17 An example of generating keyframes for a map of a scene, according to some embodiments, is shown. In the illustrated example, a first keyframe, KF1, is generated for the door on the left wall of a room. A second keyframe, KF2, is generated for the corner area where the floor, left wall, and right wall intersect. A third keyframe, KF3, is generated for the window area on the right wall of the room. A fourth keyframe, KF4, is generated for the area of the carpet on the floor, far away from the wall. A fifth keyframe, KF5, is generated for the area of the carpet closest to the user.
[0295] Figure 18 According to some embodiments, Figure 17 In some embodiments, a new persistent gesture is created when the device measures a threshold distance traveled, and / or when an application requests a new persistent gesture (PP). In some embodiments, the threshold distance can be 3 meters, 5 meters, 20 meters, or any other suitable distance. Selecting a smaller threshold distance (e.g., 1 meter) can result in an increased computational load because a larger number of PPs can be created and managed compared to a larger threshold distance. Selecting a larger threshold distance (e.g., 40 meters) can result in increased virtual content placement errors because a smaller number of PPs will be created, which will result in a smaller number of PCFs being created, meaning that virtual content attached to a PCF may be a relatively large distance away from the PCF (e.g., 30 meters), and the error increases with increasing distance from the PCF to the virtual content.
[0296] In some embodiments, a PP may be created when a new session begins. This initial PP may be considered zero and may be visualized as the center of a circle with a radius equal to the threshold distance. When the device reaches the circumference of the circle, and in some embodiments, an application requests a new PP, the new PP may be placed at the device's current location (at the threshold distance). In some embodiments, if the device can find an existing PP within the threshold distance from the device's new location, a new PP is not created at the threshold distance. In some embodiments, when a new PP is created (e.g., Figure 14In some embodiments, the device may create a PP relative to the keyframes (1150) and append the one or more closest keyframes to the PP. In some embodiments, the location of the PP relative to the keyframes may be based on the device's location at the time the PP was created. In some embodiments, a PP is not created when the device travels a threshold distance unless the application requests a PP.
[0297] In some embodiments, when an application has virtual content to display to the user, the application can request a PCF from the device. The PCF request from the application can trigger a PP request, and a new PP will be created after the device travels a threshold distance. Figure 18 A first persistent pose PP1 is shown, which may have closest keyframes (eg, KF1 , KF2 , and KF3 ) appended by, for example, computing relative poses between the keyframes and the persistent pose. Figure 18 Also shown is a second permanent pose PP2, which may have additional proximal keyframes (eg, KF4 and KF5).
[0298] Figure 19 According to some embodiments, Figure 17 1 and 2. In the example shown, PCF1 may include PP1 and PP2. As described above, PCFs may be used to display image data associated with the PCFs. In some embodiments, each PCF may have coordinates in another coordinate frame (e.g., a world coordinate frame) and a PCF descriptor, e.g., to uniquely identify the PCF. In some embodiments, the PCF descriptor may be calculated based on feature descriptors of features in a frame associated with the PCF. In some embodiments, various constellations of PCFs may be combined to represent the real world in a persistent manner that requires less data and less data transmission.
[0299] Figures 20A to 20C is a diagram illustrating an example of establishing and using a persistent coordinate frame. Figure 20A Two users 4802A, 4802B are shown with respective local tracking maps 4804A, 4804B that have not yet been localized to a canonical map. The origin 4806A, 4806B of each user is depicted by a coordinate system (e.g., a world coordinate system) in their respective regions. These origins of each tracking map may be local to each user because they depend on the orientation of their respective devices when tracking is initiated.
[0300] When the user device's sensors scan the environment, the device can capture the above combined Figure 14 The depicted images may contain features representing persistent objects such that those images may be classified as keyframes from which persistent gestures may be created. In this example, track map 4802A includes persistent gesture (PP) 4808A; track 4802B includes PP 4808B.
[0301] Likewise, as above combined Figure 14 As described above, some PPs may be classified as PCFs, which are used to determine the orientation of virtual content to render it to the user. Figure 20B It is shown that the XR devices worn by the respective users 4802A, 4802B can create local PCFs 4810A, 4810B based on the PPs 4808A, 4808B. Figure 20C It is shown that persistent content 4812A, 4812B (e.g., virtual content) can be attached to the PCF 4810A, 4810B through corresponding XR devices.
[0302] In this example, virtual content can have a virtual content coordinate frame that can be used by the application generating the virtual content regardless of how the virtual content should be displayed. For example, virtual content can be specified as surfaces, such as triangles of a mesh, at specific positions and angles relative to the virtual content coordinate frame. To render the virtual content to the user, the positions of those surfaces can be determined relative to the user who is to perceive the virtual content.
[0303] Attaching virtual content to the PCF can simplify the computations involved in determining the position of the virtual content relative to the user. The position of the virtual content relative to the user can be determined by applying a series of transformations. Some of these transformations may change and may be updated frequently. Others of these transformations may be stable and may be updated frequently or not at all. Regardless, the transformations can be applied with a relatively low computational burden, allowing the position of the virtual content relative to the user to be updated frequently, thereby providing a realistic appearance to the rendered virtual content.
[0304] exist 20A to 20C In the example shown in Figure 1, user 1's device has a coordinate system that is related to the coordinate system that defines the map origin via the transformation rig1_T_w1. User 2's device has a similar transformation rig2_T_w2. These transformations can be expressed as six degrees of transformation, specifying the translation and rotation to align the device coordinate system with the map coordinate system. In some embodiments, the transformation can be expressed as two separate transformations, one specifying the translation and the other specifying the rotation. Therefore, it should be understood that the transformations can be expressed in a form that simplifies calculations or otherwise provides advantages.
[0305] The transformations between the origin of the tracking map and the PCF identified by the corresponding user equipment are denoted as pcf1_T_w1 and pcf2_T_w2. In this example, the PCF and PP are the same, so the same transformation also characterizes the PP.
[0306] Therefore, the position of the user equipment relative to the PCF can be calculated by serial application of these transformations, for example rig1_T_pcf1=(rig1_T_w1)*(pcf1_T_w1).
[0307] like Figure 20C As shown, the virtual content is positioned relative to the PCF through the transformation of obj1_T_pcf1. This transformation can be set by the application generating the virtual content, which can receive information describing physical objects relative to the PCF from the world reconstruction system. In order to render the virtual content to the user, the transformation to the coordinate system of the user's device is calculated, which can be calculated by associating the virtual content coordinate frame to the origin of the tracking map through the transformation obj1_t_w1 = (obj1_T_pcf1) * (pcf1_T_w1). This transformation can then be related to the user's device through a further transformation rig1_T_w1.
[0308] Based on the output from the application generating the virtual content, the position of the virtual content can change. When this happens, the end-to-end transformation from the source coordinate system to the destination coordinate system can be recalculated. Additionally, the user's position and / or head pose can change as the user moves. As a result, the transformation rig1_T_w1 can change, as can any end-to-end transformations that depend on the user's position or head pose.
[0309] The transformation rig1_T_w1 can be updated as the user moves based on tracking the user's position relative to stationary objects in the physical world. This tracking can be performed by the headset positioning component or other components of the system that process the image sequence as described above. Such updates can be performed by determining the user's posture relative to a fixed reference frame (e.g., PP).
[0310] In some embodiments, because the PP is used as the PCF, the position and orientation of the user device can be determined relative to the most recent persistent gesture, or in this example, the PCF. This determination can be made by identifying feature points representing the PP in the current image captured using sensors on the device. Using image processing techniques such as stereo image analysis, the position of the device relative to those feature points can be determined. Based on this data, the system can calculate the change in transformation associated with the user's motion based on the relationship rig1_T_pcf1=(rig1_T_w1)*(pcf1_T_w1).
[0311] The system can determine and apply transformations in a computationally efficient order. For example, the need to calculate rig1_T_w1 from measurements that generate rig1_T_pcf1 can be avoided by tracking user gestures and defining the position of virtual content relative to a PP or PCF constructed based on persistent gestures. In this way, the transformation from the source coordinate system of the virtual content to the destination coordinate system of the user device can be based on a measured transformation according to the expression (rig1_T_pcf1)*(obj1_t_pcf1), where the first transformation is measured by the system and the latter transformation is provided by the application that specifies the virtual content to be rendered. In embodiments where the virtual content is positioned relative to the origin of a map, the end-to-end transformation can relate the virtual object coordinate system to the PCF coordinate system based on a further transformation between map coordinates and PCF coordinates. In embodiments where the virtual content is positioned relative to a different PP or PCF than the PP or PCF for which the user's position is tracked, a transformation can be performed between the two. Such a transformation can be fixed and can, for example, be determined from a map in which both appear.
[0312] For example, a transformation-based approach can be implemented in a device that has components that process sensor data to build a tracking map. As part of this process, these components can identify feature points that can be used as persistent gestures, which in turn can be turned into PCFs. These components can limit the number of persistent gestures generated for the map to provide appropriate spacing between persistent gestures while allowing the user to be close enough to the persistent gesture location regardless of their position in the physical environment to accurately calculate the user's gesture, as described above in conjunction with Figures 17 to 19 As the persistent gesture closest to the user is updated, any transformations used to calculate the position of virtual content relative to the user that depend on the PP (or PCF, if used) can be updated and stored for use as the user moves, as the tracking map or other refinements are made. This allows for the position of the virtual content to be updated with relatively low computational burden each time the position of the virtual content is updated, allowing it to be performed with relatively low latency.
[0313] 20A to 20C Positioning is shown relative to a tracking map, with each device having its own tracking map. However, transformations can be generated relative to any map coordinate system. Content persistence between user sessions of an XR system can be achieved through the use of a persistent map. Shared user experiences can also be achieved through the use of a map to which multiple user devices can be directed.
[0314] In some embodiments described in more detail below, the location of virtual content can be specified relative to coordinates in a canonical map that is formatted so that any of multiple devices can use the map. Each device may maintain a tracking map and can determine changes in the user's posture relative to the tracking map. In this example, the transformation between the tracking map and the canonical map can be determined by a "localization" process, which can be performed by matching structures in the tracking map (such as one or more persistent postures) to one or more structures of the canonical map (e.g., one or more PCFs).
[0315] Techniques for creating and using canonical maps in this manner are described in more detail below.
[0316] Depth Keyframes
[0317] The techniques described herein rely on comparing image frames. For example, to establish the position of the device relative to a tracking map, a new image can be captured using sensors worn by the user, and the XR system can search the image set used to create the tracking map for an image that shares at least a predetermined number of points of interest with the new image. As an example of another scenario involving image frame comparison, the tracking map can be localized to the canonical map by first finding an image frame in the tracking map associated with a persistent gesture that is similar to an image frame in the canonical map associated with a PCF. Alternatively, the transformation between the two canonical maps can be computed by first finding similar image frames in the two maps.
[0318] Depth keyframes provide a method for reducing the amount of processing required to identify similar image frames. For example, in some embodiments, the comparison can be between image features in the new 2D image (e.g., "2D features") and 3D features in the map. This comparison can be performed in any appropriate manner, such as by projecting the 3D image into a 2D plane. Conventional methods such as Bag of Words (BoW) search for 2D features of the new image in a database that includes all 2D features in the map, which can require a large amount of computing resources, especially when the map represents a large area. The conventional method then locates images that share at least one 2D feature with the new image, which may include images that are not useful for locating meaningful 3D features in the map. The conventional method then locates 3D features that are meaningless relative to the 2D features in the new image.
[0319] The inventors have recognized and understood techniques for retrieving images from a map using fewer memory resources (e.g., one-quarter the memory resources used by BoW), with greater efficiency (e.g., 2.5 ms processing time per keyframe and 100 μs for comparison of 500 keyframes), and with greater accuracy (e.g., 20% better retrieval recall than BoW for a 1024-dimensional model and 5% better retrieval recall than BoW for a 256-dimensional model).
[0320] To reduce computation, a descriptor can be calculated for an image frame, which can be used to compare the image frame with other image frames. The descriptor can be stored instead of, or in addition to, the image frame and feature points. In a map where persistent gestures and / or PCFs can be generated from image frames, the descriptors of the one or more image frames from which each persistent gesture or PCF was generated can be stored as part of the persistent gesture and / or PCF.
[0321] In some embodiments, descriptors can be calculated based on feature points in an image frame. In some embodiments, a neural network is configured to calculate a unique frame descriptor representing an image. The image can have a resolution greater than 1 megabyte, thereby capturing sufficient detail of the 3D environment within the field of view of a device worn by the user. The frame descriptor can be much shorter, such as a string of numbers, for example, in the range of 128 bytes to 512 bytes, or any number in between.
[0322] In some embodiments, a neural network is trained to compute frame descriptors that indicate similarity between images. An image in a map can be located by identifying the nearest image in a database of images used to generate the map that has a frame descriptor within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between images can be represented by the difference between the frame descriptors of the two images.
[0323] Figure 21 is a block diagram illustrating a system for generating descriptors for individual images according to some embodiments. In the example shown, a frame embedding generator 308 is shown. In some embodiments, frame embedding generator 308 may be used within server 20, but may alternatively or additionally be executed in whole or in part within one of XR devices 12.1 and 12.2, or any other device that processes an image for comparison with other images.
[0324] In some embodiments, the frame embedding generator can be configured to generate a reduced data representation of an image from an initial size (e.g., 76,800 bytes) to a final size (e.g., 256 bytes) that is indicative of the content in the image despite the reduced size. In some embodiments, the frame embedding generator can be used to generate a data representation of an image, which can be a keyframe or a frame used in other ways. In some embodiments, the frame embedding generator 308 can be configured to convert an image at a particular location and orientation into a unique string of numbers (e.g., 256 bytes). In the example shown, an image 320 captured by an XR device can be processed by a feature extractor 324 to detect points of interest 322 in the image 320. The points of interest may or may not be derived from feature points identified as described above for features 1120 ( Figure 14 ) or as otherwise described herein. In some embodiments, the point of interest may be represented by a descriptor 1130 ( Figure 14 ) described above, these points of interest can be generated using a deep sparse feature method. In some embodiments, each point of interest 322 can be represented by a digital string (e.g., 32 bytes). For example, there can be n features (e.g., 100), and each feature can be represented by a 32-byte string.
[0325] In some embodiments, the frame embedding generator 308 may include a neural network 326. The neural network 326 may include a multilayer perceptron unit 312 and a maximum (max) pooling unit 314. In some embodiments, the multilayer perceptron (MLP) unit 312 may include a multilayer perceptron, which may be trained. In some embodiments, the points of interest 322 (e.g., descriptors for the points of interest) may be reduced by the multilayer perceptron 312 and may be output as a weighted combination of descriptors 310. For example, the MLP may reduce n features to m features, which is less than n features.
[0326] In some embodiments, the MLP unit 312 can be configured to perform matrix multiplication. The multilayer perceptron unit 312 receives a plurality of points of interest 322 of the image 320 and converts each point of interest into a corresponding string of numbers (e.g., 256). For example, there may be 100 features, and each feature can be represented by a string of 256 numbers. In this example, a matrix with 100 horizontal rows and 256 vertical columns can be created. Each row can have a series of 256 numbers that vary in size, with some being smaller and others being larger. In some embodiments, the output of the MLP can be an n×256 matrix, where n represents the number of points of interest extracted from the image. In some embodiments, the output of the MLP can be an m×256 matrix, where m is the number of points of interest reduced from n.
[0327] In some embodiments, the MLP 312 may have a training phase during which model parameters for the MLP are determined and a usage phase. Figure 25 The MLP is trained as shown in [1]. The input training data may include a set of three data, each set including 1) a query image, 2) a positive sample, and 3) a negative sample. The query image can be considered as a reference image.
[0328] In some embodiments, the positive image may include an image that is similar to the query image. For example, in some embodiments, the similarity may be that the query image and the positive image have the same object, but are viewed from different angles. In some embodiments, the similarity may be that the query image and the positive image have the same object, but the object is shifted relative to the other image (e.g., to the left, right, up, down).
[0329] In some embodiments, negative samples may include images that are dissimilar to the query image. For example, in some embodiments, a dissimilar image may not contain any objects that are prominent in the query image, or may contain only a small portion (e.g., <10%, 1%) of the prominent objects in the query image. In contrast, a similar image may have a large portion (e.g., >50% or >75%) of the objects in the query image, for example.
[0330] In some embodiments, the focus points can be extracted from the images in the input training data and converted into feature descriptors. Figure 25 The training images shown and the Figure 21 These descriptors are computed using both features extracted from the operation of the frame embedding generator 308. In some embodiments, a deep sparse feature (DSF) process can be used to generate descriptors (e.g., DSF descriptors), as described in U.S. patent application Ser. No. 16 / 190,948. In some embodiments, the DSF descriptors are n×32 dimensional. The descriptors can then be passed through the model / MLP to create a 256-byte output. In some embodiments, the model / MLP can have the same structure as the MLP 312, so that once the model parameters are set through training, the resulting trained MLP can be used as the MLP 312.
[0331] In some embodiments, the feature descriptors (e.g., the 256 bytes output from the MLP model) can then be sent to a triple boundary loss module (which can be used only during the training phase and not during the use phase of the MLP neural network). In some embodiments, the triple boundary loss module can be configured to select parameters for the model to reduce the difference between the 256-byte output from the query image and the 256-byte output from the positive samples, and to increase the 256-byte output from the query image and the 256-byte output from the negative samples. In some embodiments, the training phase can include feeding multiple triplet input images into a learning process to determine the model parameters. The training process can continue, for example, until the difference for positive images is minimized and the difference for negative images is maximized, or until other appropriate exit criteria are met.
[0332] Reference again Figure 21 , the frame embedding generator 308 may include a pooling layer, shown here as a max pooling unit 314. The max pooling unit 314 may analyze each column to determine the maximum number in the corresponding column. The max pooling unit 314 may combine the maximum values of the numbers in each column of the output matrix of the MLP 312 into a global feature string 316 of, for example, 256 numbers. It will be appreciated that images processed in an XR system may desire high-resolution frames, potentially having millions of pixels. The global feature string 316 is a relatively small number that takes up relatively little memory and is easy to search compared to an image (e.g., having a resolution greater than 1 megabyte). The image can therefore be searched without having to analyze every raw frame from the camera, and it is also cheaper to store 256 bytes rather than a full frame.
[0333] Figure 22 22 is a flow chart illustrating a method 2200 for computing an image descriptor according to some embodiments. Method 2200 may begin by receiving (act 2202) a plurality of images captured by an XR device worn by a user. In some embodiments, method 2200 may include determining (act 2204) one or more keyframes from the plurality of images. In some embodiments, act 2204 may be skipped and / or may instead occur after step 2210.
[0334] Method 2200 may include identifying (act 2206) one or more points of interest in a plurality of images using an artificial neural network, and computing (act 2208) feature descriptors for each point of interest using the artificial neural network. The method may include computing (act 2210) a frame descriptor for each image, thereby representing the image based at least in part on the feature descriptors computed for the points of interest identified in the image using the artificial neural network.
[0335] Figure 2323 is a flowchart illustrating a method 2300 for localization using image descriptors, according to some embodiments. In this example, a new image frame describing the current location of an XR device can be compared to image frames stored in connection with points in a map (e.g., a persistent gesture or PCF, as described above). Method 2300 can begin by receiving (act 2302) a new image captured by an XR device worn by a user. Method 2300 can include identifying (act 2304) one or more recent keyframes in a database that includes keyframes used to generate one or more maps. In some embodiments, the recent keyframes can be identified based on coarse spatial information and / or previously determined spatial information. For example, the coarse spatial information can indicate that the XR device is located in a geographic area represented by a 50m×50m area of a map. Image matching can be performed only for points within this area. As another example, based on tracking, the XR system can know that the XR device was previously near a first persistent gesture in a map and is currently moving in the direction of a second persistent gesture in the map. This second persistent gesture can be considered the most recent persistent gesture, and the keyframes stored with it can be considered the most recent keyframes. Alternatively or additionally, other metadata such as GPS data or WiFi fingerprints may be used to select the most recent keyframe or set of most recent keyframes.
[0336] Regardless of how the nearest keyframe is selected, the frame descriptors can be used to determine whether the new image matches any of the frames selected to be associated with the nearby persistent gesture. This determination can be made by comparing the frame descriptor of the new image with the frame descriptors of the nearest keyframe, or with the frame descriptors of a subset of keyframes in the database selected in any other suitable manner, and selecting a keyframe having a frame descriptor within a predetermined distance of the frame descriptor of the new image. In some embodiments, the distance between two frame descriptors can be calculated by taking the difference between two numeric strings that can represent the two frame descriptors. In embodiments where the strings are processed as multiple strings, the difference can be calculated as a vector difference.
[0337] Once a matching image frame is identified, the orientation of the XR device relative to the image frame can be determined. Method 2300 may include performing (act 2306) feature matching on 3D features in the map corresponding to the most recently identified keyframe, and calculating (act 2308) the pose of the device worn by the user based on the feature matching results. In this way, computationally dense matching of feature points in two images can be performed for as few as one image that has been determined to be a possible match for the new image.
[0338] Figure 2424 is a flow chart illustrating a method 2400 for training a neural network according to some embodiments. Method 2400 may begin by generating (act 2402) a dataset comprising a plurality of image sets. Each of the plurality of image sets may include a query image, a positive sample image, and a negative sample image. In some embodiments, the plurality of image sets may include synthetic record pairs configured to, for example, teach the neural network basic information (such as shape). In some embodiments, the plurality of image sets may include real record pairs, which may be based on physical world records.
[0339] In some embodiments, inliers can be computed by fitting a fundamental matrix between the two images. In some embodiments, sparse overlap can be computed as the intersection over union (IoU) of the points of interest seen in the two images. In some embodiments, a positive sample can include at least twenty of the same points of interest as in the query image as inliers. A negative sample can include fewer than ten inliers. A negative sample can have fewer than half of its sparse points overlap with the sparse points of the query image.
[0340] The method 2400 may include calculating (act 2404) a loss for each image set by comparing the query image with the positive and negative images. The method 2400 may include modifying (act 2406) the artificial neural network based on the calculated loss so that a distance between a frame descriptor generated by the artificial neural network for the query image and a frame descriptor for the positive image is smaller than a distance between a frame descriptor for the query image and a frame descriptor for the negative image.
[0341] It should be understood that although the above description describes methods and apparatus configured to generate global descriptors for individual images, the methods and apparatus can be configured to generate descriptors for individual maps. For example, a map can include multiple keyframes, each of which can have a frame descriptor as described above. The max pooling unit can analyze the frame descriptors of the keyframes of a map and combine the frame descriptors into a unique map descriptor for that map.
[0342] Furthermore, it should be understood that other architectures can be used for the processing described above. For example, separate neural networks are described for generating DSF descriptors and frame descriptors. This approach is computationally efficient. However, in some embodiments, frame descriptors can be generated based on selected feature points without first generating DSF descriptors.
[0343] Ranking and merging maps
[0344] Described herein are methods and apparatus for ranking and merging multiple maps of an environment in an X-reality (XR) system. Map merging can enable maps representing overlapping portions of the physical world to be combined to represent a larger area. Ranking maps can enable efficient performance of the techniques described herein, including map merging, which involves selecting maps from a set of maps based on similarity. In some embodiments, for example, a system may maintain a set of canonical maps that are formatted in a manner that can be accessed by any of a number of XR devices. These canonical maps can be formed by merging selected tracking maps from those devices with other tracking maps or previously stored canonical maps. Canonical maps can be ranked, for example, for selecting one or more canonical maps to merge with a new tracking map and / or selecting one or more canonical maps from a set to use in a device.
[0345] To provide users with a realistic XR experience, the XR system must understand the user's physical environment so that it can correctly associate the positions of virtual objects with respect to real objects. Information about the user's actual environment can be obtained from an environment map at the user's location.
[0346] The inventors have recognized and appreciated that an XR system can provide an enhanced XR experience to multiple users sharing the same world including real and / or virtual content by enabling efficient sharing of environmental maps of the real / physical world collected by multiple users, regardless of whether the users are present in the world at the same time or at different times. However, there are significant challenges in providing such a system. Such a system may store multiple maps generated by multiple users and / or the system may store multiple maps generated at different times. For operations that may be performed using previously generated maps (such as, for example, positioning as described above), a significant amount of processing may be required to identify the relevant environmental maps of the same world (e.g., the same real-world location) from all the environmental maps collected by the XR system. In some embodiments, there may be only a small number of environmental maps that a device can access, for example, for positioning. In some embodiments, there may be a large number of environmental maps that a device can access. The inventors have recognized and appreciated that it is necessary to quickly and accurately share environmental maps from all possible environmental maps (such as, for example, Figure 28 A technique for ranking the relevance of environment maps in a universe of all canonical maps 120 is provided. Highly ranked maps can then be selected for further processing, such as rendering virtual objects on a user's display so that they interact realistically with the physical world around the user, or merging the user's collected map data with stored maps to create a larger or more accurate map.
[0347] In some embodiments, stored maps relevant to a user's task at a location in the physical world can be identified by filtering the stored maps based on multiple criteria. These criteria can indicate a comparison of a tracking map generated by the user's wearable device in the location with candidate environment maps stored in a database. The comparison can be performed based on metadata associated with the map, such as a Wi-Fi fingerprint detected by the device generating the map and / or a set of BSSIDs to which the device was connected while forming the map. The comparison can also be performed based on compressed or uncompressed content of the map. Comparisons based on compressed representations can be performed by comparing vectors calculated from the map content. For example, comparisons based on uncompressed maps can be performed by locating a tracking map within a stored map, or vice versa. Multiple comparisons can be performed in sequence based on the computation time required to reduce the number of candidate maps to be considered, where comparisons involving less computation will be performed earlier in the sequence than other comparisons requiring more computation.
[0348] Figure 26 An AR system 800 configured to rank and merge one or more environment maps according to some embodiments is depicted. The AR system may include a traversable world model 802 of the AR device. Information to populate the traversable world model 802 may come from sensors on the AR device, which may include data stored in a processor 804 (e.g., Figure 4 The computer executable instructions in the local data processing module 570 in the processor can perform some or all of the processing to convert the sensor data into a map. This map can be a tracking map, because the AR device can collect sensor data while building a tracking map as it operates in the area. Figure 1 In addition, regional attributes may be provided to indicate the area represented by the tracking map. These regional attributes may be geographic location identifiers, such as coordinates expressed as latitude and longitude, or IDs used by the AR system to represent locations. Alternatively or in addition, regional attributes may be measured characteristics that have a high probability of being unique to the area. Regional attributes may, for example, be derived from parameters of wireless networks detected in the area. In some embodiments, regional attributes may be associated with unique addresses of access points that the AR system is near and / or connected to. For example, regional attributes may be associated with MAC addresses or basic service set identifiers (BSSIDs) of 5G base stations / routers, Wi-Fi routers, and the like.
[0349] exist Figure 26 In the example of FIG, the tracking map can be merged with other maps of the environment. The map ranking portion 806 receives the tracking map from the device PW 802 and communicates with the map database 808 to select and rank the environment maps from the map database 808. The selected maps with higher rankings are sent to the map merging portion 810.
[0350] The map merging section 810 can perform a merging process on the maps sent from the map ranking section 806. This merging process may entail merging the tracking map with some or all of the ranked maps and sending the new merged map to the traversable world model 812. The map merging section can merge the maps by identifying maps that depict overlapping portions of the physical world. These overlapping portions can be aligned so that the information from both maps can be aggregated into the final map. The canonical map can be merged with other canonical maps and / or tracking maps.
[0351] Aggregation may require extending one map with information from another map. Alternatively or additionally, aggregation may require adjusting the representation of the physical world in one map based on information in another map. For example, the later map may reveal objects that have caused feature points to move, so that the map can be updated based on the later information. Alternatively, two maps may characterize the same area with different feature points, and aggregation may require selecting a set of feature points from both maps to better represent the area. Regardless of the specific processing that occurs during the merge, in some embodiments, the PCFs from all maps being merged may be retained so that applications that locate content relative to them can continue to do so. In some embodiments, the merging of maps may result in redundant persistent gestures, and some persistent gestures may be deleted. When a PCF is associated with a persistent gesture that is to be deleted, merging the maps may require modifying the PCF to be associated with the persistent gesture that remains in the map after the merge.
[0352] In some embodiments, as maps are expanded and / or updated, they may be refined. Refinement may require computations to reduce internal inconsistencies between feature points that may represent the same object in the physical world. Such inconsistencies may arise from pose inaccuracies associated with keyframes that provide feature points that represent the same object in the physical world. For example, such inconsistencies may arise from the XR device computing a pose relative to a tracking map that is in turn built based on an estimated pose, so that errors in the pose estimate accumulate, causing a "drift" in pose accuracy over time. Maps may be refined by performing bundle adjustment or other operations to reduce inconsistencies in feature points from multiple keyframes.
[0353] During refinement, the position of a persistent point relative to the map origin can change. Consequently, the transform associated with that persistent point, such as the persistent pose or PCF, may change. In some embodiments, an XR system incorporating map refinement (whether performed as part of a merge operation or for other reasons) can recalculate the transforms associated with any persistent point that has changed. These transforms may be pushed from the component that computes the transform to the component that uses it so that any use of the transform is based on the updated position of the persistent point.
[0354] The navigable world model 812 can be a cloud model that can be shared by multiple AR devices. The navigable world model 812 can store or otherwise access a map of the environment in the map database 808. In some embodiments, when a previously calculated map of the environment is updated, the previous version of the map can be deleted to remove the outdated map from the database. In some embodiments, when a previously calculated map of the environment is updated, the previous version of the map can be archived, thereby enabling the previous version of the environment to be retrieved / viewed. In some embodiments, permissions can be set so that only AR systems with certain read / write access permissions can trigger the deletion / archiving of the previous version of the map.
[0355] These environment maps, created from tracking maps provided by one or more AR devices / systems, can be accessed by AR devices in the AR system. A map ranking portion 806 can also be used to provide environment maps to AR devices. An AR device can send a message requesting an environment map for its current location, and the map ranking portion 806 can be used to select and rank environment maps relevant to the requesting device.
[0356] In some embodiments, the AR system 800 may include a downsampling unit 814 configured to receive a merged map from a cloud PW 812. The merged map received from the cloud PW 812 may be in a storage format for the cloud, which may include high-resolution information, such as a large number of PCFs or multiple image frames per square meter or a large number of feature point sets associated with PCFs. The downsampling unit 814 may be configured to downsample the cloud-formatted map to a format suitable for storage on an AR device. The device-formatted map may contain less data, such as fewer PCFs or less data stored for each PCF, to accommodate the limited local computing power and storage space of the AR device.
[0357] Figure 27is a simplified block diagram illustrating a plurality of canonical maps 120 that may be stored in a remote storage medium, such as a cloud. Each canonical map 120 may include a plurality of canonical map identifiers that indicate the location of the canonical map in physical space, such as somewhere on the Earth. These canonical map identifiers may include one or more of the following identifiers: a region identifier represented by a longitude and latitude range, a frame descriptor (e.g., Figure 21 ), Wi-Fi fingerprints, feature descriptors (e.g., Figure 21 ), and device identifications indicating one or more devices contributing to the map.
[0358] In the example shown, the canonical maps 120 are arranged geographically in a two-dimensional pattern as they would exist on the surface of the Earth. The canonical maps 120 are uniquely identifiable by corresponding longitudes and latitudes, as any canonical maps with overlapping longitudes and latitudes can be merged into a new canonical map.
[0359] Figure 28 is a schematic diagram illustrating a method for selecting a canonical map according to some embodiments, which method can be used to locate a new tracking map to one or more canonical maps. The method can begin by accessing (act 120) a world of canonical maps 120, which can be stored in a database of traversable worlds (e.g., traversable world module 538), as an example. The world of canonical maps can include canonical maps from all previously visited locations. The XR system can filter the world of all canonical maps to a small subset or just one map. It should be understood that in some embodiments, due to bandwidth limitations, it is not possible to send all canonical maps to the viewing device. Selecting a subset of those selected as possible candidates for matching the tracking map to send to the device can reduce the bandwidth and latency associated with accessing a remote database of maps.
[0360] The method may include filtering (act 300) the world of the canonical map based on regions having a predetermined size and shape. Figure 27In the example of , each square can represent an area. Each square can cover 50m×50m. Each square can have six adjacent areas. In some embodiments, action 300 can select at least one matching canonical map 120 covering longitude and latitude, where the longitude and latitude include the longitude and latitude of the location identifier received from the XR device, as long as there is at least one map at that longitude and latitude. In some embodiments, action 300 can select at least one adjacent canonical map covering longitude and latitude adjacent to the matching canonical map. In some embodiments, action 300 can select multiple matching canonical maps and multiple adjacent canonical maps. Action 300 can, for example, reduce the number of canonical maps by approximately ten times, for example from thousands to hundreds, to form a first filtered selection. Alternatively or additionally, criteria other than latitude and longitude can be used to identify adjacent maps. For example, the XR device may have previously been positioned using a canonical map in the collection as part of the same session. The cloud service can retain information about the XR device, including the maps previously positioned. In this example, the maps selected at act 300 may include those maps covering an area adjacent to the map at which the XR device is positioned.
[0361] The method may include filtering (act 302) a first filtered selection of canonical maps based on a Wi-Fi fingerprint. Act 302 may determine a latitude and longitude based on the Wi-Fi fingerprint received from the XR device as part of a location identifier. Act 302 may compare the latitude and longitude from the Wi-Fi fingerprint with the latitude and longitude of the canonical map 120 to determine one or more canonical maps that form a second filtered selection. Act 302 may reduce the number of canonical maps by approximately ten times, for example, from hundreds of canonical maps to dozens (e.g., 50) of canonical maps that form the second selection. For example, the first filtered selection may include 130 canonical maps, the second filtered selection may include 50 of the 130 canonical maps, and may not include the other 80 of the 130 canonical maps.
[0362] The method may include filtering (act 304) a second filter selection of the canonical map based on the keyframes. Act 304 may compare data representing an image captured by the XR device with data representing the canonical map 120. In some embodiments, the data representing the image and / or map may include feature descriptors (e.g., Figure 25 DSF descriptor in ) and / or global feature string (e.g., Figure 21316 in). Action 304 may provide a third filtered selection of canonical maps. In some embodiments, for example, the output of action 304 may be only five canonical maps out of the 50 canonical maps identified after the second filtered selection. The map transmitter 122 then sends one or more canonical maps based on the third filtered selection to the viewing device. Action 304 may reduce the number of canonical maps by approximately ten times, for example, from dozens of canonical maps to a single-digit number of canonical maps (e.g., 5) forming the third selection. In some embodiments, the XR device may receive the canonical map in the third filtered selection and attempt to locate within the received canonical map.
[0363] For example, action 304 may filter the canonical map 120 based on the global feature string 316 of the canonical map 120 and based on the global feature string 316 of an image captured by the viewing device (e.g., an image that may be part of a user's local tracking map). Figure 27 Each canonical map 120 in the image processing system has one or more global feature strings 316 associated therewith. In some embodiments, the global feature strings 316 can be obtained when the XR device submits images or feature details to the cloud and the cloud processes these images or feature details to generate the global feature strings 316 for the canonical map 120.
[0364] In some embodiments, the cloud can receive feature details of a live / new / current image captured by the viewing device, and the cloud can generate a global feature string 316 for the live image. The cloud can then filter the canonical map 120 based on the live global feature string 316. In some embodiments, the global feature string can be generated locally on the viewing device. In some embodiments, the global feature string can be generated remotely, for example, in the cloud. In some embodiments, the cloud can send the filtered canonical map to the XR device along with the global feature string 316 associated with the filtered canonical map. In some embodiments, when the viewing device positions its tracking map to the canonical map, it can do so by matching the global feature string 316 of the local tracking map with the global feature string of the canonical map.
[0365] It should be understood that the operation of the XR device may not perform all of the actions (300, 302, 304). For example, if the world of the canonical map is relatively small (e.g., 500 maps), the XR device attempting to perform localization may filter the world of the canonical map based on Wi-Fi fingerprints (e.g., action 302) and keyframes (e.g., action 304), but omit the region-based filtering (e.g., action 300). Moreover, it is not necessary to compare the entire map. For example, in some embodiments, the comparison of the two maps may result in the identification of common persistent points, such as persistent gestures or PCFs that appear in both the new map and the map selected from the map world. In that case, descriptors may be associated with the persistent points, and those descriptors may be compared.
[0366] Figure 29 is a flow chart illustrating a method 900 for selecting one or more ranked environment maps, according to some embodiments. In the illustrated embodiment, the ranking is performed on the AR device of the user creating the tracking map. Thus, the tracking map can be used to rank the environment maps. In embodiments where a tracking map is unavailable, some or all of the selection and ranking of the environment maps can be used that do not explicitly rely on the tracking map.
[0367] Method 900 may begin with act 902, where a set of maps in a database of environmental maps (which may be formatted as canonical maps) located near a location where a tracking map is to be formed may be accessed and then filtered for ranking. Additionally, at act 902, at least one area attribute of the area in which the user's AR device is operating is determined. In a scenario where the user's AR device is constructing a tracking map, the area attribute may correspond to the area in which the tracking map is to be created. As a specific example, the area attribute may be calculated based on signals received from an access point to a computer network while the AR device is calculating the tracking map.
[0368] Figure 30 An exemplary map ranking portion 806 of the AR system 800 is depicted in accordance with some embodiments. The map ranking portion 806 can be executed in a cloud computing environment, as it can include a portion executed on the AR device and a portion executed on a remote computing system such as the cloud. The map ranking portion 806 can be configured to perform at least a portion of the method 900.
[0369] Figure 31AAn example of area attributes AA1-AA8 of a tracking map (TM) 1102 and environment maps CM1-CM4 in a database according to some embodiments is depicted. As shown, the environment map can be associated with multiple area attributes. The area attributes AA1-AA8 can include parameters of a wireless network detected by the AR device computing the tracking map 1102, such as a basic service set identifier (BSSID) of the network to which the AR device is connected and / or the strength of a received signal to an access point of the wireless network, such as a network tower 1104. The parameters of the wireless network can conform to protocols including Wi-Fi and 5G NR. Figure 32 In the example shown in , the region attribute is a fingerprint of the area in which the user AR device collects sensor data to form a tracking map.
[0370] Figure 31B An example of a determined geographic location 1106 for a tracking map 1102 is depicted in accordance with some embodiments. In the example shown, the determined geographic location 1106 includes a centroid point 1110 and an area 1108 surrounding the centroid point. It should be understood that the determination of the geographic location of the present application is not limited to the format shown. The determined geographic location can have any suitable format, including, for example, different area shapes. In this example, the geographic location is determined from area attributes using a database that associates area attributes with geographic locations. Commercially available databases, such as those that associate Wi-Fi fingerprints with locations expressed as latitude and longitude, can be used for this operation.
[0371] exist Figure 29 In an embodiment, the map database containing environment maps may also include location data for those maps, including the latitudes and longitudes covered by the maps. Processing at action 902 may entail selecting a set of environment maps from the database that cover the same latitudes and longitudes determined for the area attributes of the tracking map.
[0372] Act 904 is a first filtering of the set of environment maps accessed in act 902. In act 902, environment maps are retained in the set based on proximity to the geographic location of the tracking map. This filtering step can be performed by comparing the latitude and longitude associated with the tracking map and the environment maps in the set.
[0373] Figure 32An example of action 904 according to some embodiments is depicted. Each area attribute can have a corresponding geographic location 1202. The set of environment maps can include an environment map having at least one area attribute with a geographic location that overlaps with the determined geographic location of the tracking map. In the example shown, the set of identified environment maps includes environment maps CM1, CM2, and CM4, each of which has at least one area attribute with a geographic location that overlaps with the determined geographic location of the tracking map 1102. CM3, which is associated with area attribute AA6, is not included in the set because it is outside the determined geographic location of the tracking map.
[0374] Other filtering steps may also be performed on the set of environment maps to reduce / rank the number of environment maps ultimately processed in the set (such as for map merging or providing navigable world information to a user device). Method 900 may include filtering (act 906) the set of environment maps based on the similarity of one or more identifiers of network access points associated with the tracking map and the environment maps in the set of environment maps. During map formation, the device collecting sensor data to generate the map may be connected to a network via a network access point (such as via Wi-Fi or a similar wireless communication protocol). The access point may be identified by a BSSID. As the user device moves through the area where data is collected to form the map, the user device may connect to multiple different access points. Similarly, when multiple devices provide information to form the map, the devices may have connected via different access points, and for this reason, multiple access points may also be used when forming the map. Therefore, there may be multiple access points associated with the map, and the set of access points may be an indication of the map's location. The signal strength from the access points may be reflected as an RSSI value, which may provide further geographic information. In some embodiments, a list of BSSIDs and RSSI values may form a regional attribute for the map.
[0375] In some embodiments, filtering the set of environment maps based on similarity of one or more identifiers of the network access points may include retaining in the set of environment maps an environment map having a highest Jaccard similarity to at least one area attribute of the tracking map based on the one or more identifiers of the network access points. Figure 33 An example of action 906 according to some embodiments is depicted. In the example shown, a network identifier associated with area attribute AA7 may be determined as the identifier of tracking map 1102. Following action 906, the set of environment maps includes: environment map CM2, which may have an area attribute with a higher Jaccard similarity than AA7; and environment map CM4, which also includes area attribute AA7. Environment map CM1 is not included in the set because it has the lowest Jaccard similarity with AA7.
[0376] The processing of actions 902-906 can be performed based on metadata associated with the map without actually accessing the content of the map stored in the map database. Other processing may involve accessing the content of the map. Action 908 indicates accessing the environment map retained in the subset after filtering based on the metadata. It should be understood that this action can be performed earlier or later in the process if subsequent operations can be performed on the accessed content.
[0377] Method 900 may include filtering (act 910) a set of environment maps based on similarity of metrics representing the content of the tracking map and the environment maps of the set of environment maps. The metrics representing the content of the tracking map and the environment maps may include vectors of values calculated from the content of the maps. For example, as described above, the depth keyframe descriptors calculated for one or more keyframes used to form the map may provide metrics for comparing maps or portions of maps. The metrics may be calculated from the maps retrieved at act 908, or may be pre-calculated and stored as metadata associated with those maps. In some embodiments, filtering the set of environment maps based on similarity of metrics representing the content of the tracking map and the environment maps of the set of environment maps may include retaining in the set of environment maps the environment map having the smallest vector distance between a feature vector of the tracking map and a vector representing an environment map in the set of environment maps.
[0378] Method 900 may include further filtering (action 912) a set of environment maps based on a degree of match between a portion of the tracking map and a portion of an environment map of the set of environment maps. The degree of match may be determined as part of a localization process. As a non-limiting example, localization may be performed by identifying critical points in the tracking map and the environment map that are sufficiently similar that they may represent the same portion of the physical world. In some embodiments, the critical points may be features, feature descriptors, keyframes, key assemblies, persistent poses, and / or PCFs. A set of critical points in the tracking map may then be aligned to produce an optimal fit with the set of critical points in the environment map. A mean squared distance may be calculated between corresponding critical points and, if below a threshold for a particular area of the tracking map, used as an indication that the tracking map and the environment map represent the same area of the physical world.
[0379] In some embodiments, filtering a set of environment maps based on a degree of match between a portion of the tracking map and a portion of an environment map in the set of environment maps may include: calculating a volume of the physical world represented by the tracking map, the tracking map also represented in an environment map in the set of environment maps; and retaining in the set of environment maps an environment map having a larger calculated volume than an environment map filtered from the set. Figure 34An example of action 912 according to some embodiments is depicted. In the example shown, the set of environment maps after action 912 includes environment map CM4, which has an area 1402 that matches an area of tracking map 1102. Environment map CM1 is not included in the set because it does not have an area that matches an area of tracking map 1102.
[0380] In some embodiments, the set of environment maps may be filtered in the order of act 906, act 910, and act 912. In some embodiments, the set of environment maps may be filtered based on act 906, act 910, and act 912, which may be performed in the order of processing required to perform the filtering, from lowest to highest. Method 900 may include loading (act 914) the set of environment maps and data.
[0381] In the example shown, the user database stores a region identifier indicating the region in which the AR device is used. The region identifier may be a region attribute that may include parameters of a wireless network detected by the AR device during use. The map database may store multiple environment maps constructed from data provided by the AR device and associated metadata. The associated metadata may include a region identifier derived from the region identifier of the AR device providing the data, and the environment map is constructed from the data. The AR device may send a message to the PW module indicating that a new tracking map has been created or is being created. The PW module may calculate a region identifier for the AR device and update the user database based on the received parameters and / or the calculated region identifier. The PW module may also determine a region identifier associated with the AR device requesting the environment map, identify the set of environment maps from the map database based on the region identifier, filter the set of environment maps, and send the filtered set of environment maps to the AR device. In some embodiments, the PW module can filter the set of environment maps based on one or more criteria, including, for example, the geographic location of the tracking map, the similarity of one or more identifiers of network access points associated with the tracking map and the environment maps of the set of environment maps, the similarity of metrics representing the content of the tracking map and the environment maps of the set of environment maps, and the degree of match between a portion of the tracking map and a portion of the environment maps of the set of environment maps.
[0382] Having thus described several aspects of some embodiments, it will be appreciated that various changes, modifications, and improvements will readily occur to those skilled in the art. As an example, embodiments are described in conjunction with an augmented reality (AR) environment. It will be appreciated that some or all of the techniques described herein may be applied in an MR environment or, more generally, in other XR and VR environments.
[0383] As another example, embodiments are described in conjunction with devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via any suitable combination of a network (such as the cloud), a separate application, and / or a device, network, and separate application.
[0384] also, Figure 29 Examples of criteria that can be used to filter candidate maps to produce a set of highly ranked maps are provided. Other criteria can be used instead of or in addition to the criteria described. For example, if multiple candidate maps have similar values for a metric used to filter out less desirable maps, the characteristics of the candidate maps can be used to determine which maps are retained as candidates or filtered out. For example, larger or more densely populated candidate maps can be prioritized over smaller candidate maps. In some embodiments, Figures 27-28 Can describe Figure 29-34 All or part of the systems and methods described in.
[0385] Figure 35 and 36 is a diagram illustrating an XR system configured to rank and merge multiple maps of an environment according to some embodiments. In some embodiments, a traversable world (PW) may determine when to trigger ranking and / or merging of maps. In some embodiments, determining which map to use may be based at least in part on the above description of the map. Figures 21 to 25 Depth keyframes described.
[0386] Figure 37 is a block diagram illustrating a method 3700 for creating an environment map of the physical world according to some embodiments. The method 3700 may be performed from positioning a tracking map captured by an XR device worn by a user (act 3702) to a canonical map (e.g., by Figure 28 900 ). Action 3702 may include positioning key assemblies of the tracking map within the group of the canonical maps. The positioning result for each key assembly may include a localized pose of the key assembly and a set of 2D to 3D feature correspondences.
[0387] In some embodiments, method 3700 may include splitting the tracking map into connected parts (act 3704), which may be used to robustly merge maps by merging connected segments. Each connected segment may include a critical assembly within a predetermined distance. Method 3700 may include: merging connected segments greater than a predetermined threshold into one or more canonical maps (act 3706); and removing the merged connected segments from the tracking map.
[0388] In some embodiments, method 3700 may include merging (act 3708) canonical maps in a group that are merged with the same connected portion of the tracking map. In some embodiments, method 3700 may include promoting (act 3710) the remaining connected portions of the tracking map that have not yet been merged with any canonical map to canonical maps. In some embodiments, method 3700 may include merging (act 3712) the persistent poses and / or PCFs of the tracking map and the canonical map, wherein the canonical map is merged with at least one connected portion of the tracking map. In some embodiments, method 3700 may include finalizing (act 3714) the canonical map, for example, by fusing map points and pruning redundant key assemblies.
[0389] Figure 38A and 38B An environment map 3800 is shown that is created by updating the canonical map 700, which can be updated with a new tracking map from the tracking map 700 ( Figure 7 ) to upgrade. Figure 7 As shown and described, the canonical map 700 may provide a floor plan 706 of a reconstructed physical object represented by point 702 in the corresponding physical world. In some embodiments, the map point 702 may represent a feature of the physical object, which may include multiple features. A new tracking map of the physical world may be captured and uploaded to the cloud to be merged with the map 700. The new tracking map may include map point 3802 and key assemblies 3804, 3806. In the example shown, the key assembly 3804 represents a physical object that has been reconstructed by, for example, establishing a correspondence with the key assembly 704 of the map 700 (e.g., Figure 38B On the other hand, key assembly 3806 represents a key assembly that has not yet been located to map 700. In some embodiments, key assembly 3806 can be promoted to a separate canonical map.
[0390] Figures 39A to 39F is a diagram illustrating an example of a cloud-based persistent coordinate system that provides a shared experience for users in the same physical space. Figure 39A For example, a canonical map 4814 from the cloud is shown. 20A to 20C The canonical map 4814 may have a canonical coordinate frame 4806C. The canonical map 4814 may have a PCF 4810C with multiple associated PPs (e.g., Figure 39C 4818A, 4818B).
[0391] Figure 39BThe relationship established between the XR devices' respective world coordinate systems 4806A, 4806B and the canonical coordinate frame 4806C is shown. This can be accomplished, for example, by positioning the tracking map on the canonical map 4814 on the respective devices. For each device, positioning the tracking map on the canonical map can result in a transformation between each device's local world coordinate system and the coordinate system of the canonical map.
[0392] Figure 39C As a result of positioning, a transformation (e.g., transformation 4816A, transformation 4816B) between a local PCF (e.g., PCF 4810A, PCF 4810B) on a respective device and a corresponding persistent pose (e.g., PP 4818A, PP 4818B) on a canonical map can be calculated. Using these transformations, each device can use its local PCF to determine where to display virtual content attached to PP 4818A, PP 4818B, or other persistent points on the canonical map relative to the local device, where the local PCF can be detected locally on the device by processing images detected using sensors on the device. Such an approach can accurately position virtual content relative to each user and can enable each user to have the same experience of virtual content in physical space.
[0393] Figure 39D A snapshot of the persistent pose from the canonical map to the local tracking map is shown. It can be seen that the local tracking maps are connected to each other through the persistent pose. Figure 39E It is shown that PCF 4810A on the device worn by user 4802A can be accessed in the device worn by user 4802B through PP 4818A. Figure 39F Tracking maps 4804A, 4804B and canonical map 4814 are shown as being able to be merged. In some embodiments, some PCFs may be removed due to the merge. In the example shown, the merged map includes PCF 4810C from canonical map 4814, but does not include PCFs 4810A and 4810B from tracking maps 4804A and 4804B. After the map merge, PPs previously associated with PCFs 4810A and 4810B may be associated with PCF 4810C.
[0394] Example
[0395] Figure 40 and Figure 41 Shown by Figure 9 An example of the first XR device 12.1 using tracking maps. Figure 40 is a three-dimensional first local tracking map (ground Figure 1 ), which can be represented by Figure 9The first XR device of the generation. Figure 41 is a diagram illustrating the Figure 9 The first XR device uploads the address to the server Figure 1 Block diagram of .
[0396] Figure 40 Shows the location on the first XR device 12.1 Figure 1 and virtual content (content 123 and content 456). Figure 1 Has an origin (origin 1). Figure 1 The first XR device 12.1 includes a number of PCFs (PCF a to PCF d). From the perspective of the first XR device 12.1, PCF a is positioned at the ground, for example. Figure 1 , and PCF b has X, Y, and Z coordinates of (0, 0, 0). Content 123 is associated with PCF a. In this example, content 123 has an X, Y, and Z relationship relative to PCF a of (1, 0, 0). Content 456 has a relationship relative to PCF b. In this example, content 456 has an X, Y, and Z relationship relative to PCF b of (1, 0, 0).
[0397] exist Figure 41 The first XR device 12.1 will be Figure 1 Uploaded to the server 20. In this example, since the server does not store a canonical map for the same area of the physical world represented by the tracking map, the tracking map is stored as the initial canonical map. The server 20 now has a map based on the map. Figure 1 The first XR device 12.1 has a canonical map that is empty at this stage. For the purposes of discussion, and in some embodiments, the server 20 has a canonical map in addition to the map. Figure 1 No other maps are included. No maps are stored on the second XR device 12.2.
[0398] The first XR device 12.1 also sends its Wi-Fi signature data to the server 20. The server 20 can use the Wi-Fi signature data to determine the approximate location of the first XR device 12.1 based on intelligence gathered from other devices that have connected to the server 20 or other servers in the past, along with the recorded GPS locations of such other devices. The first XR device 12.1 can now end the first session (see Figure 8 ), and can disconnect from the server 20.
[0399] Figure 42 is a diagram showing a method according to some embodiments Figure 16Schematic diagram of an XR system showing that after the first user 14 . 1 terminated the first session, the second user 14 . 2 has initiated a second session using a second XR device of the XR system. Figure 43A A block diagram shows a second user 14.2 initiating a second session. Because the first session of the first user 14.1 has ended, the first user 14.1 is shown in dashed lines. The second XR device 12.2 begins recording the object. The server 20 may use various systems with different granularity to determine that the second session of the second XR device 12.2 is in the same vicinity as the first session of the first XR device 12.1. For example, Wi-Fi signature data, Global Positioning System (GPS) location data, GPS data based on Wi-Fi signature data, or any other data indicative of location may be included in the first XR device 12.1 and the second XR device 12.2 to record their locations. Alternatively, the PCF identified by the second XR device 12.2 may be displayed in a manner consistent with the location. Figure 1 The similarity of PCF.
[0400] like Figure 43B As shown in , the second XR device starts up and begins collecting data, such as images 1110 from one or more cameras 44, 46. Figure 14 As shown in FIG, in some embodiments, an XR device (e.g., a second XR device 12.2) may collect one or more images 1110 and perform image processing to extract one or more features / points of interest 1120. Each feature may be converted into a descriptor 1130. In some embodiments, the descriptor 1130 may be used to describe a keyframe 1140, which may have the position and orientation of an associated image attached. One or more keyframes 1140 may correspond to a single persistent pose 1150, which may be automatically generated after a threshold distance (e.g., 3 meters) from a previous persistent pose 1150. One or more persistent poses 1150 may correspond to a single PCF 1160, which may be automatically generated after a predetermined distance (e.g., every 5 meters). Over time, as the user continues to move around the user's environment and the XR device continues to collect more data (such as images 1110), additional PCFs (e.g., PCF 3 and PCFs 4 and 5) may be created. One or more applications 1180 may run on the XR device and provide virtual content 1170 to the XR device for presentation to the user. The virtual content may have an associated content coordinate frame that may be positioned relative to one or more PCFs. Figure 43B As shown in , the second XR device 12 . 2 creates three PCFs. In some embodiments, the second XR device 12 . 2 may attempt to locate one or more canonical maps stored on the server 20 .
[0401] In some embodiments, as Figure 43C As shown in FIG, the second XR device 12.2 can download the canonical map 120 from the server 20. Figure 1 Includes PCFs a through d and origin 1. In some embodiments, server 20 may have multiple canonical maps for various locations and may determine that second XR device 12.2 is located in the same vicinity as first XR device 12.1 during the first session and send a canonical map of that vicinity to second XR device 12.2.
[0402] Figure 44 The second XR device 12.2 is shown to start identifying PCF for generating Figure 2 The second XR device 12.2 recognizes only a single PCF, PCF 1,2. The X, Y, and Z coordinates of PCF 1,2 of the second XR device 12.2 may be (1, 1, 1). Figure 2 The second XR device 12.2 may have its own origin (origin 2), which may be based on the head pose of device 2 at the start of the current head pose session. In some embodiments, the second XR device 12.2 may immediately attempt to Figure 2 In some embodiments, because the system cannot identify any or sufficient overlap between the two maps, the map Figure 2 It may not be possible to locate the standard map (map Figure 1 ) (i.e., localization may fail). Localization can be performed by identifying a portion of the physical world represented in the first map that is also represented in the second map, and computing the transformation between the first map and the second map required to align these portions. In some embodiments, the system can localize based on a PCF comparison between the local map and the canonical map. In some embodiments, the system can localize based on a persistent pose comparison between the local map and the canonical map. In some embodiments, the system can localize based on a keyframe comparison between the local map and the canonical map.
[0403] Figure 45 The second XR device 12.2 identifies the Figure 2 The other PCFs (PCF 1, 2, PCF 3, PCF 4, 5) after Figure 2 The second XR device 12.2 tries again to Figure 2 Locate to the standard map. Figure 2 has been extended to overlap at least a portion of the canonical map, so the positioning attempt will succeed. Figure 2 The overlap between the and canonical maps can be represented by PCFs, persistent poses, keyframes, or any other suitable intermediate or derived construct.
[0404] In addition, the second XR device 12.2 has shared content 123 and content 456 with the local Figure 2 Content 123 has X, Y, and Z coordinates (1, 0, 0) relative to PCFs 1, 2, and 3. Similarly, content 123 has X, Y, and Z coordinates (1, 0, 0) relative to PCFs 1, 2. Figure 2 In PCF 3, the X, Y, and Z coordinates of content 456 are (1, 0, 0).
[0405] Figure 46A and Figure 46B Shown Figure 2 The successful localization to the canonical map. Localization can be based on matching features in one map to another. By appropriate transformations, which involve translation and rotation of one map relative to the other, the overlapping area / volume / section of the maps 1410 represents the geo Figure 1 and the common part of the normative map. Figure 2 PCFs 3, 4, and 5 were created before positioning, and the normative map was created before the Figure 2 PCFs a and c were created previously, so different PCFs are created to represent the same volume in real space (e.g., in different maps).
[0406] like Figure 47 As shown in FIG, the second XR device 12.2 extends the Figure 2 , to include the PCF ad from the canonical map. Including the PCF ad indicates that the Figure 2 In some embodiments, the XR system may perform an optimization step to remove duplicate PCFs from overlapping regions, such as PCF 1410, PCF 3, and PCFs 4 and 5. Figure 2 After positioning, the placement of virtual content (such as content 456 and content 123) will be updated. Figure 2 The virtual content appears in the same real-world location relative to the user, despite the change to the content's PCF attachment, and despite the update of the map. Figure 2 PCF.
[0407] like Figure 48 As shown in FIG, the second XR device 12.2 continues to expand Figure 2 , for example, when the user moves around the real world, the second XR device 12.2 will identify other PCFs (PCF e, f, g, and h). It should also be noted that the ground Figure 1 exist Figure 47 and Figure 48 There are no extensions in .
[0408] refer to Figure 49 , the second XR device 12.2 will be Figure 2Upload to server 20. Server 20 will Figure 2 With normative Figure 1 In some embodiments, when the session for the second XR device 12.2 ends, the address Figure 2 It can be uploaded to the server 20.
[0409] The canonical map within the server 20 now includes PCF i, which is not included in the map on the first XR device 12.1. Figure 1 When a third XR device (not shown) uploads a map to the server 20 and the map includes PCF i, the canonical map on the server 20 may have been extended to include PCF i.
[0410] exist Figure 50 In the process, the server 20 will Figure 2 The server 20 determines the PCFs a to d for the normative map and the map. Figure 2 is common. The server extends the specification map to include PCF e to h and from ground Figure 2 The PCFs 1 and 2 of the first XR device 12.1 and the second XR device 12.2 are combined to form a new canonical map. Figure 1 , and is outdated.
[0411] exist Figure 51 In some embodiments, this may occur when the first XR device 12.1 and the second XR device 12.2 attempt to locate during a different or new or subsequent session. The first XR device 12.1 and the second XR device 12.2 proceed as described above to locate their respective local maps (respectively, Figure 1 peacefully Figure 2 ) to locate the new canonical map.
[0412] like Figure 52 As shown in FIG, the head coordinate frame 96 or "head pose" is Figure 2 In some embodiments, the origin of the map, origin 2, is based on the head pose of the second XR device 12.2 at the start of the session. When a PCF is created during a session, the PCF is placed relative to the world coordinate frame origin 2. Figure 2 The PCF of is used as a persistent coordinate frame relative to the canonical coordinate frame, where the world coordinate frame can be the world coordinate frame of the previous session (e.g., Figure 40 The land in Figure 1 These coordinate frames are used to Figure 2 Positioning to the same transformation related to the canonical map, as above combined Figure 46B discussed.
[0413] Previously referenced Figure 9 The transformation from the world coordinate frame to the head coordinate frame 96 is discussed. Figure 52 The head coordinate frame 96 shown in FIG. 1 has only two orthogonal axes relative to the ground. Figure 2 The PCF is in a specific coordinate position and relative to the ground Figure 2 It should be understood, however, that the head coordinate frame 96 is relative to the ground. Figure 2 The PCF is located in three dimensions and has three orthogonal axes in the three-dimensional space.
[0414] exist Figure 53 In the figure, the head coordinate frame 96 has been Figure 2 Since the second user 14.2 has moved his head, the head coordinate frame 96 has moved. The user can move his head in six degrees of freedom (6dof). The head coordinate frame 96 can therefore be moved in 6dof (i.e., from its Figure 52 The previous position in 3D, and relative to the ground Figure 2 The PCF moves around three orthogonal axes). Figure 9 The head coordinate frame 96 is adjusted as the real object detection camera 44 and the inertial measurement unit 48 in the head unit 22 detect real objects and motion, respectively. More information about head pose tracking is disclosed in U.S. patent application serial number 16 / 221,065, entitled “Enhanced Pose Determination for Display Devices,” which is incorporated herein by reference in its entirety.
[0415] Figure 54 It is shown that sounds can be associated with one or more PCFs. The user can, for example, wear headphones or earphones with stereo sound. The location of the sound through the earphones can be simulated using conventional techniques. The location of the sound can be located at a fixed position so that when the user rotates their head to the left, the location of the sound rotates to the right so that the user perceives the sound as coming from the same location in the real world. In this example, the location of the sound is represented by sound 123 and sound 456. For ease of discussion, Figure 54 In terms of analysis Figure 48 Similarly, when the first user 14.1 and the second user 14.2 are in the same room at the same or different times, they perceive the sound 123 and the sound 456 as coming from the same location in the real world.
[0416] Figure 55 and Figure 56 Another implementation of the above technology is shown. Figure 8 As mentioned above, the first user 14.1 has initiated a first session. Figure 55 As shown in FIG, the first user 14.1 has terminated the first session, as shown by the dashed line. At the end of the first session, the first XR device 12.1 Figure 1 Uploaded to the server 20. The first user 14.1 has now initiated a second session at a later time than the first session. Figure 1 is already stored on the first XR device 12.1, so the first XR device 12.1 does not download the address from the server 20. Figure 1 If you lose your place Figure 1 , the first XR device 12.1 downloads the address from the server 20 Figure 1 The first XR device 12.1 then proceeds to construct Figure 2 PCF, positioned to the ground Figure 1 , and further develop the canonical map as described above. Then, as described above, the map of the first XR device 12.1 Figure 2 Used to associate local content, head coordinate frame, local sound, etc.
[0417] refer to Figure 57 and Figure 58 It is also possible that more than one user interacts with the server in the same session. In this example, the first user 14.1 and the second user 14.2 are joined together by a third user 14.3 and a third XR device 12.3. Each XR device 12.1, 12.2 and 12.3 starts generating its own map, i.e., the map Figure 1 ,land Figure 2 peacefully Figure 3 As XR devices 12.1, 12.2, and 12.3 continue to develop Figure 1 、 2 and 3, the map is incrementally uploaded to the server 20. The server 20 merges the Figure 1 、 2 and 3 to form a canonical map. The canonical map is then sent from the server 20 to each of the XR devices 12.1, 12.2, and 12.3.
[0418] Figure 59Aspects of a viewing method for restoring and / or resetting a head pose according to some embodiments are shown. In the example shown, at act 1400, the viewing device is powered on. At act 1410, in response to the power on, a new session is initiated. In some embodiments, the new session can include establishing a head pose. One or more capture devices on a head-mounted frame secured to the user's head capture a surface of the environment by first capturing an image of the environment and then determining the surface from the image. In some embodiments, the surface data can be combined with data from a gravity sensor to establish the head pose. Other suitable methods of establishing the head pose can be used.
[0419] At action 1420, the processor of the viewing device enters a routine for tracking head pose. As the user moves their head to determine the orientation of the head mounted frame relative to the surface, the capture device continues to capture the surface of the environment.
[0420] At action 1430, the processor determines whether the head pose has been lost. The head pose may be lost due to "edge" situations, such as excessively reflective surfaces that can result in low feature acquisition, low light, blank walls, being outdoors, etc.; or due to dynamic situations, such as a crowd of people moving and forming part of a map. The routine at 1430 allows a certain amount of time, such as 10 seconds, to pass to allow sufficient time to determine whether the head pose has been lost. If the head pose has not been lost, the processor returns to 1420 and re-enters tracking of the head pose.
[0421] If the head pose has been lost at action 1430, the processor enters a routine to recover the head pose at 1440. If the head pose is lost due to low light, a message such as the following message will be displayed to the user via the display of the viewing device:
[0422] The system is detecting low-light conditions. Please move to a better-lit area.
[0423] The system will continue to monitor whether sufficient light is available and whether the head pose can be recovered. The system may alternatively determine that low texture of the surface is causing the head pose to be lost, in which case the following prompt is given to the user in the display as a suggestion to improve surface capture:
[0424] The system cannot detect enough surfaces with fine textures. Move to an area with less rough surface textures and finer textures.
[0425] At action 1450, the processor enters a routine to determine whether head pose recovery has failed. If head pose recovery has not failed (i.e., head pose recovery has been successful), the processor returns to action 1420 by re-entering tracking of the head pose. If head pose recovery has failed, the processor returns to action 1410 to establish a new session. As part of the new session, all cached data is invalidated and the head pose is re-established thereafter. Any suitable head tracking method may be used with Figure 59 U.S. Patent Application No. 16 / 221,065 describes head tracking and is hereby incorporated by reference in its entirety.
[0426] Remote positioning
[0427] Various embodiments may utilize remote resources to facilitate a persistent and consistent cross-reality experience between individuals and / or groups of users. The inventors have recognized and appreciated that the benefits of operating an XR device utilizing a canonical map as described herein may be achieved without downloading a set of canonical maps. Figure 30 An example implementation of downloading canonical maps to a device is shown. For example, the benefits of not downloading maps can be realized by sending feature and gesture information to a remote service that maintains a set of canonical maps. According to one embodiment, a device seeking to use canonical maps to position virtual content at a location specified relative to the canonical maps can receive one or more transformations between features and canonical maps from the remote service. These transformations can be used on a device that maintains information about the location of these features in the physical world to position virtual content at a location specified relative to the canonical maps, or otherwise identify a location in the physical world specified relative to the canonical maps.
[0428] In some embodiments, spatial information is captured by the XR device and transmitted to a remote service, such as a cloud-based service, which uses the spatial information to position the XR device relative to a canonical map used by applications or other components of the XR system to specify the location of virtual content relative to the physical world. Once positioned, a transformation linking a tracking map maintained by the device to the canonical map can be transmitted to the device. The transformation can be used in conjunction with the tracking map to determine the location of virtual content to be rendered relative to the canonical map, or to otherwise identify a location in the physical world relative to the canonical map.
[0429] The inventors have recognized that the data that needs to be exchanged between a device and a remote positioning service may be very small compared to the transmission of map data, which may occur when a device transmits a tracking map to a remote service and receives a set of canonical maps from the service for device-based positioning. In some embodiments, performing positioning functions on cloud resources requires only a small amount of information to be transmitted from the device to the remote service. For example, a complete tracking map need not be transmitted to the remote service to perform positioning. In some embodiments, feature and gesture information, such as may be stored in association with a persistent gesture as described above, may be transmitted to a remote server. As described above, in embodiments where features are represented by descriptors, the uploaded information may be smaller.
[0430] The result returned to the device from the positioning service can be one or more transformations that relate the uploaded features to portions of the matching canonical map. These transformations can be used in conjunction with the XR system's tracking map to identify the location of virtual content or otherwise identify locations in the physical world. In embodiments that use persistent spatial information such as the PCFs described above to specify locations relative to the canonical map, the positioning service can download the transformations between the features and one or more PCFs to the device after successful positioning.
[0431] As a result, the network bandwidth consumed by communications between the XR device and the remote service used to perform positioning can be low. The system can therefore support frequent positioning, enabling each device interacting with the system to quickly obtain information used to locate virtual content or perform other location-based functions. As the device moves through the physical environment, it may repeatedly request updated positioning information. In addition, the device may frequently obtain updates to positioning information, such as when the canonical map changes, such as by incorporating additional tracking maps to expand the map or improve its accuracy.
[0432] Additionally, uploading features and downloading transforms can enhance privacy in XR systems that share map information among multiple users by increasing the difficulty of obtaining a map through spoofing. For example, an unauthorized user can be prevented from obtaining a map from the system by sending a false request for a canonical map representing a portion of the physical world in which the unauthorized user is not located. An unauthorized user is unlikely to access features in the area of the physical world for which they are requesting map information if the unauthorized user is not physically present in that area. The difficulty of spoofing feature information in a request for map information is further compounded in embodiments where the feature information is formatted as a feature description. Additionally, when the system returns transforms that are intended to be applied to a tracking map for a device operating in the area for which location information is requested, the information returned by the system may be of little or no use to an imposter.
[0433] According to one embodiment, the positioning service is implemented as a cloud-based microservice. In some examples, implementing a cloud-based positioning service can help conserve device computing resources and enable the calculations required for positioning to be performed with very low latency. These operations can be supported by virtually unlimited computing power or other computing resources available through the provision of additional cloud resources, thereby ensuring the scalability of the XR system to support a large number of devices. In one example, many canonical maps can be maintained in memory for almost instant access or stored on high-availability devices to reduce system latency.
[0434] Furthermore, performing location tracking on multiple devices in the cloud can improve the process. Location telemetry and statistics can provide information about which canonical maps are in active memory and / or high-availability storage. For example, statistics from multiple devices can be used to identify the most frequently accessed canonical maps.
[0435] Additional accuracy can also be achieved as a result of processing in a cloud environment or other remote environment with substantial processing resources relative to the remote device. For example, localization can be performed on a higher density canonical map in the cloud relative to processing performed on the local device. Maps can be stored in the cloud, for example, with a greater number of PCFs or a higher density of feature descriptors per PCF, thereby improving the accuracy of the match between a set of features from the device and the canonical map.
[0436] Figure 61 6100 is a schematic diagram of an XR system. The user device that displays cross-reality content during a user session can take many forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As described above, these devices can be configured with software, such as applications or other components, and / or hardwired to generate local location information (e.g., a tracking map) that can be used to render virtual content on their respective displays.
[0437] Virtual content location information may be specified relative to global location information, for example, the global location information may be formatted as a canonical map containing one or more PCFs.According to some embodiments, system 6100 is configured with a cloud-based service that supports execution and display of virtual content on a user device.
[0438] In one example, positioning functionality is provided as a cloud-based service 6106, which can be a microservice. The cloud-based service 6106 can be implemented on any of a plurality of computing devices, from which computing resources can be allocated to one or more services executed in the cloud. Those computing devices can be interconnected and accessible to devices such as the wearable XR device 6102 and the handheld device 6104. Such connections can be provided via one or more networks.
[0439] In some embodiments, the cloud-based service 6106 is configured to accept descriptor information from various user devices and "locate" the device to a matching one or more canonical maps. For example, the cloud-based positioning service matches the received descriptor information with the descriptor information of the corresponding canonical map. Canonical maps can be created using the techniques described above, which create canonical maps by merging maps provided by one or more devices having image sensors or other sensors that obtain information about the physical world. However, canonical maps are not required to be created by the devices accessing them, as such maps can be created by map developers, for example, map developers can publish maps by making them available to the positioning service 6106.
[0440] According to some embodiments, the cloud service handles canonical map identification and may include operations to filter the repository of canonical maps to a set of potential matches. Filtering may be as follows: Figure 29 as shown, or by using any subset of the filter criteria and replacing Figure 29 In addition to the filter criteria shown in Figure 29 In one embodiment, geographic data may be used to limit the search for matching canonical maps to maps representing an area proximate to the device requesting location. For example, regional attributes, such as Wi-Fi signal data, Wi-Fi fingerprint information, GPS data, and / or other device location information, may be used as a coarse filter on stored canonical maps to limit the analysis of descriptors to canonical maps that are known or likely to be proximate to the user's device. Similarly, a location history for each device may be maintained by a cloud service to prioritize searches for canonical maps that are proximate to the device's last location. In some examples, filtering may include the above with respect to Figure 31B 、 Figure 32 、 Figure 33 and Figure 34 Function of discussion.
[0441] Figure 62It is an example process that can be executed by a device to use a cloud-based service to locate the position of the device using a canonical map and receive transformation information specifying one or more transformations between the local coordinate system of the device and the coordinate system of the canonical map. Various embodiments and examples describe a transformation as specifying a transformation from a first coordinate frame to a second coordinate frame. Other embodiments include a transformation from a second coordinate frame to a first coordinate frame. In any other embodiment, the transformation implements a transition from one coordinate frame to another, and the resulting coordinate frame depends only on the desired coordinate frame output (including, for example, the coordinate frame in which content is displayed). In yet another embodiment, the coordinate system transformation enables determination from the second coordinate frame to the first coordinate frame and from the first coordinate frame to the second coordinate frame.
[0442] According to some embodiments, information reflecting the transformation for each persistent gesture defined by the canonical map may be transmitted to the device.
[0443] According to one embodiment, process 6200 may begin with a new session at 6202. Starting a new session on a device may initiate the capture of image information to build a tracking map for the device. Additionally, the device may send a message to register with a location service server, prompting the server to create a session for the device.
[0444] In some embodiments, starting a new session on a device may optionally include sending adjustment data from the device to a location service. The location service returns one or more transformations calculated based on a set of features and associated poses to the device. If the poses of the features are adjusted based on device-specific information before and / or after the transformations are calculated, rather than performing those calculations on the device, the device-specific information may be sent to the location service so that the location service can apply those adjustments. As a specific example, sending device-specific adjustment information may include capturing calibration data for the sensor and / or display. The calibration data can be used, for example, to adjust the position of feature points relative to a measured position. Alternatively or additionally, the calibration data can be used to adjust the position of virtual content rendered by the display so that it appears accurately positioned for that particular device. This calibration data can be obtained, for example, from multiple images of the same scene captured using sensors on the device. The positions of the features detected in those images can be represented as functions of the sensor position, such that the multiple images yield a set of equations that can be solved for the sensor position. The calculated sensor position can be compared to the nominal position, and calibration data can be derived from any differences. In some embodiments, intrinsic information about the device's configuration can also enable the calculation of calibration data for the display.
[0445] In embodiments where calibration data is generated for a sensor and / or display, the calibration data may be applied at any point in the measurement or display process. In some embodiments, the calibration data may be sent to a positioning server, which may store the calibration data in a data structure established for each device that has registered with the positioning server and is therefore in session with the server. The positioning server may apply the calibration data to any transformations calculated as part of the positioning process for the device providing the calibration data. Thus, the computational burden of using the calibration data to improve the accuracy of the sensed and / or displayed information is borne by the calibration service, thereby providing a further mechanism to reduce the processing burden on the device.
[0446] Once the new session is established, process 6200 may continue to capture new frames of the device's environment at 6204. At 6206, each frame may be processed to generate descriptors for the captured frame (including, for example, the DSF values discussed above). These values may be calculated using some or all of the techniques described above, including those described above with respect to Figure 14 、 Figure 22 and Figure 23 Techniques discussed. As discussed, descriptors can be computed as a mapping of feature points, or in some embodiments, a mapping of image patches surrounding feature points to descriptors. The descriptors can have values that enable valid matching between newly acquired frames / images and stored maps. In addition, the number of features extracted from the images can be limited to a maximum number of feature points per image, e.g., 200 feature points per image. As described above, feature points can be selected to represent points of interest. Thus, actions 6204 and 6206 can be performed as part of a device process that forms a tracking map or otherwise periodically collects images of the physical world around the device, or can, but need not, be performed separately for localization.
[0447] The feature extraction at 6206 may include attaching gesture information to the features extracted at 6206. The gesture information may be a gesture in the local coordinate system of the device. In some embodiments, the gesture may be relative to a reference point in a tracking map, such as a persistent gesture as described above. Alternatively or additionally, the gesture may be relative to the origin of the tracking map for the device. Such an embodiment may enable a location service as described herein to provide location services for a wide range of devices, even if they do not use persistent gestures. Regardless, gesture information may be attached to each feature or group of features such that the location service may use the gesture information to calculate a transformation that may be returned to the device when matching the feature with features in a stored map.
[0448] Process 6200 may continue to decision block 6207, where a decision is made whether to request a fix. One or more criteria may be applied to determine whether to request a fix. The criteria may include the passage of time, such that the device may request a fix after a threshold amount of time. For example, if no fix is attempted within the threshold amount of time, the process may continue from decision block 6207 to action 6208, where a fix is requested from the cloud. The threshold amount of time may be between 10 and 30 seconds, e.g., 25 seconds. Alternatively or additionally, a fix may be triggered by the movement of the device. The device executing process 6200 may track its movement using an IMU and its tracking map, and initiate a fix when movement is detected that exceeds a threshold distance from the device's last location for which a fix was requested. For example, the threshold distance may be between 1 and 10 meters, e.g., between 3 and 5 meters. As another alternative, a fix may be triggered in response to an event, such as when the device creates a new persistent pose or when the device's current persistent pose changes, as described above.
[0449] In some embodiments, decision block 6207 can be implemented so that the threshold for triggering a position fix can be established dynamically. For example, in an environment where features are largely consistent, such that the confidence level in matching a set of extracted features to features of a stored map may be low, a position fix may be requested more frequently to increase the chance that at least one position fix attempt will be successful. In this case, the threshold applied at decision block 6207 can be lowered. Similarly, in an environment where features are relatively few, the threshold applied at decision block 6207 can be lowered to increase the frequency of position fix attempts.
[0450] Regardless of how positioning is triggered, when triggered, process 6200 may proceed to act 6208, where the device sends a request to the positioning service, including data used by the positioning service to perform positioning. In some embodiments, data from multiple image frames may be provided for the positioning attempt. For example, the positioning service may not consider the positioning successful unless features in multiple image frames produce consistent positioning results. In some embodiments, process 6200 may include saving feature descriptors and additional pose information to a buffer. The buffer may, for example, be a circular buffer that stores feature sets extracted from the most recently captured frames. Thus, the positioning request may be sent with multiple feature sets accumulated in the buffer. In some arrangements, the buffer size is implemented to accumulate multiple data sets that are more likely to produce a successful positioning. In some embodiments, the buffer size may be set to accumulate features from, for example, two, three, four, five, six, seven, eight, nine, or ten frames. Optionally, the buffer size may have a baseline setting that can be increased in response to a positioning failure. In some examples, increasing the buffer size and the corresponding number of feature sets transmitted reduces the likelihood that a subsequent positioning function will fail to return a result.
[0451] Regardless of how the buffer size is set, the device can transmit the contents of the buffer to the location service as part of a location request. Other information can be transmitted along with the feature points and additional pose information. For example, in some embodiments, geographic information can be transmitted. The geographic information can include, for example, GPS coordinates or a wireless signature associated with the device tracking a map or the current persistent pose.
[0452] In response to the request sent at 6208, the cloud positioning service may analyze the feature descriptors to locate the device to a canonical map or other persistent map maintained by the service. For example, the descriptors are matched to a set of features in a map where the device is located. The cloud-based positioning service may perform positioning as described above relative to device-based positioning (e.g., may rely on any of the functions discussed above for positioning (including map ranking, map filtering, position estimation, filtered map selection, Figure 44 to Figure 4 6, and / or discussed with respect to positioning module, PCF and / or PP identification and matching, etc.). However, instead of transmitting the identified canonical map to the device (e.g., in device positioning), the cloud-based positioning service can continue to generate transformations based on the matching features of the canonical map and the relative orientation of the feature set sent from the device. The positioning service can return these transformations to the device, which the device can receive at block 6210.
[0453] In some embodiments, the canonical maps maintained by the positioning service can employ PCFs, as described above. In such embodiments, feature points of the canonical map that match feature points sent from the device can have locations specified relative to one or more PCFs. Thus, the positioning service can identify one or more canonical maps and can calculate a transformation between the coordinate frame represented in the gesture sent with the positioning request and the one or more PCFs. In some embodiments, identification of one or more canonical maps is assisted by filtering potential maps based on geographic data of the corresponding device. For example, once filtered to a candidate set (e.g., by GPS coordinates, among other options), the candidate set of canonical maps can be analyzed in detail to determine matching feature points or PCFs as described above.
[0454] The data returned to the requesting device in action 6210 may be formatted as a persistent gesture transformation table. The table may be accompanied by one or more canonical map identifiers indicating the canonical map to which the device was located by the location service. However, it should be understood that the location information may be formatted in other ways, including as a transformation list with associated PCFs and / or canonical map identifiers.
[0455] Regardless of how the transforms are formatted, the device can use these transforms to calculate the position of rendered virtual content that has been specified by an application or other component of the XR system relative to any PCF in action 6212. This information is used alternatively or additionally on the device to perform any position-based operations where the position is specified based on the PCF.
[0456] In some scenarios, the location service may not be able to match the features sent from the device to any stored canonical map, or may not be able to match a sufficient number of feature sets transmitted with the request to the location service to deem a location fix successful. In such a scenario, the location service may indicate to the device that the location fix failed, rather than returning the transformation to the device as described above in conjunction with action 6210. In such a scenario, process 6200 may branch to action 6230 at decision box 6209, where the device may take one or more actions for failure handling. These actions may include increasing the size of the buffer that holds the feature sets sent for location fix. For example, if the location service does not deem a location fix successful unless three feature sets match, the buffer size may be increased from 5 to 6, thereby increasing the chance that the three transmitted feature sets will match the canonical map maintained by the location service.
[0457] Alternatively or additionally, failure handling may include adjusting operating parameters of the device to trigger more frequent positioning attempts. For example, a threshold time and / or threshold distance between positioning attempts may be reduced. As another example, the number of feature points in each feature set may be increased. A match between a feature set and features stored within a canonical map may be considered to have occurred when a sufficient number of features in the set sent from the device match features of the map. Increasing the number of features sent may increase the chances of a match. As a specific example, the initial feature set size may be 50, which may be increased to 100, 150, and then 200 upon each successive positioning failure. Upon a successful match, the size of the set may then be returned to its initial value.
[0458] Failure handling may also include obtaining positioning information in addition to obtaining positioning information from a location service. According to some embodiments, the user device may be configured to cache canonical maps. Cached maps allow the device to access and display content that is not available in the cloud. For example, cached canonical maps allow device-based positioning in the event of a communication failure or other unavailability.
[0459] According to various embodiments, Figure 62 A high-level process for device-initiated cloud-based positioning is described. In other embodiments, various one or more of the steps shown can be combined, omitted, or invoke other processes to complete the positioning and final visualization of virtual content in the corresponding device view.
[0460] Furthermore, it should be understood that while process 6200 illustrates the device determining whether to initiate positioning at decision block 6207, the trigger for initiating positioning may come from outside the device, including from a positioning service. For example, a positioning service may maintain information about each device in session with it. For example, this information may include an identifier of the canonical map to which each device was most recently positioned. The positioning service or other components of the XR system may update the canonical map, including using the above in conjunction with Figure 26 When a canonical map is updated, the location service can send a notification to each device that was recently located with respect to the map. The notification can serve as a trigger for the device to request a location fix and / or can include an updated transform recalculated using the feature set most recently sent from the device.
[0461] Figure 63A 、 Figure 63B and Figure 63C is an example process flow illustrating operations and communications between a device and a cloud service. Blocks 6350, 6352, 6354, and 6456 illustrate an example architecture and separation between components involved in a cloud-based positioning process. For example, modules, components, and / or software configured to process perception on a user device are shown at 6350 (e.g., 660, Figure 6A Device functionality for persistent world operations is shown at 6352 (including, for example, as described above and with respect to the persistent world module (e.g., 662, Figure 6A )). In other embodiments, separation between 6350 and 6352 is not required and the communications shown may be between processes executing on the device.
[0462] Similarly, shown at block 6354 is a cloud process (e.g., 802, 812, 813) configured to handle functionality associated with traversable worlds / traversable world modeling. Figure 26 ). Shown at box 6356 is a cloud process configured to handle the functions associated with positioning the device to one or more maps in the repository of stored canonical maps based on information sent from the device.
[0463] In the illustrated embodiment, process 6300 begins at 6302 when a new session begins. Sensor calibration data is obtained at 6304. The calibration data obtained may depend on the device (e.g., multiple cameras, sensors, positioning devices, etc.) represented at 6350. Once sensor calibration is obtained for a device, the calibration may be cached at 6306. If device operation results in a change in frequency parameters (e.g., collection frequency, sampling frequency, matching frequency, and other options), the frequency parameters are reset to a baseline at 6308.
[0464] Once the new session functionality is complete (e.g., calibration, steps 6302-6306), process 6300 can continue with capturing new frames 6312. At 6314, features and their corresponding descriptors are extracted from the frame. In some examples, the descriptors can include a DSF, as described above. According to some embodiments, the descriptors can have spatial information attached to them to enable subsequent processing (e.g., transform generation). At 6316, pose information generated on the device (e.g., information specified for locating features in the physical world relative to the device's tracking map, as described above) can be attached to the extracted descriptors.
[0465] At 6318, the descriptor and pose information are added to the buffer. New frames are captured and added to the buffer as shown in steps 6312-6318 in a loop until the buffer size threshold is exceeded at 6319. At 6320, in response to determining that the buffer size is met, a positioning request is transmitted from the device to the cloud. According to some embodiments, the request can be processed by a navigable world service (e.g., 6354) instantiated in the cloud. In further embodiments, the functional operations for identifying candidate canonical maps can be separated from the operations for actual matching (e.g., shown as boxes 6354 and 6356). In one embodiment, a cloud service for map filtering and / or map ranking can be executed at 6354 and process the positioning request received from 6320. According to one embodiment, the map ranking operation is configured to determine a set of candidate maps that may include the location of the device at 6322.
[0466] In one example, the map ranking functionality includes operations for identifying candidate canonical maps based on geographic attributes or other location data (e.g., observed or inferred location information). For example, the other location data may include Wi-Fi signatures or GPS information.
[0467] According to other embodiments, location data may be captured during a cross-reality session with a device and user. Process 6300 may include additional operations to populate locations for a given device and / or session (not shown). For example, location data may be stored as a device region attribute value and an attribute value for selecting candidate canonical maps proximate to the device location.
[0468] Any one or more of the location options may be used to filter the canonical map set to those that may represent an area that includes the location of the user device. In some embodiments, the canonical map may cover a relatively large area of the physical world. The canonical map may be segmented into regions, such that selection of a map may require selection of a map region. For example, a map region may be on the order of tens of square meters. Thus, the filtered canonical map set may be a set of regions of the map.
[0469] According to some embodiments, a positioning snapshot can be constructed from candidate canonical maps, posture features, and sensor calibration data. For example, an array of candidate canonical maps, posture features, and sensor calibration information can be sent along with a request to determine a specific matching canonical map. Matching with the canonical map can be performed based on descriptors received from the device and stored PCF data associated with the canonical map.
[0470] In some embodiments, a feature set from the device is compared to a feature set stored as part of the canonical map. This comparison can be based on feature descriptors and / or pose. For example, a candidate feature set for the canonical map can be selected based on the number of features in the candidate set whose descriptors are sufficiently similar to the descriptors of the feature set from the device that they are likely the same feature. For example, the candidate set can be features derived from the image frames used to form the canonical map.
[0471] In some embodiments, if the number of similar features exceeds a threshold, further processing can be performed on the candidate feature set. Further processing can determine the degree to which the gesture feature set from the device can be aligned with the candidate feature set. Posing can be performed for feature sets from the canonical map that are similar to features from the device.
[0472] In some embodiments, features are formatted as high-dimensional embeddings (e.g., DSF, etc.) and can be compared using a nearest neighbor search. In one example, the system is configured (e.g., by executing process 6200 and / or 6300) to find the first two nearest neighbors using Euclidean distance, and a ratio test can be performed. If the nearest neighbor is closer than the second nearest neighbor, the system considers the nearest neighbor to be a match. For example, "closer" in this context can be determined by the ratio of the Euclidean distance relative to the second nearest neighbor exceeding the ratio of the Euclidean distance relative to the nearest neighbor by a threshold multiple. Once a feature from the device is considered to be a "match" with a feature in the canonical map, the system can be configured to calculate a relative transformation using the posture of the matching feature. The transformation developed from the posture information can be used to indicate the transformation required to position the device to the canonical map.
[0473] The number of inliers can be used as an indicator of the quality of the match. For example, in the case of DSF matching, the number of inliers reflects the number of features that matched between the received descriptor information and the stored / canonical map. In another embodiment, the inliers determined in this embodiment can be determined by counting the number of "matched" features in each set.
[0474] Indications of the quality of the match may alternatively or additionally be determined in other ways. In some embodiments, for example, when a transformation is calculated to position a map from a device that may contain multiple features to a canonical map based on the relative poses of the matching features, the transformation statistics calculated for each of the multiple matching features may serve as an indication of quality. For example, a larger difference may indicate a poorer quality match. Alternatively or additionally, for the determined transformation, the system may calculate the average error between features with matching descriptors. The average error may be calculated for the transformation, reflecting the degree of position mismatch. Mean squared error is a specific example of an error metric. Regardless of the specific error metric, if the error is below a threshold, it may be determined that the transformation is usable for the features received from the device, and the calculated transformation is used to locate the device. Alternatively or additionally, the number of inliers may also be used to determine whether there is a map that matches the descriptors received from the device and / or the device's location information.
[0475] As described above, in some embodiments, the device may send multiple feature sets for positioning. Positioning may be considered successful when at least a threshold number of feature sets match the feature set from the canonical map with an error below a threshold and a number of inner layers above a threshold. The threshold number may be, for example, three feature sets. However, it will be appreciated that the threshold for determining whether a sufficient number of feature sets have a suitable value may be determined empirically or in another suitable manner. Similarly, other thresholds or parameters of the matching process, such as the similarity between feature descriptors considered to be matched, the number of inner layers used to select candidate feature sets, and / or the size of the mismatch error, may similarly be determined empirically or in another suitable manner.
[0476] Once a match is determined, a set of persistent map features associated with the matching canonical map or maps is identified. In embodiments where matching is based on map regions, the persistent map features may be map features within the matching region. The persistent map features may be persistent gestures or PCFs as described above. In the example of FIG63 , the persistent map features are persistent gestures.
[0477] Regardless of the format of the persistent map features, each persistent map feature can have a predetermined orientation relative to the canonical map to which it belongs. This relative orientation can be applied to a calculated transformation to align the feature set from the device with the feature set from the canonical map, thereby determining the transformation between the feature set from the device and the persistent map feature. Any adjustments, such as those that may come from calibration data, can be applied to this calculated transformation. The resulting transformation can be a transformation between the local coordinate frame of the device and the persistent map feature. This calculation can be performed for each persistent map feature that matches the map area, and the results can be stored in a table, denoted as persistent_pose_table in 6326.
[0478] In one example, block 6326 returns a table of persistent pose transformations, canonical map identifiers, and inlier levels. According to some embodiments, a canonical map ID is an identifier that uniquely identifies a canonical map and canonical map version (or region of a map, in embodiments where positioning is based on a map region).
[0479] In various embodiments, at 6328, the calculated positioning data may be used to populate positioning statistics and telemetry maintained by the positioning service. This information may be stored for each device and may be updated for each positioning attempt and may be cleared when the device's session ends. For example, maps that a device has matched to may be used to improve map ranking operations. For example, maps covering the same area that the device was previously matched to may be prioritized in the ranking. Similarly, maps covering adjacent areas may be given higher priority than more remote areas. In addition, adjacent maps may be prioritized based on the detected trajectory of the device over time, with map areas in the direction of motion being given higher priority than other map areas. The positioning service may use this information, for example, to limit the maps or map areas that are searched for a candidate feature set in the stored canonical maps based on subsequent positioning requests from the device. If matches with low error metrics and / or a large number or percentage of inliers are identified in this limited area, processing of maps outside of that area may be avoided.
[0480] Process 6300 may continue with the transfer of information from the cloud (e.g., 6354) to the user device (e.g., 6352). According to one embodiment, at 6330, the persistent gesture table and the canonical map identifier are transferred to the user device. In one example, the persistent gesture table may consist of elements including at least a string identifying the persistent gesture ID and a transformation linking the device's tracking map to the persistent gesture. In embodiments where the persistent map feature is a PCF, the table may instead indicate a transformation to a matching map's PCF.
[0481] If positioning fails at 6336, process 6300 continues by adjusting parameters that can increase the amount of data sent from the device to the positioning service to increase the chance of successful positioning. For example, a failure can be indicated when a feature set with more than a threshold number of similar descriptors cannot be found in the canonical map, or when the error metric associated with all transformed candidate feature sets is above a threshold. As an example of a parameter that can be adjusted, the size constraint of the descriptor buffer can be increased (6319). For example, in the case of a descriptor buffer size of 5, a positioning failure can trigger an increase to at least six feature sets extracted from at least six image frames. In some embodiments, process 6300 can include a descriptor buffer increment value. In one example, the increment value can be used to control the rate at which the buffer size is increased, for example, in response to a positioning failure. Other parameters, such as parameters that control the rate of positioning requests, can be changed when a matching canonical map cannot be found.
[0482] In some embodiments, the execution of 6300 may generate an error condition at 6340, which includes the execution of a positioning request failing to work rather than returning an unmatched result. For example, an error may occur due to a network error that renders the storage holding the canonical map database unavailable to the server performing the positioning service, or a request received for the positioning service containing incorrectly formatted information. In the event of an error condition, in this example, process 6300 schedules a retry of the request at 6342.
[0483] When the positioning request succeeds, any parameters adjusted in response to the failure can be reset. At 6332, process 6300 can continue to operate to reset the frequency parameters to any default values or baselines. In some embodiments, 6332 is performed regardless of any changes, thereby ensuring that a baseline frequency is always established.
[0484] The device may use the received information to update the cached location snapshot at 6334. According to various embodiments, the corresponding transformation, canonical map identifier, and other location data may be stored by the device and used to correlate locations specified relative to the canonical map, or their persistent map features such as persistent poses or PCFs, with locations determined by the device relative to its local coordinate frame (such as may be determined from its tracking map).
[0485] Various embodiments of the process for positioning in the cloud can implement any one or more of the aforementioned steps and be based on the aforementioned architecture. Other embodiments can combine various one or more of the aforementioned steps, performing the steps simultaneously, in parallel, or in another order.
[0486] According to some embodiments, the location service in the cloud in the context of a cross-reality experience may include additional functionality. For example, canonical map caching may be performed to address connectivity issues. In some embodiments, the device may periodically download and cache canonical maps of where it has been located. If the location service in the cloud is unavailable, the device may perform its own location (e.g., as described above—including information about the location of the device). Figure 26 In other embodiments, the transforms returned from a position request can be chained together and applied to subsequent sessions. For example, a device can cache a series of transforms and use the transform sequence to establish a position fix.
[0487] Various embodiments of the system may use the results of a positioning operation to update transformation information. For example, the positioning service and / or the device may be configured to maintain state information on the tracking map to a canonical map transformation. The received transformations may be averaged over time. According to one embodiment, the averaging operation may be limited to occur after a threshold number of positioning successes (e.g., three, four, five, or more times). In further embodiments, other state information may be tracked in the cloud, such as by a navigable world module. In one example, the state information may include a device identifier, a tracking map ID, a canonical map reference (e.g., version and ID), and a transformati...
Claims
1. A portable device configured to operate and display virtual content of a cross-reality system in a three-dimensional (3D) environment, the portable device comprising: at least one processor; an image sensor configured to output image data; as well as Computer executable instructions configured to, when executed by the at least one processor, perform a method comprising: forming a map of the 3D environment as the portable device moves within the 3D environment, the map comprising a first plurality of features associated with locations within the map; Maintaining a wireless fingerprint associated with the portable device by repeatedly performing the following operations: extracting a second plurality of features from the image data; obtaining network access point information at a location within the 3D environment from a network access point transmitting a wireless signal, the wireless signal being received by the portable device; selecting a location within the map that corresponds to the location of the portable device within the 3D environment based on a correspondence between features of the first plurality of features of the map and features of the second plurality of features extracted from the image data; and The obtained network access point information is linked to the selected location within the map.
2. The portable device according to claim 1, wherein Linking the obtained network access point information with the selected location includes combining the obtained network access point information with previously obtained network access point information associated with the selected location.
3. The portable device according to claim 1, wherein Obtaining the network access point information includes: triggering a scan for the network access point. The portable device according to claim 3 , wherein: A scan for network access points is triggered when an amount of time that has elapsed since a previous scan for network access points was performed exceeds a threshold. The portable device according to claim 3 , wherein: When the distance that the portable device has moved in the 3D environment exceeds a threshold, a scan for a network access point is triggered. The portable device according to claim 1 , wherein: Obtaining the network access point information includes: after scanning the network access points, receiving the network access point information pushed from the wireless hardware component.
7. The portable device according to claim 1, wherein Obtaining the network access point information includes: obtaining an access point identifier of the network access point.
8. The portable device according to claim 7, wherein: Obtaining network access point information further includes obtaining one or more signal strength indicator values of the access point identified by the access point identifier.
9. The portable device according to claim 8, wherein The access point identifier is a basic service set identifier BSSID.
10. The portable device according to claim 9, wherein The signal strength indicator value is a received signal strength indicator RSSI value. The portable device according to claim 8 , wherein: Linking the obtained network access point information with the selected location within the map includes storing the one or more signal strength indicator values in association with the access point identifier.
12. The portable device according to claim 11, wherein Linking the obtained network access point information with the selected location within the map includes identifying a subset of the access point identifiers to be excluded.
13. The portable device according to claim 12, wherein: The subset of access point identifiers to be excluded is based at least on the one or more signal strength indicator values.
14. The portable device according to claim 11, wherein Storing the one or more signal strength indication values includes storing an average of a plurality of signal strength indication values in association with the access point identifier.
15. The portable device according to claim 1, wherein The selected location within the map comprises a persistent pose or persistent coordinate frame of the map.
16. The portable device according to claim 1, wherein Obtaining the network access point information further includes filtering, clustering, and / or normalizing the network access point information.
17. The portable device according to claim 1, wherein The portable device further comprises a computer-readable medium connected to the at least one processor, and wherein the method further comprises storing the map on the computer-readable medium.
18. The portable device according to claim 17, wherein: Linking the obtained network access point information with the selected location within the map includes updating information stored on the computer-readable medium in association with the selected location within the map based on the obtained network access point information.
19. The portable device according to claim 1, wherein The method further comprises: The map of the 3D environment is located within the stored map based on the wireless fingerprint associated with the portable device and based on a coordinate transformation between the stored map and the map of the 3D environment.
20. The portable device of claim 19, further comprising an image sensor, and wherein The method includes forming a map of the 3D environment based on image data generated by the image sensor as the portable device moves within the 3D environment.
21. A method for operating a portable device, the portable device being configured to operate and display virtual content of a cross-reality system in a three-dimensional (3D) environment, the method comprising: forming a map of the 3D environment as the portable device moves within the 3D environment, the map comprising a first plurality of features associated with locations within the map; Maintaining a wireless fingerprint associated with the portable device by repeatedly performing the following operations: extracting a second plurality of features from the image data; obtaining network access point information at a location within the 3D environment from a network access point transmitting a wireless signal, the wireless signal being received by the portable device; selecting a location within the map that corresponds to the location of the portable device within the 3D environment based on a correspondence between features of the first plurality of features of the map and features of the second plurality of features extracted from the image data; as well as The obtained network access point information is linked to the selected location within the map.
22. The method according to claim 21, further comprising: A map from the portable device is located within the stored map based on the wireless fingerprint associated with the portable device, wherein locating the map includes determining a coordinate transformation between the stored map and a map of the 3D environment formed based on the image data.
23. A computer-readable medium storing computer-executable instructions, the computer-executable instructions being configured to, when executed by at least one processor, perform a method for operating a portable device configured to operate and display virtual content of a cross-reality system in a three-dimensional (3D) environment, the method comprising: forming a map of the 3D environment as the portable device moves within the 3D environment, the map comprising a first plurality of features associated with locations within the map; Maintaining a wireless fingerprint associated with the portable device by repeatedly performing the following operations: extracting a second plurality of features from the image data; obtaining network access point information at a location within the 3D environment from a network access point transmitting a wireless signal, the wireless signal being received by the portable device; selecting a location within the map that corresponds to the location of the portable device within the 3D environment based on a correspondence between features of the first plurality of features of the map and features of the second plurality of features extracted from the image data; as well as The obtained network access point information is linked to the selected location within the map.
24. The computer-readable medium of claim 23, wherein: The method further comprises: A map from the portable device is located within the stored map based on the wireless fingerprint associated with the portable device, wherein locating the map includes determining a coordinate transformation between the stored map and a map of the 3D environment formed based on the image data.
25. A computing device configured for use in a cross-reality system, wherein: A portable device operating in a three-dimensional (3D) environment renders virtual content, the computing device comprising: at least one processor; an image sensor configured to output image data; and a computer-readable medium coupled to said at least one processor, wherein a plurality of maps are stored on the computer-readable medium, wherein the plurality of maps include information identifying locations within the 3D environment, the information including stored location information and stored network access point information associated with corresponding locations in the 3D environment, and a first plurality of features associated with the locations of the plurality of maps; Computer executable instructions, which when executed by the at least one processor are configured to perform a method comprising: extracting a second plurality of features from the image data; generating position information of the portable device indicating a position within the 3D environment based on correspondences between features in the first plurality of features of the plurality of maps and features in the second plurality of features extracted from the image data; receiving network access point information; selecting one or more locations within at least one of the plurality of maps as candidate locations based at least on a comparison of the received network access point information of the portable device with stored network access point information of the plurality of maps; determining whether the generated location information of the portable device is associated with a location in the 3D environment that is the same as location information of a candidate location included in the stored location information of the plurality of maps; calculating, for a candidate position determined to be associated with the same position indicated by the generated position information of the portable device, a transformation between the generated position information of the portable device and position information of the candidate position contained in the stored position information of the plurality of maps; The received network access point information is linked with at least one of the selected candidate locations.
26. The computing device of claim 25, wherein: The received network access point information of the portable device and / or the stored network access point information of the plurality of maps includes an access point identifier and a signal strength indicator value.
27. The computing device of claim 26, wherein: The access point identifier is a BSSID, and the signal strength indicator value is an RSSI value.
28. The computing device of claim 25, wherein: Comparing the received network access point information of the portable device with the stored network access point information of the plurality of maps includes determining a Jaccard similarity between the received network access point information of the portable device and the stored network access point information of the plurality of maps.
29. The computing device of claim 28, wherein: The comparison of the received network access point information of the portable device with the stored network access point information of the plurality of maps further includes determining whether the Jaccard similarity is above a threshold.
30. The computing device of claim 25, wherein: The received location information of the portable device includes a first coordinate frame, and the stored location information of the plurality of maps includes a plurality of coordinate frames.
31. The computing device of claim 30, wherein: Calculating a transformation between the received location information of the portable device and location information of the candidate location contained in the stored location information of the plurality of maps includes calculating a transformation between the first coordinate frame and a second coordinate frame of the plurality of coordinate frames.
32. The computing device of claim 31 , wherein: Computing the transformation is based on a calculated alignment between features of the second plurality of features and matching features of the first plurality of features of the plurality of maps containing the candidate location, the second plurality of features representing features in the 3D environment of the portable device.
33. The computing device of claim 25, wherein: The plurality of maps from which the candidate location is selected comprises a filtered subset of stored maps.
34. The computing device of claim 25, wherein: The received network access point information of the portable device is stored in a first stored map on the computer-readable medium; candidate locations determined to have location information associated with the same location as the received location information of the portable device are included in the second stored map; as well as The method further comprises: Based on the first stored map and the second stored map, merging the first stored map with the second stored map to generate a merged map including location information and network access point information; as well as The merged map is stored on the computer-readable medium.
35. The computing device of claim 25, wherein: The received network access point information of the portable device is stored in a tracking map received from the portable device; candidate locations determined to have location information associated with the same location as the received location information of the portable device are included in the first stored map; as well as The method further comprises: Based on the first stored map and the tracking map, merging the first stored map with the tracking map to generate a merged map including location information and network access point information; as well as The merged map is stored on the computer-readable medium.
36. A method of operating a cross-reality system comprising a portable device and a remote computing device, wherein: The portable device operates within a 3D environment, and wherein the portable device and the remote computing device are configured to interact with each other, the method comprising: accumulating, on the portable device, network access point information associated with each of a plurality of locations in the 3D environment over time, wherein the network access point information includes a plurality of wireless network access point identifiers and associated average signal strength values; extracting a first plurality of features from the image data; selecting a location within the map of the 3D environment based at least on a correspondence between features of the first plurality of features extracted from the image data and features of a second plurality of features of the map of the 3D environment; sending a request from the portable device to the remote computing device, the request including at least network access point information for the portable device at a selected location within the map of the 3D environment and information representing the selected location within the 3D environment, the information being expressed in a coordinate frame associated with the portable device; calculating, on the remote computing device, a transformation between a coordinate frame associated with the portable device and a coordinate frame associated with at least one stored map, wherein the map includes a location matching network access point information at a selected location of the portable device within the 3D environment and information representing the selected location within the 3D environment; At least some of the network access point information is linked with selected locations within the map of the 3D environment.
37. The method of claim 36, further comprising: The transformation is sent from the remote computing device to the portable device in response to the request.
38. The method of claim 36, wherein: The network access point information includes BSSID and RSSI value.
39. The method according to claim 36, wherein The remote computing device includes a distributed server network in a cloud computing configuration.
40. The method of claim 36, wherein The method is repeatedly performed while the portable device moves within the 3D environment.
41. The method of claim 36, wherein: The portable device and the remote computing device interact with each other via a wireless communication channel.
42. The method of claim 36, wherein: The request from the portable device includes a tracking map of the 3D environment of the portable device.
43. The method according to claim 42, wherein The coordinate frame associated with the portable device is the coordinate frame of the tracking map.
44. A portable device configured to operate and display virtual content of a cross-reality system in a three-dimensional (3D) environment, the portable device comprising: at least one processor; a computer-readable medium coupled to the at least one processor; Computer-executable instructions stored on the computer-readable medium, the computer-executable instructions being configured to, when executed by the at least one processor, perform a method comprising: forming a map of the 3D environment as the portable device moves within the 3D environment, the map comprising a first plurality of features associated with locations within the map; storing the map on the computer-readable medium; Maintaining a wireless fingerprint associated with the portable device by repeatedly performing the following operations: extracting a second plurality of features from the image data; obtaining network access point information at a location within the 3D environment from a network access point transmitting a wireless signal, the wireless signal being received by the portable device; selecting a location within the map that corresponds to the location of the portable device in the 3D environment based on a correspondence between features of the first plurality of features of the map and features of the second plurality of features extracted from the image data; and The obtained network access point information is linked to the selected location within the map.
Citation Information
Patent Citations
Localization determination for mixed reality systems
US10812936B2
Fully convolutional interest point detection and description via homographic adaptation
US20190147341A1
Enhanced pose determination for display device
US20190188474A1
Methods and apparatuses for determining and / or evaluating localizing maps of image display devices
US20200034624A1
Coarse relocalization using signal fingerprints
US20190287311A1