Positioning based on detected spatial features

By combining device pose estimation with an environmental model, the device's position in the environment is updated, solving the problem of inaccurate positioning of electronic devices and achieving accurate positioning of XR content.

CN115861412BActive Publication Date: 2026-01-06APPLE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210953751.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-24
Filing Date
2022-08-10
Publication Date
2026-01-06
Estimated Expiration
2042-08-10

AI Technical Summary

Technical Problem

In the prior art, inaccurate positioning of electronic devices leads to inaccurate placement of XR content in the environment.

Method used

By obtaining the device's pose estimate in the environment and utilizing the spatial feature positions in the environment model, the device's pose estimate is updated, and the device's position in the environment is accurately determined.

Benefits of technology

It improves the positioning accuracy of electronic devices in the environment, ensuring accurate positioning and correct display of XR content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115861412B_ABST
    Figure CN115861412B_ABST
Patent Text Reader

Abstract

The present disclosure relates to "positioning based on detected spatial features." In one implementation, a method of positioning a device is performed at a device comprising one or more processors and a non-transitory memory. The method includes obtaining an estimate of a pose of the device in an environment. The method includes obtaining an environment model of the environment, the environment model comprising a spatial feature in the environment defined by a first spatial feature location. The method includes determining a second spatial feature location of the spatial feature based on the estimate of the pose of the device. The method includes determining an updated estimate of the pose of the device based on the first spatial feature location and the second spatial feature location.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 247,991, filed on September 24, 2021, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to systems, methods, and apparatus for locating devices or content in an environment based on spatial features detected in the environment. Background Technology

[0004] Determining the location of electronic devices (e.g., positioning) enables a wide range of user experiences, such as automatically turning on lights when an electronic device enters a room, adjusting speaker volume based on the distance from the speaker to the electronic device, or displaying previously placed extended reality (XR) content in an environment where an electronic device is present. However, inaccuracies in determining the device's location can lead to inaccurate placement of XR content in the environment. Attached Figure Description

[0005] Therefore, this disclosure will be understood by those skilled in the art, and a more detailed description can be made with reference to some exemplary embodiments, some of which are shown in the accompanying drawings.

[0006] Figure 1 It is a block diagram based on some specific implementations of exemplary operating environments.

[0007] Figure 2 It is a block diagram of an exemplary controller based on some specific implementations.

[0008] Figure 3 It is a block diagram of an exemplary electronic device based on some specific implementations.

[0009] Figure 4A and 4B An XR environment based on some specific implementations is shown.

[0010] Figure 5 It is a flowchart representation of a method based on some specific implementations of positioning devices.

[0011] As is customary, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Additionally, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Summary of the Invention

[0012] The various embodiments disclosed herein include devices, systems, and methods for locating a device. In various embodiments, the method is performed by a device including one or more processors and non-transitory memory. The method includes acquiring an estimate of the pose of the device in an environment. The method includes acquiring an environment model of the environment, the environment model including spatial features in the environment defined by a first spatial feature location. The method includes determining a second spatial feature location based on the estimate of the pose of the device. The method includes determining an updated estimate of the pose of the device based on the first spatial feature location and the second spatial feature location.

[0013] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors. The one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes: one or more processors, non-transitory memory, and means for performing or causing to perform any of the methods described herein. Detailed Implementation

[0014] A physical environment refers to a physical place that people can sense and / or interact with without the aid of electronic devices. A physical environment can include physical features such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and / or interact with a physical environment through senses such as sight, touch, hearing, taste, and smell. Conversely, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment can include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, etc. In the case of an XR system, a subset of a person's physical motion or a representation thereof is tracked, and in response, one or more features of one or more virtual objects simulated in the XR system are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect the movement of electronic devices (e.g., mobile phones, tablets, laptops, headsets, etc.) presenting the XR environment, and in response, adjust the graphical content and sound field presented to the person by the electronic device in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), XR systems can adapt the characteristics of graphical content in an XR environment in response to representations of physical motion (e.g., voice commands).

[0015] Many different types of electronic systems enable people to sense and / or interact with a variety of XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped like lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have an integrated opaque display and one or more speakers. Alternatively, head-mounted systems may be configured to receive external opaque displays (e.g., smartphones). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In some implementations, transparent or translucent displays can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology, which projects graphic images onto the human retina. Projection systems can also be configured to project virtual objects onto a physical environment, such as as holograms or on a physical surface.

[0016] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective aspects and / or variations do not include all the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.

[0017] As described above, a head-mounted device equipped with a scene camera captures numerous images of the user's environment over a period of several days or weeks. The device can identify objects in those images (e.g., paintings, posters, album covers) and store information about these objects in a database. To access the information efficiently, the information about these objects is stored in association with corresponding contextual information about the time each object was detected, such as time, location, or current activity. Therefore, in response to a query asking "What album cover am I looking at when I'm at Jim's house?", the electronic device can return information about a specific album cover detected at a specific time or location.

[0018] Figure 1 This is a block diagram of an exemplary operating environment 100 according to some specific implementations. Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure further relevant aspects of the exemplary specific implementations disclosed herein. Therefore, as a non-limiting example, operating environment 100 includes a controller 110 and electronic devices 120.

[0019] In some implementations, controller 110 is configured to manage and coordinate the user's XR experience. In some implementations, controller 110 includes a suitable combination of software, firmware, and / or hardware. See below for reference. Figure 2 The controller 110 is described in more detail. In some embodiments, the controller 110 is a computing device located locally or remotely relative to a physical environment 105. For example, the controller 110 is a local server located within the physical environment 105. In another example, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the physical environment 105. In some embodiments, the controller 110 is communicatively coupled to the electronic device 120 via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). Alternatively, the controller 110 may be included within the housing of the electronic device 120. In some embodiments, the functionality of the controller 110 is provided by and / or combined with the electronic device 120.

[0020] In some embodiments, electronic device 120 is configured to provide an XR experience to a user. In some embodiments, electronic device 120 includes a suitable combination of software, firmware, and / or hardware. According to some embodiments, electronic device 120 presents XR content to a user via display 122 while the user is physically present within a physical environment 105, which includes a table 107 within the field of view 111 of electronic device 120. In some embodiments, the user holds electronic device 120 in one or both of his / her hands. In some embodiments, when providing XR content, electronic device 120 is configured to display XR objects (e.g., XR cylinder 109) and implement video pass-through of the physical environment 105 (e.g., a representation 117 including table 107) on display 122. Reference is made below. Figure 3 The electronic device 120 is described in more detail.

[0021] According to some specific implementations, electronic device 120 provides an XR experience to the user while the user is virtually and / or physically present in physical environment 105.

[0022] In some embodiments, the user wears the electronic device 120 on his / her head. For example, in some embodiments, the electronic device includes a head-mounted system (HMS), a head-mounted device (HMD), or a head-mounted housing (HME). Therefore, the electronic device 120 includes one or more XR displays configured to display XR content. For example, in various embodiments, the electronic device 120 surrounds the user's field of view. In some embodiments, the electronic device 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user no longer wears the electronic device 120 but holds it in their hand, with the display facing the user's field of view and the camera facing the physical environment 105. In some embodiments, the handheld device may be placed inside a housing that can be worn on the user's head. In some embodiments, the electronic device 120 is replaced by an XR pod, housing, or chamber configured to present XR content, in which the user no longer wears or holds the electronic device 120.

[0023] Figure 2This is a block diagram of an example controller 110 according to some specific implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, controller 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, and / or similar type interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0024] In some embodiments, the one or more communication buses 204 include circuitry for communication between interconnecting system components and control system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0025] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores programs, modules, and data structures, or subsets thereof, including optional operating system 230 and XR experience module 240.

[0026] Operating system 230 includes processes for handling various basic system services and for performing hardware-related tasks. In some implementations, XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., single XR experiences for one or more users, or multiple XR experiences for corresponding groups of one or more users). To this end, in various implementations, XR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a data transmission unit 248.

[0027] In some specific implementations, the data acquisition unit 242 is configured to acquire data from at least... Figure 1 The electronic device 120 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). Therefore, in various specific embodiments, the data acquisition unit 242 includes instructions and / or logic components for the instructions, as well as heuristics and metadata for the heuristics.

[0028] In some specific implementations, the tracking unit 244 is configured to map the physical environment 105 and at least track the electronic device 120 relative to it. Figure 1 The location / positioning of the physical environment 105. Therefore, in various specific implementations, the tracking unit 244 includes instructions and / or logic components for the instructions, as well as heuristics and metadata for the heuristics.

[0029] In some implementations, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by electronic device 120. To this end, in various implementations, coordination unit 246 includes instructions and / or logic components for those instructions, as well as heuristics and metadata for those heuristics.

[0030] In some implementations, the data transmission unit 248 is configured to transmit at least data (e.g., presentation data, location data, etc.) to the electronic device 120. Therefore, in various implementations, the data transmission unit 248 includes instructions and / or logic components for the instructions, as well as heuristics and metadata for the heuristics.

[0031] Although the data acquisition unit 242, tracking unit 244, coordination unit 246 and data transmission unit 248 are shown residing on a single device (e.g., controller 110), it should be understood that in other implementations, any combination of the data acquisition unit 242, tracking unit 244, coordination unit 246 and data transmission unit 248 may reside in a separate computing device.

[0032] also, Figure 2This is used more as a functional description of various features that can exist in a specific implementation, and differs from the structural diagrams of the specific implementations described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 2 Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented in various specific implementations through one or more functional blocks. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.

[0033] Figure 3 This is a block diagram of an example of an electronic device 120 according to some specific embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure more relevant aspects of the specific embodiments disclosed herein. Therefore, as a non-limiting example, in some specific implementations, electronic device 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0034] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and communicating between system components. In some embodiments, one or more I / O devices and sensors 306 include inertial measurement units (IMUs), accelerometers, gyroscopes, thermometers, one or more physiological sensors (e.g., blood pressure monitors, heart rate monitors, blood oxygen sensors, blood glucose sensors, etc.), one or more microphones, one or more speakers, haptic engines, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0035] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, electronic device 120 includes a single XR display. Alternatively, electronic device 120 may include an XR display for each of the user's eyes. In some embodiments, one or more XR displays 312 are capable of displaying MR and VR content.

[0036] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of a user's face (including the user's eyes) (and thus may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward in order to acquire image data corresponding to a scene that the user would see when the electronic device 120 is not present (and thus may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0037] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optional operating system 330 and XR rendering module 340.

[0038] Operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some implementations, XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. Therefore, in various implementations, XR presentation module 340 includes a data acquisition unit 342, a data association unit 344, an XR presentation unit 346, and a data transmission unit 348.

[0039] In some specific implementations, the data acquisition unit 342 is configured to acquire data from at least... Figure 1 The controller 110 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). In various embodiments, the data acquisition unit 342 is configured to acquire spatial features of the environment and XR content. To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic components for instructions, as well as heuristics and metadata for heuristics.

[0040] In some implementations, the positioning unit 344 is configured to determine the location of the electronic device 120 in the XR environment, and in some implementations, to determine the location where virtual content will be displayed in the XR environment. To this end, in various implementations, the data association unit 344 includes instructions and / or logic components for the instructions, as well as heuristics and metadata for the heuristics.

[0041] In some implementations, the XR rendering unit 346 is configured to render XR content via one or more XR displays 312, for example, at a location determined by the positioning unit 344. For this purpose, in various implementations, the XR rendering unit 346 includes instructions and / or logic components for instructions, as well as heuristics and metadata for heuristics.

[0042] In some implementations, the data transmission unit 348 is configured to transmit at least data (e.g., presentation data, location data, etc.) to the controller 110. Therefore, in various implementations, the data transmission unit 348 includes instructions and / or logic components for the instructions, as well as heuristics and metadata for the heuristics.

[0043] Although the data acquisition unit 342, positioning unit 344, XR presentation unit 346 and data transmission unit 348 are shown residing on a single device (e.g., electronic device 120), it should be understood that in other embodiments, any combination of the data acquisition unit 342, positioning unit 344, XR presentation unit 346 and data transmission unit 348 may be located in a separate computing device.

[0044] also, Figure 3This serves more as a functional description of various features that may exist in a particular implementation, and differs from the structural diagrams of the specific implementations described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separate. For example, Figure 3 Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented in various specific implementations through one or more functional blocks. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.

[0045] Figure 4A It shows at least partially made of electronic devices (e.g. Figure 3 The XR environment 410 is presented by the display of the electronic device 120. The XR environment 410 is based on the first physical environment in which the electronic device exists.

[0046] XR environment 410 includes multiple objects, including one or more physical objects of the physical environment (e.g., floor 411, first wall 412, ceiling 413, and second wall 414) and one or more virtual objects (e.g., virtual navigation application window 441, virtual turn indicator 442, and virtual clock 490). In various specific implementations, some objects (e.g., physical objects and virtual turn indicator 442) are presented at positions in XR environment 410, for example, at positions defined by three coordinates in a common three-dimensional (3D) XR coordinate system, such that while some objects may exist in the physical world and others may not, spatial relationships (e.g., distance or orientation) can be defined between them. Therefore, when an electronic device moves (e.g., changes position and / or orientation) in XR environment 410, objects move on the display of the electronic device but maintain their positions in XR environment 410. Such virtual objects that move on the display in response to movement of the electronic device but maintain their positions in XR environment 410 are referred to as world-locked objects. In various specific implementations, the position of certain virtual objects (such as the virtual navigation application window 441) within the XR environment 410 changes based on the user's body pose. Such virtual objects are referred to as body-locked objects. For example, when a user is navigating within the XR environment 410, the virtual navigation application window 441 maintains a position approximately one meter in front of the user and half a meter to the user's right (e.g., position and orientation relative to the user's torso). When the user's head moves while the user's body remains stationary, the virtual navigation application window 441 appears at a fixed position within the XR environment 410.

[0047] In various specific implementations, certain virtual objects (e.g., a virtual clock 490) are displayed at a location on the display such that the objects remain stationary on the display of the electronic device as the electronic device moves within the XR environment 410. Such virtual objects that maintain their position on the display in response to movement of the electronic device are referred to as display-locked objects.

[0048] In XR environment 410, the first wall 412 intersects the floor 411 at intersection 421 and terminates at edge 422. A virtual turn indicator 441 is displayed on the floor 411 between the first wall 412 and the second wall 414 in XR environment 410, just past edge 422. The virtual navigation application window 441 provides the user with instructions to turn left at 17 meters.

[0049] As described above, the virtual turn indicator 442 is a world-locked virtual object. Therefore, in various embodiments, the virtual turn indicator 442 is associated with a position within the XR environment 410 defined by a set of three-dimensional XR coordinates in the three-dimensional coordinate system of the XR environment 410. In various embodiments, the device's pose in the XR environment 410 includes the device position defined by a set of three-dimensional XR coordinates and the device orientation defined by a set of three-degree-of-freedom angles. In various embodiments, an estimate of the device's pose is determined using visual inertial odometry (VIO) (e.g., using a camera and an inertial measurement unit (IMU)).

[0050] Based on the estimation of the device's pose and the position of the virtual turn indicator 442 in the XR environment 410, the electronic device determines the position where the virtual turn indicator 442 should be displayed on the monitor.

[0051] Figure 4B An XR environment 410 with an inaccurate estimate of the device's pose is illustrated. Specifically, the estimated position of the device differs from its actual position because the estimated position is approximately one meter closer to the second wall 414 than the actual position. Therefore, the virtual navigation application window 441 provides the user with instructions to turn left at 16 (instead of 17) meters. Furthermore, the estimated orientation of the device differs from its actual orientation because the estimated yaw angle is approximately 10 degrees to the left of the actual orientation. Therefore, the virtual turn indicator 442 is displayed to the right, partially overlapping the second wall 414 rather than between the first wall 412 and the second wall 414. This inaccurate placement of the turn indicator 442 using VIO can be caused by the timing of creating the map or three-dimensional coordinate system of the XR environment 410. Figure 4B The changes in the XR environment 410 between indicated times (e.g., placement of physical objects, lighting conditions, etc.) are formed.

[0052] In various specific implementations, electronic devices acquire an environmental model of the physical environment, which includes one or more spatial features of the physical environment. Each spatial feature is defined by its location in a three-dimensional XR coordinate system.

[0053] In various implementations, the spatial feature location includes one or more points. For example, in various implementations, the environmental model of the physical environment includes the spatial feature of the point where intersection 421 and edge 422 meet. In various implementations, the spatial feature location is defined by one or more sets of three-dimensional XR coordinates.

[0054] In various embodiments, the spatial feature location is a line, ray, or line segment. For example, in various embodiments, the environment model of the physical environment includes spatial features corresponding to intersection 421. In various embodiments, the spatial feature location is defined by linear equations. In various embodiments, the spatial feature location is defined by a first set of three-dimensional XR coordinates and a second set of three-dimensional XR coordinates. In various embodiments, the spatial feature location is defined by the first set of three-dimensional XR coordinates and orientation, and in some embodiments, by length.

[0055] In various embodiments, the location of a spatial feature is a plane or a plane segment. For example, in various embodiments, the environmental model of the physical environment includes spatial features corresponding to the first wall 412. In various embodiments, the location of a spatial feature is defined by plane equations. In various embodiments, the location of a spatial feature is defined by one or more sets of three-dimensional coordinates within a plane or plane segment.

[0056] Based on an (inaccurate) estimate of the device's pose and the spatial feature positions corresponding to intersection 421 and edge 422, the electronic device will expect the spatial feature corresponding to intersection 421 to be located at position 431 and the spatial feature corresponding to edge 422 to be located at position 432. In response to detection at... Figure 4B At the intersection 421 and edge 422 shown, the electronic device updates its estimate of the device's pose to display a virtual turn indicator 442 at the correct position on the display.

[0057] Therefore, for each spatial feature corresponding to intersection 421 and edge 422, the electronic device obtains a first spatial feature position from the environment model. The electronic device further detects the spatial features corresponding to intersection 421 and edge 422 and estimates a second spatial feature position based on an estimate of the device's pose. The electronic device determines the difference between the first and second spatial feature positions and updates the estimate of the device's pose by removing this difference. Furthermore, the electronic device determines the position for displaying a virtual turn indicator 422 on the display based on the updated estimate of the device's pose, and displays the virtual turn indicator 422 at that position on the display.

[0058] Figure 5 This is a flowchart illustrating a method 500 based on some specific implementations of a positioning device. In various specific implementations, method 500 comprises a device including one or more processors and non-transitory memories (e.g., Figure 3 The method 500 is executed by an electronic device 120. In some embodiments, the method 500 is executed by a processing logic unit (including hardware, firmware, software, or a combination thereof). In some embodiments, the method 500 is executed by a processor that executes instructions (e.g., code) stored in a non-transitory computer-readable medium (e.g., memory).

[0059] Method 500 begins in block 510, wherein the device acquires an estimate of its pose in the environment. In various embodiments, acquiring the estimate of the device's pose includes receiving data from a visual inertial odometry (VIO) system, including, for example, a camera and an inertial measurement unit (IMU). In various embodiments, the device's pose includes the device's position and orientation. In various embodiments, the device's position includes a set of three-dimensional coordinates in the three-dimensional coordinate system of the environment. In various embodiments, the device's orientation includes a set of three-degree-of-freedom angles in the three-dimensional coordinate system of the environment.

[0060] Method 500 continues in block 520, wherein the device acquires an environment model of the environment, the environment model including spatial features in the environment defined by a first spatial feature location. In various specific embodiments, the first spatial feature location is a line, ray, or line segment. For example, in Figure 4B In various specific implementations, the environmental model of the physical environment includes spatial features corresponding to the edge 422 defined by line segments. In various specific implementations, the location of the first spatial feature is a plane or a planar segment. For example, in Figure 4B In various specific implementations, the environmental model of the physical environment includes spatial features corresponding to the wall 412 defined by the planar segment.

[0061] In various specific implementations, the first spatial feature location includes at least one set of coordinates in the three-dimensional coordinate system of the environment. For example, in Figure 4BIn various specific implementations, the spatial features corresponding to the edge 422 are defined by a first set of three-dimensional coordinates corresponding to the position where the edge 422 is joined with the canopy 413 and a second set of three-dimensional coordinates corresponding to the position where the edge 422 is joined with the floor 411.

[0062] Method 500 continues in block 530, wherein the device determines a second spatial feature location of the spatial feature based on an estimate of the device's pose. In various embodiments, determining the second spatial feature location of the spatial feature includes capturing an image of the environment using an image sensor and detecting spatial features in the image of the environment. In various embodiments, the device determines a plurality of two-dimensional coordinates of a point in the image, which correspond to a first plurality of three-dimensional coordinates of the point of the spatial feature obtained from the environment model. Furthermore, based on an estimate of the device's pose (and, in various embodiments, inherent parameters of the image sensor) and the plurality of two-dimensional coordinates of the point in the image, the device determines a second plurality of three-dimensional coordinates of the point at the spatial feature location. For example, in various embodiments, the device uses pinhole camera model equations to determine the second plurality of three-dimensional coordinates.

[0063] In various implementations, determining the second spatial feature location includes receiving depth data of the environment from a depth sensor and detecting spatial features within the depth data. In various implementations, the device determines multiple depths from the device to points in the depth data, which correspond to a first plurality of three-dimensional coordinates of points of spatial features obtained from an environment model. Furthermore, based on an estimate of the device's pose and the multiple depths, the device determines a second plurality of three-dimensional coordinates of the point at the spatial feature location.

[0064] Method 500 continues in block 540, wherein the device determines an updated estimate of the device's pose based on a first spatial feature position and a second spatial feature position. In various embodiments, determining the updated estimate of the device's pose includes determining the difference between the first and second spatial feature positions and removing that difference from the estimate of the device's pose. In various embodiments, determining the difference between the first and second spatial feature positions includes determining the average difference between a plurality of points at the first spatial feature position and corresponding plurality of points at the second spatial feature position. In various embodiments, removing the difference from the estimate of the device's pose includes updating the estimate of the device's position. For example, in various embodiments, removing the difference from the estimate of the device's pose includes subtracting the difference from the device's pose. While this may be suitable for minor inaccuracies, such subtraction only affects the translation of the estimate of the device's position and not the estimate of its orientation. Therefore, in various embodiments, removing the difference from the estimate of the device's pose includes updating the estimate of the device's orientation. For example, in various specific implementations, removing discrepancies from the estimation of the device's pose involves selecting an estimate of the device's pose that minimizes a cost function of the difference between a first spatial feature position and a second spatial feature position (which is based on the estimation of the device's pose). Such optimizations can be computationally efficient for small inaccuracies.

[0065] In various specific implementations, method 500 also includes obtaining the content to be rendered at the content location in the environment. For example, in Figure 4A In this process, the electronic device acquires data regarding the virtual turn signal sign 442. Method 500 includes: determining the location of content to be displayed on a display based on an estimate of an update to the device's pose; and displaying the content at that location on the display. For example, in Figure 4A In the middle, the electronic device is displayed in the virtual turn indicator sign 442 between the first wall 412 and the second wall 414.

[0066] While various aspects of specific embodiments within the scope of the appended claims have been described above, it should be apparent that the various features of the above-described embodiments can be embodied in a wide variety of forms, and any particular structure and / or function described above are merely illustrative. Based on this disclosure, those skilled in the art will understand that the aspects described herein can be implemented independently of any other aspects, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement an apparatus and / or practice a method. Furthermore, such an apparatus and / or such a method can be implemented using other structures and / or functions besides or different from one or more aspects set forth herein.

[0067] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.

[0068] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and the appended claims, the singular forms “a” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the term “comprising” as used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0069] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when the prerequisite is detected" or "in response to detection" that the prerequisite is true, depending on the context.

Claims

1. A method comprising: at a device comprising an image sensor, one or more processors, and non-transitory memory: obtaining an estimate of a pose of the device in an environment; obtaining an environment model of the environment, the environment model comprising a spatial feature defined by a first spatial feature location in the environment; determining a second spatial feature location of the spatial feature based on the estimate of the pose of the device; and determining an updated estimate of the pose of the device based on the first spatial feature location and the second spatial feature location, at least in part by determining a difference between the first spatial feature location and the second spatial feature location and removing the difference from the estimate of the pose of the device.

2. The method of claim 1, further comprising: obtaining content to be presented at a content location in the environment; determining a location on a display to display the content based on the updated estimate of the pose of the device; and displaying the content at the location on the display.

3. The method of claim 1, wherein obtaining the estimate of the pose of the device comprises receiving data from a visual-inertial odometry system.

4. The method of claim 1, wherein the first spatial feature location is a line, a ray, or a line segment.

5. The method of claim 1, wherein the first spatial feature location is a plane or a plane segment.

6. The method of claim 1, wherein the first spatial feature location comprises at least one set of three-dimensional coordinates in a three-dimensional coordinate system of the environment.

7. The method of claim 1, wherein determining a second spatial feature location of the spatial feature comprises capturing an image of the environment using an image sensor and detecting the spatial feature in the image of the environment.

8. The method of claim 1, wherein determining the second spatial feature location of the spatial feature comprises receiving depth data of the environment from a depth sensor and detecting the spatial feature in the depth data of the environment.

9. A device comprising: non-transitory memory; and one or more processors to: obtain an estimate of a pose of the device in an environment; obtain an environment model of the environment, the environment model comprising a spatial feature defined by a first spatial feature location in the environment; determine a second spatial feature location of the spatial feature based on the estimate of the pose of the device; and determine an updated estimate of the pose of the device based on the first spatial feature location and the second spatial feature location, at least in part by determining a difference between the first spatial feature location and the second spatial feature location and removing the difference from the estimate of the pose of the device.

10. The device of claim 9, wherein the one or more processors are further to: obtain content to be presented at a content location in the environment; determine, based on the updated estimate of the pose of the device, a position on a display to display the content; and display the content at the position on the display.

11. The device of claim 9, wherein the one or more processors are to obtain the estimate of the pose of the device by receiving data from a visual-inertial odometry system.

12. The device of claim 9, wherein the first spatial feature position is a line, a ray, or a line segment.

13. The device of claim 9, wherein the first spatial feature position is a plane or a plane segment.

14. The device of claim 9, wherein the first spatial feature position comprises at least one set of three-dimensional coordinates in a three-dimensional coordinate system of the environment.

15. The device of claim 9, wherein the one or more processors are to determine the second spatial feature position of the spatial feature by capturing an image of the environment using an image sensor and detecting the spatial feature in the image of the environment.

16. The device of claim 9, wherein the one or more processors are to determine the second spatial feature position of the spatial feature by receiving depth data of the environment from a depth sensor and detecting the spatial feature in the depth data of the environment.

17. A non-transitory memory storing one or more programs, which, when executed by one or more processors of a device, cause the device to: obtain an estimate of a pose of the device in an environment; obtain an environment model of the environment, the environment model comprising a spatial feature in the environment defined by a first spatial feature position; determine, based on the estimate of the pose of the device, a second spatial feature position of the spatial feature; and determine, based on the first spatial feature position and the second spatial feature position, an updated estimate of the pose of the device at least in part by determining a difference between the first spatial feature position and the second spatial feature position and removing the difference from the estimate of the pose of the device.

18. The non-transitory memory of claim 17, wherein the one or more programs, which when executed, further cause the device to: obtain content to be presented at a content position in the environment; determine, based on the updated estimate of the pose of the device, a position on a display to display the content; and display the content at the position on the display. ​

Citation Information

Patent Citations

  • Cross reality system with buffering for localization accuracy

    US20210264674A1