Simultaneous localization and mapping using cameras capturing multiple spectra

By using a multi-spectral camera system, combining visible light and infrared spectroscopy cameras, the problem of limited positioning and mapping accuracy of a single spectroscopy camera in different environments is solved, and efficient mapping and location under multiple lighting conditions is achieved.

CN116529767BActive Publication Date: 2025-09-02QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080105593.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-01
Publication Date
2025-09-02
Estimated Expiration
2040-10-01

AI Technical Summary

Technical Problem

In the prior art, it is difficult for a single spectral camera to achieve efficient mapping and positioning at the same time under different environments, especially under lighting changes or adverse conditions, the positioning and mapping accuracy is limited.

Method used

A multi-spectral camera system is used, combining visible light and infrared spectral cameras to capture images in different environments, and using the advantages of the two cameras to generate a set of coordinates of features and update the environment map.

Benefits of technology

Under different lighting conditions, the accuracy and stability of positioning and mapping are improved, the error of a single spectral camera in disadvantaged environments is reduced, and the adaptability of the system is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116529767B_ABST
    Figure CN116529767B_ABST
Patent Text Reader

Abstract

A device that performs image processing techniques is described. The device includes a first camera and a second camera that respond to different light spectra, such as a visible light spectrum and an infrared spectrum. When the device is in a first position in an environment, the first camera captures a first image of the environment, and the second camera captures a second image of the environment. The device determines a single set of coordinates for the feature based on depictions of the feature identified in both the first image and the second image. The device generates and / or updates a map of the environment based on the set of coordinates for the feature. The device can move to another location in the environment and continue to capture images and update the map based on the images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to image processing, and more particularly to technology and techniques for simultaneous localization and mapping (SLAM) using a first camera that captures a first spectrum and a second camera that captures a second spectrum. Background Art

[0002] Simultaneous localization and mapping (SLAM) is a computational geometry technique used in devices such as robotic systems and autonomous vehicle systems. In SLAM, a device builds and updates a map of an unknown environment. The device can simultaneously track the device's position within the environment. The device typically performs mapping and localization based on sensor data collected by one or more sensors on the device. For example, a device can be activated in a specific room of a building and can move throughout the interior of the building, capturing sensor measurements. As the device moves throughout the interior of the building, the device can generate and update a map of the interior of the building based on the sensor measurements. As the device moves throughout the interior of the building and forms a map, the device can track its own position in the map. Visual SLAM (VSLAM) is a SLAM technology that performs mapping and localization based on visual data collected by one or more cameras of the device. Different types of cameras can capture images based on different light spectra, such as the visible light spectrum or the infrared spectrum. Some cameras are not suitable for use in certain environments or situations. Summary of the Invention

[0003] This document describes systems, apparatuses, methods, and computer-readable media (collectively referred to herein as the "systems and techniques") for performing visual simultaneous localization and mapping (VSLAM) using a device with multiple cameras. As the device moves throughout an environment, the device performs mapping of the environment and localizing itself within the environment based on visual data (and / or other data) collected by the device's cameras. The cameras may include a first camera that captures images by receiving light from a first spectrum and a second camera that captures images by receiving light from a second spectrum. For example, the first spectrum may be the visible light spectrum, and the second spectrum may be the infrared spectrum. Different camera types may offer advantages in certain environments and disadvantages in others. For example, a visible light camera may capture clear images in well-lit environments but may be sensitive to changes in lighting. VSLAM may not be possible using only visible light cameras when the environment is poorly lit or when the lighting changes over time (e.g., when the lighting is dynamic and / or inconsistent). Performing VSLAM using cameras that capture multiple spectrums can retain the advantages of each camera type while mitigating their disadvantages. For example, the device's first and second cameras may both capture images of the environment, and a depiction of features in the environment may appear in both images. The device can generate a set of coordinates for the features based on these delineations of the features, and can update a map of the environment based on the set of coordinates for the features. In situations where one of the cameras is at a disadvantage, the disadvantaged camera can be disabled. For example, if the lighting level of the environment is below a lighting threshold, the visible light camera can be disabled.

[0004] In another example, a device for image processing is provided. The device includes one or more memory units storing instructions. The device includes one or more processors executing the instructions, wherein execution of the instructions by the one or more processors causes the one or more processors to perform a method. The method includes receiving a first image of an environment captured by a first camera. The first camera is responsive to a first spectrum. The method includes receiving a second image of the environment captured by a second camera. The second camera is responsive to a second spectrum. The method includes identifying that a feature of the environment is depicted in both the first image and the second image. The method includes determining a set of coordinates for the feature based on the first depiction of the feature in the first image and the second depiction of the feature in the second image. The method includes updating a map of the environment based on the set of coordinates for the feature.

[0005] In one example, a method of image processing is provided. The method includes receiving image data captured by an image sensor. The method includes receiving a first image of an environment captured by a first camera. The first camera is responsive to a first spectrum. The method includes receiving a second image of the environment captured by a second camera. The second camera is responsive to a second spectrum. The method includes identifying that a feature of the environment is depicted in both the first image and the second image. The method includes determining a set of coordinates for the feature based on a first depiction of the feature in the first image and a second depiction of the feature in the second image. The method includes updating a map of the environment based on the set of coordinates for the feature.

[0006] In another example, a non-transitory computer-readable storage medium having a program thereon is provided. The program is executable by a processor to perform an image processing method. The method includes receiving a first image of an environment captured by a first camera. The first camera is responsive to a first spectrum. The method includes receiving a second image of the environment captured by a second camera. The second camera is responsive to a second spectrum. The method includes identifying that a feature of the environment is depicted in both the first image and the second image. The method includes determining a set of coordinates for the feature based on a first depiction of the feature in the first image and a second depiction of the feature in the second image. The method includes updating a map of the environment based on the set of coordinates for the feature.

[0007] In another example, an apparatus for image processing is provided. The apparatus includes components for receiving a first image of an environment captured by a first camera, the first camera responsive to a first spectrum. The apparatus includes components for receiving a second image of the environment captured by a second camera, the second camera responsive to a second spectrum. The apparatus includes components for identifying that a feature of the environment is depicted in both the first image and the second image. The apparatus includes components for determining a set of coordinates for the feature based on a first depiction of the feature in the first image and a second depiction of the feature in the second image. The apparatus includes components for updating a map of the environment based on the set of coordinates for the feature.

[0008] In some aspects, the first spectrum is at least part of a visible light (VL) spectrum, and the second spectrum is different from the VL spectrum. In some aspects, the second spectrum is at least part of an infrared (IR) spectrum, and wherein the first spectrum is different from the IR spectrum.

[0009] In some aspects, the coordinate set of the feature includes three coordinates corresponding to three spatial dimensions. In some aspects, the apparatus or device includes a first camera and a second camera. In some aspects, the apparatus and device includes at least one of a mobile handheld device, a head mounted display (HMD), a vehicle, and a robot.

[0010] In some aspects, the first camera captures a first image when the device or apparatus is in a first position, and wherein the second camera captures a second image when the device or apparatus is in the first position. In some aspects, the methods, apparatus, and computer-readable media described above further include: determining a set of coordinates for a first position of the device or apparatus within an environment based on the set of coordinates for the features. In some aspects, the methods, apparatus, and computer-readable media described above further include: determining a posture of the device or apparatus when the device or apparatus is in the first position based on the set of coordinates for the features, wherein the posture of the device or apparatus includes at least one of a pitch of the device or apparatus, a roll of the device or apparatus, and a yaw of the device or apparatus.

[0011] In some aspects, the methods, apparatuses, and computer-readable media described above further include: identifying that the device or apparatus has moved from a first position to a second position; receiving a third image of the environment captured by a second camera while the device or apparatus is in the second position; identifying that a feature of the environment is depicted in at least one of the third image and a fourth image from the first camera; and tracking the feature based on one or more depictions of the feature in at least one of the third image and the fourth image. In some aspects, the methods, apparatuses, and computer-readable media described above further include: determining a set of coordinates of the second position of the device or apparatus within the environment based on the tracking features. In some aspects, the methods, apparatuses, and computer-readable media described above further include: determining a pose of the device or apparatus while the device or apparatus is in the second position based on the tracking features, wherein the pose of the device or apparatus includes at least one of a pitch of the device or apparatus, a roll of the device or apparatus, and a yaw of the device or apparatus. In some aspects, the methods, apparatuses, and computer-readable media described above further include: generating an updated set of coordinates for the feature in the environment by updating the set of coordinates for the feature in the environment based on the tracking features; and updating a map of the environment based on the updated set of coordinates for the feature.

[0012] In some aspects, the methods, apparatus, and computer-readable media described above further include: identifying that an illumination level of the environment is above a minimum illumination threshold when the device or apparatus is in the second position; and receiving a fourth image of the environment captured by the first camera when the device or apparatus is in the second position, wherein the tracked feature is based on a third depiction of the feature in the third image and a fourth depiction of the feature in the fourth image. In some aspects, the methods, apparatus, and computer-readable media described above further include: identifying that an illumination level of the environment is below a minimum illumination threshold when the device or apparatus is in the second position, wherein the tracked feature is based on the third depiction of the feature in the third image. In some aspects, the methods, apparatus, and computer-readable media described above further include: wherein the tracked feature is also based on at least one of a set of coordinates of the feature, a first depiction of the feature in the first image, and a second depiction of the feature in the second image.

[0013] In some aspects, the methods, apparatus, and computer-readable media described above further include: identifying that the device or apparatus has moved from a first position to a second position; receiving a third image of the environment captured by a second camera while the device or apparatus is in the second position; identifying that a second feature of the environment is depicted in at least one of the third image and a fourth image from the first camera; determining a second set of coordinates for the second feature based on one or more depictions of the second feature in at least one of the third image and the fourth image; and updating a map of the environment based on the second set of coordinates for the second feature. In some aspects, the methods, apparatus, and computer-readable media described above further include: determining a set of coordinates for the second position of the device or apparatus within the environment based on the updated map. In some aspects, the methods, apparatus, and computer-readable media described above further include: determining a posture of the device or apparatus while the device or apparatus is in the second position based on the updated map, wherein the posture of the device or apparatus includes at least one of the pitch of the device or apparatus, the roll of the device or apparatus, and the yaw of the device or apparatus.

[0014] In some aspects, the methods, apparatus, and computer-readable media described above further include: identifying that an illumination level of the environment is above a minimum illumination threshold when the device or apparatus is in the second position; and receiving a fourth image of the environment captured by the first camera while the device or apparatus is in the second position, wherein determining the second set of coordinates for the second feature is based on the first depiction of the second feature in the third image and the second depiction of the second feature in the fourth image. In some aspects, the methods, apparatus, and computer-readable media described above further include: identifying that an illumination level of the environment is below a minimum illumination threshold when the device or apparatus is in the second position, wherein determining the second set of coordinates for the second feature is based on the first depiction of the second feature in the third image.

[0015] In some aspects, determining the set of coordinates for the feature includes determining a transformation between a first set of coordinates for the feature corresponding to the first image and a second set of coordinates for the feature corresponding to the second image. In some aspects, the above methods, apparatus, and computer-readable media further include: generating a map of the environment before updating the map of the environment. In some aspects, updating the map of the environment based on the set of coordinates for the feature includes adding a new map area to the map, the new map area including the set of coordinates for the feature. In some aspects, updating the map of the environment based on the set of coordinates for the feature includes revising a map area of ​​the map, the map area including the set of coordinates for the feature. In some aspects, the feature is at least one of an edge and a corner.

[0016] In some aspects, the device or apparatus comprises a camera, a mobile device or apparatus (e.g., a mobile phone or so-called "smartphone" or other mobile device or apparatus), a wireless communication device or apparatus, a mobile handheld device, a wearable device or apparatus, a head-mounted display (HMD), an extended reality (XR) device or apparatus (e.g., a virtual reality (VR) device or apparatus, an augmented reality (AR) device or apparatus, or a mixed reality (MR) device or apparatus), a robot, a vehicle, an unmanned vehicle, an autonomous vehicle, a personal computer, a laptop computer, a server computer, or other device or apparatus. In some aspects, one or more processors comprise an image signal processor (ISP). In some aspects, the device or apparatus comprises a first camera. In some aspects, the device or apparatus comprises a second camera. In some aspects, the device or apparatus comprises one or more additional cameras for capturing one or more additional images. In some aspects, the device or apparatus comprises an image sensor that captures image data corresponding to the first image, the second image, and / or the one or more additional images. In some aspects, the device or apparatus further includes a display for displaying the first image, the second image, another image, a map, one or more notifications associated with image processing, and / or other displayable data.

[0017] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification, any or all drawings, and each claim.

[0018] The foregoing and other features and embodiments will become more fully apparent upon reference to the following description, claims, and appended drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The illustrative embodiments of the present application are described in detail below with reference to the following drawings:

[0020] Figure 1 is a block diagram illustrating an example of the architecture of an image capture and processing device according to some examples;

[0021] Figure 2 is a conceptual diagram illustrating an example of a technique for performing visual simultaneous localization and mapping (VSLAM) using a camera of a VSLAM device according to some examples;

[0022] Figure 3 is a conceptual diagram illustrating an example of a technique for performing VSLAM using a visible light (VL) camera and an infrared (IR) camera of a VSLAM device according to some examples;

[0023] Figure 4 is a conceptual diagram illustrating an example of a technique for performing VSLAM using an infrared (IR) camera of a VSLAM device according to some examples;

[0024] Figure 5 is a conceptual diagram illustrating two images of the same environment captured under different lighting conditions according to some examples;

[0025] Figure 6A is a perspective diagram illustrating an unmanned ground vehicle (UGV) performing VSLAM according to some examples;

[0026] Figure 6B is a perspective diagram illustrating an unmanned aerial vehicle (UAV) performing VSLAM according to some examples;

[0027] Figure 7A is a perspective view illustrating a head mounted display (HMD) performing VSLAM according to some examples;

[0028] Figure 7B is a diagram illustrating some examples Figure 7A A perspective view of a head-mounted display (HMD) worn by a user;

[0029] Figure 7C is a perspective view illustrating a front surface of a mobile handheld device using a forward-facing camera to perform VSLAM according to some examples;

[0030] Figure 7D is a perspective view illustrating a rear surface of a mobile handheld device using a rear-facing camera to perform VSLAM according to some examples;

[0031] Figure 8 is a conceptual diagram illustrating extrinsic calibration of a VL camera and an IR camera according to some examples;

[0032] Figure 9 is a conceptual diagram illustrating a transformation between coordinates of a feature detected by an IR camera and coordinates of the same feature detected by a VL camera according to some examples;

[0033] Figure 10A is a conceptual diagram illustrating feature associations between coordinates of features detected by an IR camera and coordinates of the same features detected by a VL camera according to some examples;

[0034] Figure 10B is a conceptual diagram illustrating example descriptor styles for features according to some examples;

[0035] Figure 11 is a conceptual diagram illustrating an example of joint map optimization according to some examples;

[0036] Figure 12 is a conceptual diagram illustrating feature tracking and stereo matching according to some examples;

[0037] Figure 13A is a conceptual diagram illustrating stereo matching between coordinates of a feature detected by an IR camera and coordinates of the same feature detected by a VL camera according to some examples;

[0038] Figure 13B is a conceptual diagram illustrating triangulation between coordinates of a feature detected by an IR camera and coordinates of the same feature detected by a VL camera according to some examples;

[0039] Figure 14A is a conceptual diagram illustrating monocular matching between coordinates of a feature detected by a camera in an image frame and coordinates of the same feature detected by the camera in a subsequent image frame according to some examples;

[0040] Figure 14B is a conceptual diagram illustrating triangulation between coordinates of a feature detected by a camera in an image frame and coordinates of the same feature detected by the camera in a subsequent image frame according to some examples;

[0041] Figure 15 is a conceptual diagram illustrating keyframe-based fast relocalization;

[0042] Figure 16 is a conceptual diagram illustrating fast relocalization based on keyframes and centroid points according to some examples;

[0043] Figure 17 is a flow chart illustrating an example of an image processing technique according to some examples; and

[0044] Figure 18 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. DETAILED DESCRIPTION

[0045] Provide certain aspects and embodiments of the present disclosure below. As will be apparent to those skilled in the art, some of these aspects and embodiments can be applied independently, and some of them can be applied in combination. In the following description, for the purpose of explanation, specific details are set forth in order to provide a comprehensive understanding of the embodiments of the application. However, it will be apparent that various embodiments can be put into practice without these specific details. The accompanying drawings and description are not intended to be restrictive.

[0046] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Specifically, the following description of the exemplary embodiments will provide those skilled in the art with a description of implementations for implementing the exemplary embodiments. It should be understood that various changes may be made in the functionality and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0047] An image capture device (e.g., a camera) is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). The terms "image," "image frame," and "frame" are used interchangeably herein. An image capture device typically includes at least one lens that receives light from a scene and bends the light toward the image sensor of the image capture device. The light received by the lens passes through an aperture controlled by one or more control mechanisms and is received by the image sensor. The one or more control mechanisms can control exposure, focus, and / or zoom based on information from the image sensor and / or based on information from an image processor (e.g., a host or application process and / or an image signal processor). In some examples, the one or more control mechanisms include a motor or other control mechanism that moves the lens of the image capture device to a target lens position.

[0048] Simultaneous localization and mapping (SLAM) is a computational geometry technique used in devices such as robotic systems, autonomous vehicle systems, extended reality (XR) systems, head-mounted displays (HMDs), and the like. As mentioned above, XR systems can include, for example, augmented reality (AR) systems, virtual reality (VR) systems, and mixed reality (MR) systems. An XR system can be a head-mounted display (HMD) device. Using SLAM, a device can build and update a map of an unknown environment while tracking the device's position within that environment. The device can typically perform these tasks based on sensor data collected by one or more sensors on the device. For example, a device can be activated in a specific room of a building and can be moved throughout the building, mapping the entire interior of the building while tracking its own position within the map as the device forms the map.

[0049] Visual SLAM (VSLAM) is a SLAM technology that performs mapping and positioning based on visual data collected by one or more cameras of a device. In some cases, a monocular VSLAM device can use a single camera to perform VLAM. For example, a monocular VSLAM device can utilize a camera to capture one or more images of an environment and can determine distinctive visual features, such as corners or other points in one or more images. The device can move through the environment and can capture more images. The device can track the movement of those features in the continuous images captured when the device is in different positions, orientations and / or postures in the environment. The device can use these tracked features to generate a three-dimensional (3D) map and determine its own positioning within the map.

[0050] VSLAM can be performed using a visible light (VL) camera that detects light within the spectrum visible to the human eye. Some VL cameras only detect light within the spectrum visible to the human eye. An example of a VL camera is a camera that captures red (R), green (G), and blue (B) image data (referred to as RGB image data). The RGB image data can then be combined into a full color image. A VL camera that captures RGB image data can be referred to as an RGB camera. The camera can also capture other types of color images, such as images with luminance (Y) and chrominance (chrominance blue referred to as U or Cb, and chrominance red referred to as V or Cr) components. Such images can include YUV images, YC b C r Images, etc.

[0051] VL cameras typically capture clear images of well-lit environments. Features such as edges and corners are easily recognizable in clear images of well-lit environments. However, VL cameras typically have difficulty capturing clear images of poorly lit environments (such as environments captured at night and / or in dim light). Images of poorly lit environments captured by VL cameras may be unclear. For example, in unclear images of poorly lit environments, features such as edges and corners may be difficult or even impossible to recognize. A VSLAM device using a VL camera may not be able to detect certain features in a poorly lit environment that the VSLAM device can detect when the environment is well lit. In some cases, because the environment may appear different to the VL camera depending on the lighting of the environment, a VSLAM device using a VL camera may sometimes be unable to recognize parts of the environment that the VSLAM device has already observed due to changes in lighting conditions in the environment. The inability to recognize parts of the environment that the VSLAM device has already observed may cause errors in the positioning and / or mapping of the VSLAM device.

[0052] As described in more detail below, this document describes systems and techniques for performing VSLAM using a VSLAM device with multiple types of cameras. For example, systems and techniques can perform VSLAM using a VSLAM device including a VL camera and an infrared (IR) camera (or multiple VL cameras and / or multiple IR cameras). The VSLAM device can use the VL camera to capture one or more images of an environment, and can use the IR camera to capture one or more images of an environment. In some examples, the VSLAM device can detect one or more features in the VL image data from the VL camera and in the IR image data from the IR camera. The VSLAM device can determine a single coordinate set (e.g., three-dimensional coordinates) for a feature in one or more features based on the depiction of the features in the VL image data and in the IR image data. The VSLAM device can generate and / or update a map of the environment based on a coordinate set for a feature.

[0053] Further details regarding the systems and techniques are provided herein with respect to various figures. Figure 1 1 is a block diagram illustrating an example of the architecture of an image capture and processing system 100. Image capture and processing system 100 includes various components for capturing and processing an image of a scene (e.g., an image of scene 110). Image capture and processing system 100 can capture individual frames (or photographs) and / or can capture a video comprising multiple images (or video frames) in a particular sequence. Lens 115 of system 100 faces scene 110 and receives light from scene 110. Lens 115 bends the light toward image sensor 130. The light received by lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by image sensor 130.

[0054] The one or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. The one or more control mechanisms 120 may include a plurality of mechanisms and components; for example, the control mechanisms 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. The one or more control mechanisms 120 may also include additional control mechanisms beyond those illustrated, such as controls for analog gain, flash, HDR, depth of field, and / or other image capture properties.

[0055] Focus control mechanism 125B in control mechanism 120 can obtain a focus setting. In some examples, focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, focus control mechanism 125B can adjust the position of lens 115 relative to the position of image sensor 130. For example, based on the focus setting, focus control mechanism 125B can move lens 115 closer to or further away from image sensor 130 by actuating a motor or servo (or other lens mechanism), thereby adjusting the focus. In some cases, system 100 may include additional lenses, such as one or more microlenses on each photodiode of image sensor 130, each microlens bending light received from lens 115 toward the corresponding photodiode (before the light reaches the photodiode). The focus setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The focus setting can be determined using control mechanism 120, image sensor 130, and / or image processor 150. The focus setting can be referred to as an image capture setting and / or an image processing setting.

[0056] Exposure control mechanism 125A in control mechanism 120 can obtain an exposure setting. In some cases, exposure control mechanism 125A stores the exposure setting in a memory register. Based on this exposure setting, exposure control mechanism 125A can control the size of the aperture (e.g., aperture size or f-stop), the duration of time the aperture is open (e.g., exposure time or shutter speed), the sensitivity of image sensor 130 (e.g., ISO speed or film speed), the analog gain applied to image sensor 130, or any combination thereof. The exposure setting can be referred to as an image capture setting and / or an image processing setting.

[0057] Zoom control mechanism 125C of control mechanism 120 can obtain a zoom setting. In some examples, zoom control mechanism 125C stores the zoom setting in a memory register. Based on the zoom setting, zoom control mechanism 125C can control the focal length of a group of lens elements (lens group) including lens 115 and one or more additional lenses. For example, zoom control mechanism 125C can control the focal length of the lens group by actuating one or more motors or servos (or other lens mechanisms) to move one or more of the lenses relative to each other. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens group can include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens group can include a focusing lens (which in some cases can be lens 115) that first receives light from scene 110, and then the light passes through an afocal zoom system between the focusing lens (e.g., lens 115) and image sensor 130 before reaching image sensor 130. In some cases, an afocal zoom system can include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference of each other) and a negative (e.g., diverging, concave) lens between them. In some cases, zoom control mechanism 125C moves one or more lenses in the afocal zoom system, such as one or both of the negative lens and the positive lens.

[0058] Image sensor 130 includes one or more photodiode arrays or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by image sensor 130. In some cases, different photodiodes may be covered by different color filters and, therefore, may measure light that matches the color of the filter covering the photodiode. For example, a Bayer color filter includes a red filter, a blue filter, and a green filter, where each pixel of an image is generated based on red light data from at least one photodiode covered by a red filter, blue light data from at least one photodiode covered by a blue filter, and green light data from at least one photodiode covered by a green filter. Other types of color filters may use yellow, magenta, and / or cyan (also known as "emerald") filters instead of or in addition to the red, blue, and / or green filters. Some image sensors (e.g., image sensor 130) may lack color filters entirely and, instead, may use different photodiodes (in some cases stacked vertically) throughout the pixel array. Different photodiodes throughout the pixel array may have different spectral sensitivity curves and, therefore, respond to different wavelengths of light. Monochrome image sensors may also lack color filters and, therefore, lack color depth.

[0059] In some cases, image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which may be used for phase detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier to amplify the analog signal output by the photodiode, and / or an analog-to-digital converter (ADC) to convert the analog signal output by the photodiode (and / or amplified by the analog gain amplifier) ​​to a digital signal. In some cases, certain components or functions discussed with respect to one or more of control mechanism 120 may alternatively or additionally be included in image sensor 130. Image sensor 130 may be a charge coupled device (CCD) sensor, an electron multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal oxide semiconductor (CMOS), an N-type metal oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0060] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more of any other type of processor 1810 discussed with respect to the computing device 1800. The host processor 152 may be a digital signal processor (DSP) and / or other type of processor. In some implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system on a chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth TM, Global Positioning System (GPS), etc.), any combination thereof, and / or other components. The I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a serial general-purpose input / output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In one illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port, and the ISP 154 may communicate with the image sensor 130 using a MIPI port.

[0061] The image processor 150 may perform a number of tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. The image processor 150 may store image frames and / or processed images in a random access memory (RAM) 140 / 1020, a read-only memory (ROM) 145 / 1025, a cache, a memory unit, another storage device, or some combination thereof.

[0062] Various input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a trackpad, a touch-sensitive surface, a printer, any other output device 1835, any other input device 1845, or some combination thereof. In some cases, captions may be entered into the image processing device 105B via a physical keyboard or keypad of the I / O device 160, or via a virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O 160 may include one or more ports, jacks, or other connectors that enable wired connections between the system 100 and one or more peripheral devices, through which the system 100 may receive data from and / or send data to the one or more peripheral devices. The I / O 160 may also include one or more wireless transceivers that enable wireless connections between the system 100 and one or more peripheral devices, through which the system 100 may receive data from and / or send data to the one or more peripheral devices. Peripheral devices may include any of the types of I / O devices 160 discussed previously, and may themselves be considered I / O devices 160 once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector.

[0063] In some cases, the image capture and processing system 100 can be a single device. In some cases, the image capture and processing system 100 can be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some implementations, the image capture device 105A and the image processing device 105B can be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some implementations, the image capture device 105A and the image processing device 105B can be disconnected from each other.

[0064] like Figure 1 As shown, the vertical dotted line will Figure 1 1. The image capture and processing system 100 is divided into two parts, respectively representing an image capture device 105A and an image processing device 105B. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, some of the components shown in the image capture device 105A, such as the ISP 154 and / or the host processor 152, may be included in the image capture device 105A.

[0065] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smartphone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device and the image processing device 105B may include a computing device, such as a mobile handset, a desktop computer, or other computing device.

[0066] Although the image capture and processing system 100 is shown as including certain components, one of ordinary skill will understand that the image capture and processing system 100 may include more than Figure 1 . The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of the image capture and processing system 100 may include and / or be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device that implements the image capture and processing system 100.

[0067] In some cases, the image capture and processing system 100 can be part of or implemented by a device (referred to as a VSLAM device) capable of performing VSLAM. For example, the VSLAM device can include one or more image capture and processing systems 100, image capture system 105A, image processing system 105B, computing system 1800, or any combination thereof. For example, the VSLAM device can include a visible light (VL) camera and an infrared (IR) camera. The VL camera and the IR camera can each include at least one of the image capture and processing system 100, image capture device 105A, image processing device 105B, computing system 1800, or some combination thereof.

[0068] Figure 2 2 is a conceptual diagram 200 illustrating an example of a technology for performing VSLAM using a camera 210 of a visual simultaneous localization and mapping (VSLAM) device 205. In some examples, the VSLAM device 205 can be a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, an extended reality (XR) device, a head-mounted display (HMD), or some combination thereof. In some examples, the VSLAM device 205 can be a wireless communication device, a mobile device (e.g., a mobile phone or so-called "smart phone" or other mobile device), a wearable device, an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR0 device), a head-mounted display (HMD), a personal computer, a laptop computer, a server computer, an unmanned ground vehicle, an unmanned aerial vehicle, an unmanned water vehicle, an unmanned underwater vehicle, an unmanned vehicle, an autonomous vehicle, a vehicle, a robot, any combination thereof, and / or other devices.

[0069] The VSLAM device 205 includes a camera 210. The camera 210 can respond to light from a specific spectrum. The spectrum can be a subset of the electromagnetic (EM) spectrum. For example, the camera 210 can be a VL camera that responds to the visible light (VL) spectrum, an IR camera that responds to the infrared (IR) spectrum, a UV camera that responds to the ultraviolet (UV) spectrum, a camera that responds to light from another spectrum of another part of the electromagnetic spectrum, or some combination thereof. In some cases, the camera 210 can be a near-infrared (NIR) camera that responds to the NIR spectrum. The NIR spectrum can be a subset of the IR spectrum that is close to and / or adjacent to the VL spectrum.

[0070] Camera 210 can be used to capture one or more images, including image 215. VSLAM system 270 can use feature extraction engine 220 to perform feature extraction. Feature extraction engine 220 can use image 215 to perform feature extraction by detecting one or more features within the image. Features can be, for example, edges, corners, areas of color change, areas of luminosity change, or a combination thereof. In some cases, when feature extraction engine 220 fails to detect any features in image 215, feature extraction engine 220 may not be able to perform feature extraction on image 215. In some cases, feature extraction engine 220 may fail when it fails to detect at least a predetermined minimum number of features in image 215. If feature extraction engine 220 fails to successfully perform feature extraction on image 215, VSLAM system 270 does not continue further and can wait for the next image frame captured by camera 210.

[0071] When feature extraction engine 220 detects at least a predetermined minimum number of features in image 215, feature extraction engine 220 may successfully perform feature extraction on image 215. In some examples, the predetermined minimum number of features may be one, in which case feature extraction engine 220 successfully performs feature extraction by detecting at least one feature in image 215. In some examples, the predetermined minimum number of features may be greater than one and may be, for example, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, a number greater than 100, or a number between any two of the previously listed numbers. Images with one or more clearly delineated features may be maintained as keyframes in a map database, and their depictions of features may be used to track those features in other images.

[0072] Once the feature extraction engine 220 successfully performs feature extraction on one or more images 215, the VSLAM system 270 can use the feature tracking engine 225 to perform feature tracking. The feature tracking engine 225 can perform feature tracking by identifying features in the image 215 that have already been identified in one or more previous images. The feature tracking engine 225 can also track changes in one or more locations of features between different images. For example, the feature extraction engine 220 can detect the face of a specific person as a feature depicted in a first image. The feature extraction engine 220 can detect the same features (e.g., the face of the same person) depicted in a second image captured by the camera 210 after the first image and received therefrom. The feature tracking 225 can identify that these features detected in the first image and the second image are two depictions of the same features (e.g., the face of the same person). The feature tracking engine 225 can identify that the features have moved between the first image and the second image. For example, the feature tracking engine 225 can identify that the features are depicted on the right side of the first image and are depicted in the center of the second image.

[0073] The movement of the feature between the first image and the second image may be caused by movement of the subject within the captured scene between the time the camera 210 captured the first image and the time the camera 210 captured the second image. For example, if the feature is a person's face, the person may have walked across a portion of the captured scene between the time the camera 210 captured the first image and the time the camera 210 captured the second image, resulting in the feature being positioned differently in the second image than in the first image. The movement of the feature between the first image and the second image may be caused by movement of the camera 210 between the time the camera 210 captured the first image and the time the camera 210 captured the second image. In some examples, the VSLAM device 205 may be a robot or a vehicle and may move itself and / or its camera 210 between the time the camera 210 captured the first image and the time the camera 210 captured the second image. In some examples, the VSLAM device 205 may be a head-mounted display (HMD) worn by a user (e.g., an XR headset), and the user may move his or her head and / or body between the time the camera 210 captured the first image and the time the camera 210 captured the second image.

[0074] The VSLAM system 270 can identify a set of coordinates for each feature identified by the VSLAM system 270 using the feature extraction engine 220 and / or the feature tracking engine 225, which can be referred to as a map point. The set of coordinates for each feature can be used to determine a map point 240. The local map engine 250 can use the map points 240 to update a local map. The local map can be a map of a local area of ​​the environment. The local area can be the area where the VSLAM device 205 is currently located. The local area can be, for example, a room or a collection of rooms within the environment. The local area can be, for example, a collection of one or more rooms visible in the image 215. The set of coordinates for the map points corresponding to the features can be updated by the VSLAM system 270 using the map optimization engine 235 to increase accuracy. For example, by tracking features across multiple images captured at different times, the VSLAM system 270 can generate a set of coordinates for the map points of the features based on each image. An accurate set of coordinates for the map points of the features can be determined by triangulating or generating an average coordinate based on multiple map points for the features determined from different images. The map optimization engine 235 can update the local map using the local mapping engine 250 to update the coordinate set for the feature to use the accurate coordinate set determined using triangulation and / or averaging. Observing the same feature from different angles can provide additional information about the true location of the feature, which can be used to increase the accuracy of the map point 240.

[0075] The local map 250, together with the global map 255, can be part of the mapping system 275. The global map 255 can map the global area of ​​the environment. The VSLAM device 205 can be located in the global area of ​​the environment and / or in the local area of ​​the environment. The local area of ​​the environment can be smaller than the global area of ​​the environment. The local area of ​​the environment can be a subset of the global area of ​​the environment. The local area of ​​the environment can overlap with the global area of ​​the environment. In some cases, the local area of ​​the environment can include a portion of the environment that has not yet been merged into the global map by the map merging engine 257 and / or the global mapping engine 255. In some examples, the local map can include map points within such a portion of the environment that has not yet been merged into the global map. In some cases, the global map 255 can map the entire environment observed by the VSLAM device 205. The updates of the local map by the local mapping engine 250 can be merged into the global map using the map merging engine 257 and / or the global mapping engine 255, thereby keeping the global map up to date. In some cases, after a local map has been optimized using the map optimization engine 235, the local map may be merged with the global map using the map merging engine 257 and / or the global mapping engine 255 so that the global map is an optimized map. Map points 240 may be fed into a local map by the local mapping engine 250 and / or may be fed into a global map using the global mapping engine 255. The map optimization engine 235 may improve the accuracy of the map points 240 and the local and / or global maps. In some cases, such as in Figure 11 As illustrated in and discussed with respect to conceptual diagram 1100 of , the map optimization engine 235 may simplify the local map and / or the global map by replacing bundles of map points with centroid map points.

[0076] The VSLAM system 270 may also determine a pose 245 of the device 205 based on feature extraction and / or feature tracking performed by the feature extraction engine 220 and / or the feature tracking engine 225. The pose 245 of the device 205 may refer to the position of the device 205, the pitch of the device 205, the roll of the device 205, the yaw of the device 205, or some combination thereof. The pose 245 of the device 205 may refer to the pose of the camera 210 and, therefore, may include the position of the camera 210, the pitch of the camera 210, the roll of the camera 210, the yaw of the camera 210, or some combination thereof. The pose 245 of the device 205 may be determined relative to a local map and / or a global map. The pose 245 of the device 205 may be marked on the local map by the local mapping engine 250 and / or on the global map by the global mapping engine 255. In some cases, a history of the pose 245 may be stored by the local mapping engine 250 and / or by the global mapping engine 255 within the local map and / or the global map. The history of gestures 245 together may indicate a path that the VSLAM device 205 has traveled.

[0077] In some cases, when a feature that has previously been identified in a set of earlier captured images is not identified in image 215, feature tracking engine 225 may not be able to successfully perform feature tracking on image 215. In some examples, the set of earlier captured images may include all images captured during a time period that ends before the capture of image 215 and begins at a predetermined start time. The predetermined start time may be an absolute time, such as a specific time and date. The predetermined start time may be a relative time, such as a predetermined amount of time (e.g., 30 minutes) before the capture of image 215. The predetermined start time may be the time when VSLAM device 205 was most recently initialized. The predetermined start time may be the time when VSLAM device 205 most recently received an instruction to start the VSLAM process. The predetermined start time may be the time when VSLAM device 205 most recently determined that it entered a new room or new area of ​​the environment.

[0078] If the feature tracking engine 225 fails to successfully perform feature tracking of the image, the VSLAM system 270 can use the relocation engine 230 to perform relocation. The relocation engine 230 attempts to determine where the VSLAM device 205 is located in the environment. For example, the feature tracking engine 225 may not be able to identify any features from one or more previously captured images and / or from the local map 250. The relocation engine 230 can try to see if any features identified by the feature extraction engine 220 match any features in the global map. If one or more features identified by the feature extraction engine 220 of the VSLAM system 270 match one or more features in the global map 255, the relocation engine 230 successfully performs relocation by determining the map points 240 for one or more features and / or determining the posture 245 of the VSLAM device 205. The relocation engine 230 can also compare any features identified by the feature extraction engine 220 in the image 215 with the features in the keyframes stored with the local map and / or global map. Each keyframe can be an image that clearly depicts a specific feature, so that the image 230 can be compared with the keyframe to determine whether the image 230 also depicts the specific feature. If none of the features identified by the VSLAM system 270 during feature extraction 220 matches any feature in the global map and / or in any keyframe, the relocalization engine 230 has failed to successfully perform relocalization. If the relocalization engine 230 has failed to successfully perform relocalization, the VSLAM system 270 can exit and reinitialize the VSLAM process. Exiting and reinitializing can include generating the local map 250 and / or the global map 255 from scratch.

[0079] VSLAM equipment 205 can include transport, and VSLAM equipment 205 can move itself around the environment by this transport.For example, VSLAM equipment 205 can include one or more motors, one or more actuators, one or more wheels, one or more propellers, one or more turbines, one or more rotors, one or more wings, one or more wings, one or more gliders, one or more pedals, one or more legs, one or more feet, one or more pistons, one or more nozzles, one or more propellers, one or more sails, one or more other modes of transportation discussed herein or its combination.In some examples, VSLAM equipment 205 can be the equipment of vehicle, robot or any other type discussed herein.VSLAM equipment 205 including transport can use path planning engine 260 to perform path planning, to plan the path for VSLAM equipment 205 to move.Once path planning engine 260 has planned path for VSLAM equipment 205, VSLAM equipment 205 just can use mobile actuator 265 to perform mobile actuation to actuate transport, and move VSLAM equipment 205 along the path planned by path planning engine 260. In some examples, path planning engine 260 may use Dijkstra's algorithm to plan a path. In some examples, path planning engine 260 may include stationary obstacle avoidance and / or moving obstacle avoidance when planning a path. In some examples, path planning engine 260 may include determining how to optimally move from a first pose to a second pose when planning a path. In some examples, path planning engine 260 may plan a path optimized to reach and observe every part of each room before continuing to move to other rooms. In some examples, path planning engine 260 may plan a path optimized to reach and observe every room in the environment as quickly as possible. In some examples, path planning engine 260 may plan a path that returns to a previously observed room to re-observe a particular feature in order to refine one or more map points corresponding to the feature in the local map and / or global map. In some examples, path planning engine 260 may plan a path that returns to a previously observed room to observe a portion of the previously observed room that lacks a map point in the local map and / or global map to determine whether any features can be observed in that portion of the room.

[0080] Although various elements of conceptual diagram 200 are shown separately from VSLAM device 205, it should be understood that VSLAM device 205 may include any combination of elements of conceptual diagram 200. For example, at least a subset of VSLAM system 270 may be part of VSLAM device 205. At least a subset of mapping system 275 may be part of VSLAM device 205. For example, VSLAM device 205 may include camera 210, feature extraction engine 220, feature tracking engine 225, relocalization engine 230, map optimization engine 235, local mapping engine 250, global mapping engine 255, map merging engine 257, path planning engine 260, motion actuator 255, or some combination thereof. In some examples, the VSLAM device 205 can capture an image 215, identify features in the image 215 via a feature extraction engine 220, track features via a feature tracking engine 225, optimize a map using a map optimization engine 235, perform relocalization using a relocalization engine 230, determine a map point 240, determine a device pose 245, generate a local map using a local mapping engine 250, update a local map using a local mapping engine 250, perform map merging using a map merging engine 257, generate a global map using a global mapping engine 255, update a global map using a global mapping engine 255, plan a path using a path planning engine 260, actuate movement using a motion actuator 265, or some combination thereof. In some examples, the feature extraction engine 220 and / or the feature tracking engine 225 are part of the front end of the VSLAM device 205. In some examples, the relocalization engine 230 and / or the map optimization engine 235 are part of the back end of the VSLAM device 205. Based on the image 215 and / or previous images, the VSLAM device 205 can identify features through feature extraction 220, track features through feature tracking 225, perform map optimization 235, perform relocalization 230, determine map points 240, determine poses 245, generate a local map 250, update the local map 250, perform map merging, generate a global map 255, update the global map 255, perform path planning 260, or some combination thereof.

[0081] In some examples, the map points 240, the device pose 245, the local map, the global map, the path planned by the path planning engine 260, or a combination thereof are stored at the VSLAM device 205. In some examples, the map points 240, the device pose 245, the local map, the global map, the path planned by the path planning engine 260, or a combination thereof are stored remotely from the VSLAM device 205 (e.g., on a remote server), but are accessible to the VSLAM device 205 via a network connection. The mapping system 275 can be part of the VSLAM device 205 and / or the VSLAM system 270. The mapping system 275 can be part of a device (e.g., a remote server) that is remote from the VSLAM device 205 but in communication with the VSLAM device 205.

[0082] In some cases, the VSLAM device 205 can communicate with a remote server. The remote server can include at least one subset of the VSLAM system 270. The remote server can include at least one subset of the mapping system 275. For example, the VSLAM device 205 can include a camera 210, a feature extraction engine 220, a feature tracking engine 225, a relocalization engine 230, a map optimization engine 235, a local mapping engine 250, a global mapping engine 255, a map merging engine 257, a path planning engine 260, a mobile actuator 255, or some combination thereof. In some examples, the VSLAM device 205 can capture an image 215 and send the image 215 to the remote server. Based on the image 215 and / or the previous image, the remote server may identify features via the feature extraction engine 220, track features via the feature tracking engine 225, optimize the map using the map optimization engine 235, perform relocalization using the relocalization engine 230, determine map points 240, determine the device pose 245, generate a local map using the local mapping engine 250, update the local map using the local mapping engine 250, perform map merging using the map merging engine 257, generate a global map using the global mapping engine 255, update the global map using the global mapping engine 255, plan a path using the path planning engine 260, or some combination thereof. The remote server may send the results of these processes back to the VSLAM device 205.

[0083] Figure 3 is a conceptual diagram 300 illustrating an example of a technique for performing visual simultaneous localization and mapping (VSLAM) using a visible light (VL) camera 310 and an infrared (IR) camera 315 of a VSLAM device 305 . Figure 3 The VSLAM device 305 can be any type of VSLAM device, including Figure 2The VSLAM device 305 may be any of the types of VSLAM devices discussed above. The VSLAM device 305 includes a VL camera 310 and an IR camera 315. In some cases, the IR camera 315 may be a near infrared (NIR) camera. The IR camera 315 may capture an IR image 325 by receiving and capturing light in the NIR spectrum. The NIR spectrum may be a subset of the IR spectrum that is close to and / or adjacent to the VL spectrum.

[0084] The VSLAM device 305 can use the VL camera 310 and / or the ambient light sensor to determine whether the environment in which the VSLAM device 305 is located is well-lit or poorly lit. For example, if the average brightness in the VL image 320 captured by the VL camera 310 exceeds a predetermined brightness threshold, the VSLAM device 305 can determine that the environment is well-lit. If the average brightness in the VL image 320 captured by the VL camera 310 is lower than a predetermined brightness threshold, the VSLAM device 305 can determine that the environment is poorly lit. Figure 3 As shown in the conceptual diagram 300, if the VSLAM device 305 determines that the environment is well-lit, the VSLAM device 305 can use both the VL camera 310 and the IR camera 315 for the VSLAM process. Figure 4 As shown in the conceptual diagram 400 , if the VSLAM device 305 determines that the environment is poorly lit, the VSLAM device 305 may disable the use of the VL camera 310 for the VSLAM process and may use only the IR camera 315 for the VSLAM process.

[0085] The VSLAM device 305 can move throughout the environment and arrive at multiple locations along the path through the environment. The path planning engine 395 can plan at least a subset of the path as discussed herein. The VSLAM device 305 can move itself along the path by using a mobile actuator 397 to actuate a motor or other means of transport. For example, if the VSLAM device 305 is a robot or a vehicle, the VSLAM device 305 can move itself along the path. The VSLAM device 305 can be moved along the path by the user. For example, if the VSLAM device 305 is a head-mounted display (HMD) (e.g., an XR headset) worn by the user, the VSLAM device 305 can be moved along the path by the user. In some cases, the environment can be a virtual environment or a partially virtual environment rendered at least in part by the VSLAM device 305. For example, if the VSLAM device 305 is an AR, VR, or XR headset, at least part of the environment can be virtual.

[0086] At each of the several locations along the path through the environment, the VL camera 310 of the VSLAM device 305 captures a VL image 320 of the environment, and the IR camera 315 of the VSLAM device 305 captures one or more IR images of the environment. In some cases, the VL image 320 and the IR image 325 are captured simultaneously. In some examples, the VL image 320 and the IR image 325 are captured within the same time window. The time window can be short, such as 1 second, 2 seconds, 3 seconds, less than 1 second, greater than 3 seconds, or a duration between any of the previously listed durations. In some examples, the time between the capture of the VL image 320 and the capture of the IR image 325 is less than a predetermined threshold time. The short predetermined time threshold can be a short duration, such as 1 second, 2 seconds, 3 seconds, less than 1 second, greater than 3 seconds, or a duration between any of the previously listed durations.

[0087] Before the VSLAM device 305 is used to perform the VSLAM process, the extrinsic calibration engine 385 of the VSLAM device 305 can perform extrinsic calibration 385 of the VL camera 310 and the IR camera 315. The extrinsic calibration engine 385 can determine a transformation by which the coordinates in the IR image 325 captured by the IR camera 315 can be converted into coordinates in the VL image 320 captured by the VL camera 310, and vice versa. In some examples, the transformation is a direct linear transform (DLT). In some examples, the transformation is a stereo matching transform. The extrinsic calibration engine 385 can determine a transformation by which the coordinates in the VL image 320 and / or the IR image 325 can be converted into three-dimensional map points. Figure 8 Conceptual diagram 800 illustrates an example of extrinsic calibration as performed by extrinsic calibration engine 385. Transformation 840 may be an example of a transformation determined by extrinsic calibration engine 385.

[0088] The VL camera 310 of the VSLAM device 305 captures a VL image 320. In some examples, the VL camera 310 of the VSLAM device 305 can capture the VL image 320 in grayscale. In some examples, the VL camera 310 of the VSLAM device 305 can capture the VL image 320 in color, and the VL image 320 can be converted from color to grayscale at the ISP 154, the host processor 152, or the image processor 150. The IR camera 315 of the VSLAM device 305 captures an IR image 325. In some cases, the IR image 325 can be a grayscale image. For example, the grayscale IR image 325 can represent objects that emit or reflect a large amount of IR light as white or light gray, and can represent objects that emit or reflect a small amount of IR light as black or dark gray, or vice versa. In some cases, the IR image 325 can be a color image. For example, the color IR image 325 may represent an object that emits or reflects a large amount of IR light with a color close to one end of the visible color spectrum (e.g., red), and may represent an object that emits or reflects a small amount of IR light with a color close to the other end of the visible color spectrum (e.g., blue or purple), or vice versa. In some examples, the IR camera 315 of the VSLAM device 305 may convert the IR image 325 from color to grayscale at the ISP 154, the host processor 152, or the image processor 150. In some cases, after the VL image 320 and / or the IR image 325 are captured, the VSLAM device 305 sends the VL image 320 and / or the IR image 325 to another device (such as a remote server).

[0089] The VL feature extraction engine 330 can perform feature extraction of the VL image 320. The VL feature extraction engine 330 can be part of the VSLAM device 305 and / or the remote server. The VL feature extraction engine 330 can identify one or more features as depicted in the VL image 320. Using the VL feature extraction engine 330 to identify features can include determining the two-dimensional (2D) coordinates of the features as depicted in the VL image 320. The 2D coordinates can include rows and columns in the pixel array of the VL image 320. The VL image 320 with many features clearly depicted can be maintained in a map database as a VL keyframe, and its depiction of features can be used to track those features in other VL images and / or IR images.

[0090] The IR feature extraction engine 335 can perform feature extraction on the IR image 325. The IR feature extraction engine 335 can be part of the VSLAM device 305 and / or the remote server. The IR feature extraction engine 335 can identify one or more features as depicted in the IR image 325. Identifying features using the IR feature extraction engine 335 can include determining the two-dimensional (2D) coordinates of the features as depicted in the IR image 325. The 2D coordinates can include rows and columns in the pixel array of the IR image 325. The IR image 325 with many features clearly depicted can be maintained in the map database as an IR keyframe, and its depiction of features can be used to track those features in other IR images and / or VL images. Features can include, for example, corners or other distinctive features of objects in the environment. The VL feature extraction engine 330 and the IR feature extraction engine 335 can also perform any of the processes discussed with respect to the feature extraction engine 220 of the conceptual diagram 200.

[0091] Either or both of the VL / IR feature association engine 365 and / or the stereo matching engine 367 can be part of the VSLAM device 305 and / or the remote server. The VL feature extraction engine 330 and the IR feature extraction engine 335 can identify one or more features depicted in both the VL image 320 and the IR image 325. The VL / IR feature association engine 365 identifies these features depicted in both the VL image 320 and the IR image 325 based on, for example, a transformation determined using an extrinsic calibration performed by the extrinsic calibration engine 385. The transformation can transform a 2D coordinate in the IR image 325 into a 2D coordinate in the VL image 320, and vice versa. The stereo matching engine 367 can further determine a set of three-dimensional (3D) map coordinates—map points—based on the 2D coordinates in the IR image 325 and the 2D coordinates in the VL image 320 captured from slightly different angles. Stereo constraints may be determined by the stereo matching engine 367 between the views of the VL camera 310 and the IR camera 315 to speed up feature search and matching performance for feature tracking and / or relocalization.

[0092] The VL feature tracking engine 340 may be part of the VSLAM device 305 and / or the remote server. The VL feature tracking engine 340 tracks features identified in the VL image 320 using the VL feature extraction engine 330 that were also depicted and detected in previously captured VL images captured by the VL camera 310 prior to capturing the VL image 320. In some cases, the VL feature tracking engine 330 may also track features identified in the VL image 320 that were also depicted and detected in previously captured IR images captured by the IR camera 315 prior to capturing the VL image 320. The IR feature tracking engine 345 may be part of the VSLAM device 305 and / or the remote server. The IR feature tracking engine 345 tracks features identified in the IR image 325 using the IR feature extraction engine 335 that were also depicted and detected in previously captured IR images captured by the IR camera 315 prior to capturing the IR image 325. In some cases, IR feature tracking engine 335 may also track features identified in IR image 325 that were also depicted and detected in a previously captured IR image captured by IR camera 315 prior to capturing VL image 320. Features determined to be depicted in both VL image 320 and IR image 325 using VL / IR feature correlation engine 365 and / or stereo matching engine 367 may be tracked using VL feature tracking engine 340, IR feature tracking engine 345, or both. VL feature tracking engine 340 and IR feature tracking engine 345 may also perform any of the processes discussed with respect to feature tracking engine 225 of conceptual diagram 200.

[0093] Each of VL map points 350 is a set of coordinates in a map that is determined using mapping system 390 based on features extracted using VL feature extraction engine 330, features tracked using VL feature tracking engine 340, and / or common features identified using VL / IR feature correlation engine 365 and / or stereo matching engine 367. Each of IR map points 355 is a set of coordinates in a map that is determined using mapping system 390 based on features extracted using IR feature extraction engine 335, features tracked using IR feature tracking engine 345, and / or common features identified using VL / IR feature correlation engine 365 and / or stereo matching engine 367. VL map points 350 and IR map points 355 can be three-dimensional (3D) map points, e.g., having three spatial dimensions. In some examples, each of VL map points 350 and / or IR map points 355 can have an X coordinate, a Y coordinate, and a Z coordinate. Each coordinate can represent a location along a different axis. Each axis can extend to a different spatial dimension that is perpendicular to the other two spatial dimensions. Determining VL map points 350 and IR map points 355 using mapping engine 390 may also include any of the processes discussed with respect to determining map points 240 of conceptual map 200. Mapping engine 390 may be part of VSLAM device 305 and / or part of a remote server.

[0094] The joint map optimization engine 360 ​​adds the VL map points 350 and the IR map points 355 to the map and / or optimizes the map. The joint map optimization engine 360 ​​may merge the VL map points 350 and the IR map points 355 corresponding to features determined to be depicted in both the VL image 320 and the IR image 325 (e.g., using the VL / IR feature association engine 365 and / or the stereo matching engine 367) into a single map point. The joint map optimization engine 360 ​​may also merge the VL map points 350 corresponding to features determined to be depicted in previous IR map points from one or more previous IR images and / or previous VL map points from one or more previous VL images into a single map point. The joint map optimization engine 360 ​​may also merge the IR map points 350 corresponding to features determined to be depicted in previous VL map points from one or more previous VL images and / or previous IR map points from one or more previous IR images into a single map point. As more VL images 320 and IR images 325 are captured depicting a particular feature, the joint map optimization engine 360 ​​can update the location of the map points corresponding to the feature in the map to make them more accurate (e.g., based on triangulation). For example, an updated set of coordinates for the map points for the feature can be generated by updating or revising a previous set of coordinates for the map points for the feature. The map can be a local map as discussed with respect to the local mapping engine 250. In some cases, the map is merged with a global map using the map merging engine 257 of the mapping system 290. The map can be a global map as discussed with respect to the global mapping engine 255. In some cases, as in Figure 11 As illustrated in and discussed with respect to conceptual diagram 1100 of FIG. , joint map optimization engine 360 ​​can simplify a map by replacing map point bundles with centroid map points. Joint map optimization engine 360 ​​can also perform any of the processes discussed with respect to map optimization engine 235 in conceptual diagram 200.

[0095] The mapping system 290 can generate a map of the environment based on a set of coordinates determined by the VSLAM device 305 for all map points (including VL map points 350 and IR map points 355) for all detected and / or tracked features. In some cases, when the mapping system 390 initially generates a map, the map can start as a map of a small portion of the environment. As more features are detected from more images, and as more features are converted into map points that the mapping system updates the map to include, the mapping system 390 can expand the map to map larger and larger portions of the environment. The map can be sparse or semi-dense. In some cases, the selection criteria used by the mapping system 390 for the map points corresponding to features may be demanding to support robust tracking of features using the VL feature tracking engine 340 and / or the IR feature tracking engine 345.

[0096] The device pose determination engine 370 can determine the pose of the VSLAM device 305. The device pose determination engine 370 can be part of the VSLAM device 305 and / or the remote server. The pose of the VSLAM device 305 can be determined based on: feature extraction by the VL feature extraction engine 330, feature extraction by the IR feature extraction engine 335, feature association by the VL / IR feature association engine 365, stereo matching by the stereo matching engine 367, feature tracking by the VL feature tracking engine 340, feature tracking by the IR feature tracking engine 345, determination of VL map points 350 by the mapping system 390, determination of IR map points 355 by the mapping system 390, map optimization by the joint map optimization engine 360, generation of a map by the mapping system 390, updating of the map by the mapping system 390, or some combination thereof. The pose of the device 305 may refer to the position of the VSLAM device 305, the pitch of the VSLAM device 305, the roll of the VSLAM device 305, the yaw of the VSLAM device 305, or some combination thereof. The pose of the VSLAM device 305 may refer to the pose of the VL camera 310 and may therefore include the position of the VL camera 310, the pitch of the VL camera 310, the roll of the VL camera 310, the yaw of the VL camera 310, or some combination thereof. The pose of the VSLAM device 305 may refer to the pose of the IR camera 315 and may therefore include the position of the IR camera 315, the pitch of the IR camera 315, the roll of the IR camera 315, the yaw of the IR camera 315, or some combination thereof. The device pose determination engine 370 may, in some cases, use the mapping system 390 to determine the pose of the VSLAM device 305 relative to the map. The device pose determination engine 370 may, in some cases, use the mapping system 390 to mark the pose of the VSLAM device 305 on the map. In some cases, the device pose determination engine 370 may determine and store a history of poses within a map or otherwise. The history of poses may represent the path of the VSLAM device 305. The device pose determination engine 370 may also perform any of the processes discussed with respect to determining the pose 245 of the VSLAM device 205 of the conceptual diagram 200. In some cases, the device pose determination engine 370 may determine the pose of the VSLAM device 305 by determining the pose of the body of the VSLAM device 305, determining the pose of the VL camera 310, determining the pose of the IR camera 315, or some combination thereof. One or more of the three poses may be separate outputs of the device pose determination engine 370. In some cases, the device pose determination engine 370 may merge or combine two or more of the three poses into a single output of the device pose determination engine 370, for example, by averaging the pose values ​​corresponding to two or more of the three poses.

[0097] The relocalization engine 375 can determine the position of the VSLAM device 305 within the map. For example, if the VL feature tracking engine 340 and / or the IR feature tracking engine 345 fail to identify any features from the features identified in the previous VL and / or IR images in the VL image 320 and / or the IR image 325, the relocalization 375 can relocate the VSLAM device 305 within the map. The relocalization engine 375 can determine the position of the VSLAM device 305 in the map by matching the features identified in the VL image 320 and / or the IR image 325 via the VL feature extraction engine 330 and / or the IR feature extraction engine 335 with the features corresponding to the map points in the map, with the features depicted in the VL keyframes, with the features depicted in the IR keyframes, or some combination thereof. The relocalization engine 375 can be part of the VSLAM device 305 and / or the remote server. The relocalization engine 375 can also perform any process discussed with respect to the relocalization engine 230 of the conceptual diagram 200.

[0098] The loop closure detection engine 380 can be part of the VSLAM device 305 and / or the remote server. The loop closure detection engine 380 can identify when the VSLAM device 305 has completed traveling along a path that is shaped like a closed loop or another closed shape without gaps or openings. For example, the loop closure detection engine 380 can identify that at least some of the features depicted and detected in the VL image 320 and / or in the IR image 325 match features identified earlier during the travel along the path that the VSLAM device 305 is traveling on. The loop closure detection engine 380 can detect loop closure based on a map as generated and updated by the mapping system 390 and based on a posture determined by the device posture determination engine 370. The loop closure detection performed by the loop closure detection engine 380 prevents the VL feature tracking engine 340 and / or the IR feature tracking engine 345 from incorrectly treating certain features depicted and detected in the VL image 320 and / or the IR image 325 as new features (when such features match features previously detected earlier in the same location and / or area during travel along the path along which the VSLAM device 305 has already traveled).

[0099] The VSLAM device 305 can include any type of transportation vehicle discussed with respect to the VSLAM device 205. A path planning engine 395 can plan the path that the VSLAM device 305 will travel using the transportation vehicle. The path planning engine 395 can plan the path based on a map, based on the posture of the VSLAM device 305, based on repositioning by the repositioning engine 375, and / or based on loop closure detection by the loop closure detection engine 380. The path planning engine 395 can be part of the VSLAM device 305 and / or a remote server. The path planning engine 395 can also perform any process discussed with respect to the path planning engine 260 of the conceptual diagram 200. A motion actuator 397 can be part of the VSLAM device 305 and can be activated by the VSLAM device 305 or a remote server to actuate the transportation vehicle to move the VSLAM device 305 along the path planned by the path planning engine 395. For example, the motion actuator 397 can include one or more actuators that actuate one or more motors of the VSLAM device 305. Movement actuator 397 may also perform any of the processes discussed with respect to movement actuator 265 of conceptual diagram 200 .

[0100] The VSLAM device 305 can use a map to perform various functions about the positioning depicted or defined in the map. For example, using a robot as an example of the VSLAM device 305 utilizing the technology described herein, the robot can actuate a motor to move the robot from a first position to a second position via a mobile actuator 397. The second position can be determined using a map of the environment, for example, to ensure that the robot avoids bumping into a wall or other obstacle that has been identified in the map for positioning, or to avoid inadvertently revisiting the positioning that the robot has visited. In some cases, the VSLAM device 305 can plan to revisit the positioning that the VSLAM device 305 has visited. For example, the VSLAM device 305 can revisit a previous positioning to verify a previous measurement, to correct the drift in the measurement after terminating a loop path or otherwise arriving at the end of a long path, to improve the accuracy of a map point that appears inaccurate (for example, an outlier) or has a low weight or confidence value, to detect more features in the area comprising a small amount and / or sparse map points, or some combination thereof. The VSLAM device 305 can actuate motors to move itself from an initial location to a target location to achieve a goal, such as food delivery, package delivery, package retrieval, capturing image data, mapping an environment, finding and / or reaching a charging station or power outlet, finding and / or reaching a base station, finding and / or reaching an exit from an environment, finding and / or reaching an entrance to an environment or another environment, or some combination thereof.

[0101] Once the VSLAM device 305 is successfully initialized, the VSLAM device 305 may repeat many of the processes shown in the conceptual diagram 300 at each new location of the VSLAM device 305. For example, the VSLAM device 305 may iteratively initiate the VL feature extraction engine 330, the IR feature extraction engine 335, the VL / IR feature association engine 365, the stereo matching engine 367, the VL feature tracking engine 340, the IR feature tracking engine 345, the mapping system 390, the joint map optimization system 360, the device pose determination engine 370, the relocalization engine 375, the loop closure detection engine 380, the path planning engine 395, the motion actuator 397, or some combination thereof at each new location of the VSLAM device 305. Features detected in each VL image 320 and / or each IR image 325 at each new location of the VSLAM device 305 may include features also observed in previously captured VL and / or IR images. The VSLAM device 305 may track the movement of these features from the previously captured image to the most recent image to determine the pose of the VSLAM device 305. The VSLAM device 305 may update the 3D map point coordinates corresponding to each of the features.

[0102] The mapping system 390 can assign a specific weight to each map point in the map. Different map points in the map can have different weights associated with them. Due to the reliability of the transformations calibrated using the extrinsic calibration engine 385, map points generated from VL / IR feature association 365 and stereo matching 367 can generally have good accuracy and therefore can have higher weights than map points seen using only the VL camera 310 or only the IR camera 315. Features depicted in a higher number of VL and / or IR images generally have improved accuracy compared to features depicted in a lower number of VL and / or IR images. Therefore, map points for features depicted in a higher number of VL and / or IR images can have greater weights in the map than map points depicted in a lower number of VL and / or IR images. The joint map optimization engine 360 ​​can include global optimization and / or local optimization algorithms that can correct the positioning of lower weight map points based on the positioning of higher weight map points, thereby improving the overall accuracy of the map. For example, if the long edge of a wall includes several high-weight map points that form a roughly straight line and low-weight map points that slightly disrupt the linearity of the line, the positioning of the low-weight map points can be adjusted to bring them into (or close to) the line so as to no longer disrupt the linearity of the line (or to a lesser extent disrupt the linearity of the line). In some cases, the joint map optimization engine 360 ​​can remove or move certain map points with low weights (e.g., if future observations appear to indicate that those map points are incorrectly positioned). The features identified in the VL image 320 and / or IR image 325 captured when the VSLAM device 305 arrives at the new position can also include new features that were not previously identified in any previously captured VL and / or IR image. The mapping system 390 can update the map to incorporate these new features, effectively expanding the map.

[0103] In some cases, the VSLAM device 305 can communicate with a remote server. The remote server can execute some of the processes discussed above as performed by the VSLAM device 305. For example, the VSLAM device 305 can capture a VL image 320 and / or an IR image 325 of the environment as described above, and send the VL image 320 and / or the IR image 325 to the remote server. The remote server can then identify the features depicted in the VL image 320 and the IR image 325 by a VL feature extraction engine 330 and an IR feature extraction engine 335. The remote server can include and can run a VL / IR feature association engine 365 and / or a stereo matching engine 367. The remote server can perform feature tracking using the VL feature tracking engine 340, perform feature tracking using the IR feature tracking engine 345, generate VL map points 350, generate IR map points 355, perform map optimization using the joint map optimization engine 360, generate a map using the mapping system 390, update the map using the mapping system 390, determine the device pose of the VSLAM device 305 using the device pose determination engine 370, perform relocalization using the relocalization engine 375, perform loop closure detection using the loop closure detection engine 380, plan a path using the path planning engine 395, issue a motion actuation signal to activate the motion actuator 397 and thereby trigger movement of the VSLAM device 305, or some combination thereof. The remote server can send the results of any of these processes back to the VSLAM device 305. By offloading computationally resource-intensive tasks to the remote server, the VSLAM device 305 can be smaller, can include a less powerful processor, can conserve battery power and therefore last longer between battery charges, perform tasks more quickly and efficiently, and be less resource-intensive.

[0104] If the environment is well-lit, both the VL image 320 of the environment captured by the VL camera 310 and the IR image 325 captured by the IR camera 315 are clear. When the environment is poorly lit, the VL image 320 of the environment captured by the VL camera 310 may not be clear, but the IR image 325 captured by the IR camera 315 may still remain clear. Therefore, the lighting level of the environment may affect the usefulness of the VL image 320 and the VL camera 310.

[0105] Figure 4 is a conceptual diagram 400 illustrating an example of a technique for performing visual simultaneous localization and mapping (VSLAM) using an infrared (IR) camera 315 of a VSLAM device. Figure 4 The VSLAM technology shown in the conceptual diagram 400 is similar to Figure 3 The VSLAM technology is shown in the conceptual diagram 300. However, in Figure 4In the VSLAM technique shown in conceptual diagram 400 of FIGURE 4, visible light camera 310 may be disabled 420 by lighting check engine 405 because lighting check engine 405 detects that the environment in which VSLAM device 305 is located is poorly lit. In some examples, disabling 420 visible light camera 310 means that visible light camera 310 is turned off and no longer captures VL images. In some examples, disabling 420 visible light camera 310 means that visible light camera 310 still captures VL images, such as for use by lighting check engine 405 to check whether lighting conditions in the environment have changed, but those VL images are not originally used for VSLAM.

[0106] In some examples, the lighting check engine 405 can use the VL camera 310 and / or the ambient light sensor 430 to determine whether the lighting level of the environment in which the VSLAM device 305 is located is well-lit or poorly lit. The lighting level can be referred to as a lighting condition. In order to check the lighting level of the environment, the VSLAM device 305 can capture a VL image and / or can use the ambient light sensor 430 to perform ambient light sensor measurements. If the average brightness in the VL image captured by the VL camera exceeds a predetermined brightness threshold 410, the VSLAM device 305 can determine that the environment is well-lit. If the average brightness in the VL image captured by the VL camera is lower than the predetermined brightness threshold 410, the VSLAM device 305 can determine that the environment is poorly lit. Average brightness can refer to the mean brightness in the VL image, the median brightness in the VL image, the mode brightness in the VL image, the median brightness in the VL image, or some combination thereof. In some cases, determining the average brightness can include reducing the VL image by one or more times and determining the average brightness of the reduced image. Similarly, if the brightness measured by the ambient light sensor exceeds a predetermined brightness threshold 410, the VSLAM device 305 can determine that the environment is well-lit. If the brightness measured by the ambient light sensor is below the predetermined brightness threshold 410, the VSLAM device 305 can determine that the environment is poorly lit. The predetermined brightness threshold 410 can be referred to as a predetermined lighting threshold, a predetermined lighting level, a predetermined minimum lighting level, a predetermined minimum lighting threshold, a predetermined brightness level, a predetermined minimum brightness level, a predetermined minimum brightness threshold, or some combination thereof.

[0107] Different areas of the environment can have different lighting levels (e.g., well-lit or poorly lit). The lighting check engine 405 can check the lighting level of the environment each time the VSLAM device 305 moves from one posture of the VSLAM device 305 to another posture. The lighting level in the environment can also change over time, for example, due to sunrise or sunset, blinds or curtains changing position, artificial light sources being turned on or off, dimmer switches of artificial light sources modifying how much light the artificial light sources output, artificial light sources being moved or pointed in different directions, or some combination thereof. The lighting check engine 405 can periodically check the lighting level of the environment based on specific time intervals. The lighting check engine 405 can check the lighting level of the environment each time the VSLAM device 305 uses the VL camera 310 to capture a VL image 320 and / or each time the VSLAM device 305 uses the IR camera 315 to capture an IR image 325. Since the last time the lighting check engine 405 checked the lighting level, the lighting check engine 405 can periodically check the lighting level of the environment each time the VSLAM device 305 captures a specific number of VL images and / or IR images.

[0108] Figure 4 The VSLAM technology shown in conceptual diagram 400 may include capture of IR images 325 by IR camera 315, feature detection using IR feature extraction engine 335, feature tracking using IR feature tracking engine 345, generation of IR map points 355 using mapping system 390, performance map optimization using joint map optimization engine 360, generation of a map using mapping system 390, updating of a map using mapping system 390, determination of a device pose of VSLAM device 305 using device pose determination engine 370, relocalization using relocalization engine 375, loop closure detection using loop closure detection engine 380, path planning using path planning engine 395, motion actuation using motion actuator 397, or some combination thereof. In some cases, Figure 4 The VSLAM technology shown in the conceptual diagram 400 can be used in Figure 3 The VSLAM technique shown in the conceptual diagram 300 of FIG is then executed. For example, an initially well-lit environment may become poorly lit over time, such as when the sun sets after a period of time and day turns to night.

[0109] exist Figure 4 Before the VSLAM technology shown in the conceptual diagram 400 is activated, the map may have been used by the mapping system 390 Figure 3 The VSLAM technology shown in the conceptual diagram 300 generates and / or updates. Figure 4 The VSLAM technology shown in the conceptual diagram 400 can be used to Figure 3The VSLAM technique shown in the conceptual diagram 300 partially or completely generates a map. Figure 4 The mapping system 390 shown in the conceptual diagram 400 can continue to update and refine the map even if the lighting of the environment suddenly changes. Figure 3 and Figure 4 The VSLAM device 305 of the VSLAM technology shown in the conceptual diagrams 300 and 400 can still work well, reliably and resiliently. Figure 3 The initial portion of the map generated by the VSLAM technology shown in conceptual diagram 300 can be reused instead of rebuilding the map from scratch to save computing resources and time.

[0110] The VSLAM device 305 may identify a set of 3D coordinates for an IR map point 355 for a new feature depicted in the IR image 325. For example, the VSLAM device 305 may triangulate the 3D coordinates for the IR map point 355 for the new feature based on the depiction of the new feature in the IR image 325 and the depiction of the new feature in other IR images and / or other VL images. The VSLAM device 305 may update an existing set of 3D coordinates for map points for previously identified features based on the depiction of the feature in the IR image 325.

[0111] IR camera 315 is used to Figure 3 and Figure 4 The transformations determined by the extrinsic calibration engine 385 during extrinsic calibration can be used during both VSLAM techniques. Figure 4 The VSLAM technology shown in the conceptual diagram 400 determines that new map points in the map and updates to existing map points are accurate and consistent with the use of Figure 3 The VSLAM technology shown in the conceptual diagram 300 determines new map points and updates to existing map points.

[0112] For a certain environmental area, if the ratio of new features (not previously identified in the map) to existing features (previously identified in the map) is low, then this means that for this environmental area, the map is basically complete. If the map is basically complete for a certain environmental area, the VSLAM device 305 can abandon updating the map for this environmental area and, instead, only focus on tracking its location, orientation, and posture within the map at least when the VSLAM device 305 is in this environmental area. As more maps are updated, the environmental area can include the entire environment.

[0113] In some cases, the VSLAM device 305 can communicate with a remote server. The remote server can perform Figure 4The VSLAM technology shown in the conceptual diagram 400 of FIG. Figure 3 Any of the processes performed by the remote server in the VSLAM technology shown in the conceptual diagram 300 of . In addition, the remote server may include an illumination check engine 405 that checks the illumination level of the environment. For example, the VSLAM device 305 can capture a VL image using a VL camera 310 and / or capture an ambient light measurement using an ambient light sensor 430. The VSLAM device 305 can send the VL image and / or the ambient light measurement to the remote server. The illumination check engine 405 of the remote server can determine whether the environment is well-lit or poorly lit based on the VL image and / or the ambient light measurement, for example by determining the average brightness of the VL image and comparing the average brightness of the VL image with a predetermined brightness threshold 410 and / or by comparing the brightness of the ambient light measurement with a predetermined brightness threshold 410.

[0114] Figure 4 The VSLAM technology shown in conceptual diagram 400 may be referred to as “night mode” VSLAM technology, “night mode” VSLAM technology, “dark mode” VSLAM technology, “low light” VSLAM technology, “poor lighting environment” VSLAM technology, “poor lighting” VSLAM technology, “dim lighting” VSLAM technology, “poor lighting” VSLAM technology, “dim lighting” VSLAM technology, “IR only” VSLAM technology, “IR mode” VSLAM technology, or some combination thereof. Figure 3 The VSLAM technology shown in conceptual diagram 300 may be referred to as “day mode” VSLAM technology, “day mode” VSLAM technology, “light mode” VSLAM technology, “bright mode” VSLAM technology, “high light” VSLAM technology, “well-lit environment” VSLAM technology, “good lighting” VSLAM technology, “bright lighting” VSLAM technology, “good lighting” VSLAM technology, “bright lighting” VSLAM technology, “VL-IR” VSLAM technology, “hybrid” VSLAM technology, “hybrid VL-IR” VSLAM technology, or some combination thereof.

[0115] Figure 5 is a conceptual diagram illustrating two images of the same environment captured under different lighting conditions. Specifically, first image 510 is an example of a VL image of an environment captured by VL camera 310 when the environment is well-lit. Various features, such as edges and corners between various walls and points on a star 540 in a painting hanging on the wall, are clearly visible and can be extracted by VL feature extraction engine 330.

[0116] On the other hand, second image 520 is an example of a VL image of an environment captured by VL camera 310 when the ambient lighting is poor. Due to the poor ambient lighting in second image 520, many features clearly visible in first image 510 are not visible at all or are not clearly visible in second image 520. For example, very dark area 530 in the lower right corner of second image 520 is almost pitch black, making it impossible to see any features at all in very dark area 530. For example, this very dark area 530 covers three of the five points of a star 540 in a painting hanging on the wall. The remainder of second image 520 is still slightly illuminated. However, due to the poor ambient lighting, there is a significant risk that many features will not be detected in second image 520. Due to the poor ambient lighting, there is also a significant risk that some features detected in second image 520 will not be recognized as matching previously detected features, even if they do match. For example, even if the VL feature extraction engine 330 detects two points of a star 540 that are still faintly visible in the second image 520, the VL feature tracking engine 340 may not be able to identify the two points of the star 540 as belonging to the same star 540 detected in one or more other images (such as the first image 510).

[0117] The first image 510 may also be an example of an IR image of an environment captured by the IR camera 315, while the second image 520 is an example of a VL image of the same environment captured by the VL camera 310. Even in poor lighting, the IR image may be clear.

[0118] Figure 6A is a perspective view 600 illustrating an unmanned ground vehicle (UGV) 610 performing visual simultaneous localization and mapping (VSLAM). Figure 6A The UGV 610 shown in the perspective view 600 may be implemented Figure 2 The VSLAM device 205 of the VSLAM technology shown in the conceptual diagram 200 executes Figure 3 The VSLAM device 305 and / or implementation of the VSLAM technology shown in the conceptual diagram 300 Figure 4 An example of a VSLAM device 305 of VSLAM technology is shown in conceptual diagram 400 of FIG. The UGV 610 includes a VL camera 310 located adjacent to an IR camera 315 along the front surface of the UGV 610. The UGV 610 includes a plurality of wheels 615 located along the bottom surface of the UGV 610. The wheels 615 can serve as a means of transportation for the UGV 610 and can be motorized using one or more motors. The motors, and therefore the wheels 615, can be actuated via the motion actuator 265 and / or the motion actuator 397 to move the UGV 610.

[0119] Figure 6Bis a perspective view 650 illustrating an unmanned aerial vehicle (UAV) 620 performing visual simultaneous localization and mapping (VSLAM). Figure 6B The UAV 620 shown in the perspective view 650 may be performing Figure 2 The VSLAM device 205 of the VSLAM technology shown in the conceptual diagram 200 executes Figure 3 The VSLAM device 305 and / or implementation of the VSLAM technology shown in the conceptual diagram 300 Figure 4 An example of a VSLAM device 305 of VSLAM technology is shown in conceptual diagram 400 of UAV 620. UAV 620 includes a VL camera 310 located along the front of the body of the UGV 610, adjacent to an IR camera 315. UAV 620 includes multiple propellers 625 located along the top of the UAV 620. Propellers 625 may be spaced apart from the body of the UAV 620 by one or more attachments to prevent propellers 625 from becoming lodged in circuitry on the body of the UAV 620 and / or to prevent propellers 625 from obstructing the field of view of the VL camera 310 and / or IR camera 315. Propellers 625 may serve as a means of transport for the UAV 620 and may be motorized using one or more motors. The motors, and therefore the propellers 625, may be actuated via motion actuator 265 and / or motion actuator 397 to move the UAV 620.

[0120] In some cases, the propeller 625 of the UAV 620 or another part of the VSLAM device 205 / 305 (e.g., an antenna) may partially block the field of view of the VL camera 310 and / or the IR camera 315. In some examples, this partial occlusion can be removed from any VL image and / or IR image in which this partial occlusion occurs before feature extraction is performed. In some examples, this partial occlusion is not removed from the VL image and / or IR image in which this partial occlusion occurs before feature extraction is performed, but the VSLAM algorithm is configured to ignore this partial occlusion for the purpose of feature extraction and, therefore, does not consider any portion of this partial occlusion as a feature of the environment.

[0121] Figure 7A 7 is a perspective diagram 700 illustrating a head-mounted display (HMD) 710 performing visual simultaneous localization and mapping (VSLAM). The HMD 710 may be an XR headset. Figure 7A The HMD 710 shown in the perspective view 700 may be an implementation Figure 2 The VSLAM device 205 of the VSLAM technology shown in the conceptual diagram 200 executes Figure 3 The VSLAM device 305 and / or implementation of the VSLAM technology shown in the conceptual diagram 300 Figure 44. An example of a VSLAM device 305 of VSLAM technology is shown in conceptual diagram 400 of FIG. An HMD 710 includes a VL camera 310 and an IR camera 315 along the front of the HMD 710. The HMD 710 may be, for example, an augmented reality (AR) headset, a virtual reality (VR) headset, a mixed reality (MR) headset, or some combination thereof.

[0122] Figure 7B The diagram shows Figure 7A 7. A perspective view 730 of a head-mounted display (HMD) worn by a user 720. The user 720 wears the HMD 710 on the user's 720 head, overlying the user's 720 eyes. The HMD 710 can capture VL images with the VL camera 310 and / or IR images with the IR camera 315. In some examples, the HMD 710 displays one or more images based on the VL images and / or IR images to the user's 720 eyes. For example, the HMD 710 can provide the user 720 with overlay information on their view of the environment. In some examples, the HMD 710 can generate two images for display to the user 720—one for the user's 720 left eye and one for the user's 720 right eye. Although HMD 710 is shown as having only one VL camera 310 and one IR camera 315, in some cases, HMD 710 (or any other VSLAM device 205 / 305) can have more than one VL camera 310 and / or more than one IR camera 315. For example, in some examples, HMD 710 can include a pair of cameras on either side of HMD 710, where each pair of cameras includes a VL camera 310 and an IR camera 315. Thus, stereo VL and IR views can be captured by the camera and / or displayed to the user. In some cases, other types of VSLAM devices 205 / 305 can also include more than one VL camera 310 and / or more than one IR camera 315 for stereo image capture.

[0123] HMD 710 does not include wheels 615, propellers 625, or other means of transportation of its own. Instead, HMD 710 relies on the movements of user 720 to move HMD 710 around the environment. Therefore, in some cases, HMD 710 may skip path planning using path planning engine 260 / 395 and / or movement actuation using motion actuators 265 / 397 when performing VSLAM techniques. In some cases, HMD 710 may still perform path planning using path planning engine 260 / 395 and may indicate directions to user 720 to follow a proposed path to guide the user along the proposed path planned using path planning engine 260 / 395. In some cases, such as when HMD 710 is a VR headset, the environment may be fully virtual or partially virtual. If the environment is at least partially virtual, movement through the virtual environment may also be virtual. For example, movement through the virtual environment may be controlled by one or more joysticks, buttons, video game controllers, mice, keyboards, trackpads, and / or other input devices. Mobile actuator 265 / 397 can include any such input device. Movement through the virtual environment can not need wheels 615, propellers 625, legs or any other form of transport. If the environment is a virtual environment, then HMD 710 can still use path planning engine 260 / 395 to perform path planning and / or perform mobile actuation 265 / 397. If the environment is a virtual environment, then HMD 710 can use mobile actuator 265 / 397 to perform mobile actuation by performing virtual movement in the virtual environment. Even if the environment is virtual, VSLAM technology can still be valuable because the virtual environment may be unmapped and / or generated by a device (such as a remote server or console associated with a video game or video game platform) other than the VSLAM device 205 / 305. In some cases, VSLAM can even be performed in a virtual environment by a VSLAM device 205 / 305 with its own physical transport system (allowing it to physically move around the physical environment). For example, VSLAM may be performed in a virtual environment to test whether the VSLAM device 205 / 305 is working correctly without wasting time or energy on movement and without wearing out the physical transport system of the VSLAM device 205 / 305 .

[0124] Figure 7C7 is a perspective view 740 of a front surface 755 of a mobile handheld device 750 illustrating the use of forward-facing cameras 310 and 315 to perform VSLAM according to some examples. The mobile handheld device 750 can be, for example, a cellular phone, a satellite phone, a portable game console, a music player, a health tracking device, a wearable device, a wireless communication device, a laptop computer, a mobile device, or a combination thereof. The front surface 755 of the mobile handheld device 750 includes a display screen 745. The front surface 755 of the mobile handheld device 750 includes a VL camera 310 and an IR camera 315. The VL camera 310 and the IR camera 315 are shown as being in a bezel around the display screen 745 on the front surface 755 of the mobile device 750. In some examples, the VL camera 310 and / or the IR camera 315 can be positioned in a notch or cutout cut out of the display screen 745 on the front surface 755 of the mobile device 750. In some examples, the VL camera 310 and / or the IR camera 315 can be under-display cameras located between the display screen 210 and the rest of the mobile handset 750, such that light passes through a portion of the display screen 210 before reaching the VL camera 310 and / or the IR camera 315. The VL camera 310 and the IR camera 315 of the perspective view 740 are forward-facing. The VL camera 310 and the IR camera 315 face a direction perpendicular to the planar surface of the front surface 755 of the mobile device 750.

[0125] Figure 7D 7 is a perspective view 760 illustrating a rear surface 765 of a mobile handheld device 750 using rear-facing cameras 310 and 315 to perform VSLAM according to some examples. The VL camera 310 and IR camera 315 of the perspective view 760 are rear-facing. The VL camera 310 and IR camera 315 face a direction perpendicular to the planar surface of the rear surface 765 of the mobile handheld device 750. Although the rear surface 765 of the mobile handheld device 750 does not have a display screen 745 as shown in the perspective view 760, in some examples, the rear surface 765 of the mobile handheld device 750 may have a display screen 745. If the rear surface 765 of the mobile handheld device 750 has a display screen 745, any positioning of the VL camera 310 and IR camera 315 relative to the display screen 745 as discussed with respect to the front surface 755 of the mobile handheld device 750 may be used.

[0126] Similar to HMD 710, mobile handheld device 750 does not include wheels 615, propellers 625, or other means of transportation of its own. Instead, mobile handheld device 750 relies on the movement of the user holding or wearing mobile handheld device 750 to move handheld device 750 around the environment. Therefore, in some cases, mobile handheld device 750 can skip path planning using path planning engine 260 / 395 and / or mobile actuation using mobile actuator 265 / 397 when performing VSLAM technology. In some cases, mobile handheld device 750 can still perform path planning using path planning engine 260 / 395 and can indicate to the user the direction to follow the suggested path to guide the user along the suggested path planned using path planning engine 260 / 395. In some cases, such as when mobile handheld device 750 is used for AR, VR, MR or XR, the environment can be completely virtual or partially virtual. In some cases, the mobile handheld device 750 can be inserted into the head mounted device so that the mobile handheld device 750 serves as the display of the HMD 710, wherein the display screen 745 of the mobile handheld device 750 serves as the display of the HMD 710. If the environment is at least partially virtual, movement through the virtual environment can also be virtual. For example, movement through the virtual environment can be controlled by one or more joysticks, buttons, video game controllers, mice, keyboards, trackpads, and / or other input devices coupled to the mobile handheld device 750 in a wired or wireless manner. The motion actuator 265 / 397 can include any such input device. Movement through the virtual environment can not require wheels 615, propellers 625, legs, or any other form of transportation. If the environment is a virtual environment, the mobile handheld device 750 can still use the path planning engine 260 / 395 to perform path planning and / or perform motion actuation 265 / 397. If the environment is a virtual environment, the mobile handset 750 may perform movement actuations using the movement actuator 265 / 397 by performing virtual movements within the virtual environment.

[0127] like Figure 3 、 Figure 4 、 Figure 6A 、 Figure 6B 、 Figure 7A 、 Figure 7B 、 Figure 7C and Figure 7D The VL camera 310 shown in FIG can be referred to as the first camera 310. Figure 3 、 Figure 4 、 Figure 6A 、 Figure 6B 、 Figure 7A 、 Figure 7B 、 Figure 7C and Figure 7DThe IR camera 315 shown in FIG3 may be referred to as the second camera 315. The first camera 310 may be responsive to a first spectrum, while the second camera 315 is responsive to a second spectrum. Although the first camera 310 is labeled as a VL camera throughout the figures and the description herein, it should be understood that the VL spectrum is merely one example of a first spectrum to which the first camera 310 is responsive. Although the second camera 315 is labeled as an IR camera throughout the figures and the description herein, it should be understood that the IR spectrum is merely one example of a second spectrum to which the second camera 315 is responsive. The first spectrum may include at least one of: at least a portion of the VL spectrum, at least a portion of the IR spectrum, at least a portion of the ultraviolet (UV) spectrum, at least a portion of the microwave spectrum, at least a portion of the radio spectrum, at least a portion of the X-ray spectrum, at least a portion of the gamma spectrum, at least a portion of the electromagnetic (EM) spectrum, or a combination thereof. The second spectrum can include at least one of: at least a portion of a VL spectrum, at least a portion of an IR spectrum, at least a portion of an ultraviolet (UV) spectrum, at least a portion of a microwave spectrum, at least a portion of a radio spectrum, at least a portion of an X-ray spectrum, at least a portion of a gamma spectrum, at least a portion of an electromagnetic (EM) spectrum, or a combination thereof. The first spectrum can be different from the second spectrum. In some examples, the first spectrum and the second spectrum can, in some cases, not have any overlapping portions. In some examples, the first spectrum and the second spectrum can at least partially overlap.

[0128] Figure 8 8 is a conceptual diagram illustrating extrinsic calibration of a visible light (VL) camera 310 and an infrared (IR) camera 315. When the VSLAM device is located in a calibration environment, the extrinsic calibration engine 385 performs extrinsic calibration of the VL camera 310 and the IR camera 315. The calibration environment includes a pattern surface 830 having a known pattern, which has one or more features at known locations. In some examples, such as Figure 8 As shown in conceptual diagram 800 of , pattern surface 830 can have a checkerboard pattern. A checkerboard surface can be useful because it has regularly spaced features, such as the corners of each square on the checkerboard surface. The checkerboard pattern can be referred to as a chessboard pattern. In some examples, pattern surface 830 can have another pattern, such as a crosshair, a quick response (QR) code, an ArUco mark, a pattern of one or more alphanumeric characters, or some combination thereof.

[0129] VL camera 310 captures a VL image 810 depicting a pattern surface 830. IR camera 315 captures an IR image 820 depicting the pattern surface 830. Features of the pattern surface 830 (such as the square corners of the checkerboard pattern) are detected within the depictions of the pattern surface 830 in VL image 810 and IR image 820. A transformation 840 is determined that converts the 2D pixel coordinates (e.g., row and column) of each feature as depicted in IR image 820 to the 2D pixel coordinates (e.g., row and column) of the same feature as depicted in VL image 810. Transformation 840 can be determined based on the known actual location of the same feature in actual pattern surface 830 and / or based on the known relative location of the feature with respect to other features in pattern surface 830. In some cases, the transformation 840 may also be used to map the 2D pixel coordinates (e.g., rows and columns) of each feature as depicted in the IR image 820 and / or the VL image 810 to a three-dimensional (3D) coordinate set of map points in the environment, having three coordinates corresponding to the three spatial dimensions.

[0130] In some examples, the extrinsic calibration engine 385 constructs a world coordinate system for extrinsic calibration at the upper left corner of the checkerboard pattern. The transformation 840 can be a direct linear transformation (DLT). Based on the 3D-2D correspondence between the known 3D positioning of features on the pattern surface 830 and the 2D pixel coordinates (e.g., rows and columns) in the VL image 810 and the IR image 820, certain parameters can be identified. For clarity, parameters or variables representing matrices are referenced herein within square brackets ("[" and "["). The brackets themselves should be understood not to indicate an equivalence class or any other mathematical concept. The camera intrinsic parameters [K of the VL camera 310 are shown in Figure 10). VL ] and the camera internal parameters of IR camera IR 315 [K IR ] may be determined based on the properties of the VL camera 310 and the IR camera 315 and / or based on the 3D-2D correspondence. The camera pose of the VL camera 310 during the capture of the VL image 810 and the camera pose of the IR camera 315 during the capture of the IR image 820 may be determined based on the 3D-2D correspondence. The variable p VL The variable p may represent a set of 2D coordinates of points in the VL image 810. IR A set of 2D coordinates of corresponding points in the IR image 820 may be represented.

[0131] Determining the transformation 840 may include using the equation To solve the rotation matrix R and / or translation t. IR and p VL can be homogeneous coordinates. The values ​​for [R] and t can be determined so that the transformation 840 successfully transforms the point p in the IR image 820 IRis consistently converted to point p in the VL image 810 VL (For example, by solving this equation multiple times for different features of the pattern surface 830, using singular value decomposition (SVD), and / or using iterative optimization.) Because the extrinsic calibration engine 385 can perform extrinsic calibration before the VSLAM device 205 / 305 is used to perform VSLAM, time and computing resources are generally not an issue in determining the transformation 840. In some cases, the transformation 840 can be similarly used to transform the point p in the VL image 810 to the point p in the VL image 810. VL Transformed into point p in IR image 820 IR .

[0132] Figure 9 9 is a conceptual diagram 900 illustrating a transformation 840 between the coordinates of a feature detected in an infrared (IR) image 920 captured by an infrared (IR) camera 315 and the coordinates of the same feature detected in a visible light (VL) image 910 captured by a visible light (VL) camera 310. The conceptual diagram illustrates several features in an environment observed by the VL camera 310 and the IR camera 315. Three gray-style shaded circles represent commonly observed features 930 depicted in both the VL image 910 and the IR image 920. Commonly observed features 930 may be depicted, observed, and / or detected in both the VL image 910 and the IR image 920 by the feature extraction engine 220 / 330 / 335 during feature extraction. Three white-shaded circles represent VL features 940 depicted, observed, and / or detected in the VL image 910 but not in the IR image 920. VL features 940 may be detected in the VL image 910 during VL feature extraction 330. The three black shaded circles represent IR features 945 that are depicted, observed, and / or detected in the IR image 920 but not in the VL image 910. The IR features 945 may be detected in the IR image 920 during the IR feature extraction 335.

[0133] The 3D coordinate set for the map point for the commonly observed feature in commonly observed feature 930 can be determined based on the depictions of the commonly observed feature in VL image 910 and IR image 920. For example, the 3D coordinate set for the map point for the commonly observed feature can be triangulated using a midpoint algorithm. Point O represents IR camera 315. Point O' represents VL camera 310. Point U along the arrow from point O to the commonly observed feature in commonly observed feature 930 represents the depiction of the commonly observed feature in IR image 920. Point U' along the arrow from point O' to the commonly observed feature in commonly observed feature 930 represents the depiction of the commonly observed feature in VL image 910.

[0134] The set of 3D coordinates for the map points for the VL features in VL feature 940 can be determined based on the depiction of the VL features in VL image 910 and one or more other depictions of the VL features in one or more other VL images and / or in one or more IR images. For example, the set of 3D coordinates for the map points for the VL features can be triangulated using a midpoint algorithm. Point W' along the arrow from point O' to the VL feature in VL feature 940 represents the depiction of the VL feature in VL image 910.

[0135] The set of 3D coordinates for the map point for the IR feature in IR feature 945 can be determined based on the depiction of the IR feature in IR image 920 and one or more other depictions of the IR feature in one or more other IR images and / or in one or more VL images. For example, the set of 3D coordinates for the map point for the IR feature can be triangulated using a midpoint algorithm. Point W along the arrow from point O to the IR feature in IR feature 945 represents the depiction of the IR feature in IR image 920.

[0136] In some examples, transformation 840 may transform the 2D location of features detected in IR image 920 into a 2D location in the field of view of VL camera 310. The 2D location in the field of view of VL camera 310 may be transformed into a set of 3D coordinates of map points used in the map based on the pose of VL camera 310. In some examples, the pose of VL camera 310 associated with the first VL keyframe may be initialized by mapping system 390 as the origin of the world coordinate system of the map. Using the VSLAM techniques illustrated in at least one of conceptual diagrams 200, 300, and / or 400, a second VL keyframe captured by VL camera 310 after the first VL keyframe is registered in the world coordinate system of the map. The IR keyframe may be captured by IR camera 315 simultaneously with or within the same time window as the second VL keyframe. The time window may last for a predetermined duration, such as one or more picoseconds, one or more nanoseconds, one or more milliseconds, or one or more seconds. The IR keyframes are used for triangulation to determine a set of 3D coordinates for map points (or portions of map points) corresponding to the commonly observed feature 930 .

[0137] Figure 10A10 is a conceptual diagram 1000 illustrating feature associations between coordinates of features detected in an infrared (IR) image 1020 captured by an infrared (IR) camera 315 and coordinates of the same features detected in a visible light (VL) image 1010 captured by a visible light (VL) camera 310. A gray-style shaded circle labeled P represents a commonly observed feature P. Points u along the arrow from point O to the commonly observed feature P represent a depiction of the commonly observed feature P in the IR image 1020. Points u′ along the arrow from point O′ to the commonly observed feature P represent a depiction of the commonly observed feature P in the VL image 1010.

[0138] Transformation 840 may be applied to point u in IR image 1020, which may produce point u as shown in VL image 1010. In some examples, VL / IR feature association 365 may identify that points u and u′ represent a commonly observed feature P by searching for a match for point u′ in IR image 1020 within region 1030 around the location of point u′ in VL image 1010 based on the point being transformed from IR image 1020 to VL image 1010 using transformation 840, and determining that points within region 1030 Matches point u'. In some examples, VL / IR feature association 365 may identify that points u and u' represent a common observed feature P by: Search for points within the area 1030 surrounding the location and determine the matching of point uu and point match.

[0139] Figure 10B is a conceptual diagram 1050 illustrating an example descriptor pattern for a feature. Points u' and Whether the match can be based on the point u' and The descriptor pattern is determined by whether the associated descriptor patterns match within a predetermined maximum percentage variation of each other. The descriptor pattern includes a feature pixel 1060, which is a point representing a feature. The descriptor pattern includes a number of pixels surrounding the feature pixel 1060. The example descriptor pattern shown in conceptual diagram 1050 takes the form of a 5-pixel by 5-pixel square of pixels, with the feature pixel 1060 at the center of the descriptor pattern. Different descriptor pattern shapes and / or sizes can be used. In some examples, the descriptor pattern can be a 3-pixel by 3-pixel square of pixels, with the feature pixel 1060 at the center. In some examples, the descriptor pattern can be a 7-pixel by 7-pixel square of pixels, or a 9-pixel by 9-pixel square of pixels, with the feature pixel 1060 at the center. In some examples, the descriptor pattern can be a circle, an ellipse, a rectangular, or another shape of pixels, with the feature pixel 1060 at the center.

[0140] The descriptor pattern includes 5 black arrows, each of which passes through the feature pixel 1060. Each of the black arrows passes from one end of the descriptor pattern to the opposite end of the descriptor pattern. The black arrows represent the intensity gradient around the feature pixel 1060 and can be derived along the direction of the arrow. The intensity gradient can correspond to the difference in luminosity of the pixel along each arrow. If the VL image is in color, then each intensity gradient can correspond to the difference in color intensity of the pixel along each arrow on one of the colors in the color set (e.g., red, green, blue). The intensity gradient can be normalized to fall within the range between 0 and 1. The intensity gradients can be sorted according to the direction their corresponding arrows face and can be connected into a histogram distribution. In some examples, the histogram distribution can be stored in a binary string of 256 bits in length.

[0141] As mentioned above, point u' and Whether the match can be based on the point u' and In some examples, a binary string storing a histogram distribution corresponding to the descriptor pattern for point u' may be combined with a binary string storing a histogram distribution corresponding to the descriptor pattern for point u'. In some examples, if the binary string corresponding to point u' is compared with the binary string corresponding to point The difference between the binary strings of is less than the maximum percentage change, then the point u' and ' is determined to match and therefore depict the same feature P. In some examples, the maximum percentage change can be 5%, 10%, 15%, 20%, 25%, less than 5%, greater than 25%, or a percentage value between any two of the previously listed percentage values. If the binary string corresponding to point u' is the same as the binary string corresponding to point If the binary string of the point u' differs by more than the maximum percentage change, then is determined to be a mismatch, and therefore does not depict the same feature P.

[0142] Figure 11 1 is a conceptual diagram 1100 illustrating an example of joint map optimization. Conceptual diagram 1100 illustrates a bundle of points 1110. Bundle 1110 includes points shaded in gray, representing commonly observed features observed by both VL camera 310 and IR camera 315 at the same time or at different times, as determined using VL / IR feature associations 365. Bundle 1110 includes points shaded in white, representing features observed by VL camera 310 but not by IR camera 315. Bundle 1110 includes points shaded in black, representing features observed by IR camera 315 but not by VL camera 310.

[0143] Bundle adjustment (BA) is an example technique for performing joint map optimization 360. A cost function (such as the reprojection error of a 2D point to a 3D map point) can be used for BA as the optimization target. The joint map optimization engine 360 ​​can use BA to modify the keyframe poses and / or map point information based on the residual gradient to minimize the reprojection error. In some examples, the VL map points 350 and the IR map points 355 can be optimized separately. However, map optimization using BA can be computationally intensive. Therefore, the VL map points 350 and the IR map points 355 can be optimized together by the joint map optimization engine 360 ​​instead of optimizing them separately. In some examples, the reprojection error terms generated from the IR, RGB channels, or both will be placed in the target loss function for BA.

[0144] In some cases, the local search window represented by the bundle 1110 can be determined based on map points corresponding to commonly observed features in the bundle 1110 that are colored in gray. Other map points (such as map points colored in white that are only observed by the VL camera 310 or map points colored in black that are only observed by the IR camera 315) can be ignored or discarded in the loss function, or can be given less weight than the commonly observed features. After BA optimization, if the map points in the bundle are distributed very close to each other, then the centroid 1120 of these map points in the bundle 1110 can be calculated. In some examples, the location of the centroid 1120 is calculated to be at the center of the bundle 1110. In some examples, the location of the centroid 1120 is calculated based on an average of the locations of the points in the bundle 1110. In some examples, the location of the centroid 1120 is calculated based on a weighted average of the locations of the points in the bundle 1110, where some points (e.g., commonly observed points) are weighted more than other points (e.g., non-commonly observed points). The centroid 1120 is at Figure 11 1100. The centroid 1120 can then be used as a map point for the map by the mapping system 390, and the other points in the bundle can be discarded from the map by the mapping system 390. The use of the centroid 1120 supports consistent spatial optimization and avoids redundant computations for points with similar descriptors or points that are narrowly distributed (e.g., points that are distributed within a predetermined range of each other).

[0145] Figure 121 is a conceptual diagram 1200 illustrating feature tracking 1250 / 1255 and stereo matching 1240 / 1245. Conceptual diagram 1200 illustrates VL image frame t 1220 captured by VL camera 310. Conceptual diagram 1200 illustrates VL image frame t+1 1230 captured by VL camera 310 after capturing VL image frame t 1220. One or more features are depicted in VL image frame t 1220 and VL image frame t+1 1230, and feature tracking 1250 tracks changes in the positioning of the one or more features from VL image frame t 1220 to VL image frame t+1 1230.

[0146] Conceptual diagram 1200 illustrates IR image frame t 1225 captured by IR camera 315. Conceptual diagram 1200 illustrates IR image frame t+1 1235 captured by IR camera 315 subsequent to capturing IR image frame t 1225. One or more features are depicted in IR image frame t 1225 and IR image frame t+1 1235, and feature tracking 1255 tracks changes in the positioning of the one or more features from IR image frame t 1225 to IR image frame t+1 1235.

[0147] The VL image frame t 1220 may be captured simultaneously with the IR image frame t 1225. The VL image frame t 1220 may be captured within the same time window as the IR image frame t 1225. Stereo matching 1240 matches one or more features depicted in the VL image frame t 1220 with matching features depicted in the IR image frame t 1225. Stereo matching 1240 identifies features that are commonly observed in the VL image frame t 1220 and the IR image frame t 1225. Stereo matching 1240 may use a method such as Figure 10A and Figure 10B 1050 and 1050. Transformation 840 may be used in either or both directions to transform points corresponding to features (their representations) in VL image frame t 1220 to corresponding representations in IR image frame t 1225, and vice versa.

[0148] The VL image frame t+1 1230 may be captured simultaneously with the IR image frame t+1 1235. The VL image frame t+1 1230 may be captured within the same time window as the IR image frame t+1 1235. Stereo matching 1245 matches one or more features depicted in the VL image frame t+1 1230 with matching features depicted in the IR image frame t+1 1235. Stereo matching 1240 may use a method such as Figure 10A and Figure 10B1050 of conceptual diagrams 1000 and 1050. Transformation 840 may be used in either or both directions to transform points corresponding to features (their representations) in VL image frame t+1 1230 to corresponding representations in IR image frame t+1 1235, or vice versa.

[0149] The correspondence between the VL map point 350 and the IR map point 355 may be established during stereo matching 1240 / 1245. Similarly, the correspondence between the VL keyframe and the IR keyframe may be established during stereo matching 1240 / 1245.

[0150] Figure 13A is a conceptual diagram 1300 illustrating stereo matching between the coordinates of a feature detected in an IR image 1320 captured by an infrared (IR) camera 315 and the coordinates of the same feature detected in a visible light (VL) image 1310 captured by a VL camera 310. 3D points P' and P" represent observed sample locations of the same feature. A more accurate location P of the feature is later determined by Figure 13B is determined by the triangulation shown in the conceptual diagram 1350.

[0151] 3D point P" represents a feature observed in the VL camera frame O' 1310. Because the depth scale of the feature is unknown, P" is uniformly sampled along the line O'U' in front of the VL image frame 1310. Points in the IR image 1320 Denotes a point U′, C transformed into the IR channel via transformation 840 ([R] and t) VL is the 3D VL camera positioning output by VSLAM, [T VL ] is the transformation matrix derived from the VSLAM output, including both orientation and positioning. [K IR ] is the intrinsic parameter matrix for the IR camera. Many P" samples are projected onto the IR image frame 1320, and then these projected samples A search within a window of is performed to find corresponding feature observations with similar descriptors in the IR image frame 1320. The best sample is then and its corresponding 3D point P'' are selected. Thus, from point P'' in the VL camera frame 1310 to point P'' in the IR image 1320 The final transformation can be written as follows:

[0152]

[0153] 3D point P' represents a feature observed in the IR camera frame 1320. Point P' in the VL image 1310 Denotes a point U, C in the VL channel via the inverse transformation 840 ([R] and t) IR is the 3D IR camera positioning output by VSLAM, [T IR ] is the transformation matrix derived from the VSLAM output, including both orientation and positioning. [K VL ] is the intrinsic parameter matrix for the VL camera. Many P' samples are projected onto the VL image frame 1310, and then these projected samples A search within a window of is performed to find corresponding feature observations with similar descriptors in the VL image frame 1310. The best sample is then and its corresponding 3D sample point P' is selected. Therefore, from point P' in the IR camera frame 1320 to point P' in the VL image 1310 The final transformation can be written as follows:

[0154]

[0155] The set of 3D coordinates for the location point P' for the feature is based on the first line drawn from point O to point U and from point O' to point The 3D coordinate set for the location point P" for the feature is based on the first line drawn from point O' to point U' and the line drawn from point O to point The intersection point is determined by drawing a second line.

[0156] Figure 13B is a conceptual diagram 1350 illustrating triangulation between the coordinates of a feature detected in an infrared (IR) image captured by an IR camera and the coordinates of the same feature detected in a visible light (VL) image captured by a VL camera. Figure 13A In the stereo matching transformation shown in conceptual diagram 1300 , a position point P′ for a feature is determined. Based on the stereo matching transformation, a position point P″ for the same feature is determined. In the triangulation operation shown in conceptual diagram 1350 , a line segment is drawn from point P′ to point P″. In conceptual diagram 1350 , the line segment is represented by a dashed line. A more accurate position P for the feature is determined as the midpoint along the line segment.

[0157] Figure 14A 14 is a conceptual diagram 1400 illustrating monocular matching between the coordinates of a feature detected by a camera in image frame t 1410 and the coordinates of the same feature detected by the camera in a subsequent image frame t+1 1420. The camera may be a VL camera 310 or an IR camera 315. Image frame t 1410 is captured by the camera when the camera is in a pose C' shown by coordinates O'. Image frame t+1 1420 is captured by the camera when the camera is in a pose C shown by coordinates O.

[0158] Point P" represents a feature observed by the camera during the capture of image frame t 1410. Point U' in image frame t 1410 represents the feature observation of point P" within image frame t 1410. Point U' in image frame t+1 1420 Denotes point U' transformed to image frame t+1 1420 via transformation 1440 (including [R] and t). Transformation 1440 can be similar to transformation 840. C is the camera position of image frame t 1410, [T] is the transformation matrix generated from motion prediction, including both orientation and position. [K] is the intrinsic parameter matrix of the corresponding camera. A number of P" samples are projected onto image frame t+1 1420, and then these projected samples A search within a window of is performed to find a corresponding feature observation with the same descriptor in image frame t+1 1420. The best sample is then and its corresponding 3D point P'' are selected. Therefore, from point P'' in camera frame t 1410 to point P'' in image frame t+1 1420 The final transformation 1440 can be written as follows:

[0159]

[0160] Unlike transformation 840 used for stereo matching, R and t for transformation 1440 can be determined based on predictions through a constant velocity model v×Δt based on the velocity of the camera between the previous image frame t-1 (not shown) and the capture of image frame t1410.

[0161] Figure 14B is a conceptual diagram 1450 illustrating triangulation between coordinates of a feature detected by a camera in an image frame and coordinates of the same feature detected by the camera in a subsequent image frame.

[0162] The set of 3D coordinates for the location point P' for the feature is based on the first line drawn from point O to point U and from point O' to point The 3D coordinate set for the location point P" for the feature is based on the first line drawn from point O' to point U' and the line drawn from point O to point In the triangulation operation shown in conceptual diagram 1450 , a line segment is drawn from point P′ to point P″. In conceptual diagram 1450 , the line segment is represented by a dashed line. A more accurate position P for the feature is determined as the midpoint along the line segment.

[0163] Figure 15 1500 is a conceptual diagram illustrating fast relocalization based on keyframes. As shown in the conceptual diagram 1500, relocalization using keyframes speeds up relocalization and improves night mode ( Figure 4 The success rate of the VSLAM technology shown in the conceptual diagram 400 is shown in FIG. 1500. The relocalization using key frames maintains the daytime mode ( Figure 3 speed and high success rate in VSLAM technology shown in the conceptual diagram 300).

[0164] The circles shaded with a gray pattern in conceptual diagram 1500 represent 3D map points for features observed by IR camera 315 during night mode. The black shaded circles in conceptual diagram 1500 represent 3D map points for features observed by VL camera 310, IR camera 315, or both during day mode. To help overcome feature sparsity in night mode, unobserved map points within the range of the map point currently observed by IR camera 315 can also be retrieved to assist in relocalization.

[0165] In the relocalization algorithm shown in conceptual diagram 1500, the current IR image captured by the IR camera 315 is compared with other IR camera keyframes to find matching candidates with the most common descriptors in the keyframe images, as indicated by bag-of-words scores (BoWs) above a predetermined threshold. For example, all map points belonging to the current IR camera keyframe 1510 are matched against a submap in conceptual diagram 1500, the submap consisting of map points of a candidate keyframe (not shown) and map points of neighboring keyframes (not shown) of the candidate keyframe. These submaps include points that are observed and unobserved in the keyframe view. The map points of each subsequent consecutive IR camera keyframe 1515 and the nth IR camera keyframe 1520 are matched against this submap map points in conceptual diagram 1500. The submap map points can include both map points of the candidate keyframe and map points of neighboring keyframes of the candidate keyframe. In this way, the relocalization algorithm can verify a candidate keyframe by consistent matches against the submap between multiple consecutive IR keyframes. Here, the search algorithm retrieves a specific range area (such as Figure 15 Finally, the best candidate keyframe is selected when its submap can be consistently matched with map points of consecutive IR keyframes. This matching can be performed at any time. Since the matching process uses more 3D map point information, the relocalization can be more accurate than without this additional map point information (IR camera keyframe, a later IR camera keyframe after the fifth IR camera keyframe, or another IR camera keyframe).

[0166] Figure 1616 is a conceptual diagram 1600 illustrating fast relocalization based on a keyframe (e.g., IR camera keyframe m 1610) and a centroid 1620 (also referred to as a centroid point). As in conceptual diagram 1500, circles 1650 shaded in gray in conceptual diagram 1600 represent 3D map points for features observed by IR camera 315 during night mode in IR camera keyframe m 1610. Black shaded circles in conceptual diagram 1600 represent 3D map points for features observed by VL camera 310, IR camera 315, or both during day mode.

[0167] The white-colored star represents a centroid 1620 generated based on the four black points in inner circle 1625 of conceptual map 1600. Centroid 1620 may be generated based on the four black points in inner circle 1625 because the four black points in inner circle 1625 are not very close to each other in 3D space and these map points all have similar descriptors.

[0168] The relocalization algorithm can compare the feature corresponding to circle 1650 with other features in outer circle 1630. Since centroid 1620 has already been generated, the relocalization algorithm can discard the four black dots in inner circle 1625 for relocalization purposes, as considering all four black dots in inner circle 1625 would be duplicates. In some examples, the relocalization algorithm can compare the feature corresponding to circle 1650 with centroid 1620 rather than any of the four black dots in inner circle 1625. In some examples, the relocalization algorithm can compare the feature corresponding to circle 1650 with only one of the four black dots in inner circle 1625 rather than all four black dots in inner circle 1625. In some examples, the relocalization algorithm can compare the feature corresponding to circle 1650 with neither centroid 1620 nor any of the four black dots in inner circle 1625. In any of these examples, the relocalization algorithm uses fewer computational resources.

[0169] Figure 15 Concept map 1500 and Figure 16 The fast relocation technique shown in the conceptual diagram 1600 may be Figure 2 The relocalization 230 of the VSLAM technology shown in the conceptual diagram 200, Figure 3 The VSLAM technology shown in the conceptual diagram 300 is used to relocalize 375 and / or Figure 4 An example of relocalization 375 using VSLAM technology is shown in conceptual diagram 400 .

[0170] Figure 8 、 Figure 9 、 Figure 10A 、 Figure 12 、 Figure 13A and Figure 13BThe various VL images (810, 910, 1010, 1220, 1230, 1310) in FIG. 3 may each be referred to as a first image or a first type of image. Each of the first type of images may be an image captured by the first camera 310. Figure 8 、 Figure 9 、 Figure 10A 、 Figure 12 、 Figure 13A 、 Figure 13B 、 Figure 15 and Figure 16 The various IR images (820, 920, 1020, 1225, 1235, 1320, 1510, 1515, 1520, 1610) in FIG. 1 may each be referred to as a second image or a second type of image. Each of the second type of images may be an image captured by second camera 315. First camera 310 may be responsive to a first spectrum, while second camera 315 may be responsive to a second spectrum. While first camera 310 is sometimes referred to herein as a VL camera 310, it should be understood that the VL spectrum is merely one example of a first spectrum to which first camera 310 is responsive. While second camera 315 is sometimes referred to herein as an IR camera 315, it should be understood that the IR spectrum is merely one example of a second spectrum to which second camera 315 is responsive. The first spectrum may include at least one of: at least a portion of a VL spectrum, at least a portion of an IR spectrum, at least a portion of an ultraviolet (UV) spectrum, at least a portion of a microwave spectrum, at least a portion of a radio spectrum, at least a portion of an X-ray spectrum, at least a portion of a gamma spectrum, at least a portion of an electromagnetic (EM) spectrum, or a combination thereof. The second spectrum can include at least one of: at least a portion of a VL spectrum, at least a portion of an IR spectrum, at least a portion of an ultraviolet (UV) spectrum, at least a portion of a microwave spectrum, at least a portion of a radio spectrum, at least a portion of an X-ray spectrum, at least a portion of a gamma spectrum, at least a portion of an electromagnetic (EM) spectrum, or a combination thereof. The first spectrum can be different from the second spectrum. In some examples, the first spectrum and the second spectrum can, in some cases, not have any overlapping portions. In some examples, the first spectrum and the second spectrum can at least partially overlap.

[0171] Figure 17 is a flowchart 1700 illustrating an example of an image processing technique. Figure 17 The image processing technique shown in flowchart 1700 may be performed by a device. The device may be the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the VSLAM device 205, the VSLAM device 305, the UGV 610, the UAV 620, the XR headset 710, one or more remote servers, one or more network servers of a cloud service, the computing system 1800, or some combination thereof.

[0172] At operation 1705, the device receives a first image of an environment captured by a first camera. The first camera is responsive to a first spectrum. At operation 1710, the device receives a second image of the environment captured by a second camera. The second camera is responsive to a second spectrum. The device may include the first camera, the second camera, or both. The device may include one or more additional cameras and / or sensors in addition to the first and second cameras. In some aspects, the device includes at least one of a mobile handheld device, a head-mounted display (HMD), a vehicle, and a robot.

[0173] The first spectrum may be different from the second spectrum. In some examples, the first spectrum and the second spectrum may not have any overlapping portions in some cases. In some examples, the first spectrum and the second spectrum may at least partially overlap. In some examples, the first camera is the first camera 310 discussed herein. In some examples, the first camera is the VL camera 310 discussed herein. In some aspects, the first spectrum is at least a portion of the visible light (VL) spectrum, and the second spectrum is different from the VL spectrum. In some examples, the first camera is the second camera 315 discussed herein. In some examples, the first camera is the IR camera 315 discussed herein. In some aspects, the second spectrum is at least a portion of the infrared (IR) spectrum, and the first spectrum is different from the IR spectrum. Either the first spectrum and the second spectrum may include at least one of: at least a portion of the VL spectrum, at least a portion of the IR spectrum, at least a portion of the ultraviolet (UV) spectrum, at least a portion of the microwave spectrum, at least a portion of the radio spectrum, at least a portion of the x-ray spectrum, at least a portion of the gamma spectrum, at least a portion of the electromagnetic (EM) spectrum, or a combination thereof.

[0174] In some examples, a first camera captures a first image when the device is in a first position, and wherein a second camera captures a second image when the device is in the first position. The device may determine a set of coordinates for a first position of the device within an environment based on a set of coordinates for a feature. The set of coordinates for the first position of the device within the environment may be referred to as a position of the device in the first position or a position of the first position. The device may determine a pose of the device when the device is in the first position based on the set of coordinates for the feature. The pose of the device may include at least one of the pitch of the device, the roll of the device, the yaw of the device, or a combination thereof. In some cases, the pose of the device may also include a set of coordinates for the first position of the device within the environment.

[0175] At operation 1715, the device identifies a feature of the environment depicted in both the first image and the second image. The feature can be a feature of the environment that is visually detectable and / or recognizable in the first image and the second image. For example, the feature can include at least one of an edge or a corner.

[0176] At operation 1720, the device determines a set of coordinates for the feature based on a first depiction of the feature in the first image and a second depiction of the feature in the second image. The set of coordinates for the feature may include three coordinates corresponding to three spatial dimensions. Determining the set of coordinates for the feature may include determining a transformation between the first set of coordinates for the feature corresponding to the first image and the second set of coordinates for the feature corresponding to the second image.

[0177] At operation 1725, the device updates a map of the environment based on the coordinate set for the feature. The device may generate a map of the environment before updating the map of the environment at operation 1725 (e.g., if the map has not yet been generated). Updating the map of the environment based on the coordinate set for the feature may include adding a new map area to the map. The new map area may include the coordinate set for the feature. Updating the map of the environment based on the coordinate set for the feature may include revising the map area of ​​the map (e.g., revising an existing map area that is already at least partially represented in the map). The map area may include the coordinate set for the feature. Revising the map area may include revising a previous coordinate set for the feature based on the coordinate set for the feature. For example, if the coordinate set for the feature is more accurate than a previous coordinate set for the feature, revising the map area may include replacing the previous coordinate set for the feature with the coordinate set for the feature. Revising the map area may include replacing the previous coordinate set for the feature with an average coordinate set for the feature. The device may determine the average coordinate set for the feature by averaging the previous coordinate set for the feature with the coordinate set for the feature (and / or one or more additional coordinate sets for the feature).

[0178] In some cases, the device may identify that the device has moved from a first location to a second location. The device may receive a third image of the environment captured by the second camera while the device is in the second location. The device may identify that a feature of the environment is depicted in at least one of the third image and a fourth image from the first camera. The device may track the feature based on one or more depictions of the feature in at least one of the third image and the fourth image. The device may determine a set of coordinates of the second location of the device within the environment based on the tracked feature. The device may determine a pose of the device while the device is in the second location based on the tracked feature. The pose of the device may include at least one of the pitch of the device, the roll of the device, the yaw of the device, or a combination thereof. In some cases, the pose of the device may include a set of coordinates of the second location of the device within the environment. The device may generate an updated set of coordinates of the feature in the environment by updating the set of coordinates of the feature in the environment based on the tracked feature. The device may update a map of the environment based on the updated set of coordinates of the feature. The tracked feature may be based on at least one of the set of coordinates of the feature, a first depiction of the feature in the first image, and a second depiction of the feature in the second image.

[0179] The environment can be well-lit, for example, by sunlight, moonlight, and / or artificial lighting. The device can identify that the lighting level of the environment is above a minimum lighting threshold when the device is in the second position. Based on the lighting level being above the minimum lighting threshold, the device can receive a fourth image of the environment captured by the first camera while the device is in the second position. In such a case, tracking the feature is based on a third depiction of the feature in the third image and a fourth depiction of the feature in the fourth image.

[0180] The environment can be poorly lit, for example, due to a lack of sunlight, a lack of moonlight, dim moonlight, a lack of artificial lighting, and / or dim artificial lighting. The device can identify that the lighting level of the environment is below a minimum lighting threshold when the device is in the second position. Based on the lighting level being below the minimum lighting threshold, tracking the feature can be based on a third depiction of the feature in the third image.

[0181] The device may identify that the device has moved from a first location to a second location. The device may receive a third image of the environment captured by the second camera while the device is in the second location. The device may identify that a second feature of the environment is depicted in at least one of the third image and a fourth image from the first camera. The device may determine a second set of coordinates for the second feature based on one or more depictions of the second feature in at least one of the third image and the fourth image. The device may update a map of the environment based on the second set of coordinates for the second feature. The device may determine a set of coordinates for a second location of the device within the environment based on the updated map. The device may determine a posture of the device while the device is in the second location based on the updated map. The posture of the device may include at least one of the pitch of the device, the roll of the device, the yaw of the device, or a combination thereof. In some cases, the posture of the device may also include a set of coordinates for the second location of the device within the environment.

[0182] The environment may be well-lit. The device may identify that the lighting level of the environment is above a minimum lighting threshold when the device is in the second position. Based on the lighting level being above the minimum lighting threshold, the device may receive a fourth image of the environment captured by the first camera while the device is in the second position. In such a case, determining the second set of coordinates for the second feature is based on the first depiction of the second feature in the third image and the second depiction of the second feature in the fourth image.

[0183] The environment may be poorly lit. The device may identify that an illumination level of the environment is below a minimum illumination threshold when the device is in the second position. Determining a second set of coordinates for the second feature based on the illumination level being below the minimum illumination threshold may be based on the first depiction of the second feature in the third image.

[0184] The first camera may have a first frame rate, and the second camera may have a second frame rate. The first frame rate may be different from (e.g., greater than or less than) the second frame rate. The first frame rate may be the same as the second frame rate. The effective frame rate of a device may refer to how many frames come in from all activated cameras per second (or per other time unit). When both the first camera and the second camera are activated, for example, when the lighting level of the environment exceeds a minimum lighting threshold, the device may have a first effective frame rate. When only one of the two cameras (e.g., only the first camera or only the second camera) is activated, for example, when the lighting level of the environment is below the minimum lighting threshold, the device may have a second effective frame rate. The first effective frame rate of the device may exceed the second effective frame rate of the device.

[0185] In some cases, at least a subset of the techniques illustrated in flowchart 1700 and conceptual diagrams 200, 300, 400, 800, 900, 1000, 1050, 1100, 1200, 1300, 1350, 1400, 1450, 1500, and 1600 may be implemented by a method related to Figure 17In some cases, at least a subset of the techniques illustrated in flowchart 1700 and conceptual diagrams 200 , 300 , 400 , 800 , 900 , 1000 , 1050 , 1100 , 1200 , 1300 , 1350 , 1400 , 1450 , 1500 , and 1600 may be performed by one or more network servers of a cloud service. In some examples, at least a subset of the techniques illustrated in flowchart 1700 and conceptual diagrams 200, 300, 400, 800, 900, 1000, 1050, 1100, 1200, 1300, 1350, 1400, 1450, 1500, and 1600 may be performed by image capture and processing system 100, image capture device 105A, image processing device 105B, VSLAM device 205, VSLAM device 305, UGV 610, UAV 620, XR headset 710, one or more remote servers, one or more network servers of a cloud service, computing system 1800, or some combination thereof. The computing system may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or a computing device of an autonomous vehicle, a robotic device, a television, and / or any other computing device with the resource capabilities to perform the processes described herein. In some cases, the computing system, device, or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing system, device, or apparatus may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on an Internet Protocol (IP) or other types of data.

[0186] Components of a computing system, device, or apparatus may be implemented in circuitry. For example, a component may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0187] The processes illustrated in flowchart 1700 and conceptual diagrams 200, 300, 400, and 1200 are organized as a logical flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, these operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the operations. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.

[0188] In addition, at least a subset of the techniques shown in flowchart 1700 and conceptual diagrams 200, 300, 400, 800, 900, 1000, 1050, 1100, 1200, 1300, 1350, 1400, 1450, 1500, and 1600 described herein can be performed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed on one or more processors, through hardware, or a combination thereof. As described above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.

[0189] Figure 18 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. Specifically, Figure 18 The diagram illustrates an example of a computing system 1800, which can be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using a connection 1805. The connection 1805 can be a physical connection using a bus, or a direct connection (such as in a chipset architecture) to the processor 1810. The connection 1805 can also be a virtual connection, a network connection, or a logical connection.

[0190] In some embodiments, computing system 1800 is a distributed system, in which the functionality described in this disclosure can be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represents a plurality of such components, each of which performs some or all of the functionality described for that component. In some embodiments, a component can be a physical or virtual device.

[0191] Example system 1800 includes at least one processing unit (CPU or processor) 1810 and connections 1805 that couple various system components to processor 1810, including system memory 1815, such as read-only memory (ROM) 1820 and random access memory (RAM) 1825. Computing system 1800 may include a cache 1812 of high-speed memory directly connected to, immediately adjacent to, or integrated as part of processor 1810.

[0192] Processor 1810 may include any general-purpose processor and hardware or software services configured to control processor 1810, such as services 1832, 1834, and 1836 stored in storage device 1830, as well as dedicated processors where software instructions are incorporated into the actual processor design. Processor 1810 may essentially be a completely independent computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0193] To enable user interaction, the computing system 1800 includes an input device 1845, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, and the like. The computing system 1800 can also include an output device 1835, which can be one or more of several output mechanisms. In some instances, a multimodal system can enable a user to provide multiple types of input / output to communicate with the computing system 1800. The computing system 1800 can include a communication interface 1840, which can generally govern and manage user input and system output. The communication interface can use wired and / or wireless transceivers to perform or facilitate receiving and / or sending wired or wireless communications, including those using the following: audio jacks / plugs, microphone jacks / plugs, universal serial bus (USB) ports / plugs, / plug, Ethernet port / plug, fiber optic port / plug, proprietary wired port / plug, Wireless signal transmission, Low energy (BLE) wireless signal transmission, Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. The communication interface 1840 may also include one or more global navigation satellite system (GNSS) receivers or transceivers that are used to determine the location of the computing system 1800 based on receiving one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There is no restriction to operate on any particular hardware arrangement, and thus the basic features herein can be easily replaced with improved hardware or firmware arrangements as they are developed.

[0194] The storage device 1830 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other type of computer-readable medium (which may store data accessible by a computer), such as a magnetic cartridge, a flash memory card, a solid-state memory device, a digital versatile disk, a cassette, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage medium, flash memory, a memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) optical disc, a rewritable compact disc (CD), a digital video disc (DVD), a Blu-ray disc (BDD), a holographic optical disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASH EPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.

[0195] Storage devices 1830 may include software services, servers, services, etc. that, when the code defining such software is executed by processor 1810, cause the system to perform a function. In some embodiments, hardware services that perform a particular function may include software components stored in a computer-readable medium that interface with the necessary hardware components (such as processor 1810, connection 1805, output device 1835, etc.) to perform the function.

[0196] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transient media in which data can be stored and do not include carrier waves and / or transient electronic signals that are propagated wirelessly or by wired connections. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent any combination of a process, function, subroutine, program, routine, subroutine, module, software package, class, or instruction, data structure, or program statement. A code segment may be coupled to another code segment or hardware circuit by transmitting and / or receiving information, data, independent variables, parameters, or memory contents. Any suitable means may be used to transmit, forward, or send information, independent variables, parameters, data, etc., including memory sharing, message passing, token passing, network transmission, etc.

[0197] In certain embodiments, computer-readable storage devices, media, and memories may include wired or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0198] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that embodiments can be put into practice without these specific details. For clarity of explanation, in some cases, the present technology can be presented as including separate functional blocks, which include functional blocks that include steps or routines in the method embodied by devices, device components, software or a combination of hardware and software. Additional components other than those shown in the figures and / or described herein can be used. For example, circuits, systems, networks, processes and other components can be shown as components in block diagram form, so as not to obscure the embodiments in unnecessary details. In other instances, well-known circuits, processes, algorithms, structures and techniques can be shown without unnecessary details, so as not to obscure the embodiments.

[0199] Individual embodiments may be described above as processes or methods, which are depicted as flow charts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flow chart may describe an operation as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. A process may be terminated when its operation is complete, but may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or main function.

[0200] The processes and methods according to the above examples can be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform certain functions or groups of functions. Parts of the computer resources used may be accessed over a network. Computer-executable instructions may be, for example, binary, intermediate format instructions (such as assembly language), firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the method according to the described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, and the like.

[0201] Devices implementing the processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may employ any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) that perform the necessary tasks may be stored in a computer-readable or machine-readable medium. (Multiple) processors may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, and the like. The functions described herein may also be embodied in peripheral devices or add-on cards. As a further example, such functions may also be implemented on circuit boards between different chips or on different processes performed in a single device.

[0202] The instructions, the media for carrying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0203] In the foregoing description, aspects of the present application have been described with reference to their specific embodiments, but those skilled in the art will recognize that the present application is not limited thereto. Therefore, although the illustrative embodiments of the present application have been described in detail herein, it should be understood that the inventive concept can be embodied and adopted differently in other ways, and the appended claims are intended to be interpreted as including such variants except those limited by the prior art. The various features and aspects of the above-mentioned application can be used alone or in combination. In addition, the embodiments can be utilized in any number of environments and applications exceeding the environments and applications described herein without departing from the broader spirit and scope of this specification. Accordingly, the description and drawings should be considered to be illustrative rather than restrictive. For illustrative purposes, the method is described in a particular order. It should be understood that in alternative embodiments, the method can be performed in an order different from the described order.

[0204] A person of ordinary skill will understand that the less than ("<") and greater than (">") symbols or terms used in this document can be replaced by the less than or equal to ("≤") and greater than or equal to ("≥") symbols respectively without departing from the scope of this specification.

[0205] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform those operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuit) to perform those operations, or any combination thereof.

[0206] The phrase “coupled to” means any component is directly or indirectly physically connected to another component, and / or any component is in direct or indirect communication with another component (e.g., connected to another component via a wired or wireless connection, and / or other suitable communication interface).

[0207] Claim language that states “at least one of” a set or “one or more” a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language that states “at least one of A and B” means A, B, or A and B. In another example, claim language that states “at least one of A, B, and C” means A, B, C, or A and B, or A and C, or B and C, or A, B, and C. The language “at least one of” a set and / or “one or more” a set does not limit the set to the items listed in the set. For example, claim language that states “at least one of A and B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0208] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as causing a departure from the scope of this application.

[0209] The techniques described herein can be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device handheld device, or an integrated circuit device with multiple uses, including applications in wireless communication device handheld devices and other devices. Any features described as modules or components can be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code, which includes instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which can include packaging materials. The computer-readable medium can include a memory or data storage medium, such as a random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. The technology may additionally or alternatively be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0210] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Accordingly, the term "processor" as used herein can refer to any one of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the technology described herein. In addition, in some aspects, the functions described herein can be provided within a dedicated software module or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (codec).

Claims

1. A device for processing image data, the device comprising: one or more memory units that store instructions; as well as one or more processors that execute the instructions, wherein execution of the instructions by the one or more processors causes the one or more processors to: receiving a first image of an environment captured by a first camera, the first camera responsive to a first spectrum; receiving a second image of the environment captured by a second camera, the second camera responsive to a second light spectrum, wherein the first camera captured the first image while the device was in a first position, and wherein the second camera captured the second image while the device was in the first position; a first feature identifying the environment being depicted in both the first image and the second image; determining a set of coordinates of the first feature based on a first depiction of the first feature in the first image and a second depiction of the first feature in the second image; updating a map of the environment based on the set of coordinates for the first feature; identifying that the device has moved from the first position to a second position; receiving a third image of the environment captured by the second camera while the apparatus is in the second position; identifying whether a lighting level of the environment is above or below a minimum lighting threshold when the apparatus is in the second position; receiving, in response to the illumination level being above the minimum illumination threshold, a fourth image of the environment captured by the first camera while the apparatus is in the second position, and identifying that a second feature of the environment is depicted in both the third image and the fourth image; identifying that the device has moved from the second position to a third position; receiving a fifth image of the environment captured by the second camera while the apparatus is in the third position; identifying whether a lighting level of the environment is above or below the minimum lighting threshold when the apparatus is in the third position; as well as In response to the illumination level being below the minimum illumination threshold, receiving a sixth image of the environment captured by the first camera to check whether the illumination level of the environment has changed, and identifying a third feature of the environment is depicted in the fifth image. 2 . The device of claim 1 , wherein the device is at least one of a mobile handheld device, a head-mounted display (HMD), a vehicle, and a robot. 3 . The apparatus of claim 1 , wherein the apparatus comprises at least one of the first camera and the second camera. The apparatus of claim 1 , wherein the first spectrum is at least a portion of a visible light (VL) spectrum, and the second spectrum is different from the VL spectrum.

5. The apparatus of claim 1, wherein the second spectrum is at least a portion of an infrared (IR) spectrum, and the first spectrum is different from the IR spectrum. The apparatus of claim 1 , wherein the set of coordinates of the first feature comprises three coordinates corresponding to three spatial dimensions.

7. The apparatus of claim 1 , wherein execution of the instructions by the one or more processors causes the one or more processors to: A set of coordinates of the first position of the device within the environment is determined based on the set of coordinates for the first feature.

8. The apparatus of claim 1 , wherein execution of the instructions by the one or more processors causes the one or more processors to: A pose of the device when the device is in the first position is determined based on the set of coordinates for the first feature, wherein the pose of the device comprises at least one of a pitch of the device, a roll of the device, and a yaw of the device.

9. The apparatus of claim 1 , wherein execution of the instructions by the one or more processors causes the one or more processors to: identifying the first feature of the environment as depicted in at least one of the third image and a fourth image from the first camera; and The first feature is tracked based on one or more depictions of the first feature in at least one of the third image and the fourth image.

10. The apparatus of claim 9, wherein execution of the instructions by the one or more processors causes the one or more processors to: A set of coordinates of the second position of the device within the environment is determined based on tracking the first feature.

11. The apparatus of claim 9, wherein execution of the instructions by the one or more processors causes the one or more processors to: A pose of the device when the device is in the second position is determined based on tracking the first feature, wherein the pose of the device includes at least one of a pitch of the device, a roll of the device, and a yaw of the device.

12. The apparatus of claim 9, wherein execution of the instructions by the one or more processors causes the one or more processors to: generating an updated set of coordinates for the first feature in the environment by updating the set of coordinates for the first feature in the environment based on tracking the first feature; and The map of the environment is updated based on the updated set of coordinates of the first feature.

13. The apparatus of claim 9 , wherein execution of the instructions by the one or more processors causes the one or more processors to: In response to an illumination level of the environment being above the minimum illumination threshold when the apparatus is in the second position, the first feature is tracked based on a third depiction of the first feature in the third image and a fourth depiction of the first feature in the fourth image.

14. The apparatus of claim 9, wherein execution of the instructions by the one or more processors causes the one or more processors to: Identifying that an illumination level of the environment is below a minimum illumination threshold when the apparatus is in the second position, wherein tracking the first feature is based on a third depiction of the first feature in the third image.

15. The apparatus of claim 9, wherein tracking the first feature is further based on at least one of the set of coordinates of the first feature, the first depiction of the first feature in the first image, and the second depiction of the first feature in the second image.

16. The apparatus of claim 1 , wherein execution of the instructions by the one or more processors causes the one or more processors to: identifying a second feature of the environment depicted in at least one of the third image and a fourth image from the first camera; determining a second set of coordinates for the second feature based on one or more depictions of the second feature in at least one of the third image and the fourth image; and The map of the environment is updated based on the second set of coordinates for the second feature.

17. The apparatus of claim 16, wherein execution of the instructions by the one or more processors causes the one or more processors to: A set of coordinates of the second location of the device within the environment is determined based on updating the map.

18. The apparatus of claim 16, wherein execution of the instructions by the one or more processors causes the one or more processors to: A pose of the device when the device is in the second position is determined based on updating the map, wherein the pose of the device includes at least one of a pitch of the device, a roll of the device, and a yaw of the device.

19. The apparatus of claim 16, wherein execution of the instructions by the one or more processors causes the one or more processors to: In response to the illumination level of the environment being above the minimum illumination threshold when the device is in the second position, determining the second set of coordinates of the second feature based on the first depiction of the second feature in the third image and the second depiction of the second feature in the fourth image.

20. The apparatus of claim 16, wherein execution of the instructions by the one or more processors causes the one or more processors to: An identification is made that an illumination level of the environment is below a minimum illumination threshold when the apparatus is in the second position, wherein determining the second set of coordinates for the second feature is based on a first depiction of the second feature in the third image.

21. The apparatus of claim 1, wherein determining the set of coordinates for the first feature comprises determining a transformation between a first set of coordinates for the first feature corresponding to the first image and a second set of coordinates for the first feature corresponding to the second image.

22. The apparatus of claim 1 , wherein execution of the instructions by the one or more processors causes the one or more processors to: The map of the environment is generated before updating the map of the environment.

23. The apparatus of claim 1, wherein updating the map of the environment based on the set of coordinates for the first feature comprises adding a new map area to the map, the new map area comprising the set of coordinates for the first feature.

24. The apparatus of claim 1, wherein updating the map of the environment based on the set of coordinates for the first feature comprises revising a map area of ​​the map that includes the set of coordinates for the first feature.

25. The apparatus of claim 1, wherein the first feature is at least one of an edge and a corner.

26. A method for processing image data, the method comprising: receiving a first image of an environment captured by a first camera, the first camera responsive to a first spectrum; receiving a second image of the environment captured by a second camera, the second camera responsive to a second light spectrum, wherein the first camera captured the first image while the device was in a first position, and wherein the second camera captured the second image while the device was in the first position; a first feature identifying the environment being depicted in both the first image and the second image; determining a set of coordinates of the first feature based on a first depiction of the first feature in the first image and a second depiction of the first feature in the second image; as well as updating a map of the environment based on the set of coordinates for the first feature; identifying that the device has moved from the first position to a second position; receiving a third image of the environment captured by the second camera while the apparatus is in the second position; identifying whether a lighting level of the environment is above or below a minimum lighting threshold when the apparatus is in the second position; as well as receiving, in response to the illumination level being above the minimum illumination threshold, a fourth image of the environment captured by the first camera while the apparatus is in the second position, and identifying that a second feature of the environment is depicted in both the third image and the fourth image; identifying that the device has moved from the second position to a third position; receiving a fifth image of the environment captured by the second camera while the apparatus is in the third position; identifying whether a lighting level of the environment is above or below the minimum lighting threshold when the apparatus is in the third position; as well as In response to the illumination level being below the minimum illumination threshold, receiving a sixth image of the environment captured by the first camera to check whether the illumination level of the environment has changed, and identifying a third feature of the environment is depicted in the fifth image.

27. The method of claim 26, wherein the first spectrum is at least a portion of a visible light (VL) spectrum, and the second spectrum is different from the VL spectrum.

28. The method of claim 26, wherein the second spectrum is at least a portion of an infrared (IR) spectrum, and the first spectrum is different from the IR spectrum.

29. The method of claim 26, wherein the set of coordinates of the first feature includes three coordinates corresponding to three spatial dimensions.

30. The method of claim 26, further comprising: A set of coordinates of the first position of the device within the environment is determined based on the set of coordinates for the first feature.

31. The method of claim 26, further comprising: A pose of the device when the device is in the first position is determined based on the set of coordinates for the first feature, wherein the pose of the device comprises at least one of a pitch of the device, a roll of the device, and a yaw of the device.

32. The method of claim 26, further comprising: identifying the first feature of the environment as depicted in at least one of the third image and a fourth image from the first camera; as well as The first feature is tracked based on one or more depictions of the first feature in at least one of the third image and the fourth image.

33. The method of claim 26, further comprising: identifying a second feature of the environment depicted in at least one of the third image and a fourth image from the first camera; determining a second set of coordinates for the second feature based on one or more depictions of the second feature in at least one of the third image and the fourth image; as well as The map of the environment is updated based on the second set of coordinates for the second feature.

34. A computer program product comprising computer-readable instructions, which, when executed by a processor, cause the processor to perform the method of any one of claims 26 to 33.

35. A computer-readable medium having program code recorded thereon, wherein the program code is executable by one or more processors to cause the processors to perform the method of any one of claims 26-33.

Citation Information

Patent Citations

  • SLAM (simultaneous localization and mapping) system based on fusion of color images and infrared images

    CN107204015A