System and method for motion blur compensation for feature tracking
By receiving image sensor and motion sensor data, estimating the motion blur level and determining feature weights, the problem of feature tracking inaccuracy caused by motion blur in the imaging system is solved, and the accuracy and reliability of the system are improved.
Patent Information
- Application Number
- CN202480010973.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-16
- Filing Date
- 2024-01-31
- Publication Date
- 2025-09-16
AI Technical Summary
Existing imaging and feature tracking systems are not accurate enough when dealing with motion blur, especially when the motion of the camera and the environment during image capture causes blur that affects the performance of feature tracking.
By receiving images captured by an image sensor and motion data captured by a motion sensor, the motion blur level is estimated, the weights of the features are determined based on the level, and the features are tracked according to the weights, and the weights are used to perform feature matching and mapping between images.
The accuracy and reliability of feature tracking are improved, the impact of motion blur on the imaging system is reduced, and the accuracy of the system in environment mapping and pose estimation is enhanced.
Smart Images

Figure CN120660119A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to imaging and feature tracking. For example, aspects of the present disclosure relate to systems and techniques for voxel-based mapping of an environment based on image data and depth data. Background Art
[0002] A camera is a device that includes an image sensor that receives light from a scene and captures image data depicting the scene, such as still images or video frames of a video. A depth sensor is a sensor that obtains depth data that indicates how far different points in the scene are from the depth sensor. The depth data may include a depth map, a depth image, a point cloud, or another indication of depth, range, and / or distance. A depth sensor may also be referred to as a range sensor or a distance sensor. A depth sensor may have limitations in the depth data it obtains. For example, depth data captured by a depth sensor may identify the depth of points along the edges of a surface in a scene, but not the depth of other portions of the surface that lie between those edges. Summary of the Invention
[0003] Systems and techniques for imaging are described. In some examples, a system receives an image of an environment captured using an image sensor according to image capture settings, and receives motion data captured using a motion sensor. The system determines a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level for the at least one feature of the environment in the image. The estimated motion blur level is based on the motion data and the image capture settings. The system tracks a feature of the environment across a plurality of images (including the received image) according to the respective weights for the feature of the environment across the plurality of images (including the received image), including the determined weight. For example, images in a set of images captured when the system was stationary (and corresponding features in those images) may have higher respective weights due to a higher confidence in their accuracy (e.g., because there is a low probability of motion blur) and thus play a greater role in tracking the feature, while images in the set of images captured when the system was moving may have lower respective weights due to a lower confidence in their accuracy (e.g., because there is a high probability of motion blur) and thus play a smaller role in tracking the feature. The system may use the tracked features to map the environment, localize the system, and / or determine a pose of the system.
[0004] According to at least one example, an apparatus for imaging is provided. The apparatus includes a memory and at least one processor (e.g., implemented in a circuit) coupled to the memory. The at least one processor is configured and capable of: receiving an image of an environment captured using at least one image sensor according to image capture settings; receiving motion data captured using a motion sensor; determining a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level of the at least one feature of the environment in the image, wherein the estimated motion blur level is based on the motion data and the image capture settings; and tracking the feature of the environment across a plurality of images according to the respective weights of the feature of the environment across the plurality of images, wherein the plurality of images includes the image, and wherein the respective weights include the weight.
[0005] In another example, an imaging method is provided. The method includes receiving an image of an environment captured using at least one image sensor according to image capture settings; receiving motion data captured using a motion sensor; determining a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level of the at least one feature of the environment in the image, wherein the estimated motion blur level is based on the motion data and the image capture settings; and tracking the feature of the environment across a plurality of images according to respective weights for the feature of the environment across the plurality of images, wherein the plurality of images includes the image, and wherein the respective weights include the weight.
[0006] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: receive an image of an environment captured using at least one image sensor according to image capture settings; receive motion data captured using a motion sensor; determine a weight associated with at least one of multiple features of the environment in the image based on an estimated motion blur level of the at least one feature of the environment in the image, wherein the estimated motion blur level is based on the motion data and the image capture settings; and track the feature of the environment across multiple images according to the corresponding weights for the feature of the environment across the multiple images, wherein the multiple images include the image, and wherein the corresponding weights include the weight.
[0007] In another example, an apparatus for imaging is provided. The apparatus includes: means for receiving an image of an environment captured using at least one image sensor according to image capture settings; means for receiving motion data captured using a motion sensor; means for determining a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level of the at least one feature of the environment in the image, wherein the estimated motion blur level is based on the motion data and the image capture settings; and means for tracking the feature of the environment across a plurality of images according to the respective weights for the feature of the environment across the plurality of images, wherein the plurality of images includes the image, and wherein the respective weights include the weight.
[0008] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: determining a pose of the device in the environment based on the features tracked in the multiple images and according to the corresponding weights of the features of the environment in the multiple images, wherein the device includes the at least one image sensor and the motion sensor. In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: minimizing a weighted least squares reprojection error according to the corresponding weights of the features of the environment in the multiple images to determine the pose of the device in the environment. In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: outputting an indication of the pose of the device.
[0009] In some aspects, one or more of the methods, devices, and computer-readable media described above further include mapping the environment based on the features of the environment tracked in the multiple images and according to the corresponding weights of the features of the environment in the multiple images to generate a map of the environment. In some aspects, one or more of the methods, devices, and computer-readable media described above further include determining a position of the device within the map of the environment based on the features of the environment tracked in the multiple images and according to the corresponding weights of the features of the environment in the multiple images, wherein the device includes the at least one image sensor and the motion sensor. In some aspects, one or more of the methods, devices, and computer-readable media described above further include outputting at least a portion of the map of the environment.
[0010] In some aspects, the respective weights of the feature of the environment across the multiple images correspond to respective error variance values of the multiple images, further comprising: tracking the feature of the environment across the multiple images based on the respective error variance values of the multiple images to track the feature of the environment across the multiple images based on the respective weights of the feature of the environment across the multiple images.
[0011] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include determining the estimated motion blur level of the at least one of the features of the environment in the image.
[0012] In some aspects, the estimated motion blur level of the at least one of the features of the environment in the image is based on a distance from the at least one image sensor to the at least one of the features of the environment, wherein the weight associated with the at least one of the features of the environment in the image is based on the distance from the at least one image sensor to the at least one of the features of the environment.
[0013] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include determining a ratio of a constant divided by the estimated motion blur level of the image based on the estimated motion blur level of at least one of the features of the environment in the image to determine the weight associated with the image.
[0014] In some aspects, determining that the estimated level of motion blur is less than a predetermined threshold; and in response to determining that the estimated level of motion blur is less than the predetermined threshold, setting the weight associated with the at least one of the features of the environment in the image to a predetermined value to determine the weight associated with the at least one of the features of the environment in the image based on the estimated level of motion blur of the at least one of the features of the environment in the image. In some aspects, the predetermined threshold represents a motion blur magnitude no greater than a pixel.
[0015] In some aspects, the estimated motion blur level is an estimated motion blur magnitude.
[0016] In some aspects, the image capture settings include an exposure time.
[0017] In some aspects, the apparatus is part of and / or includes a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a head-mounted display (HMD) device, a wireless communication device, a mobile device (e.g., a mobile phone and / or mobile handset and / or so-called "smartphone" or other mobile device), a camera, a personal computer, a laptop computer, a server computer, a vehicle or a computing device or a component of a vehicle, another device, or a combination thereof. In some aspects, the apparatus includes one or more cameras for capturing one or more images. In some aspects, the apparatus further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatus described above may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors).
[0018] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. This subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0019] The foregoing and other features and aspects will become more apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Illustrative aspects of the present application are described in detail below with reference to the following drawings:
[0021] Figure 1 is a block diagram illustrating an example architecture of an image capture and processing system according to some examples;
[0022] Figure 2 is a block diagram illustrating an example architecture of an imaging system that performs imaging and feature tracking processes according to some examples;
[0023] Figure 3A is a perspective view illustrating a head-mounted display (HMD) used as part of an imaging system according to some examples;
[0024] Figure 3B is an example based on some examples Figure 3A A perspective view of a head-mounted display (HMD) being worn by a user;
[0025] Figure 4Ais a perspective view illustrating a front surface of a mobile phone including a front-facing camera and usable as part of an imaging system according to some examples;
[0026] Figure 4B is a perspective view illustrating a rear surface of a mobile phone including a rear camera and usable as part of an imaging system according to some examples;
[0027] Figure 5 is a perspective view illustrating a vehicle including various sensors according to some examples;
[0028] Figure 6 is a conceptual diagram illustrating feature tracking in an image with motion blur according to some examples;
[0029] Figure 7 is a block diagram illustrating a process for pose estimation according to some examples, the process taking into account exposure time;
[0030] Figure 6 is a conceptual diagram illustrating feature tracking in an image with motion blur;
[0031] Figure 7 is a block diagram illustrating a process for pose estimation that takes exposure time into account;
[0032] Figure 8 is a block diagram illustrating a process for automatic exposure control according to some examples that takes motion data into account;
[0033] Figure 9 is a flow diagram illustrating a process of pose estimation according to some examples, the process taking into account estimated motion blur;
[0034] Figure 10 is a flow diagram illustrating a process of pose estimation according to some examples, the process taking into account estimated motion blur;
[0035] Figure 11 is a conceptual diagram illustrating image reprojection according to some examples;
[0036] Figure 12 is a block diagram illustrating a process for automatic exposure control according to some examples that takes into account angular velocity, motion information, scene depth, and / or optical flow;
[0037] Figure 13 is a block diagram illustrating an example of a neural network that may be used for environment mapping according to some examples;
[0038] Figure 14 is a flow diagram illustrating a process for mapping an environment according to some examples; and
[0039] Figure 15is a diagram illustrating an example of a computing system for implementing certain aspects described herein. DETAILED DESCRIPTION
[0040] The following provides certain aspects of the present disclosure. Some of these aspects can be applied independently, and some of them can be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for explanation purposes to provide a thorough understanding of various aspects of the application. However, it will be apparent that various aspects can be implemented without these specific details. Each drawing and description is not intended to be restrictive.
[0041] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. On the contrary, the following description of exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0042] A camera is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). The terms "image," "image frame," and "frame" are used interchangeably herein. A camera can be configured with various image capture and image processing settings. Different settings produce images with different appearances. Some camera settings, such as ISO, exposure time, aperture size, f-stop, shutter speed, focus, and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters can be applied to the image sensor used to capture one or more image frames. Other camera settings can configure post-processing of one or more image frames, such as changes to contrast, brightness, saturation, sharpness, levels, curves, or color. For example, settings or parameters can be applied to a processor (e.g., an image signal processor or ISP) used to process one or more image frames captured by the image sensor.
[0043] Exposure time refers to how long a camera's aperture is open (during image capture) to expose the camera's image sensor (or, in some cameras, a sheet of film) to light from the environment. Longer exposure times allow the image sensor to receive more light from the environment, resulting in images captured by the camera depicting the environment as brighter. However, if the camera moves (and / or if objects in the environment move) during the time the aperture is open and the image sensor receives light during image capture, longer exposure times can also result in motion blur in the image. Motion blur occurs because the image sensor receives light from certain objects in the scene at different times and from different directions during the time the aperture is open and the image sensor receives light during image capture, corresponding to the objects' relative positions to the camera. Shorter exposure times can reduce motion blur if the camera (and / or objects in the environment) are moving, but can result in images captured by the camera being more dimly lit and, in some cases, noisier and / or grainier.
[0044] A motion sensor is a sensor that obtains data regarding position (e.g., lateral position and / or altitude), orientation (e.g., pitch, yaw, and / or roll), pose (e.g., position and / or orientation), motion (e.g., change in position and / or orientation), rate, velocity, acceleration, or a combination thereof. A motion sensor may include a global navigation satellite system (GNSS) receiver, an inertial measurement unit (IMU), an accelerometer, a gyroscope, a gyrometer, a barometer, an altimeter, or a combination thereof. A motion sensor may be coupled to a device such as a camera or a device that includes a camera (e.g., a mobile phone, a head-mounted display (HMD) device, a vehicle, a computer, or any other device discussed herein). In the case where a motion sensor is coupled to a camera, the motion sensor may indicate when the camera is stationary, moving, accelerating, decelerating, turning, etc.
[0045] Imaging and feature tracking systems and techniques are described. In some examples, a system receives an image of an environment captured using an image sensor according to image capture settings, and receives motion data captured using a motion sensor. The system determines a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level of the at least one feature of the environment in the image. The estimated motion blur level is based on the motion data and the image capture settings. The system tracks the feature of the environment across a plurality of images (including the received image) according to the respective weights for the feature of the environment across the plurality of images (including the received image), including the determined weight. For example, images in a set of images captured when the system was stationary may have higher respective weights due to a higher confidence in their accuracy (e.g., because there is a low probability of motion blur) and thus play a greater role in feature tracking, while images in the set of images captured when the system was moving may have lower respective weights due to a lower confidence in their accuracy (e.g., because there is a high probability of motion blur) and thus play a smaller role in feature tracking. The system may use the tracked features to map the environment, localize the system, and / or determine a pose of the system.
[0046] The imaging and feature tracking systems and techniques described herein provide a number of technical improvements over existing imaging and feature tracking systems. For example, the imaging and feature tracking systems and techniques described herein can provide more accurate feature tracking by using weights to account for motion blur in different images, compared to feature tracking systems and techniques that do not account for motion blur. For example, if the imaging and feature tracking systems and techniques described herein determine that an image is estimated to have a high level of motion blur (e.g., due to a high exposure time and / or motion detected using a motion sensor), the system may choose not to use features extracted from that image for feature tracking, or choose to weight features extracted from that image less than features extracted from an image estimated to have a low or non-existent level of motion blur (e.g., due to a low exposure time and / or a lack of motion detected using a motion sensor). The imaging and feature tracking systems and techniques described herein can also be used to set image capture settings (such as exposure time) to reduce the estimated level of motion blur of an image to be used in feature tracking, while keeping noise / grain levels low enough to avoid affecting feature tracking. In this way, imaging and feature tracking systems and techniques provide more reliable image capture for feature tracking and systems that rely on feature tracking, such as imaging systems, pose tracking systems, visual simultaneous localization and mapping (VSLAM) systems, other simultaneous localization and mapping (SLAM) systems, etc.
[0047] Various aspects of the application will be described with respect to the accompanying drawings. Figure 1is a block diagram illustrating the architecture of image capture and processing system 100. Image capture and processing system 100 includes various components for capturing and processing images of one or more scenes (e.g., images of scene 110). Image capture and processing system 100 can capture individual images (or photographs) and / or can capture a video comprising multiple images (or video frames) in a particular sequence. Lens 115 of system 100 faces scene 110 and receives light from scene 110. Lens 115 bends the light toward image sensor 130. The light received by lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by image sensor 130. In some examples, scene 110 is a scene in an environment. In some examples, image capture and processing system 100 is coupled to vehicle 190 and / or a portion of vehicle 190, and scene 110 is a scene in the environment surrounding vehicle 190. In some examples, scene 110 is a scene of at least a portion of a user. For example, scene 110 may be a scene of one or both of a user's eyes and / or at least a portion of a user's face.
[0048] The one or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. The one or more control mechanisms 120 may include a plurality of mechanisms and components; for example, the control mechanisms 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. The one or more control mechanisms 120 may also include additional control mechanisms beyond those illustrated, such as controls for analog gain, flash, HDR, depth of field, and / or other image capture attributes.
[0049] Focus control mechanism 125B of control mechanism 120 may obtain a focus setting. In some examples, focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, focus control mechanism 125B may adjust the position of lens 115 relative to the position of image sensor 130. For example, based on the focus setting, focus control mechanism 125B may actuate a motor or servo to move lens 115 closer to or further away from image sensor 130, thereby adjusting the focus. In some cases, system 100 may include additional lenses, such as one or more microlenses above each photodiode of image sensor 130, each of which bends light received from lens 115 toward the corresponding photodiode before it reaches the photodiode. The focus setting may be determined using contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus setting may be determined using control mechanism 120, image sensor 130, and / or image processor 150. The focus setting may be referred to as an image capture setting and / or an image processing setting.
[0050] Exposure control mechanism 125A of control mechanism 120 may obtain an exposure setting. In some cases, exposure control mechanism 125A stores the exposure setting in a memory register. Based on the exposure setting, exposure control mechanism 125A may control the size of the aperture (e.g., aperture size or number of stops), the duration that the aperture is open (e.g., exposure time or shutter speed), the sensitivity of image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by image sensor 130, or any combination thereof. The exposure setting may be referred to as image capture settings and / or image processing settings. In some examples, the exposure setting may be provided to exposure control mechanism 125A of control mechanism 120 from image processor 150, host processor 152, ISP 154, or a combination thereof, for example, as described with respect to Figure 8 and / or Figure 12 Further illustration and discussion.
[0051] Zoom control mechanism 125C of control mechanism 120 may obtain a zoom setting. In some examples, zoom control mechanism 125C stores the zoom setting in a memory register. Based on the zoom setting, zoom control mechanism 125C may control the focal length of an assembly of lens elements (lens assembly) including lens 115 and one or more additional lenses. For example, zoom control mechanism 125C may control the focal length of the lens assembly by actuating one or more motors or servos to move one or more of the lenses relative to each other. The zoom setting may be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, the focusing lens may be lens 115) that first receives light from scene 110, where the light then passes through an afocal zoom system between the focusing lens (e.g., lens 115) and image sensor 130 before reaching image sensor 130. In some cases, an afocal zoom system may include two positive (e.g., converging, convex) lenses having equal or similar focal lengths (e.g., within a threshold difference), with a negative (e.g., diverging, concave) lens between them. In some cases, zoom control mechanism 125C moves one or more of the lenses in the afocal zoom system, such as the negative lens and one or both of the positive lenses.
[0052] Image sensor 130 includes one or more arrays of photodiodes or other light-sensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by image sensor 130. In some cases, different photodiodes may be covered by different color filters and, therefore, may measure light that matches the color of the color filter covering the photodiode. For example, a Bayer color filter includes a red filter, a blue filter, and a green filter, where each pixel of an image is generated based on red light data from at least one photodiode covered in the red filter, blue light data from at least one photodiode covered in the blue filter, and green light data from at least one photodiode covered in the green filter. Other types of color filters may use yellow, magenta, and / or cyan (also known as "emerald") filters as an alternative to or in addition to red, blue, and / or green filters. Some image sensors may have no color filters at all and, instead, use different photodiodes (in some cases stacked vertically) throughout the pixel array. Different photodiodes throughout the pixel array may have different spectral sensitivity curves, thereby responding to different wavelengths of light.Monochrome image sensors may also lack color filters and, therefore, lack color depth.
[0053] In some cases, image sensor 130 may alternatively or additionally include an opaque mask and / or a reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which may be used for phase detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output by the photodiode and / or an analog-to-digital converter (ADC) for converting the analog signal output by the photodiode (and / or the analog signal amplified by the analog gain amplifier) into a digital signal. In some cases, certain components or functions discussed with respect to one or more control mechanisms in control mechanism 120 may alternatively or additionally be included in image sensor 130. Image sensor 130 may be a charge coupled device (CCD) sensor, an electron multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal oxide semiconductor (CMOS), an N-type metal oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0054] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more of any other type of processors 1510 discussed with respect to the computing system 1500. The host processor 152 may be a digital signal processor (DSP) and / or other type of processor. In some implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system on a chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth tM, Global Positioning System (GPS), etc.), any combination thereof, and / or other components. The I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a serial general-purpose input / output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In one illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port, and the ISP 154 may communicate with the image sensor 130 using a MIPI port.
[0055] The image processor 150 may perform a number of tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. The image processor 150 may store image frames and / or processed images in random access memory (RAM) 140 and / or 1520, read-only memory (ROM) 145 and / or 1525, a cache, a memory unit, another storage device, or some combination thereof.
[0056] Various input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device 1535, any other input device 1545, or some combination thereof. In some cases, subtitles may be entered into the image processing device 105B via a physical keyboard or keypad of the I / O device 160, or via a virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O 160 may include one or more ports, jacks, or other connectors that enable wired connections between the system 100 and one or more peripheral devices, via which the system 100 may receive data from the one or more peripheral devices and / or transmit data to the one or more peripheral devices. The I / O 160 may also include one or more wireless transceivers that enable wireless connections between the system 100 and one or more peripheral devices, via which the system 100 may receive data from the one or more peripheral devices and / or transmit data to the one or more peripheral devices. Peripheral devices may include any of the types of I / O devices 160 discussed previously, and may themselves be considered I / O devices 160 once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector.
[0057] In some cases, the image capture and processing system 100 can be a single device. In some cases, the image capture and processing system 100 can be two or more independent devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some implementations, the image capture device 105A and the image processing device 105B can be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some implementations, the image capture device 105A and the image processing device 105B can be disconnected from each other.
[0058] like Figure 1 As shown, the vertical dotted line will Figure 1 1. The image capture and processing system 100 is divided into two parts, representing an image capture device 105A and an image processing device 105B. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, some components illustrated in the image capture device 105A (such as the ISP 154 and / or the host processor 152) may be included in the image capture device 105A.
[0059] The image capture and processing system 100 may include an electronic device, such as a mobile or landline telephone handset (e.g., a smartphone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 1502.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile phone, a desktop computer, or other computing device.
[0060] Although the image capture and processing system 100 is shown as including certain components, one of ordinary skill will appreciate that the image capture and processing system 100 may include more than Figure 1 . The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of the image capture and processing system 100 may include, and / or be implemented using, electronic circuits or other electronic hardware that may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits) and / or may include, and / or be implemented using, computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device that implements the image capture and processing system 100.
[0061] Figure 2100 , 1200 , and / or 1400 ), a computing system 1500 , a processor 1510 , or a combination thereof. In some examples, imaging system 200 may include or may be part of, for example, one or more laptop computers, phones, tablet computers, mobile phones, video game consoles, vehicle computers, vehicles, desktop computers, wearable devices, televisions, media centers, XR systems, head-mounted display (HMD) devices, other types of computing devices discussed herein, or combinations thereof.
[0062] The imaging system 200 includes one or more image sensors 205 configured to capture image data 220 (e.g., one or more images or portions thereof) according to one or more image capture settings 210 (e.g., exposure time and / or any other image capture settings discussed herein). In some examples, the image sensors 205 include one or more image sensors or one or more cameras. In some examples, the image data 220 captured using the image sensors 205 according to the image capture settings 210 includes raw image data, image data, pixel data, image frames, raw video data, video data, video frames, or a combination thereof. In some examples, at least one of the image sensors 205 may be oriented toward a user and / or vehicle (e.g., may face the user and / or vehicle) and, therefore, may capture sensor data (e.g., image data) of (e.g., depicting or otherwise representing) at least a portion of the user and / or vehicle. In some examples, at least one of the image sensors 205 may be facing away from the user and / or vehicle (e.g., facing away from the user and / or vehicle) and / or toward the environment in which the user and / or vehicle is located, and thus may capture (e.g., depict or otherwise represent) sensor data (e.g., image data) of at least a portion of the environment. In some examples, the sensor data captured by the at least one of the image sensors 205 that is facing away from the user (and / or vehicle) and / or toward the environment may have a field of view (FoV) that includes, is included in, overlaps with, and / or otherwise corresponds to the FoV of the user's eyes (and / or the FoV of the location of the vehicle).
[0063] exist Figure 2 In FIG, a graphic representing image sensor 205 illustrates image sensor 205 as including a camera facing an environment that includes two people standing and a table with a laptop computer and three chairs surrounding it. Figure 2 In FIG, the graphic representing image data 220 illustrates an image depicting the environment illustrated in the graphic representing image sensor 205. Figure 2 , a graphic representing image capture settings 210 illustrates an hourglass near the camera, which represents exposure time.
[0064] The imaging system 200 also includes one or more motion sensors 215 that capture motion data 270. Motion sensors 215 are sensors that capture data regarding position (e.g., lateral position and / or altitude), orientation (e.g., pitch, yaw, and / or roll), pose (e.g., position and / or orientation), motion (e.g., change in position and / or orientation), rate, velocity, acceleration, or a combination thereof. Motion sensors may include a global navigation satellite system (GNSS) receiver, an inertial measurement unit (IMU), an accelerometer, a gyroscope, a gyrometer, a barometer, an altimeter, or a combination thereof. Because both the motion sensors 215 and the image sensor 205 are part of or coupled to the imaging system 200, the motion data 270 captured by the motion sensors 215 may indicate when the image sensor 205 of the imaging system 200 is stationary, moving, accelerating, decelerating, turning, or a combination thereof. The motion sensors 215 may also be referred to as positioning sensors, acceleration sensors, orientation sensors, pose sensors, or a combination thereof. In some examples, for example, motion data 270 may include angular velocity 1242 from an IMU and / or gyroscope, motion information 1244 and / or scene depth information 1246 from a six degrees of freedom (6DoF) sensor, optical flow data 1248 from a previous image in image data 220 (and / or previous image data captured by image sensor 205 prior to image data 220), or a combination thereof.
[0065] exist Figure 2 In FIG, a graphic representing a motion sensor 215 illustrates a gear and a slider representing motion detection. Figure 2 , a graphic representing motion data 270 illustrates a camera device (representing imaging system 200) surrounded by arrows pointing diagonally downward and rightward, indicating that motion data 270 indicates that the camera device (and therefore imaging system 200) is moving downward and rightward in the direction of the arrows.
[0066] The imaging system 200 includes a feature tracker 225 that detects, extracts, identifies, and / or tracks features 230 in the image data 220. The features 230 that the feature tracker 225 detects, extracts, identifies, and / or tracks in the image data 220 may include, for example, edges, corners, and spots. The feature tracker 225 may generate a descriptor corresponding to the feature 230, which may be a vector-based indicator of the feature 230, such as a scale-invariant feature transform (SIFT) descriptor or a variant thereof. In some examples, the feature tracker 225 may use one or more trained machine learning (ML) models 290 to detect, extract, identify, and / or track the feature 230 in the image data 220. For example, the trained ML model 290 may be trained using training data that includes images (such as the image data 220) having previously identified features (such as the features 230). Figure 2 , the graphic representing feature tracker 225 and features 230 includes a set of white triangles with black outlines, representing features 230 , overlaid on an image representing image data 220 .
[0067] Imaging system 200 includes a motion blur estimator 235 that estimates a level of motion blur in a given image of image data 220 based on image capture settings 210 used to capture the image and / or motion data 270 captured by motion sensor 215 concurrently with (e.g., concurrently with and / or within a threshold amount of time before or after) the image (of image data 220). In some examples, motion blur estimator 235 also estimates a level of motion blur in the given image of image data 220 based on image analysis of the given image of image data 220, such as by using trained ML model 290 to determine a level of motion blur (or, more generally, blur) detected in the image itself. In some examples, motion blur estimator 235 also estimates a level of motion blur in the given image of image data 220 based on estimates of motion blur levels in one or more other images in a series of images in image data 220, such as one or more images preceding the given image in the series and / or one or more images following the given image in the series. The series of images may be video frames of a video, in which case image data 220 may include video frames, video, or both. Figure 2 , a graph representing motion blur estimator 235 is illustrated as including a graph representing motion data 270 and corresponding arrows along the sides of the graph representing image data 220 to indicate the direction of motion blur in image data 220 .
[0068] In some examples, motion blur estimator 235 may also provide feedback to the image capture device including image sensor 205 (e.g., based on motion data 270 from motion sensor 215) to modify image capture settings 210, such as to reduce exposure time, when the estimated motion blur level exceeds a threshold. For example, at fast motion speeds, exposure times exceeding 2 milliseconds (ms) may produce motion blur, so motion blur estimator 235 may reduce the maximum allowed exposure time to 2 ms or less, for example by modifying other image capture settings 210 (e.g., ISO, gain) to increase brightness to compensate for the shorter exposure time. Once motion data 270 indicates that imaging system 200 has slowed below a threshold speed or has come to a standstill, motion blur estimator 235 may allow exposure time to increase again, for example, by increasing the maximum allowed exposure time to above 2 ms, without causing motion blur. In some cases, exposure time may be gradually scaled (e.g., linearly, logarithmically, or according to some other relationship) based on changes to the estimated motion blur level and / or motion data 270.
[0069] Imaging system 200 (e.g., motion blur estimator 235 and / or feature tracker 225) may assign respective weights 240 to different images of image data 220 based on their respective estimated levels of motion blur. For example, imaging system 200 may set weights 240 such that images of image data 220 with low estimated levels of motion blur (because they were captured while imaging system 200 was stationary) may have higher respective weights due to a higher confidence that they are accurate. The higher weights in these images mean that the features 230 extracted, detected, and / or identified by the feature tracker 225 in these images are considered more important, more accurate, more reliable, and / or more heavily characterized (e.g., multiplied by a multiplier greater than 1 and / or offset) in the calculations (e.g., performed by the feature tracker 225) for feature tracking (e.g., performed by the feature tracker 225), environment mapping (e.g., performed by the mapping engine 250), and / or pose estimation (e.g., performed by the pose engine 260) (e.g., performed by the feature tracker 225 and / or the SLAM engine 245). The imaging system 200 can also set the weights 240 so that, conversely, images with image data 220 having a high estimated level of motion blur (because they were captured while the imaging system 200 was in motion) can have lower corresponding weights due to a lower confidence that they are accurate. The lower weights in these images mean that the features 230 extracted, detected, and / or identified in these images by the feature tracker 225 may be ignored, skipped, deleted, or treated as less important, less accurate, less reliable, and / or less heavily characterized (e.g., multiplied by a multiplier less than 1 and / or having an offset subtracted) in calculations (e.g., performed by the feature tracker 225) for feature tracking (e.g., performed by the feature tracker 225), environment mapping (e.g., performed by the mapping engine 250), and / or pose estimation (e.g., performed by the pose engine 260). Figure 2 Within, a graph representing weights 240 is illustrated as a scale to indicate that features 230 from different images are treated differently (e.g., ignored, skipped, have their values multiplied by a multiplier, have an offset applied, etc.) based on whether the image from which those features 230 were detected / extracted / identified was found to have a high or low estimated level of motion blur.
[0070] In some examples, imaging system 200 may further assign different weights 240 to different features within a given image (e.g., of image data 220). If imaging system 200 undergoes translational motion, sometimes only portions of the image captured during that motion are affected by motion blur, or some portions are affected by motion blur more than other portions. For example, typically, portions of a scene farther from image sensor 205 move less within the field of view of image sensor 205 than portions of the scene closer to image sensor 205. Therefore, in some examples, imaging system 200 may assign weights 240 to features 230 of an image based on their respective depths (e.g., as detected using image data 220 and / or depth data from a depth sensor). The depth sensor of imaging system 200 may include, for example, a radio detection and ranging (RADAR) sensor, a light detection and ranging (LIDAR) sensor, a sound detection and ranging (SODAR) sensor, a sound navigation and ranging (SONAR) sensor, a time-of-flight (ToF) sensor, a structured light sensor, a stereo camera, a laser rangefinder, or a combination thereof. In some examples, imaging system 200 may assign weights 240 to features 230 of an image such that portions of the scene farther from imaging system 200 (e.g., greater than a threshold depth) receive higher weights than portions of the scene closer to imaging system 200 (e.g., less than a threshold depth), particularly when the estimated level of motion blur is above a threshold level. If imaging system 200 experiences rotational motion exceeding a threshold, the entire image may be blurred due to motion blur, which may cause imaging system 200 to assign lower weights 240 to features 230 than to translational motion, and / or assign lower weights 240 to features 230 more uniformly than to translational motion (e.g., without much difference between near and far features).
[0071] Imaging system 200 includes a simultaneous localization and mapping (SLAM) engine 245 that receives features 230 and corresponding weights 240 from feature tracker 225 and / or motion blur estimator 235, and in some cases, image data 220 from image sensor 205 and / or motion data 270 from motion sensor 215. SLAM engine 245 includes a mapping engine 250 that generates a map 255 of the environment depicted in image data 220 based on tracking features 230 as imaging system 200 moves through different locations in the environment, wherein features 230 corresponding to high estimated motion blur levels are ignored or play a smaller role in the mapping calculations based on weights 240. In some examples, SLAM engine 245 and / or mapping engine 250 can use a trained ML model 290 to generate map 255 based on features 230 and weights 240, and in some cases, also based on image data 220 and / or motion data 270. For example, the trained ML model 290 may be trained using training data that includes a pre-generated map of an environment as well as an image, features extracted from the image, weights for the features, and / or motion data from a device that captured the image. Figure 2 , a graphic representing mapping engine 250 and map 255 is illustrated as a top-down view of the environment depicted in the graphic representing image data 220 , in which a table, a chair, a laptop computer, and two people are visible.
[0072] The SLAM engine 245 of the imaging system 200 also includes a pose engine 260 that estimates a pose 265 of the imaging system 200 (e.g., the image sensor 205 and / or the motion sensor 215), for example, the pose 265 of the imaging system 200 relative to the environment (and / or its map 255), based on the features 230 and weights 240, and in some cases, the image data 220 and / or the motion data 270. The pose engine 260 can ignore features 230 corresponding to high estimated motion blur levels or give these features 230 a smaller role in the pose estimation calculation based on the weights 240. In some examples, the SLAM engine 245 and / or the pose engine 260 can use a trained ML model 290 to estimate the pose 265 based on the features 230 and weights 240, and in some cases, the image data 220 and / or the motion data 270. For example, the trained ML model 290 may be trained using training data that includes a predetermined pose of an environment and an image of the environment captured by a device at the predetermined pose, features extracted from the image, weights for the features, and / or motion data from the device that captured the image. Figure 2Within, a graphic representing pose engine 260 and pose 265 is illustrated as arrows overlaid on a graphic representing map 255 that represents the location of image sensor 205 to capture image data 220 as depicted in the graphic representing image data 220 .
[0073] In some examples, imaging system 200 includes an output processor 275 that generates output data 280 based on map 255, pose 265, features 230, weights 240, image data 220, motion data 270, or a combination thereof. In some examples, output data 280 includes map 255, pose 265, features 230, weights 240, image data 220, motion data 270, or a combination thereof. In some examples, output data 280 includes information or images derived from or otherwise based on map 255, pose 265, features 230, weights 240, image data 220, motion data 270, or a combination thereof. For example, in some examples, output data 280 includes a route through map 255 along a path including pose 265 (e.g., from pose 265 to a destination) (and / or directions for following the route), an augmented reality (AR) view based on map 255 and / or pose 265 and / or image data 220, a virtual reality (VR) view based on map 255 and / or pose 265 and / or image data 220, a mixed reality (MR) view based on map 255 and / or pose 265 and / or image data 220, an extended reality (XR) view based on map 255 and / or pose 265 and / or image data 220, or a combination thereof.
[0074] In some examples, output processor 275 can use trained ML model 290 to generate output data 280 based on map 255, pose 265, features 230, weights 240, image data 220, and / or motion data 270. For example, trained ML model 290 can be trained using training data that includes pre-generated output data and corresponding maps, poses, features, weights, image data, and / or motion data. Figure 2 , a graphic representing output processor 275 and output data 280 is illustrated as a route overlaid on a graphic representing map 255 , the route including the positions of arrows representing poses, the graphic representing pose 265 .
[0075] Imaging system 200 includes one or more output devices 285 configured to output output data 280 and / or poses 265 and / or map 255. Output device 285 may include one or more visual output devices, such as displays or connectors therefor. Output device 285 may include one or more audio output devices, such as speakers, headphones, and / or connectors therefor. Output device 285 may include one or more of output device 1535 and / or communication interface 1540 of computing system 1500. In some examples, imaging system 200 causes a display of output device 285 to display output data 280 and / or poses 265 and / or map 255.
[0076] In some examples, output device 285 includes one or more transceivers. A transceiver may include a wired transmitter, receiver, transceiver, or a combination thereof. A transceiver may include a wireless transmitter, receiver, transceiver, or a combination thereof. The transceiver may include one or more of output device 1535 and / or communication interface 1540 of computing system 1500. In some examples, imaging system 200 causes the transceiver to transmit output data 280 and / or pose 265 and / or map 255 to a recipient device. In some examples, the recipient device may include HMD 310, mobile phone 410, vehicle 510, computing system 1500, or a combination thereof. In some examples, the recipient device may include a display, and data transmitted from the transceiver of output device 285 to the recipient device may cause the display of the recipient device to display output data 280 and / or pose 265 and / or map 255.
[0077] In some examples, the display of output device 285 of imaging system 200 functions as an optical "see-through" display that allows light from the real-world environment (scene) surrounding imaging system 200 to pass across (e.g., through) the display of output device 285 to one or both eyes of the user. For example, the display of output device 285 may be at least partially transparent, translucent, light-transmissive, light-transmissive, or a combination thereof. In an illustrative example, the display of output device 285 includes a transparent, translucent, and / or light-transmissive lens and a projector. The display of output device 285 may include a projector that projects virtual content (e.g., output data 280 and / or poses 265 and / or map 255) onto the lens. The lens may be, for example, a lens of a pair of glasses, a lens of goggles, a contact lens, a lens of a head-mounted display (HMD) device, or a combination thereof. Light from the real-world environment passes through the lens and reaches one or both eyes of the user. The projector can project virtual content (e.g., output data 280 and / or pose 265 and / or map 255) onto the lens so that the virtual content appears overlaid on the user's view of the environment from the perspective of one or both of the user's eyes. In some examples, the projector can project the virtual content onto one or both retinas of one or both of the user's eyes instead of onto the lens, which can be referred to as a virtual retinal display (VRD), a retinal scanning display (RSD), or a retinal projector (RP) display.
[0078] In some examples, the display of the output device 285 of the imaging system 200 is a digital "pass-through" display that allows a user of the imaging system 200 and / or a recipient device to view a view of the environment by displaying the view of the environment on the display of the output device 285. The view of the environment displayed on the digital pass-through display can be a view of the real-world environment surrounding the imaging system 200, for example, based on sensor data (e.g., images, videos, depth images, point clouds, other depth data, or a combination thereof) captured by one or more environment-facing sensors in the image sensors 205 (e.g., output data 280 and / or poses 265 and / or map 255). The view of the environment displayed on the digital pass-through display can be a virtual environment (e.g., as in VR), which in some cases may include elements based on the real-world environment (e.g., the boundaries of a room). The view of the environment displayed on the digital pass-through display can be an augmented environment based on the real-world environment (e.g., as in AR). The view of the environment displayed on the digital pass-through display can be a hybrid environment based on the real-world environment (e.g., as in MR). The view of the environment displayed on the digital pass-through display may include virtual content (eg, output data 280 and / or poses 265 and / or map 255 ) overlaid on or otherwise incorporated into the view of the environment.
[0079] exist Figure 2 Within, a graphic illustrating a display, a speaker, a wireless transceiver, and a vehicle representing an output device 285 is used to output a graphic representing output data 280 and / or a pose 265 and / or a map 255 using the display, the speaker, the wireless transceiver, and / or a system associated with the vehicle (e.g., a control computing system such as an ADAS for the vehicle, an IVI system for the vehicle, a control system for the vehicle, or a combination thereof).
[0080] The trained ML model 290 may include one or more neural networks (NNs) (e.g., neural network 1300), one or more convolutional neural networks (CNNs), one or more trained time delay neural networks (TDNNs), one or more deep networks, one or more autoencoders, one or more deep belief networks (DBNs), one or more recurrent neural networks (RNNs), one or more generative adversarial networks (GANs), one or more conditional generative adversarial networks (cGANs), one or more other types of neural networks, one or more trained support vector machines (SVMs), one or more trained random forests (RFs), one or more computer vision systems, one or more deep learning systems, one or more classifiers, one or more transformers, or a combination thereof. Figure 2, a graphic representing a trained ML model 290 illustrates a set of circles connected to another set of circles. Each of these circles can represent a node (e.g., node 1316), a neuron, a perceptron, a layer, a portion thereof, or a combination thereof. The circles are arranged in columns. The white circles in the leftmost column represent the input layer (e.g., input layer 1310). The white circles in the rightmost column represent the output layer (e.g., output layer 1314). The two columns of shaded circles between the leftmost and rightmost columns of white circles each represent a hidden layer (e.g., hidden layers 1312A to 1312N).
[0081] In some examples, the imaging system 200 includes a feedback engine 295 for the imaging system 200. The feedback engine 295 can detect feedback received from a user interface of the imaging system 200. The feedback can include feedback regarding the output of the output device 285 (e.g., output data 280 and / or pose 265 and / or map 255). The feedback engine 295 can detect feedback received from another subsystem of the imaging system 200 regarding one subsystem of the imaging system 200, such as whether one subsystem decides to use data from another subsystem. For example, the feedback engine 295 can detect whether the SLAM engine 245 decides to use features 230 and / or weights 240 generated by the feature tracker 225 and / or the motion blur estimator 235 based on whether the features 230 and / or weights 240 meet the needs of the SLAM engine 245 to generate the map 255 and / or pose 265, and can provide feedback regarding the functionality of the trained ML model 290 used by the SLAM engine 245 to generate the map 255 and / or pose 265. Similarly, the feedback engine 295 can detect whether the output processor 275 decides to use the map 255 and / or pose 265 generated by the SLAM engine 245 based on whether the map 255 and / or pose 265 meet the needs of the output processor 275 to generate the output data 280, and can provide feedback on the functionality of the trained ML model 290 used by the output processor 275 to generate the output data 280. Similarly, the feedback engine 295 can detect whether the output device 285 decides to output the output data 280 and / or pose 265 and / or map 255 generated by the output processor 275 and / or the SLAM engine 245 based on whether the output data 265 and / or pose 265 and / or map 255 meet the needs of the output device 285 for output, and can provide feedback on the functionality of the trained ML model 290 used by the SLAM engine 245 and / or the output processor 275 to generate the pose 280 and / or map 255 and / or output data 280.
[0082] Feedback received by the feedback engine 295 can be positive feedback or negative feedback. For example, if one subsystem of the imaging system 200 uses data from another subsystem of the imaging system 200, or if positive feedback is received from the user through the user interface or from one of the subsystems, the feedback engine 295 can interpret this as positive feedback. If one subsystem of the imaging system 200 refuses to use data from another subsystem of the imaging system 200, or if negative feedback is received from the user through the user interface or from one of the subsystems, the feedback engine 295 can interpret this as negative feedback. Positive feedback can also be based on attributes of sensor data from sensors of the imaging system 200 (e.g., image sensor 205, motion sensor 215, and / or other sensors, such as microphones or depth sensors), such as detecting that the user smiles, laughs, nods, utters a positive statement (e.g., "yes," "confirm," "ok," "next," "confirm," "approve," "I like this"), or otherwise reacts positively to an output of one of the subsystems described herein or its indication. Negative feedback may also be based on attributes of the sensor data from the image sensor 205, such as the user frowning, crying, shaking their head (e.g., making a "no" gesture), uttering a negative statement (e.g., "no," "rejection," "no," "not this," "I hate this," "this is useless," "this is not what I want"), or otherwise reacting negatively to an output of one of the subsystems described herein or an indication thereof.
[0083] In some examples, the feedback engine 295 provides the feedback as training data to the trained ML model 290 and / or one or more subsystems of the imaging system 200 that use the trained ML model 290 (e.g., the feature tracker 225, the motion blur estimator 235, the SLAM engine 245, the mapping engine 250, the pose engine 260, the output processor 275, and / or the trained ML model 290) to update the one or more trained ML models 290 of the imaging system 200. Positive feedback can be used to strengthen and / or enhance weights associated with the output of the ML system and / or the trained ML model 290, and / or weaken or remove weights other than the weights associated with the output of the ML system and / or the trained ML model 290. Negative feedback can be used to weaken and / or remove weights associated with the output of the ML system and / or the trained ML model 290, and / or strengthen and / or enhance weights other than the weights associated with the output of the ML system and / or the trained ML model 290.
[0084] It should be understood that references herein to the image sensor 205 and other sensors described herein (as image sensors) should be understood to also include other types of sensors that can produce output in the form of images, such as depth sensors (e.g., RADAR, LIDAR, SONAR, SODAR, ToF, structured light) that produce depth maps, depth images, and / or point clouds that can be represented in the form of images (e.g., semi-dense point clouds), and / or rendered images of 3D models. It should be understood that references herein to image data and / or images produced by such sensors can include any sensor data that can be output in the form of images, such as depth maps, depth images, point clouds that can be represented in the form of images (e.g., semi-dense point clouds), and / or rendered images of 3D models.
[0085] In some examples, certain elements of imaging system 200 (e.g., image sensor 205, motion sensor 215, feature tracker 225, motion blur estimator 235, SLAM engine 245, mapping engine 250, pose engine 260, output processor 275, trained ML model 290, feedback engine 295, or a combination thereof) include software elements, such as an instruction set corresponding to a program (e.g., a hardware driver, a user interface (UI), an application programming interface (API), an operating system (OS), etc.) that runs on a processor (e.g., processor 1510 of computing system 1500, image processor 150, host processor 152, ISP 154, a microcontroller, a controller, or a combination thereof). In some examples, one or more of these elements of imaging system 200 may include one or more hardware elements, such as a dedicated processor (e.g., processor 1510 of computing system 1500, image processor 150, host processor 152, ISP 154, a microprocessor, a controller, or a combination thereof). In some examples, one or more of these elements of imaging system 200 may include a combination of one or more software elements and one or more hardware elements.
[0086] Figure 3A300 is a perspective view illustrating a head-mounted display (HMD) 310 used as part of imaging system 200. HMD 310 can be, for example, an augmented reality (AR) headset, a virtual reality (VR) headset, a mixed reality (MR) headset, an extended reality (XR) headset, or some combination thereof. HMD 310 can be an example of imaging system 200. HMD 310 includes a first camera 330A and a second camera 330B along the front of HMD 310. First camera 330A and second camera 330B can be examples of image sensor 205 of imaging system 200. HMD 310 includes a third camera 330C and a fourth camera 330D that face the user's eyes when the user's eyes face display 340. Third camera 330C and fourth camera 330D can be examples of image sensor 205 of imaging system 200. In some examples, HMD 310 may have only a single camera with a single image sensor. In some examples, the HMD 310 may include one or more additional cameras in addition to the first camera 330A, the second camera 330B, the third camera 330C, and the fourth camera 330D. In some examples, the HMD 310 may include one or more additional sensors in addition to the first camera 330A, the second camera 330B, the third camera 330C, and the fourth camera 330D, which may also include other types of image sensors 205 of the imaging system 200. In some examples, the first camera 330A, the second camera 330B, the third camera 330C, and / or the fourth camera 330D may be examples of the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof. In some examples, any of the first camera 330A, the second camera 330B, the third camera 330C, and / or the fourth camera 330D may be or may include a depth sensor.
[0087] HMD 310 may include one or more displays 340 visible to user 320 (who is wearing HMD 310 on their head). The one or more displays 340 of HMD 310 may be examples of one or more displays of output device 285 of imaging system 200. In some examples, HMD 310 may include one display 340 and two viewfinders. The two viewfinders may include a left viewfinder for user 320's left eye and a right viewfinder for user 320's right eye. The left viewfinder may be oriented so that user 320's left eye sees the left side of the display. The right viewfinder may be oriented so that user 320's right eye sees the right side of the display. In some examples, HMD 310 may include two displays 340, including a left display that displays content to user 320's left eye and a right display that displays content to user 320's right eye. The one or more displays 340 of HMD 310 may be digital "pass-through" displays or optical "see-through" displays.
[0088] The HMD 310 may include one or more earpieces 335, which may function as speakers and / or headphones that output audio to one or both ears of a user of the HMD 310 and may be an example of an output device 285. Figure 3A and Figure 3B 335, it should be understood that the HMD 310 may include two earpieces, one for each ear of the user (left and right). In some examples, the HMD 310 may also include one or more microphones (not shown). The one or more microphones may be examples of the image sensor 205 of the imaging system 200. In some examples, the audio output by the HMD 310 to the user through the one or more earpieces 335 may include or be based on audio recorded using the one or more microphones.
[0089] Figure 3B This is an example Figure 3A310 is a perspective view 350 of a head-mounted display (HMD) being worn by user 320. User 320 wears HMD 310 on user 320's head, above user 320's eyes. HMD 310 may capture images using first camera 330A and second camera 330B. In some examples, HMD 310 displays one or more output images toward user 320's eyes using display 340. In some examples, the output images may include output data 280 and / or pose 265 and / or map 255. The output images may be based on the images captured by first camera 330A and second camera 330B (e.g., image data 220), for example, overlaid with virtual content (e.g., output data 280 and / or pose 265 and / or map 255). The output images may provide a stereoscopic view of the environment, in some cases overlaid with virtual content and / or with other modifications. For example, HMD 310 may display a first display image based on an image captured by first camera 330A to the right eye of user 320. HMD 310 may display a second display image based on an image captured by second camera 330B to the left eye of user 320. For example, HMD 310 may provide overlay virtual content in the display image overlaid on the images captured by first camera 330A and second camera 330B. Third camera 330C and fourth camera 330D may capture images of the user's eyes before, during, and / or after viewing the display image displayed by display 340. In this way, sensor data from third camera 330C and / or fourth camera 330D may capture the user's eye (and / or other parts of the user) reacting to the virtual content. Earpiece 335 of HMD 310 is illustrated as being in the ear of user 320. The HMD 310 may output audio to the user 320 through the earpiece 335 and / or through another earpiece (not shown) of the HMD 310 in the other ear of the user 320 (not shown).
[0090] Figure 4A is a perspective view 400 illustrating a front surface of a mobile phone 410 that includes a front-facing camera and can be used as part of the imaging system 200. The mobile phone 410 can be an example of the imaging system 200. The mobile phone 410 can be, for example, a cellular phone, a satellite phone, a portable game console, a music player, a fitness tracking device, a wearable device, a wireless communication device, a laptop computer, a mobile device, any other type of computing device or computing system discussed herein, or a combination thereof.
[0091] The front surface 420 of the mobile phone 410 includes a display 440. The front surface 420 of the mobile phone 410 includes a first camera 430A and a second camera 430B. The first camera 430A and the second camera 430B can be examples of the image sensor 205 of the imaging system 200. The first camera 430A and the second camera 430B can face the user, including the user's eyes, while content (e.g., output data 280 and / or pose 265 and / or map 255) is displayed on the display 440. The display 440 can be an example of a display of the output device 285 of the imaging system 200.
[0092] A first camera 430A and a second camera 430B are illustrated in a bezel surrounding a display 440 on the front surface 420 of a mobile handset 410. In some examples, the first camera 430A and the second camera 430B may be positioned in a notch or cutout cut out of the display 440 on the front surface 420 of the mobile handset 410. In some examples, the first camera 430A and the second camera 430B may be under-display cameras positioned between the display 440 and the rest of the mobile handset 410 so that light passes through a portion of the display 440 before reaching the first camera 430A and the second camera 430B. The first camera 430A and the second camera 430B of the perspective view 400 are front-facing cameras. The first camera 430A and the second camera 430B face a direction perpendicular to the planar surface of the front surface 420 of the mobile handset 410. The first camera 430A and the second camera 430B may be two of one or more cameras of the mobile handset 410. In some examples, the front surface 420 of the mobile handset 410 may have only a single camera.
[0093] In some examples, display 440 of mobile phone 410 displays one or more output images to a user using mobile phone 410. In some examples, the output images may include output data 280 and / or pose 265 and / or map 255. The output images may be based on images (e.g., image data 220) captured by first camera 430A, second camera 430B, third camera 430C, and / or fourth camera 430D, for example, overlaid with virtual content (e.g., output data 280 and / or pose 265 and / or map 255).
[0094] In some examples, the front surface 420 of the mobile handset 410 may include one or more additional cameras in addition to the first camera 430A and the second camera 430B. The one or more additional cameras may also be examples of the image sensor 205 of the imaging system 200. In some examples, the front surface 420 of the mobile handset 410 may include one or more additional sensors in addition to the first camera 430A and the second camera 430B. The one or more additional sensors may also be examples of the image sensor 205 of the imaging system 200. In some cases, the front surface 420 of the mobile handset 410 includes more than one display 440. The one or more displays 440 of the front surface 420 of the mobile handset 410 may be examples of displays of the output device 285 of the imaging system 200. For example, the one or more displays 440 may include one or more touch screen displays.
[0095] The mobile handset 410 may include one or more speakers 435A and / or other audio output devices (eg, earphones or headphones or connectors thereof) that may output audio to one or more ears of a user of the mobile handset 410 . Figure 4A 4. Although one speaker 435A is illustrated in the mobile handset 410, it should be understood that the mobile handset 410 may include more than one speaker and / or other audio device. In some examples, the mobile handset 410 may also include one or more microphones (not shown). The one or more microphones may be examples of the image sensor 205 of the imaging system 200. In some examples, the mobile handset 410 may include one or more microphones along and / or adjacent to the front surface 420 of the mobile handset 410, where these microphones are examples of the image sensor 205 of the imaging system 200. In some examples, the audio output to the user by the mobile handset 410 via the one or more speakers 435A and / or other audio output devices may include or be based on audio recorded using the one or more microphones.
[0096] Figure 4B is a perspective view 450 illustrating a rear surface 460 of a mobile phone that includes a rear-facing camera and can be used as part of the imaging system 200. The mobile phone 410 includes a third camera 430C and a fourth camera 430D on the rear surface 460 of the mobile phone 410. The third camera 430C and the fourth camera 430D of the perspective view 450 are rear-facing. The third camera 430C and the fourth camera 430D can be Figure 2 4. An example of the image sensor 205 of the imaging system 200. The third camera 430C and the fourth camera 430D face a direction perpendicular to the planar surface of the rear surface 460 of the mobile phone 410.
[0097] Third camera 430C and fourth camera 430D may be two of the one or more cameras of mobile phone 410. In some examples, rear surface 460 of mobile phone 410 may have only a single camera. In some examples, rear surface 460 of mobile phone 410 may include one or more additional cameras in addition to third camera 430C and fourth camera 430D. The one or more additional cameras may also be examples of image sensor 205 of imaging system 200. In some examples, rear surface 460 of mobile phone 410 may include one or more additional sensors in addition to third camera 430C and fourth camera 430D. The one or more additional sensors may also be examples of image sensor 205 of imaging system 200. In some examples, first camera 430A, second camera 430B, third camera 430C, and / or fourth camera 430D may be examples of image capture and processing system 100, image capture device 105A, image processing device 105B, or a combination thereof. In some examples, any of first camera 430A, second camera 430B, third camera 430C, and / or fourth camera 430D may be or may include a depth sensor.
[0098] The mobile handset 410 may include one or more speakers 435B and / or other audio output devices (eg, earphones or headphones or connectors thereof) that may output audio to one or more ears of a user of the mobile handset 410 . Figure 4B 4. Although one speaker 435B is illustrated in the mobile handset 410, it should be understood that the mobile handset 410 may include more than one speaker and / or other audio device. In some examples, the mobile handset 410 may also include one or more microphones (not shown). The one or more microphones may be examples of the image sensor 205 of the imaging system 200. In some examples, the mobile handset 410 may include one or more microphones along and / or adjacent to the rear surface 460 of the mobile handset 410, where these microphones are examples of the image sensor 205 of the imaging system 200. In some examples, the audio output to the user by the mobile handset 410 via the one or more speakers 435B and / or other audio output devices may include or be based on audio recorded using the one or more microphones.
[0099] Mobile phone 410 can use display 440 on front surface 420 as a pass-through display. For example, display 440 can display an output image, such as output data 280 and / or position 265 and / or map 255. The output image can be based on an image (e.g., image data 220) captured by third camera 430C and / or fourth camera 430D, for example, overlaid with virtual content (e.g., output data 280 and / or position 265 and / or map 255). First camera 430A and / or second camera 430B can capture images of the user's eyes (and / or other parts of the user) before, during, and / or after the output image with virtual content is displayed on display 440. In this way, sensor data from first camera 430A and / or second camera 430B can capture the user's eyes (and / or other parts of the user) reacting to the virtual content.
[0100] Figure 5 5 is a perspective view 500 illustrating a vehicle 510 including various sensors. Vehicle 510 can be an example of imaging system 200. Vehicle 510 is illustrated as an automobile, but can be, for example, an automobile, a truck, a bus, a train, a ground vehicle, an airplane, a helicopter, an aircraft, an air vehicle, a boat, a submarine, a water vehicle, an underwater vehicle, a hovercraft, other types of vehicles discussed herein, or combinations thereof. In some examples, the vehicle can be controlled and / or used at least in part with a subsystem of vehicle 510, such as an ADAS system of vehicle 510, an IVI system of vehicle 510, a control system of vehicle 510, a vehicle electronic control unit (ECU) 630 of vehicle 510, or combinations thereof.
[0101] Vehicle 510 includes a display 520. Vehicle 510 includes various sensors, all of which can be examples of image sensor 205. Vehicle 510 includes a first camera 530A and a second camera 530B located at the front, a third camera 530C and a fourth camera 530D located at the rear, and a fifth camera 530E and a sixth camera 530F located at the top. Vehicle 510 includes a first microphone 535A located at the front, a second microphone 535B located at the rear, and a third microphone 535C located at the top. Vehicle 510 includes a first sensor 540A located on one side (e.g., adjacent to one rearview mirror) and a second sensor 540B located on the other side (e.g., adjacent to the other rearview mirror). The first sensor 540A and the second sensor 540B may include a camera, a microphone, a depth sensor (e.g., a radio detection and ranging (RADAR) sensor, a light detection and ranging (LIDAR) sensor, a sound detection and ranging (SODAR) sensor, a sound navigation and ranging (SONAR) sensor, a time of flight (ToF) sensor, a structured light sensor, a stereo camera, etc.), or any other type of sensor 205 described herein. In some examples, in addition to Figure 5 In addition to the illustrated sensors, the vehicle 510 may also include additional image sensors 205. In some examples, the vehicle 510 may lack Figure 5 Some of the sensors illustrated.
[0102] In some examples, the display 520 of the vehicle 510 displays one or more output images to a user of the vehicle 510 (e.g., the driver of the vehicle 510 and / or one or more passengers). In some examples, the output images may include the output data 280 and / or the pose 265 and / or the map 255. The output images may be based on images (e.g., image data 220) captured by the first camera 530A, the second camera 530B, the third camera 530C, the fourth camera 530D, the fifth camera 530E, the sixth camera 530F, the first sensor 540A, and / or the second sensor 540B, for example, overlaid with virtual content (e.g., output data 280 and / or the pose 265 and / or the map 255). In some examples, any of the first camera 530A, the second camera 530B, the third camera 530C, the fourth camera 530D, the fifth camera 530E, the sixth camera 530F, the first sensor 540A, and / or the second sensor 540B may be or may include a depth sensor.
[0103] Figure 6is a conceptual diagram illustrating feature tracking in image 600 with motion blur. The image is affected by motion blur due to the motion of the camera, making the entire scene appear blurry. Features, such as feature 230, are illustrated as overlaid on image 600 as white triangles with black outlines. Certain features identified and tracked from previous frames are shown with a second shaded triangle with a black outline, and the line between the two triangles indicates the direction and distance the detected feature moved. Time is written near certain features, indicating how long those features were tracked in the image sequence (e.g., video). Because image 600 is affected by motion blur, features in image 600 appear to move in various directions and over various distances due to the uncertainty introduced by motion blur. Therefore, images that are strongly affected by motion blur, such as image 600, can negatively impact feature tracking, making features extracted, detected, identified, and / or tracked (e.g., using feature tracker 225) unreliable for feature tracking purposes. This, in turn, can negatively impact environment mapping (e.g., performed using mapping engine 250), pose estimation (e.g., performed using pose engine 260), other SLAM functionality (e.g., performed using SLAM engine 245), other output generation functionality (e.g., performed using output processor 275), and / or other functionality of imaging system 200. Accordingly, imaging system 200 that compensates for motion blur by giving less weight to features 230 extracted from such images (e.g., using weights 240 generated using motion blur estimator 235) can be more accurate and reliable in feature tracking (e.g., performed using feature tracker 225), environment mapping (e.g., performed using mapping engine 250), pose estimation (e.g., performed using pose engine 260), other SLAM functionality (e.g., performed using SLAM engine 245), other output generation functionality (e.g., performed using output processor 275), and / or other functionality of imaging system 200.
[0104] In addition, conventional automatic exposure control (AEC) systems control exposure based on image statistics (e.g., average brightness, luminosity, and / or brightness) of previously captured images and do not take motion data into account. An imaging system 200 (e.g., using motion blur estimator 235, AEC engine 805, and / or AEC engine 1205) that uses motion data 270 from motion sensor 215 to also influence the determination of image capture settings 210 (such as exposure time) can also reduce instances of images with motion blur (such as image 600) by first reducing the exposure time of images captured when imaging system 200 is moving (e.g., exceeding a threshold speed or acceleration) (e.g., instead increasing the gain of those images) and increasing the exposure time back to a normal image statistics-based level for images captured when imaging system 200 is stationary (or moving less than a threshold speed).
[0105] Figure 7 7 is a block diagram illustrating a process 700 for pose estimation that takes into account exposure time. The process 700 for imaging can be performed by an imaging system (e.g., a chipset, a processor or multiple processors (such as an ISP, HP, or other processors), or other components). In some examples, the imaging system may include, for example, the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the imaging system 200, the image sensor 205, the motion sensor 215, the feature tracker 225, the motion blur estimator 235, the SLAM engine 245, the mapping engine 250, the pose engine 260, the output processor 275 and / or the trained ML model 290, the trained ML model 290, the feedback engine 295, the HMD 310, mobile phone 410, vehicle 510, imaging system that performs any process described herein (e.g., process 800, process 900, process 1000, process 1200, process 1400), camera 1105, neural network 1300, computing system 1500, processor 1510, system and apparatus, device, non-transitory computer-readable medium having stored thereon a program to be executed using a processor, or a combination thereof.
[0106] The imaging system receives an image 705 (e.g., image data 220) and uses the image 705 to perform feature tracking 715 (e.g., using feature tracker 225). The imaging system uses the feature tracking 715 for pose estimation 720 to estimate a pose 725 of the imaging system, for example, using triangulation to estimate the pose 725. The imaging system may also receive an exposure time 710 setting used to capture the image 705, which the imaging system may use to determine a confidence level 730 associated with the pose 725 estimated via the pose estimation 720. For example, a higher exposure time may result in a lower confidence level 730 due to an increased likelihood of motion blur, while a lower exposure time may result in a higher confidence level 730 due to a decreased likelihood of motion blur.
[0107] The imaging system also receives accelerometer data 735 and / or gyroscope data 740 from the motion sensor 215 as motion data 270. The imaging system uses the accelerometer data 735 and / or gyroscope data 740 to perform state propagation 745 to determine an IMU propagation state 750. The imaging system includes a pose estimator 755 that receives the pose 725, the confidence 730, and the IMU propagation state 750 to determine an output pose 760. The pose estimator 755 can use a Kalman filter, an extended Kalman filter, another pose estimation function, or a combination thereof. The pose estimation 720 and / or the pose estimator 755 can be examples of portions of the pose engine 260. The pose 725 and / or the output pose 760 can be examples of the pose 265.
[0108] Figure 88 is a block diagram illustrating a process 800 for automatic exposure control that takes motion data into account. The process 800 for imaging can be performed by an imaging system (e.g., a chipset, a processor or multiple processors (such as an ISP, HP, or other processor), or other components). In some examples, the imaging system may include, for example, the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the imaging system 200, the image sensor 205, the motion sensor 215, the feature tracker 225, the motion blur estimator 235, the SLAM engine 245, the mapping engine 250, the pose engine 260, the output processor 275 and / or the trained ML model 290, the trained ML model 290, the feedback engine 295, the HMD 310, mobile phone 410, vehicle 510, imaging system that performs any process described herein (e.g., process 700, process 900, process 1000, process 1200, process 1400), camera 1105, neural network 1300, computing system 1500, processor 1510, system and apparatus, device, non-transitory computer-readable medium having stored thereon a program to be executed using a processor, or a combination thereof.
[0109] The imaging system includes an AEC engine 805 having an exposure / gain table 810 that identifies exposure time settings and gain settings to be used by a camera 830. The AEC engine 805 may output image capture settings 820 (e.g., including exposure time and / or gain) (e.g., image capture settings 210) to a camera 830 (e.g., image sensor 205). A delay 825 (e.g., N frames when the camera 830 is capturing video) may occur before the camera 830 applies the image capture settings 820. After the delay 825, the camera 830 captures a raw image 835 with the image capture settings 820 applied. The AEC engine 805 may determine an adjustment 815 (e.g., an increase or decrease) to the maximum exposure time based on image statistics of the raw image 835 to generate a new maximum exposure time 840. The adjustment 815 may be made to match the image luminance of the raw image 835 (e.g., the average luminance of the raw image 835) to a target luminance. The imaging system uses the new maximum exposure time 840 and a motion-aware AEC setting 845 based on the motion data 270 received from the motion sensor 215 to determine updates 850 to the exposure / gain table 810. For example, if the imaging system is determined to be in motion (e.g., translation and / or rotation speed and / or acceleration exceeds a threshold), the motion-aware AEC setting 845 may adjust the new maximum exposure time 840 downward, or if the imaging system is determined to be stationary (e.g., translation and / or rotation speed and / or acceleration is below a threshold), the motion-aware AEC setting may leave the new maximum exposure time 840 at its current position. The updated exposure / gain table 810 may then be used to determine additional image capture settings 820 to be used by the camera 830 to capture additional raw images 835, and so on. The AEC engine 805 and / or other image processing subsystems of the imaging system may process each raw image 835 to produce an image 855, which may be used by the feature tracker 860 (e.g., feature tracker 225, feature tracking 715).
[0110] Figure 9is a flow diagram illustrating a process 900 for pose estimation that takes into account estimated motion blur. The process 900 for imaging can be performed by an imaging system (e.g., a chipset, a processor or multiple processors (such as an ISP, HP, or other processors), or other components). In some examples, the imaging system may include, for example, the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the imaging system 200, the image sensor 205, the motion sensor 215, the feature tracker 225, the motion blur estimator 235, the SLAM engine 245, the mapping engine 250, the pose engine 260, the output processor 275 and / or the trained ML model 290, the trained ML model 290, the feedback engine 295, the HMD 310, mobile phone 410, vehicle 510, imaging system that performs any process described herein (e.g., process 700, process 800, process 1000, process 1200, process 1400), camera 1105, neural network 1300, computing system 1500, processor 1510, system and apparatus, device, non-transitory computer-readable medium having stored thereon a program to be executed using a processor, or a combination thereof.
[0111] At operation 905, the imaging system (or components thereof) is configured to and may determine initial poses, 3D points, and their correspondences (observations) in the image. At operation 910, the imaging system (or components thereof) is configured to and may determine exposure time and motion information from 6DoF (e.g., from the AEC engine, motion sensor 215, and / or optical flow).
[0112] At operation 915, the imaging system (or its components) is configured to estimate motion blur of the feature points based on exposure time, depth, and motion. The imaging system may recalculate the measurement variance of all tracking points with motion blur. At operation 920, the imaging system (or its components) is configured to generate pose estimation results 925 based on the data from operations 905 and 915 using an optimization framework. The pose estimation results 925 may include pose measurements and / or covariance of the Kalman filter.
[0113] Figure 101 is a flow diagram illustrating a process 1000 for pose estimation that takes into account estimated motion blur. The process 1000 for imaging can be performed by an imaging system (e.g., a chipset, a processor or multiple processors (such as an ISP, HP, or other processors), or other components). In some examples, the imaging system may include, for example, the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the imaging system 200, the image sensor 205, the motion sensor 215, the feature tracker 225, the motion blur estimator 235, the SLAM engine 245, the mapping engine 250, the pose engine 260, the output processor 275 and / or the trained ML model 290, the trained ML model 290, the feedback engine 295, the HMD 310, mobile phone 410, vehicle 510, imaging system that performs any process described herein (e.g., process 700, process 800, process 900, process 1200, process 1400), camera 1105, neural network 1300, computing system 1500, processor 1510, system and apparatus, device, non-transitory computer-readable medium having stored thereon a program to be executed using a processor, or a combination thereof.
[0114] At operation 1005, the imaging system (or its components) is configured to and may determine initial poses, 3D points and their correspondences (observations) in the image. At operation 1010, the imaging system (or its components) is configured to and may calculate the reprojection error in the reprojection (e.g., see Figure 11 At operation 1015 , the imaging system (or components thereof) is configured to and may calculate weights (eg, weights 240 ) based on feature tracking statistics (such as reprojection and feature tracking quality).
[0115] At operation 1020 , the imaging system (or components thereof) is configured to and may determine exposure time and motion information from 6DoF (eg, from the AEC engine, motion sensor 215 , and / or optical flow) and / or from an Extended Kalman Filter (EKF).
[0116] At operation 1025, the imaging system (or a component thereof) is configured to and may estimate motion blur of the feature point based on exposure time, depth, and motion. The imaging system may recalculate the weights based on the motion blur using a weighting function (provided below in Equation 1) as the calculated weights:
[0117]
[0118] In Equation 1, a and b are constants. r represents the motion blur amplitude. δ represents the threshold motion blur amplitude, e.g., 1 pixel or less. In some examples, the threshold is more than one pixel (e.g., 2 pixels, 3 pixels, etc.). A weighting function (Equation 1) is used to account for the additional error variance due to motion blur.
[0119] At operation 1030, the imaging system (or a component thereof) is configured to and may estimate a pose (e.g., using the recalculated weights of operation 1025) so as to minimize a weighted least squares reprojection error (e.g., the reprojection error of operation 1010). The imaging system may loop back to operation 1010, each time checking at decision point 1035 whether a maximum number of iterations has been reached. If the maximum number of iterations has not been reached, the imaging system loops back to operation 1010. If the maximum number of iterations has been reached, the imaging system generates a pose estimate 1040 based on the pose estimate of operation 1030. The pose estimate 1040 may include pose measurements and / or a covariance of a Kalman filter.
[0120] Figure 11 1 is a conceptual diagram illustrating image reprojection. Points representing the position of camera 1105 are illustrated. Dark markers are illustrated along image plane 1110, representing observations 1115 of features along image plane 1110. Additional markers are illustrated in 3D space outside image plane 1110, representing 3D point estimates 1120 corresponding to each of the features observed in observations 1115. 3D point estimates 1120 include nearby points 1125 (illustrated as white markers with black outlines) and more distant points 1130 (illustrated as black markers with dashed outlines). As discussed with respect to weights 240, in some examples, features representing more distant points 1130 may be assigned a higher weight than features representing nearby points 1125.
[0121] The pose of the camera 1105 is estimated by tracking feature points in the image represented by the image plane 1110. A motion blurred image formation model may be used according to the following equations 2 to 5:
[0122] Equation 2:
[0123] Equation 3:
[0124] Equation 4:
[0125] Equation 5:
[0126] In Equations 2 to 5, Δx represents the motion blur of feature point i as calculated using Equation 5. irepresents the feature point i in the camera. T = [R t] represents the incremental pose of the camera during the exposure window. R = exp[ω*dT] × Means every t = v × dT + 0.5a × dT 2 The incremental rotation of the camera. dT represents the exposure time. B(x)∈R W×H is the captured image (e.g., a motion blurred image). t (x) is the virtual sharp image captured at time t. τ is the exposure time. The estimated motion blur (e.g., by motion blur estimator 235) is used for pose estimation (e.g., by pose engine 260) to compute reliable poses from tracked features, and / or in some cases for feature tracking (e.g., by feature tracker 225), environment mapping (e.g., by mapping engine 250), and / or other SLAM functions (e.g., by SLAM engine 245).
[0127] If motion and scene information is available from 6DoF, the imaging system calculates motion blur at several control points in the image at different exposure levels. The imaging system selects these control points whose motion blur is less than the maximum exposure time of 1 to 2 pixels. The imaging system calculates motion blur as the movement of pixels in the image during the exposure window. When depth information is available, Equation 6 is used together with Equations 2 to 5:
[0128]
[0129] In Equation 6, x i Represents the position of the control point at t. represents the position of the control point at t+dT. π represents the camera function. T = [R t] represents the incremental pose of the camera during the exposure window. R = exp[ω*dT] × Indicates the incremental rotation of the camera. t = v × dT + 0.5a × dT 2 dT represents the exposure time to be estimated. Indicates the amount of motion blur.
[0130] In some examples, angular velocity is provided from an IMU and / or gyroscope. In some examples (e.g., extended reality (XR) applications), many fast motions are rotational. In XR applications, assuming pure rotation is a reasonable approximation. Depth information is not required in pure rotation. Equations 2 to 6 can be used, modified to T = [R] and d i =1.0m.
[0131] In some examples, motion is provided via optical flow. Optical flow gives the apparent motion of each pixel on the image plane 1110. The imaging system can calculate the median displacement of the image points based on the optical flow information across frames. The maximum exposure time can be calculated as shown in Equation 7 below:
[0132] dT=a / (b*FPS)
[0133] Equation 7
[0134] In Equation 7, dT represents the maximum exposure time. a represents the maximum allowed motion blur amplitude. b represents the median displacement of the image points from the optical flow data. FPS represents the image capture rate in frames per second, measured in Hertz (Hz).
[0135] Figure 12 1 is a block diagram illustrating a process 1200 for automatic exposure control that takes into account angular velocity, motion information, scene depth, and / or optical flow. The process 1200 for imaging may be performed by an imaging system (e.g., a chipset, a processor or processors (such as an ISP, HP, or other processors), or other components). In some examples, the imaging system may include, for example, the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the imaging system 200, the image sensor 205, the motion sensor 215, the feature tracker 225, the motion blur estimator 235, the SLAM engine 245, the mapping engine 250, the pose engine 260, the output processor 275, and / or the trained ML model 290, the trained ML model 290, the feedback engine 295, the HMD 310, mobile phone 410, vehicle 510, imaging system that performs any process described herein (e.g., process 700, process 800, process 900, process 1000, and / or process 1400), camera 1105, neural network 1300, computing system 1500, processor 1510, systems and devices, equipment, non-transitory computer-readable media having stored thereon a program to be executed using a processor, or a combination thereof.
[0136] The imaging system includes an AEC engine 1205 (e.g., AEC engine 805) having an exposure / gain table 1210 (e.g., exposure / gain table 810) that identifies exposure time settings and gain settings to be used by a camera 1230 (e.g., camera 830, image sensor 205). The AEC engine 1205 can output image capture settings 1220 (e.g., including exposure time and / or gain) (e.g., image capture settings 820, image capture settings 210) to the camera 1230. A delay 1225 (e.g., delay 825) (e.g., of N frames when the camera 1230 is capturing video) may occur before the camera 1230 applies the image capture settings 1220. After the delay 1225, the camera 1230 captures a raw image 1235 (e.g., raw image 835, image data 220) with the image capture settings 1220 applied.
[0137] The imaging system may receive motion data 270 from the motion sensor 215 and / or from optical flow analysis of previous images, including, for example, angular velocity 1242 from an IMU and / or gyroscope, motion information 1244 and / or scene depth information 1246 from a six-degree-of-freedom (6DoF) sensor, and optical flow data 1248 from previous images (e.g., previous image data captured by the camera 1230 prior to the raw image 1235). The imaging system may use this data, along with information about the geometry of the camera 1230 and / or adjustments to the exposure determined by the AEC engine 1205 to match the image luminance of the raw image 1235 to the target luminance (e.g., adjustment 815), to determine a maximum exposure time 1240, which the imaging system may use to generate an update 1250 to the exposure / gain table 1210. For example, if the imaging system is determined to be in motion (e.g., translation and / or rotation speed and / or acceleration exceeds a threshold), the motion-aware AEC settings 1245 may reduce the maximum exposure time 1240, or if the imaging system is determined to be stationary (e.g., translation and / or rotation speed and / or acceleration is below a threshold), the motion-aware AEC settings may increase the maximum exposure time 1240 or (e.g., based on adjustment 815) leave the maximum exposure time 1240 at the position suggested by the AEC engine 1205. The updated exposure / gain table 1210 may then be used to determine additional image capture settings 1220 to be used by the camera 1230 to capture additional raw images 1235, etc. The AEC engine 1205 and / or other image processing subsystems of the imaging system may process each raw image 1235 to produce an image 1255 (e.g., image 855, image data 220), which may be used by the feature tracker 1260 (e.g., feature tracker 225, feature tracking 715).
[0138] In extremely low light (e.g., less than 5 lux) or low light (e.g., 15 lux to 30 lux), the imaging system may allow the maximum exposure time to be higher during slower motion, which helps capture better images with sufficient contrast required for 6DoF tracking. Under slower motion, the imaging system may allow the maximum exposure to reach 8ms to help capture images with less noise. Example low exposure settings (e.g., during motion and extremely low light) may include exposure time = 2ms; gain = maximum gain. Using motion-aware AEC, exposure settings (e.g., even during motion in low light) may include exposure time = 8ms; gain = maximum gain.
[0139] Figure 13 1 is a block diagram illustrating an example of a neural network (NN) 1300 that can be used for media processing operations. The neural network 1300 can include any type of deep network, such as a convolutional neural network (CNN), an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), a generative adversarial network (GAN), and / or other types of neural networks. The neural network 1300 can be an example of a trained ML model in the trained ML model 290. The neural network 1300 can be used by the feature tracker 225, the motion blur estimator 235, the SLAM engine 245, the mapping engine 250, the pose engine 260, the output processor 275, the feedback engine 295, or a combination thereof.
[0140] Input layer 1310 of neural network 1300 includes input data. The input data of input layer 1310 may include data representing pixels of one or more input image frames. In some examples, the input data of input layer 1310 includes data representing pixels of image data (e.g., image data 220, an image captured by one of cameras 330A-330D, an image captured by one of cameras 430A-430D, an image captured by one of cameras 530A-530F, image 600, image 705, original image 835, image 855, image plane 1110, original image 1235, image 1255, image data of operation 1405, an image captured using input device 1545, or a combination thereof). In some examples, the input data of input layer 1310 includes motion data 270 captured by motion sensor 215. In some examples, the input data of input layer 1310 includes depth data captured by a depth sensor. In some examples, the input data to the input layer 1310 includes processed data to be further processed, such as features 230, weights 240, maps 255, poses 265, output data 280, or a combination thereof.
[0141] The image may include image data from an image sensor, including raw pixel data (including a single color per pixel based on, for example, a Bayer color filter) or processed pixel values (e.g., RGB pixels of an RGB image). Neural network 1300 includes a plurality of hidden layers 1312, 1312B through 1312N. Hidden layers 1312, 1312B through 1312N include "N" hidden layers, where "N" is an integer greater than or equal to 1. The plurality of hidden layers may include as many layers as required for a given application. Neural network 1300 further includes an output layer 1314 that provides output resulting from the processing performed by hidden layers 1312, 1312B through 1312N.
[0142] In some examples, output layer 1314 may provide output data, such as features 230, weights 240, map 255, pose 265, output data 280, or intermediate data used to generate any of these (e.g., by feature tracker 225, motion blur estimator 235, SLAM engine 245, mapping engine 250, pose engine 260, output processor 275, feedback engine 295, or a combination thereof).
[0143] Neural network 1300 is a multi-layer neural network of interconnected filters. Each filter can be trained to learn features that represent input data. Information associated with these filters is shared between different layers, and each layer retains information as it processes it. In some cases, neural network 1300 may comprise a feedforward network, in which case there are no feedback connections where the output of the network is fed back into itself. In some cases, network 1300 may comprise a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read in.
[0144] In some cases, information can be exchanged between layers through node-to-node interconnections between the layers. In some cases, the network may include a convolutional neural network, which may not link every node in one layer to every other node in the next layer. In a network in which information is exchanged between layers, a node in input layer 1310 may activate a set of nodes in the first hidden layer 1312A. For example, as shown in the figure, each of the input nodes in input layer 1310 may be connected to each of the nodes in the first hidden layer 1312A. The nodes in the hidden layer may transform the information of each input node by applying an activation function (e.g., a filter) to the information. The information derived from this transformation may then be passed to the nodes in the next hidden layer 1312B and activated, which may perform their own designated functions. Example functions include convolution, reduction, amplification, data transformation, and / or any other suitable function. The output of hidden layer 1312B may then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 1312N may activate one or more nodes in output layer 1314, providing the processed output image. In some cases, although a node in neural network 1300 (e.g., node 1316) is shown as having multiple output lines, the node has a single output, and all lines shown as output from the node represent the same output value.
[0145] In some cases, each node or the interconnection between nodes may have a weight, which is a set of parameters derived from the training of neural network 1300. For example, an interconnection between nodes may represent a piece of information learned about the interconnected nodes. The interconnection may have a tunable digital weight that can be tuned (e.g., based on a training data set), thereby allowing neural network 1300 to adapt to the input and learn as more data is processed.
[0146] The neural network 1300 is pre-trained to process features of the data from the input layer 1310 using different hidden layers 1312 , 1312B to 1312N to provide an output via an output layer 1314 .
[0147] Figure 141400 is a flow diagram illustrating a process 1400 for imaging. The process 1400 for imaging may be performed by an imaging system (e.g., a chipset, a processor or processors (such as an ISP, HP, or other processor), or other components). In some examples, the imaging system may include, for example, image capture and processing system 100, image capture device 105A, image processing device 105B, image processor 150, ISP 154, host processor 152, imaging system 200, image sensor 205, motion sensor 215, feature tracker 225, motion blur estimator 235, SLAM engine 245, mapping engine 250, pose engine 260, output processor 275 and / or trained ML model 290, trained ML model 290, feedback engine 295, HMD 310, mobile phone 410, vehicle 510, imaging system that performs any process described herein (e.g., process 700, process 800, process 900, process 1000, and / or process 1200), camera 1105, neural network 1300, computing system 1500, processor 1510, system and apparatus, device, non-transitory computer-readable medium having stored thereon a program to be executed using a processor, or a combination thereof. In some examples, the imaging system includes a display. In some examples, the imaging system includes a transceiver and / or other communication interface.
[0148] At operation 1405 , the imaging system (or components thereof) is configured to and may receive an image of an environment captured using at least one image sensor according to image capture settings.
[0149] Illustrative examples of image sensors include image sensor 130, image sensor 205, first camera 330A, second camera 330B, third camera 330C, fourth camera 330D, first camera 430A, second camera 430B, third camera 430C, fourth camera 430D, first camera 530A, second camera 530B, third camera 530C, fourth camera 530D, fifth camera 530E, sixth camera 530F, first sensor 540A, second sensor 540B, sensor 625, an image sensor for capturing images used as input data for input layer 1310 of NN 1300, input device 1545, another image sensor described herein, another sensor described herein, or a combination thereof. Examples of depth sensors include image sensor 205, first sensor 540A, second sensor 540B, sensor 625, a depth sensor for capturing depth data used as input data for input layer 1310 of NN 1300, input device 1545, another depth sensor described herein, another sensor described herein, or a combination thereof. Examples of image data include image data 220 and / or image data captured by any of the previously listed image sensors.
[0150] In some aspects, the image capture settings include an exposure time. Examples of image capture settings include image capture settings 210, exposure time 710, image capture settings 820, the exposure time of operations 910 and / or 1020, and image capture settings 1220.
[0151] At operation 1410, the imaging system (or a component thereof) is configured to receive motion data captured using a motion sensor. Examples of motion sensors include motion sensor 215. Examples of motion data include motion data 270, accelerometer data 735, gyroscope data 740, motion information of operations 910 and / or 1020, angular velocity 1242, and / or motion information 1244.
[0152] At operation 1415, the imaging system (or a component thereof) is configured to and may determine a weight (e.g., weight 240) associated with at least one of the features of the environment in the image (e.g., feature 230) based on an estimated motion blur level (e.g., determined using motion blur estimator 235) of the at least one of the plurality of features of the environment in the image. The estimated motion blur level is based on the motion data and the image capture settings.
[0153] In some aspects, the imaging system (or components thereof) is configured to and may determine an estimated level of motion blur for at least one of the features of the environment in the image (eg, using the motion blur estimator 235 ).
[0154] In some aspects, the estimated level of motion blur of at least one of the features of the environment in the image is based on a distance from the at least one image sensor to the at least one of the features of the environment, wherein a weight associated with the at least one of the features of the environment in the image is based on the distance from the at least one image sensor to the at least one of the features of the environment. For example, the distance can be closer than a threshold (e.g., nearby point 1125) or farther than a threshold (e.g., farther point 1130).
[0155] In some aspects, the imaging system (or components thereof) is configured and may be based on (eg, as in the ratio in Equation 1) The weight associated with the image is determined by determining a ratio of a constant divided by the estimated motion blur level of the image based on an estimated motion blur level of at least one feature of the features of the environment in the image.
[0156] In some aspects, the estimated motion blur level is an estimated motion blur magnitude.
[0157] At operation 1420, the imaging system (or components thereof) is configured to and may track features of the environment (e.g., of image data 220) across multiple images according to respective weights (e.g., weights 240) of the features of the environment (e.g., features 230) across the multiple images, the multiple images including the image, and the respective weights including the weights.
[0158] In some aspects, the imaging system (or its components) is configured and may determine a pose of the imaging system in the environment based on the features tracked in the multiple images and according to the respective weights of the features of the environment in the multiple images (e.g., pose 265 determined by pose engine 260). The imaging system includes at least one image sensor and a motion sensor. In some aspects, the imaging system (or its components) is configured and may minimize a weighted least squares reprojection error according to the respective weights of the features of the environment in the multiple images to determine the pose of the device in the environment (e.g., as in operation 1030). In some aspects, the imaging system (or its components) is configured and may output an indication of the pose of the imaging system (e.g., outputting pose 265 and / or output data 280 via output device 285).
[0159] In some aspects, the imaging system (or its components) is configured and may map the environment based on features of the environment tracked in multiple images and according to respective weights of the features of the environment in the multiple images to generate a map of the environment (e.g., map 255 generated using mapping engine 250). In some aspects, the imaging system (or its components) is configured and may determine a position of the imaging system within the map of the environment (e.g., pose 265) based on features of the environment tracked in multiple images and according to respective weights of the features of the environment in the multiple images, wherein the imaging system includes at least one image sensor and a motion sensor. In some aspects, the imaging system (or its components) is configured and may output at least a portion of the map of the environment (e.g., outputting map 255 and / or outputting data 280 via output device 285).
[0160] In some aspects, the respective weights of the features of the environment across the multiple images correspond to the respective error variance values for the multiple images. In some aspects, the imaging system (or components thereof) is configured to and may track the features of the environment across the multiple images based on the respective error variance values for the multiple images to track the features of the environment across the multiple images based on the respective weights of the features of the environment across the multiple images (e.g., as shown in Equation 1).
[0161] In some aspects, an imaging system (or component thereof) is configured and may: determine that an estimated level of motion blur is less than a predetermined threshold (e.g., threshold δ in Equation 1); and, in response to determining that the estimated level of motion blur is less than the predetermined threshold, set a weight associated with at least one of the features of the environment in the image to a predetermined value (e.g., constant a in Equation 1) to determine the weight associated with at least one of the features of the environment in the image based on the estimated level of motion blur of the at least one of the features of the environment in the image. In some aspects, the predetermined threshold represents a motion blur magnitude no greater than a pixel.
[0162] In some examples, the processes described herein (e.g., Figure 1 process, Figure 2 process, Figure 7 The process of 700 Figure 8 The process of 800 Figure 9 The process of 900 Figure 10 The process of 1000 Figure 11 process, Figure 12 The process of 1200 Figure 13 process, Figure 14Process 1400 and / or other processes described herein) may be performed by a computing device or apparatus. In some examples, the processes described herein may be performed by the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the imaging system 200, the image sensor 205, the motion sensor 215, the feature tracker 225, the motion blur estimator 235, the SLAM engine 245, the mapping engine 250, the pose engine 260, the output processor 275 and / or the trained ML model 290, the trained ML model 290, the feedback engine 295, the HMD 310, the mobile phone 410, the vehicle 510, an imaging system performing any process described herein (e.g., process 700, process 800, process 900, process 1000, process 1200, and / or process 1400), the camera 1105, the neural network 1300, the computing system 1500, the processor 1510, or a combination thereof.
[0163] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a web-connected watch or smartwatch or other wearable device), a server computer, a vehicle or a computing device of a vehicle, a robotic device, a television, and / or any other computing device with the resource capacity to perform the processes described herein. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on an Internet Protocol (IP) or other types of data.
[0164] A component of a computing device can be implemented in circuitry. For example, a component may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0165] The processes described herein are illustrated as logic flow diagrams, block diagrams, or conceptual diagrams, the operations of which represent sequences of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operations. Generally speaking, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0166] Additionally, the processes described herein may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes together on one or more processors, implemented by hardware, or a combination of the above. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0167] Figure 15 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. Specifically, Figure 15 An example of a computing system 1500 is illustrated, which can be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 1505. Connection 1505 can be a physical connection using a bus, or a direct connection to processor 1510, such as in a chipset architecture. Connection 1505 can also be a virtual connection, a networked connection, or a logical connection.
[0168] In some aspects, computing system 1500 is a distributed system, wherein the functionality described in this disclosure may be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a plurality of such components, each of which performs some or all of the functionality of the described component. In some aspects, each component may be a physical or virtual device.
[0169] Example system 1500 includes at least one processing unit (CPU or processor) 1510 and connections 1505 that couple various system components including system memory 1515, such as read-only memory (ROM) 1520 and random access memory (RAM) 1525, to processor 1510. Computing system 1500 may include a cache 1512 of high-speed memory directly connected to, in close proximity to, or integrated as part of processor 1510.
[0170] Processor 1510 may include any general-purpose processor and hardware or software services, such as services 1532, 1534, and 1536 stored in storage device 1530, configured to control processor 1510 as well as a dedicated processor where software instructions are incorporated into the actual processor design. Processor 1510 may essentially be a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0171] To enable user interaction, the computing system 1500 includes an input device 1545 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, and the like. The computing system 1500 may also include an output device 1535 that can be one or more of a plurality of output mechanisms. In some instances, a multimodal system may enable a user to provide multiple types of input / output to communicate with the computing system 1500. The computing system 1500 may include a communication interface 1540, which generally may govern and manage user input and system output. The communication interface may perform or facilitate receiving and / or sending wired or wireless communications using wired and / or wireless transceivers, including utilizing an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs, Wireless signal transmission, Low energy (BLE) wireless signal transmission, The communication interface 1540 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 1500 based on receiving one or more signals from one or more satellites associated with the one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There is no restriction on operating on any particular hardware arrangement, and thus the base features herein may be readily substituted for improved hardware or firmware arrangements as they are developed.
[0172] The storage device 1530 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other type of computer-readable medium that can store data that can be accessed by a computer, such as a magnetic tape cartridge, a flash memory card, a solid-state memory device, a digital versatile disk, a magnetic cassette, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage medium, a flash memory, a memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, a digital video disc (DVD) optical disc, a Blu-ray disc (BDD) optical disc, a holographic optical disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASH EPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin-transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.
[0173] Storage devices 1530 may include software services, servers, services, etc. that, when code defining such software is executed by processor 1510, cause the system to perform functions. In some aspects, hardware services that perform specific functions may include software components for performing functions stored in a computer-readable medium connected to the necessary hardware components (such as processor 1510, connection 1505, output devices 1535, etc.).
[0174] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transient media in which data can be stored and does not include carrier waves and / or transient electronic signals that are propagated wirelessly or on a wired connection. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer-readable media may store thereon code and / or machine-executable instructions that may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, categories, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be passed, forwarded, or sent using any suitable means, including memory sharing, message passing, token passing, network sending, etc.
[0175] In some aspects, computer-readable storage devices, media, and memories may include wired or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media specifically excludes media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.
[0176] Specific details are provided in the above description to provide a detailed understanding of the aspects and examples provided herein. However, it will be understood by those skilled in the art that these aspects can be practiced without these specific details. For clarity of explanation, in some cases, the present technology can be presented as comprising separate functional blocks, including functional blocks comprising devices, device components, steps in the method embodied in software or a combination of hardware and software or routines. Additional components other than those components shown in the accompanying drawings and / or described herein can be used. For example, circuits, systems, networks, processes and other components can be shown as components in block diagram form to avoid confusing these aspects in unnecessary details. In other cases, known circuits, processes, algorithms, structures and techniques can be shown without unnecessary details to avoid confusing various aspects.
[0177] Various aspects may be described above as processes or methods, which may be depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or concurrently. Furthermore, the order of the operations may be rearranged. A process is terminated when its operations are completed, but a process may have additional steps not included in the accompanying figures. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, termination of the process may correspond to the function returning to the calling function or the main function.
[0178] The processes and methods according to the examples described above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtained from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0179] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. By way of further example, such functionality may also be implemented on circuit boards among different chips or different processes executed on a single device.
[0180] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0181] In the foregoing description, various aspects of the present application are described with reference to specific aspects thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the exemplary aspects of the present application have been described in detail herein, it is to be understood that each inventive concept can be implemented and adopted in various other ways, and the appended claims are not intended to be interpreted as including these variations, unless limited by the prior art. The various features and aspects of the application described above can be used individually or in combination. In addition, the various aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader essence and scope of this specification. Therefore, the description and drawings should be considered as illustrative rather than restrictive. For illustrative purposes, each method is described in a specific order. It should be understood that, in alternative aspects, each method can be performed in a different order than described.
[0182] It should be understood by those of ordinary skill in the art that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of the present description.
[0183] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0184] The phrase “coupled to” refers to any component being physically connected directly or indirectly to another component, and / or any component being in communication directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0185] Claim language or other language that recites "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language that recites "at least one of A and B" means A, B, or A and B. In another example, claim language that recites "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language that recites "at least one of A and B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0186] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in conjunction with the various aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A skilled person may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as departing from the scope of this application.
[0187] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices, or integrated circuit devices with multiple uses, including applications in wireless communication devices and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques may be implemented at least in part by a computer-readable data storage medium comprising program code, the program code comprising instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, or the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0188] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in an alternative embodiment, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within a dedicated software module or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
[0189] Illustrative aspects of the present disclosure include:
[0190] Aspect 1. An apparatus for mapping an environment, the apparatus comprising: a memory; and at least one processor (e.g., implemented in a circuit), the at least one processor coupled to the memory and configured to: receive an image of an environment captured using at least one image sensor according to image capture settings; receive motion data captured using a motion sensor; determine a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level of the at least one feature of the environment in the image, wherein the estimated motion blur level is based on the motion data and the image capture settings; and track the feature of the environment across a plurality of images according to the corresponding weight of the feature of the environment across the plurality of images, wherein the plurality of images includes the image, and wherein the corresponding weight includes the weight.
[0191] Aspect 2. An apparatus according to Aspect 1, wherein the at least one processor is configured to: determine a pose of the apparatus in the environment based on the features tracked in the multiple images and according to the corresponding weights of the features of the environment in the multiple images.
[0192] Aspect 3. An apparatus according to any one of Aspects 1 to 2, wherein the at least one processor is configured to: minimize the weighted least squares of the reprojection error based on the corresponding weights of the features of the environment in the multiple images to determine the pose of the apparatus in the environment.
[0193] Aspect 4. An apparatus according to any one of Aspects 1 or 3, wherein the at least one processor is configured to: output an indication of the pose of the apparatus.
[0194] Aspect 5. An apparatus according to any one of Aspects 1 to 4, wherein the at least one processor is configured to: map the environment based on the features of the environment tracked in the multiple images and according to the corresponding weights of the features of the environment in the multiple images to generate a map of the environment.
[0195] Aspect 6. An apparatus according to any one of Aspects 1 to 5, wherein the at least one processor is configured to: determine the position of the apparatus within the map of the environment based on the features of the environment tracked in the multiple images and based on the corresponding weights of the features of the environment in the multiple images.
[0196] Aspect 7. The apparatus according to any one of aspects 1 to 6, wherein the at least one processor is configured to: output at least a portion of the map of the environment.
[0197] Aspect 8. An apparatus according to any one of Aspects 1 to 7, wherein the corresponding weights of the features of the environment across the multiple images correspond to corresponding error variance values of the multiple images, and wherein the at least one processor is configured to: track the features of the environment across the multiple images according to the corresponding error variance values of the multiple images, so as to track the features of the environment across the multiple images according to the corresponding weights of the features of the environment across the multiple images.
[0198] Aspect 9. An apparatus according to any one of aspects 1 to 8, wherein the at least one processor is configured to: determine the estimated level of motion blur of the at least one of the features of the environment in the image.
[0199] Aspect 10. An apparatus according to any one of Aspects 1 to 9, wherein the estimated motion blur level of at least one of the features of the environment in the image is based on the distance from the at least one image sensor to the at least one of the features of the environment, and wherein the weight associated with the at least one of the features of the environment in the image is based on the distance from the at least one image sensor to the at least one of the features of the environment.
[0200] Aspect 11. An apparatus according to any one of Aspects 1 to 10, wherein the at least one processor is configured to: determine a ratio of a constant divided by the estimated motion blur level of the image based on the estimated motion blur level of at least one of the features of the environment in the image to determine the weight associated with the image.
[0201] Aspect 12. An apparatus according to any one of Aspects 1 to 11, wherein the at least one processor is configured to: determine that the estimated motion blur level is less than a predetermined threshold; and in response to determining that the estimated motion blur level is less than the predetermined threshold, set the weight associated with at least one of the features of the environment in the image to a predetermined value to determine the weight associated with at least one of the features of the environment in the image based on the estimated motion blur level of the at least one of the features of the environment in the image.
[0202] Clause 13. The apparatus according to any one of clauses 1 to 12, wherein the predetermined threshold represents a motion blur magnitude no greater than a pixel.
[0203] Clause 14. The apparatus of any one of clauses 1 to 13, wherein the estimated motion blur level is an estimated motion blur magnitude.
[0204] Aspect 15. The apparatus of any one of aspects 1 to 14, wherein the image capture settings include an exposure time.
[0205] Aspect 16. The apparatus of any one of aspects 1 to 15, wherein the apparatus is at least one of a mobile device, a wireless communication device, or an extended reality device.
[0206] Aspect 17. A method for imaging, the method comprising: receiving an image of an environment captured using at least one image sensor according to an image capture setting; receiving motion data captured using a motion sensor; determining a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level of the at least one feature of the features of the environment in the image, wherein the estimated motion blur level is based on the motion data and the image capture setting; and tracking the features of the environment across a plurality of images according to their corresponding weights, wherein the plurality of images includes the image, and wherein the corresponding weights include the weights.
[0207] Aspect 18. The method according to Aspect 17 further comprises: determining a posture of the device in the environment based on the features tracked in the multiple images and according to the corresponding weights of the features of the environment in the multiple images, wherein the device includes the at least one image sensor and the motion sensor.
[0208] Aspect 19. The method according to any one of Aspects 17 to 18, further comprising: minimizing weighted least squares of reprojection error based on the corresponding weights of the features of the environment in the multiple images to determine the pose of the device in the environment.
[0209] Aspect 20. The method according to any one of aspects 17 to 19, further comprising: outputting an indication of the pose of the device.
[0210] Aspect 21. The method according to any one of Aspects 17 to 20, further comprising: mapping the environment based on the features of the environment tracked in the multiple images and according to the corresponding weights of the features of the environment in the multiple images to generate a map of the environment.
[0211] Aspect 22. The method according to any one of Aspects 17 to 21, further comprising: determining the position of the device within the map of the environment based on the features of the environment tracked in the multiple images and according to the corresponding weights of the features of the environment in the multiple images, wherein the device includes the at least one image sensor and the motion sensor.
[0212] Aspect 23. The method according to any one of aspects 17 to 22, further comprising: outputting at least a portion of the map of the environment.
[0213] Aspect 24. A method according to any one of Aspects 17 to 23, wherein the corresponding weights of the features of the environment across the multiple images correspond to corresponding error variance values of the multiple images, and the method further includes: tracking the features of the environment across the multiple images according to the corresponding error variance values of the multiple images, so as to track the features of the environment across the multiple images according to the corresponding weights of the features of the environment across the multiple images.
[0214] Aspect 25. The method according to any one of aspects 17 to 24, further comprising: determining the estimated level of motion blur of the at least one of the features of the environment in the image.
[0215] Aspect 26. A method according to any one of Aspects 17 to 25, wherein the estimated motion blur level of at least one of the features of the environment in the image is based on the distance from the at least one image sensor to the at least one of the features of the environment, and wherein the weight associated with the at least one of the features of the environment in the image is based on the distance from the at least one image sensor to the at least one of the features of the environment.
[0216] Aspect 27. A method according to any one of Aspects 17 to 26, the method further comprising: determining a ratio of a constant divided by the estimated motion blur level of the image based on the estimated motion blur level of at least one of the features of the environment in the image to determine the weight associated with the image.
[0217] Aspect 28. The method according to any one of Aspects 17 to 27, further comprising: determining that the estimated motion blur level is less than a predetermined threshold; and in response to determining that the estimated motion blur level is less than the predetermined threshold, setting the weight associated with at least one of the features of the environment in the image to a predetermined value to determine the weight associated with at least one of the features of the environment in the image based on the estimated motion blur level of the at least one of the features of the environment in the image.
[0218] Aspect 29. The method according to any one of aspects 17 to 28, wherein the predetermined threshold represents a motion blur magnitude no greater than a pixel.
[0219] Aspect 30. The method of any one of aspects 17 to 29, wherein the estimated motion blur level is an estimated motion blur magnitude.
[0220] Aspect 31. The method of any one of aspects 17 to 30, wherein the image capture settings include an exposure time.
[0221] Aspect 32. The method according to any one of aspects 17 to 31, wherein the method is configured to be performed using an apparatus comprising at least one of a mobile device, a wireless communication device, and an extended reality device.
[0222] Aspect 33. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform the operations of any one of aspects 1 to 32.
[0223] Aspect 34. An apparatus for imaging, the apparatus comprising one or more components for performing the operations according to any one of aspects 1 to 32.
Claims
1. An apparatus for imaging, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receiving an image of an environment captured using at least one image sensor according to image capture settings; receiving motion data captured using a motion sensor; determining a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level of the at least one feature of the plurality of features of the environment in the image, wherein the estimated motion blur level is based on the motion data and the image capture settings; as well as The feature of the environment across a plurality of images is tracked according to respective weights of the feature of the environment across the plurality of images, wherein the plurality of images includes the image, and wherein the respective weights include the weight.
2. The apparatus of claim 1 , wherein the at least one processor is configured to: A pose of the device in the environment is determined based on the features tracked in the plurality of images and as a function of the respective weights of the features of the environment in the plurality of images.
3. The apparatus of claim 2, wherein the at least one processor is configured to: The pose of the device in the environment is determined by minimizing a weighted least squares reprojection error according to the respective weights of the features of the environment in the plurality of images.
4. The apparatus of claim 2, wherein the at least one processor is configured to: An indication of the pose of the device is output.
5. The apparatus of claim 1 , wherein the at least one processor is configured to: The environment is mapped based on the features of the environment tracked in the plurality of images and according to the respective weights of the features of the environment of the plurality of images to generate a map of the environment.
6. The apparatus of claim 5, wherein the at least one processor is configured to: A position of the device within the map of the environment is determined based on the features of the environment tracked in the plurality of images and in accordance with the respective weights of the features of the environment of the plurality of images.
7. The apparatus of claim 5, wherein the at least one processor is configured to: At least a portion of the map of the environment is output.
8. An apparatus according to claim 1, wherein the corresponding weights of the multiple images correspond to corresponding error variance values of the features of the environment across the multiple images, and wherein the at least one processor is configured to: track the features of the environment across the multiple images based on the corresponding error variance values of the multiple images, and to track the features of the environment across the multiple images based on the corresponding weights of the features of the environment across the multiple images.
9. The apparatus of claim 1 , wherein the at least one processor is configured to: The estimated level of motion blur of the at least one of the features of the environment in the image is determined.
10. The apparatus of claim 1 , wherein the estimated level of motion blur of the at least one of the features of the environment in the image is based on a distance from the at least one image sensor to the at least one of the features of the environment, and wherein the weight associated with the at least one of the features of the environment in the image is based on the distance from the at least one image sensor to the at least one of the features of the environment.
11. The apparatus of claim 1 , wherein the at least one processor is configured to: A ratio of a constant divided by the estimated motion blur level of the image is determined based on the estimated motion blur level of the at least one of the features of the environment in the image to determine the weight associated with the image.
12. The apparatus of claim 1 , wherein the at least one processor is configured to: determining that the estimated motion blur level is less than a predetermined threshold; and In response to determining that the estimated level of motion blur is less than the predetermined threshold, setting the weight associated with the at least one of the features of the environment in the image to a predetermined value to determine the weight associated with the at least one of the features of the environment in the image based on the estimated level of motion blur of the at least one of the features of the environment in the image.
13. The apparatus of claim 12, wherein the predetermined threshold represents a motion blur magnitude no greater than a pixel. The apparatus of claim 1 , wherein the estimated motion blur level is an estimated motion blur magnitude.
15. The apparatus of claim 1, wherein the image capture settings include an exposure time.
16. The apparatus of claim 1, wherein the apparatus is at least one of a mobile device, a wireless communication device, or an extended reality device.
17. A method for imaging, the method comprising: receiving an image of an environment captured using at least one image sensor according to image capture settings; receiving motion data captured using a motion sensor; determining a weight associated with at least one of a plurality of features of the environment in the image based on an estimated motion blur level of the at least one feature of the plurality of features of the environment in the image, wherein the estimated motion blur level is based on the motion data and the image capture settings; as well as The feature of the environment across a plurality of images is tracked according to respective weights of the feature of the environment across the plurality of images, wherein the plurality of images includes the image, and wherein the respective weights include the weight.
18. The method according to claim 17, further comprising: A pose of a device in the environment is determined based on the features tracked in the plurality of images and according to the respective weights of the features of the environment in the plurality of images, wherein the device includes the at least one image sensor and the motion sensor.
19. The method according to claim 18, further comprising: A weighted least squares reprojection error is minimized based on the respective weights of the features of the environment in the plurality of images to determine the pose of the device in the environment.
20. The method according to claim 18, further comprising: An indication of the pose of the device is output.
21. The method according to claim 17, further comprising: The environment is mapped based on the features of the environment tracked in the plurality of images and according to the respective weights of the features of the environment of the plurality of images to generate a map of the environment.
22. The method according to claim 21, further comprising: A position of a device within the map of the environment is determined based on the features of the environment tracked in the plurality of images and in accordance with the respective weights of the features of the environment of the plurality of images, wherein the device includes the at least one image sensor and the motion sensor.
23. The method according to claim 21, further comprising: At least a portion of the map of the environment is output.
24. The method of claim 17, wherein the respective weights of the features of the environment across the plurality of images correspond to respective error variance values for the plurality of images, the method further comprising: The features of the environment across the multiple images are tracked according to the respective error variance values of the multiple images to track the features of the environment across the multiple images according to the respective weights of the features of the environment across the multiple images.
25. The method of claim 17, further comprising: The estimated level of motion blur of the at least one of the features of the environment in the image is determined.
26. The method of claim 17, wherein the estimated level of motion blur of the at least one of the features of the environment in the image is based on a distance from the at least one image sensor to the at least one of the features of the environment, and wherein the weight associated with the at least one of the features of the environment in the image is based on the distance from the at least one image sensor to the at least one of the features of the environment.
27. The method of claim 17, further comprising: A ratio of a constant divided by the estimated motion blur level of the image is determined based on the estimated motion blur level of the at least one of the features of the environment in the image to determine the weight associated with the image.
28. The method of claim 17, further comprising: determining that the estimated motion blur level is less than a predetermined threshold; as well as In response to determining that the estimated level of motion blur is less than the predetermined threshold, setting the weight associated with the at least one of the features of the environment in the image to a predetermined value to determine the weight associated with the at least one of the features of the environment in the image based on the estimated level of motion blur of the at least one of the features of the environment in the image.
29. The method of claim 28, wherein the predetermined threshold represents a motion blur magnitude no greater than a pixel.
30. The method of claim 17, wherein the image capture settings include exposure time.