Hybrid system for feature detection and descriptor generation
Through the non-machine learning-based feature detector and machine learning-based descriptor generator in the hybrid system, the problem of inconsistency in feature descriptors in images under different conditions is solved, and robust feature detection and descriptor generation is realized, improving the stability of positioning and map construction.
Patent Information
- Application Number
- CN202380074439.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-14
- Filing Date
- 2023-09-13
- Publication Date
- 2025-06-03
AI Technical Summary
When capturing images under different conditions, it is difficult to generate common descriptors for the same features detected in different images, resulting in reduced robustness of positioning and map construction.
A hybrid system is adopted, combining non-machine learning-based feature detectors and machine learning-based descriptor generators. A non-machine learning-based feature detector is used to detect feature points, while a machine learning-based descriptor generator generates feature descriptors through a machine learning system, and uses a transformer neural network to determine a unique signature across different types of input data.
The ability to generate common or unique descriptors for detected features in images captured under different conditions is achieved, improving the robustness of positioning and map construction without manual labeling of data for training.
Smart Images

Figure CN120092272A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to processing sensor data (e.g., images, radar data, light detection and ranging (LIDAR) data, etc.). For example, aspects of the present disclosure relate to hybrid systems for performing feature (e.g., key point) detection and descriptor generation. Background Art
[0002] Many devices and systems allow the capture of characteristics of a scene based on sensor data such as an image (or frame) of the scene, video data of the scene (including multiple frames), radar data, LIDAR data, etc. For example, a camera or a device including a camera can capture a sequence of frames of a scene (e.g., a video of the scene). In some cases, the sequence of frames can be processed to perform one or more functions, can be output for display, can be output for processing and / or consumption by other devices, and for other uses.
[0003] The degrees of freedom (DoF) refer to the number of basic ways in which a rigid object can move in three-dimensional (3D) space. In some examples, six different DoFs of an object can be tracked, including three translational DoFs and three rotational DoFs. Some devices can track some or all of these degrees of freedom. In some cases, tracking (e.g., 6DoF tracking) can be used to perform positioning and mapping functions. For example, to perform positioning and mapping functions, a device or system can perform feature analysis (e.g., extraction, tracking, etc.) and other complex functions. Summary of the Invention
[0004] In some examples, techniques for using a hybrid system to perform feature (e.g., key point) detection and descriptor generation are described. According to at least one illustrative example, a method for processing image data is provided, the method comprising: obtaining input data; processing the input data using a machine learning-based feature detector to determine one or more feature points in the input data; and using a machine learning system to determine a corresponding feature descriptor for each respective feature point among the one or more feature points.
[0005] In another example, an apparatus for processing image data is provided, which includes at least one memory and at least one processor (e.g., implemented in a circuit) coupled to the at least one memory. The at least one processor is configured to: obtain input data; process the input data using a machine learning-based feature detector to determine one or more feature points in the input data; and use a machine learning system to determine a corresponding feature descriptor for each respective feature point among the one or more feature points.
[0006] In another example, a non-transitory computer-readable medium is provided, having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: obtain input data; process the input data using a non-machine learning-based feature detector to determine one or more feature points in the input data; and use a machine learning system to determine a corresponding feature descriptor for each respective feature point among the one or more feature points.
[0007] In another example, an apparatus for processing image data is provided. The apparatus includes: means for obtaining input data; means for processing the input data using a non-machine learning-based feature detector to determine one or more feature points in the input data; and means for using a machine learning system to determine a corresponding feature descriptor for each respective feature point among the one or more feature points.
[0008] In some aspects, one or more of the apparatuses described herein are or are part of the following: a camera, a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other devices. In some aspects, an apparatus includes one or more cameras for capturing one or more images. In some aspects, the apparatus further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatus may include one or more sensors that may be used to determine the position and / or orientation of the apparatus, the state of the apparatus, and / or for other purposes.
[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0010] The foregoing and other features and embodiments will become more apparent when reference is made to the following specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Exemplary embodiments of the present application are described in detail below with reference to the following drawings, in which:
[0012] Figure 1 is a block diagram illustrating the architecture of an image capture and processing device according to some examples;
[0013] Figure 2is a block diagram illustrating an architecture of an example extended reality (XR) system according to some examples;
[0014] Figure 3 is a block diagram illustrating an architecture of a simultaneous localization and mapping (SLAM) device according to some examples;
[0015] Figure 4 is an example frame captured by a SLAM system according to some aspects;
[0016] Figure 5 is a diagram illustrating an example of a hybrid system 500 for detecting features (e.g., key points or feature points) and generating descriptors of the detected features according to some aspects;
[0017] Figure 6 is a flowchart illustrating an example of a process for processing image data according to some examples of the present disclosure; and
[0018] Figure 7 is a block diagram illustrating an example of a computing system for implementing certain aspects described herein. Detailed Description
[0019] Certain aspects and implementations of the present disclosure are provided below. Some of these aspects and implementations may be applied independently, and some of them may be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of the implementations of the present application. However, it will be apparent that the various implementations may be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.
[0020] The following description only provides example implementations and is not intended to limit the scope, applicability, or configuration of the present disclosure. On the contrary, the following description of the example implementations will provide those skilled in the art with an enabling description for implementing the example implementations. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the essence and scope of the present application as set forth in the appended claims.
[0021] As described above, the devices and systems may determine or capture characteristics of a scene based on sensor data associated with the scene. The sensor data may include an image (or frame) of the scene, video data of the scene (including multiple frames), radar data, LIDAR data, any combination thereof, and / or other data.
[0022] For example, an image capture device (e.g., a camera) is a device that uses an image sensor to receive light and capture image frames such as still images or video frames. The terms "image", "image frame", "video frame", and "frame" may be used interchangeably herein. An image capture device typically includes at least one lens that receives light from a scene and bends the light toward the image sensor of the image capture device. The light received by the lens passes through an aperture controlled by one or more control mechanisms and is received by the image sensor. The one or more control mechanisms may control exposure, focus, and / or zoom based on information from the image sensor and / or based on information from an image processor (e.g., a host or application process and / or an image signal processor). In some examples, the one or more control mechanisms include a motor or other control mechanism for moving the lens of the image capture device to a target lens position.
[0023] Degree of freedom (DoF) refers to the number of fundamental ways in which a rigid object can move in three-dimensional (3D) space. In some examples, six different DoFs of an object can be tracked. The six DoFs can include three translational DoFs corresponding to translational movement along three perpendicular axes, which may be referred to as the x-axis, y-axis, and z-axis. The six DoFs can also include three rotational DoFs corresponding to rotational movement about three axes, which may be referred to as pitch, yaw, and roll. Some devices (e.g., extended reality (XR) devices such as virtual reality (VR) or augmented reality (AR) headsets, mobile devices, transportation vehicles or transportation systems, robotic devices, etc.) can track some or all of these degrees of freedom. For example, a 3DoF tracker (e.g., of an XR headset) can track three rotational DoFs. A 6DoF tracker (e.g., of an XR headset) can track all six DoFs.
[0024] In some cases, tracking (e.g., 6DoF tracking) can be used to perform localization and mapping functions. Mapping can include the process of constructing or generating a map of a particular environment. Localization can include the process of determining the position of an object (e.g., a vehicle, an XR device, a robotic device, a mobile phone, etc.) within a map (e.g., a map generated using the mapping process). An example of a technique for localization and mapping is Visual Simultaneous Localization and Mapping (VSLAM). VSLAM is a computational geometry technique used in devices with cameras, such as a vehicle or a vehicle system (e.g., an autonomous driving system), an XR device (e.g., a head-mounted display (HMD), an AR head-mounted headset, etc.), a robotic device or system, a mobile phone, etc. In VSLAM, the device can construct and update a map of an unknown environment based on frames captured by the device's camera. The device can track its pose (e.g., the pose of the device's image sensor, such as the camera pose, which can be determined using 6DOF tracking) (e.g., position and / or orientation) within the environment as the device updates the map. For example, the device can be activated in a particular room of a building and can move throughout the interior of the building, capturing image frames. The device can construct a map of the environment and track its position within the environment based on tracking the positions where different objects in the environment appear in different image frames. Other types of sensor data besides image frames can also be used for VSLAM, such as radar and / or LIDAR data.
[0025] In the context of a system that tracks movement through an environment (e.g., an XR system, a robotic system, a vehicle such as an autonomous vehicle, a VSLAM system, etc.), degrees of freedom can refer to which of the six degrees of freedom the system is capable of tracking. As mentioned above, a 3DoF tracking system typically tracks three rotational DoFs (e.g., pitch, yaw, and roll). For example, a 3DoF headset can track a user of the headset turning their head left or right, tilting their head up or down, and / or tilting their head left or right. A 6DoF system can track three translational DoFs as well as three rotational DoFs. Thus, for example, a 6DoF headset can track a user moving forward, backward, laterally, and / or vertically in addition to tracking the three rotational DoFs.
[0026] To perform localization and mapping functions, a device (e.g., an XR device, a mobile device, etc.) may perform feature analysis (e.g., extraction, tracking, etc.) and other complex functions. For example, key-point features (also referred to as feature points or key points) may be determined from an image captured by a camera. Descriptors for the key-point features may also be generated to provide semantic meanings of the key-point features. Key-point features are features in an image that do not change under different conditions (e.g., different illumination and / or lighting, different views, different weather conditions, etc.), such as points associated with corners of an object in the image, distinctive features of an object, etc. Key-point features may be used as non-semantic features to improve the robustness of localization and mapping. Key-point features may include distinctive features extracted from one or more images, such as points associated with corners of a table, edges of a street sign, etc. Generating stable key-point features and descriptors (e.g., that do not change over time when the conditions and / or views of a scene change) is important so that localization and mapping using such features can be accurately performed.
[0027] However, when images are captured under different conditions (e.g., different illumination and / or lighting, different views, different weather conditions, etc.), it may be difficult to generate a common descriptor for the same features detected in different images. For example, a descriptor generated for a feature associated with a traffic sign detected in a first image captured during the day may be different from a descriptor generated for the same feature associated with the traffic sign detected in a second image captured in the dark (e.g., at night). Similarly, a descriptor generated for a feature associated with a distinctive part of a building detected in a first image captured under clear conditions (e.g., sunny, no fog or clouds, etc.) may be different from a descriptor generated for the same feature associated with the same part of the building detected in a second image captured under cloudy or rainy conditions.
[0028] In some cases, a machine learning-based system (e.g., using a deep learning neural network) may be used to detect key-point features (e.g., key points or feature points) for localization and mapping and generate descriptors for the detected features. However, it may be difficult to obtain ground truth and annotations (or labels) for training the machine learning-based key-point feature (e.g., key point or feature point) detector and descriptor generator. For example, the benefit of using a machine learning-based system to generate descriptors is that manual (e.g., by a human) annotation of features is not required. However, a human may still be needed to generate a descriptor that will be used as ground truth data (e.g., labeled data) for training the machine learning-based system.
[0029] This document describes systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to as "systems and technologies") for providing a hybrid system for performing feature detection (e.g., detecting key points, also referred to as feature points) and descriptor generation. For example, the hybrid system may include a machine learning-based descriptor generator and a non-machine learning-based feature detector. In some aspects, the non-machine learning-based feature detector (e.g., a feature point detector) may be based on, for example, computer vision algorithms. The non-machine learning-based feature detector may detect or generate feature points (or key points) from input sensor data (e.g., one or more input images, LIDAR data, radar data, etc.).
[0030] The machine learning-based descriptor generator may include or be a machine learning system (e.g., a deep learning neural network) that may generate descriptors for the feature points (or key points) detected by the non-machine learning-based feature detector. The machine learning-based descriptor generator may generate descriptors (also referred to as feature descriptors) at least in part by generating a description of the features detected or depicted by the non-machine learning-based feature detector in the input sensor data (e.g., local image patches extracted around features in an image). In some cases, the feature descriptor may describe the feature as a feature vector or a set of feature vectors.
[0031] In some aspects, the machine learning system for descriptor generation may include a Transformer neural network architecture. For example, a Transformer-based neural network may use Transformer cross-attention (e.g., cross-view attention) to determine a unique signature (which may be used as a feature descriptor) across different types of input data (e.g., images captured during the day, radar data, and / or LIDAR data, images captured during the night, radar data, and / or LIDAR data, images captured when it is raining, radar data, and / or LIDAR data, images captured when it is foggy, radar data, and / or LIDAR data, etc.), thus providing robustness to varying input data. Generating a common or unique descriptor across such varying input data is more difficult to do manually.
[0032] Such hybrid systems allow for the use of machine learning in an unsupervised manner to perform feature detection and descriptor generation (in which case, training does not require labeling). Additionally, the above-described Transformer-based solution (e.g., using cross-attention to generate a unique signature for different types of input data) may scale with more data and may be trained using unsupervised learning (thus not requiring labeled data).
[0033] Various aspects of this application will be described with reference to the drawings. Figure 1is a block diagram illustrating the architecture of an exemplary image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing an image of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture individual images (or photos), and / or can capture video including multiple images (or video frames) in a particular sequence. The lens 115 of the image capture and processing system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the light towards the image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.
[0034] One or more control mechanisms 120 can control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. One or more control mechanisms 120 can include multiple mechanisms and components; for example, one or more control mechanisms 120 can include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 can also include additional control mechanisms other than those illustrated, such as control mechanisms for controlling analog gain, flash, HDR, depth of field, and / or other image capture attributes.
[0035] One or more focus control mechanisms 125B of one or more control mechanisms 120 can obtain a focus setting. In some examples, one or more focus control mechanisms 125B store the focus setting in a memory register. Based on the focus setting, one or more focus control mechanisms 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, one or more focus control mechanisms 125B can move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or a servo system (or other lens mechanism) to adjust the focus. In some cases, additional lenses can be included in the image capture and processing system 100, such as one or more microlenses located over each photodiode of the image sensor 130, each of the one or more microlenses bending the light towards the corresponding photodiode before the light received from the lens 115 reaches the corresponding photodiode. The focus setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The focus setting can be determined using one or more control mechanisms 120, the image sensor 130, and / or the image processor 150. The focus setting can be referred to as an image capture setting and / or an image processing setting.
[0036] One or more exposure control mechanisms 125A of one or more control mechanisms 120 may obtain an exposure setting. In some cases, one or more exposure control mechanisms 125A store the exposure setting in a memory register. Based on the exposure setting, one or more exposure control mechanisms 125A may control the size of the aperture (e.g., aperture size or f-number), the duration for which the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting may be referred to as an image capture setting and / or an image processing setting.
[0037] One or more zoom control mechanisms 125C of one or more control mechanisms 120 may obtain a zoom setting. In some examples, one or more zoom control mechanisms 125C store the zoom setting in a memory register. Based on the zoom setting, one or more zoom control mechanisms 125C may control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, one or more zoom control mechanisms 125C may control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom setting may be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, the focusing lens may be the lens 115), which first receives light from the scene 110, and the light then passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before reaching the image sensor 130. In some cases, the afocal zoom system may include two positive (e.g., converging, convex) lenses having equal or similar focal lengths (e.g., within a threshold difference of each other), with a negative (e.g., diverging, concave) lens therebetween. In some cases, one or more zoom control mechanisms 125C move one or more of the lenses in the afocal zoom system, such as the negative lens and one or two of the positive lenses.
[0038] The image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image produced by the image sensor 130. In some cases, different photodiodes may be covered by different color filters and thus may measure light that matches the color of the color filter covering the photodiode. For example, a Bayer color filter includes a red color filter, a blue color filter, and a green color filter, where each pixel of the image is generated based on red light data from at least one photodiode covered by the red color filter, blue light data from at least one photodiode covered by the blue color filter, and green light data from at least one photodiode covered by the green color filter. Other types of color filters may use yellow, magenta, and / or cyan (also known as "emerald") color filters to replace or supplement the red, blue, and / or green color filters. Some image sensors (e.g., the image sensor 130) may have no color filter at all and may instead use different photodiodes (vertically stacked in some cases) throughout the pixel array. Different photodiodes throughout the pixel array may have different spectral sensitivity curves and thus respond to different wavelengths of light. Monochrome image sensors may also lack a color filter and thus lack color depth.
[0039] In some cases, the image sensor 130 may alternatively or additionally include an opaque mask and / or a reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which can be used for phase detection autofocus (PDAF). The image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output by the photodiode and / or an analog-to-digital converter (ADC) for converting the analog signal output by the photodiode (and / or the analog signal amplified by the analog gain amplifier) into a digital signal. In some cases, certain components or functions discussed for one or more of the control mechanisms 120 may alternatively or additionally be included in the image sensor 130. The image sensor 130 can be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), complementary metal-oxide-semiconductor (CMOS), N-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0040] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more of any other type of processors 1610 discussed for the computing system 1600. The host processor 152 may be a digital signal processor (DSP) and / or other types of processors. In some embodiments, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth TM , Global Positioning System (GPS), etc.), any combination thereof, and / or other components. The I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General-Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 Physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port, and the ISP 154 may communicate with the image sensor 130 using a MIPI port.
[0041] The image processor 150 may perform multiple tasks, such as demosaicking, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving inputs, managing outputs, managing memory, or some combination thereof. The image processor 150 may store the image frames and / or the processed images in a random access memory (RAM) 140 / 1620, a read-only memory (ROM) 145 / 1625, a cache, a memory cell, another storage device, or some combination thereof.
[0042] A variety of input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device 1635, any other input device 1645, or some combination thereof. In some cases, captions may be input into the image processing device 105B through the physical keyboard or keypad of the I / O device 160, or through the virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O port 156 may include one or more ports, jacks, or other connectors that implement a wired connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The I / O port 156 may include one or more wireless transceivers that implement a wireless connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The peripheral devices may include any type of I / O device 160 discussed previously, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they may themselves be considered I / O devices 160.
[0043] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more independent devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some specific embodiments, the image capture device 105A and the image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly via one or more wireless transceivers. In some specific embodiments, the image capture device 105A and the image processing device 105B may be disconnected from each other.
[0044] As Figure 1 shown, the vertical dashed line will Figure 1The image capture and processing system 100 is divided into two parts, representing the image capture device 105A and the image processing device 105B, respectively. The image capture device 105A includes a lens 115, one or more control mechanisms 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, certain components illustrated in the image capture device 105A (such as the ISP 154 and / or the host processor 152) may be included in the image capture device 105A.
[0045] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop computer or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 wi-fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some specific implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile phone, a desktop computer, or other computing device.
[0046] Although the image capture and processing system 100 is shown as including certain components, those of ordinary skill in the art should understand that the image capture and processing system 100 may include more components than Figure 1 those shown therein. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some specific implementations, the components of the image capture and processing system 100 may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a GPU, a DSP, a CPU, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the image capture and processing system 100.
[0047] Figure 2 FIG. 1 is a diagram illustrating the architecture of an example system 200 in accordance with some aspects of the present disclosure. The system 200 can be an XR system (e.g., running (or executing) XR applications and / or implementing XR operations), a vehicle system, a robotic system, or other types of systems. The system 200 can perform tracking and positioning, mapping of the environment (e.g., scene) in the physical world, and / or positioning and rendering of content on the display 209 (e.g., positioning and rendering of virtual content on a screen, visible plane / area, and / or other display as part of an XR experience). For example, the system 200 can generate a map of the environment in the physical world (e.g., a three-dimensional (3D) map), track the pose (e.g., location and orientation) of the system 200 relative to the environment (e.g., relative to the 3D map of the environment), and / or determine positions and / or anchors at specific locations on the map of the environment. In one example, the system 200 can position and / or anchor virtual content at a specific location on the map of the environment and can render the virtual content on the display 209 such that the virtual content appears to be at a location in the environment corresponding to the specific location on the map of the scene where the virtual content is positioned and / or anchored. The display 209 can include a monitor, glasses, a screen, a lens, a projector, and / or other display mechanisms. For example, in the context of an XR system, the display 209 can allow a user to see the real-world environment and can also allow XR content to be overlaid, overlapped, blended, or otherwise displayed thereon.
[0048] In this illustrative example, the system 200 includes one or more image sensors 202, an accelerometer 204, a gyroscope 206, a storage device 207, a computing component 210, a pose engine 220, an image processing engine 224, and a rendering engine 226. It should be noted that Figure 2 the components 202 - 126 shown are non-limiting examples provided for illustrative and explanatory purposes, and other examples can include more, fewer, or different components compared to Figure 2 the components shown. For example, in some cases, the system 200 can include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 2 one or more other software and / or hardware components not shown. Although the various components of the system 200 (such as the image sensor 202) may be referred to herein in the singular form, it should be understood that the system 200 can include multiple of any of the components discussed herein (e.g., multiple image sensors 202).
[0049] System 200 includes an input device 208 or communicates (wired or wirelessly) with the input device. The input device 208 may include any suitable input device, such as a touch screen, a pen or other pointing device, a keyboard, a mouse, buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, a video game controller, a steering wheel, a joystick, a set of buttons, a trackball, a remote control, any other input device 1645 discussed herein, or any combination thereof. In some cases, the image sensor 202 may capture an image that can be processed to interpret a gesture command.
[0050] In some specific implementations, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, pose engine 220, image processing engine 224, and rendering engine 226 may be part of the same computing device. For example, in some cases, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, pose engine 220, image processing engine 224, and rendering engine 226 may be integrated into a device or system, such as an HMD, XR glasses (e.g., AR glasses), a vehicle or vehicle system, a smart phone, a laptop computer, a tablet computer, a gaming system, and / or any other computing device. However, in some specific implementations, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, pose engine 220, image processing engine 224, and rendering engine 226 may be part of two or more separate computing devices. For example, in some cases, some of the components 202-126 may be part of or implemented by one computing device, and the remaining components may be part of or implemented by one or more other computing devices.
[0051] The storage device 207 can be any storage device for storing data. Additionally, the storage device 207 can store data from any of the components in the system 200. For example, the storage device 207 can store data from the image sensor 202 (e.g., image or video data), data from the accelerometer 204 (e.g., measurements), data from the gyroscope 206 (e.g., measurements), data from the computing component 210 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, face recognition data, occlusion data, etc.), data from the pose engine 220, data from the image processing engine 224, and / or data from the rendering engine 226 (e.g., output frames). In some examples, the storage device 207 can include a buffer for storing frames to be processed by the computing component 210.
[0052] One or more computing components 210 can include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, an image signal processor (ISP) 218, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). The computing component 210 can perform various operations such as image enhancement, computer vision, graphics rendering, tracking, positioning, pose estimation, map building, content anchoring, content rendering, image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, the computing component 210 can implement (e.g., control, operate, etc.) the pose engine 220, the image processing engine 224, and the rendering engine 226. In other examples, the computing component 210 can also implement one or more other processing engines.
[0053] The image sensor 202 can include any image and / or video sensor or capture device. In some examples, the image sensor 202 can be part of a multi-camera assembly (such as a dual-camera assembly). The image sensor 202 can capture image and / or video content (e.g., raw image and / or video data), which can then be processed by the computing component 210, the pose engine 220, the image processing engine 224, and / or the rendering engine 226 as described herein. In some examples, the image sensor 202 can include an image capture and processing system 100, an image capture device 105A, an image processing device 105B, or a combination thereof.
[0054] In some examples, the image sensor 202 may capture image data and may generate an image (also referred to as a frame) based on the image data, and / or may provide the image data or frame to the pose engine 220, the image processing engine 224, and / or the rendering engine 226 for processing. The image or frame may include a video frame or a static image in a video sequence. The image or frame may include an array of pixels representing a scene. For example, the image may be a Red-Green-Blue (RGB) image having red, green, and blue color components per pixel; a Luminance, Chroma Red, Chroma Blue (YCbCr) image having a luminance component and two chrominance (color) components (chroma red and chroma blue) per pixel; or any other suitable type of color or monochrome image.
[0055] In some cases, the image sensor 202 (and / or other cameras of the system 200) may also be configured to capture depth information. For example, in some embodiments, the image sensor 202 (and / or other cameras) may include a Red-Green-Blue Depth (RGB-D) camera. In some cases, the system 200 may include one or more depth sensors (not shown) that are separate from the image sensor 202 (and / or other cameras) and may capture depth information. For example, such depth sensors may obtain depth information independently of the image sensor 202. In some examples, the depth sensor may be physically mounted in the same general location as the image sensor 202, but may operate at a different frequency or frame rate than the image sensor 202. In some examples, the depth sensor may take the form of a light source that projects a structured or textured light pattern (which may include one or more narrowband lights) onto one or more objects in a scene. Depth information may then be obtained by exploiting the geometric deformation of the projected pattern caused by the surface shape of the object. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).
[0056] System 200 may also include other sensors in one or more of its sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 204), one or more gyroscopes (e.g., gyroscope 206), and / or other sensors. The one or more sensors may provide speed, orientation, and / or other position-related information to the computing component 210. For example, accelerometer 204 may detect the acceleration of system 200 and generate an acceleration measurement based on the detected acceleration. In some cases, accelerometer 204 may provide one or more translation vectors (e.g., up / down, left / right, front / back), which may be used to determine the position or pose of system 200. Gyroscope 206 may detect and measure the orientation and angular velocity of system 200. For example, gyroscope 206 may be used to measure the pitch, roll, and yaw of system 200. In some cases, gyroscope 206 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, image sensor 202 and / or pose engine 220 may use the measurements obtained by accelerometer 204 (e.g., one or more translation vectors) and / or the measurements obtained by gyroscope 206 (e.g., one or more rotation vectors) to calculate the pose of system 200. As previously mentioned, in other examples, system 200 may also include other sensors, such as an inertial measurement unit (IMU), magnetometer, gaze and / or eye tracking sensors, machine vision sensors, intelligent scene sensors, speech recognition sensors, impact sensors, vibration sensors, position sensors, tilt sensors, etc.
[0057] As described above, in some cases, the one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure the specific force, angular velocity, and / or orientation of system 200. In some examples, the one or more sensors may output information measured in association with the capture of an image captured by image sensor 202 (and / or other cameras of system 200) and / or depth information obtained using one or more depth sensors of system 200.
[0058] The pose engine 220 can determine the pose of the system 200 (also referred to as the head pose) and / or the pose of the image sensor 202 (or other cameras of the system 200) using the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors). In some cases, the pose of the system 200 and the pose of the image sensor 202 (or other cameras) can be the same. The pose of the image sensor 202 refers to the position and orientation of the image sensor 202 relative to a reference frame (e.g., for an object). In some specific implementations, the camera pose can be determined for six degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by the X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as the image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some specific implementations, the camera pose can be determined for three degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).
[0059] In some cases, a device tracker (not shown) can use measurements from one or more sensors and image data from the image sensor 202 to track the pose of the system 200 (e.g., 6DoF pose). For example, the device tracker can fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and motion of the system 200 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the system 200, the device tracker can generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate an update to the 3D map for the scene. The 3D map update can include, for example but not limited to, new or updated features and / or landmarks or fiducial points associated with the scene and / or the 3D map of the scene, and / or a localization update that identifies or updates the position of the system 200 within the scene and the 3D map of the scene. The 3D map can provide a digital representation of the scene in the real / physical world. In some examples, the 3D map can anchor location-based objects and / or content to real-world coordinates and / or objects. The system 200 can use the map-built scene (e.g., the scene in the physical world represented by and / or associated with the 3D map) to merge the physical world and the virtual world and / or to merge virtual content or objects with the physical environment.
[0060] In some aspects, the computing component 210 may determine and / or track the pose (also referred to as the camera pose) of the image sensor 202 and / or the system 200 as a whole based on images captured by the image sensor 202 (and / or other cameras of the system 200) using a visual tracking solution. For example, in some examples, the computing component 210 may use computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques to perform the tracking. For example, the computing component 210 may perform SLAM or may communicate (wired or wirelessly) with a SLAM system ( Figure 2 not shown in the figure)(such as Figure 3 SLAM system 300). SLAM refers to a class of techniques that create a map of the environment (e.g., a map of the environment modeled by the system 200) while tracking the camera (e.g., the image sensor 202) and / or the pose of the system 200 relative to the map. This map may be referred to as a SLAM map and may be three-dimensional (3D). SLAM techniques may use color or grayscale image data captured by the image sensor 202 (and / or other cameras of the system 200) and may be used to generate an estimate of the 6DoF pose measurements of the image sensor 202 and / or the system 200. Such SLAM techniques configured to perform 6DoF tracking may be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) may be used to estimate, correct, and / or otherwise adjust the estimated pose.
[0061] In some cases, 6DoF SLAM (e.g., 6DoF tracking) may associate features (e.g., key points) observed from certain input images from the image sensor 202 (and / or other cameras or sensors) to the SLAM map. For example, 6DoF SLAM may use feature point associations from the input images (or other sensor data, such as radar sensors, LIDAR sensors, etc.) to determine the pose (position and orientation) of the image sensor 202 and / or the system 200 of the input images. 6DoF mapping may also be performed to update the SLAM map. In some cases, the SLAM map maintained using 6DoF SLAM may contain 3D feature points (e.g., key points) triangulated from two or more images. For example, key frames may be selected from the input images or video stream to represent the observed scene. For each key frame, the corresponding 6DoF camera pose associated with the image may be determined. The pose of the image sensor 202 and / or the system 200 may be determined by projecting features (e.g., feature points or key points) from the 3D SLAM map into the image or video frame and updating the camera pose based on the verified 2D-3D correspondences.
[0062] In an illustrative example, the computing component 210 may extract feature points (e.g., key points) from certain input images (e.g., each input image, a subset of input images, etc.) or from each key frame. Feature points (also referred to as key points or registration points) as used herein are distinctive or identifiable portions of an image, such as a part of a hand, an edge of a table, etc. Features extracted from the captured images may represent different feature points along a three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and each feature point may have an associated feature location. Feature points in a key frame match (are the same as or correspond to) or fail to match feature points of a previously captured input image or key frame. Feature detection may be used to detect feature points. Feature detection may include image processing operations for examining one or more pixels of an image to determine whether a feature exists at a particular pixel. Feature detection may be used to process the entire captured image or certain portions of the image. For each image or key frame, once features have been detected, local image patches around the features may be extracted. Any suitable technique may be used to extract features, such as Scale-Invariant Feature Transform (SIFT) (which localizes features and generates their descriptions), Learned Invariant Feature Transform (LIFT), Speeded-Up Robust Features (SURF), Gradient Location-Orientation Histogram (GLOH), Oriented FAST and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoint (FREAK), KAZE, Accelerated KAZE (AKAZE), Normalized Cross-Correlation (NCC), descriptor matching, another suitable technique, or a combination thereof.
[0063] In some cases, the system 200 may also track a user's hand and / or fingers to allow the user to interact with and / or control virtual content in a virtual environment. For example, the system 200 may track the pose and / or movement of a user's hand and / or fingertips to identify or translate the user's interaction with the virtual environment. User interaction may include, for example but not limited to, moving a virtual content item, resizing a virtual content item, selecting an input interface element in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interfaces), providing input through a virtual user interface, etc.
[0064] Figure 3 is a block diagram illustrating the architecture of a Simultaneous Localization and Mapping (SLAM) system 300. In some examples, the SLAM system 300 may be Figure 2The system 200 can include the system, or can be a part of the system. In some examples, the SLAM system 300 can be, can include, or can be a part of the following: an XR device, an autonomous vehicle, a vehicle, a computing system of a vehicle, a wireless communication device, a mobile device or a cell phone (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device (e.g., a connected watch), a personal computer, a laptop computer, a server computer, a portable video game console, a portable media player, a camera device, a manned or unmanned ground vehicle, a manned or unmanned aerial vehicle, a manned or unmanned water vehicle, a manned or unmanned underwater vehicle, a manned or unmanned vehicle, a robot, another device, or any combination thereof.
[0065] Figure 3 Each sensor of the one or more sensors 305 of the SLAM system 300 is included in or coupled to each sensor. The one or more sensors 305 can include one or more cameras 310. Each camera of the one or more cameras 310 can include an image capture device 105A, an image processing device 105B, an image capture and processing system 100, another type of camera, or a combination thereof. Each camera of the one or more cameras 310 can respond to light from a specific spectrum. The spectrum can be a subset of the electromagnetic (EM) spectrum. For example, each camera of the one or more cameras 310 can be a VL camera that responds to the visible light (VL) spectrum, an IR camera that responds to the infrared (IR) spectrum, a UV camera that responds to the ultraviolet (UV) spectrum, a camera that responds to light of another spectrum from another part of the electromagnetic spectrum, or some combination thereof.
[0066] The one or more sensors 305 can include one or more other types of sensors in addition to the cameras 310, such as one or more of the following: an accelerometer, a gyroscope, a magnetometer, an inertial measurement unit (IMU), an altimeter, a barometer, a thermometer, a radio detection and ranging (RADAR) sensor, a light detection and ranging (LIDAR) sensor, a sound navigation and ranging (SONAR) sensor, a sound detection and ranging (SODAR) sensor, a global navigation satellite system (GNSS) receiver, a global positioning system (GPS) receiver, a Beidou navigation satellite system (BDS) receiver, a Galileo receiver, a Globalnaya Navigazionnaya Sputnikovaya Sistema (GLONASS) receiver, a Navigation Indian Constellation (NavIC) receiver, a quasi-zenith satellite system (QZSS) receiver, a Wi-Fi positioning system (WPS) receiver, a cellular network positioning system receiver, A beacon positioning receiver, a short-range wireless beacon positioning receiver, a personal area network (PAN) positioning receiver, a wide area network (WAN) positioning receiver, a wireless local area network (WLAN) positioning receiver, other types of positioning receivers, other types of sensors discussed herein, or combinations thereof. In some examples, one or more sensors 305 may include Figure 2 any combination of sensors of the system 200.
[0067] Figure 3 The SLAM system 300 of Figure 3 includes a visual-inertial odometry (VIO) tracker 315. The term visual-inertial odometry may also be referred to as visual odometry herein. The VIO tracker 315 receives sensor data 365 from one or more sensors 305. For example, the sensor data 365 may include one or more images captured by one or more cameras 310. The sensor data 365 may include other types of sensor data from one or more sensors 305, such as data from any of the types of sensors 305 listed herein. For example, the sensor data 365 may include IMU data from one or more inertial measurement units (IMU) of one or more sensors 305.
[0068] After receiving the sensor data 365 from one or more sensors 305, the VIO tracker 315 performs feature detection, extraction, and / or tracking using the feature tracking engine 320 of the VIO tracker 315. For example, in the case where the sensor data 365 includes one or more images captured by one or more cameras 310 of the SLAM system 300, the VIO tracker 315 may identify, detect, and / or extract features in each image. Features may include visually distinct points in the image, such as portions of the image depicting edges and / or corners. The VIO tracker 315 may periodically and / or continuously receive the sensor data 365 from one or more sensors 305, such as by continuously receiving more images from one or more cameras 310 as the one or more cameras 310 capture video, where the images are video frames of the video. The VIO tracker 315 may generate descriptors of the features. The feature descriptors may be generated at least in part by generating a description of the features, as depicted in local image patches extracted around the features. In some examples, the feature descriptors may describe the features as a set of one or more feature vectors. In some cases, the VIO tracker 315 may be implemented using the hybrid system 500 discussed below for Figure 5 discussed.
[0069] The VIO tracker 315 (which may have a map building engine 330 and / or a relocalization engine 355 in some cases) can associate multiple features with a map of the environment based on such feature descriptors. The feature tracking engine 320 of the VIO tracker 315 can perform feature tracking by identifying features in each image that the VIO tracker 315 has previously identified in one or more previous images (in some cases, based on identifying features with matching feature descriptors in different images). The feature tracking engine 320 can track changes in the one or more positions depicting the features in each different image. For example, the feature extraction engine can detect a specific corner of a room depicted on the left side of a first image captured by the first camera in the camera 310. The feature extraction engine can detect the same feature (e.g., the same specific corner of the same room) depicted on the right side of a second image captured by the first camera. The feature tracking engine 320 can identify that the features detected in the first image and the second image are two depictions of the same feature (e.g., the same specific corner of the same room), and that the feature appears at two different positions in the two images. The VIO tracker 315 can determine that the first camera has moved based on the same feature appearing on the left side of the first image and the right side of the second image, for example, if the feature (e.g., a specific corner of a room) depicts a static part of the environment.
[0070] The VIO tracker 315 can include a sensor integration engine 325. The sensor integration engine 325 can use sensor data from other types of sensors 305 (in addition to the camera 310) to determine information that the feature tracking engine 320 can use when performing feature tracking. For example, the sensor integration engine 325 can receive IMU data from the IMU of one or more sensors 305 (which may be included as part of the sensor data 365). The sensor integration engine 325 can determine, based on the IMU data in the sensor data 365, that from the acquisition or capture of the first image by the first camera in the camera 310 to the acquisition or capture of the second image, the SLAM system 300 has rotated 15 degrees in the clockwise direction. Based on this determination, the sensor integration engine 325 can identify that a feature depicted at a first position in the first image is expected to appear at a second position in the second image, and that the second position is expected to be a predetermined distance to the left of the first position (e.g., a predetermined number of pixels, inches, centimeters, millimeters, or another distance metric). The feature tracking engine 320 can consider this expectation when tracking features between the first image and the second image.
[0071] Based on feature tracking performed by the feature tracking engine 320 and / or sensor integration performed by the sensor integration engine 325, the VIO tracker 315 can determine the 3D feature positions 372 of specific features. The 3D feature positions 372 can include one or more 3D feature positions and can also be referred to as 3D feature points. The 3D feature positions 372 can be a set of coordinates along three different axes perpendicular to each other, such as an X coordinate along the X axis (e.g., in the horizontal direction), a Y coordinate along the Y axis perpendicular to the X axis (e.g., in the vertical direction), and a Z coordinate along the Z axis perpendicular to both the X axis and the Y axis (e.g., in the depth direction). In some aspects, the VIO tracker 315 can also determine one or more key frames 370 (hereinafter referred to as key frames 370) corresponding to a specific feature. A key frame corresponding to a specific feature (from one or more key frames 370) can be an image that clearly depicts the specific feature. In some examples, a key frame corresponding to a specific feature (from one or more key frames 370) can be an image that clearly depicts the specific feature. In some examples, a key frame corresponding to a specific feature can be an image that reduces the uncertainty of the 3D feature position 372 of the specific feature when considered by the feature tracking engine 320 and / or the sensor integration engine 325 for determining the 3D feature position 372. In some examples, a key frame corresponding to a specific feature also includes data about the pose 385 of the SLAM system 300 and / or the camera 310 during the capture of the key frame. In some examples, the VIO tracker 315 can transmit the 3D feature positions 372 and / or the key frames 370 corresponding to one or more features to the map building engine 330. In some examples, the VIO tracker 315 can receive map slices 375 from the map building engine 330. The VIO tracker 315 can use the information within the map slices 375 as features for feature tracking using the feature tracking engine 320.
[0072] Based on feature tracking performed by the feature tracking engine 320 and / or sensor integration performed by the sensor integration engine 325, the VIO tracker 315 can determine the pose 385 of the SLAM system 300 and / or the camera 310 during the capture of each image in the sensor data 365. The pose 385 can include the position of the SLAM system 300 and / or the camera 310 in 3D space, such as a set of coordinates along three different axes perpendicular to each other (e.g., an X coordinate, a Y coordinate, and a Z coordinate). The pose 385 can include the orientation of the SLAM system 300 and / or the camera 310 in 3D space, such as pitch, roll, yaw, or some combination thereof. In some examples, the VIO tracker 315 can transmit the pose 385 to the relocalization engine 355. In some examples, the VIO tracker 315 can receive the pose 385 from the relocalization engine 355.
[0073] The SLAM system 300 further includes a map construction engine 330. The map construction engine 330 can generate a 3D map of the environment based on the 3D feature positions 372 and / or key frames 370 received from the VIO tracker 315. The map construction engine 330 can include a map densification engine 335, a key frame remover 340, a beam adjuster 345, and / or a loop closure detector 350. The map densification engine 335 can perform map densification, in some examples, increasing the number and / or density of 3D coordinates that describe the map geometry. The key frame remover 340 can remove key frames, and / or in some cases add key frames. In some examples, the key frame remover 340 can remove key frames 370 corresponding to regions in the map to be updated and / or with low corresponding confidence values. In some examples, the beam adjuster 345 can refine the 3D coordinates that describe the scene geometry, the parameters of the relative motion, and / or the optical characteristics of the image sensor used to generate the frame according to an optimality criterion involving the corresponding image projections of all points. The loop closure detector 350 can identify when the SLAM system 300 has returned to a previously mapped area and can use such information to update map slices and / or reduce the uncertainty of certain 3D feature points or other points in the map geometry.
[0074] The map construction engine 330 can output map slices 375 to the VIO tracker 315. The map slices 375 can represent a 3D portion or subset of the map. The map slices 375 can include map slices 375 representing new, previously unmapped areas of the map. The map slices 375 can include map slices 375 representing updates (or modifications or revisions) to previously mapped areas of the map. The map construction engine 330 can output map information 380 to the relocalization engine 355. The map information 380 can include at least a portion of the map generated by the map construction engine 330. The map information 380 can include one or more 3D points that make up the geometry of the map, such as one or more 3D feature positions 372. The map information 380 can include one or more key frames 370 corresponding to certain features and certain 3D feature positions 372.
[0075] The SLAM system 300 further includes a relocalization engine 355. The relocalization engine 355 can perform relocalization, for example, when the VIO tracker 315 fails to identify a threshold number of features in an image, and / or when the VIO tracker 315 loses track of the pose 385 of the SLAM system 300 within the map generated by the map construction engine 330. The relocalization engine 355 can perform relocalization by performing extraction and matching using the extraction and matching engine 360. For example, the extraction and matching engine 360 can extract features from an image captured by the camera 310 of the SLAM system 300 when the SLAM system 300 is at the current pose 385, and can match the extracted features with features depicted in different keyframes 370, identified by 3D feature positions 372, and / or identified in the map information 380. By matching these extracted features with previously identified features, the relocalization engine 355 can identify that the pose 385 of the SLAM system 300 is the pose 385 at which the previously identified features are visible to the camera 310 of the SLAM system 300, and thus is similar to one or more previous poses 385 at which the previously identified features were visible to the camera 310. In some cases, the relocalization engine 355 can perform relocalization based on wide-baseline map construction or the distance between the current camera position and the camera position at which the features were originally captured. The relocalization engine 355 can receive information about the pose 385 (e.g., information about one or more recent poses of the SLAM system 300 and / or the camera 310) from the VIO tracker 315, and the relocalization engine 355 can base its relocalization determination on this information. Once the relocalization engine 355 relocalizes the SLAM system 300 and / or the camera 310 and thus determines the pose 385, the relocalization engine 355 can output the pose 385 to the VIO tracker 315.
[0076] Figure 4 An example frame 400 of a scene is illustrated. Frame 400 provides an illustrative example of feature information that can be captured and / or processed by a system (e.g., Figure 2 the system 200 shown) during tracking and / or map construction. In Figure 4In the example shown, the example feature 402 is illustrated as circles of different diameters. In some cases, the center of each of the features 402 may be referred to as the feature center position. In some cases, the diameter of the circle may represent the feature scale (also referred to as the blob size) associated with each of the example features 402. Each of the features 402 may also include a dominant orientation vector 403 illustrated as a radial line segment. In one illustrative example, the dominant orientation vector 403 (also referred to herein as the dominant orientation) may be determined based on the pixel gradients within a block (also referred to as a blob or region). For example, the dominant orientation vector 403 may be determined based on the orientation of edge features within a neighborhood around the center of the feature (e.g., a block of nearby pixels). Another example feature 404 is shown as having a dominant orientation 406. In some implementations, a feature may have multiple dominant orientations. For example, if no single orientation is clearly dominant, the feature may have two or more dominant orientations associated with the most prominent orientations. Another example feature 408 is illustrated as having two dominant orientation vectors 410 and 412. In addition to the feature center position, blob size, and dominant orientation, each of the features 402, 404, 408 may also be associated with a descriptor that can be used to correlate features between different frames. For example, if the pose of the camera capturing frame 400 changes, the x-y coordinates of each of the feature center positions of each of the features 402, 404, 408 may also change, and the descriptors assigned to each feature can be used to match features between two different frames. In some cases, the tracking and mapping operations of the XR system may utilize different types of descriptors for the features 402, 404, 408. Examples of descriptors for the features 402, 404, 408 may include SIFT, FREAK, and / or other descriptors. In some cases, the tracker may operate directly on the image blocks or may operate on the descriptors (e.g., SIFT descriptors, FREAK descriptors, etc.).
[0077] As previously mentioned, in some cases, machine learning-based systems (e.g., using deep learning neural networks) may be used to detect features (e.g., key points or feature points) for localization and mapping and generate descriptors for the detected features. However, it may be difficult to obtain the ground truth and annotations (or labels) for training machine learning-based feature (e.g., key point or feature point) detectors and descriptor generators.
[0078] The systems and techniques described herein provide a hybrid system for performing feature detection to detect feature points (or key points) and descriptor generation to generate feature descriptors. As described herein, the hybrid system can include a non-machine learning based feature detector for detecting feature points (e.g., a feature point detector based on, for example, computer vision algorithms) and a machine learning based descriptor generator for generating descriptors (e.g., feature descriptors) for the detected feature points (or key points) (e.g., a deep learning neural network). The machine learning based descriptor generator can generate descriptors at least in part by generating a description of the features detected or depicted in the input sensor data.
[0079] Figure 5 is a diagram illustrating an example of a hybrid system 500 for detecting features (e.g., key points or feature points) and generating descriptors for the detected features (e.g., feature descriptors). As described above, in some cases, the above for Figure 4 the VIO tracker 315 described can use Figure 5 the hybrid system 500 to implement. The hybrid system 500 includes a feature point detector 504 and a machine learning (ML) based descriptor generator 506.
[0080] The feature point detector 504 is a non-machine learning based feature detector. For example, the feature point detector 504 can use, for example, one or more computer vision algorithms to detect feature points (or key points) from the input data 502. The input data 502 can include image data, radar data (e.g., radar images), LIDAR data (e.g., LIDAR point clouds), and / or other sensor data. In one illustrative example, as Figure 5 shown, the input data 502 can include an image 503 of a scene under a first lighting condition (e.g., during the day), an image 505 of the same scene under a second lighting condition (e.g., at night), and an image 507 of the same scene when there are specific weather conditions (e.g., fog, rain, etc.). Multiple sets of images with similar differences can also be included in the input data.
[0081] In some illustrative examples, the input data 502 can include other images of the same scene but from different angles during a first lighting condition (e.g., during the day), during a second lighting condition (e.g., at night), and under the same or different weather conditions. Additionally or alternatively, in some illustrative examples, the input data 502 can include images of different scenes during the day, at night, and under the same or different weather conditions.
[0082] The feature point detector 504 can detect key points in images 503, 505, 507. The feature point detector 504 can output patches 509 around the feature points (or key points) detected in image 503, patches 511 around the same feature points (or key points) detected in image 505, and patches 513 around the same feature points (or key points) detected in image 507. Similar patches can be generated for other feature points detected in images 503, 505, 507 and in other images and / or sensor data.
[0083] Patches 509, 511, and 513 can be output to the ML-based descriptor generator 506 for descriptor generation. The ML-based descriptor generator 506 can process patches 509, 511, and 513 to generate feature descriptors of the features in each of the patches 509, 511, and 513. Each feature descriptor can describe the corresponding feature as a feature vector or a set of feature vectors, as described above for Figure 3 described.
[0084] In some aspects, the ML-based descriptor generator 506 can include a transformer-based neural network having a transformer neural network architecture. The transformer-based neural network can use transformer cross-attention (e.g., cross-view attention across sensor data from different views or perspectives of a common feature) to determine a unique signature across different types of input data 502. This unique signature can then be used as a feature descriptor.
[0085] A loss function can be used to train the ML-based descriptor generator 506 (e.g., via backpropagation gradients determined based on the loss determined by the loss function). The loss function can enforce the same descriptors across different characteristics of the input data 502 (e.g., across all lighting conditions, weather, etc.). In such cases, no labeling is required, and thus the ML model of the ML-based descriptor generator 506 can be trained in an unsupervised manner.
[0086] The hybrid system 500 provides advantages over traditional systems for detecting features and generating detectors. For example, as previously mentioned, when sensor data is captured under different conditions (e.g., different lighting and / or illumination, different views, different weather conditions, etc.), it may be difficult to generate a common descriptor for the same features detected in different sensor data (e.g., images, radar data, LIDAR data, etc.). In one example, the descriptor generated for a feature associated with the edge of a traffic sign detected in a first image captured during the day may be different from the descriptor generated for the same feature associated with the edge of a traffic sign detected in a second image captured in the dark (e.g., at night). For example, during the day, the background behind the traffic sign may be blue (e.g., corresponding to the blue sky), while at night, the background behind the traffic sign may be black (e.g., corresponding to the dark sky). Traditional computer vision algorithms take the neighboring pixels around the edge or other distinguishing parts of an object (e.g., a traffic sign) as input and encode the information associated with the neighboring pixels (e.g., based on gradients, how the color changes, etc.) to generate a descriptor. However, if the color and / or brightness of the neighboring pixels during the day (e.g., blue and high illumination) is different from the color and / or brightness during the night (e.g., black and low illumination), the two different descriptors for the same feature may be different. If a device or system (e.g., a vehicle, an XR device, a robotic device, etc.) receives two descriptors corresponding to the feature of the edge of a traffic sign, the device or system will determine that the two descriptors correspond to different locations, while the feature (e.g., the edge of the traffic sign) actually corresponds to the same location. Thus, using such techniques will result in inaccurate positioning by the device or system.
[0087] The hybrid system 500 allows the use of non-ML techniques to perform feature detection and uses machine learning in an unsupervised manner to perform descriptor generation (in which case, training does not require manual labeling). The ML-based descriptor generator 506 (e.g., utilizing a transformer-based architecture) provides robustness to such varying input data by generating a common or unique descriptor across the varying input data. Additionally, the ability of the ML-based descriptor generator 506 (e.g., using cross-attention based on a transformer-based architecture) to generate unique signatures for features detected in different types of input data can scale with more data and does not require labeled data.
[0088] Figure 6 is a flowchart illustrating an example of a process 600 for processing image and / or video data. The process 600 may be executed by a computing device (or apparatus) or by a component or system of a computing device (e.g., a chipset). The computing device (or its component or system) may include Figure 5The hybrid system 500 or can be the hybrid system. The operations of process 1000 can be implemented as software components executed and run on one or more processors (e.g., Figure 7 processor 710 or other processors). In addition, the first network entity can be enabled to transmit and receive signals in process 600, for example, via one or more antennas and / or one or more transceivers (such as a wireless transceiver).
[0089] At block 602, a computing device (or its components or systems) can obtain input data (e.g., input data 502). In some cases, the input data includes one or more images, radar data, light detection and ranging (LIDAR) data, any combination thereof, and / or other data. In an illustrative example, the input data includes a first image of a scene having a first characteristic, a second image of a scene having a second characteristic, and a third image of a scene having a third characteristic. For example, the first characteristic can be a daytime characteristic, the second characteristic can be a nighttime characteristic, and the third characteristic can be a weather condition, as Figure 5 shown in the illustrative example of
[0090] At block 604, a computing device (or its components or systems) can process the input data using a non-machine learning-based feature detector (e.g., feature point detector 504) to determine one or more feature points in the input data. In an illustrative example, the non-machine learning-based feature detector is configured to determine one or more feature points based on computer vision algorithms, as described herein (e.g., for Figure 5 ).
[0091] At block 606, a computing device (or its components or systems) may use a machine learning system (e.g., the ML-based descriptor generator 506) to determine a respective feature descriptor for each respective feature point among one or more feature points. In one illustrative example, the machine learning system is a neural network. For example, the neural network can be a transformer neural network or can include a transformer neural network. In some cases, as described herein, the transformer neural network is configured to perform cross-attention. For example, the computing device (or its components or systems) may utilize the transformer neural network to perform transformer cross-attention (e.g., cross-view attention) to determine a unique signature across the obtained input data (e.g., different types of input data 502). This unique signature can be used as a feature descriptor. For example, the respective feature descriptor for each respective feature point can be based on this unique signature. The computing device (or its components or systems) may apply a loss function to train the transformer neural network (e.g., via backpropagation gradients determined based on the loss determined by the loss function). As described above, the loss function enforces the same descriptor across different characteristics of the input data 502 (e.g., across all lighting conditions, weather, etc.), thereby providing robustness to different inputs with varying characteristics.
[0092] The computing device (or apparatus) can include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a connected watch or smartwatch or other wearable device), a server computer, a vehicle (e.g., an autonomous vehicle or a semi-autonomous vehicle) or a computing device or system of a vehicle, a robotic device, a laptop computer, a smart TV, a camera, and / or any other computing device having the resource capabilities to perform the processes described herein (including process 600 and / or any other process described herein). In some cases, the computing device or apparatus can include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device can include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface can be configured to communicate and / or receive data based on Internet Protocol (IP) or other types of data.
[0093] Components of a computing device may be implemented in circuitry. For example, a component may include or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0094] Process 600 is illustrated as a logic flow diagram, and the operations of the logic flow diagram represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. In general, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0095] Additionally, process 600 and / or any other process described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, e.g., in the form of a computer program including multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0096] Figure 7 Example computing device architecture 700 of an example computing device that may implement the various techniques described herein is illustrated. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device of a vehicle), or other devices. For example, computing device architecture 700 may implement Figure 5System 500. The components of computing device architecture 700 are shown to communicate electrically with each other using a connector 705 (such as a bus). An example computing device architecture 700 includes a processing unit (CPU or processor) 710 and a computing device connector 705 that couples various computing device components including computing device memory 715 (such as read only memory (ROM) 720 and random access memory (RAM) 725) to the processor 710.
[0097] Computing device architecture 700 may include a cache of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 710. Computing device architecture 700 may copy data from memory 715 and / or storage device 730 to cache 712 for quick access by the processor 710. In this way, the cache can provide a performance boost that avoids delays while the processor 710 waits for data. These engines and other engines may control or be configured to control the processor 710 to perform various actions. Other computing device memory 715 may also be used. Memory 715 may include a variety of different types of memory with different performance characteristics. The processor 710 may include any general-purpose processor and hardware or software services (such as service 1 732, service 2 734, and service 3 736 stored in storage device 730) configured to control the processor 710, as well as a dedicated processor in which software instructions are incorporated into the processor design. The processor 710 may be a self-contained system that includes multiple cores or processors, buses, memory controllers, caches, etc. A multi-core processor may be symmetric or asymmetric.
[0098] To enable user interaction with computing device architecture 700, input device 745 may represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. Output device 735 may also be one or more of a number of output mechanisms known to those skilled in the art, such as a display, a projector, a television, a speaker device, etc. In some cases, a multimodal computing device may enable a user to provide multiple types of input to communicate with computing device architecture 700. Communication interface 740 generally may govern and manage user input and computing device output. There are no limitations on operating on any particular hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0099] The storage device 730 is a non-volatile memory and can be a hard disk or other type of computer-readable medium that can store computer-accessible data, such as a magnetic tape cassette, a flash memory card, a solid-state memory device, a digital versatile disc, a cassette tape, a random access memory (RAM) 725, a read-only memory (ROM) 720, and hybrid forms thereof. The storage device 730 may include services 732, 734, 736 for controlling the processor 710. Other hardware or software modules or engines are contemplated. The storage device 730 may be connected to the computing device connector 705. In one aspect, a hardware module that performs a specific function may include software components stored in a computer-readable medium connected to necessary hardware components (such as the processor 710, the connector 705, the output device 735, etc.) to perform the function.
[0100] Aspects of the present disclosure are applicable to any suitable electronic device (such as a security system, a smart phone, a tablet computer, a laptop computer, a vehicle, a drone, or other devices) that includes or is coupled to one or more active depth sensing systems. Although the following is described with respect to a device having or coupled to one light projector, aspects of the present disclosure are applicable to devices having any number of light projectors and are thus not limited to a particular device.
[0101] The term "device" is not limited to one or a particular number of physical objects (such as a smart phone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that can implement at least some portions of the present disclosure. Although the following description and examples use the term "device" to describe aspects of the present disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or a particular implementation. For example, a system can be implemented on one or more printed circuit boards or other substrates and can have movable or static components. Although the following description and examples use the term "system" to describe aspects of the present disclosure, the term "system" is not limited to a particular configuration, type, or number of objects.
[0102] Specific details were provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, one of ordinary skill in the art will understand that the embodiments may be practiced without these specific details. For clarity of illustration, in some instances, the present technology may be presented as including separate functional blocks, including functional blocks that include devices, device components, steps or routines in a method embodied as software or a combination of hardware and software. Additional components other than those shown and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components to avoid obscuring these embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0103] Individual embodiments may be described above as a process or method depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. A process is terminated when the operations of the process are complete, but the process may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.
[0104] The processes and methods according to the above examples may be implemented using computer-executable instructions stored or otherwise obtained from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or a group of functions. Portions of the computer resources used may be accessed over a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.
[0105] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. A computer-readable medium may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media (such as flash memory), memories or memory devices, magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, compact discs (CDs) or digital versatile discs (DVDs), any suitable combination thereof, and the like. A computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent a process, function, subroutine, program, routine, subroutine, module, engine, software package, class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network transmission, and the like.
[0106] In some embodiments, computer-readable storage devices, media, and memories may include wired or wireless signals that include bitstreams and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.
[0107] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and may take any form factor among a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks (e.g., a computer program product) may be stored in a computer-readable or machine-readable medium. A processor may execute the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, and the like. The functionality described herein may also be embodied in peripheral devices or plug-in cards. By additional example, such functionality may also be implemented on a circuit board among different chips or different processes executed on a single device.
[0108] Instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.
[0109] In the foregoing description, aspects of the present application have been described with reference to specific embodiments of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although exemplary embodiments of the present application have been described in detail herein, it is to be understood that the inventive concept can be embodied and employed in various other ways, and the appended claims are intended to be construed to include such variations, unless limited by the prior art. The various features and aspects of the foregoing application may be used singly or in combination. Additionally, the embodiments can be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods can be performed in a different order than that described.
[0110] One of ordinary skill in the art will appreciate that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced, respectively, with the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols without departing from the scope of the specification.
[0111] Where a component is described as “configured to” perform certain operations, such configuration can be achieved, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0112] The phrase “coupled to” means that any component is directly or indirectly physically connected to another component, and / or any component directly or indirectly communicates with another component (e.g., is connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0113] Claim language or other language reciting "at least one" of a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one" of a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.
[0114] Claim language or other language reciting "at least one processor, the at least one processor being configured to", "at least one processor being configured to", "one or more processors, the one or more processors being configured to", "one or more processors being configured to", etc. indicates that one processor or multiple processors (in any combination) can perform the associated operations. For example, claim language reciting "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or multiple processors are each assigned tasks for a particular subset of operations X, Y, and Z such that the multiple processors together perform X, Y, and Z; or a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting "at least one processor, the at least one processor being configured to: X, Y, and Z" can mean that any single processor can perform at least a subset of operations X, Y, and Z.
[0115] When referring to one or more elements that perform functions (e.g., the steps of a method), one element may perform all the functions, or more than one element may perform these functions jointly. When more than one element performs these functions jointly, each function does not need to be performed by each of these elements (e.g., different functions may be performed by different elements), and / or each function does not need to be performed entirely by only one element (e.g., different elements may perform different sub - functions of a function). Similarly, when referring to one or more elements that are configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all the functions, or more than one element may be jointly configured to cause another element to perform these functions.
[0116] When referring to an entity (e.g., any entity or device described herein) that performs functions or is configured to perform functions (e.g., the steps of a method), the entity may be configured to cause one or more elements (individually or jointly) to perform these functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of these functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all the functions, or to cause more than one component to perform these functions jointly. When the entity is configured to cause more than one component to perform these functions jointly, each function does not need to be performed by each of these components (e.g., different functions may be performed by different components), and / or each function does not need to be performed entirely by only one component (e.g., different components may perform different sub - functions of a function).
[0117] The various illustrative logical blocks, modules, engines, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0118] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device such as a cellular phone, or an integrated circuit device having multiple uses, including applications in wireless communication devices such as cellular phones and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium including program code, the program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially implemented by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0119] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.
[0120] Exemplary aspects of the present disclosure include:
[0121] Aspect 1. An apparatus for processing image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain input data; process the input data using a machine learning - based feature detector to determine one or more feature points in the input data; and use a machine learning system to determine a corresponding feature descriptor for each respective feature point among the one or more feature points.
[0122] Aspect 2. The apparatus according to aspect 1, wherein the input data includes at least one of one or more images, radar data, or light detection and ranging (LIDAR) data.
[0123] Aspect 3. The apparatus according to any one of aspects 1 or 2, wherein the input data includes a first image of a scene having a first characteristic, a second image of a scene having a second characteristic, and a third image of a scene having a third characteristic.
[0124] Aspect 4. The apparatus according to aspect 3, wherein the first characteristic is a daytime characteristic, the second characteristic is a nighttime characteristic, and the third characteristic is a weather condition.
[0125] Aspect 5. The apparatus according to any one of aspects 1 to 4, wherein the machine learning - based feature detector is configured to determine the one or more feature points based on a computer vision algorithm.
[0126] Aspect 6. The apparatus according to any one of aspects 1 to 5, wherein the machine learning system is a neural network.
[0127] Aspect 7. The apparatus according to aspect 6, wherein the neural network is a transformer neural network.
[0128] Aspect 8. A method for processing image data, the method comprising: obtaining input data; processing the input data using a machine learning - based feature detector to determine one or more feature points in the input data; and using a machine learning system to determine a corresponding feature descriptor for each respective feature point among the one or more feature points.
[0129] Aspect 9. The method according to aspect 8, wherein the input data includes at least one of one or more images, radar data, or light detection and ranging (LIDAR) data.
[0130] Aspect 10. The method according to any one of aspects 8 or 9, wherein the input data includes a first image of a scene having a first characteristic, a second image of a scene having a second characteristic, and a third image of a scene having a third characteristic.
[0131] Aspect 11. The method according to aspect 10, wherein the first characteristic is a daytime characteristic, the second characteristic is a nighttime characteristic, and the third characteristic is a weather condition.
[0132] Aspect 12. The method according to any one of aspects 8 to 11, the method further comprising determining the one or more feature points using the non-machine learning based feature detector based on a computer vision algorithm.
[0133] Aspect 13. The method according to any one of aspects 8 to 12, wherein the machine learning system is a neural network.
[0134] Aspect 14. The method according to aspect 13, wherein the neural network is a transformer neural network.
[0135] Aspect 15. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform any of the operations according to any one of aspects 8 to 14.
[0136] Aspect 16. An apparatus comprising components for performing any of the operations according to any one of aspects 8 to 14.
Claims
1. An apparatus for processing image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain input data; process the input data using a machine learning - based feature detector to determine one or more feature points in the input data; and use a machine learning system to determine a corresponding feature descriptor for each respective feature point among the one or more feature points.
2. The apparatus according to claim 1, wherein the input data comprises at least one of one or more images, radar data, or light detection and ranging (LIDAR) data.
3. The apparatus according to claim 1, wherein the input data comprises a first image of a scene having a first characteristic, a second image of a scene having a second characteristic, and a third image of a scene having a third characteristic.
4. The apparatus according to claim 3, wherein the first characteristic is a daytime characteristic, the second characteristic is a nighttime characteristic, and the third characteristic is a weather condition.
5. The apparatus according to claim 1, wherein the machine learning - based feature detector is configured to determine the one or more feature points based on a computer vision algorithm.
6. The apparatus according to claim 1, wherein the machine learning system is a neural network.
7. The apparatus according to claim 6, wherein the neural network is a transformer neural network.
8. The apparatus according to claim 7, wherein the transformer neural network is configured to perform transformer cross - attention to determine a unique signature across the obtained input data, and the corresponding feature descriptor for each respective feature point is based on the unique signature.
9. A method for processing image data, the method comprising: obtaining input data; processing the input data using a machine learning - based feature detector to determine one or more feature points in the input data; and using a machine learning system to determine a corresponding feature descriptor for each respective feature point among the one or more feature points.
10. The method according to claim 9, wherein the input data comprises at least one of one or more images, radar data, or light detection and ranging (LIDAR) data.
11. The method according to claim 9, wherein the input data comprises a first image of a scene having a first characteristic, a second image of a scene having a second characteristic, and a third image of a scene having a third characteristic.
12. The method according to claim 11, wherein the first characteristic is a daytime characteristic, the second characteristic is a nighttime characteristic, and the third characteristic is a weather condition.
13. The method according to claim 9, the method further comprising using the machine learning - based feature detector to determine the one or more feature points based on a computer vision algorithm.
14. The method according to claim 9, wherein the machine learning system is a neural network.
15. The method according to claim 14, wherein the neural network is a transformer neural network.
16. The method according to claim 15, wherein the transformer neural network is configured to perform transformer cross-attention to determine a unique signature across the obtained input data, and the corresponding feature descriptor of each corresponding feature point is based on the unique signature.
17. A non-transitory computer-readable medium storing instructions thereon, which when executed by one or more processors cause the one or more processors to: Obtain input data; Process the input data using a machine learning-based feature detector to determine one or more feature points in the input data; and Use a machine learning system to determine a corresponding feature descriptor for each corresponding feature point among the one or more feature points.
18. The non-transitory computer-readable medium according to claim 17, wherein the input data includes at least one of one or more images, radar data, or light detection and ranging (LIDAR) data.
19. The non-transitory computer-readable medium according to claim 17, wherein the input data includes a first image of a scene having a first characteristic, a second image of a scene having a second characteristic, and a third image of a scene having a third characteristic.
20. The non-transitory computer-readable medium according to claim 19, wherein the first characteristic is a daytime characteristic, the second characteristic is a nighttime characteristic, and the third characteristic is a weather condition.
21. The non-transitory computer-readable medium according to claim 17, wherein the machine learning-based feature detector is configured to determine the one or more feature points based on a computer vision algorithm.
22. The non-transitory computer-readable medium according to claim 17, wherein the machine learning system is a neural network.
23. The non-transitory computer-readable medium according to claim 22, wherein the neural network is a transformer neural network.
24. The non-transitory computer-readable medium according to claim 23, wherein the transformer neural network is configured to perform transformer cross-attention to determine a unique signature across the obtained input data, and the corresponding feature descriptor of each corresponding feature point is based on the unique signature.