Scaling for depth estimation

By using the depth values ​​provided by the camera tracker to scale the output of the self-supervised monocular depth network, the problem of fuzzy depth prediction value scale is solved, and the accurate unit of measurement conversion of depth information is achieved.

CN120226045APending Publication Date: 2025-06-27QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080359.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-04
Filing Date
2023-10-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The depth prediction value output from the self-supervised monocular camera system is blurred and cannot be directly used for applications such as autonomous vehicles and extended real-world systems that require real-world units.

Method used

Use depth values ​​from the camera tracker, especially sparse depth values, to scale the depth prediction data generated by the supervised monocular depth network to obtain scale-corrected depth prediction values.

Benefits of technology

Through post-scaling technology, the depth prediction value can be converted from relative scales to specific units of measurement, solving the problem of fuzzy depth value scale and improving the application reliability of depth information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226045A_ABST
    Figure CN120226045A_ABST
Patent Text Reader

Abstract

Systems and techniques for processing sensor data are provided. For example, a process may include determining a predicted depth map of an image using a trained machine learning system, the predicted depth map including a respective predicted depth value for each pixel of the image. The process may also include obtaining depth values of the image from a tracker, the depth values including less than all pixels of the image, the tracker configured to determine the depth values based on one or more feature points between frames. The process may also include scaling the predicted depth map of the image using the depth values. The output of the process may be a scale corrected depth predictor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to processing sensor data (e.g., images, radar data, light detection and ranging (LIDAR) data, etc.). For example, aspects of the present disclosure relate to performing scaling for depth estimation, such as using sparse depth values from a camera tracking engine to scale self-supervised monocular depth data to obtain scale-corrected (e.g., metric-corrected) depth prediction values. Background Art

[0002] Many devices and systems allow capturing characteristics of a scene based on sensor data, such as an image (or frame) of the scene, video data of the scene (including multiple frames), radar data, etc. For example, a camera or a device including a camera can capture a sequence of frames of a scene (e.g., a video of the scene). In some cases, the sequence of frames can be processed to perform one or more functions, can be output for display, can be output for processing and / or consumption by other devices, and other uses.

[0003] For applications such as mixed reality (XR), autonomous driving, camera image / video processing, and robotics, depth perception is valuable. However, self-supervised monocular (single) camera systems output depth prediction values with scale ambiguity. Summary of the Invention

[0004] In some examples, techniques for scaling depth prediction data with ambiguous scale are described, where the depth prediction data is generated from a trained depth network such as a self-supervised monocular depth network. According to at least one illustrative example, a method for processing image data is provided. The method includes: using a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including corresponding predicted depth values for each pixel of the image. The method may further include: obtaining depth values of the image from a tracker configured to determine the depth values based on one or more feature points between frames, the depth values including depth values of less than all pixels of the image; and using the depth values to scale the predicted depth map of the image.

[0005] In another example, a device for processing image data is provided, which includes at least one memory and at least one processor (e.g., implemented in a circuit) coupled to the at least one memory. The at least one processor is configured to: use a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including corresponding predicted depth values for each pixel of the image. The at least one processor may be configured to: obtain depth values of the image from a tracker, the tracker being configured to determine these depth values based on one or more feature points between frames, these depth values including depth values for less than all pixels of the image; and use these depth values to scale the predicted depth map of the image.

[0006] In another example, a non-transitory computer-readable medium is provided, on which instructions are stored that, when executed by one or more processors (e.g., implemented in a circuit), cause the one or more processors to: use a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including corresponding predicted depth values for each pixel of the image. The at least one processor may be caused to: obtain depth values of the image from a tracker, the tracker being configured to determine these depth values based on one or more feature points between frames, these depth values including depth values for less than all pixels of the image; and use these depth values to scale the predicted depth map of the image.

[0007] In another example, a device for processing image data is provided. The device includes: means for using a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including corresponding predicted depth values for each pixel of the image. The device may further include: means for obtaining depth values of the image from a tracker, the tracker being configured to determine these depth values based on one or more feature points between frames, these depth values including depth values for less than all pixels of the image; and means for using these depth values to scale the predicted depth map of the image.

[0008] In some aspects, one or more of the apparatuses described herein are or are part of the following: a camera, a mobile device (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other devices. In some aspects, an apparatus includes one or more cameras for capturing one or more images. In some aspects, the apparatus further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatus may include one or more sensors that may be used to determine the location and / or orientation of the apparatus, the status of the apparatus, and / or for other purposes.

[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0010] The foregoing and other features and aspects will become more apparent when reference is made to the following specification, claims, and appended drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Exemplary aspects of the present application are described in detail below with reference to the following drawings:

[0012] Figure 1A is a block diagram illustrating the architecture of an image capture and processing device according to some examples;

[0013] Figure 1B is a block diagram illustrating the transformation from a two-dimensional image to three-dimensional or depth prediction data according to some examples;

[0014] Figure 2A illustrates an example architecture of a neural network that may be used according to some aspects of the present disclosure;

[0015] Figure 2B is a block diagram illustrating an ML engine according to aspects of the present disclosure;

[0016] Figure 2C is a block diagram illustrating self-supervised training of monocular depth estimation according to aspects of the present disclosure;

[0017] Figure 3A illustrates an example workflow for training a depth model based on a depth-to-segmentation model according to some examples;

[0018] Figure 3BIllustrates an example workflow for performing inference using a trained deep model according to some examples;

[0019] Figure 3C Illustrates an example workflow for training a deep model based on photometric loss according to some examples;

[0020] Figure 3D Illustrates an example workflow for training a deep model based on a ground truth map according to some examples;

[0021] Figure 4 Is an example frame captured by a SLAM system according to some aspects;

[0022] Figure 5 Is a diagram illustrating an example of a hybrid system 500 for detecting features (e.g., key points or feature points) and generating descriptors for the detected features according to some aspects;

[0023] Figure 6 Is a flowchart illustrating an example of a process for processing image data according to some examples of the present disclosure;

[0024] Figure 7 Illustrates a flowchart related to performing post-scaling of scale-ambiguous depth prediction data using depth values (e.g., sparse depth values) according to some examples of the present disclosure;

[0025] Figure 8 Illustrates a flowchart of an example method for performing graphics processing according to some examples of the present disclosure; and

[0026] Figure 9 Is a block diagram illustrating an example of a computing system for implementing certain aspects described herein. Detailed Description

[0027] Certain aspects of the present disclosure are provided below. Some of these aspects may be applied independently, and some of them may be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for the purpose of explanation to provide a thorough understanding of the aspects of the present application. However, it is obvious that the various aspects may be implemented without these specific details. The accompanying drawings and description are not intended to be restrictive.

[0028] The following description only provides example aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the example aspects will provide those skilled in the art with a description that can be used to implement the example aspects. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the essence and scope of the present application as set forth in the appended claims.

[0029] As described above, devices and systems can determine or capture characteristics of a scene based on sensor data associated with the scene. The sensor data can include images (or frames) of the scene, video data of the scene (including multiple frames), radar data, LIDAR data, any combination thereof, and / or other data.

[0030] For example, an image capture device (e.g., a camera) is a device that uses an image sensor to receive light and capture image frames such as still images or video frames. The terms “image,” “image frame,” “video frame,” and “frame” may be used interchangeably herein. An image capture device typically includes at least one lens that receives light from the scene and bends the light toward the image sensor of the image capture device. The light received by the lens passes through an aperture controlled by one or more control mechanisms and is received by the image sensor. The one or more control mechanisms can control exposure, focus, and / or zoom based on information from the image sensor and / or based on information from an image processor (e.g., a host or application process and / or an image signal processor). In some examples, the one or more control mechanisms include a motor or other control mechanism for moving the lens of the image capture device to a target lens position.

[0031] Performing depth estimation can be a fundamental problem in three-dimensional (3D) applications. For example, a system can use depth estimation to transform two-dimensional (2D) image capture into 3D space. Many systems and applications (e.g., extended reality (XR) systems, autonomous driving systems, camera image / video processing systems, robotic systems, etc.) can benefit from such 2D-to-3D transformation. Monocular estimation (estimation using images from a single camera) involves estimating the distance of each pixel from the camera given an image captured by a single camera. The following Figure 1B illustrates monocular depth estimation for several images.

[0032] Self-supervised learning can involve using recorded video as training data for a neural network. For example, self-supervised learning / training can be used to train a neural network (referred to as a depth network) configured to estimate depth from images. In some cases, self-supervised training techniques can be used instead of supervised training techniques (which utilize ground truth data) to train the depth network because it may be difficult to obtain ground truth for the data (e.g., by having humans manually label many images for improved training). For example, there are limitations to collecting large-scale ground truth depth annotations. In some cases, expensive sensors (e.g., LIDAR) and laborious data collection efforts may be required to collect accurate depth ground truth data, which may make it difficult to scale the collection of data.

[0033] One problem with using self-supervision is that deep networks trained with self-supervision can only output relative depth maps. For example, relative depth maps look correct in a relative sense (e.g., a table in the foreground of a scene appears in front of a person in the background of the scene), but the depth values are not in a specific standard measurement scale such as meters or feet. In some cases, deep networks trained using self-supervised learning can use structure-from-motion equations that define how two frames (or images) are related to each other and how to transform pixels from one frame to another. In one example, a specific point on a person's body (e.g., a pixel corresponding to the person's nose tip) can move from a first frame to a second frame. The depth value of the pixel can be transformed from a first position in the first frame to a second position in the second frame as it changes from frame to frame. The system can use such transformed positions to derive a loss function to train the self-supervised network.

[0034] The problem with relative depth maps occurs when the depth values are multiplied by some factor. For example, when applying a factor from one frame to the next, the system can only output relative depth values between the frames. The depth values are relative within the corresponding frame itself or between frames, rather than according to an independent measurement system (e.g., a metric system). For example, data from an algorithm can indicate that a pixel is a certain unit distance from another pixel in the same frame or has moved a specific number of units from one frame to the next. However, the system does not necessarily know what the "unit" equals in terms of the measurement system (e.g., in inches or centimeters). The predicted depth values are only valid within the multiplicative factor.

[0035] Such relative depth maps are thus ambiguous with respect to scale. When a depth map is ambiguous with respect to scale, systems such as autonomous vehicles, XR systems, and other systems cannot directly use the depth values because they cannot project the pixels into real-world units (and thus learn how far an object is from the system such as a vehicle, XR system, etc.).

[0036] This document describes systems, devices, processes (also referred to as methods), and computer-readable media (collectively "systems and techniques") that provide solutions to the problem of scale-ambiguous depth prediction outputs from trained deep networks as described above. For example, as described in more detail herein, an example method includes a system that uses six degrees of freedom (6DoF) data to generate depth values. In some cases, the depth values can include sparse depth values, where not every pixel in the corresponding image or frame is represented by a depth value (e.g., depth values are generated for fewer than all pixels in the image or frame). The system can then perform post hoc scaling of the scale-ambiguous depth prediction data to generate scale-corrected (e.g., metric-corrected) depth prediction values.

[0037] The degree of freedom (DoF) refers to the number of fundamental ways in which a rigid object can move through three-dimensional (3D) space. In some examples, six different DoFs of an object can be tracked. The six DoFs can include three translational DoFs corresponding to translational movement along three perpendicular axes, which can be referred to as the x-axis, y-axis, and z-axis. The six DoFs can also include three rotational DoFs corresponding to rotational movement about the three axes, which can be referred to as pitch, yaw, and roll. Some devices (e.g., XR devices such as virtual reality (VR) or augmented reality (AR) headsets or head-mounted displays (HMDs), mobile devices, vehicles or vehicle systems, robotic devices, etc.) can track some or all of these degrees of freedom. For example, a 3DoF tracker (e.g., of an XR headset) can track three rotational DoFs. A 6DoF tracker (e.g., of an XR headset) can track all six DoFs.

[0038] In some cases, tracking (e.g., 6DoF tracking) can be used to perform localization and mapping functions. For example, visual simultaneous localization and mapping (VSLAM) is a computational geometry technique used in devices with cameras (such as XR devices (e.g., VR HMDs, AR headsets, etc.), robotic devices or systems, mobile phones, vehicles or vehicle systems, etc.). In VSLAM, the device can construct and update a map of an unknown environment based on frames captured by the device's camera. As the device updates the map, the device can track the pose of the device within the environment (e.g., position and / or orientation), such as the pose of the device's image sensor (such as the camera pose), which can be determined using 6DoF tracking. For example, the device can be activated in a specific room of a building and can move throughout the interior of the building, capturing image frames. The device can map the environment and track its position within the environment based on the positions where different objects in the environment appear in different image frames. Other types of sensor data besides image frames can also be used for VSLAM, such as radar and / or LIDAR data.

[0039] In the context of systems that track movement through an environment (e.g., XR systems, robotic systems, vehicles such as autonomous vehicles, VSLAM systems, etc.), degrees of freedom may refer to which of the six degrees of freedom the system is capable of tracking. As noted above, 3DoF tracking systems typically track three rotational DoF (e.g., pitch, yaw, and roll). For example, a 3DoF headset may track a user of the headset turning their head left or right, tilting their head up or down, and / or tilting their head left or right. A 6DoF system may track three translational DoF as well as three rotational DoF. Thus, for example, a 6DoF headset may track a user moving forward, backward, laterally, and / or vertically in addition to tracking the three rotational DoF.

[0040] To perform localization and mapping functions, a device (e.g., an XR device, a mobile device, etc.) may perform feature analysis (e.g., extraction, tracking, etc.) and other complex functions. For example, camera keypoint features (also referred to as feature points) may be used as non-semantic features to improve localization and mapping robustness. Keypoint features may include unique features extracted from one or more images, such as points associated with the corners of a table, the edges of a street sign, etc. The previously mentioned depth values (e.g., sparse depth values) may be combined with these feature points to obtain. The depth values may then be used to generate scale-corrected (e.g., metric-corrected) depth prediction values, as disclosed herein.

[0041] In some cases, machine learning-based systems (e.g., using deep learning neural networks) may be used to detect features (e.g., keypoints) for localization and mapping and generate descriptors for the detected features. However, it may be difficult to obtain ground truth and annotations (or labels) for training machine learning-based feature (e.g., keypoint or feature point) detectors and descriptor generators.

[0042] As previously noted, this document describes systems and techniques for providing post hoc scaling of scale-ambiguous depth prediction values using depth values (e.g., sparse depth values) from a camera tracker. For example, the system can include a non-machine learning-based feature detector or camera tracker and a machine learning-based depth network. In some aspects, the feature detector (e.g., feature point detector) can be based on, e.g., computer vision algorithms, and the machine learning system (e.g., deep learning neural network) can be used to generate descriptors of the detected feature points or key points and thus depth values (e.g., sparse depth values). Descriptors (also referred to as feature descriptors) can be generated at least in part by generating a description of the features detected or depicted in the input sensor data (e.g., local image patches extracted around features in an image). In some cases, the feature descriptor can describe the feature as a feature vector or a collection of feature vectors. The depth values (e.g., sparse depth values) associated with the feature descriptor can be used for post hoc scaling to ultimately generate scale-corrected (e.g., metric-corrected) depth prediction values.

[0043] The systems and techniques described in this document provide various advantages. For example, by performing post hoc scaling of depth predictions, the systems and techniques can allow the system to identify the accurate distances between the system and other objects in the scene. In one illustrative example, the benefit of obtaining the correct scaling of depth values in autonomous driving is that the appropriate known distance of an object from the camera is valuable for avoiding collisions.

[0044] Various aspects of the present application will be described with reference to the accompanying drawings. Figure 1A FIG. 8 is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing an image of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture individual images (or photos), and / or can capture a video including multiple images (or video frames) in a particular sequence. The lens 115 of the image capture and processing system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the light towards the image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.

[0045] One or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. One or more control mechanisms 120 may include multiple mechanisms and components; for example, one or more control mechanisms 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 may also include additional control mechanisms other than those illustrated, such as control mechanisms for controlling analog gain, flash, HDR, depth of field, and / or other image capture attributes.

[0046] One or more focus control mechanisms 125B of one or more control mechanisms 120 may obtain a focus setting. In some examples, one or more focus control mechanisms 125B store the focus setting in a memory register. Based on the focus setting, one or more focus control mechanisms 125B may adjust the positioning of the lens 115 relative to the positioning of the image sensor 130. For example, based on the focus setting, one or more focus control mechanisms 125B may move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or a servo system (or other lens mechanism) to adjust the focus. In some cases, additional lenses may be included in the image capture and processing system 100, such as one or more microlenses located over each photodiode of the image sensor 130, each of the one or more microlenses bending the light received from the lens 115 toward the corresponding photodiode before the light reaches the corresponding photodiode. The focus setting may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. One or more control mechanisms 120, the image sensor 130, and / or the image processor 150 may be used to determine the focus setting. The focus setting may be referred to as an image capture setting and / or an image processing setting.

[0047] One or more exposure control mechanisms 125A of one or more control mechanisms 120 may obtain an exposure setting. In some cases, one or more exposure control mechanisms 125A store the exposure setting in a memory register. Based on the exposure setting, one or more exposure control mechanisms 125A may control the size of the aperture (e.g., aperture size or aperture number), the duration the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting may be referred to as an image capture setting and / or an image processing setting.

[0048] One or more zoom control mechanisms 125C of one or more control mechanisms 120 may obtain a zoom setting. In some examples, one or more zoom control mechanisms 125C store the zoom setting in a memory register. Based on the zoom setting, one or more zoom control mechanisms 125C may control the focal length of an assembly of lens elements (lens assembly) including lens 115 and one or more additional lenses. For example, one or more zoom control mechanisms 125C may control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more of the lenses in the lens relative to each other. The zoom setting may be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, the focusing lens may be lens 115), which first receives light from scene 110, and the light then passes through an afocal zoom system between the focusing lens (e.g., lens 115) and image sensor 130 before reaching the image sensor 130. In some cases, the afocal zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference of each other), with a negative (e.g., diverging, concave) lens therebetween. In some cases, one or more zoom control mechanisms 125C move one or more of the lenses in the afocal zoom system, such as the negative lens and one or two of the positive lenses.

[0049] Image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a particular pixel in the image produced by image sensor 130. In some cases, different photodiodes may be covered by different color filters and may thus measure light that matches the color of the color filter covering the photodiode. For example, a Bayer color filter includes a red color filter, a blue color filter, and a green color filter, where each pixel of the image is generated based on red light data from at least one photodiode covered by the red color filter, blue light data from at least one photodiode covered by the blue color filter, and green light data from at least one photodiode covered by the green color filter. Other types of color filters may use yellow, magenta, and / or cyan (also referred to as "emerald") color filters to replace or supplement the red, blue, and / or green color filters. Some image sensors (e.g., image sensor 130) may have no color filter at all and may instead use different photodiodes (vertically stacked in some cases) throughout the pixel array. The different photodiodes throughout the pixel array may have different spectral sensitivity curves and thus respond to light of different wavelengths. Monochromatic image sensors may also lack color filters and thus lack color depth.

[0050] In some cases, the image sensor 130 may alternatively or additionally include an opaque mask and / or a reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which can be used for phase detection autofocus (PDAF). The image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output by the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output by the photodiodes (and / or the analog signal amplified by the analog gain amplifier) into a digital signal. In some cases, certain components or functions discussed with respect to one or more of the control mechanisms 120 may alternatively or additionally be included in the image sensor 130. The image sensor 130 can be a charge-coupled device (CCD) sensor, an electron multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), complementary metal-oxide-semiconductor (CMOS), N-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0051] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more processors of any other type of processor 910 discussed with respect to the computing system 900. The host processor 152 can be a digital signal processor (DSP) and / or other types of processors. In some specific implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth TM, global positioning systems (GPS), etc.), any combination thereof, and / or other components. The I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General-Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 Physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an illustrative example, the host processor 152 may communicate with the image sensor 130 using the I2C port, and the ISP 154 may communicate with the image sensor 130 using the MIPI port.

[0052] The image processor 150 may perform multiple tasks, such as demosaicking, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving inputs, managing outputs, managing memory, or some combination thereof. The image processor 150 may store the image frames and / or the processed images in a Random Access Memory (RAM) 140 / 1620, a Read-Only Memory (ROM) 145 / 1625, a cache, a memory cell, another storage device, or some combination thereof.

[0053] A variety of input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device 935, any other input device 945, or some combination thereof. In some cases, captions may be input into the image processing device 105B via the physical keyboard or keypad of the I / O device 160, or via the virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O port 156 may include one or more ports, jacks, or other connectors that implement a wired connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The I / O port 156 may include one or more wireless transceivers that implement a wireless connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The peripheral devices may include any type of I / O device 160 discussed previously, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they themselves may be considered I / O devices 160.

[0054] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some specific embodiments, the image capture device 105A and the image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly via one or more wireless transceivers. In some specific embodiments, the image capture device 105A and the image processing device 105B may be disconnected from each other.

[0055] As Figure 1A shown, the vertical dashed line divides Figure 1A the image capture and processing system 100 into two parts, representing the image capture device 105A and the image processing device 105B, respectively. The image capture device 105A includes a lens 115, one or more control mechanisms 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, certain components illustrated in the image capture device 105A (such as the ISP 154 and / or the host processor 152) may be included in the image capture device 105A.

[0056] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 wi-fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some embodiments, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile phone, a desktop computer, or other computing device.

[0057] Although the image capture and processing system 100 is shown as including certain components, those of ordinary skill in the art should understand that the image capture and processing system 100 may include more components than Figure 1A those shown therein. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, the components of the image capture and processing system 100 may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a GPU, a DSP, a CPU, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the image capture and processing system 100.

[0058] Figure 1B A first set of images 170 is illustrated, where, for example, the raw image 172 is processed by the image capture and processing device 100 of Figure 1A and a depth prediction or depth map 174 is generated. Another set of images 176 shows the raw image 178 and the depth map 180, including mapping of the depth of the object in the image 182.

[0059] As part of the image processing, a neural network or machine learning model may be used to generate data about the image. Figure 2AIllustrates an example architecture of a neural network 200 that can be used according to some aspects of the present disclosure. The example architecture of the neural network 200 can be defined by an example neural network description 202 in a neural controller 201. The neural network 200 is an example of a machine learning model that can be deployed and implemented on any device such as an autonomous vehicle or an XR system. The neural network 200 can be a feedforward neural network or any other known or to-be-developed neural network or machine learning model.

[0060] The neural network description 202 can include a complete specification of the neural network 200, including Figure 2A the neural architecture shown in. For example, the neural network description 202 can include: a description or specification of the architecture of the neural network 200 (e.g., layers, layer interconnects, number of nodes in each layer, etc.); input and output descriptions indicating how the inputs and outputs are formed or processed; indications of activation functions, operations, or filters in the neural network, etc.; neural network parameters such as weights, biases, etc.; and so on.

[0061] The neural network 200 can reflect the neural architecture defined in the neural network description 202. The neural network 200 can include any suitable neural or deep learning type of network. In some cases, the neural network 200 can include a feedforward neural network. In other cases, the neural network 200 can include a recurrent neural network, which can have loops that allow information to be carried across nodes when reading inputs. The neural network 200 can include any other suitable neural network or machine learning model. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input layer and the output layer. The hidden layers of the CNN include a series of hidden layers as described below, such as convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. In other examples, the neural network 200 can represent any other neural network or deep learning network, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), etc.

[0062] In Figure 2AIn a non-limiting example, the neural network 200 includes an input layer 203 that can receive one or more sets of input data. The input data can be any type of data (e.g., image data, video data, network parameter data, user data, etc.). The neural network 200 can include hidden layers 204A through 204N (collectively referred to hereinafter as "204"). The hidden layer 204 can include a number n of hidden layers, where n is an integer greater than or equal to one. The number n of hidden layers can include many layers required for the desired processing result and / or presentation intent. In an illustrative example, any one of the hidden layers 204 can include data representing one or more of the data provided at the input layer 203. The neural network 200 also includes an output layer 206 that provides an output generated by the processing performed by the hidden layer 204. The output layer 206 can provide output data based on the input data.

[0063] In Figure 2A an example, the neural network 200 is a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains the information when processing the information. Information can be exchanged between the nodes through node-to-node interconnections between the various layers. The nodes of the input layer 203 can activate a set of nodes in the first hidden layer 204A. For example, as shown, each input node of the input layer 203 is connected to each node of the first hidden layer 204A. The nodes of the hidden layer 204A can transform the information by applying an activation function to the information of each input node. Then, the information derived from the transformation can be passed to the nodes of the next hidden layer (e.g., 204B) and can activate those nodes, which can perform their own specified functions. Example functions include convolution, upsampling, data transformation, pooling, and / or any other suitable function. The output of the hidden layer (e.g., 204B) can then activate the nodes of the next hidden layer (e.g., 204N), and so on. The output of the last hidden layer can activate one or more nodes of the output layer 206, at which point the output is provided. In some cases, although the nodes in the neural network 200 (e.g., nodes 208A, 208B, 208C) are shown as having multiple output lines, a node can have a single output and all the lines shown as output from the node can represent the same output value.

[0064] In some cases, each node or the interconnections between the nodes can have weights, which are a set of parameters derived from training the neural network 200. For example, the interconnections between the nodes can represent a piece of information learned about the interconnected nodes. The interconnections can have numerical weights that can be tuned (e.g., based on a training data set) to allow the neural network 200 to adapt to the input and be able to learn as it processes more data.

[0065] The neural network 200 can be pre-trained to process the features of data from the input layer 203 using different hidden layers 204, so as to provide an output through the output layer 206. For example, in some cases, the neural network 200 can use a training process called backpropagation to adjust the weights of the nodes. Backpropagation can include forward pass, loss function, backward pass, and weight update. The forward pass, loss function, backward pass, and parameter update can be performed for one training iteration. This process can be repeated for each training data set for a certain number of iterations until the weights of these layers are accurately tuned (e.g., a configurable threshold determined based on experimental and / or empirical studies is met).

[0066] An increasing number of ML (e.g., AI) algorithms (e.g., models) are being incorporated into various techniques including those for image processing. Figure 2B FIG. 220 is a block diagram illustrating an ML engine 224 including an input 222 and an output 226 according to aspects of the present disclosure. As an example, one or more devices performing image processing can include the ML engine 224. In some cases, the ML engine 224 can be similar to the neural network 200. In this example, the ML engine 224 includes three parts: the input 222 of the ML engine 224, the ML engine 224, and the output 226 from the ML engine 224. The input 222 of the ML engine 224 can be data that the ML engine 224 can use to make predictions or otherwise operate on it.

[0067] Figure 2C Illustrates the self-supervised training of a monocular depth estimation system 230. The first input image 232 (I s ) is a neighbor of the second input image 234 (I t ). The pixel p t is included in I t , and the pixel p s is included in I s . It is assumed that these different pixels are different views of the same point of an object in the images. Then, p t and p s are geometrically related as follows:

[0068] d(p s )h(p s ) = K(R t→s K -1d (p t )h(p t ) + t t→s ),

[0069] where h(p) = [h, w, l] represents the homogeneous coordinates of the pixel p, and h and w are the vertical position and horizontal position of the pixel p in the image, and d(p) is the depth at p, is the camera intrinsic matrix, and is the 6Dof F relative camera motion pose from t to s, where and are the rotation matrix and the translation vector.

[0070] System 230 may include a pose network 236 and a depth network 240, which are trainable units or components that can provide pose output and depth output to a view synthesis engine 238, and the view synthesis engine can be a non-trainable unit or component that can output I’ t The photometric loss component 242 can use the structural similarity index measure to compare I t and I’ t . System 230 may also include a smooth loss component 244, which can be used to prevent drastic changes in the predicted depth map. This can utilize different kinds of loss functions to reduce more drastic changes in the depth map. As pointed out above, self-supervised training of monocular depth estimation does not implement or provide an absolutely known unit for determining depth (such as a metric unit), but the output data is relative to the objects in the frame or relative to the movement between frames. Therefore, the solution provided herein addresses the problem of such scale-ambiguous depth values output from system 230.

[0071] Figure 3A depicts an example workflow 300A for training a depth model 310 and a depth-to-segmentation model 325. The depth model 310 can be applied to, for example, the trained depth network 704 introduced below in Figure 7 .

[0072] In workflow 300A, the depth model 310 is used to process the input image 305A. The depth model 310 is generally a machine learning model configured to generate a depth map based on the input image. As discussed above, the depth map can indicate the depth of the corresponding object covered by each pixel in the input image 305A (e.g., the distance from the camera). In some aspects, the depth model is a neural network.

[0073] As shown, during the training process, the generated depth map can be used to calculate the depth loss 315A. In some aspects, a synthetic version of the input image 305A is generated using the generated depth map, and the synthetic version can be compared with the original input image 305A to generate the depth loss 315A, as discussed in more detail below with reference to Figure 3A . In some cases, the generated depth map is compared with a ground truth depth map to generate the depth loss 315A, as discussed in more detail below with reference to Figure 3CFor more detailed discussion. Generally speaking, the depth loss 315A can be used (e.g., using backpropagation) to iteratively refine the depth model 310.

[0074] In the illustrated workflow 300A, the cross-task distillation module 320 is used to transfer semantic knowledge from the pre-trained segmentation model 330 to the depth model 310. That is, if the depth model 310 is represented as f D and the pre-trained semantic segmentation model is represented as f S then the cross-task distillation module 320 enables the transfer of knowledge from the teacher model f S to the student model f D . However, unlike conventional knowledge distillation where the teacher and student networks are used for the same visual task, f D and f S are used for two different tasks and their outputs cannot be directly compared. In other words, given an input, the system cannot directly measure the difference between the outputs of f D (depth map) and f S (segmentation map) in order to generate the loss required to train f D .

[0075] Therefore, in the illustrated aspect, a depth-to-segmentation model 325 (which can be represented as h D2S ) is used. In one aspect, the depth-to-segmentation model 325 is a neural network. In some aspects, the depth-to-segmentation model 325 is a small neural network (e.g., having only two conventional convolutional layers and one pointwise convolutional layer, or having a pointwise convolutional layer before zero to four convolutional layers), enabling it to be trained effectively and at minimal computational cost.

[0076] For example, in one aspect, the depth-to-segmentation model 325 consists of two 3×3 convolutional layers (or any number of convolutional layers), each followed by a BatchNorm layer and a ReLu layer, and a final pointwise convolutional layer that outputs the segmentation map. In some aspects, using a deeper network for the depth-to-segmentation model 325 can lead to improved accuracy of the depth model 310, but these improvements are not as significant as those achieved using a smaller depth-to-segmentation model 325. Additionally, a larger or deeper depth-to-segmentation model 325 can play an out-of-scope role in learning, thus weakening the knowledge flow to the depth model 310.

[0077] In the workflow 300A, the depth-to-segmentation model 325 receives the depth map generated by the depth model 310 and translates the depth map into a semantic segmentation map. In other words, the depth-to-segmentation model 325 generates a segmentation map based on the input depth map.

[0078] Additionally, as shown, a pre-trained segmentation model 330 is used to generate a segmentation map based on the original input image 105. Although the pre-trained segmentation model 330 is depicted, in some aspects, the cross-task distillation module 320 may use a ground truth segmentation map of the input image 305A (e.g., provided by a user) instead of using the pre-trained segmentation model 330 to process the input image 305A to generate a ground truth segmentation map. As used herein, the segmentation map used to generate the segmentation loss 335 (which may be provided by a user or generated by the pre-trained segmentation model 330) may be referred to as a "ground truth" segmentation map to reflect that the segmentation map is used as the ground truth when calculating the loss, even if the segmentation map may actually be a fake ground truth map generated by a trained model. As used herein, the term "ground truth segmentation map" may include a true or actual segmentation map (e.g., provided by a user), as well as a pseudo or generated ground truth segmentation map (e.g., generated by the pre-trained segmentation model 330).

[0079] Given the segmentation map generated by the depth-to-segmentation model 325 (based on the predicted depth map generated by the depth model 310) and the segmentation map generated by the pre-trained segmentation model 330 (based on the input image 105), the system is able to construct the segmentation loss 335. This segmentation loss 335 can then be used to distill semantic knowledge from f S to f D . In one aspect, the following Equation 1 is used to define this new segmentation loss 335, where is the segmentation loss 335. is the semantic segmentation map generated by the depth-to-segmentation model 325 based on the predicted depth map generated by the depth model 310 given the input image 105. That is, where I t is the input image 105. Additionally, S t is the semantic segmentation output generated by the pre-trained semantic segmentation model 330, represents the cross-entropy loss, and H and W are the height and width of the input image 105.

[0080]

[0081] In the illustrated workflow 300A, the segmentation loss 335 can be used to allow the depth-to-segmentation model 325 to be jointly trained with the depth model 310. This enables the pre-trained segmentation model 330 to provide semantic supervision to the depth model 310 by backpropagating the segmentation loss 335 through the depth-to-segmentation model 325. That is, the segmentation loss 335 can be backpropagated through the depth-to-segmentation model 325 (e.g., generating gradients for each layer), and the resulting tensor or gradients output from the depth-to-segmentation model 325 can be backpropagated through the depth model 310.

[0082] Although, for clarity of concept, the illustrated workflow 300A depicts a single input image 305A (suggesting stochastic gradient descent), in all respects, the training workflow 300A can be used to provide training in batches of input images 105.

[0083] In some aspects, as discussed above, the semantic classes used by the pre-trained segmentation model 330 can be combined or grouped to achieve improved training of the deep model 310. Semantic segmentation typically can include more fine-grained visual recognition information that is not present or realistic in the depth map. For example, road objects and sidewalk objects are typically considered two different semantic classes, but the depth map typically does not contain such classification information because the road and the sidewalk are both on the ground plane and have similar depth variations. Thus, there is no need to distinguish them on the depth map. On the other hand, the depth map does contain information for distinguishing certain classes. For example, given their different patterns of depth values, road participants (e.g., pedestrians, vehicles) can be easily separated from the background (e.g., road, building).

[0084] Thus, in some aspects, the semantic classes can be grouped or combined such that the semantic information is retained while unnecessary complexity is removed from the distillation. In one such aspect, the classes are combined into a first group for objects in the foreground (e.g., vehicles, pedestrians, signs, etc.) and a second group for objects in the background (e.g., buildings, the ground itself, etc.). In at least one aspect, the objects in the foreground are depicted as two subgroups based on (e.g.) their shape. For example, the system can use a first group (or subgroup) for thin structures (e.g., traffic lights and signs, poles, etc.) and a second group (or subgroup) for wider shapes (such as people, vehicles, etc.).

[0085] Similarly, the objects in the background can be divided into a third group (or subgroup) and a fourth group (or subgroup), where the third group contains background objects (e.g., buildings, vegetation, etc.) and the fourth group includes the ground (e.g., road, sidewalk, etc.). Such class merging applied to the segmentation map generated by the pre-trained segmentation model 330 can improve the resulting accuracy of the deep model 310. In some aspects, such merging is performed based on a user-specified configuration (e.g., indicating which classes should be merged into a given group). In one aspect, merging the classes includes relabeling the segmentation map based on the grouping of the classes. For example, if light poles and signs are merged into the same class, they will be assigned the same value in the (new) segmentation map. The depth-to-segmentation model 325 is typically configured to output a segmentation map based on the merged classes.

[0086] As discussed above, since the depth-to-segmentation model 325 can be small, the distillation method adds only a small amount of computation to training. Additionally, for each training input image 105, the segmentation map from the teacher network (the pre-trained segmentation model 330) needs to be computed only once and can then be reused as needed. This represents an improvement over existing systems that co-train a segmentation model with a depth model, which may need to process images with the segmentation model multiple times during training.

[0087] Figure 3B An example workflow 300B for performing inference using a trained depth model 310 is depicted. The workflow 300B can be used, for example Figure 7 for a trained deep network 704.

[0088] In the illustrated aspect, the depth model 310 has been trained using a cross-task distillation module (such as the cross-task distillation module 320 discussed above with reference to Figure 3A That is, the depth model 310 can be trained at least in part based on a segmentation loss generated with the help of a depth-to-segmentation model (such as the depth-to-segmentation model 325 discussed above with reference to Figure 3A In this way, the depth model 310 learns segmentation knowledge that can enable significantly improved depth estimation.

[0089] Once training is complete (e.g., determined based on a termination criterion such as sufficient accuracy or otherwise determined that the model is sufficiently trained), the depth model 310 can operate in an independent manner without any additional computation for semantic information during inference. That is, the input image 305B can be processed by the depth model 310 to generate an accurate depth map 345 without passing any data through Figure 3A the depth-to-segmentation model 325 or the pre-trained segmentation model 330. Thus, in some aspects, the cross-task distillation module 320 can be discarded after training. In some aspects, the depth-to-segmentation model 325 can be stored for use with future refinement or training.

[0090] However, in some aspects, only the depth model 310 is used during inference. Since the depth model 310 is trained using cross-task distillation in a more semantically aware manner, it exhibits superior accuracy compared to existing systems. Additionally, since the workflow 300B does not use a separate segmentation network or depth-to-segmentation model during inference, the computational resources required (e.g., power consumption, latency, memory footprint, number of operations, etc.) are significantly reduced compared to existing systems.

[0091] Figure 3CDepicts an example workflow 300C for training a deep model using photometric loss and segmentation loss. Workflow 300C generally provides more details for the calculation of the depth loss 315A discussed above (shown as depth loss 315C in Figure 3A ). Specifically, workflow 300C uses a self-supervised method to implement the calculation of the depth loss 315C and the training of the deep model 310 without the need for a ground truth depth map. The example workflow 300C can be used for Figure 3C the trained deep network 704 of Figure 7 .

[0092] In the illustrated workflow 300C, the input images 305A and 305B are adjacent (e.g., neighboring) or close (e.g., within a defined number of frames or timestamps) frames from a video. Both are provided to a pose model 346 that is configured to determine the relative camera motion between the input images 305A and 305B. Generally, the pose model 346 is a machine learning model (e.g., a neural network) that infers the camera pose (e.g., position and orientation in six degrees of freedom) of the input images.

[0093] For example, consider two adjacent or close video frames, I t and I s (e.g., input images 305A and 305B). Assume that pixel p t ∈ I t and pixel p s ∈ I s are two different views of the same point of an object. In this case, p t and p s are geometrically related as indicated by Equation 2 below, where h(p)=[h, w, 1] represents the homogeneous coordinates of pixel p, and h and w are the vertical and horizontal positions of the pixel in the image, d(p) is the depth at p, K is the camera intrinsic matrix, and T t→s is the six-degree-of-freedom relative camera motion / pose from t to s.

[0094]

[0095] The determined pose generated by the pose model 205 is provided to the view synthesizer 355. Additionally, the input image 305B can be provided to the deep model 310 to generate a predicted depth map of the image 305B. As depicted in workflow 300C, the generated depth map is then provided to the view synthesizer 355.

[0096] Given the generated depth map I t output by the deep model 310 and which can be represented as D t(Image 305B) and the relative camera pose from I t (Image 305B) to I s (Image 305A), assuming that the points captured in I t also exist in I s In, the view synthesizer 355 can synthesize I from I based on Equation 2 s synthesized I t . The synthesized version 305B (I t ) of the input image can be represented as

[0097] As shown, by minimizing the difference between the synthesized image and the actual image 305B (indicated by the depth loss 315A), the system can train the pose model 346 and the depth model 310. In some aspects, this depth loss 315A is called the photometric loss (represented as ), and can be defined using Equation 3 below, where ||·||1 represents norm and SSIM is the structural similarity index measure. Note is calculated on a per-pixel basis.

[0098]

[0099] In some aspects, the system can further include a smooth regularization or loss to prevent drastic changes in the predicted depth map. Additionally, in some aspects, not all 3D points in I t can be found in I s (e.g., due to occlusion and objects (fully or partially) moving out of the frame). Some objects can also be moving (e.g., cars), and these objects are not considered in the geometric model of Equation 2. In one such aspect, to correctly measure the photometric loss and train the network, the system can mask out the pixel points that violate the geometric model.

[0100] In the illustrated workflow 300C, the depth model 310 is also refined based on the segmentation loss 335 (propagated through the depth-to-segmentation model 325), as discussed above. In one aspect, therefore, the total loss of the depth model 310 ) can be defined using Equation 4 below, where the self-supervised depth loss is calculated at N s scales, is the photometric loss at the k th th scale, λ SM,k and are the weights and losses for smooth regularization at the k th th scale, and λ D2S is the weight of the cross-task distillation loss

[0101]

[0102] Figure 3D Depicts an example workflow 300D for training a depth model using a ground truth depth map and a segmentation loss. Workflow 300D generally provides more details for the calculation of depth loss 315A, as referenced above Figure 3A in the discussion. Specifically, workflow 300D uses ground truth depth map 345 to calculate depth loss 315D. Workflow 300D can be used to train Figure 7 the trained depth network 704.

[0103] In the illustrated workflow 300D, a depth-to-segmentation model 325 can be used to calculate a segmentation loss 335, which is used to refine depth model 310, as referenced above Figure 3A in the discussion.

[0104] Further as shown, for each input image 305D, the corresponding ground truth depth map 345 is used to calculate depth loss 315D. For example, the system can use cross entropy to calculate depth loss 315D based on the ground truth depth map 345 and the predicted depth map generated by depth model 310. Then, this depth loss 315D can be used together with segmentation loss 335 to refine depth model 310.

[0105] Figure 4 Is a diagram illustrating the architecture of an example system 400 in accordance with some aspects of the present disclosure. System 400 can be used in conjunction with Figure 7 the camera tracker 714 shown therein.

[0106] In some cases, the system can be a tracking system that does not have a mapping function, such that the system generates features (e.g., sparse features) and may not include various mapping / localization engines / functions. System 400 can be an XR system (e.g., running (or executing) XR applications and / or implementing XR operations), a system of a vehicle, a robotic system, or other types of systems. System 400 can perform tracking and localization, mapping of an environment in the physical world (e.g., a scene), and / or localization and rendering of virtual content on display 409 (e.g., as part of an XR experience, localization and rendering of content on a screen, a visible plane / area, and / or other displays). For example, System 400 can generate a map of an environment in the physical world (e.g., a three-dimensional (3D) map), track the pose (e.g., position and orientation) of System 400 relative to the environment (e.g., relative to the 3D map of the environment), and / or determine a localization and / or an anchor point at a specific location on the map of the environment. In one example, System 400 can localize and / or anchor virtual content at a specific location on the map of the environment and can present the virtual content on display 409 such that the virtual content appears to be located at a position in the environment corresponding to the specific location on the map of the scene where the virtual content is localized and / or anchored. Display 409 can include a monitor, glass, a screen, a lens, a projector, and / or other display mechanisms. For example, in the context of an XR system, display 409 can allow a user to see the real-world environment and also allow XR content to be overlaid on, overlapped with, blended with, or otherwise displayed on the real-world environment.

[0107] In this illustrative example, System 400 includes one or more image sensors 402, an accelerometer 404, a gyroscope 406, a storage device 407, a computing component 410, a pose engine 420, an image processing engine 424, and a rendering engine 426. It should be noted that Figure 4 the components 402 to 126 shown are provided as non-limiting examples for illustrative and explanatory purposes, and other examples can include more, fewer, or different components compared to Figure 4 the components shown. For example, in some cases, System 400 can include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 4One or more other software and / or hardware components not shown. Although the various components of system 400 (such as image sensor 402) may be referred to in the singular form herein, it should be understood that system 400 may include multiple of any of the components discussed herein (e.g., multiple image sensors 402).

[0108] System 400 includes an input device 408 or communicates (wired or wirelessly) with the input device. Input device 408 may include any suitable input device, such as a touch screen, a pen or other pointer device, a keyboard, a mouse, buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, a video game controller, a steering wheel, a joystick, a set of buttons, a trackball, a remote control, any other input device 1645 discussed herein, or any combination thereof. In some cases, image sensor 402 may capture images that can be processed to interpret gesture commands.

[0109] In some specific implementations, one or more of image sensors 402, accelerometer 404, gyroscope 406, storage device 407, computing component 410, pose engine 420, image processing engine 424, and rendering engine 426 may be part of the same computing device. For example, in some cases, one or more of image sensors 402, accelerometer 404, gyroscope 406, storage device 407, computing component 410, pose engine 420, image processing engine 424, and rendering engine 426 may be integrated into a device or system (such as an HMD, XR glasses (e.g., AR glasses), a vehicle or a vehicle system, a smart phone, a laptop computer, a tablet computer, a gaming system, and / or any other computing device). However, in some specific implementations, one or more of image sensors 402, accelerometer 404, gyroscope 406, storage device 407, computing component 410, pose 420, image processing engine 424, and rendering engine 426 may be part of two or more separate computing devices. For example, in some cases, some of components 402 to 426 may be part of or implemented by one computing device, and the remaining components may be part of or implemented by one or more other computing devices.

[0110] The storage device 407 can be any storage device for storing data. Additionally, the storage device 407 can store data from any of the components in the system 400. For example, the storage device 407 can store data from the image sensor 402 (e.g., image or video data), data from the accelerometer 404 (e.g., measurements), data from the gyroscope 406 (e.g., measurements), data from the computing component 410 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, face recognition data, occlusion data, etc.), data from the pose engine 420, data from the image processing engine 424, and / or data from the rendering engine 426 (e.g., output frames). In some examples, the storage device 407 can include a buffer for storing frames to be processed by the computing component 410.

[0111] One or more computing components 410 can include a central processing unit (CPU) 412, a graphics processing unit (GPU) 414, a digital signal processor (DSP) 416, an image signal processor (ISP) 418, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). The computing component 410 can perform various operations, such as image enhancement, computer vision, graphics rendering, tracking, positioning, pose estimation, mapping, content anchoring, content rendering, image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, the computing component 410 can implement (e.g., control, operate, etc.) the pose engine 420, the image processing engine 424, and the rendering engine 426. In other examples, the computing component 410 can also implement one or more other processing engines.

[0112] The image sensor 402 can include any image and / or video sensor or capture device. In some examples, the image sensor 402 can be part of a multi-camera assembly (such as a dual-camera assembly). The image sensor 402 can capture image and / or video content (e.g., raw image and / or video data), which can then be processed by the computing component 410, the pose engine 420, the image processing engine 424, and / or the rendering engine 426, as described herein. In some examples, the image sensor 402 can include an image capture and processing system 100, an image capture device 105A, an image processing device 105B, or a combination thereof.

[0113] In some examples, the image sensor 402 may capture image data and generate an image (also referred to as a frame) based on the image data and / or provide the image data or frame to the pose engine 420, the image processing engine 424, and / or the rendering engine 426 for processing. The image or frame may include a video frame or a static image in a video sequence. The image or frame may include an array of pixels representing a scene. For example, the image may be a red, green, and blue (RGB) image having red, green, and blue color components per pixel; a luminance, chroma red, chroma blue (YCbCr) image having a luminance component and two chroma (color) components (chroma red and chroma blue) per pixel; or any other suitable type of color or monochrome image.

[0114] In some cases, the image sensor 402 (and / or other cameras of the system 400) may also be configured to capture depth information. For example, in some embodiments, the image sensor 402 (and / or other cameras) may include an RGB-depth (RGB-D) camera. In some cases, the system 400 may include one or more depth sensors (not shown) that are separate from the image sensor 402 (and / or other cameras) and may capture depth information. For example, such a depth sensor may obtain depth information independently of the image sensor 402. In some examples, the depth sensor may be physically mounted in the same general location as the image sensor 402 but may operate at a different frequency or frame rate than the image sensor 402. In some examples, the depth sensor may take the form of a light source that projects a structured or textured light pattern (which may include one or more narrowband lights) onto one or more objects in a scene. Then, depth information may be obtained by exploiting the geometric deformation of the projected pattern caused by the surface shape of the object. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).

[0115] System 400 may also include other sensors in one or more of its sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 404), one or more gyroscopes (e.g., gyroscope 406), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other position-related information to computing component 410. For example, accelerometer 404 may detect the acceleration of system 400 and may generate an acceleration measurement based on the detected acceleration. In some cases, accelerometer 404 may provide one or more translation vectors (e.g., up / down, left / right, front / back) that may be used to determine the position or pose of system 400. Gyroscope 406 may detect and measure the orientation and angular velocity of system 400. For example, gyroscope 406 may be used to measure the pitch, roll, and yaw of system 400. In some cases, gyroscope 406 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, image sensor 402 and / or pose engine 420 may use the measurements obtained by accelerometer 404 (e.g., one or more translation vectors) and / or the measurements obtained by gyroscope 406 (e.g., one or more rotation vectors) to calculate the pose of system 400. As previously noted, in other examples, system 400 may also include other sensors such as an inertial measurement unit (IMU), magnetometer, gaze and / or eye tracking sensors, machine vision sensors, intelligent scene sensors, speech recognition sensors, impact sensors, vibration sensors, position sensors, tilt sensors, etc.

[0116] As noted above, in some cases, one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure the specific force, angular velocity, and / or orientation of system 400. In some examples, one or more sensors may output information measured in association with the capture of an image captured by image sensor 402 (and / or other cameras of system 400) and / or depth information obtained using one or more depth sensors of system 400.

[0117] The pose engine 420 can use the outputs of one or more sensors (e.g., accelerometer 404, gyroscope 406, one or more IMUs, and / or other sensors) to determine the pose of the system 400 (also referred to as the head pose) and / or the pose of the image sensor 402 (or other cameras of the system 400). In some cases, the pose of the system 400 and the pose of the image sensor 402 (or other cameras) can be the same. The pose of the image sensor 402 refers to the positioning and orientation of the image sensor 402 relative to a reference frame (e.g., with respect to an object). In some specific implementations, the camera pose can be determined for six degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by the X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as the image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some specific implementations, the camera pose can be determined for three degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).

[0118] In some cases, a device tracker (not shown) can use measurements from one or more sensors and image data from the image sensor 402 to track the pose of the system 400 (e.g., 6DoF pose). For example, the device tracker can fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and motion of the system 400 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the system 400, the device tracker can generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate an update to the 3D map for the scene. The 3D map update can include, for example but not limited to, new or updated features and / or feature or landmark points associated with the scene and / or the 3D map of the scene, and / or a localization update that identifies or updates the position of the system 400 within the scene and the 3D map of the scene. The 3D map can provide a digital representation of the scene in the real / physical world. In some examples, the 3D map can anchor location-based objects and / or content to real-world coordinates and / or objects. The system 400 can use the mapped scene (e.g., the scene in the physical world represented and / or associated with the 3D map) to merge the physical and virtual worlds and / or to merge virtual content or objects with the physical environment.

[0119] In some aspects, the computing component 410 may use a visual tracking solution to determine and / or track the pose (also referred to as the camera pose) of the image sensor 402 and / or the system 400 as a whole based on images captured by the image sensor 402 (and / or other cameras of the system 400). For example, in some examples, the computing component 410 may use computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques to perform the tracking. For example, the computing component 410 may perform SLAM or may communicate (wired or wirelessly) with a SLAM system (not shown in Figure 4 such as Figure 5 the SLAM system 500. SLAM refers to a class of techniques that create a map of the environment (e.g., a map of the environment modeled by the system 400) while simultaneously tracking the camera (e.g., the image sensor 402) and / or the pose of the system 400 relative to the map. This map may be referred to as a SLAM map and may be three-dimensional (3D). SLAM techniques may use color or grayscale image data captured by the image sensor 402 (and / or other cameras of the system 400) and may be used to generate an estimate of the 6DoF pose measurement of the image sensor 402 and / or the system 400. Such SLAM techniques configured to perform 6DoF tracking may be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., the accelerometer 404, the gyroscope 406, one or more IMUs, and / or other sensors) may be used to estimate, correct, and / or otherwise adjust the estimated pose.

[0120] In some cases, 6DoF SLAM (e.g., 6DoF tracking) may associate features (keypoints) observed from certain input images from the image sensor 402 (and / or other cameras or sensors) to the SLAM map. For example, 6DoF SLAM may use feature point associations from the input images (or other sensor data, such as radar sensors, LIDAR sensors, etc.) to determine the pose (position and orientation) of the image sensor 402 and / or the system 400 for the input images. 6DoF mapping may also be performed to update the SLAM map. In some cases, the SLAM map maintained using 6DoF SLAM may contain 3D feature points (e.g., keypoints) triangulated from two or more images. For example, key frames may be selected from the input images or video stream to represent the observed scene. For each key frame, the corresponding 6DoF camera pose associated with the image may be determined. The pose of the image sensor 402 and / or the system 400 may be determined by projecting features (e.g., feature points or keypoints) from the 3D SLAM map into the image or video frame and updating the camera pose based on the verified 2D-3D correspondences.

[0121] In an illustrative example, the computing component 410 may extract feature points (e.g., key points) from certain input images (e.g., each input image, a subset of the input images, etc.) or from each key frame. Feature points (also referred to as registration points) as used herein are unique or identifiable portions of an image, such as a part of a hand, an edge of a table, and other examples. Features extracted from the captured images may represent different feature points along a three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and each feature point may have an associated feature location. Feature points in a key frame match (are the same as or correspond to) or fail to match feature points of a previously captured input image or key frame. Feature detection may be used to detect feature points. Feature detection may include image processing operations for examining one or more pixels of an image to determine whether a feature exists at a particular pixel. Feature detection may be used to process the entire captured image or certain portions of the image. For each image or key frame, once features have been detected, local image patches around the features may be extracted. Any suitable technique may be used to extract features, such as Scale-Invariant Feature Transform (SIFT) (which localizes features and generates their descriptions), Learned Invariant Feature Transform (LIFT), Speeded-Up Robust Features (SURF), Gradient Location-Orientation Histogram (GLOH), Oriented FAST and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoint (FREAK), KAZE, Accelerated-KAZE (AKAZE), Normalized Cross-Correlation (NCC), descriptor matching, another suitable technique, or a combination thereof.

[0122] In some cases, the system 400 may also track the user's hand and / or fingers to allow the user to interact with and / or control virtual content in the virtual environment. For example, the system 400 may track the pose and / or movement of the user's hand and / or fingertips to identify or translate the user's interaction with the virtual environment. User interaction may include, for example but not limited to, moving virtual content items, resizing virtual content items, selecting input interface elements in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interfaces), providing input through the virtual user interface, etc.

[0123] Figure 5 is a block diagram illustrating the architecture of a Simultaneous Localization and Mapping (SLAM) system 500. In some examples, the SLAM system 500 may be, may include Figure 4 the system 400 or may be a part thereof, and may be Figure 7Part of the camera tracker 714. In some examples, the SLAM system 500 can be, can include an XR device, an autonomous vehicle, a vehicle, a computing system of a vehicle, a wireless communication device, a mobile device or a cellular phone (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device (e.g., a network-connected watch), a personal computer, a laptop computer, a server computer, a portable video game console, a portable media player, a camera device, a manned or unmanned ground vehicle, a manned or unmanned aerial vehicle, a manned or unmanned water vehicle, a manned or unmanned underwater vehicle, a manned or unmanned vehicle, a robot, another device, or any combination thereof, or can be a part of them.

[0124] Figure 5 The SLAM system 500 includes or is coupled to each of one or more sensors 505. The one or more sensors 505 can include one or more cameras 510. Each of the one or more cameras 510 can include an image capture device 105A (as Figure 1A shown), an image processing device 105B (as Figure 1A shown), an image capture and processing system 100 (as Figure 1A shown), another type of camera, or a combination thereof. Each of the one or more cameras 510 can respond to light from a specific spectrum. The spectrum can be a subset of the electromagnetic (EM) spectrum. For example, each of the one or more cameras 510 can be a VL camera that responds to the visible light (VL) spectrum, an IR camera that responds to the infrared (IR) spectrum, a UV camera that responds to the ultraviolet (UV) spectrum, a camera that responds to light of another spectrum from another part of the electromagnetic spectrum, or some combination thereof.

[0125] One or more sensors 505 may include one or more other types of sensors in addition to camera 510, such as one or more of each of the following: accelerometers, gyroscopes, magnetometers, inertial measurement units (IMUs), altimeters, barometers, thermometers, radio detection and ranging (RADAR) sensors, light detection and ranging (LIDAR) sensors, sound navigation and ranging (SONAR) sensors, sound detection and ranging (SODAR) sensors, global navigation satellite system (GNSS) receivers, global positioning system (GPS) receivers, beidou navigation satellite system (BDS) receivers, galileo receivers, Globalnaya Navigazionnaya Sputnikovaya Sistema (GLONASS) receivers, navigation Indian constellation (NavIC) receivers, quasi-zenith satellite system (QZSS) receivers, Wi-Fi positioning system (WPS) receivers, cellular network positioning system receivers, beacon positioning receivers, short-range wireless beacon positioning receivers, personal area network (PAN) positioning receivers, wide area network (WAN) positioning receivers, wireless local area network (WLAN) positioning receivers, other types of positioning receivers, other types of sensors discussed herein, or combinations thereof. In some examples, one or more sensors 505 may include Figure 4 any combination of the sensors of system 400.

[0126] Figure 5 The SLAM system 500 of Figure 5 includes a visual-inertial odometry (VIO) tracker 515. The term visual-inertial odometry may also be referred to as visual odometry herein. The VIO tracker 515 receives sensor data 565 from one or more sensors 505. For example, the sensor data 565 may include one or more images captured by one or more cameras 510. The sensor data 565 may include other types of sensor data from one or more sensors 505, such as data from any type of sensor 505 listed herein. For example, the sensor data 565 may include IMU data from one or more inertial measurement units (IMUs) of one or more sensors 505.

[0127] After receiving sensor data 565 from one or more sensors 505, the VIO tracker 515 performs feature detection, extraction, and / or tracking using the feature tracking engine 520 of the VIO tracker 515. For example, in the case where the sensor data 565 includes one or more images captured by one or more cameras 510 of the SLAM system 500, the VIO tracker 515 can identify, detect, and / or extract features in each image. Features can include visually distinct points in an image, such as portions of the image depicting edges and / or corners. The VIO tracker 515 can periodically and / or continuously receive sensor data 565 from one or more sensors 505, such as by continuously receiving more images from one or more cameras 510 while the one or more cameras 510 capture video, where the images are video frames of the video. The VIO tracker 515 can generate descriptors for the features. The feature descriptors can be generated at least in part by generating a description of the features, as depicted in the local image patches surrounding the feature extraction. In some examples, the feature descriptors can describe the features as a collection of one or more feature vectors. In some cases, the VIO tracker 515 can be implemented using the hybrid system 500 discussed below with respect to Figure 5 as discussed.

[0128] In some cases, the VIO tracker 515, together with the mapping engine 530 and / or the relocalization engine 555, can associate multiple features with a map of the environment based on such feature descriptors. The feature tracking engine 520 of the VIO tracker 515 can perform feature tracking by identifying features in each image that the VIO tracker 515 has previously identified in one or more previous images, in some cases based on identifying features with matching feature descriptors in different images. The feature tracking engine 520 can track changes in the one or more positions depicting the features in each different image. For example, the feature extraction engine can detect a specific corner of a room depicted in the left side of a first image captured by a first camera of the cameras 510. The feature extraction engine can detect the same feature (e.g., the same specific corner of the same room) depicted in the right side of a second image captured by the first camera. The feature tracking engine 520 can identify that the features detected in the first image and the second image are two depictions of the same feature (e.g., the same specific corner of the same room), and that the feature appears in two different positions in the two images. The VIO tracker 515 can determine that the first camera has moved based on the same feature appearing on the left side of the first image and the right side of the second image, such as if the feature (e.g., the specific corner of the room) is a static part of the environment.

[0129] The VIO tracker 515 may include a sensor integration engine 525. The sensor integration engine 525 may use sensor data from other types of sensors 505 (in addition to the camera 510) to determine information that the feature tracking engine 520 may use when performing feature tracking. For example, the sensor integration engine 525 may receive IMU data from the IMU in one or more of the sensors 505 (e.g., it may be included as part of the sensor data 565). The sensor integration engine 525 may determine, based on the IMU data in the sensor data 565, that the SLAM system 500 has rotated 15 degrees in the clockwise direction from the acquisition or capture of the first image by the first camera in the camera 510 to the acquisition or capture of the second image. Based on this determination, the sensor integration engine 525 may identify that a feature depicted at a first position in the first image is expected to appear at a second position in the second image, and the second position is expected to be a predetermined distance (e.g., a predetermined number of pixels, inches, centimeters, millimeters, or another distance metric) to the left of the first position. The feature tracking engine 520 may consider this expectation when tracking features between the first image and the second image.

[0130] Based on feature tracking performed by the feature tracking engine 520 and / or sensor integration performed by the sensor integration engine 525, the VIO tracker 515 may determine the 3D feature position 572 of a particular feature. The 3D feature position 572 may include one or more 3D feature positions and may also be referred to as 3D feature points. The 3D feature position 572 may be a set of coordinates along three different axes that are perpendicular to each other, such as an X coordinate along the X axis (e.g., in the horizontal direction), a Y coordinate along the Y axis perpendicular to the X axis (e.g., in the vertical direction), and a Z coordinate along the Z axis perpendicular to both the X axis and the Y axis (e.g., in the depth direction). In some aspects, the VIO tracker 515 may also determine one or more key frames 570 (hereinafter referred to as key frames 570) corresponding to a particular feature. A key frame corresponding to a particular feature (from one or more key frames 570) may be an image in which the particular feature is clearly depicted. In some examples, a key frame corresponding to a particular feature (from one or more key frames 570) may be an image in which the particular feature is clearly depicted. In some examples, a key frame corresponding to a particular feature may be an image that reduces the uncertainty of the 3D feature position 572 of the particular feature when considered by the feature tracking engine 520 and / or the sensor integration engine 525 for determining the 3D feature position 572. In some examples, a key frame corresponding to a particular feature also includes data regarding the pose 585 of the SLAM system 500 and / or the camera 510 during the capture of the key frame. In some examples, the VIO tracker 515 may send the 3D feature position 572 and / or the key frames 570 corresponding to one or more features to the mapping engine 530. In some examples, the VIO tracker 515 may receive map tiles 575 from the mapping engine 530. The VIO tracker 515 may use the information within the map tiles 575 as features for feature tracking using the feature tracking engine 520.

[0131] Based on feature tracking performed by the feature tracking engine 520 and / or sensor integration performed by the sensor integration engine 525, the VIO tracker 515 may determine the pose 585 of the SLAM system 500 and / or the camera 510 during the capture of each image in the sensor data 565. The pose 585 may include the position of the SLAM system 500 and / or the camera 510 in 3D space, such as a set of coordinates along three different axes that are perpendicular to each other (e.g., an X coordinate, a Y coordinate, and a Z coordinate). The pose 585 may include the orientation of the SLAM system 500 and / or the camera 510 in 3D space, such as pitch, roll, yaw, or some combination thereof. In some examples, the VIO tracker 515 may send the pose 585 to the relocalization engine 555. In some examples, the VIO tracker 515 may receive the pose 585 from the relocalization engine 555.

[0132] The SLAM system 500 further includes a mapping engine 530. The mapping engine 530 can generate a 3D map of the environment based on the 3D feature positions 572 and / or key frames 570 received from the VIO tracker 515. The mapping engine 530 can include a map densification engine 535, a key frame remover 540, a beam adjuster 545, and / or a loop closure detector 550. The map densification engine 535 can perform map densification, in some examples, increasing the number and / or density of 3D coordinates that describe the map geometry. The key frame remover 540 can remove key frames, and / or in some cases add key frames. In some examples, the key frame remover 540 can remove key frames 570 corresponding to regions in the map to be updated and / or having a low corresponding confidence value. In some examples, the beam adjuster 545 can refine the 3D coordinates that describe the scene geometry, the parameters of the relative motion, and / or the optical characteristics of the image sensor used to generate the frame according to an optimality criterion involving the corresponding image projections of all points. The loop closure detector 550 can identify when the SLAM system 500 has returned to a previously mapped area, and can use this information to update map slices and / or reduce the uncertainty of certain 3D feature points or other points in the map geometry.

[0133] The mapping engine 530 can output map slices 575 to the VIO tracker 515. The map slices 575 can represent a 3D portion or subset of the map. The map slices 575 can include map slices 575 representing new, previously unmapped areas of the map. The map slices 575 can include map slices 575 representing updates (or modifications or revisions) to previously mapped areas of the map. The mapping engine 530 can output map information 580 to the relocalization engine 555. The map information 580 can include at least a portion of the map generated by the mapping engine 530. The map information 580 can include one or more 3D points that make up the geometry of the map, such as one or more 3D feature positions 572. The map information 580 can include one or more key frames 570 corresponding to certain features and certain 3D feature positions 572.

[0134] The SLAM system 500 also includes a relocalization engine 555. The relocalization engine 555 can perform relocalization, for example, when the VIO tracker 515 fails to identify a threshold number of features in an image, and / or when the VIO tracker 515 loses track of the pose 585 of the SLAM system 500 within the map generated by the mapping engine 530. The relocalization engine 555 can perform relocalization by performing extraction and matching using the extraction and matching engine 560. For example, the extraction and matching engine 560 can extract features from an image captured by the camera 510 of the SLAM system 500 when the SLAM system 500 is at the current pose 585, and can match the extracted features with features depicted in different key frames 570, identified by 3D feature positions 572, and / or recognized in the map information 580. By matching these extracted features with previously identified features, the relocalization engine 555 can identify that the pose 585 of the SLAM system 500 is the pose 585 at which the previously identified features are visible to the camera 510 of the SLAM system 500, and thus is similar to one or more previous poses 585 at which the previously identified features were visible to the camera 510. In some cases, the relocalization engine 555 can perform relocalization based on wide baseline mapping or the distance between the current camera position and the camera position at which the features were initially captured. The relocalization engine 555 can receive information about the pose 585 (e.g., information about one or more recent poses of the SLAM system 500 and / or the camera 510) from the VIO tracker 515, and the relocalization engine 555 can base its relocalization determination on this pose information. Once the relocalization engine 555 relocalizes the SLAM system 500 and / or the camera 510 and thus determines the pose 585, the relocalization engine 555 can output the pose 585 to the VIO tracker 515.

[0135] The SLAM system 500 can generate depth values (e.g., sparse depth values) as a byproduct of the processing by the other systems of the SLAM system described above. The depth values (e.g., sparse depth values) are used to perform post hoc scaling, as described in more detail below with reference to Figure 7 described in more detail.

[0136] Figure 6 An example frame 600 of a scene is illustrated. The frame 600 provides an illustrative example of feature information that can be captured and / or processed by a system (e.g., the system 400 shown in Figure 4 or the system 500 of Figure 5 ) during tracking and / or mapping. In the Figure 6 illustrated example, example features 602 are illustrated as circles of different diameters. In some cases, the center of each of the features 602 can be referred to as the feature center position. In some cases, the diameter of the circle can represent the feature scale (also referred to as the blob size) associated with each of the example features 602.

[0137] Each of the features 602 may also include a dominant orientation vector 603 illustrated as a radial segment. In one illustrative example, the dominant orientation vector 603 (also referred to herein as the dominant orientation) may be determined based on the pixel gradients within a block (also referred to as a blob or region). For example, the dominant orientation vector 603 may be determined based on the orientation of edge features in the neighborhood around the center of the feature (e.g., a block of nearby pixels). Another example feature 604 is shown as having a dominant orientation 606. In some embodiments, a feature may have multiple dominant orientations. For example, if no single orientation is clearly dominant, the feature may have two or more dominant orientations associated with the most prominent orientations. Another example feature 608 is illustrated as having two dominant orientation vectors 610 and 612.

[0138] In addition to the feature center location, blob size, and dominant orientation, each of the features 602, 604, 608 may also be associated with a descriptor that can be used to associate features between different frames. For example, if the pose of the camera capturing frame 600 changes, the x-y coordinates of each of the feature center locations of each of the features 602, 604, 608 may also change, and the descriptors assigned to each feature can be used to match features between two different frames. In some cases, the tracking and mapping operations of the XR system may utilize different types of descriptors for the features 602, 604, 608. Examples of descriptors for the features 602, 604, 608 may include SIFT, FREAK, and / or other descriptors. In some cases, the tracker may operate directly on the image blocks or may operate on the descriptors (e.g., SIFT descriptors, FREAK descriptors, etc.).

[0139] As previously described, machine learning-based systems (e.g., using deep learning neural networks) may be used in some cases to detect features (e.g., key points or feature points) for localization and mapping and generate descriptors for the detected features. However, it may be difficult to obtain ground truth and annotations (or labels) for training machine learning-based feature (e.g., key point or feature point) detectors and descriptor generators.

[0140] Figure 7A flowchart illustrating a system 700 is shown, which depicts an example of performing post - scaling of scale - ambiguous depth prediction data using a depth value 715 (e.g., a sparse depth value). As shown, a trained depth network 704 receives an image 702. The trained depth network 704 produces a scale - ambiguous depth prediction 706 of the image 702. As noted above, the challenge with self - supervised machine - learning depth network 704 is that its depth prediction values are typically based on the relative distances between objects in the image 702 or between different image frames (in general “units” rather than on a known specific scale such as meters or feet). In this regard, the depth prediction values are scale - ambiguous. Figure 7 The scale - ambiguous depth prediction 706 in Figure 7 shows lighter shading for the parts of the image 702 closer to the camera and darker parts for the parts farther from the camera. However, the shading in the scale - ambiguous depth prediction 706 does not indicate the actual value of the distance of the chair or table shown in the image 702 in a specific scale (e.g., meters in the metric system).

[0141] As previously described, the systems and techniques described herein provide post - scaling (using a post - scaling engine 708) of the scale - ambiguous depth prediction 706 to generate a scale - corrected depth prediction 716 of the image. The term post - hoc (Latin) means “after this” or “after the event”. For example, after the trained depth network 704 has produced a scale - ambiguous depth prediction, the scaling engine or post - scaling engine 708 will use a depth value (e.g., a sparse depth value) to correct the scale of the depth prediction. Post - hoc can also refer to post - analysis or post - testing or statistical analysis that was not specified before seeing the data. Post - hoc theorizing and generating hypotheses can be based on the data that has already been observed. In some cases, the data that has been observed is the scale - ambiguous depth prediction 706 obtained from the trained depth network 704.

[0142] The depth value 715 (e.g., a sparse depth value) can be obtained by using multiple images 710, 712 obtained from a camera tracker 714 that can generate the depth value 715. In one example, 100 pixels can have depth values for significant portions of the images 710, 712. The camera tracking engine or camera tracker 714 can use a six-degree-of-freedom (6DoF) tracking algorithm that can rely on significant feature points across the images 710, 712 and can match these significant points to solve for camera motion. In the 6DoF algorithm, there are three variables for rotation and three variables for translation. In some aspects, the camera tracker 714 can use a computer vision algorithm that does not have a deep learning component. In such aspects, the camera tracker 714 does not rely on training data or machine learning techniques. The camera tracker 714 can determine significant features in the image data of the images 710, 712 and can match these significant features such that the system 700 can solve an optimization to find the camera motion from one image 710 to another image 712. The depth of the significant points (shown as depth value 715, such as a sparse depth value) is a byproduct of the algorithm. Thus, the camera tracker 714 obtains the depth values of the significant points (e.g., a set of sparse significant points corresponding to sparse depth values). The depth value 715 can be obtained in meters or another measurement system such as feet via the camera tracker 714.

[0143] In some aspects, for each frame, the system can use one or more representative values (e.g., one or more statistical measures such as median, mean, or average or other representative values) of the depth value 715 (e.g., a sparse depth value) to scale the predicted depth map. For example, in some cases the following equation can be used: depth final = depth initial * median sparse / median pred . In such an equation, the system can use a representative value (e.g., the median) of the depth value 715 (e.g., a sparse depth value) for the corresponding frame (relative to one or more other frames). The depth initial value can be related to the initial depth prediction for a particular pixel. The median sparse value can represent the value of the entire frame as a median derived from the depth value 715 (e.g., a sparse depth value). The median sparse value can alternatively cover a region or sub-region of the entire frame that may or may not include the pixels associated with the depth initial value. The median pred can refer to across the entire frame or, in an alternative method, in a region of the entire frame that may or may not include the pixels associated with the depth initialThe predicted median on a sub-region of pixels associated with a value. This represents an example of scaling that converts the prediction to an appropriate unit such as metric or feet. The scale of the depth prediction is correct and is thus more useful for applications such as autonomous driving.

[0144] On the other hand, if a stereo camera is available, the system can use a feature matching algorithm to find depth values 715 (e.g., sparse depth values) by using a stereo image pair. These depth values 715 can then be used to scale the predictions from the machine learning network 704. The system 700 can compute the median of the depth values 715 obtained from the stereo image pair. Note that the difference is that the two images from the stereo camera are simultaneous in time, rather than sequential in time like images 710, 712. A camera tracking algorithm can be applied to the two stereo images in a similar manner as applying the algorithm to two temporally consecutive images to obtain the depth values 715. The application of the algorithm can be even simpler because the camera motion in this scenario is fixed for the stereo images. A camera tracking system with more than two images can also be deployed and depth values 715 (e.g., sparse depth values) can be obtained through a similar analysis.

[0145] As another illustrative example of a representative value, the system 700 can use a representative value (e.g., a statistical measure such as an average) between frames. On the other hand, instead of performing frame-level scaling (one scalar per frame), the system 700 can implement a method where different scalars can be used for different parts of the frame. For example, the system 700 can obtain depth values (e.g., sparse depth values) and perform regional scaling. Sparse feature points are scattered across the frame. The system 700 can divide the frame into a grid of multiple blocks such that the sparse values fall within a part or block of the grid, and the system 700 can obtain the scaling of the block based on the sparse values within the block of the grid.

[0146] In some aspects, for example, the system 700 can perform object detection and divide the image into regions based on the detected objects and generate foreground and background. The system 700 can collect the sparse values in the corresponding regions and determine one or more representative values (e.g., one or more statistical measures such as median, average, etc.) of the sparse values in the corresponding regions and use the one or more representative values for the post-scaling engine 708. This process can also be performed for a group of regions (e.g., a group of objects) that may all fall within the foreground (e.g., a tractor and a tree or other grouping), where the median, mean, or other representative value of the sparse values associated with the group of regions is obtained, while the background regions (mountains and sky or other grouping) may have different median / mean / other representative values of the sparse values associated with the background regions. To determine the median / mean / other value of the corresponding group of objects, different types of objects can be grouped together into various different types of groups.

[0147] In one aspect, the post hoc scaling engine 708 may provide correct metric (or other unit) scaling to the depth predictions 706 from the self-supervised network 704. Computationally, this benefit may come for free since 6DoF camera tracking algorithms operating on the camera tracker 714 will typically need to run on the device. The system 700 may be evaluated on an internal XR benchmark, which consists of multiple scenarios such as eight scenarios. The inventors have found that when using the post hoc median scaling as described above, there is a significant improvement in the absolute relative error associated with depth predictions.

[0148] Figure 7 The system 700 may support different configurations. For example, in one example configuration, the system 700 may include a camera tracker 714 and a post hoc scaling module engine 708. The camera tracker 714 and / or the post hoc scaling engine 708 may then receive a scale-ambiguous depth prediction 706 from an external device and generate a scale-corrected depth prediction 716. In another example configuration, the system 700 may include a trained depth network 704 as well as a camera tracker 714 and a post hoc scaling engine 708. In yet another example configuration, the system 700 may include a post hoc scaling engine 708 that receives a scale-ambiguous depth prediction 706 and a depth value 715 (e.g., a sparse depth value), and performs a post hoc scaling operation to produce a scale-corrected depth prediction 716.

[0149] Figure 8 is a flow chart illustrating an example of a process 800 for processing image and / or video data. The process 800 may be executed by a computing device (or apparatus) or by a component or system of a computing device (e.g., a chipset). The computing device (or its component or system) may include or may be Figure 5 the system 500, Figure 7 the system 700 or any of their components. The operations of the process 800 may be implemented as software components that execute and run on one or more processors (e.g., Figure 9 the processor 910 or other processors). Additionally, the transmission and reception of signals by a first network entity in the process 800 may be enabled, for example, by one or more antennas and / or one or more transceivers such as a wireless transceiver.

[0150] At block 802, a computing device (or its components or systems) may use a trained machine learning system to determine a predicted depth map of an image. The predicted depth map includes corresponding predicted depth values for each pixel of the image (e.g., input images 702, 710, 712). In an illustrative example, the trained machine learning system is a trained neural network. In some cases, the input data includes one or more images, radar data, light detection and ranging (LIDAR) data, any combination thereof, and / or other data. In one illustrative example, the input data includes a first image of a scene having a first characteristic, a second image of a scene having a second characteristic, and a third image of a scene having a third characteristic. These characteristics may relate to movement of the camera or movement within a single image.

[0151] At block 804, a computing device (or its components or systems) may obtain depth values of an image (e.g., depth values 715, such as sparse depth values) from a tracker (e.g., using camera tracker 714), which is configured to determine depth values based on one or more feature points between frames. In some cases, the tracker is six degrees of freedom (6DoF). The depth values 715 include depth values for fewer than all pixels of the image (e.g., sparse depth values). In one aspect, the camera tracker 714 may be configured to use a 6DoF tracking algorithm to generate depth values 715 based on matching identified significant feature values across multiple frames (e.g., which may be a series of frames in time or a set of stereo images) and solving for camera motion. In some cases, these frames include one or more stereo image pairs. For example, a computing device (or its components or systems) may obtain depth values 715 from a feature tracker (e.g., camera tracker 714), which is configured to determine depth values 715 from one or more stereo image pairs.

[0152] At block 806, using a post-scaling engine 708, a computing device (or its components or systems) may scale the predicted depth map of the image using the depth values 715 (e.g., sparse depth values). In an illustrative example, a computing device (or its components or systems) may scale the predicted depth map using a representative value of the depth values 715. In an illustrative example, a computing device (or its components or systems) may scale the predicted depth map associated with the image using a first representative value of the depth values 715 and a second representative value of the predicted depth values of the predicted depth map.

[0153] In another example, the first representative value may include a first statistical measure (e.g., mean) of the depth values or a second statistical measure (e.g., median) of the depth values. The second representative value may include the mean of the predicted depth values of the predicted depth map or the median of the predicted depth values of the predicted depth map.

[0154] In another illustrative example, a computing device (or its components or systems) may scale a predicted depth map by determining a final depth map based on multiplying the predicted depth map by a scaling factor. In an illustrative example, the scaling factor may include a relationship between a first representative value of depth values and a second representative value of predicted depth values of the predicted depth map. In one example, the relationship includes a ratio of the first representative value of depth values to the second representative value of predicted depth values of the predicted depth map.

[0155] In another example, the first representative value may include a first statistical measure (e.g., an average value) of depth values or a second statistical measure (e.g., a median value) of depth values. The second representative value may include an average value of predicted depth values of the predicted depth map or a median value of predicted depth values of the predicted depth map.

[0156] The computing device (or apparatus) may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a connected watch or a smart watch, or other wearable devices), a server computer, a vehicle (e.g., an autonomous vehicle or a semi-autonomous vehicle) or a computing device or system of a vehicle, a robotic device, a laptop computer, a smart TV, a camera, and / or any other computing device having the resource capabilities to perform the processes described herein (including process 800 and / or other processes described herein). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on Internet Protocol (IP) or other types of data.

[0157] The components of the computing device may be implemented in circuitry. For example, the components may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0158] Process 800 is illustrated as a logical flow diagram, and the operations of the logical flow diagram represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally speaking, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0159] Additionally, process 800 and / or any other process described herein can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, e.g., in the form of a computer program that includes a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.

[0160] Figure 9 Example computing device architecture 900 of an example computing device that can implement the various techniques described herein is illustrated. In some examples, the computing device can include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device of a vehicle), or other devices. For example, computing device architecture 900 can implement Figure 7 system 700 or any component thereof as a separate aspect. The components of computing device architecture 900 are shown to communicate electrically with each other using a connection member 905 (such as a bus). Example computing device architecture 900 includes a processing unit (CPU or processor) 910 and a computing device connection member 905 that couples various computing device components including a computing device memory 915 (such as a read-only memory (ROM) 920 and a random access memory (RAM) 925) to the processor 910.

[0161] The computing device architecture 900 may include a cache that is directly connected to the processor 910, very close to the processor, or integrated as part of the processor, which is a high-speed memory. The computing device architecture 900 may copy data from the memory 915 and / or the storage device 930 to the cache 912 for quick access by the processor 910. In this way, the cache can provide a performance boost to avoid delays when the processor 910 is waiting for data. These engines and other engines may control or be configured to control the processor 910 to perform various actions. Other computing device memories 915 may also be used. The memory 915 may include various different types of memories with different performance characteristics. The processor 910 may include any general-purpose processor and hardware or software services (such as service 1 932, service 2 934, and service 3 936 stored in the storage device 930) configured to control the processor 910, as well as a dedicated processor in which software instructions are incorporated into the processor design. The processor 910 may be a self - contained system, including multiple cores or processors, buses, memory controllers, caches, etc. The multi - core processor may be symmetric or asymmetric.

[0162] To enable user interaction with the computing device architecture 900, the input device 945 may represent any number of input mechanisms, such as a microphone for voice, a touch - sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The output device 935 may also be one or more of the multiple output mechanisms known to those skilled in the art, such as a display, a projector, a television, a speaker device, etc. In some instances, a multimodal computing device may enable a user to provide multiple types of input to communicate with the computing device architecture 900. The communication interface 940 generally may govern and manage user input and computing device output. There is no limitation on operating on any specific hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0163] The storage device 930 is non - volatile memory and may be a hard disk or other types of computer - readable media that can store computer - accessible data, such as a magnetic tape cassette, a flash memory card, a solid - state memory device, a digital versatile disc, a cassette tape, a random - access memory (RAM) 925, a read - only memory (ROM) 920, and their hybrid forms. The storage device 930 may include services 932, 934, 936 for controlling the processor 910. Other hardware or software modules or engines are contemplated. The storage device 930 may be connected to the computing device connection 905. In one aspect, a hardware module that performs a specific function may include software components stored in a computer - readable medium connected to the necessary hardware components (such as the processor 910, the connection 905, the output device 935, etc.) to perform the function.

[0164] Aspects of the present disclosure are applicable to any suitable electronic device (such as a security system, a smart phone, a tablet computer, a laptop computer, a vehicle, a drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although described below with respect to a device having or coupled to one light projector, aspects of the present disclosure may be applicable to devices having any number of light projectors and are thus not limited to a particular device.

[0165] The term "device" is not limited to one or a particular number of physical objects (such as a smart phone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that can implement at least some portions of the present disclosure. Although the following description and examples use the term "device" to describe various aspects of the present disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or a particular aspect. For example, a system can be implemented on one or more printed circuit boards or other substrates and can have components that are movable or static. Although the following description and examples use the term "system" to describe various aspects of the present disclosure, the term "system" is not limited to a particular configuration, type, or number of objects.

[0166] Specific details are provided in the above description to provide a thorough understanding of the aspects and examples provided herein. However, one of ordinary skill in the art will understand that these aspects can be practiced without these specific details. For clarity, in some instances, the techniques may be presented as including separate functional blocks, including functional blocks that contain devices, device components, steps or routines in a method embodied in software or hardware and software combinations. Additional components other than those shown and / or described herein can be used. For example, circuits, systems, networks, processes, and other components can be shown in block diagram form as components to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques can be shown without unnecessary detail to avoid obscuring the aspects.

[0167] Aspects can be described above as a process or method, which is depicted as a flowchart, a process diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe operations as a sequential process, many of the operations can be performed in parallel or concurrently. Additionally, the order of the operations can be rearranged. A process is terminated when the operations of the process are completed, but a process can have additional steps not included in the figures. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the termination of the process can correspond to the function returning to the calling function or the main function.

[0168] The processes and methods according to the above examples can be implemented using computer-executable instructions stored or otherwise obtained from a computer-readable medium. Such instructions can include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions. Portions of the computer resources used can be accessed via a network. The computer-executable instructions can be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.

[0169] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium can include non-transitory media in which data can be stored and that do not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transitory media can include, but are not limited to, magnetic disks or tapes, optical storage media (such as flash memory), memories or memory devices, magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, compact discs (CDs) or digital versatile discs (DVDs), any suitable combination thereof, etc. The computer-readable medium can have code and / or machine-executable instructions stored thereon that can represent a process, function, subroutine, program, routine, subroutine, module, engine, software package, class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. The information, arguments, parameters, data, etc. can be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0170] In some aspects, computer-readable storage devices, media, and memories can include wires or wireless signals, such as bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.

[0171] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may execute the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or plug-in cards. By way of additional example, such functionality may also be implemented on a circuit board among different chips or different processes executed on a single device.

[0172] Instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0173] In the above description, aspects of the present application are described with reference to their specific aspects, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the illustrative aspects of the present application have been described in detail herein, it is to be understood that the various inventive concepts may be implemented and adopted in other various ways, and the appended claims are not to be construed as including such variations, unless limited by the prior art. The various features and aspects of the above applications may be used singly or in combination. In addition, without departing from the broader spirit and scope of this specification, the aspects may be utilized in any number of environments and applications beyond those described herein. Therefore, the specification and drawings are to be regarded as illustrative rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be understood that in alternative aspects, the methods may be performed in a different order than that described.

[0174] Those of ordinary skill in the art will appreciate that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of this specification.

[0175] In cases where a component is described as “configured to” perform certain operations, such configuration may be implemented, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0176] The phrase "coupled to" means that any component is directly or indirectly physically connected to another component, and / or any component directly or indirectly communicates with another component (e.g., is connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0177] Claim language or other language that recites "at least one of" a set and / or "one or more" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language that recites "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language that recites "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" a set and / or "one or more" in a set does not limit the set to the items listed in the set. For example, claim language that recites "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.

[0178] Claim language or other language that recites "at least one processor, the at least one processor being configured to", "at least one processor being configured to", "one or more processors, the one or more processors being configured to", "one or more processors being configured to", etc. indicates that one processor or multiple processors (in any combination) can perform the associated operations. For example, claim language that recites "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or multiple processors are each assigned tasks for a particular subset of operations X, Y, and Z such that the multiple processors together perform X, Y, and Z; or a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language that recites "at least one processor, the at least one processor being configured to: X, Y, and Z" can mean that any single processor can perform at least a subset of operations X, Y, and Z.

[0179] In the case of referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may perform these functions jointly. When more than one element performs these functions jointly, each function does not need to be performed by each of those elements (e.g., different functions may be performed by different elements), and / or each function does not need to be performed only by one element as a whole (e.g., different elements may perform different sub - functions of a function). Similarly, in the case of referring to one or more elements that are configured to cause another element (e.g., a device) to perform a function, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.

[0180] In the case of referring to an entity (e.g., any entity or device described herein) that performs a function or is configured to perform a function (e.g., a step of a method), the entity may be configured to cause one or more elements (individually or jointly) to perform these functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of these functions, and / or any combination thereof. When referring to an entity that performs a function, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform these functions jointly. When the entity is configured to cause more than one component to perform these functions jointly, each function does not need to be performed by each of those components (e.g., different functions may be performed by different components), and / or each function does not need to be performed only by one component as a whole (e.g., different components may perform different sub - functions of a function).

[0181] The various illustrative logical blocks, modules, engines, circuits, and algorithmic steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0182] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device such as a cellular phone, or an integrated circuit device having multiple uses, including applications in wireless communication devices such as cellular phones and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially realized by a computer-readable data storage medium including program code, the program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially realized by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0183] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.

[0184] Exemplary aspects of the present disclosure include:

[0185] Aspect 1. A device for scaling depth prediction, the device comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: use a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including corresponding predicted depth values for each pixel of the image; obtain depth values of the image, the depth values including depth values for less than all pixels of the image; and use the depth values to scale the predicted depth map of the image.

[0186] Aspect 2. The device according to aspect 1, wherein the at least one processor is configured to obtain the depth values from a tracker (e.g., a six-degree-of-freedom (6DoF) tracker), the tracker being configured to determine the depth values based on one or more feature points between frames.

[0187] Aspect 3. The device according to aspect 2, wherein the 6DoF tracker is configured to use a 6DoF tracking algorithm to generate the depth values based on significant feature values identified by matching across multiple frames and solving for camera motion.

[0188] Aspect 4. The device according to any one of aspects 1 to 3, wherein the at least one processor is configured to obtain the depth values from a feature tracker, the feature tracker being configured to determine the depth values from one or more stereo image pairs.

[0189] Aspect 5. The device according to any one of aspects 1 to 4, wherein the at least one processor is configured to: scale the predicted depth map using a representative value of the depth values.

[0190] Aspect 6. The device according to aspect 5, wherein the representative value includes an average value of the depth values or a median value of the depth values.

[0191] Aspect 7. The device according to either aspect 5 or 6, wherein the at least one processor is configured to: scale the predicted depth map associated with the image using a first representative value of the depth values and a second representative value of the predicted depth values of the predicted depth map.

[0192] Aspect 8. The device according to aspect 7, wherein the first representative value includes a first statistical measure of the depth values or a second statistical measure value of the depth values, and wherein the second representative value includes a first statistical measure of the predicted depth values of the predicted depth map or a second statistical measure of the predicted depth values of the predicted depth map.

[0193] Aspect 9. The apparatus according to any one of aspects 1 to 8, wherein, in order to scale the predicted depth map, the at least one processor is configured to: determine a final depth map based on multiplying the predicted depth map by a scaling factor.

[0194] Aspect 10. The apparatus according to aspect 9, wherein the scaling factor includes a relationship between a first representative value of the depth values and a second representative value of the predicted depth values of the predicted depth map.

[0195] Aspect 11. The apparatus according to aspect 10, wherein the first representative value includes a first statistical measure of the depth values or a second statistical measure value of the depth values, and wherein the second representative value includes a first statistical measure of the predicted depth values of the predicted depth map.

[0196] Aspect 12. The apparatus according to any one of aspects 1 to 11, wherein the trained machine learning system is a trained neural network.

[0197] Aspect 13. A method for processing image data, the method comprising: using a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including respective predicted depth values for each pixel of the image; obtaining depth values of the image, the depth values including depth values for less than all pixels of the image; and using the depth values to scale the predicted depth map of the image.

[0198] Aspect 14. The method according to aspect 13, the method further comprising: obtaining the depth values from a six degrees of freedom (6DoF) tracker configured to determine the depth values at least in part by identifying one or more salient feature point frames.

[0199] Aspect 15. The method according to aspect 14, wherein the 6DoF tracker is configured to use a 6DoF tracking algorithm to generate the depth values based on matching identified salient feature values across multiple frames and solving for camera motion.

[0200] Aspect 16. The method according to any one of aspects 13 to 15, the method further comprising: obtaining the depth values from a feature tracker configured to determine the depth values from one or more stereo image pairs.

[0201] Aspect 17. The method according to any one of aspects 13 to 17, the method further comprising: scaling the predicted depth map using a representative value of the depth values.

[0202] Aspect 18. The method according to aspect 17, wherein the representative value includes an average value of the depth values or a median value of the depth values.

[0203] Aspect 19. The method according to any one of aspects 17 or 18, the method further comprising: scaling the predicted depth map associated with the image using a first representative value of the depth values and a second representative value of the predicted depth values of the predicted depth map.

[0204] Aspect 20. The method according to aspect 19, wherein the first representative value includes a first statistical measure of the depth values or a second statistical measure value of the depth values, and wherein the second representative value includes a first statistical measure of the predicted depth values of the predicted depth map or a second statistical measure of the predicted depth values of the predicted depth map.

[0205] Aspect 21. The method according to any one of aspects 13 to 20, wherein in order to scale the predicted depth map, the method further comprises: determining a final depth map based on multiplying the predicted depth map by a scaling factor.

[0206] Aspect 22. The method according to aspect 21, wherein the scaling factor includes a relationship between a first representative value of the depth values and a second representative value of the predicted depth values of the predicted depth map.

[0207] Aspect 23. The method according to aspect 22, wherein the first representative value includes a first statistical measure of the depth values or a second statistical measure value of the depth values, and wherein the second representative value includes a first statistical measure of the predicted depth values of the predicted depth map.

[0208] Aspect 24. The method according to any one of aspects 13 to 23, wherein the trained machine learning system is a trained neural network.

[0209] Aspect 25. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform any of the operations according to any one of aspects 13 to 24.

[0210] Aspect 26. An apparatus, the apparatus comprising components for performing any of the operations according to any one of aspects 13 to 24.

Claims

1. An apparatus for depth prediction scaling, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: use a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including corresponding predicted depth values for each pixel of the image; obtain depth values of the image from a tracker, the tracker configured to determine the depth values based on one or more feature points between frames, the depth values including depth values for less than all pixels of the image; and use the depth values to scale the predicted depth map of the image.

2. The apparatus according to claim 1, wherein the tracker is a six - degree - of - freedom (6DoF) tracker.

3. The apparatus according to claim 2, wherein the 6DoF tracker is configured to use a 6DoF tracking algorithm to generate the depth values based on significant feature values identified by matching across multiple frames and solving for camera motion.

4. The apparatus according to claim 1, wherein the frames include one or more stereo image pairs.

5. The apparatus according to claim 1, wherein the at least one processor is configured to: scale the predicted depth map using a representative value of the depth values.

6. The apparatus according to claim 5, wherein the at least one processor is configured to: scale the predicted depth map associated with the image using a first representative value of the depth values and a second representative value of the predicted depth values of the predicted depth map.

7. The apparatus according to claim 6, wherein the first representative value includes a first statistical measure of the depth values or a second statistical measure value of the depth values, and wherein the second representative value includes a first statistical measure of the predicted depth values of the predicted depth map or a second statistical measure of the predicted depth values of the predicted depth map.

8. The apparatus according to claim 1, wherein, to scale the predicted depth map, the at least one processor is configured to: determine a final depth map based on multiplying the predicted depth map by a scaling factor.

9. The apparatus according to claim 8, wherein the scaling factor includes a relationship between a first representative value of the depth values and a second representative value of the predicted depth values of the predicted depth map.

10. The apparatus according to claim 9, wherein the first representative value includes a first statistical measure of the depth values or a second statistical measure value of the depth values, and wherein the second representative value includes a first statistical measure of the predicted depth values of the predicted depth map or a second statistical measure of the predicted depth values of the predicted depth map.

11. A method for processing image data, the method comprising: using a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including corresponding predicted depth values for each pixel of the image; Obtain the depth values of the image from a tracker, the tracker being configured to determine the depth values based on one or more feature points between frames, the depth values including depth values of less than all pixels of the image; and Use the depth values to scale the predicted depth map of the image.

12. The method according to claim 11, wherein the tracker is a six-degree-of-freedom (6DoF) tracker.

13. The method according to claim 12, wherein the 6DoF tracker is configured to use a 6DoF tracking algorithm to generate the depth values based on significant feature values identified by matching across multiple frames and solving for camera motion.

14. The method according to claim 11, wherein the frames include one or more stereo image pairs.

15. The method according to claim 11, the method further comprising: Using a representative value of the depth values to scale the predicted depth map.

16. The method according to claim 15, the method further comprising: Using a first representative value of the depth values and a second representative value of the predicted depth values of the predicted depth map to scale the predicted depth map associated with the image.

17. The method according to claim 16, wherein the first representative value includes a first statistical measure of the depth values or a second statistical measure value of the depth values, and wherein the second representative value includes a first statistical measure of the predicted depth values of the predicted depth map or a second statistical measure of the predicted depth values of the predicted depth map.

18. The method according to claim 11, wherein scaling the predicted depth map includes: Determining a final depth map based on multiplying the predicted depth map by a scaling factor.

19. The method according to claim 18, wherein the scaling factor includes a relationship between a first representative value of the depth values and a second representative value of the predicted depth values of the predicted depth map.

20. The method according to claim 19, wherein the first representative value includes a first statistical measure of the depth values or a second statistical measure value of the depth values, and wherein the second representative value includes a first statistical measure of the predicted depth values of the predicted depth map or a second statistical measure of the predicted depth values of the predicted depth map.

21. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following operations: Use a trained machine learning system to determine a predicted depth map of an image, the predicted depth map including respective predicted depth values for each pixel of the image; Obtain the depth values of the image from a tracker, the tracker being configured to determine the depth values based on one or more feature points between frames, the depth values including depth values of less than all pixels of the image; and Use the depth values to scale the predicted depth map of the image.

22. The non-transitory computer-readable storage medium according to claim 21, wherein the tracker is a six-degree-of-freedom (6DoF) tracker.

23. The non-transitory computer-readable storage medium according to claim 22, wherein the 6DoF tracker is configured to use a 6DoF tracking algorithm to generate the depth value based on significant eigenvalue identified by matching across multiple frames and solving for camera motion.

24. The non-transitory computer-readable storage medium according to claim 21, wherein the frame includes one or more stereo image pairs.

25. The non-transitory computer-readable storage medium according to claim 21, wherein the instructions, when executed by the one or more processors, cause the one or more processors to: scale the predicted depth map using a representative value of the depth value.

26. The non-transitory computer-readable storage medium according to claim 21, wherein the instructions, when executed by the one or more processors, cause the one or more processors to: scale the predicted depth map associated with the image using a first representative value of the depth value and a second representative value of the predicted depth value of the predicted depth map.

27. The non-transitory computer-readable storage medium according to claim 21, wherein, in order to scale the predicted depth map, the instructions, when executed by the one or more processors, cause the one or more processors to: determine a final depth map based on multiplying the predicted depth map by a scaling factor.

28. The non-transitory computer-readable storage medium according to claim 27, wherein the scaling factor includes a relationship between a first representative value of the depth value and a second representative value of the predicted depth value of the predicted depth map.