Depth estimation using image and sparse depth input

By using neural network models to process images and sparse depth inputs in monocular depth estimation, the problem of depth scale blur is solved and the accuracy of depth estimation is achieved.

CN120019409APending Publication Date: 2025-05-16QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073747.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-12
Filing Date
2023-09-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the fuzzy problem of depth scale in monocular depth estimation, which affects the accuracy of depth prediction.

Method used

Deep output is generated by using machine learning systems, especially neural network models, combining images and sparse depth inputs. The system includes an encoder and a decoder, which processes image and depth information to generate feature representations, and the decoder generates corresponding depth outputs.

Benefits of technology

This method can effectively eliminate the blur of depth scale, improve the accuracy of depth estimation, and generate high-quality depth output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019409A_ABST
    Figure CN120019409A_ABST
Patent Text Reader

Abstract

Systems and techniques are provided for generating depth information from one or more images. For example, a method may include obtaining an image of a scene and obtaining depth information associated with one or more objects in the scene. The method may include processing the image and the depth information using an encoder of a neural network model to generate a feature representation of the image and the depth information. The method may also include processing the image and the feature representation of the depth information using a decoder of the neural network model to generate a depth output corresponding to the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to depth estimation based on one or more images. For example, aspects of the present disclosure relate to systems and techniques for performing depth estimation based on images and sparse depth input using a machine learning system. Background Art

[0002] Machine learning models (e.g., deep learning models such as neural networks) can be used to perform a variety of tasks, including depth estimation, detection and / or recognition (e.g., scene or object detection and / or recognition), pose estimation, image reconstruction, classification, three-dimensional (3D) modeling, dense regression tasks, data compression and / or decompression, image processing, etc. Machine learning models can be general and can achieve high-quality results in a variety of tasks. Summary of the invention

[0003] The following presents a simplified summary of the invention related to one or more aspects disclosed herein. Therefore, the following summary of the invention should neither be considered as an exhaustive overview related to all conceived aspects, nor should it be considered to identify key or decisive elements related to all conceived aspects or to delineate the scope associated with any particular aspect. Therefore, the sole purpose of the following summary of the invention is to present certain concepts related to one or more aspects related to the mechanisms disclosed herein in a simplified form before the detailed embodiments presented below.

[0004] Described herein are systems and techniques for performing depth estimation based on images (e.g., one or more dense grayscale or color images) and sparse depth input using a machine learning system (e.g., a neural network system or model). In some cases, the machine learning system can be trained using self-supervised learning.

[0005] According to at least one example, a method for generating depth information based on one or more images is provided. The method may include: obtaining an image of a scene; obtaining depth information associated with one or more objects in the scene; processing the image and the depth information using an encoder of a neural network model to generate a feature representation of the image and the depth information; and processing the feature representation of the image and the depth information using a decoder of the neural network model to generate a depth output corresponding to the image.

[0006] In another example, a device for generating depth information based on one or more images is provided, the device comprising at least one memory and at least one processor, the at least one processor being coupled to the at least one memory. The at least one processor may be configured to: obtain an image of a scene; obtain depth information associated with one or more objects in the scene; process the image and the depth information using an encoder of a neural network model to generate a feature representation of the image and the depth information; and process the feature representation of the image and the depth information using a decoder of the neural network model to generate a depth output corresponding to the image.

[0007] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: obtain an image of a scene; obtain depth information associated with one or more objects in the scene; process the image and the depth information using an encoder of a neural network model to generate a feature representation of the image and the depth information; and process the feature representation of the image and the depth information using a decoder of the neural network model to generate a depth output corresponding to the image.

[0008] In another example, an apparatus for generating depth information from one or more images is provided. The apparatus may include: a component for obtaining an image of a scene; a component for obtaining depth information associated with one or more objects in the scene; a component for processing the image and the depth information using an encoder of a neural network model to generate a feature representation of the image and the depth information; and a component for processing the feature representation of the image and the depth information using a decoder of the neural network model to generate a depth output corresponding to the image.

[0009] In some aspects, one or more of the devices described herein are, are part of, and / or include an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device or wireless communication device (e.g., a mobile phone or other mobile device), a wearable device (e.g., a connected watch or other wearable device), a camera, a personal computer, a laptop, a vehicle or a computing device or component of a vehicle, a server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the above-mentioned device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors).

[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.

[0011] The foregoing and other features and aspects will become more apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are presented to aid in describing various aspects of the present disclosure and are provided solely for illustration and not limitation of the various aspects:

[0013] Figure 1A is an image of the scene;

[0014] Figure 1B yes Figure 1A an image of a scene having virtual content added based on a depth of the scene depicted in the image;

[0015] Figure 1C is an example of an image having virtual content added based on the depth of a scene depicted in the image;

[0016] Figure 1D is another example of an image with virtual content based on the depth of the scene depicted in the image;

[0017] Figure 2 is a diagram illustrating an example of depth scale blur when determining depth from a single image according to aspects of the present disclosure;

[0018] Figure 3 is a diagram illustrating an example of a machine learning system configured to generate a depth output based on an image and a sparse depth input in accordance with aspects of the present disclosure;

[0019] Figure 4 is an example of an image with a seed point projected to an image plane of the image according to aspects of the present disclosure;

[0020] Figure 5A is an example of a method that can be provided as input according to various aspects of the present disclosure Figure 3 An illustration of an example of color channels of an image of a machine learning system;

[0021] Figure 5B is an example of a method that can be provided as input according to various aspects of the present disclosure Figure 3 An illustration of an example of a sparse depth map for a machine learning system;

[0022] Figure 5C is an example of a method that can be provided as input according to various aspects of the present disclosure Figure 3 A diagram of an example of an effectiveness graph for a machine learning system;

[0023] Figure 5D is an example of an input that can be provided to Figure 3 An illustration of an example of a sparse inverse depth map of a machine learning system;

[0024] Figure 6 is a flow chart illustrating an example process for generating depth information from one or more images in accordance with aspects of the present disclosure;

[0025] Figure 7 is a block diagram illustrating an example of a deep learning network according to some examples;

[0026] Figure 8 is a block diagram illustrating an example of a convolutional neural network according to some examples; and

[0027] Fig. 9 is a diagram illustrating an example system architecture for implementing certain aspects described herein. DETAILED DESCRIPTION

[0028] Some aspects and examples of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and examples can be applied independently, and some of them can be applied in combination. In the following description, specific details are set forth for explanation purposes to provide a thorough understanding of aspects and examples of the present disclosure. However, it is apparent that various aspects and examples can be practiced without these specific details. Each drawing and description are not intended to be restrictive.

[0029] The following description provides only exemplary aspects and examples, and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of exemplary aspects and examples will provide an enabling description for implementing aspects and examples of the present disclosure to those skilled in the art. It should be understood that various changes may be made to the functions and arrangements of the elements without departing from the scope of the present application as set forth in the appended claims.

[0030] As described above, machine learning systems (e.g., deep neural network systems or models) can be used to perform a variety of tasks, such as, but not limited to, detection and / or recognition (e.g., scene or object detection and / or recognition, face detection and / or recognition, etc.), depth estimation, pose estimation, image reconstruction, classification, three-dimensional (3D) modeling, dense regression tasks, data compression and / or decompression, and image processing, etc. In addition, machine learning models can be general and can achieve high-quality results in a variety of tasks.

[0031] In some cases, a machine learning system may perform depth prediction based on a single image (e.g., based on receiving a single image as input). Depth prediction based on a single input image may be referred to as monocular depth estimation. Monocular depth estimation may be used in many applications (e.g., XR applications, transportation applications, etc.). In some cases, monocular depth estimation may be used to perform occlusion rendering, such as rendering virtual objects in a 3D environment based on using depth and object segmentation information. In some cases, monocular depth prediction may be used to perform 3D reconstruction, such as based on creating a mesh of a scene using depth information and one or more poses. In some cases, monocular depth prediction may be used to perform collision avoidance, such as based on estimating the distance to one or more objects using depth information.

[0032] Depth estimation (e.g., such as monocular depth estimation) can be used to generate three-dimensional content (e.g., such as XR content) with higher accuracy. For example, monocular depth estimation can be used to generate XR content that combines a baseline image or video with one or more enhanced overlays of rendered 3D objects. The baseline image data (e.g., image or video frame) enhanced or overlaid by the XR system can be a two-dimensional (2D) representation of a 3D scene. A natural approach to generating XR content can be to overlay rendered objects onto the baseline image data without compensating for 3D depth information that can be represented in the 2D baseline image data.

[0033] For example, Figure 1A An example image 102 depicts a scene including a first person in the foreground of the scene and a table behind the first person. Figure 1A The example image 102 may be used or otherwise provided as 2D baseline image data to be enhanced with one or more virtual object overlays. For example, the image 102 may be provided to an XR system as baseline image data for generating rendered XR content.

[0034] For example, Figure 1B Describes the use of Figure 1A Here, the enhanced frame may be generated as the baseline image 102 superimposed with the virtual content item 103. For example, the virtual content item 103 may be a rendered 3D object, a rendered 2D object, etc. As shown, the virtual content item 103 is a drinking glass, which is superimposed on top of the baseline image 102 so that the virtual content item 103 (e.g., the drinking glass) appears to be located on top of the table in the foreground of the scene represented in the baseline image 102. Figure 1B As shown, when virtual content item 103 is superimposed on top of example image 102, a portion of virtual content item 103 blocks a portion of a first person (e.g., the person depicted in example image 102), while the portion of virtual content item 103 should be blocked by the portion of the first person (e.g., based on the first person being closer to the camera than the virtual content item). Based on depth information of the scene depicted in the image, the machine learning system can perform occlusion rendering such that the portion of virtual content item 103 is rendered behind the portion of the first person.

[0035] Figure 1Cis an example of an image 104 with a virtual blanket 105 added based on depth information of a scene depicted in the image 104. For example, the image 104 may be a frame of video data included in a plurality of frames of video data. As previously described, the depth information of the scene depicted in the image 104 may be used to generate XR content that combines a baseline image or video with one or more augmented overlays of rendered 3D objects. For example, the depth information may be determined for some (or all) of the plurality of frames of video data (e.g., including the image 104) and used by a machine learning system to perform occlusion rendering to accurately render the virtual blanket 105 added to the scene depicted in the image 104. The depth information may be used to render the virtual blanket 105 at a constant and / or consistent position within the scene of the example image 104 for different frames and different perspectives. For example, as the perspective or point of view (POV) changes between frames of video data, the depth information may be used to render the virtual blanket 105 at the same relative position relative to other foreground objects included in the scene depicted in the image 104. Figure 1D 1 is an example of an image 106 having a virtual block 107 added based on depth information of a scene depicted in the image 106. For example, the virtual block 107 may include a plurality of virtual blocks. The depth information may be determined for the image 106 (and / or may be determined for another image included in the same video data or frame group as the image 106) and used to render the virtual block 107 in the scene depicted in the image 106. For example, the depth information may be used to render physically accurate or realistic interactions between objects included in the scene of the image 106 and corresponding virtual blocks in the virtual blocks 107. As previously described, the depth information may be used to render the virtual block 107 at a constant and / or consistent position relative to other foreground objects depicted in the scene of the image 106.

[0036] In some examples, a machine learning system may perform monocular depth estimation to determine depth information associated with an image (e.g., image data or video data). Monocular depth estimation can be efficient because it can be performed based on receiving a single image or frame as input. However, when generating a monocular depth prediction (e.g., when performing monocular depth estimation), the neural network may have difficulty eliminating the ambiguity of the depth scale of one or more objects in the scene depicted in the image. For example, a first object may be associated with a first depth scale, and a second object may be associated with a second depth scale that is different from the first depth scale. Such problems may be referred to as scale ambiguity or depth scale ambiguity for monocular depth estimation. Figure 2200 is an illustration of an example of such depth scale ambiguity when determining depth from a single image 202. As shown, image 202 depicts a scene including object 204. However, a neural network trained to predict the depth of object 204 may have difficulty determining the exact scale of the scene (and, therefore, object 204). For example, a representation of object 204 in image 202 may be associated with a first scale 206 having a first distance from the plane of image 202, may be associated with a second scale 208 having a second distance from the plane of image 202, and / or may be associated with a third scale 210 having a third distance from the plane of image 202. Figure 2 In the example of , depth scale ambiguity may occur based on an image that includes the same representation of a larger object at a greater distance from an imaging plane (e.g., the plane of image 202) and a smaller object at a closer distance from the image plane. A neural network trained to predict the depth of object 204 may have difficulty determining the scale associated with object 204 and its surroundings (e.g., the scene of image 202).

[0037] There is a need for systems and techniques that can be used to determine depth scale information associated with depth estimates. There is also a need for systems and techniques that can be used to determine depth scale information associated with a monocular depth estimate (e.g., depth scale information associated with a depth estimate based on a single image or frame) and / or that can be used to disambiguate a depth scale associated with a monocular depth estimate.

[0038] Described herein are systems, apparatus, processes (also referred to as methods), and computer-readable media (collectively, "systems and techniques") for performing depth estimation to predict the depth of an image based on an image and a sparse depth input using a machine learning system (e.g., a neural network system or model). For example, the systems and techniques described herein can generate a feature representation of the image and the depth information using depth information including one or more sparse depth inputs. The generated feature representation can eliminate the ambiguity of a depth scale associated with a depth output (e.g., a depth estimate) corresponding to the image. The depth output corresponding to the image can be a monocular depth estimate.

[0039] In some examples, the sparse depth input may include sparse seed points associated with depth information. For example, the sparse depth input may be included in a sparse depth map including multiple locations that correspond to corresponding multiple locations in an input image associated with a depth estimate (e.g., a monocular depth estimate). The sparse depth map may include values ​​representing corresponding depths of corresponding pixels at corresponding locations in the input image, or may include zero values ​​indicating that corresponding pixels at corresponding locations in the input image lack sparse depth information. In some cases, systems and techniques may use the sparse depth input (e.g., included in a sparse depth map) to scale and / or rescale one or more depth predictions associated with the input image. For example, the sparse depth map may be used to rescale one or more depth predictions corresponding to a depth output of the input image.

[0040] In one illustrative example, the sparse depth map may be provided as input to an encoder of a machine learning network. For example, the sparse depth map may be provided as input to an encoder of a neural network model for generating a depth output corresponding to an image. Based on the sparse depth map being provided as input to the encoder of the neural network, scale information (e.g., associated with the sparse depth map) may be propagated to an output of the neural network (e.g., may be propagated to a depth output corresponding to the image) without requiring subsequent rescaling of depth predictions to be performed based on the sparse depth input.

[0041] In some aspects, a sparse depth input (e.g., included in a sparse depth map) may be fused with an input image. For example, the sparse depth input may be fused with one or more local features determined for the input image. In some cases, the sparse depth map and the input image may be provided as inputs to the same encoder of the neural network. In some aspects, the encoder of the neural network may include a validity channel, which may also be referred to as a sparse validity channel. The input to the encoder of the neural network may include an input image of a scene and sparse depth information associated with the input image (e.g., a sparse depth map). In some examples, the input to the encoder of the neural network may additionally include validity information, such as a validity mask associated with the sparse depth map. In some aspects, the input to the encoder of the neural network may include an input image and a sparse depth map, and the validity mask associated with the sparse depth map may be generated based on validity information included in or indicated by the input.

[0042] Aspects of the systems and techniques described herein will be described with reference to the accompanying figures.

[0043] Figure 3308 is a diagram illustrating an example of a neural network system 300 configured to generate a depth output 308 based on image and depth inputs. For example, neural network system 300 may generate depth output 308 based on one or more image and depth inputs 302. For example, image and depth inputs 302 may include one or more image inputs, one or more depth inputs, and / or a combination thereof. In some aspects, image and depth inputs 302 may also be referred to herein as “input 302”.

[0044] The neural network system 300 includes an encoder 304 and a decoder 306. The encoder 304 obtains an input 302, which includes an image and a depth input. In some cases, the input 302 may include additional inputs, such as a validity map (e.g., as described in detail below). The encoder 304 may include one or more convolutional layers configured to determine features representing the input data. For example, the encoder 304 may generate one or more feature representations of the image and depth information of the input 302. In some cases, the encoder 304 may include other additional layers, such as one or more normalization layers, activation layers (e.g., one or more rectified linear units (ReLU)), pooling layers (e.g., one or more maximum pooling layers), and / or other layers. In an illustrative example, the encoder 304 may be a ResNet encoder (e.g., a ResNet-34 encoder, a ResNet-18 encoder, etc.).

[0045] The features output by the encoder 304 (e.g., image and depth features representing the input 302) may be processed by the decoder 306 to generate or predict a depth output 308. For example, the decoder 306 may use the features generated by the encoder 304 to generate or predict a depth output for the image included in the input 302. The depth output 308 may include one or more depth estimates or depth predictions for corresponding positions in the image included in the input 302. For example, the depth output 308 may be a depth map including depth estimates or predictions for some (or all) pixel positions in the image included in the input 302. In some aspects, the decoder 306 may include one or more deconvolution layers and in some cases other additional layers configured to generate output features (e.g., one or more non-linear activation functions, such as rectified linear units (ReLU), one or more pooling layers, etc.). The architecture of the decoder 306 may regress the output to a final value of the predicted depth output 308 (e.g., using a regression layer, which in some cases may calculate a metric for the regression task of predicting the depth output 308, such as a semi-mean squared error loss). In some cases, the decoder 306 may include one or more final layers (e.g., fully connected layers, softmax layers, multi-layer perceptrons (MLPs), and / or other layers) configured to generate or predict a depth output 308 for an image from the input 302. For example, the final layer may include a softmax layer that may perform a softmax function to transform the output of the fully connected layer (e.g., represented as a vector of K elements) into a probability distribution (e.g., also represented as a vector of K elements that sum to 0). In one illustrative example, the encoder 302 and the decoder 304 may have a ResNet-UNet architecture (e.g., where the encoder 302 is a ResNet-34 encoder, a ResNet-18 encoder, or other type of ResNet encoder). In other aspects, the decoder 306 may include one or more transformer models or other neural network architectures.

[0046] The depth information of input 302 includes a sparse depth input. For example, the sparse depth input can be based on a seed point, and the neural network system 300 can use the seed point to eliminate the blur of the depth scale of the predicted depth of the image. Figure 44 is an example of an image 402 with seed points projected onto the image plane of the image 402. As shown, the seed points include a seed point 404 for a point where the floor meets the wall, a seed point 406 for a corner point on a table, and so on. Each corresponding seed point associated with an image (e.g., such as a seed point associated with the image 402) can provide depth information for a point within the image. The seed points are sparse because a seed point is not provided for every point in the scene. However, each seed point provides a very accurate depth estimate. For example, the seed point 404 may include or indicate depth information for a point where the floor meets the wall (e.g., within the image 402), the seed point 406 may include or indicate depth information for a point representing a corner point on a table (e.g., within the image 402), and so on.

[0047] The seed points may be provided from a system configured to generate or determine 6-degrees-of-freedom (6-DOF) data representing the position and depth of one or more objects depicted in an image. In one example, the system may be a 6-DOF tracker. In some cases, 6-DOF data may be obtained when performing simultaneous localization and mapping (SLAM). The 6DOF data may include three-dimensional rotation data (e.g., including pitch, roll, and yaw) and three-dimensional translation data (e.g., horizontal displacement, vertical displacement, and depth displacement relative to a reference point). In some examples, the three-dimensional rotation data may be about a camera used to capture the image (e.g., including the pitch, roll, and yaw of the camera when the image was captured). In some aspects, the seed points may include only the three-dimensional translation data of each point (e.g., horizontal displacement, vertical displacement, and depth displacement of each point). In other aspects, the seed points may include three-dimensional translation data and three-dimensional rotation data.

[0048] The neural network system 300 may use the depth information to generate or predict a depth output 308 for an image from the input 302. In some aspects, the neural network system 300 may use the depth information based on one or more seed points included in the image and depth input 302 to generate or predict a depth output 308 for an image from the input 302. For example, the seed points may be projected (e.g., based on a re-projection operation) to a two-dimensional image plane of the image of the input 302 to generate a depth map. In one illustrative example, sparse seed points (e.g., providing depth information for some but not all points in a scene of the image from the input 302) may be projected to a 2D image plane of the image of the input 302 to generate a sparse depth map.

[0049] Using depth information based on the above-described seed points as an input to a neural network model of the neural network system 300 may improve the accuracy of the neural network system 300. For example, depth information associated with a seed point included in the input 302 may be used to propagate scale information through the neural network model of the neural network system 300. In some aspects, using depth information based on the seed point of the input 302 may also eliminate the need to perform rescaling (e.g., in accordance with scale ambiguity of monocular depth estimation as described above). For example, rescaling may be eliminated based on the neural network system 300 propagating scale information (e.g., associated with the seed point of the input 302) to the depth output 308. The use of seed points included in the input 302 may also enable the neural network model of the neural network system 300 to be implemented using a less complex and / or computationally less expensive network architecture (e.g., such as a transformer), which would otherwise be effective without sparse seed points.

[0050] The image of input 302 may be a grayscale image or a color image. The image may be considered a "dense" image because each corresponding position (or pixel) in the image includes one or more values ​​indicating color information of the corresponding position or pixel (e.g., for a color image, a red (R) value, a green (G) value, and a blue (B) value for each pixel, for a grayscale image, a grayscale value for each pixel, etc.). Figure 5A is instantiated and can be provided as input to Figure 3 300 . For example, color channels 502, 504, 506 may be associated with an RGB or other color image included in input 302 of neural network system 300. Each pixel (or position) of color channels 502, 504, and 506 corresponds to a color value for that pixel (or position). For example, a pixel from an image may include a triplet of values ​​between 0 and 255 (e.g., represented in 8 bits) indicating the color of the pixel in an RGB color system, such as (0, 0, 0) corresponding to black and (255, 255, 255) corresponding to white. Each pixel value may be normalized to a value between 0 and 1 to obtain Figure 5A The values ​​of each pixel (or position) of color channels 502, 504, and 506 are shown.

[0051] Figure 5B is instantiated and can be provided as input to Figure 3 300 . For example, depth map 508 may be included in input 302 of neural network system 300. In some examples, depth map 508 may be a sparse depth map. Each position or location in depth map 508 corresponds to a pixel or position in an image of input 302 (e.g., Figure 5A506). For example, position 509 in depth map 508 corresponds to pixel or position 505 (e.g., the top left position) in color channel 504, and also corresponds to the top left pixel or position in color channels 502 and 506. As shown, depth map 508 is sparse in that it only includes depth information (or depth data) for certain pixels or positions (e.g., color channels 502, 504, and 506) of the image. For example, depth map 508 includes depth information (value 0.5, value 1.4, and value 2.3) for three corresponding pixels (or positions) in color channels 502, 504, and 506. In some examples, one or more positions or locations in a sparse depth map (e.g., such as depth map 508) may be empty or otherwise include no depth information. In some examples, positions or locations in a sparse depth map (e.g., such as depth map 508) that do not include depth information may have a value of 0.

[0052] As previously described, in some cases, input 302 may include validity data. For example, input 302 may include a validity map indicating locations or positions in the sparse depth map where depth information is included and / or indicating locations or positions in the sparse depth map where depth information is not included. Figure 5C is an example that may be provided as part of input 302 Figure 3 508 . Validity data (e.g., validity map 512) may be used to provide an additional validity channel for input 302 and may be used (e.g., by neural network system 300) to distinguish valid depth information from invalid (or missing) values ​​in sparse depth map 508. For example, a value of 0 in sparse depth map 508 may represent valid depth information (e.g., indicating a depth of 0) or may represent a missing value (e.g., indicating that no depth information is provided for a corresponding location or position in sparse depth map 508). In one illustrative example, validity map 512 may be used to mitigate the tendency of neural network system 300 to interpret 0 (e.g., a depth of 0) in depth information as missing. Figure 5B 0) shown in depth map 508 is interpreted as zero depth rather than a missing value.

[0053] Each position or location in the effectiveness map 512 corresponds to a position in the depth input (e.g., Figure 5A508). For example, position 513 in the validity map 512 corresponds to position 509 (the upper left position) in the depth map 508. As shown, the validity map 512 includes a value of 0 or a value of 1 for each position in the validity map 512. A value of 0 is generated for a position in the validity map 512 that corresponds to a position in the depth map 508 that does not include depth information. For example, based on position 513 in the depth map 508 not including depth information (e.g., a depth value of 0), position 509 in the validity map 512 includes a value of 0. A value of 1 is generated for a position in the validity map 512 that corresponds to a position in the depth map 508 that includes depth information. For example, based on position 511 in the depth map 508 including depth information (e.g., a depth value of 0.5), position 515 in the validity map 512 includes a value of 1.

[0054] In some cases, the depth information may include information based on information from a depth map (e.g., Figure 5B The depth map 508 may include inverse depth information of the inverse of the depth value of the depth map 508 . Figure 5D is a diagram illustrating an example of a sparse inverse depth map 528 that can be generated based on the sparse depth map 518. Each position in the inverse sparse depth map 528 has a value that is the inverse of the corresponding value in the corresponding position of the sparse depth map 518. For example, a value of 0.5 at position 519 of the sparse depth map 518 results in a value of 2.0 at the corresponding position 529 of the inverse sparse depth map 528. In some examples, for numerical reasons, the sparse inverse depth map 528 can be provided to the sparse inverse depth map 528 as part of the input 302. Figure 3 The neural network system 300 may be used instead of a sparse depth map (e.g., sparse depth map 508). For example, the neural network system 300 may predict the inverse of the actual depth value, in which case using the inverse of the depth information from the depth map (e.g., inverse depth information such as sparse inverse depth map 528) may yield more accurate results.

[0055] In one illustrative example, Figure 3 The input 302 of the neural network system 300 may include an RGB image having three color channels (e.g., a color channel 502 which may be a red (R) channel, a color channel 504 which may be a green (G) channel, and a color channel 506 which may be a blue (B) channel), a sparse depth map based on sparse seed points (e.g., Figure 5B sparse depth map 508) and validity map (e.g., Figure 5C The values ​​of the different inputs may be combined (e.g., concatenated) and then input to the neural network system 300.

[0056] In such examples, input 302 has a certain shape or size (B, 5, H, W), where B refers to the batch size (e.g., the number of images in a batch of images for each training iteration or epoch, for which parameters such as weights of neural network system 300 are updated during training); "5" refers to the number of channels of input 302 (e.g., three channels for 3 color channels of the input image, one channel for the depth map, and one channel for the validity map); H refers to the height of the input (e.g., they are all the same resolution); and W refers to the width of the input. In some aspects, the height parameter and the width parameter (e.g., H and W, respectively) can be given in pixels.

[0057] The depth output 308 of the neural network system 300 includes a certain shape or size (B, 1, H, W), where a single channel "1" corresponds to a depth output 308 including a single channel having the same size (H×W) as the RGB image, sparse depth map, and validity map of the input 302. For example, the depth output 308 can be a dense depth map having an estimated depth value for each position of the sparse depth map included in the input 302 (e.g., having an estimated depth value for each position of the image included in the input 302).

[0058] RGB images can be represented as . Given an RGB image , with camera intrinsic parameters and posture , the system including the neural network system 300 may Given sparse 3D seed points Based on the projection of N sparse 3D seed points, the system can generate or obtain the projected sparse inverse depth and validity mask. and :

[0059] = = K

[0060] =

[0061] and zero at other locations

[0062] and zero at other locations

[0063] here, Represents N sparse 3D points p(j) The projection of the jth point in is the RGB image I included in the input 302 i The projection point in the image plane. represents the sparse inverse depth points, and Represents a sparse validity mask.

[0064] In one illustrative example, an input to a neural network model of neural network system 300 (e.g., input 302) may be determined as , where @ represents a channel-by-channel cascade operation. For example, input 302 may be an RGB input image I i With sparse validity mask V i The inverse sparse depth map D i Channel-by-channel cascading.

[0065] Based on the channel-by-channel concatenation of input 302, the systems and techniques described herein can ensure that The local features and corresponding sparse depth values ​​in Alignment, where Indicates whether the sparse depth value is valid. The neural network system 300 can therefore ensure that mutually beneficial representations (e.g., local appearance and corresponding depth values ​​from input 302) are aligned, fused, and processed together at the highest resolution and as early as possible.

[0066] In some cases, the neural network system 300 can be trained using self-supervised learning. For example, instead of using ground truth depth information (e.g., as labels and / or annotations associated with training data inputs) to calculate a loss and perform back propagation through the network for training (e.g., as performed for supervised learning techniques), the systems and techniques described herein can perform training based on relative pose information and depth information. For example, the relative pose between two frames (e.g., two frames of training data) may be known. The two frames may be a source frame (e.g., an RGB or grayscale source image) and a target frame (e.g., a target image). Depth information (e.g., one or more depth estimates) may also be known for the source frame. In an illustrative example, using the relative pose and depth information of the source image, the source image may be warped into a target image. The image may then be rendered in a target view, and the rendered view (in the target view) may be compared to an image observed at the target view (e.g., compared to the target frame) to determine a loss. Back propagation may then be performed based on the loss.

[0067] Figure 3The neural network system 300 provides benefits over other types of machine learning systems that predict image depth (e.g., by performing monocular depth estimation). For example, in one example neural network system, separate encoders may be used to process image data and sparse depth inputs, such as a first encoder for processing image data and a second encoder (which may be referred to as a sparse encoder) for processing sparse depth inputs. A shared decoder may receive the output of the first encoder and the output of the second encoder and may determine a predicted depth. However, such models assume direct supervision (e.g., supervised learning), whereas Figure 3 The neural network is trained in a fully self-supervised manner, as described above. In addition, the sparse depth input will be processed by an additional encoder (e.g., a sparse encoder) that reduces the spatial resolution of the depth-based features before fusing them in a shared decoder. Such neural network model architectures have several disadvantages, including separation from RGB, where the model is less able to associate local depth information to pixels in an image (e.g., a color image or a grayscale image) because the two inputs are processed in separate independent encoder modules. Compared to dual encoder approaches that fuse image and depth values ​​at a downsampled resolution only after the encoder, the neural network 300 ensures that mutually beneficial representations (local appearance and corresponding depth values) are aligned, fused, and processed together at the highest resolution and as early as possible.

[0068] The dual encoder neural network model architecture is also better than Figure 3 The neural network system 300 is more expensive because it has higher computational and memory requirements due to the added parameters (e.g., weights, biases, etc.) in the sparse encoder. There is also a lack of validity specifications in such neural network model architectures. For example, sparse depth information will be densely processed by the same set of network weights. This may degrade learning because numerically 0 does not indicate zero depth but rather an invalid / missing value, while non-zero values ​​indicate pixel depth. The validity depth map described herein can address such issues.

[0069] use Figure 3 The architecture of the neural network system 300 and the Figure 4 and FIG. 5A to FIG. 5D The neural network system 300 can process and learn the mutual information between local image features and depth information (e.g., based on the seed point) based on the input information described above. In addition, by processing the image and depth information earlier in the encoder 302 (at the input), the neural network system 300 ensures that high spatial resolution is maintained when fusing the depth information with the image (e.g., RGB) information. In addition, this does not confuse the locations where sparse point values ​​apply.

[0070] Such solutions also require minimal computational requirements because lightweight neural networks can be used due to the simple nature of the architecture required to generate the features used to generate the depth output 308. For example, the architecture of the present invention is more lightweight than the dual encoder neural network model architecture described above. For example, as described above, the same encoder 304 is used to process images (e.g., color images or grayscale images) and auxiliary sparse inputs (e.g., sparse depth information), which solves the locality problem and reduces computational and memory requirements. In addition, as described above, a validity channel (e.g., a validity map) can be provided as an input to the neural network system 300. Adding an additional validity channel can alleviate the problem of interpreting 0 as zero depth instead of a missing value.

[0071] A system or device including or in communication with neural network system 300 may perform one or more operations based on depth output 308. For example, the system or device may use depth output 308 to generate or render virtual content with an input image, perform detection and / or recognition (e.g., scene or object detection and / or recognition, facial detection and / or recognition, etc.), perform pose estimation, perform image reconstruction, perform three-dimensional (3D) modeling, etc. Highly accurate depth predictions of depth output 308 may improve such operations, such as improving the quality of the reconstructed 3D mesh (e.g., less noise for walls and flat surfaces, more orthogonal angles, clearer object boundaries, etc.).

[0072] Figure 6 600 is a flow chart illustrating an example process 600 for generating depth information from one or more images using one or more of the techniques described herein. Process 600 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., a chipset, a processor such as a neural processing unit (NPU), a digital signal processor (DSP), etc.) utilizing or implementing a neural network model (e.g., Figure 3 The neural network system 300) is used to execute.

[0073] At block 602, a computing device (or a component thereof) may obtain an image of a scene (e.g., as Figure 3 At block 604, the computing device (or a component thereof) may obtain depth information associated with one or more objects in the scene (e.g., as a Figure 3 302 shown in FIG. 30). The image may include a plurality of pixels having a resolution. In some aspects, the depth information may include a sparse depth map (e.g., Figure 5B 508 ), the sparse depth map comprising a plurality of locations having a resolution, wherein each location in a first subset of the plurality of locations in the sparse depth map comprises a value representing a corresponding depth of a corresponding pixel having a corresponding location in the image (e.g., Figure 5B), and each position in a second subset of positions of the plurality of positions in the sparse depth map includes a zero value corresponding to a lack of depth information for a corresponding pixel having a corresponding position in the image (e.g., corresponding upper left position 505 in color channel 504 and upper left positions in color channels 502 and 506). Figure 5B Additionally or alternatively, the depth information may include a sparse depth map (e.g., Figure 5D A sparse inverse depth map 528 is provided, the sparse depth map comprising a plurality of positions having a resolution, wherein each position in a first subset of the plurality of positions in the sparse depth map comprises an inverse of a value (e.g., value 519 in the sparse depth map 518) representing a corresponding depth of a corresponding pixel having the corresponding position in the image (e.g., value 2.0 at position 529 of the inverse sparse depth map 528), and each position in a second subset of the plurality of positions in the sparse depth map comprises a zero value corresponding to a lack of depth information for the corresponding pixel having the corresponding position in the image.

[0074] In some aspects, a computing device (or component thereof) may obtain a validity map having a resolution (e.g., Figure 5C Each position in the validity map includes a first value (eg, a value of 1, such as Figure 5C 515 of the validity map 512) or a second value indicating whether the corresponding position in the sparse depth map includes a zero value (e.g., a value of 0, such as Figure 5C In such aspects, the computing device (or a component thereof) may process the image, the depth information, and the validity map using an encoder of the neural network model to generate a feature representation. In such aspects, the feature representation represents the image, the depth information, and the validity map.

[0075] At block 606, the computing device (or a component thereof) may use an encoder (e.g., Figure 3 The encoder 304 of the neural network system 300 processes the image and depth information to generate feature representations of the image and depth information. For example, the computing device (or a component thereof) may generate a channel-by-channel concatenation of the image, the depth information, and the validity map. Based on generating the validity map, the encoder of the neural network model (e.g., Figure 3The encoder 304 of the neural network system 300 of the embodiment of the present invention can be used to generate feature representations of the image and depth information based on the channel-by-channel concatenation. In some cases, the channel-by-channel concatenation may include one or more color channels of the image, a depth channel associated with the sparse depth map, and a validity channel associated with the validity map. The encoder of the neural network model uses the feature representation generated by the channel-by-channel concatenation (e.g., Figure 3 The feature representation generated by the encoder 304 may include multiple features having a resolution (e.g., the same resolution as the channel-by-channel concatenation and / or the same resolution as the image, the sparse depth map, and / or the validity map).

[0076] For example, the input to the encoder of a neural network model can be Figure 3 302 of the same or similar. In some cases, the input to the encoder of the neural network model may include a neural network having three color channels (e.g., such as Figure 5A The red (R) channel 502, Figure 5A The green (G) channel 504 and Figure 5A The input of the encoder of the neural network model may include a sparse depth map based on one or more sparse seed points. For example, the input of the encoder of the neural network model may include Figure 5B 508 of the sparse depth map. In some cases, the input to the encoder of the neural network model may include a validity map, such as Figure 5C 512 of the validity map. As described above, the computing device (or a component thereof) may generate a channel-by-channel cascade of the image information, the depth information, and the validity map. For example, the values ​​of the RGB color channel inputs, the values ​​of the sparse depth map (and / or the sparse depth map seed points), and / or the values ​​of the validity map may be combined in a channel-by-channel cascade to generate an input to an encoder of the neural network model.

[0077] In some cases, a channel-by-channel concatenation of the image, depth information, and validity map can be generated as , where @ indicates a channel-by-channel cascade operation. For example, the encoder of a neural network model (e.g., Figure 3 The input of the encoder 304 can be an RGB input image I i With sparse validity mask V i The inverse sparse depth map D i In some examples, the channel-by-channel cascade can be used to ensure that local features in the image (e.g., RGB input image I i ) and D i The corresponding sparse depth values ​​in are aligned, where the validity map V i Indicates whether the corresponding sparse depth value is valid or invalid.

[0078] At block 608, the computing device (or a component thereof) may use a decoder (e.g., Figure 3 The decoder 306 of the neural network system 300 processes the feature representation of the image and the depth information to generate a depth output corresponding to the image (e.g., Figure 3 The depth output 308 may include a depth output of the image. For example, the depth output may include a depth map having a resolution. Each position in the depth map may include a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image (e.g., the top left position in the depth map includes a depth value indicating the depth of the top left pixel in the image). In some aspects, the computing device (or components thereof) may process the depth output to generate a 3-dimensional mesh of the scene depicted in the image or to perform one or more other operations or functions.

[0079] In some aspects, a computing device (or a component thereof) may obtain a plurality of seed points. For example, each of the plurality of seed points may indicate a corresponding position and depth of a corresponding point in a scene. In such aspects, a computing device (or a component thereof) may generate depth information based on the plurality of seed points. For example, to generate depth information based on the plurality of seed points, the computing device (or a component thereof) may project the plurality of seed points onto a two-dimensional image plane associated with the image.

[0080] As described above, the processes described herein (e.g., process 600 and / or any other process described herein) can be implemented by utilizing or implementing a neural network model (e.g., Figure 3 In one example, process 600 may be performed by electronic device 100 of FIG. 1 . In another example, process 600 may be performed by a computing device or apparatus having a neural network system 300 of FIG. 1 . Fig. 9 The use or implementation of a neural network model (e.g., Figure 3 The neural network system 300) is executed by a computing system of a computing device architecture of a computing system 900. For example, Fig. 9 The computing device of the computing device architecture of the computing system 900 shown may implement Figure 6 The operation and / or this article about Figures 3 to 5D Components and / or operations described in any of the figures (e.g., Figure 3 Neural network system 300).

[0081] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, an XR device (e.g., a VR headset, an AR headset, AR glasses, etc.), a wearable device (e.g., a connected watch or smart watch, or other wearable device), a server computer, a vehicle (e.g., an autonomous vehicle) or a computing device of a vehicle, a robotic device, a laptop, a smart TV, a camera, and / or any other computing device having resource capabilities to perform the processes described herein (including process 600 and / or other processes described herein). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on an Internet Protocol (IP) or other types of data.

[0082] The components of the computing device may be implemented in circuits. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0083] Process 600 is illustrated as a logic flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally speaking, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.

[0084] Additionally, process 600 and / or any other process described herein may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, through hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program that includes a plurality of instructions that may be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0085] As stated in this article, Figure 3 The neural network system 300 may be implemented using one neural network or multiple neural networks. Figure 7 It can be Figure 3 illustrative example of a deep learning neural network 700 used by the neural network system 300 of FIG. 700. An input layer 720 includes input data. In an illustrative example, the input layer 720 may include data representing pixels of an input video frame. The neural network 700 includes a plurality of hidden layers 722a, 722b to 722n. The hidden layers 722a, 722b to 722n include "n" hidden layers, where "n" is an integer greater than or equal to one. The plurality of hidden layers may include as many layers as required for a given application. The neural network 700 also includes an output layer 724 that provides outputs resulting from processing performed by the hidden layers 722a, 722b to 722n. In an illustrative example, the output layer 724 may provide a classification of an object in an input video frame. The classification may include a category that identifies a type of object (e.g., a person, dog, cat, or other object).

[0086] Neural network 700 is a multi-layer neural network composed of interconnected nodes. Each node can represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains the information as it processes it. In some cases, neural network 700 may include a feedforward network, in which case there is no feedback connection in which the output of the network is fed back into itself. In some cases, neural network 700 may include a recursive neural network, which may have loops that allow information to be carried across nodes when reading in input.

[0087] Information can be exchanged between nodes by interconnecting nodes to nodes between each layer. The nodes of input layer 720 can activate the set of nodes in the first hidden layer 722a. For example, as shown in the figure, each input node in the input node of input layer 720 is connected to each node in the node of the first hidden layer 722a. The nodes of hidden layers 722a, 722b to 722n can transform the information by applying activation function to the information of each input node. The information derived from the transformation can then be transferred to the node of the next hidden layer 722b and can activate the node of the next hidden layer, and the node of the next hidden layer can perform their own specified functions. Example functions include convolution, upsampling, data conversion and / or any other suitable function. Then, the output of hidden layer 722b can activate the node of the next hidden layer, and so on. The output of last hidden layer 722n can activate one or more nodes of output layer 724, and output is provided at the one or more nodes. In some cases, although a node in neural network 700 (eg, node 726 ) is shown as having multiple output lines, the node has a single output, and all lines shown as outputting from the node represent the same output value.

[0088] In some cases, each node or interconnection between nodes may have a weight, which is a set of parameters derived from the training of neural network 700. Once neural network 700 is trained, it may be referred to as a trained neural network, which may be used to classify one or more objects. For example, an interconnection between nodes may represent a piece of information learned about the interconnected nodes. The interconnection may have a tunable digital weight that may be tuned (e.g., based on a training data set), allowing neural network 700 to adapt to the input and be able to learn as more and more data is processed.

[0089] The neural network 700 is pre-trained to process features from the data in the input layer 720 using different hidden layers 722a, 722b to 722n to provide an output through the output layer 724. In examples where the neural network 700 is used to identify objects in images, the neural network 700 may be trained using training data that includes both images and labels. For example, training images may be input into the network, where each training image has a label indicating the class of one or more objects in each image (essentially, indicating to the network what the objects are and what features they have). In one illustrative example, the training images may include an image of the number 2, in which case the label for the image may be [0010000000].

[0090] In some cases, the neural network 700 may use a training process known as back propagation to adjust the weights of the nodes. Back propagation may include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. This process may be repeated for a certain number of iterations for each set of training images until the neural network 700 is trained well enough so that the weights of each layer are accurately tuned.

[0091] For the example of identifying an object in an image, a forward pass may include passing a training image through the neural network 700. Prior to training the neural network 700, the weights are initially randomized. The image may include, for example, an array of numbers representing pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 digital array having 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or a brightness and two chrominance components, etc.).

[0092] For the first training iteration of the neural network 700, the output will likely include values ​​that do not favor any particular class due to the random selection of weights at initialization. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each of the different classes may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 700 is unable to determine low-level features and therefore cannot make an accurate determination of what the classification of the object may be. A loss function may be used to analyze the errors in the output. Any suitable loss function definition may be used. An example of a loss function includes the mean squared error (MSE). The MSE is defined as , which computes the sum of the ground truth outputs (e.g., actual answers) minus one-half the square of the predicted outputs (e.g., predicted answers). The loss can be set equal to The value of .

[0093] For the first training images, the loss (or error) will be high because the actual values ​​will be very different from the predicted outputs. The goal of training is to minimize the amount of loss so that the predicted outputs are the same as the training labels. The neural network 700 can perform a backward pass by determining which inputs (weights) contribute most to the network's loss, and the weights can be adjusted so that the loss is reduced and ultimately minimized.

[0094] The derivative of the loss with respect to the weight (expressed as dL / dW, where W is the weight at a particular layer) can be calculated to determine the weight that contributes most to the loss of the network. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be expressed as , where w represents the weight, w i Represents the initial weight, and η represents the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes a larger weight update, while a lower value indicates a smaller weight update.

[0095] In some cases, neural network 700 may use self-supervised learning to train the neural network, as described above with respect to Figure 3 The neural network system described above.

[0096] Neural network 700 may include any suitable deep network. One example includes a convolutional neural network (CNN) that includes an input layer and an output layer with multiple hidden layers between the input layer and the output layer. Figure 8 An example of a CNN is described. The hidden layer of the CNN includes a series of convolutional layers, nonlinear layers, pooling layers (for downsampling), and fully connected layers. The neural network 700 may include any other deep network other than a CNN, such as an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), etc.

[0097] Figure 8 is an illustrative example of a convolutional neural network 800 (CNN 800). An input layer 820 of CNN 800 includes data representing an image. For example, the data may include an array of numbers representing pixels of the image, where each number in the array includes a value from 0 to 255 describing the intensity of the pixel at the location of the number in the array. Using the previous example from above, the array may include a 28×28×3 array of numbers having 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chrominance components, etc.). The image may be passed through a convolutional hidden layer 822a, an optional non-linear activation layer, a pooling hidden layer 822b, and a fully connected hidden layer 822c to obtain an output at an output layer 824. Although Figure 8 Only one hidden layer among the hidden layers is shown in , but one of ordinary skill will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in CNN 800. As previously described, the output may indicate a single category of an object, or may include a probability of a category that best describes an object in the image.

[0098] The first layer of CNN 800 is a convolutional hidden layer 822a. The convolutional hidden layer 822a analyzes the image data of the input layer 820. Each node of the convolutional hidden layer 822a is connected to an area of ​​nodes (pixels) of the input image called a receptive field. The convolutional hidden layer 822a can be considered as one or more filters (each filter corresponds to a different activation or feature map), where each convolution iteration of the filter is a node or neuron of the convolutional hidden layer 822a. For example, the area of ​​the input image covered by the filter at each convolution iteration will be the receptive field of the filter. In an illustrative example, if the input image includes a 28×28 array and each filter (and corresponding receptive field) is a 5×5 array, there will be 24×24 nodes in the convolutional hidden layer 822a. Each connection between a node and the receptive field of that node learns a weight, and in some cases learns an overall bias, so that each node learns to analyze its specific local receptive field in the input image. Each node of the hidden layer 822a will have the same weights and biases (referred to as shared weights and shared biases). For example, the filter has an array of weights (numbers) and the same depth as the input. For the video frame example, the filter would have a depth of 3 (according to the three color components of the input image). An illustrative example size of the filter array is 5×5×3, corresponding to the size of the node's receptive field.

[0099] The convolutional nature of the convolutional hidden layer 822a is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filter of the convolutional hidden layer 822a may start at the upper left corner of the input image array and may be convolved around the input image. As noted above, each convolution iteration of the filter may be considered as a node or neuron of the convolutional hidden layer 822a. In each convolution iteration, the value of the filter is multiplied by the original pixel values ​​of the corresponding number of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​at the upper left corner of the input image array). The multiplications from each convolution iteration may be added together to obtain the sum for that iteration or node. Next, the process is continued at the next position in the input image according to the receptive field of the next node in the convolutional hidden layer 822a.

[0100] For example, the filter may move a step amount to the next receptive field. The step amount may be set to 1 or other suitable amount. For example, if the step amount is set to 1, the filter will move 1 pixel to the right at each convolution iteration. Processing the filter at each unique position in the input volume produces a number representing the filter result at that position, thereby producing a sum value determined for each node of the convolution hidden layer 822a.

[0101] The mapping from the input layer to the convolutional hidden layer 822a is called an activation map (or feature map). The activation map includes a value for each node that represents the result of the filter at each location of the input volume. The activation map may include an array that includes various summed values ​​produced by each iteration of the filter over the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map will include a 24×24 array. The convolutional hidden layer 822a may include several activation maps in order to identify multiple features in the image. Figure 8 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 822a can detect three different types of features, each of which is detectable across the entire image.

[0102] In some examples, a nonlinear hidden layer may be applied after the convolutional hidden layer 822a. A nonlinear layer may be used to introduce nonlinearity into a system that already computes linear operations. An illustrative example of a nonlinear layer is a rectified linear unit (ReLU) layer. The ReLU layer may apply a function f(x)=max(0, x) to all values ​​in the input volume, which changes all negative activations to 0. Thus, the ReLU may increase the nonlinear properties of the CNN 800 without affecting the receptive field of the convolutional hidden layer 822a.

[0103] A pooling hidden layer 822b may be applied after the convolutional hidden layer 822a (and after the nonlinear hidden layer when used). The pooling hidden layer 822b is used to simplify the information in the output from the convolutional hidden layer 822a. For example, the pooling hidden layer 822b may take each activation map output from the convolutional hidden layer 822a and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by a pooling hidden layer. The pooling hidden layer 822a uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. A pooling function (e.g., a maximum pooling filter, an L2 norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 822a. In Figure 8 In the example shown, three pooling filters are used for the three activation maps in the convolutional hidden layer 822a.

[0104] In some examples, max pooling may be used by applying a max pooling filter (e.g., having a size of 2×2) with a stride (e.g., equal to the dimension of the filter, such as a stride of 2) to the activation map output from the convolutional hidden layer 822a. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer may summarize a region of 2×2 nodes in the previous layer (each node is a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by a 2×2 max pooling filter at each iteration of the filter, where the maximum of the four values ​​is output as the “maximum” value. If such a max pooling filter is applied to an activation filter having a dimension of 24×24 nodes from the convolutional hidden layer 822a, the output from the pooling hidden layer 822b will be an array of 12×12 nodes.

[0105] In some examples, an L2-norm pooling filter may also be used. The L2-norm pooling filter involves computing the square root of the sum of the squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (rather than computing the maximum value as done in max pooling), and using the computed value as output.

[0106] Intuitively, a pooling function (e.g., max pooling, L2 norm pooling, or other pooling functions) determines whether a given feature is found anywhere in a region of an image, and discards the exact location information. This can be done without affecting the results of feature detection, because once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max pooling (and other pooling methods) provides the following benefits: there are far fewer pooled features, thereby reducing the number of parameters required in the later layers of CNN 800.

[0107] The last layer of connections in the network is a fully connected layer that connects each node from the pooling hidden layer 822b to each output node in the output layer 824. Using the above example, the input layer includes 28×28 nodes that encode the pixel intensities of the input image, the convolutional hidden layer 822a includes 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for the filter) to the three activation maps, and the pooling layer 822b includes a layer of 3×12×12 hidden feature nodes based on applying a max pooling filter to a 2×2 region on each of the three feature maps. Extending this example, the output layer 824 may include ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 822b is connected to each node of the output layer 824.

[0108] The fully connected layer 822c can obtain the output of the previous pooling layer 822b (which should represent an activation map of high-level features) and determine the features that are most relevant to a particular category. For example, the fully connected layer 822c layer can determine the high-level features that are most strongly associated with a particular category and can include weights (nodes) for the high-level features. The product between the weights of the fully connected layer 822c and the pooling hidden layer 822b can be calculated to obtain probabilities for different categories. For example, if the CNN 800 is being used to predict that an object in a video frame is a person, there will be high values ​​in the activation map representing high-level features of a person (e.g., there are two legs, a face at the top of the object, two eyes at the upper left and upper right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0109] In some examples, the output from the output layer 824 may include an M-dimensional vector (in the previous example, M=10), where M may include the number of categories that the program must choose from when classifying an object in an image. Other example outputs may also be provided. Each number in the N-dimensional vector may represent the probability that an object belongs to a certain category. In an illustrative example, if the 10-dimensional output vector representing ten different categories of objects is [000.050.800.150000], then the vector indicates that there is a 5% probability that the image is an object of the third category (e.g., a dog), an 80% probability that the image is an object of the fourth category (e.g., a person), and a 15% probability that the image is an object of the sixth category (e.g., a kangaroo). The probability for a category can be considered as a confidence level that the object is part of that category.

[0110] Fig. 9 is a diagram illustrating an example of a system for implementing certain aspects of the present disclosure. Specifically, Fig. 9 An example of a computing system 900 is illustrated, which may be, for example, any computing device constituting a computing system, a camera system, or any component thereof, wherein the components of the system communicate with each other using a connection 905. The connection 905 may be a physical connection using a bus, or a direct connection into the processor 910, such as in a chipset architecture. The connection 905 may also be a virtual connection, a networked connection, or a logical connection.

[0111] In some examples, computing system 900 is a distributed system, where the functionality described in the present disclosure may be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some examples, one or more of the described system components represent a number of such components, each performing some or all of the functionality for which the component is described. In some examples, each component may be a physical or virtual device.

[0112] The example system 900 includes at least one processing unit (CPU or processor) 910 and connections 905 that couple various system components including system memory 915, such as read only memory (ROM) 920 and random access memory (RAM) 925, to the processor 910. The computing system 900 may include a cache 912 of high-speed memory directly connected to the processor 910, in close proximity to the processor, or integrated as part of the processor.

[0113] Processor 910 may include any general purpose processor and hardware or software services, such as services 932, 934, and 936 stored in storage device 930, that are configured to control processor 910 as well as a dedicated processor where software instructions are incorporated into the actual processor design. Processor 910 may essentially be a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0114] To enable user interaction, the computing system 900 includes an input device 945 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The computing system 900 may also include an output device 935, which may be one or more of a plurality of output mechanisms. In some cases, a multimodal system may enable a user to provide multiple types of input / output to communicate with the computing system 900. The computing system 900 may include a communication interface 940, which may generally govern and manage user input and system output.

[0115] The communication interface can perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including utilizing an audio jack / plug, a microphone jack / plug, a Universal Serial Bus (USB) port / plug, an Apple ® Lightning ® Ports / plugs, Ethernet ports / plugs, Fiber optic ports / plugs, Dedicated wired ports / plugs, Bluetooth ® Wireless signal transmission, Bluetooth ® Low energy (BLE) wireless signal transmission, IBEACON ®Wireless signaling, radio frequency identification (RFID) wireless signaling, near field communication (NFC) wireless signaling, dedicated short range communication (DSRC) wireless signaling, 802.11 Wi-Fi wireless signaling, wireless local area network (WLAN) signaling, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), infrared (IR) communication wireless signaling, public switched telephone network (PSTN) signaling, integrated services digital network (ISDN) signaling, 3G / 4G / 5G / LTE cellular data network wireless signaling, ad hoc network signaling, radio wave signaling, microwave signaling, infrared signaling, visible light signaling, ultraviolet light signaling, wireless signaling along the electromagnetic spectrum, or some combination thereof.

[0116] The communication interface 940 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 900 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the Global Positioning System (GPS) of the United States, the Global Navigation Satellite System (GLONASS) of Russia, the Beidou Navigation Satellite System (BDS) of China, and the Galileo GNSS of Europe. There is no limitation to operating on any particular hardware arrangement, and thus the basic features herein may be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0117] The storage device 930 may be a non-volatile and / or non-transitory and / or computer-readable memory device, and may be a hard disk or other type of computer-readable medium that can store data that can be accessed by a computer, such as a magnetic tape cartridge, a flash memory card, a solid-state memory device, a digital versatile disk, a cassette, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage medium, a flash memory, a memristor memory, any other solid-state memory, a compact disk-read only memory (CD-ROM) optical disk, a rewritable compact disk (CD) optical disk, a digital video disk (DVD) optical disk, a Blu-ray disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick ®card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin-transfer torque RAM (STT-RAM), another memory chip or box, and / or a combination thereof.

[0118] The storage device 930 may include software services, servers, services, etc., which, when the code defining such software is executed by the processor 910, causes the system to perform a function. In some examples, a hardware service that performs a particular function may include a software component stored in a computer-readable medium connected to the necessary hardware components (such as the processor 910, the connection 905, the output device 935, etc.) to perform the function. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data may be stored and does not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer readable media may store thereon code and / or machine executable instructions, which may represent a procedure, function, subprogram, program, routine, subroutine, module, software package, category, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by transmitting and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be transmitted, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network sending, etc.

[0119] In some examples, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media specifically exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.

[0120] Specific details are provided in the above description to provide a detailed understanding of the aspects and examples provided herein. However, it will be understood by those skilled in the art that these aspects and examples can also be practiced without these specific details. For the sake of clarity of explanation, in some instances, the present technology may be presented as including a separate functional block, which includes a device, device component, step or routine in a method embodied in a combination of software or hardware and software. Additional components other than those components shown in the accompanying drawings and / or described herein may be used. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid confusing various aspects and examples in unnecessary details. In other instances, known circuits, processes, algorithms, structures and techniques may be shown without unnecessary details to avoid confusing various aspects and examples.

[0121] Various aspects and examples may be described above as a process or method, which is depicted as a flow chart, a flowchart, a data flow diagram, a structure diagram, or a block diagram. Although a flow chart may describe an operation as a sequential process, many operations in the operation may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. The process is terminated when the operation of the process is completed, but the process may have additional steps not included in the accompanying drawings. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.

[0122] The processes and methods according to the above examples can be implemented using stored computer executable instructions or otherwise obtained from computer readable media. Such instructions may include, for example, instructions and data that configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or function group. Parts of the computer resources used can be accessed through a network. Computer executable instructions can be, for example, binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer readable media that can be used to store instructions, information used, and / or information created during the method according to the described examples include disks or optical disks, flash memories, USB devices with non-volatile memory, networked storage devices, etc.

[0123] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. By way of additional examples, such functionality may also be implemented on circuit boards among different chips or different processes executed on a single device.

[0124] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0125] In the foregoing description, various aspects of the present application are described with reference to the specific examples of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Therefore, although the illustrative aspects and examples of the present application have been described in detail herein, it should be understood that the inventive concept can be implemented and adopted in various other ways, and the appended claims are intended to be interpreted as including these variations, unless limited by the prior art. Various features and aspects of the above-mentioned applications can be used individually or in combination. In addition, without departing from the broader spirit and scope of this specification, various aspects and examples can be utilized in any number of environments and applications beyond those described herein. Therefore, the description and the accompanying drawings should be considered as illustrative rather than restrictive. For the purpose of illustration, each method is described in a specific order. It should be understood that in alternative aspects and examples, the method can be performed in an order different from the described order.

[0126] It should be understood by those skilled in the art that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to (" ”) and greater than or equal to (“ ”) symbol instead.

[0127] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0128] The phrase "coupled to" means that any component is directly or indirectly physically connected to another component, and / or any component is directly or indirectly in communication with another component (eg, connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0129] Claim language or other language in this disclosure that states "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language that states "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language that states "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language that states "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0130] Various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, frames, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints proposed for the entire system. The technician may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as departing from the scope of the present application.

[0131] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices, or integrated circuit devices with multiple uses, including applications in wireless communication devices, mobile phones, and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the technology may be implemented at least in part by a computer-readable data storage medium including a program code, which includes instructions for executing one or more of the above methods, algorithms, and / or operations when executed. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as a synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0132] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.

[0133] Exemplary aspects of the present disclosure include:

[0134] Aspect 1. A device for generating depth information based on one or more images, the device comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory, the at least one processor being configured to: obtain an image of a scene; obtain depth information associated with one or more objects in the scene; process the image and the depth information using an encoder of a neural network model to generate a feature representation of the image and the depth information; and process the image and the feature representation of the depth information using a decoder of the neural network model to generate a depth output corresponding to the image.

[0135] Aspect 2. The apparatus according to aspect 1, wherein the image comprises a plurality of pixels having a resolution.

[0136] Aspect 3. An apparatus according to Aspect 2, wherein the depth information comprises a sparse depth map, the sparse depth map comprising a plurality of positions having the resolution, and wherein: each position in a first subset of positions of the plurality of positions in the sparse depth map comprises a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image; and each position in a second subset of positions of the plurality of positions in the sparse depth map comprises a zero value corresponding to a lack of depth information for a corresponding pixel having a corresponding position in the image.

[0137] Aspect 4. An apparatus according to any one of Aspects 2 to 3, wherein the depth information comprises a sparse depth map, the sparse depth map comprising a plurality of positions having the resolution, and wherein: each position in a first subset of positions of the plurality of positions in the sparse depth map comprises an inverse of a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image; and each position in a second subset of positions of the plurality of positions in the sparse depth map comprises a zero value corresponding to a lack of depth information for a corresponding pixel having a corresponding position in the image.

[0138] Aspect 5. An apparatus according to any one of Aspects 3 to 4, wherein the at least one processor is configured to: obtain a validity map having the resolution, each of the multiple positions in the validity map including a corresponding first value indicating that the corresponding position in the sparse depth map includes a valid depth value or including a corresponding second value indicating that the corresponding position in the sparse depth map includes a zero value; and process the image, the depth information and the validity map using the encoder of the neural network model to generate the feature representation, wherein the feature representation represents the image, the depth information and the validity map.

[0139] Aspect 6. An apparatus according to Aspect 5, wherein the at least one processor is configured to: generate a channel-by-channel cascade of the image, the depth information, and the validity map; and process the channel-by-channel cascade using the encoder of the neural network model to generate the feature representation.

[0140] Aspect 7. An apparatus according to Aspect 6, wherein the channel-by-channel cascade includes one or more color channels of the image, a depth channel associated with the sparse depth map, and a validity channel associated with the validity map; and the feature representation includes multiple features having the resolution.

[0141] Aspect 8. An apparatus according to any one of Aspects 2 to 7, wherein the depth output comprises a depth map having the resolution, each position in the depth map comprising a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image.

[0142] Aspect 9. An apparatus according to any one of aspects 1 to 8, wherein the at least one processor is configured to: process the depth output to generate the scene 3-dimensional mesh.

[0143] Aspect 10. An apparatus according to any one of Aspects 1 to 9, wherein the at least one processor is configured to: obtain a plurality of seed points, each of the plurality of seed points indicating a corresponding position and depth of a corresponding point in the scene; and generate the depth information based on the plurality of seed points.

[0144] Aspect 11. The apparatus according to aspect 10, wherein, in order to generate the depth information based on the plurality of seed points, the at least one processor is configured to: project the plurality of seed points to a two-dimensional image plane associated with the image.

[0145] Aspect 12. A method for generating depth information based on one or more images, the method comprising: obtaining an image of a scene; obtaining depth information associated with one or more objects in the scene; processing the image and the depth information using an encoder of a neural network model to generate a feature representation of the image and the depth information; and processing the image and the feature representation of the depth information using a decoder of the neural network model to generate a depth output corresponding to the image.

[0146] Clause 13. The method according to clause 12, wherein the image comprises a plurality of pixels having a resolution.

[0147] Aspect 14. A method according to Aspect 13, wherein the depth information includes a sparse depth map, the sparse depth map includes multiple positions having the resolution, and wherein: each position in a first subset of positions of the multiple positions in the sparse depth map includes a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image; and each position in a second subset of positions of the multiple positions in the sparse depth map includes a zero value corresponding to a lack of depth information for a corresponding pixel having a corresponding position in the image.

[0148] Aspect 15. A method according to any one of Aspects 13 to 14, wherein the depth information includes a sparse depth map, the sparse depth map includes multiple positions having the resolution, and wherein: each position in a first subset of positions of the multiple positions in the sparse depth map includes an inverse of a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image; and each position in a second subset of positions of the multiple positions in the sparse depth map includes a zero value corresponding to a lack of depth information for a corresponding pixel having a corresponding position in the image.

[0149] Aspect 16. The method according to any one of Aspects 14 to 15, further comprising: obtaining a validity map having the resolution, each of a plurality of positions in the validity map comprising a corresponding first value indicating that the corresponding position in the sparse depth map comprises a valid depth value or comprising a corresponding second value indicating that the corresponding position in the sparse depth map comprises a zero value; and processing the image, the depth information and the validity map using the encoder of the neural network model to generate the feature representation, wherein the feature representation represents the image, the depth information and the validity map.

[0150] Aspect 17. The method according to Aspect 16 further comprises: generating a channel-by-channel cascade of the image, the depth information, and the validity map; and processing the channel-by-channel cascade using the encoder of the neural network model to generate the feature representation.

[0151] Aspect 18. A method according to Aspect 17, wherein the channel-by-channel cascade includes one or more color channels of the image, a depth channel associated with the sparse depth map, and a validity channel associated with the validity map; and the feature representation includes multiple features having the resolution.

[0152] Aspect 19. A method according to any one of aspects 13 to 18, wherein the depth output comprises a depth map having the resolution, each position in the depth map comprising a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image.

[0153] Aspect 20. The method according to any one of aspects 12 to 19, further comprising: processing the depth output to generate a 3-dimensional mesh of the scene.

[0154] Aspect 21. According to any one of Aspects 12 to 20, the method further includes: obtaining a plurality of seed points, each of the plurality of seed points indicating a corresponding position and depth of a corresponding point in the scene; and generating the depth information based on the plurality of seed points.

[0155] Aspect 22. The method according to aspect 21, wherein generating the depth information based on the plurality of seed points comprises: projecting the plurality of seed points to a two-dimensional image plane associated with the image.

[0156] Aspect 23. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform operations including: obtaining an image of a scene; obtaining depth information associated with one or more objects in the scene; processing the image and the depth information using an encoder of a neural network model to generate a feature representation of the image and the depth information; and processing the image and the feature representation of the depth information using a decoder of the neural network model to generate a depth output corresponding to the image.

[0157] Clause 24. The non-transitory computer-readable medium of Clause 23, wherein the image comprises a plurality of pixels having a resolution.

[0158] Aspect 25. A non-transitory computer-readable medium according to Aspect 24, wherein the depth information comprises a sparse depth map, the sparse depth map comprising a plurality of positions having the resolution, and wherein: each position in a first subset of positions of the plurality of positions in the sparse depth map comprises a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image; and each position in a second subset of positions of the plurality of positions in the sparse depth map comprises a zero value corresponding to a lack of depth information for a corresponding pixel having a corresponding position in the image.

[0159] Aspect 26. A non-transitory computer-readable medium according to Aspect 25, wherein the sparse depth map is an inverse sparse depth map, and wherein each position in the first subset of positions comprises an inverse of the corresponding depth of the corresponding pixel having the corresponding position in the image.

[0160] Aspect 27. A non-transitory computer-readable medium according to any one of Aspects 25 to 26, wherein the instructions further cause the one or more processors to perform operations including: obtaining a validity map having the resolution, each of a plurality of positions in the validity map including a corresponding first value indicating that the corresponding position in the sparse depth map includes a valid depth value or including a corresponding second value indicating that the corresponding position in the sparse depth map includes a zero value; and processing the image, the depth information, and the validity map using the encoder of the neural network model to generate the feature representation, wherein the feature representation represents the image, the depth information, and the validity map.

[0161] Aspect 28. A non-transitory computer-readable medium according to Aspect 27, wherein the instructions further cause the one or more processors to perform operations including: generating a channel-by-channel cascade of the image, the depth information, and the validity map; and processing the channel-by-channel cascade using the encoder of the neural network model to generate the feature representation.

[0162] Aspect 29. A non-transitory computer-readable medium according to Aspect 28, wherein the channel-by-channel cascade includes one or more color channels of the image, a depth channel associated with the sparse depth map, and a validity channel associated with the validity map; and the feature representation includes a plurality of features having the resolution.

[0163] Aspect 30. A non-transitory computer-readable medium according to any one of Aspects 24 to 29, wherein the depth output comprises a depth map having the resolution, each position in the depth map comprising a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image.

[0164] Aspect 31. A non-transitory computer-readable medium according to any one of aspects 23 to 30, wherein the instructions further cause the one or more processors to perform operations comprising: processing the depth map to generate a 3-dimensional mesh of the scene.

[0165] Aspect 32. A non-transitory computer-readable medium according to any one of Aspects 23 to 31, wherein the instructions further cause the one or more processors to perform operations including: obtaining a plurality of seed points, each of the plurality of seed points indicating a corresponding position and depth of a corresponding point in the scene; and generating the depth information based on the plurality of seed points.

[0166] Aspect 33. A non-transitory computer-readable medium according to Aspect 32, wherein in order to generate the depth information based on the multiple seed points, the instructions further cause the one or more processors to perform operations including: projecting the multiple seed points to a two-dimensional image plane associated with the image.

[0167] Aspect 34. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform the operations of any one of Aspects 1 to 11.

[0168] Aspect 35. A non-transitory computer-readable medium having instructions stored thereon, which when executed by one or more processors cause the one or more processors to perform operations according to any one of aspects 12 to 22.

[0169] Aspect 36. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform operations according to any one of Aspects 23 to 33.

[0170] Aspect 37. An apparatus comprising one or more components for performing the operations according to any one of aspects 1 to 11.

[0171] Aspect 38. An apparatus comprising one or more components for performing the operations according to any one of aspects 12 to 22.

[0172] Aspect 39. An apparatus comprising one or more components for performing the operations according to any one of aspects 23 to 33.

Claims

1. An apparatus for generating depth information based on one or more images, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory, the at least one processor configured to: obtaining an image of the scene; obtaining depth information associated with one or more objects in the scene; Processing the image and the depth information using an encoder of a neural network model to generate feature representations of the image and the depth information; as well as The image and the feature representation of the depth information are processed using a decoder of the neural network model to generate a depth output corresponding to the image.

2. The apparatus of claim 1, wherein the image comprises a plurality of pixels having a resolution.

3. The apparatus of claim 2, wherein the depth information comprises a sparse depth map comprising a plurality of locations having the resolution, and wherein: Each position in a first subset of the plurality of positions in the sparse depth map comprises a value representing a respective depth of a respective pixel having a corresponding position in the image; and Each position in a second subset of positions of the plurality of positions in the sparse depth map includes a zero value corresponding to a lack of depth information for a corresponding pixel having the corresponding position in the image.

4. The apparatus of claim 2, wherein the depth information comprises a sparse depth map comprising a plurality of locations having the resolution, and wherein: Each position in a first subset of the plurality of positions in the sparse depth map comprises an inverse of a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image; and Each position in a second subset of positions of the plurality of positions in the sparse depth map includes a zero value corresponding to a lack of depth information for a corresponding pixel having the corresponding position in the image.

5. The apparatus of claim 3, wherein the at least one processor is configured to: obtaining a significance map having the resolution, each of a plurality of positions in the significance map comprising a respective first value indicating that the corresponding position in the sparse depth map comprises a valid depth value or comprising a respective second value indicating that the corresponding position in the sparse depth map comprises a zero value; and The encoder using the neural network model processes the image, the depth information, and the significance map to generate the feature representation, wherein the feature representation represents the image, the depth information, and the significance map.

6. The apparatus of claim 5, wherein the at least one processor is configured to: generating a channel-by-channel concatenation of the image, the depth information, and the significance map; and The channel-by-channel cascade is processed using the encoder of the neural network model to generate the feature representation.

7. The device according to claim 6, wherein: The channel-by-channel cascade includes one or more color channels of the image, a depth channel associated with the sparse depth map, and a significance channel associated with the significance map; and The feature representation includes a plurality of features having the resolution.

8. The apparatus of claim 2, wherein the depth output comprises a depth map having the resolution, each position in the depth map comprising a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image.

9. The apparatus of claim 1 , wherein the at least one processor is configured to: The depth output is processed to generate a 3-dimensional mesh of the scene.

10. The apparatus of claim 1, wherein the at least one processor is configured to: Obtaining a plurality of seed points, each seed point of the plurality of seed points indicating a corresponding position and depth of a corresponding point in the scene; and The depth information is generated based on the plurality of seed points.

11. The apparatus according to claim 10, wherein to generate the depth information based on the plurality of seed points, the at least one processor is configured to: The plurality of seed points are projected to a two-dimensional image plane associated with the image.

12. A method for generating depth information based on one or more images, the method comprising: obtaining an image of the scene; obtaining depth information associated with one or more objects in the scene; Processing the image and the depth information using an encoder of a neural network model to generate feature representations of the image and the depth information; as well as The image and the feature representation of the depth information are processed using a decoder of the neural network model to generate a depth output corresponding to the image. The method of claim 12 , wherein the image comprises a plurality of pixels having a resolution.

14. The method of claim 13, wherein the depth information comprises a sparse depth map comprising a plurality of locations having the resolution, and wherein: Each position in a first subset of the plurality of positions in the sparse depth map comprises a value representing a respective depth of a respective pixel having a corresponding position in the image; and Each position in a second subset of positions of the plurality of positions in the sparse depth map includes a zero value corresponding to a lack of depth information for a corresponding pixel having the corresponding position in the image.

15. The method of claim 13, wherein the depth information comprises a sparse depth map comprising a plurality of locations having the resolution, and wherein: Each position in a first subset of the plurality of positions in the sparse depth map comprises an inverse of a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image; and Each position in a second subset of positions of the plurality of positions in the sparse depth map includes a zero value corresponding to a lack of depth information for a corresponding pixel having the corresponding position in the image.

16. The method according to claim 14, further comprising: obtaining a significance map having the resolution, each of a plurality of positions in the significance map comprising a respective first value indicating that the corresponding position in the sparse depth map comprises a valid depth value or comprising a respective second value indicating that the corresponding position in the sparse depth map comprises a zero value; and The encoder using the neural network model processes the image, the depth information, and the significance map to generate the feature representation, wherein the feature representation represents the image, the depth information, and the significance map.

17. The method according to claim 16, further comprising: generating a channel-by-channel concatenation of the image, the depth information, and the significance map; as well as The channel-by-channel cascade is processed using the encoder of the neural network model to generate the feature representation.

18. The method of claim 17, wherein: The channel-by-channel cascade includes one or more color channels of the image, a depth channel associated with the sparse depth map, and a significance channel associated with the significance map; and The feature representation includes a plurality of features having the resolution.

19. The method of claim 13, wherein the depth output comprises a depth map having the resolution, each position in the depth map comprising a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image.

20. The method according to claim 12, further comprising: The depth output is processed to generate a 3-dimensional mesh of the scene.

21. The method according to claim 12, further comprising: Obtaining a plurality of seed points, each of the plurality of seed points indicating a corresponding position and depth of a corresponding point in the scene; as well as The depth information is generated based on the plurality of seed points.

22. The method of claim 21 , wherein generating the depth information based on the plurality of seed points comprises: The plurality of seed points are projected to a two-dimensional image plane associated with the image.

23. A non-transitory computer readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform operations comprising: obtaining an image of the scene; obtaining depth information associated with one or more objects in the scene; Processing the image and the depth information using an encoder of a neural network model to generate feature representations of the image and the depth information; as well as The image and the feature representation of the depth information are processed using a decoder of the neural network model to generate a depth output corresponding to the image.

24. The non-transitory computer readable medium of claim 23, wherein the image comprises a plurality of pixels having a resolution.

25. The non-transitory computer readable medium of claim 24, wherein the depth information comprises a sparse depth map comprising a plurality of locations having the resolution, and wherein: Each position in a first subset of the plurality of positions in the sparse depth map comprises a value representing a respective depth of a respective pixel having a corresponding position in the image; and Each position in a second subset of positions of the plurality of positions in the sparse depth map includes a zero value corresponding to a lack of depth information for a corresponding pixel having the corresponding position in the image.

26. The non-transitory computer-readable medium of claim 25, wherein the sparse depth map is an inverse sparse depth map, and wherein each position in the first subset of positions comprises an inverse of the corresponding depth of the corresponding pixel having the corresponding position in the image.

27. The non-transitory computer readable medium of claim 25, wherein the instructions further cause the one or more processors to perform operations comprising: obtaining a significance map having the resolution, each of a plurality of positions in the significance map comprising a respective first value indicating that the corresponding position in the sparse depth map comprises a valid depth value or comprising a respective second value indicating that the corresponding position in the sparse depth map comprises a zero value; and The encoder using the neural network model processes the image, the depth information, and the significance map to generate the feature representation, wherein the feature representation represents the image, the depth information, and the significance map.

28. The non-transitory computer readable medium of claim 27, wherein the instructions further cause the one or more processors to perform operations comprising: generating a channel-by-channel concatenation of the image, the depth information, and the significance map, wherein the channel-by-channel concatenation includes one or more color channels of the image, a depth channel associated with the sparse depth map, and a significance channel associated with the significance map; and The channel-by-channel cascade is processed using the encoder of the neural network model to generate the feature representation, wherein the feature representation includes a plurality of features having the resolution.

29. The non-transitory computer readable medium of claim 24, wherein: The depth output comprises a depth map having the resolution, each position in the depth map comprising a value representing a corresponding depth of a corresponding pixel having a corresponding position in the image; and The instructions further cause the one or more processors to perform operations including processing the depth map to generate a 3-dimensional mesh of the scene.

30. The non-transitory computer readable medium of claim 23, wherein the instructions further cause the one or more processors to perform operations comprising: Obtaining a plurality of seed points, each seed point of the plurality of seed points indicating a corresponding position and depth of a corresponding point in the scene; and The depth information is generated based on the plurality of seed points.