Monocular depth supervision from 3D bounding boxes
By combining 3D object detection information and ground-based true depth information during the deep training phase, the depth estimation network is improved, solving the problem of insufficient depth estimation for dynamic objects. This enhances the autonomous navigation and 3D object detection capabilities of the autonomous agent and reduces its reliance on LiDAR sensors.
Patent Information
- Application Number
- CN202180044325.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-23
- Filing Date
- 2021-06-23
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2041-06-23
AI Technical Summary
Existing depth estimation systems lack sufficient accuracy in depth estimation for dynamic objects, especially in autonomous agents, which affects the accuracy of downstream tasks such as 3D object detection.
By incorporating 3D object detection information as part of the training loss during the deep training phase, and combining 3D bounding box information with ground truth depth information, the training method of the depth estimation network is improved, thereby enhancing the accuracy of depth estimation for dynamic objects.
It improves the accuracy of depth estimation for dynamic objects, enhances the autonomous agent's autonomous navigation and 3D object detection capabilities in complex environments, reduces reliance on LiDAR sensors, and lowers costs.
Smart Images

Figure CN115867940B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Certain aspects of the present disclosure generally relate to depth estimation systems. BACKGROUND
[0002] Autonomous agents (e.g., vehicles, robots, etc.) rely on depth estimation to perform various tasks. These various tasks can include constructing a three-dimensional (3D) representation of a surrounding environment or identifying 3D objects. The 3D representation can be used for various tasks, such as localization and / or autonomous navigation. Improving the accuracy of depth estimation can improve the accuracy of downstream tasks, such as producing a 3D representation or 3D object detection. It is desirable to improve the accuracy of depth estimation obtained from images captured by sensors of an autonomous agent. SUMMARY
[0003] In one aspect of the disclosure, a method is disclosed. The method includes capturing a two-dimensional (2D) image of an environment adjacent to a ego vehicle. The environment includes at least a dynamic object and a static object. The method also includes producing, via a depth estimation network, a depth map of the environment based on the 2D image. An accuracy of a depth estimation for the dynamic object in the depth map is greater than an accuracy of a depth estimation for the static object in the depth map. The method further includes producing a three-dimensional (3D) estimate of the environment based on the depth map. The method still further includes controlling an action of the ego vehicle based on the identified location.
[0004] In another aspect of the disclosure, a non-transitory computer-readable medium having non-transitory program code recorded thereon is disclosed. The program code is executed by a processor and includes program code to capture a 2D image of an environment adjacent to a ego vehicle, the environment including at least a dynamic object and a static object. The program code also includes program code to produce, via a depth estimation network, a depth map of the environment based on the 2D image. An accuracy of a depth estimation for the dynamic object in the depth map is greater than an accuracy of a depth estimation for the static object in the depth map. The program code further includes program code to produce a 3D estimate of the environment based on the depth map. The program code still further includes program code to control an action of the ego vehicle based on the identified location.
[0005] Another aspect of the disclosure relates to a device. The device has a memory, one or more processors coupled to the memory, and instructions stored in the memory. The instructions, when executed by the processors, are operable to cause the device to capture a 2D image of an environment adjacent to a ego vehicle, the environment including at least a dynamic object and a static object. The instructions further cause the device to generate, via a depth estimation network, a depth map of the environment based on the 2D image. An accuracy of a depth estimation for the dynamic object in the depth map is greater than an accuracy of a depth estimation for the static object in the depth map. The instructions additionally cause the device to generate a 3D estimate of the environment based on the depth map. The instructions further cause the device to control an action of the ego vehicle based on the identified location.
[0006] This summary of the feature and technical advantages of the present disclosure has been presented quite broadly in order to provide those skilled in the art with a general understanding of the features and advantages of the present disclosure. Additional features and advantages of the present disclosure will be described in the detailed description that follows. It should be appreciated that the present disclosure can be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. It should also be realized that such equivalent constructions do not depart from the scope of the present disclosure as set forth in the appended claims. The novel features of the present disclosure, both as to its organization and manner of operation, together with further objects and advantages thereof, will be better understood from the following description when considered in connection with the accompanying drawings. It is explicitly BRIEF DESCRIPTION OF DRAWINGS
[0007] The features, nature, and advantages of the present disclosure will become more apparent from the detailed description set forth below when considered in conjunction with the drawings in which like reference characters identify correspondingly throughout.
[0008] Figure 1 An example of a vehicle in an environment in accordance with aspects of the present disclosure is illustrated.
[0009] Figure 2A An example of a single image in accordance with aspects of the present disclosure.
[0010] Figure 2B An example of a depth map in accordance with aspects of the present disclosure.
[0011] Figure 2C An example of a reconstructed target image in accordance with aspects of the present disclosure.
[0012] Figure 3 An example of a depth network in accordance with aspects of the present disclosure is illustrated.
[0013] Figure 4An example of a pose network is illustrated in accordance with aspects of the present disclosure.
[0014] Figure 5 An example of a training pipeline is illustrated in accordance with aspects of the present disclosure.
[0015] Figure 6 FIG. 1 is an example diagram illustrating a hardware implementation for a depth estimation system in accordance with aspects of the present disclosure.
[0016] Figure 7 A flow diagram of a method is illustrated in accordance with aspects of the present disclosure. DETAILED DESCRIPTION
[0017] The detailed description set forth below intends as an description of various configurations and is not intended to represent the only configurations in which the concepts described herein can be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form in order to avoid obscuring the concepts being presented.
[0018] Agents, such as autonomous agents, can perform various tasks based on depth estimation. For example, an agent can produce a 3D representation of a scene based on images obtained from sensors. The 3D representation can also be referred to as a 3D model, 3D scene, or 3D map. The 3D representation can facilitate various tasks, such as scene understanding, motion planning, and / or obstacle avoidance. For example, an agent can navigate through an environment autonomously based on the 3D representation. Additionally, or alternatively, an agent can identify 3D objects based on depth estimation.
[0019] Artificial neural networks, such as deep networks, can be trained to estimate depth from sensor measurements. Unlike improving downstream tasks, such as 3D object detection, conventional systems for depth training emphasize depth performance. Depth training refers to a training phase in which a deep network is trained to estimate depth from images. Aspects of the present disclosure relate to improving depth estimation for downstream tasks by incorporating 3D object detection information as part of the training loss for depth training.
[0020] Figure 1 An example of a ego vehicle 100 (e.g., ego agent) in an environment 150 is illustrated in accordance with aspects of the present disclosure. As Figure 1As shown, ego vehicle 100 is traveling on road 110. First vehicle 104 (e.g., other agent) can be in front of ego vehicle 100, and second vehicle 116 can be adjacent to ego vehicle 100. In this example, ego vehicle 100 can include 2D camera 108 (such as a 2D RGB camera) and second sensor 106. Second sensor 106 can be another RGB camera or another type of sensor (such as RADAR and / or ultrasound). Additionally, or alternatively, ego vehicle 100 can include one or more additional sensors. For example, the additional sensors can be side-facing and / or rear-facing sensors.
[0021] In one configuration, 2D camera 108 captures a 2D image that includes objects in the field of view 114 of 2D camera 108. Second sensor 106 can produce one or more output streams. The 2D image captured by 2D camera includes a 2D image of first vehicle 104 because first vehicle 104 is in the field of view 114 of 2D camera 108.
[0022] Information obtained from second sensor 106 and 2D camera 108 can be used to navigate ego vehicle 100 along a route when ego vehicle 100 is in an autonomous mode. Second sensor 106 and 2D camera 108 can be powered by power provided from a battery (not shown) of vehicle 100. The battery can also power an engine of the vehicle. Information obtained from second sensor 106 and 2D camera 108 can be used to produce a 3D representation of the environment.
[0023] Aspects of the present disclosure improve depth estimation for objects in an environment. The improved depth estimation can improve downstream tasks, such as 3D object detection. A downstream task can refer to a task performed based on the depth estimation. In some implementations, ground-truth points corresponding to an object can be selected based on 3D bounding box information. A weight can be increased for each ground-truth point (e.g., pixel) corresponding to the object. As the weight is increased, a depth network improves the depth estimation for the object.
[0024] The improved depth estimation can refer to depth estimation with improved accuracy. The objects can include objects that are not salient in an input image, such as cars and pedestrians. Improving the accuracy of the depth estimation for objects that are not salient in an input image can be at the expense of the accuracy of the depth estimation for representative objects that are salient in the input image, such as roads and buildings. Improving the accuracy of the depth estimation for objects that are not salient in an input image improves the model for downstream tasks, such as 3D object detection, because the identification of these objects can be improved.
[0025] Supervised monocular depth networks learn an estimation function by regressing an input image to an estimated depth output. Supervised training refers to learning from labeled ground truth information. For example, a conventional supervised monocular depth network can use ground truth depth (e.g., LIDAR data) to train a neural network as a regression model. In supervised depth networks, a convolutional neural network produces an initial coarse estimate, and another neural network is used to refine the prediction to generate a more accurate result. As supervised techniques for depth estimation develop, the availability of target depth labels decreases due to the cost of producing labeled data. For example, labeling outdoor scenes is a time-consuming task.
[0026] Some conventional monocular depth solutions replace LiDAR information with vision-based information. That is, the depth network does not use point clouds directly from LiDAR, but estimates point clouds from multiple single images. In such examples, conventional monocular depth solutions use the estimated point clouds for 3D bounding box detection. As described, cameras (e.g., vision-based systems) are ubiquitous in most systems and are less expensive compared to LiDAR sensors. Therefore, camera-based solutions can be applied to a wider range of platforms. However, LiDAR systems outperform vision-based systems. Improving the accuracy of depth estimation can close the gap between LiDAR systems and vision-based systems. Conventional systems can also close the gap between LiDAR systems and vision-based systems by including information from sparse LiDAR sensors when training and testing. Sparse LiDAR information can correct misalignments. These conventional systems reduce the use of LiDAR information.
[0027] Due to their cost, LIDAR systems can be economically infeasible. Cameras, such as red-green-blue (RGB) cameras, can provide dense information. Additionally, cameras can be more economically feasible compared to LIDAR sensors. Aspects of the present disclosure improve monocular depth estimation produced from depth networks trained in a supervised manner. The improved monocular depth estimation closes the gap between LIDAR and vision solutions such that cameras can augment, supplement, or replace ranging sensors. In some implementations, depth training (e.g., training for image-based depth estimation) can be self-supervised by initiating intrinsic geometric constraints in a robot, or via sparse depth labels from a calibrated LiDAR sensor.
[0028] Eliminating the disparity between depth estimation from monocular cameras and depth estimation from LiDAR sensors can reduce cost and increase a robust solution because the cameras supplement the functionality of the ranging sensors. For example, cameras can perform better in some environments (such as rainy environments) compared to LIDAR sensors. Conversely, LIDAR sensors can perform better in other environments (such as low light conditions) compared to cameras. Thus, monocular depth estimation can improve the ability of an agent to perform various tasks.
[0029] Furthermore, an agent can generate a larger amount of image data compared to LIDAR data. The image data can be used as training data for a depth network. As such, the use of monocular sensors can increase the amount of training data, thereby improving self-supervised monocular depth estimation.
[0030] As described, aspects of the present disclosure improve 3D object detection from monocular images (e.g., pseudo LiDAR point clouds). 3D object detection is a component that enables autonomous agents to navigate autonomously. Currently, LiDAR information can be used for 3D object detection. It is desirable to improve 3D object detection by processing monocular point clouds, rather than LiDAR information.
[0031] Accurate depth estimation can improve autonomous navigation throughout an environment. For example, accurate depth estimation can improve collision avoidance with objects, such as cars or pedestrians. Aspects of the present disclosure are not limited to autonomous agents. Aspects of the present disclosure also contemplate agents operating in a manual mode or a semi-autonomous mode. In a manual mode, a human driver manually operates (e.g., controls) the agent. In an autonomous mode, an agent control system operates the agent without human intervention. In a semi-autonomous mode, a human can operate the agent, and the agent control system can override or assist the human. For example, the agent control system can override the human to prevent a collision or to comply with one or more traffic rules.
[0032] In some examples, a conventional system obtains a pseudo point cloud produced by a pre-trained depth network. The pre-trained depth network transforms an input image into information needed for 3D bounding box detection (e.g., a pseudo point cloud). The pre-trained depth network includes a monocular, stereo, or multi-view network trained in a semi-supervised or supervised manner. The pre-trained depth network can be a general-purpose depth network that is not trained for a particular task. The pre-trained depth network can be referred to as an off-the-shelf network.
[0033] A pre-trained deep network can learn strong priors of the environment (e.g., ground plane, vertical walls, structures). The learned priors can improve the overall depth results. However, the learned priors do not improve the depth estimation for objects related to 3D object detection. The objects related to 3D object detection are typically dynamic objects, such as cars and pedestrians. As known to those skilled in the art, dynamic objects are a difficult problem for monocular depth estimation in a semi-supervised setting. For example, the motion of dynamic objects violates the static world assumption that forms the basis of the photometric loss used to train the depth network.
[0034] As described, conventional off-the-shelf depth networks are trained using metrics and losses that are different from the downstream task, such as 3D object detection. Thus, conventional off-the-shelf depth networks can reduce the accuracy of the related task. For example, a conventional off-the-shelf depth network can accurately recover the ground plane because the ground plane covers a large portion (e.g., a large number of pixels) of the image. In contrast, a large number of pixels representing a pedestrian can be less than the number of pixels representing the ground plane (such as a road). Thus, the relevance of a pedestrian for depth estimation can be less because the goal of a monocular depth network is to maximize the accuracy of the prediction for all pixels.
[0035] Aspects of the present disclosure do not use off-the-shelf depth networks developed for depth estimation only, but rather train the depth network to improve the downstream task, such as 3D object detection. In some implementations, 3D object detection information can be incorporated as part of the training loss for depth training. The 3D object detection information can be designated for training the downstream task training. Thus, the 3D object information is already available to the depth network. Conventional systems do not use such information for depth training.
[0036] In some implementations, if 3D bounding box information is not available at training time, the depth network falls back to depth training without 3D bounding box information. As such, the training phase can learn from images with 3D bounding boxes and images without annotated 3D bounding boxes. Such implementations improve the flexibility of depth training to use different sources of information so that available labels are not discarded.
[0037] In some implementations, based on the training, the depth network performs depth estimation at test time based on an input image. That is, the depth network is trained using 3D bounding box information as well as ground truth depth information, if available. However, the depth network can only use the input image for depth estimation at test time.
[0038] Figure 2AAn example of a target image 200 of a scene 202 is illustrated in accordance with aspects of the present disclosure. The target image 200 can be captured by a monocular camera. The monocular camera can capture a forward-facing view of an agent (e.g., a vehicle). In one configuration, the monocular camera is integrated with the vehicle. For example, the monocular camera can be defined in a roof structure, a windshield, a guardrail, or other portion of the vehicle. The vehicle can have one or more cameras and / or other types of sensors. The target image 200 can also be referred to as a current image. The target image 200 captures a 2D representation of the scene.
[0039] Figure 2B An example of a depth map 220 of the scene 202 is illustrated in accordance with aspects of the present disclosure. The depth map 220 can be estimated from the target image 200 and one or more source images. The source images can be images captured at a previous time step relative to the target image 200. The depth map 220 provides a depth of the scene. The depth can be represented as a color or other characteristic.
[0040] Figure 2C An example of a 3D reconstruction 240 of the scene 202 is illustrated in accordance with aspects of the present disclosure. The 3D reconstruction can be generated from the depth map 220 and poses of the source images and the target image 200. As shown in Figure 2A and 2C The perspective of the scene 202 in the 3D reconstruction 240 is different than the perspective of the scene 202 in the target image 200. Because the 3D reconstruction 240 is a 3D view of the scene 202, the perspective can be changed as desired. The 3D reconstruction 240 can be used to control one or more actions of the agent.
[0041] Figure 3 An example of a deep network 300 is illustrated in accordance with aspects of the present disclosure. As shown in Figure 3 The deep network 300 includes an encoder 302 and a decoder 304. The deep network 300 generates a per-pixel depth map of an input image 320, such as the depth map 220 of Figure 2B .
[0042] The encoder 302 includes a plurality of encoder layers 302a-d. Each of the encoder layers 302a-d can be a packing layer to downsample features during an encoding process. The decoder 304 includes a plurality of decoder layers 304a-d. In Figure 3 each of the decoder layers 304a-d can be an unpacking layer to upsample features during a decoding process. That is, each of the decoder layers 304a-d can unpack a received feature map.
[0043] The skip connections 306 transmit activations and gradients between the encoder layers 302a-d and the decoder layers 304a-d. The skip connections 306 facilitate resolving details at higher resolutions. For example, the gradients can be propagated directly back to the layers via the skip connections 306, improving training. Additionally, the skip connections 306 transmit image details (e.g., features) directly from the convolutional layers to the deconvolutional layers, improving image recovery at higher resolutions.
[0044] The decoder layers 302a-d can produce intermediate inverse depth maps 310. Each intermediate inverse depth map 310 can be upsampled before concatenating with the feature maps unpacked from the corresponding skip connection 306 and the corresponding decoder layer 302a-d. The inverse depth maps 310 also serve as the output of the depth network from which the loss is computed. Unlike conventional systems that hyper resolve each inverse depth map 310 incrementally. Aspects of the present disclosure upsample each inverse depth map 310 to the highest resolution using bilinear interpolation. Upsampling to the highest resolution reduces copy-based artifacts and photometric ambiguities, improving depth estimation.
[0045] Figure 4 An example of a pose network 400 for ego-motion estimation according to aspects of the present disclosure is illustrated. In contrast to conventional pose networks, Figure 4 the pose network 400 does not use an explainability mask. In conventional systems, the explainability mask removes objects that do not conform to the static world assumption.
[0046] As Figure 4 shown, the pose network 400 includes a plurality of convolutional layers 402, a final convolutional layer 404, and a multi-channel (e.g., six-channel) average pooling layer 406. The final convolutional layer 404 can be a 1x1 layer. The multi-channel layer 406 can be a six-channel layer.
[0047] In one configuration, a target image (I t ) 408 and a source image (I s ) 410 are input to the pose network 400. The target image 408 and the source image 410 can be concatenated together such that the concatenated target image 408 and source image 410 are input to the pose network 400. During training, one or more source images 410 can be used during different training epochs. The source images 410 can include an image at a previous time step (t-1) and an image at a future time step (t+1). The output is a set of six degrees of freedom (DoF) transforms between the target image 408 and the source image 410. If more than one source image 410 is considered, the process can be repeated for each source image 410.
[0048] Different training methods can be used to train the monocular depth estimation network. The training methods can include, for example, supervised training, semi-supervised training, and self-supervised training. Supervised training refers to a network that regresses ground truth depth information by applying a loss, such as an LI loss (e.g., absolute error). In self-supervised training, depth information and pose information are warped to produce a reconstructed image (e.g., a 3D representation of a 2D image). Photometric loss minimizes the difference between the original image and the reconstructed image. Semi-supervised training can be a combination of self-supervised and supervised training.
[0049] As described, in some implementations, the weight of ground truth points corresponding to a first set of objects in an image is greater than the weight of ground truth points corresponding to a second set of objects in the image. The set of objects can be determined based on a desired task. For example, for 3D object detection, the first set of objects can be dynamic objects and / or objects with a lower occurrence rate in the input image. As an example, the first set of objects can include vehicles and / or pedestrians. Additionally, the second set of objects can be static objects and / or objects with a higher occurrence rate in the input image. As an example, the second set of objects can include buildings, roads, and / or sidewalks. In most cases, a series of images includes buildings, roads, and / or sidewalks that occur more often than the occurrence of people and / or vehicles.
[0050] Aspects of the disclosure add 3D bounding boxes to monocular depth estimation training data. During training, the 3D bounding boxes can be used in addition to the depth information. That is, the monocular depth network can be trained with 3D bounding boxes and ground truth depth information.
[0051] In one configuration, the weight of pixels within a 3D bounding box is adjusted. A pseudo point cloud can be generated by increasing the relevance of pixels within the 3D bounding box. Increasing the weight (e.g., relevance) of the pixels can cause the depth metric to decrease. Also, increasing the weight of the pixels improves 3D object detection.
[0052] 3D object detection can be improved by focusing on portions of an image that are considered relevant for an assigned task (e.g., 3D object detection). That is, the depth network can be trained for a 3D object detection task. Task-specific training is different from conventional systems that use a pre-trained general-purpose depth network that is not conceived or trained for a particular task.
[0053] Aspects of the disclosure can improve supervised training, self-supervised training, and semi-supervised training. For supervised training, different weights can be applied to image pixels that contain ground truth information. These pixels can be identified from an annotated depth map. For example, the depth map can be annotated with a bounding box. The weight for pixels within the bounding box can be adjusted.
[0054] In some aspects, for self-supervised training, the 3D bounding box is projected back to the input image, generating a 2D projection. Different weights can be applied to pixels that fall within the 2D reprojected bounding box. Figure 5 An example training pipeline 500 for training a depth estimation network 504 is illustrated in accordance with aspects of the present disclosure. As Figure 5 shown, a depth estimation network 300, as described in Figure 3 , can produce a depth estimate 506 from a two-dimensional input image 502. The training pipeline 500 is not limited to using a depth estimation network 300, as described in Figure 3 , other types of depth estimation neural networks can be implemented.
[0055] The depth estimation network 504 can be used by a view synthesis module 508 to produce a reconstructed image 510 (e.g., a warped source image). In some implementations, the current image 502 and the source image 504 are input to a pose network, as described in Figure 4 , to estimate the pose of a sensor (e.g., a monocular camera). The current image 502 can be an image at time step t, and the source image 502 can be an image at time step t-1. The view synthesis module 508 can produce the reconstructed image 510 based on the estimated depth and the estimated pose. The view synthesis module 508 can also be referred to as a scene reconstruction network. The view synthesis module 508 can be trained with respect to the difference between the target image 502 and the reconstructed image 510. The network can be trained to minimize a loss, such as a photometric loss 520.
[0056] The photometric loss 520 is computed based on the difference between the target image 502 and the reconstructed image 510 (e.g., a warped source image that approximates the target image). The photometric loss 520 can be used to update the depth network 300, the view synthesis module 508, the pose network 500, and / or the weights of the pixels.
[0057] The photometric loss 520 (L p ) can be determined as follows:
[0058]
[0059] where SSIM() is a function for estimating the structural similarity (SSIM) between the target image 502 and the reconstructed image 510. The SSIM can be determined as follows:
[0060] SSIM(x, y) = [l(x, y)] α [c(x, y)] β [s(x, y)] γ , (2)
[0061] where s() determines structural similarity, c() determines contrast similarity, and I() determines brightness similarity. a, b, and g are parameters used to adjust the relative importance of each component, each parameter being greater than zero.
[0062] During the testing phase, the training pipeline 500 can produce the reconstructed image 510 as described above. The photometric loss 516 can not be computed during the testing phase. The reconstructed image 510 can be used for localization and / or other vehicle navigation tasks.
[0063] For example, the view estimation module 508 can project each point (e.g., pixel) in the current image 502 to a location in the source image 506 based on the estimated depth 506 and the sensor pose. After projecting the point to the source image 504, bilinear interpolation can be used to warp the point to the warped source image 510. That is, bilinear interpolation obtains a value (e.g., RGB value) for the point in the warped source image 510 based on the source image 504.
[0064] That is, the location (e.g., x, y coordinates) of a point in the warped source image 510 can correspond to the location of that point in the target image 502. Also, the color of a point in the warped source image 510 can be based on the colors of neighboring pixels in the source image 504. The warped source image 510 can be a 3D reconstruction of the 2D target image.
[0065] In some implementations, the 3D object detection network 518 can estimate the location of the object 514 in the warped source image 510. The location of the object 514 can be annotated with the 3D bounding box 512. For illustrative purposes, Figure 5 The 3D bounding box 512 in the warped source image 510 is shown as a 2D bounding box. In one configuration, the 3D bounding box 512 is projected back to the current image 502. A 2D bounding box 516 can be generated from the projected 3D bounding box 512. Different weights can be applied to pixels that fall within the 2D bounding box. For example, the weights can be increased such that pixels that fall within the 2D bounding box have a greater contribution to the depth estimation than pixels with lower weights. In one implementation, pixels with increased weights have a greater contribution to minimizing the loss during training than pixels with decreased weights.
[0066] The processes discussed for supervised and self-supervised training are performed for semi-supervised training. In some implementations, additional weights balance the supervised training process and the self-supervised training process. In some implementations, multiple passes can be performed during the training phase, each pass can be performed with different parameter values (e.g., weights). After each pass, performance (e.g., accuracy) is measured. The weights can be ordered and adjusted (e.g., improved) accordingly based on the size of the loss.
[0067] Pixelization can be caused by the supervised loss. The self-supervised loss mitigates pixelization. Scale inaccuracies can be caused by the self-supervised loss. The supervised loss mitigates scale inaccuracies. An additional weight can reduce pixelization of the depth map that can be caused by the supervised loss. The additional weight can also reduce scale inaccuracies of the depth map that can be caused by the supervised depth error loss.
[0068] In one configuration, when determining the pixel weight adjustment, the network determines a total number of valid pixels (NVP). A valid pixel refers to a pixel that has corresponding ground truth depth information, or when using photometric loss, a valid re-projection value. For example, for supervised training, the neural network identifies pixels in the image that have depth information in the ground truth depth image.
[0069] Additionally, the network determines a number of valid pixels with a bounding box (NBP). The weight for valid pixels with a bounding box is determined based on the following pixel ratio: ((NVP-NBP) / NBP). For example, if an image includes 100,000 valid pixels, and 2000 are within a bounding box, the weight for pixels in the bounding box would be 49 (e.g., (100000-2000) / 2000). In this example, the weight outside of the bounding box is normalized to 1. Differently, the weight inside of the bounding box is determined as the described ratio (49) (e.g., (NVP-NBP) / NBP). The pixel ratio varies depending on the different images.
[0070] The pixel ratio enables learning in areas that do not belong to the bounding box, so that the structure and geometry of the scene is preserved. If a bounding box is not annotated in a particular image, the weight is zero and is not applied to any pixel. As such, aspects of the disclosure can train a depth network from images with annotated 3D bounding boxes and images without annotated bounding boxes, thereby improving the robustness of the training data.
[0071] As described, ground truth points can be obtained from the 3D bounding box information. Due to the weight adjustment, detection of the first set of objects by the neural network (e.g., a depth estimation neural network) can be improved. The improved detection can come at the expense of detection of the second set of objects. That is, reducing the detection accuracy for the second set of objects improves the detection accuracy for the first set of objects. In one aspect, a model for a downstream task of 3D object detection is improved.
[0072] Figure 6 FIG. 1 is an example diagram illustrating a hardware implementation for a depth estimation system 600 according to aspects of the present disclosure. The depth estimation network 600 can be a component of a vehicle, a robotic device, or other device. For example, as described in more detail below, the depth estimation network 600 can be a component of a vehicle, such as a vehicle 1000 of FIG. 10. Figure 6As shown, the depth estimation network 600 is a component of a vehicle 628. Aspects of the disclosure are not limited to the depth estimation network 600 being a component of a vehicle 628, as other types of agents, such as a bus, a ship, a drone, or a robot, are also contemplated for using the depth estimation network 600.
[0073] The vehicle 628 can operate in one or more of an autonomous mode of operation, a semi-autonomous mode of operation, and a manual mode of operation. Further, the vehicle 628 can be an electric vehicle, a hybrid vehicle, or other type of vehicle.
[0074] The depth estimation network 600 can be implemented with a bus architecture, which is generically represented with the bus 660. The bus 660 is comprised of any number of interconnecting buses and bridges, depending on the specific application of the depth estimation network 600 and the overall design constraints. The bus 660 links together various circuits including the processor 620, the communication module 622, the positioning module 618, the sensor module 602, the motion module 626, the navigation module 626, and the computer-readable medium 616 represented one or more processors and / or hardware modules. The bus 660 can also link various other circuits well known in the art, thus no further description will be provided as to such circuits, such as timing sources, peripherals, voltage regulators, and power management circuits.
[0075] The depth estimation network 600 includes the transceiver 616, the sensor module 602, the depth estimation module 608, the communication module 622, the positioning module 618, the motion module 626, the navigation module 626, and the computer-readable medium 616 coupled to the processor 620. The transceiver 616 is coupled to the antenna 666. The transceiver 616 communicates with various other devices over one or more communication networks, such as an infrastructure network, a V2V network, a V2I network, a V2X network, a V2P network, or other type of network.
[0076] The depth estimation network 600 includes the processor 620 coupled to the computer-readable medium 616. The processor 620 executes processes including execution of software stored on the computer-readable medium 616 to provide functionality in accordance with this disclosure. The software, when executed by the processor 620, causes the depth estimation network 600 to perform various functions described for a particular device, such as the vehicle 628, or any of the modules 602, 608, 614, 616, 618, 620, 622, 624, 626. The computer-readable medium 616 can also be used for storing data that is manipulated by the processor 620 when executing software.
[0077] The sensor module 602 can be used to obtain measurements via different sensors, such as the first sensor 606 and the second sensor 606. The first sensor 606 can be a vision sensor, such as a stereo camera or a red-green-blue (RGB) camera for capturing 2D images. The second sensor 606 can be a ranging sensor, such as a light detection and ranging (LIDAR) sensor or a radio detection and ranging (RADAR) sensor. Of course, aspects of the present disclosure are not limited to the aforementioned sensors, as other types of sensors, such as, for example, thermal, sonar, and / or laser, are also contemplated for use in any of the sensors 606, 606.
[0078] Measurements of the first sensor 606 and the second sensor 606 can be processed by the processor 620, the sensor module 602, the depth estimation module 608, the communication module 622, the localization module 618, the motion module 626, the navigation module 626, the computer-readable medium 616 to implement the functionality described herein. In one configuration, data captured by the first sensor 606 and the second sensor 606 can be transmitted to an external device via the transceiver 616. The first sensor 606 and the second sensor 606 can be coupled to the vehicle 628 or can be in communication with the vehicle 628.
[0079] The localization module 618 can be used to determine a location of the vehicle 628. For example, the localization module 618 can use a global positioning system (GPS) to determine a location of the vehicle 628. The communication module 622 can be used to facilitate communication via the transceiver 616. For example, the communication module 622 can be configured to provide communication capabilities via different wireless protocols, such as WiFi, long term evolution (LTE), 6G, etc. The communication module 622 can also be used to communicate with other components of the vehicle 628 that are not modules of the depth estimation network 600.
[0080] The motion module 626 can be used to facilitate motion of the vehicle 628. As an example, the motion module 626 can control movement of wheels. As another example, the motion module 626 can communicate with one or more power sources of the vehicle 628, such as an engine and / or a battery. Of course, aspects of the present disclosure are not limited to providing motion via wheels, but are contemplated for other types of components that provide motion, such as thrusters, pedals, fins, and / or jet engines.
[0081] The depth estimation network 600 also includes a navigation module 626 for planning a route or controlling the vehicle 628 via the motion of the motion module 626. In one configuration, the navigation module 626 employs a defensive driving mode when the depth estimation module 608 identifies a hazardous agent. The navigation module 626 can override user input when the user input is predicted to cause a collision. The module can be a software module running in the processor 620, a hardware module coupled to the processor 620, or some combination thereof, resident / stored in the computer-readable medium 616.
[0082] The depth estimation module 608 can communicate with the sensor module 602, the transceiver 616, the processor 620, the communication module 622, the localization module 618, the motion module 626, the navigation module 626, and the computer-readable medium 616. In one configuration, the depth estimation module 608 receives sensor data from the sensor module 602. The sensor module 602 can receive sensor data from the first sensor 606 and the second sensor 606. In accordance with aspects of the disclosure, the sensor module 602 can filter the data to remove noise, encode the data, decode the data, merge the data, extract frames, or perform other functions. In an alternative configuration, the depth estimation module 608 can receive sensor data directly from the first sensor 606 and the second sensor 606.
[0083] In one configuration, the depth estimation module 608 can communicate with and / or work in conjunction with one or more of the sensor module 602, the transceiver 616, the processor 620, the communication module 622, the localization module 618, the motion module 626, the navigation module 624, the first sensor 606, the second sensor 604, and the computer-readable medium 614. The depth estimation module 608 can be configured to receive a two-dimensional (2D) image of an environment adjacent to the ego vehicle. The environment includes dynamic objects and static objects. The 2D image can be captured by the first sensor 606 and the second sensor 606.
[0084] The depth estimation module 608 can be configured to produce a depth map of the environment based on the 2D image. In one implementation, the accuracy of the depth estimation for dynamic objects in the depth map is greater than the accuracy of the depth estimation for static objects in the depth map. The depth estimation module 608 can work in conjunction with a view synthesis module (not shown in Figure 6 , such as the view synthesis module 508 described for Figure 5 Additionally, the depth estimation module 608 can work in conjunction with a 3D object detection network (not shown in Figure 6 , such as the 3D object detection network 510 described for Figure 5Working with the 3D object detection network 518 described above, the depth estimation module 608 can identify the position of the dynamic object in the 3D estimation. Finally, working in conjunction with at least the localization module 618, the motion module 626, and the navigation module 624, the depth estimation module 608 can control the action of the ego vehicle based on the identified position.
[0085] The depth estimation module 608 can be implemented by referring to Figure 3 The deep network described (such as the deep network 300), the posture network (such as the reference Figure 4 ), a view synthesis module, and / or a 3D object detection network.
[0086] Figure 7 A flow chart illustrating a process 700 for identifying objects and controlling a vehicle based on depth estimation according to aspects of the present disclosure is shown. The process 700 may be performed by a vehicle (such as a vehicle with a reference to FIG. Figure 1 The vehicle 100 described), deep networks (such as those referring to Figure 3 The deep network 300 described in Figure 6 The depth estimation module 608 described, as for Figure 4 The gesture network 400 described, and / or as described with reference to Figure 5 One or more of the training pipelines 500 described are executed.
[0087] like Figure 7 As shown, process 700 includes capturing a two-dimensional (2D) image of an environment adjacent to an ego vehicle. The 2D image may be a target image 200 of a scene 202 as described with reference to FIG. 2 . The environment may include dynamic objects and static objects (block 702 ). The environment may include one or more dynamic objects, such as vehicles, pedestrians, and / or cyclists. The environment may also include one or more static objects, such as roads, sidewalks, and / or buildings. The 2D image may be captured via a monocular camera integrated with the ego vehicle.
[0088] Process 700 may also include generating a depth map of the environment via a depth estimation network based on the 2D image (block 704). The depth map may be depth map 220 of scene 202 as described with reference to FIG. 2 . In one implementation, the accuracy of depth estimation for dynamic objects in the depth map is greater than the accuracy of depth estimation for static objects in the depth map. Figure 5 The accuracy can be higher with the described training.
[0089] For example, during training, the weights of each pixel corresponding to the location of the object in the 2D image can be adjusted. The depth estimation network (e.g., the depth network) can be trained based on the ground truth information and the adjusted weights. Additionally, during training, the location of the dynamic object can be identified based on the annotated ground truth information. For example, the location can be identified based on a 3D bounding box that identifies the location of the object in the 3D estimate. During training, the 3D bounding box can be converted to a 2D bounding box to identify the location of the object in the 2D image.
[0090] Additionally, in this example, the weights can be adjusted based on the first number of pixels that include depth information and the number of pixels corresponding to the location of the object in the 2D image. Alternatively, the weights can be adjusted based on photometric loss and supervised depth error loss. The photometric loss can be the photometric loss 520 as described with reference to Figure 5
[0091] The process 700 can also include generating a 3D estimate of the environment based on the depth map (block 706). The 3D estimate can be a reconstructed image (e.g., a warped source image), such as the reconstructed image 510 as described with reference to Figure 5
[0092] Based on these teachings, one skilled in the art will appreciate that the scope of the disclosure is intended to cover any aspect of the disclosure, whether implemented independently of, or combined with, any other aspect of the disclosure. For example, an apparatus can be implemented or a method can be practiced using any number of the aspects set forth. In addition, the scope of the disclosure is intended to cover an apparatus or method practiced using, as an alternative to some or all of the aspects set forth, other structure, functionality, or structure and functionality. It should be understood that any aspect of the disclosure can be implemented alone, or in combination with any other aspect or aspects. Further, it is intended that the scope of the disclosure is to cover any and all modifications, variations, combinations or equivalents that fall within the spirit and scope of the present disclosure. Accordingly, proper scope of the disclosure is to be indicated by the appended claims rather than by the foregoing description, and all changes that come within the meaning and range of equivalents are intended to be embraced therein.
[0093] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects.
[0094] Although specific aspects are described herein, many variations and permutations of these aspects fall within the scope of the disclosure. Although some benefits and advantages of the preferred aspects are described herein, it is to be understood that a specific aspect can not include all of the benefits and advantages described herein. Indeed, some aspects can not implement a specific benefit or advantage, but in doing so an additional benefit or advantage not specifically described herein can be implemented. Different aspects can have different advantages, and these aspects can be used individually or in any combination.
[0095] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure), ascertaining and the like. Additionally, "determining" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory) and the like. Furthermore, "determining" can include resolving, selecting, choosing, establishing and the like.
[0096] As used herein, the phrase "at least one of a list of items" means any single one of the items in the list and any combination of two or more of the items in the list. As an example, "at least one of: a, b, or c" means: a, b, c, a-b, a-c, b-c, and a-b-c.
[0097] The various illustrative logical blocks, modules, and circuits described in connection with the disclosure can be implemented or performed with a processor specially configured to perform the functions described in the disclosure. The processor can be a neural network processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof specially designed to perform the functions described herein. Alternatively, the processing system can include one or more neuromorphic processors for implementing the models of neurons and nervous systems described herein. The processor can be a microprocessor, a controller, a microcontroller, or a state machine specially configured to perform the functions described herein. A processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or other specific configurations described herein.
[0098] The step of the method or algorithm described in conjunction with the present disclosure can be directly implemented with the software module executed by hardware, processor or the combination of the two.Software module can reside in storage or machine-readable medium, including random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), register, hard disk drive, removable disk, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage device or any other medium that can be used for carrying or storing the desired program code of the form of instruction or data structure and can be accessed by computer.Software module can include single instruction or many instructions, and can be distributed on several different code segments, between different programs and across multiple storage media distribution.Storage medium can be coupled to processor so that processor can read information from storage medium and write information to storage medium.In an alternative, storage medium can constitute an integral body with processor.
[0099] The methods disclosed herein include one or more steps or actions for implementing the described methods. The method steps and / or actions may be interchangeable with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.
[0100] The functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in the device. The processing system may be implemented using a bus architecture. The bus may include any number of interconnecting buses and bridges, depending on the specific application and overall design constraints of the processing system. The bus may link various circuits (including processors, computer-readable media, and bus interfaces) together. The bus interface may be used to connect a network adapter to the processing system via the bus, among other things. The network adapter may be used to implement signal processing functions. For certain aspects, a user interface (e.g., a keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits that are well known in the art and therefore will not be described in any further detail, such as timing sources, peripherals, voltage regulators, power management circuits, etc.
[0101] The processor may be responsible for managing the bus and processing, including the execution of software stored on a machine-readable medium.Software shall be construed to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, pseudocode, hardware description language, or otherwise.
[0102] In a hardware implementation, the machine-readable medium can be a part of the processing system that is separate from the processor. However, as will be readily appreciated by those skilled in the art, the machine-readable medium or any portion thereof can be external to the processing system. For example, the machine-readable medium can include a transmission line, a carrier wave modulated by data, and / or a computer product separate from the device, all of which can be accessed by the processor through a bus interface. Alternatively, or in addition, the machine-readable medium or any portion thereof can be integrated into the processor, such as may be the case for a cache and / or a specialized register file. Although the various components discussed can be described as having specific locations, such as local components, they can also be configured in various ways, such as with certain components being configured as part of a distributed computing system.
[0103] The machine-readable medium may include several software modules. The software modules may include a sending module and a receiving module. Each software module may reside in a single storage device or be distributed across multiple storage devices. For example, when a triggering event occurs, the software module may be loaded from a hard drive into RAM. During execution of the software module, the processor may load some of the instructions into a cache to increase access speed. One or more cache lines may then be loaded into a special-purpose register file for execution by the processor. When referring to the functionality of a software module below, it will be understood that such functionality is implemented by the processor when executing instructions from the software module. In addition, it should be appreciated that various aspects of the present disclosure result in improvements in the functionality of a processor, computer, machine, or other system that implements such aspects.
[0104] If implemented in software, the functions may be stored on or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media includes both computer storage media and communication media including any storage media that facilitates transfer of a computer program from one place to another.
[0105] In addition, it should be appreciated that the modules and / or other appropriate means for performing the methods and techniques described herein may be downloaded and / or otherwise obtained by the user terminal and / or base station as appropriate. For example, such a device may be coupled to a server to facilitate the transmission of the means for performing the methods described herein. Alternatively, the various methods described herein may be provided via a storage means so that the user terminal and / or base station may obtain the various methods after being coupled to the device or providing the storage means to the device. Furthermore, any other suitable technology for providing the methods and techniques described herein to the device may be utilized.
[0106] It is to be understood that the claims are not limited to the precise configurations and components illustrated above. Various modifications, changes and variations can be made in the arrangement, operation and details of the methods and apparatus described above without departing from the scope of the claims.
Claims
1. A method for navigating a vehicle through an environment, comprising: identifying a first location of a dynamic object in a two-dimensional (2D) image of the environment based on identifying a previous location of the dynamic object in a first three-dimensional (3D) estimate of the environment, the 2D image including the dynamic object and static objects; assigning a first weight to each pixel in the 2D image associated with the dynamic object based on identifying the first location of the dynamic object; assigning a second weight to each pixel in the 2D image associated with the static objects, the first weight being greater than the second weight; generating, via a depth estimation network, a depth map of the environment based on the 2D image, the depth map including: a dynamic object depth estimate for the dynamic object, the dynamic object depth estimate being associated with a first accuracy based on the first weight; and a static object depth estimate for the static objects, the static object depth estimate being associated with a second accuracy based on the second weight, the first accuracy of the dynamic object depth estimate being greater than the second accuracy of the static object depth estimate; generating a second 3D estimate of the environment based on the depth map; identifying a second location of the dynamic object in the second 3D estimate; and controlling an action of the vehicle based on identifying the second location.
2. The method of claim 1, wherein: the dynamic object includes a pedestrian, a neighboring vehicle, or a cyclist; and the static objects include a road, a sidewalk, or a building.
3. The method of claim 1, further comprising training a depth estimation network of the vehicle by: adjusting the weight of each pixel in the 2D image associated with the dynamic object; and training the depth estimation network based on ground truth information and the adjusted weight.
4. The method of claim 3, further comprising: During training, the second location is identified based on annotated ground truth information.
5. The method of claim 3, further comprising: identifying the second location of the dynamic object based on a 3D bounding box identifying the second location of the dynamic object in the second 3D estimate; converting the 3D bounding box to a 2D bounding box; and identifying the first location of the dynamic object in the 2D image based on the 2D bounding box.
6. The method of claim 3, further comprising: adjusting the weight of each pixel associated with the dynamic object based on a first number of pixels including depth information and a number of pixels corresponding to a location of the dynamic object in the 2D image; or adjusting the weight of each pixel associated with the dynamic object based on photometric loss and supervised depth error loss. The 2D image is captured via a monocular camera integrated with the vehicle.
8. An apparatus for navigating a vehicle through an environment, comprising:
7. The method of claim 1, further comprising: a processor; a memory coupled to the processor; and instructions stored in the memory and operable, when executed by the processor, to cause the apparatus to perform the following processing: identify a first location of a dynamic object in a two-dimensional (2D) image of the environment based on identifying a previous location of the dynamic object in a first three-dimensional (3D) estimate of the environment, the 2D image including the dynamic object and a static object; assign a first weight to each pixel in the 2D image associated with the dynamic object based on identifying the first location of the dynamic object; assign a second weight to each pixel in the 2D image associated with the static object, the first weight being greater than the second weight; generate, via a depth estimation network, a depth map of the environment based on the 2D image, the depth map including: a dynamic object depth estimate for the dynamic object, the dynamic object depth estimate being associated with a first accuracy based on the first weight; and a static object depth estimate for the static object, the static object depth estimate being associated with a second accuracy based on the second weight, the first accuracy of the dynamic object depth estimate being greater than the second accuracy of the static object depth estimate; generate a second 3D estimate of the environment based on the depth map; identify a second location of the dynamic object in the second 3D estimate; and control an action of the vehicle based on identifying the second location.
9. The device of claim 8, wherein: the dynamic object includes a pedestrian, a neighboring vehicle, or a cyclist; and the static object includes a road, a sidewalk, or a building.
10. The device of claim 8, wherein the instructions further cause the device to train the depth estimation network of the vehicle by: adjusting the weight of each pixel in the 2D image associated with the dynamic object; and training the depth estimation network based on ground truth information and the adjusted weight.
11. The device of claim 10, wherein the instructions further cause the device to train the depth estimation network of the vehicle by identifying the second location based on annotated ground truth information during training.
12. The device of claim 10, wherein the instructions further cause the device to train the depth estimation network of the vehicle by: identifying the second location of the dynamic object based on a 3D bounding box identifying a second location of the dynamic object in a second 3D estimate; converting the 3D bounding box to a 2D bounding box; and identifying the first location of the dynamic object in the 2D image based on the 2D bounding box.
13. The device of claim 10, wherein the instructions further cause the device to train the depth estimation network of the vehicle by: adjusting the weight of each pixel associated with the dynamic object based on a first number of pixels including depth information and a number of pixels corresponding to a location of the dynamic object in the 2D image; or adjusting the weight of each pixel associated with the dynamic object based on a photometric loss and a supervised depth error loss. 14. The device of claim 8, wherein the instructions further cause the device to capture the 2D image via a mono camera integrated with the vehicle.
15. A non-transitory computer-readable medium having program code recorded thereon for navigating a vehicle through an environment, the program code being executed by a processor and comprising: program code to identify a first location of a dynamic object in a two-dimensional (2D) image of the environment based on identifying a previous location of the dynamic object in a first three-dimensional (3D) estimate of the environment, the 2D image including the dynamic object and a static object; program code to assign a first weight to each pixel in the 2D image associated with the dynamic object based on identifying the first location of the dynamic object; program code to assign a second weight to each pixel in the 2D image associated with the static object, the first weight being greater than the second weight; program code to generate, via a depth estimation network, a depth map of the environment based on the 2D image, the depth map including a dynamic object depth estimate for the dynamic object, the dynamic object depth estimate being associated with a first accuracy based on the first weight; and a static object depth estimate for the static object, the static object depth estimate being associated with a second accuracy based on the second weight, the first accuracy of the dynamic object depth estimate being greater than the second accuracy of the static object depth estimate; program code to generate a second 3D estimate of the environment based on the depth map; program code to identify a second location of the dynamic object in the second 3D estimate; and program code to control an action of the vehicle based on identifying the second location.
16. The non-transitory computer-readable medium of claim 15, wherein: the dynamic object includes a pedestrian, a neighboring vehicle, or a cyclist; and the static object includes a road, a sidewalk, or a building.
17. The non-transitory computer-readable medium of claim 15, wherein the program code further comprises program code to train a depth estimation network of the vehicle by: adjusting the weight of each pixel in the 2D image associated with the dynamic object; and training the depth estimation network based on ground truth information and the adjusted weight.
18. The non-transitory computer-readable medium of claim 17, wherein the program code further comprises program code to train a depth estimation network of the vehicle by identifying the second location based on annotated ground truth information during training.
19. The non-transitory computer-readable medium of claim 17, wherein the program code further comprises program code to train a depth estimation network of the vehicle by: identifying the second location of the dynamic object based on a 3D bounding box identifying a second location of the dynamic object in a second 3D estimate; converting the 3D bounding box to a 2D bounding box; and identifying the first location of the dynamic object in the 2D image based on the 2D bounding box.
20. The non-transitory computer-readable medium of claim 17, wherein the program code further comprises program code to train a depth estimation network of the vehicle by: adjusting the weight of each pixel associated with the dynamic object based on a first number of pixels that include depth information and a number of pixels corresponding to a location of the dynamic object in the 2D image; or adjusting the weight of each pixel associated with the dynamic object based on photometric loss and supervised depth error loss.
Citation Information
Patent Citations
Visual odometer method based on dynamic and static scene separation
CN110910447A