Domain Adaptation for Deep Densification
By training the machine learning system through synthetic data models, the depth sensor noise and data mode in the real environment are simulated, and combined with semi-supervised learning and adversarial loss function, the sparseness problem of sparse depth sensor data is solved, improving the effectiveness of the depth estimation model and the accuracy of autonomous driving.
Patent Information
- Application Number
- CN201980101422.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2039-10-24
AI Technical Summary
The depth data provided by existing sparse depth sensors in environmental scanning is sparse and difficult to effectively combine with the data of optical sensors, resulting in the sparseness of depth knowledge in computer vision applications.
The machine learning system is trained by synthesizing data models to simulate the depth sensor noise and data modes of the real environment, and combined with semi-supervised learning and adversarial loss functions, the effectiveness of the depth estimation model is improved.
The models trained on synthetic data can better adapt to the real environment, improve the conversion effect of sparse depth data to dense depth maps, and enhance the accuracy of autonomous driving and other applications.
Smart Images

Figure CN114600151B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image processing: for example, processing image data and relatively sparse depth data to form denser depth data, and training a model to perform such processing. Background Art
[0002] Sensing techniques such as RADAR, LiDAR, ultrasound, or other time-of-flight techniques utilize the fact that a wave with known characteristics is emitted and then reflected back from an object with specific density characteristics. If the speed of travel of the wave and the environmental characteristics are known, the time taken for the wave to travel through the medium can be determined using the echo or reflection, and then the distance to the point where the signal was reflected can be calculated. According to this technique, these waves can be electromagnetic waves or sound waves. These waves can be sent at various frequencies.
[0003] Although many sensors of this type, such as LiDAR, can be used to determine the distance of an object within a specified range with reasonable accuracy, they retrieve a relatively sparse sampling of the environment, especially if the object being scanned is far from the sensor. In other words, the data near the depth of the object at the distance sensing points they provide is discrete, and there are significant gaps between the vectors along which they provide depth data. The effect becomes apparent when comparing the distance measurements of a sparse depth sensor with the values passively obtained by an optical sensor. Figure 1 A situation of the LiDAR case is shown, where an RGB image is used as a reference and the measured distances are projected onto the image (see the paper "Are we ready for autonomous driving? The Kitti vision benchmark suite" by Andreas Geiger, Philip Lenz, and Raquel Urtasun, published at the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR 2012) in 2012). The RGB image shown at 101 (the image is presented in grayscale in this figure) is obtained from a camera mounted on top of a car near a LiDAR scanner that rotates 360 degrees to obtain the data shown at 102 (bird's-eye view). After projecting the acquired distances onto the RGB view, as shown at 103, the sparsity of the signal becomes apparent.
[0004] Many computer vision applications can benefit from depth knowledge and reduced sparsity. To derive the depth in all pixels of an RGB image (or at least more pixels than the subset actually measurable), the information from the original RGB image can be combined with the relatively sparse depth measurements from, for example, LiDAR, and a full-resolution depth map can be estimated. This task is generally referred to as depth completion. Figure 2 An example of a general depth completion pipeline is shown.
[0005] It is known to train a deep neural network to perform depth completion, where the training is supervised from real data, or is self-supervised from additional modalities or sensors such as a second camera, an inertial measurement unit, or continuously acquired video frames. (Self-supervised means training a neural network without explicit real supervision). The advantage of this method is that the sensed data naturally represents the data that can be expected in real-world applications, but there are problems such as noisy data and data acquisition costs.
[0006] There is a desire to develop a system and method for training a depth estimator that at least partially addresses these problems. Summary of the Invention
[0007] According to one aspect, there is provided a method for training an environmental analysis system, the method comprising: receiving a data model of an environment; forming a first training input according to the data model, the first training input comprising a visual stream representing the environment viewed from a plurality of positions; forming a second training input according to the data model, the second training input comprising a depth stream representing the depth of an object in the environment relative to the plurality of positions; forming a third training input, the third training input comprising a depth stream representing the depth of an object in the environment relative to the plurality of positions, the third training input being sparser than the second training input; estimating, by means of the analysis system, a series of depths having a sparsity lower than that of the third training input according to the first training input and the third training input; and adapting the analysis system according to a comparison between the estimated depth sequence and the second training input.
[0008] The data model is a synthetic data model, for example, a synthetic data model of a synthetic environment. Thus, the steps of forming the first training input and the second training input may include inferring those inputs in a self-consistent manner based on the environment defined by the data model. By using synthetic data to train the model, the efficiency and effectiveness of training can be improved.
[0009] The analysis system may be a machine learning system having a series of weights, and the step of adapting the analysis system may include adapting the weights. The use of a machine learning system can help form a model that provides good results for real data.
[0010] The third training input may be filtered and / or enhanced to simulate data generated by a physical depth sensor.
[0011] This can improve the effectiveness of the resulting model for real data.
[0012] The third training input may be filtered and / or enhanced to simulate data generated by a scanning depth sensor.
[0013] This can improve the effectiveness of the resulting model for real data.
[0014] The third training input can be filtered and / or enhanced by adding noise. This can improve the effectiveness of the resulting model for real data.
[0015] The method can include forming the third training input by filtering and / or enhancing the second training input. This can provide an effective way to generate the third training input.
[0016] The second training input and the third training input can represent depth maps. This can help form a model effective for real depth map data.
[0017] The third training input can be increased such that for each of a plurality of positions, depth data for vectors extending from the respective position at a common angle to the vertical direction is preferentially included. This helps simulate data obtained from a scanning sensor.
[0018] The third training input can be filtered by preferentially excluding data for vectors extending from one of the positions to an object at a greater or lesser depth that has been determined to be further from the estimate than a predetermined threshold. This can help simulate data obtained from a real sensor.
[0019] The third training input can be filtered by excluding data for vectors according to the color represented in the visual flow of the object towards which the respective vector is directed. This can help simulate data obtained from a real sensor.
[0020] The data model can be a model of a synthetic environment.
[0021] The method can include repeatedly adapting the analysis system. The method can include performing most of this adaptation according to data describing one or more synthetic environments. This can provide an effective way to train the model.
[0022] The step of training the system can include training the system by a semi-supervised learning algorithm. This can provide effective training results.
[0023] The method can include training the system by the steps of: providing a view of an environment oriented and centered for translation at a first reference frame as an input to the system, and in response to the input, estimating the depth associated with pixels in the view by means of the system; forming an estimated view of the environment oriented and centered for translation at a second reference frame different from the first reference frame according to the view and the estimated depth; estimating the visual authenticity of the estimated view; and adjusting the system according to the estimate. This helps effectively train the system.
[0024] The method can be performed by a computer executing code stored in a non - transient form. The code can be a stored computer program.
[0025] The method can include: sensing an image of a real environment by an image sensor; sensing a first depth map of the real environment by a depth sensor, the first depth map having a first sparsity; and forming a second depth map of the real environment by the system and based on the image and the first depth map, the second depth map having a lower sparsity than the first depth map. This helps to increase the data from the real depth sensor based on data from a real camera.
[0026] The method can include controlling an autonomous vehicle based on the second depth map. This helps to improve the accuracy of vehicle driving.
[0027] According to a second aspect, there is provided an environmental analysis system formed by the above - mentioned method.
[0028] According to a third aspect, there is provided an environmental analysis engine, including: an image sensor for sensing an image of the environment; a time - of - flight depth sensor; and the environmental analysis system as described in the previous paragraph; wherein the environmental analysis system is configured to receive the image sensed by the image sensor and the depth sensed by the depth sensor, and thereby form an estimate of the depth of an object depicted in the image.
[0029] According to a fourth aspect, there is provided an autonomous vehicle including the environmental analysis engine as described in the previous paragraph, the vehicle being configured to drive based on an estimate of the depth of an object depicted in the image.
[0030] According to a fifth aspect, there is provided a cellular communication terminal including the environmental analysis engine as described in the third aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present disclosure will now be described by way of example with reference to the accompanying drawings. In the drawings:
[0032] Figure 1 Distance measurements using a LiDAR sensor are shown.
[0033] Figure 2 The effect of a general depth completion pipeline is shown.
[0034] Figure 3 The domain gap between synthetic data and real data is shown.
[0035] Figure 3 (a) shows a rendering created synthetically using a driving simulator.
[0036] Figure 3 (b) shows a real driving scenario.
[0037] Figure 4 Shows geometric sensor simulation.
[0038] Figure 5 Shows sensor - specific data patterns.
[0039] Figure 6 Shows a stereo image pair of an autonomous vehicle.
[0040] Figure 7 Shows an image with mis - projections.
[0041] Figure 8 Shows an adversarial method using virtual projection correction.
[0042] Figure 9 Shows a multi - mode sensing system on a vehicle.
[0043] Figure 10 Shows a synthesized RGB image and its depth map.
[0044] Figure 11 Shows depth sparsification with Bernoulli point dropping.
[0045] Figure 12 Shows RGB projection from left to right.
[0046] Figure 13 Shows a student - teacher training scheme.
[0047] Figure 14 Shows a domain adaptation pipeline.
[0048] Figure 15 Shows an example of a camera for implementing the method described in this application.
[0049] Figure 16 Shows the results of the synthetic domain.
[0050] Figure 17 Shows the results on the real domain.
[0051] Figure 18 Shows the results of quantitative analysis and ablation study. Detailed implementation
[0052] This specification relates to training a machine learning system (also known as an artificial intelligence model) to form a relatively dense depth map from a sparse or less dense depth map and an image of a scene (e.g., an RGB image). A depth map is a set of data that describes the depth of an object from a location along a series of vectors extending in different directions from that location. If the depth map is directly derived from a real sensor, the data can be depth measurements. Conveniently, the AI model can be trained using a modern rendering engine and by formulating a pipeline that is fully trained on synthetic data without the need for real ground truth or additional sensors. Then, the trained model can be used to estimate depth from real data, for example, in autonomous driving or auto-pilot cars or smartphones.
[0053] A domain adaptation pipeline for sparse-to-dense depth image completion is fully trained on synthetic data without real data (i.e., real training data sensed from a real environment, such as via a depth sensor for sensing real depth data) or additional sensors (e.g., a second camera or an inertial measurement unit (IMU)). Although the pipeline itself is agnostic to sparse sensor hardware, the system is illustrated with LiDAR data, which is commonly used in driving scenarios where an RGB camera is mounted on top of a car together with other sensors. This system can be used in other scenarios.
[0054] Figure 1 An RGB, black-and-white, or grayscale image 101 of a scene can be processed together with depth data of the scene 102 to produce a map 103 that indicates the depth of objects in the scene from a given point. Figure 2 An RGB, black-and-white, or grayscale image 202 of a scene can be processed together with a relatively sparse depth map 201 of the scene to infer a more detailed depth map 203.
[0055] Domain adaptation is used to simulate real sensor noise. The solution described in this application includes four modules: geometric sensor simulation, data-driven sensor simulation, semi-supervised consistency, and virtual projection.
[0056] This system can be trained using real data specifically derived from a synthetic domain or can be used with self-supervised methods.
[0057] Geometric Sensor Simulation
[0058] We distinguish between two domains: synthetically created images and real acquisitions, as Figure 3As shown, it shows the domain gap between synthetic data and real data. Pairs created synthetically consisting of RGB images and depth maps can be retrieved densely in the synthetic domain, while if its 3D point cloud is projected onto the reference view of the RGB image, the LiDAR scanner of the scanned environment creates a sparse signal. To create these, the synthetic environment can be modeled. Then, an RGB image of this environment from a selected viewpoint can be formed, as well as a depth map from this viewpoint. For training purposes, the same steps can be repeated from multiple viewpoints to form a larger set of training data.
[0059] Figure 3 (a) shows the use of a driving simulator (see the paper "CARLA: An Open Urban Driving Simulator" by Dosovitskiy, Alexey, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun, arXiv number: 1711.03938, 2017), where the RGB image is at the top and the intensity-encoded depth map is at the bottom (farther regions are brighter). Figure 3 (b) shows a real driving scene (from the paper "Are We Ready for Autonomous Driving? The Kitti Vision Benchmark Suite" by Geiger, Andreas, Philip Lenz, and Raquel Urtasun at the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR 2012) in 2012), which has RGB acquisition (top) and projected LiDAR point cloud (bottom). While the synthetic depth data is dense, the LiDAR projection creates a sparse depth map on the RGB reference view.
[0060] For a human observer, the inherent domain gap is obvious. By simulating real noise patterns on the virtual domain, the effectiveness of the AI model can be improved. To simulate the LiDAR pattern on synthetic data, different sparsification methods can be selected. Past methods (such as the paper "Sparse to Dense: Depth Prediction from Sparse Depth Samples and a Single Image" by Ma, Fangchang, and Sertac Karaman at the IEEE International Conference on Robotics and Automation (ICRA 2018) in 2018) draw from a Bernoulli distribution independent of surrounding pixels with probability p Perform drawing. In this system, we utilize geometric information from real scenes to simulate signals in the real domain. We actually place depth sensors at relative spatial positions similar to RGB images in the real domain (LiDAR reference). To simulate the sparsity of sensor signals, we use binary projection masks on the LiDAR reference to reduce the sampling rate and project onto our synthesized views. The resulting synthesized sparse signals are visually closer to the real domain than in at least some prior art methods (as Figure 4 shown).
[0061] Figure 4 shows a virtual depth map (top image) created at a spatial position similar to the RGB image (LiDAR reference). The real mask on the LiDAR reference is used for sampling. The projected point cloud on the RGB reference (lower part) is visually closer to the real domain than random sampling of points (e.g., using naive Bernoulli sampling).
[0062] Data - specific sensor simulation
[0063] The geometrically corrected sparse signal from the previous step is closer to the real domain than in at least some other synthesis methods. However, it has been found that further processing to match the actual distribution is beneficial. This can be achieved by modeling two additional filters.
[0064] One effect is that in a real system, dark areas reflect less LiDAR signal and thus contain fewer 3D points. Another effect is that the correction process causes depth blurring in self - occlusion regions visible, for example, at thin structures. The LiDAR sensor can see beyond the RGB view and measure some objects in the "shadow" of thin structures. Due to the sparsity of the signal, these measurements do not have to be consistent along a ray in the RGB view to appear simultaneously on its projection. Figure 5 shows these two cases.
[0065] In Figure 5 , at the top, a projection from a scene with a dark object (black car) is shown. While the LiDAR projection is uniform on the surrounding structures, the dark car has significantly fewer 3D measurements on its surface. In the bottom image (see the paper "Noise - Aware Unsupervised Deep LiDAR - Stereo Image Fusion" by Cheng, Xuelian, Yiran Zhong, Yuchao Dai, Pan Ji, and Hongdong Li in CVPR 2019), the shadow of the pillar (black area) is the reason for the LiDAR sensor being mounted on the right side of the camera, and we project points based on the reference view for observation. The sparse sampling in the boxed area shows the depth blurring caused by projecting onto another reference where the points are actually occluded.
[0066] To mimic this sensor behavior, we perform data cleaning in the real domain by removing sparse signals that may be misaligned and noisy, such as Figure 5 shown. We achieve this by using a point dropout filter on data with a hard threshold. If the difference between the predicted depth and the sparse input is greater than a specified value, we do not use that point for sparse monitoring.
[0067] Additional selective sparsification is performed in the synthetic domain, where points from the sparse input are removed based on the RGB image. Although a naive method of removing points given a specific dropout distribution would discard points independently, we learn the probability distribution of actual point dropout in the synthetic domain. We use real LiDAR-RGB pairs to learn a model for RGB conditioning to discard more likely points (e.g., in dark regions such as Figure 5 shown).
[0068] In addition, random point dropout and consecutive recovery in the input LiDAR are used to provide a sparse monitoring signal in the real domain.
[0069] Virtual Projection
[0070] Most current models with self-monitoring training assume the existence of multiple available sensors. Typically, a second RGB camera is used together with the photometric loss between two views - these may be from a physical left-right stereo image pair or from consecutive frames in a video sequence - in order to train a model for estimating depth.
[0071] Figure 6 Two example images (top and bottom) from the Kitti dataset are shown. Even when spatially placed at different positions, the scene content is very similar. (See the paper "Are we ready for autonomous driving? The Kitti vision benchmark suite" by Geiger, Andreas, Philip Lenz, and Raquel Urtasun at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2012) in 2012). Simultaneous acquisitions have been made using two similar cameras placed spatially at different positions. The left RGB image is shown at the top of Figure 6 (a), and the right stereo image is shown at the bottom of Figure 6 (b).
[0072] For example, given the depth map of the top image, the RGB values can be projected onto the bottom image (and vice versa). As long as the dense depth map is correct, the resulting image will look very similar to the view from the left, despite view-point related occlusions. However, in the case of an incorrect depth estimate, the error becomes clearly visible, e.g., in Figure 7As shown, where the correctly estimated regions (such as streets and signs) do produce correctly viewed image regions, while the part around the car in the middle creates highly incorrect image content that emphasizes the flying pixels in the free space around the depth discontinuity at the edge of the car.
[0073] We utilize the observation that by using the synthesized new views and the adversarial loss on the new views after warping, the depth regions where the projection does not have problems are obtained. Thus, the adversarial method helps to align the projections from simulated and real data. Although the projection can be performed with any camera pose, the method does not require additional sensing: Figure 8 The pipeline is schematically shown. To further improve the quality, the method can also be combined with Figure 1 view consistency, for example, through the cycle consistency check of back-projecting to the origin.
[0074] Figure 8 The adversarial method using virtual projection correction is shown. The depth prediction 301 on the RGB input (left) together with the input camera pose (left) 304 are used to warp 305 the input RGB image (top) 302 to virtually create a new view 306, where the view 306 is evaluated by the adversarial loss 303 to help align the problem regions that can be strongly penalized after warping.
[0075] Semi-supervised consistency
[0076] One way to utilize domain adaptation and the closed-domain gap is by creating pseudo-labels on the target domain. We achieve this by leveraging the consistency with self-generated labels during training. By creating depth maps in the real domain that act as pseudo-labels, a semi-supervised method is applied to the depth completion task. During the training process, we incorporate noisy pseudo-predictions to draw out our noise model.
[0077] Although there are various ways to achieve semi-supervised consistency, in practice we follow the method of Tarvainen and Valpola (see the paper "Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results" by Antti Tarvainen and Harri Valpola in NeurIPS 2017).
[0078] In other words, we propose a domain adaptation pipeline for depth completion, which has the following remarkable stages:
[0079] 1. Geometric sensor simulation
[0080] ■ Project the real data patterns and artifacts on the synthetic domain.
[0081] 2. Data Dedicated Sensor Simulation
[0082] ■ Selective Sparsification and Data Refinement.
[0083] 3. Predict Virtual Projections onto Other Views to Reveal Errors in Predictions
[0084] ■ Perspective Changes Reveal Incorrect Depths. This Observation is Used for Adversarial Loss.
[0085] 4. Apply Semi-Supervised Method to Implement Self-Generated Depth Figure 1 Consistency
[0086] ■ Use a Harsh Teacher for Semi-Supervised Learning with Sparse Point Supervision.
[0087] The following describes an implementation of this method in a specific exemplary use case of RGB and LiDAR fusion in a driving scenario environment. For data generation, we use the driving simulator CARLA (see the paper "CARLA: An Open Urban Driving Simulator" by Dosovitskiy, Alexey, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun, arXiv:1711.03938, 2017), as well as real driving scenarios from the Kitti dataset (see the paper "Are We Ready for Autonomous Driving? The Kitti Vision Benchmark Suite" by Geiger, Andreas, Philip Lenz, and Raquel Urtasun at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2012), 2012).
[0088] Sensor and Data Simulation
[0089] The Kitti dataset assumes that the vehicle is equipped with multiple sensing systems as Figure 9 shown. We actually place the cameras in the simulated environment at the same relative spatial positions as the stereo camera pair and the sparse LiDAR sensor on the Kitti vehicle. For convenience, we add this information to the graphical illustration.
[0090] Figure 9Shows a multi - mode sensing system on a vehicle (in this case, a car). The car is equipped with multiple sensors at different spatial positions. Two stereo cameras (Camera 0 - Camera 1 for grayscale imaging and Camera 2 - Camera 3 for RGB) are placed in front of the Velodyne LiDAR scanner between the car axes. An additional IMU / GPS module is set at the back of the car. We place a virtual camera at the same relative position on the synthetic domain as the RGB camera shown by the arrow.
[0091] The real data of the driving dataset (see the paper "Are we ready for autonomous driving? The Kitti vision benchmark suite" by Andreas Geiger, Philip Lenz, and Raquel Urtasun at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2012) in 2012) is retrieved based on time and car pose information using intensive processing stages. Our simulator can project the depth information of the environment onto the same reference frame as the synthetic layout of the RGB camera to retrieve the real monitoring on the synthetic domain.
[0092] As a first step, to simulate the LiDAR pattern, we retrieve the depth map on the virtual LiDAR reference, as Figure 10 shown. Figure 10 Shows the synthetic RGB and depth maps. The top image shows the scene in the simulator where the RGB image has been rendered. The lower image (brighter color - coded) is the depth map of the same scene seen from the LiDAR reference frame. To sparsify the input depth map, sensor simulation is used. Bernoulli sampling is proposed in the paper "Sparse to Dense: Predicting Depth from Sparse Depth Samples and a Single Image" by Fangchang Ma and Sertac Karaman at the IEEE International Conference on Robotics and Automation (ICRA 2018). Meanwhile, Figure 11 provides sparse data with an appearance different from the real domain. We use the LiDAR mask from the real domain and learn RGB - adjusted point dropping to generate the real - looking LiDAR data as shown in Figure 4 shown.
[0093] Figure 11 Shows depth sparsity with Bernoulli point dropping. The depth map shown by Bernoulli (brighter) is sparsified with random point dropping. If the probability distribution of the random dropping is a Bernoulli distribution, the result is shown in the figure (lower left). The magnified area (lower right) shows the result of this environment - independent point dropping for a smaller area.
[0094] Virtual projection
[0095] The virtual projection of the RGB image from left to right makes it easier to notice depth estimation errors. Figure 12 The projection from the synthetic left RGB image to the right view with the ground truth depth is shown. Although the line of sight can be seen to be occluded, the color placement is correct, while the same process with an incorrectly estimated depth causes obvious artifacts, as Figure 7 shown. This is used together with the adversarial loss on the distorted image and identifies failures in updating the network parameters during training.
[0096] Semi-supervised consistency
[0097] Semi-supervision can be achieved through different approaches. A teacher-student model is implemented, providing pseudo-dense supervision for the real domain and using a weighted average consistency objective to improve the network's output. In our understanding, a mean teacher (see the paper "Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results" by Antti Tarvainen and Harri Valpola in NeurIPS 2017). is used to improve the prediction based on the pull of a noisy pseudo-ensemble of the student's predictions. Figure 13 The process architecture is shown in, where two copies of the same network (teacher and student) are used during training. Figure 13 The student-teacher training scheme is shown. The input is fed into two copies of the same network, namely the teacher and the student. For each branch, different augmentations and noise are added. The teacher creates a pseudo-supervision signal to train the student, and the exponential moving average of the student's weights is used to update the teacher's parameters during training. At each training step, the same mini-batch is fed into the two models with different augmentations and separately added noise. While the student's weights are updated in the normal way, the teacher's weights are the exponential moving average of the student's weights to make the predictions consistent between iterations. In this way, the teacher makes small step updates in the direction of the student. After the training phase, the teacher's pseudo-labels are no longer needed, and the student network is used to estimate the dense depth.
[0098] Figure 14 The entire pipeline including all the mentioned domains and modules is shown in.
[0099] Figure 14Shows the domain adaptation pipeline. These three boxes illustrate the synthetic domain and the real domain, as well as the new views created synthetically. On the synthetic domain (lower left box), sensor simulation is used to create pairs of RGB and LiDAR input signals from the presented full ground truth (GT) depth map. The depth estimation model is managed by the ground truth and shared with the real domain with RGB and LiDAR inputs. These points are filtered on the real domain (upper left box) to create additional sparse signals from the LiDAR for monitoring. Semi-supervision is achieved using a student-teacher model (the student in this figure is described as the depth estimation model), where the teacher network is only used during training. The camera pose defines a distinguishable distortion of the input RGB image to a new reference (e.g., the right view), where an adversarial loss or a consistency loss can be applied to identify errors in the predictions and help update the weights.
[0100] Comparison of results with other pipelines
[0101] Figure 16 and Figure 17 Show the qualitative results of the synthetic image and the real domain respectively. Interestingly, in cases where the network has learned generalization and completed simulation caused by RGB information, the network also outputs possible values in areas where there is no depth information in the real data. Precise predictions can also be made with possible information in areas that do not include LiDAR data to see a similar effect on the real domain.
[0102] Figure 18 Shows a quantitative analysis of the ablation study of different modules, where the method is compared with self-supervised methods in the prior art. By continuously adding different modules, improvements can be seen at each step, which supports the view that each component helps to complete the task of depth. It is possible to have a system that uses only some components, but it is expected to be less efficient. The error assessment also shows that the proposed method is comparable to modern self-supervised methods, even though they use additional sensor information from a second virtual camera (see the paper "Self-Supervised Sparse to Dense: Self-Supervised Depth Completion from LiDAR and Monocular Camera" by Ma, Fangchange, Guilherme Venturelli Cavalheiro, and Sertac Karaman in the International Conference on Robotics and Automation (ICRA 2019) in 2019 and another approach (see the paper "Heartbeats: Depth Completion from Inertial Odometry and Vision" (arXiv number: 1905.08616 2019) by Wong, Alex, Xiaohan Fei, and Stefano Soatto).
[0103] Figure 17Shows the results in the real domain. The top image shows the RGB input, the middle image shows the sparse LiDAR signal projected onto the RGB reference, and the bottom image shows the network output. It can be seen that thin structures are retrieved precisely, and depth completion fills the holes between relatively sparse LiDAR points and line-of-sight occlusions (e.g., around the column in the right center), and provides depth estimates for distant objects where only RGB information exists. In addition, in the absence of LiDAR information, dark areas such as the car on the right are filled with possible depth values.
[0104] Figure 15 Shows an example of a camera for implementing an image processor to process images captured by an image sensor 1102 in the camera 1101. Such a camera 1101 typically includes some built-in processing capabilities. This can be provided by a processor 1104. The processor 1104 can also be used to implement the basic functions of the device. The camera typically also includes a memory 1103.
[0105] A transceiver 1105 is capable of communicating with other entities 1110, 1111 via a network. These entities can be physically remote from the camera 1101. The network can be a publicly accessible network such as the Internet. The entities 1110, 1111 can be deployed in the cloud. In one example, the entity 1110 is a computing entity and the entity 1111 is a command and control entity. These entities are logical entities. In practice, each of these entities can be provided by one or more physical devices such as servers and data storage, and the functions of two or more entities can be provided by a single physical device. Each physical device implementing an entity includes a processor and a memory. The device can also include a transceiver for sending data to the transceiver 1105 of the camera 1101 and receiving data from the transceiver 1105 of the camera 1101. The memory stores code executable by the processor in a non-transitory manner to implement the corresponding entity in the manner described in this application.
[0106] The command and control entity 1111 can train the artificial intelligence models used in each module of the system. This is typically a computationally intensive task, even though the resulting model can be effectively described, and thus may be effective for the development of algorithms executed in the cloud where a large amount of energy and computing resources are available. It is expected that this is more efficient than forming such a model on a typical camera.
[0107] In one implementation, once a deep learning algorithm is developed in the cloud, the command and control entity can automatically form the corresponding model and have it sent to the relevant camera device. In this example, the processor 1104 implements the system at the camera 1101.
[0108] In another possible implementation, an image can be captured by the camera sensor 1102, and the image data can be sent to the cloud by the transceiver 1105 for processing in the system. Then, the resulting target image can be sent back to the camera 1101, as shown at 1112 in Figure 15 .
[0109] The method can be deployed in various ways, such as in the cloud, on the device, or in dedicated hardware. As described above, the cloud facility can perform training to develop new algorithms or improve existing algorithms. Depending on the computing power close to the data corpus, training can be performed close to the source data, or it can be performed in the cloud, for example, using an inference engine. The system can also be implemented in the camera, in dedicated hardware, or in the cloud.
[0110] The vehicle can be equipped with a processor programmed to implement the model trained as described above. The model can obtain inputs from the images and depth sensors carried by the vehicle and can output a denser depth map. This denser depth map can be used as an input for controlling the vehicle's autonomous driving or collision avoidance system.
[0111] The applicant hereby separately discloses each individual feature described in this application and any combination of two or more such features. With the ordinary knowledge of those skilled in the art, such features or combinations can be implemented as a whole based on this specification, regardless of whether such features or combinations of features can solve any problem disclosed in this application, and without limiting the scope of the claims. The applicant points out that various aspects of this disclosure can consist of any such individual feature or combination of features. In view of the above description, it will be apparent to those skilled in the art that various modifications can be made within the scope of this disclosure.
Claims
1. A method for training an environmental analysis system, characterized in that The method includes: Receiving a data model of an environment; Forming a first training input according to the data model, the first training input including a visual stream representing the environment viewed from multiple positions; Forming a second training input according to the data model, the second training input including a depth stream representing the depth of an object in the environment relative to the multiple positions; Forming a third training input, the third training input including a depth stream representing the depth of an object in the environment relative to the multiple positions, the third training input being sparser than the second training input; Estimating a depth sequence, which is denser than the third training input, according to the first training input and the third training input by means of the analysis system; and Adapting the analysis system according to a comparison between the estimated depth sequence and the second training input; Filtering the third training input by excluding data of vectors according to the color represented in the visual stream of the object towards which the corresponding vectors are directed.
2. The method according to claim 1, characterized in that, The analysis system is a machine learning system with a series of weights, and the step of adapting the analysis system includes adapting the weights.
3. The method according to claim 1, characterized in that, Filtering the third training input to simulate data generated by a physical depth sensor.
4. The method according to claim 1, characterized in that, Filtering the third training input to simulate data generated by a scanning depth sensor.
5. The method according to claim 1 above, characterized in that, Enhancing the third training input by adding noise.
6. The method according to claim 1 above, characterized in that, Including forming the third training input by filtering the second training input.
7. The method according to claim 1 above, characterized in that, The second training input and the third training input represent depth maps.
8. The method according to claim 7, wherein Increasing the third training input so that for each of the multiple positions, it preferentially includes depth data for vectors extending at a common angle extending vertically from the corresponding position.
9. The method according to claim 1, wherein Filtering the third training input by preferentially excluding data of vectors extending from one of the positions to an object at a depth greater or smaller than a predetermined threshold that has been determined to be further away from the estimate.
10. The method according to any one of the preceding claims 1-9, characterized in that, The data model is a model of a synthetic environment.
11. The method according to any one of the preceding claims 1-9, characterized in that The method includes repeatedly adapting the analysis system, and the method includes performing most of this adaptation according to data describing one or more synthetic environments.
12. The method according to any one of the preceding claims 1-9, characterized in that, The step of training the system includes training the system by a semi-supervised learning algorithm.
13. The method according to any one of the preceding claims 1-9, characterized in that, Including training the system by the following steps: Providing a view of the environment oriented and centered at a first reference frame as an input to the system, and in response to the input, estimating the depth associated with the pixels in the view by means of the system; Forming an estimated view of the environment oriented and centered at a second reference frame different from the first reference frame according to the view and the estimated depth; Estimating the visual authenticity of the estimated view; Adjusting the system according to the visual authenticity of the estimated view.
14. The method according to any one of the preceding claims 1-9, characterized in that, The method is executed by a computer executing code stored in a non-transitory form.
15. The method according to any one of the preceding claims 1-9, characterized in that, Including: An image sensor senses an image of the real environment; A depth sensor senses a first depth map of the real environment, the first depth map having a first sparsity; Form a second depth map of the real environment through the system and based on the image and the first depth map, the second depth map having a lower sparsity than the first depth map.
16. The method according to claim 15, wherein Comprising: Control an autonomous vehicle based on the second depth map.
17. An environmental analysis system formed by the method according to any one of claims 1 to 16.
18. An environmental analysis engine, characterized in that, Comprising: An image sensor for sensing an image of the environment; A time-of-flight depth sensor; The environmental analysis system according to claim 17; Wherein, the environmental analysis system is configured to receive the image sensed by the image sensor and the depth sensed by the depth sensor, and thereby form an estimate of the depth of the object depicted in the image.
19. An autonomous vehicle comprising the environmental analysis engine according to claim 18, characterized in that, The vehicle is configured to drive based on the estimate of the depth of the object depicted in the image.
20. A cellular communication terminal, characterized in that, Comprising the environmental analysis engine according to claim 18.