A method for jointly estimating scene depth using a combined lightweight ToF sensor and an RGB image sensor
Through the deep learning method combined with the feature fusion of lightweight ToF sensors and RGB image sensors, the problem of difficult to obtain high-resolution and dense depth maps in the prior art is solved, and high-quality depth map estimation at low hardware cost and power consumption is achieved, with wide application potential.
Patent Information
- Application Number
- CN202211064992.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-09-01
AI Technical Summary
The prior art has not used a lightweight ToF sensor and RGB image sensor to jointly estimate the scene depth, which leads to higher prices and power consumption than RGB image sensors, and it is difficult to obtain high-resolution, dense depth maps comparable to the quality of existing depth sensors.
A deep learning-based method is designed to combine the depth distribution measured by lightweight ToF sensors and the RGB images acquired by the RGB image sensor to establish a neural network structure to extract features with different resolutions, and fuse the two features through the fusion module to finally predict a high-resolution and dense depth map.
It realizes the depth sensor hardware settings of one to two orders of magnitude lower in hardware power consumption and price, and obtains high-resolution, dense depth maps comparable to the quality of existing depth sensors, with wide application potential in mixed reality and robotics fields.
Smart Images

Figure CN115482264B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a depth estimation method in the technical fields of mixed reality and robot vision perception, and particularly relates to a method for jointly estimating the depth of a scene based on a lightweight ToF sensor and an RGB image sensor. Background Art
[0002] Depth sensors play an important role in fields such as mixed reality applications and robots. The existing implementation methods of depth sensors mainly include binocular, structured light, and time-of-flight (ToF) methods. Among them, the binocular-based method is not commonly used in small mobile devices because the distance between the two cameras cannot be too close, so it has a relatively large size. For depth sensors based on structured light and ToF, high power consumption and high price are relatively big problems. Compared with RGB image sensors, depth sensors based on structured light and ToF are at least one to two orders of magnitude higher in price and power consumption.
[0003] There is a type of lightweight ToF sensor widely used in mobile devices, such as VL53L0, VL53L1, VL53L3, and VL53L5 in the FlightSense series of STMicroelectronics. These sensors are designed for applications such as simple gesture recognition, autofocus, and simple obstacle detection, and their price and power consumption can reach the same level as RGB image sensors. Compared with general ToF sensors, the main difference of this type of lightweight ToF sensor is that due to its lightweight design concept, its resolution is extremely low (usually not exceeding 10x10), while general ToF sensors usually have a resolution of tens of thousands or more. At the same time, each pixel of the lightweight ToF sensor measures the depth distribution of a region of the scene, so there is a large uncertainty, while each pixel of a general ToF sensor measures the depth value corresponding to an accurate three-dimensional position.
[0004] There is no report in the prior art on jointly applying a lightweight ToF sensor and an RGB imager to estimate the depth of a scene. Therefore, if a lightweight ToF sensor with price and power consumption advantages and an RGB imager can be jointly used to obtain a high-resolution and dense depth map with quality comparable to existing depth sensors, it has great significance in applications. Summary of the Invention
[0005] The object of the present invention is to overcome the deficiencies of the prior art. For a hardware setup that simultaneously incorporates a lightweight ToF sensor and an RGB image sensor, a deep learning-based method is designed to jointly estimate a high-resolution and dense depth map using the depth distribution measured by the lightweight ToF sensor and the RGB image captured by the RGB image sensor. The present invention jointly estimates the scene depth by combining the lightweight ToF sensor and the RGB image. Such a hardware setup can reduce the hardware power consumption and price by one to two orders of magnitude compared to existing depth sensors. However, through the method proposed in the present invention, a high-resolution and dense depth map comparable in quality to existing depth sensors can be obtained.
[0006] The technical solution of the present invention is as follows. First, the present invention provides a method for jointly estimating the scene depth by combining a lightweight ToF sensor and an RGB image sensor, including the following steps:
[0007] (1) Establish a neural network structure. Extract features of different resolutions from the scene depth distribution data measured by the lightweight ToF sensor and the RGB image data captured by the RGB image sensor through the neural network, and fuse the two types of features through the fusion module in the network. The fusion starts from the feature map with the lowest resolution and gradually progresses to the feature map with a higher resolution, and finally predicts a high-resolution and dense depth map;
[0008] (2) Construct a dataset to train the neural network structure. The dataset contains the scene depth distribution data measured by the lightweight ToF sensor and the RGB image data captured by the RGB image sensor as the input of the neural network, and at the same time contains the ground truth depth map as the supervised data; set a loss function to supervise the network training using the ground truth depth map, optimize the parameters of the neural network structure, and obtain the parameter values of all parameters in the neural network structure;
[0009] (3) Use the lightweight ToF sensor and the RGB image sensor to respectively collect the scene depth distribution data and the RGB image data of the same scene area, and align the measured depth distribution data to the RGB image area; load the parameter values of all parameters after training into the neural network structure, and then input the aligned depth distribution data and RGB image data into the neural network structure to output the finally predicted high-resolution and dense depth map.
[0010] As a preferred embodiment of the present invention, in the step (1), the neural network structure includes a depth distribution feature extraction module, an RGB image feature extraction module, a fusion module for depth distribution features and RGB image features, and a depth prediction module; the depth distribution feature extraction module first samples a number of depth hypotheses from the depth distribution, and then uses a depth encoder to extract corresponding depth hypothesis features from the number of depth hypotheses. The RGB image feature extraction module uses an image encoder to extract image features at multiple resolutions from the image. The fusion module fuses the previously extracted depth hypothesis features and image features at different resolutions. The depth prediction module receives the finally fused features and predicts the final high-resolution and dense depth map.
[0011] As a preferred embodiment of the present invention, for the depth distribution feature extraction module, first, a number of depth hypotheses are obtained from the depth distribution according to the sampling strategy, and the sampled depth hypotheses pass through a depth encoder to obtain depth hypothesis features with gradually increasing feature dimensions.
[0012] As a preferred embodiment of the present invention, the RGB image feature extraction module includes a number of image encoders, and each time it passes through an image encoder, a 2-fold downsampling process is performed; the original RGB image passes through a continuous number of image encoders to obtain image feature maps with gradually increasing downsampling multiples, and at the same time, their feature dimensions also gradually increase.
[0013] As a preferred embodiment of the present invention, the fusion module for depth distribution features and RGB image features includes a cross-modal feature fuser and a decoder;
[0014] The fusion module for depth distribution features and RGB image features fuses the depth hypothesis features with the same corresponding feature dimensions from the feature maps with higher downsampling multiples to those with lower downsampling multiples obtained by the RGB image feature extraction module in sequence. The specific processing is as follows:
[0015] S1. Obtain the decoded feature map by passing the lowest-resolution RGB feature map obtained from the RGB image feature extraction module through the decoder. Denote the lowest resolution as HxW, where H is the image height and W is the image width;
[0016] S2. Upsample the decoded feature map to obtain a decoded feature map with a resolution of (Hx2)x(Wx2);
[0017] S3. Fuse the RGB feature map with the same resolution as the decoded feature map in S2 obtained from the feature extraction module and the depth hypothesis features with the same feature dimensions as the RGB feature map through the cross-modal feature fuser to obtain a fused feature map with a resolution of (Hx2)x(Wx2);
[0018] S4. Concatenate the fused feature map of S3 and the decoded feature map in S2, and then process it through a decoder to obtain a new decoded feature map;
[0019] S5. Use the decoded feature map fused in S4 as the undecimated decoded feature map and return it to S2 for processing. Continuously repeat the processing steps of S2 - S4 to finally obtain a high-resolution fused feature map.
[0020] As a preferred solution of the present invention, the depth prediction module obtains a final high-resolution and dense depth prediction map by passing the fused feature map output by the fusion module through a depth decoder.
[0021] As a preferred solution of the present invention, the dataset in step (2) should include depth distribution measurement data and RGB image data of a lightweight ToF sensor as inputs to the neural network, and at the same time include a ground truth depth map as supervised data. However, these data can be derived or simulated from the original data in the dataset; at the same time, the dataset is not limited to known datasets or private datasets collected using a hardware device equipped with both a lightweight ToF sensor and an RGB image sensor.
[0022] As a preferred solution of the present invention, in step (3), the measurement data of the lightweight ToF sensor of the scene to be measured and the image data collected by the RGB image sensor are input into the neural network structure loaded with the parameter settings trained in step (2) to output a final predicted high-resolution and dense depth map.
[0023] The present invention estimates a high-resolution and dense depth map based on a hardware setup equipped with both a lightweight ToF sensor and an RGB image sensor. Compared with the existing depth sensors mentioned in the background art, the hardware power consumption and price of this device are one to two orders of magnitude lower. The method proposed by the present invention can improve the output of this device to a high-resolution and dense depth map comparable to the quality of existing depth sensors, and is expected to make this device a new design idea for depth sensors. The depth map obtained by the present invention can be widely applied in fields such as mixed reality and robotics. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic flowchart of the present invention;
[0025] Figure 2 is a schematic diagram of the overall network framework of the neural network structure of the present invention;
[0026] Figure 3 is a schematic diagram of the fusion module of the neural network structure of the present invention;
[0027] Figure 4 is an actual effect diagram of the depth prediction structure of the present invention. Detailed Implementation Manner
[0028] Now, the representative implementation manners shown in the accompanying drawings will be further refined. It should be understood that the following description is not intended to limit the implementation manners to a preferred one. On the contrary, it is intended to cover alternative forms, modified forms, and equivalent forms that may be included within the essence and scope of the described implementation manners defined by the appended claims.
[0029] As Figure 1 shown in the flowchart of
[0030] (1) Establish a neural network structure. Extract features of different resolutions from the depth distribution and RGB images through the neural network, and fuse the two types of features through the fusion module in the network. The fusion starts from the feature map of the lowest resolution and gradually ascends to the feature map of higher resolutions, and finally predicts a high-resolution and dense depth map.
[0031] (2) Construct a data set to train the neural network structure. The data set contains the depth distribution and RGB image data measured by the lightweight ToF sensor as the input of the neural network, and also contains the ground truth depth map as the supervised data. Set a loss function to supervise the network training using the ground truth depth map, optimize the parameter values of all parameters in the neural network structure, and obtain the parameter values of all parameters in the neural network structure.
[0032] (3) Use the lightweight ToF sensor and RGB image sensor to collect the scene depth distribution data and RGB image data of the same scene area respectively, and align the measured depth distribution data to the RGB image area. Load the parameter values of all trained parameters into the neural network structure, then input the aligned depth distribution data and RGB image data into the neural network structure, and output the finally predicted high-resolution and dense depth map.
[0033] The overall network framework of the neural network structure of the present invention is as Figure 2 shown. The neural network structure includes a depth distribution feature extraction module, an RGB image feature extraction module, a fusion module for depth distribution features and RGB image features, and a depth prediction module. The depth distribution feature extraction module first samples several depth hypotheses from the distribution, and then uses a depth encoder to extract corresponding depth features from the several depth hypotheses. The RGB image feature extraction module uses an image encoder to extract image feature maps at multiple resolutions for the image. The fusion module fuses the previously extracted depth hypothesis features and image feature maps at different resolutions. The depth prediction module receives the finally fused feature map and uses a depth decoder to predict the final high-resolution and dense depth map.
[0034] In a specific embodiment of the present invention, the depth distribution feature extraction module first samples a number of depth hypotheses from the depth distribution according to the probability density. The sampling method is uniform sampling in the inverse cumulative distribution function of the depth distribution. The characteristic of the sampling is that the higher the probability density of the depth range, the more the number of samples. The depth hypotheses obtained by sampling pass through three multi-layer perceptrons (MLPs) to obtain depth hypothesis features with gradually increasing feature dimensions.
[0035] Preferably, the RGB image feature extraction module includes five convolutional modules, and each convolutional module is processed with 2-fold downsampling. The original RGB image passes through five consecutive convolutional modules to obtain image feature maps with downsampling factors of 2, 4, 8, 16, and 32 times respectively, and at the same time, the feature dimensions increase sequentially.
[0036] Preferably, the fusion module of the depth distribution feature and the RGB image feature includes three Transformer modules based on the attention mechanism and a decoding module. Among them, two attention mechanisms based on cross-data modality are used for information exchange and fusion between the RGB image features and the depth hypothesis features. One of these two attention mechanisms is to learn the contribution degree of different depth hypothesis features to the image features, and then fuse the depth hypothesis features into the image features according to the contribution degree. The other is the opposite, fusing the image features into the depth hypothesis features. The third attention mechanism captures long-range context relationships through the attention mechanism on the single-modal image data for depth information propagation. The decoding module is a convolutional module.
[0037] Preferably, the depth prediction module includes a depth decoder, which consists of two sets of consecutive convolution-activation operations. Considering the operation efficiency, usually the fusion feature map obtained by the last fusion module only has half or a quarter of the resolution of the input RGB image. Therefore, the depth decoding module also has several upsampling operations to increase the resolution to the original image size.
[0038] In a preferred embodiment of the present invention, the dataset in step (2) should include the depth distribution measurement data and RGB image data of the lightweight ToF sensor as the input of the neural network, and at the same time include the ground truth depth map as the supervised data. However, these data can be derived or simulated from the original data in the dataset. At the same time, the dataset is not limited to known datasets or private datasets collected by using a hardware device equipped with both a lightweight ToF sensor and an RGB image sensor. For example, when training the neural network structure using a known dataset, if the known dataset does not include the depth distribution measurement data of the lightweight ToF sensor, the depth distribution measured by the lightweight ToF sensor can be simulated through the depth map data in the known dataset. The specific method is as follows: According to the depth map data in the existing data, first divide several regions to simulate the measurement range of the lightweight ToF sensor, and then count the depth histogram in each region. The depth data in this region is approximately regarded as conforming to a Gaussian distribution, and the mean and variance of the depth data in this region are calculated and regarded as the measurement data of the lightweight ToF sensor. If a private dataset collected by using a device including a lightweight ToF sensor, an RGB image sensor, and a depth camera is used, the depth distribution obtained by the lightweight ToF sensor and the RGB image obtained by the RGB image sensor are directly used as the input of the neural network, and then the depth map collected by the depth camera is used as the ground truth to supervise the network training.
[0039] To further illustrate the present invention, the following introduces an embodiment implemented according to the complete method of the present invention, and its implementation process is as follows:
[0040] Taking NYU-Depth V2 as an example of a known dataset, the idea and specific implementation steps of jointly estimating the scene depth by using a lightweight ToF sensor and an RGB image sensor are described. The RGB images and the ground truth depth maps of the embodiment are all from the NYU-Depth V2 known dataset, and the depth distribution measured by the lightweight ToF sensor in the embodiment comes from the simulation result based on the ground truth depth map.
[0041] Step 1: Using the division of the NYU-Depth V2 known dataset, the training set contains 24,000 RGB images and the corresponding ground truth depth maps. For the ground truth depth maps provided in the training set, step 2 is implemented.
[0042] Step 2: For the ground truth depth maps in the training set described in step 1, divide several regions to simulate the measurement range of the lightweight ToF sensor, and then count the depth histogram in each region. The depth data in this region is approximately regarded as conforming to a Gaussian distribution, and the mean and variance of the depth data in this region are calculated to obtain the simulated depth distribution measurement data of the lightweight ToF sensor.
[0043] Step 3: Input the RGB images in the training set described in Step 1 and the simulated depth distribution measurement data described in Step 2 into the neural network structure. The overall network framework of the neural network structure of the present invention is as shown in Figure 2 . Input the RGB images and depth distribution data into the RGB image feature extraction module and the depth distribution feature extraction module of the neural network structure respectively. The RGB image feature extraction module includes five convolutional modules, and 2-fold downsampling is performed after passing through each convolutional module. The original RGB images pass through five consecutive convolutional modules to obtain image feature maps with downsampling factors of 2, 4, 8, 16, and 32 times respectively, and the feature dimensions increase to 24, 40, 64, 176, and 2048 in sequence. For the depth distribution feature extraction module, first, a number of depth hypotheses are sampled from the depth distribution according to the probability density. The sampling method is uniform sampling in the inverse cumulative distribution function of the depth distribution. The characteristic of the sampling is that the number of samples in the depth range with higher probability density will be more. The sampled depth hypotheses pass through three multi-layer perceptrons to obtain depth hypothesis features with feature dimensions of 40, 60, and 176.
[0044] Input the RGB feature maps of each size and the corresponding depth hypothesis feature maps into the feature fusion module to fuse the two-way feature maps. The main framework structure of the fusion module of the present invention is as shown in Figure 3 . First, obtain the decoded feature map by passing the RGB feature map with 32-fold downsampling obtained from the RGB image feature extraction module through the decoding module, and then upsample the decoded feature map to obtain the decoded feature map with 16-fold downsampling. The RGB feature map with 16-fold downsampling obtained from the feature extraction module and the depth hypothesis features with the same feature dimension as the RGB feature map are jointly fused through three Transformer modules based on the attention module to obtain the fused feature with 16-fold downsampling. Concatenate the fused feature with 16-fold downsampling and the decoded feature with 16-fold downsampling, and then process it through the decoding module to obtain the decoded feature with 8-fold downsampling. Use the same fusion module to fuse the features with 8-fold and 4-fold downsampling, and finally obtain the fused feature with 2-fold downsampling.
[0045] Input the fused feature with 2-fold downsampling output by the fusion module into the depth prediction module. The depth prediction module includes a convolutional module and an operation of upsampling by 2 times. Finally, obtain the final high-resolution and dense depth prediction map.
[0046] Step 4: For the depth prediction map output in Step 3, use the ground truth depth map contained in the training set to set the total loss function, calculate the loss for each valid pixel in the ground truth depth map, and the valid pixels are the pixels with depth values. Add up all the losses of the valid pixels to obtain the total loss of each frame prediction result, and train each parameter in the neural network structure to minimize the total loss to achieve the effect of supervised learning.
[0047] The total loss function is a depth loss function with scale invariance, and its calculation method is as follows:
[0048]
[0049]
[0050] where d i represents the true depth map provided by the known data set, represents the depth map predicted by the neural network structure, T represents the total number of valid pixels in the true depth map, and λ is a coefficient that balances the root mean square error and the error variance.
[0051] The specific method for training the neural network structure is to perform iterative update of parameters according to the method of gradient backpropagation, and use the GPU for acceleration until the total loss is reduced to within the set threshold or the number of network iterations meets the requirements, then stop training. The final result using the above method is as Figure 4 shown. In the figure, the lighter the color of the prediction result and the true value depth map of the present invention, the farther the distance represents, and vice versa, the darker the color, the closer the distance represents. It can be seen that the prediction result of the present invention is basically the same as the true value depth map. However, the true value depth map is full of noise. For example, the true value depth map corresponding to the white wall in the distance has several dark patches. In contrast, the prediction result of the present invention is closer to the true depth of the flat wall.
[0052] The above embodiments are used to explain and illustrate the present invention, rather than to limit the present invention. Within the spirit and scope of the claims of the present invention, any modification and change to the present invention fall within the protection scope of the present invention.
Claims
1. A method for jointly estimating the depth of a scene using a combined lightweight ToF sensor and an RGB image sensor, characterized in that, The method includes the following steps: Step (1): Establish a neural network structure. Through the neural network, extract features of different resolutions from the scene depth distribution data measured by the lightweight ToF sensor and the RGB image data collected by the RGB image sensor, and fuse the two types of features through the fusion module in the network. The fusion starts from the feature map with the lowest resolution and gradually ascends to the feature map with a higher resolution, and finally predicts a high-resolution and dense depth map; Step (2): Construct a dataset to train the neural network structure. The dataset contains the scene depth distribution data measured by the lightweight ToF sensor and the RGB image data collected by the RGB image sensor as the input of the neural network, and also contains the ground truth depth map as the supervised data; Set the loss function to supervise the network training using the ground truth depth map, optimize the parameters of the neural network structure, and obtain the parameter values of all parameters in the neural network structure; Step (3): Use the lightweight ToF sensor and the RGB image sensor to collect the scene depth distribution data and the RGB image data of the same scene area respectively, and align the measured depth distribution data to the RGB image area; Load the parameter values of all trained parameters into the neural network structure, and then input the aligned depth distribution data and RGB image data into the neural network structure to output the finally predicted high-resolution and dense depth map; Step (4): For the depth prediction map output in step (3), use the ground truth depth map contained in the training set to set the total loss function, calculate the loss for each valid pixel in the ground truth depth map, and the valid pixel is the pixel with a depth value; Add up all the losses of the valid pixels to obtain the total loss of each frame of prediction result, and train each parameter in the neural network structure to minimize the total loss to achieve the effect of supervised learning; The total loss function is a depth loss function with scale invariance, and the calculation method is: Among them, d i represents the true depth map provided by the known data set, represents the depth map predicted by the neural network structure, T represents the total number of valid pixels in the true depth map, and λ is the coefficient that balances the root mean square error and the error variance.
2. The method for jointly estimating the scene depth by the combined lightweight ToF sensor and RGB image sensor according to claim 1, wherein: In the step (1), the neural network structure includes a depth distribution feature extraction module, an RGB image feature extraction module, a fusion module for depth distribution features and RGB image features, and a depth prediction module; The depth distribution feature extraction module first samples several depth hypotheses from the depth distribution, and then uses the depth encoder to extract the corresponding depth hypothesis features from the several depth hypotheses. The RGB image feature extraction module uses the image encoder to extract image features at multiple resolutions for the image. The fusion module fuses the previously extracted depth hypothesis features and image features at different resolutions. The depth prediction module receives the finally fused features and predicts the final high-resolution and dense depth map.
3. The method for jointly estimating the scene depth by the combined lightweight ToF sensor and RGB image sensor according to claim 2, characterized in that: For the depth distribution feature extraction module, first obtain several depth hypotheses from the depth distribution according to the sampling strategy, and the depth hypotheses obtained by sampling pass through the depth encoder to obtain depth hypothesis features with gradually increasing feature dimensions.
4. The method for jointly estimating the scene depth by the combined lightweight ToF sensor and RGB image sensor according to claim 2, characterized in that: The RGB image feature extraction module includes several image encoders, each of which is subjected to a 2-fold downsampling process; after the original RGB image passes through several consecutive image encoders, image feature maps with successively increasing downsampling multiples are obtained, and the feature dimensions are Also increase successively.
5. The method for jointly estimating the scene depth by the combined lightweight ToF sensor and RGB image sensor according to claim 2, characterized in that: The fusion module of the depth distribution feature and the RGB image feature includes a cross-modal feature fuser and a decoder; The fusion module of the depth distribution feature and the RGB image feature obtains the feature maps of different downsampling multiples from the RGB image feature extraction module, and sequentially fuses the depth hypothesis features with the same feature dimension from the high downsampling multiple to the low downsampling multiple. The specific processing is as follows: S1, the lowest resolution RGB feature map obtained from the RGB image feature extraction module is obtained through the decoder to obtain a decoded feature map, and the lowest resolution is calculated as H×W, where H is the image height and W is the image width; S2, upsampling the decoded feature map to obtain a decoded feature map with a resolution of (H×2)×(W×2); S3, the RGB feature map obtained from the feature extraction module with the same resolution as the decoded feature map in S2 and the deep hypothesis feature with the same feature dimension as the RGB feature map are fused through a cross-modal feature fuser to obtain a fused feature map with a resolution of (H×2)×(W×2); S4, concatenate the fusion feature map of S3 and the decoding feature map in S2, and then process it through the decoder to obtain a new decoding feature map; S5: The decoded feature map obtained by fusion in S4 is returned to S2 as the decoded feature map that has not been upsampled, and the process of steps S2 to S4 is repeated continuously to finally obtain a high-resolution fused feature map.
6. The method for jointly estimating the scene depth by the combined lightweight ToF sensor and RGB image sensor according to claim 5, characterized in that: The depth prediction module passes the fusion feature map output by the fusion module through a depth decoder to obtain a final high-resolution, dense depth prediction map.
7. The method for jointly estimating the scene depth by the combined lightweight ToF sensor and RGB image sensor according to claim 1, characterized in that: In the step (2), the data set should include the depth distribution measurement data and RGB image data of the lightweight ToF sensor as the input of the neural network, and also include the true value depth map as the supervision data, but these data can be derived or simulated from the original data in the data set; at the same time, the data set is not limited to a known data set or a private data set collected by a hardware device that is equipped with a lightweight ToF sensor and an RGB image sensor.
8. The method for jointly estimating the scene depth by the combined lightweight ToF sensor and RGB image sensor according to claim 1, characterized in that: In the step (3), the lightweight ToF sensor measurement data of the scene to be measured and the image data collected by the RGB image sensor are input into the neural network structure loaded with the parameter settings obtained by training in step (2), and the final predicted high-resolution, dense depth map is output.
Citation Information
Patent Citations
Image processing method and device and electronic equipment
CN112258528A
Indoor RGB-D image semantic segmentation method based on wavelet transform
CN114842216A