Lightweight monocular image depth estimation method and device
By introducing multiple local perception modules, space-channel dual feature fusion modules and subpixel convolution upsampling modules in the monocular depth estimation technology, the problems of large computing volume and low inference efficiency in the prior art are solved, and efficient depth estimation on low-computing equipment is achieved.
Patent Information
- Application Number
- CN202510431459.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing monocular depth estimation technology has problems with large calculation volume, low inference efficiency, and high requirements for equipment resources, especially in edge devices or real-time scenarios.
A lightweight monocular image depth estimation method is designed to expand the receptive field and enhance feature interaction through multiple local perception modules and space-channel dual feature fusion modules, combine with the sub-pixel convolution upsampling module to reduce the computational amount, and optimize the network through view reconstruction loss and depth smoothing loss.
It realizes efficient monocular depth estimation on low-computing equipment, maintains high accuracy and significantly reduces the calculation amount and model size, so that the model has faster inference speed on resource-constrained devices.
Smart Images

Figure CN119941817A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a lightweight monocular image depth estimation method and device. Background Art
[0002] Depth estimation technology is of great significance in applications such as autonomous driving, robot navigation, and augmented reality. This task aims to use single-frame two-dimensional image information to predict the depth value of each pixel in the scene. Monocular depth estimation is more challenging than binocular or multi-camera depth estimation. In recent years, many deep learning methods have been applied to the field of monocular depth estimation, but they still face problems such as large computational workload, low reasoning efficiency, and high requirements for device resources, especially in edge devices or real-time scenarios. These problems are particularly prominent. Currently, the widely used monocular depth estimation networks are mostly based on convolutional neural networks or architectures combining convolution and Transformer. These networks can effectively extract feature information of images and enhance the prediction ability of the network through multi-scale feature fusion. However, when dealing with complex scenes or depth information of different scales, such networks still have technical bottlenecks such as insufficient receptive field, high computational complexity, and lack of adaptability.
[0003] The Chinese authorization announcement number is "CN113870335B", and the name is "A study on monocular depth estimation based on multi-scale feature fusion". This method uses the Non-Local attention mechanism to enhance the global feature capture capability, introduces the attention mechanism between shallow features, local features and deep features, and realizes cross-space and cross-feature layer interaction. It combines the multi-scale feature fusion module and the hollow space pyramid pooling module in the decoding network to expand the receptive field and enhance the ability to learn local details. This method has a high computational cost and is difficult to meet real-time requirements or be efficiently applied in resource-constrained scenarios.
[0004] In summary, how to design a model that can take into account both lightweight and high precision while realizing real-time monocular depth estimation on low-computing power devices is a problem that technical personnel in this field urgently need to solve. Summary of the invention
[0005] The technical solution of the present invention to solve the above technical problem is to provide a lightweight monocular image depth estimation method, comprising the following steps:
[0006] Step 1: Extract shallow features of images: prepare image datasets and extract shallow features through convolutional layers;
[0007] Step 2: Multi-scale feature modeling: The extracted shallow features are converted into multi-scale features through convolutional layers and pooling layers to capture objects and scene information of different scales.
[0008] Step 3, local-global feature interaction: Continuously expand the depth of convolution with large convolution kernels to build a larger receptive field, and perform spatial-channel dual feature aggregation through a lightweight attention mechanism to achieve local-global feature interaction;
[0009] Step 4, multi-level feature fusion: The feature maps of different scales in the encoder are fused with the high-resolution feature maps in the decoder through skip connections to maintain local details and global consistency in the depth map;
[0010] Step 5: The decoder gradually upsamples: the low-resolution feature map is gradually restored to the original resolution of the input image using a sub-pixel convolution upsampling module;
[0011] Step 6, depth map prediction: Generate a depth map of corresponding resolution through the output of the decoder;
[0012] Step 7, pose estimation: The pose estimation network receives adjacent frame images of the input frame image, performs pose estimation, and generates pose information;
[0013] Step 8, view reconstruction: reconstruct the current frame using the estimated depth map and the poses of the adjacent frames;
[0014] Step 9, model optimization: The network is optimized through view reconstruction loss and depth smoothing loss, and finally a lightweight model with high-precision depth estimation is generated.
[0015] Furthermore, in step 1, shallow feature extraction is performed through three convolutional blocks, each of which consists of a 3×3 convolution, a batch normalization layer, and a GELU activation function.
[0016] Furthermore, in step 2, multi-scale feature modeling downsamples the extracted shallow features through a 3×3 convolutional layer, and performs channel splicing with the feature map after average pooling of the feature image of the previous layer to generate multi-scale features.
[0017] Furthermore, in step 3, the local-global feature interaction constructs a larger receptive field through multiple local perception modules for the spliced feature maps of different resolutions, helping the network capture long-distance spatial information while retaining local details. Then, the spatial-channel dual feature fusion module uses a lightweight attention mechanism to perform spatial-channel dual feature fusion to achieve local-global feature interaction.
[0018] Furthermore, the multiple local perception module is composed of three local perception modules, each of which includes 1×1 convolution, normalization and SiLU activation function, 7×7 large convolution kernel depth expansion convolution, ECA lightweight attention mechanism and depth separable convolution.
[0019] Furthermore, the spatial-channel dual feature fusion module includes a channel attention branch and a spatial attention branch. The channel attention branch is used to capture the global dependencies between different positions in the input data, and the spatial attention branch is used to improve spatial relationship modeling.
[0020] Furthermore, in step 5, the sub-pixel convolution upsampling module is composed of two 3×3 convolutions, a sub-pixel convolution operation and a Sigmoid function.
[0021] Further, in step 9, the network is optimized by minimizing the image reconstruction loss and the edge-aware smoothing loss, where the image reconstruction loss is based on the photometric error composed of the structural similarity index and the L1 loss, and the edge-aware smoothing loss is used to smooth the generated inverse depth map.
[0022] In order to solve the above technical problems, the present invention also proposes a lightweight monocular image depth estimation device, which is used to execute instructions to implement the lightweight monocular image depth estimation method as described above, including:
[0023] Image acquisition unit: used to process the input visible light image;
[0024] Image processing unit: including feature encoding module, feature fusion module and feature decoding module, used to extract feature information of the image, generate feature maps of different resolutions, and perform feature fusion and decoding;
[0025] Prediction result output unit: used to output the predicted depth map;
[0026] Pose estimation network image processing unit: used to process adjacent frame images of the input image of the depth estimation network image to generate relative pose information;
[0027] Image reconstruction unit: used to reconstruct the current frame using the predicted depth map and the pose of the adjacent frames, and output the final predicted depth image.
[0028] Furthermore, the feature encoding module includes multiple convolution blocks, the feature fusion module adopts jump connection, and the feature decoding module uses a sub-pixel convolution upsampling module for decoding.
[0029] Compared with the prior art, this application has the following beneficial effects:
[0030] 1. The present invention designs multiple local perception modules, in which large convolution kernels with different expansion rates are used to deeply expand the convolution superposition, expand the receptive field, and adaptively select features through lightweight attention, thereby enhancing key local details. It can effectively alleviate the edge blur problem and more accurately detect small targets and clear edges, thereby improving the detail performance of the depth estimation map.
[0031] 2. The present invention designs a space-channel dual feature fusion module, which aggregates features in the spatial and channel dimensions. This module can model the global contextual relationship of the image and accurately perceive the structural information of complex scenes through a simplified attention mechanism, thus solving the problem of low accuracy and clarity of the network when processing complex scene structures.
[0032] 3. While maintaining high accuracy, the technical solution of the present invention reduces the amount of calculation through lightweight convolution and attention modules and multi-scale feature fusion structure, and adopts sub-pixel convolution for upsampling operation, which reduces the number of feature channels after upsampling while retaining image details, so that the model has a faster inference speed on resource-constrained devices. The final model size is 4.1M and the calculation amount is 5.9G.
[0033] 4. A lightweight monocular image depth estimation device provided by the present invention can implement the designed lightweight method. A novel lightweight monocular image depth estimation framework is embedded in the image processing unit, which can effectively improve the effect of depth estimation, so that the proposed device can quickly obtain high-precision depth estimation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0035] Figure 1 This is a flowchart of the steps of the lightweight monocular image depth estimation method and device of the present invention;
[0036] Figure 2 A network structure diagram of the processing method of the present invention;
[0037] Figure 3 A network structure diagram of the multiple local perception modules of the present invention;
[0038] Figure 4 It is a network structure diagram of the space-channel dual feature fusion module of the present invention;
[0039] Figure 5 This is a network structure diagram of the sub-pixel convolution upsampling module of the present invention;
[0040] Figure 6 It is a schematic diagram of the structure of a lightweight monocular image depth estimation device of the present invention. DETAILED DESCRIPTION
[0041] The present invention proposes a lightweight monocular image depth estimation method and device, aiming to design a novel lightweight depth estimation method so that the depth estimation task can be applied at high speed on mobile devices while maintaining high estimation accuracy.
[0042] The lightweight monocular image depth estimation method and device proposed by the present invention will be described in a specific embodiment as follows:
[0043] Embodiment 1:
[0044] A lightweight monocular image depth estimation method, such as Figure 1 As shown, the following steps are included:
[0045] Step 1: Extract shallow features of images: prepare image datasets and extract shallow features through convolutional layers;
[0046] Specifically, shallow feature extraction is performed through three convolutional blocks, each of which consists of a 3×3 convolution, a batch normalization layer, and a GELU activation function.
[0047] Step 2: Multi-scale feature modeling: The extracted shallow features are converted into multi-scale features through convolutional layers and pooling layers to capture objects and scene information of different scales.
[0048] Specifically, multi-scale feature modeling downsamples the extracted shallow features through a 3×3 convolutional layer, and performs channel splicing with the feature map after average pooling of the feature image of the previous layer to generate multi-scale features to capture objects and scene information of different scales;
[0049] Step 3, local-global feature interaction: Continuously expand the depth of convolution with large convolution kernels to build a larger receptive field, and perform spatial-channel dual feature aggregation through a lightweight attention mechanism to achieve local-global feature interaction;
[0050] Specifically, the local-global feature interaction module uses multiple local perception modules to construct a larger receptive field for the feature maps of different resolutions after the splicing operation, helping the network capture long-distance spatial information while retaining local details. Then, the spatial-channel dual feature fusion module uses a lightweight attention mechanism to perform spatial-channel dual feature fusion to achieve local-global feature interaction.
[0051] The multiple local perception module consists of three local perception modules, each of which has the same composition. First, 1x1 convolution, normalization and SiLU activation function are used to adjust the number of channels, which can effectively reduce the amount of calculation, select and compress important features, balance feature distribution and enhance nonlinear expression ability. Then, the receptive field is expanded through deep dilation convolution with a large 7×7 convolution kernel. The dilation rates of the three deep dilation convolutions are 1, 2, and 3 respectively. Then, batch normalization layer and SiLU activation are performed. Through the ECA lightweight attention mechanism, the selectivity and expression ability of features are enhanced while almost no increase in the amount of calculation, and the "hole" problem of the receptive field of the dilated convolution is reduced. Then, the number of channels is adjusted through 1×1 convolution layer, batch normalization layer and SiLU activation function, and the input features are used to improve the expression ability of the model through deep separable convolution. Finally, the input and features containing rich local information are added together to obtain the output of the local perception module.
[0052] All spatial-channel dual feature fusion modules have the same structure, which consists of a spatial attention branch and a channel attention branch. The channel attention branch first normalizes the input features, then generates queries, keys and values from the input features through linear layers and deep convolutions, calculates the similarity between queries and keys to generate attention weights, uses the attention weights to weighted sum the values, and finally maps the output back to the original space through a linear layer. The channel attention branch captures the global dependencies between different positions in the input data, enabling the model to dynamically select and focus on important information; the spatial attention branch first improves spatial relationship modeling through 3×3 deep convolution and batch normalization, then uses 1×1 convolution, batch normalization layer and GELU activation function to adjust the number of channels, compresses the input feature map in the channel dimension through maximum pooling and average pooling, so that the spatial features of each position are integrated, and splices the results of average pooling and maximum pooling in the channel dimension to form a feature map containing multiple spatial information. Then, a 7×7 convolution function is used to efficiently fuse the information of the multi-layer feature map to capture image details, and then the output of the convolution is passed through Sigmoid Activate, limit the attention weight to the range of [0, 1], and finally add residual connection to add the weighted feature map to the input feature map, so as to alleviate the gradient vanishing problem and enhance the training stability of the model; finally, add the feature maps after the channel attention branch and the spatial attention branch, and then perform layer normalization, 1×1 point-by-point convolution, GELU activation function and 1×1 point-by-point convolution to integrate information, introduce nonlinear transformation and standardize the features, which improves the feature extraction ability, computational efficiency and training stability of the model.
[0053] Step 4, multi-level feature fusion: The feature maps of different scales in the encoder are fused with the high-resolution feature maps in the decoder through skip connections to maintain local details and global consistency in the depth map;
[0054] Step 5: The decoder gradually upsamples: the low-resolution feature map is gradually restored to the original resolution of the input image using a sub-pixel convolution upsampling module;
[0055] Specifically, the sub-pixel convolution upsampling module consists of two 3×3 convolutions, a sub-pixel convolution operation, and a Sigmoid function. Sub-pixel convolution can reduce the number of feature channels after upsampling while retaining image details, which effectively reduces the computational burden of subsequent convolution operations and achieves lightweight decoder.
[0056] Step 6, depth map prediction: Generate a depth map of corresponding resolution through the output of the decoder;
[0057] Step 7, pose estimation: The pose estimation network receives adjacent frame images of the input frame image, performs pose estimation, and generates pose information;
[0058] Step 8, view reconstruction: reconstruct the current frame using the estimated depth map and the poses of the adjacent frames;
[0059] Step 9, model optimization: The network is optimized through view reconstruction loss and depth smoothing loss, and finally a lightweight model with high-precision depth estimation is generated.
[0060] Specifically, the network is optimized by minimizing the image reconstruction loss, which is based on the photometric error composed of the structural similarity index and the L1 loss, and the edge-aware smoothing loss is used to smooth the generated inverse depth map.
[0061] Embodiment 2:
[0062] A lightweight monocular image depth estimation method, the network structure diagram is as follows Figure 2 As shown, the specific steps include:
[0063] Step 1: The deep network receives a single RGB image as input and then performs shallow feature extraction on it through three convolution blocks (main pole feature extraction). Each convolution block consists of a 3×3 convolution, a batch normalization layer, and a GELU activation function.
[0064] Step 2: Downsample the extracted shallow features through a 3×3 convolutional layer, and perform channel splicing with the feature map after average pooling of the feature image of the previous layer to generate multi-scale features to capture objects and scene information of different scales;
[0065] Step 3: After the splicing operation, the feature maps of different resolutions are respectively constructed into a larger receptive field through multiple local perception modules to help the network capture long-distance spatial information while retaining local details. Then, the spatial-channel dual feature fusion module is used to perform spatial-channel dual feature fusion using a lightweight attention mechanism to achieve local-global feature interaction.
[0066] like Figure 3 As shown in the figure, the multiple local perception module consists of three local perception modules, and the composition of each local perception module is the same. First, 1x1 convolution, normalization and SiLU activation function are used to adjust the number of channels, which can effectively reduce the amount of calculation, select and compress important features, balance the feature distribution and enhance the nonlinear expression ability. Then, the receptive field is expanded by deep expansion convolution with a large 7×7 convolution kernel. The expansion rates of the three deep expansion convolutions are 1, 2, and 3 respectively. Then, batch normalization layer and SiLU activation are performed. Through the ECA lightweight attention mechanism, the selectivity and expression ability of the features are enhanced while almost no increase in the amount of calculation, and the "hole" problem of the receptive field of the expansion convolution is reduced. Then, the number of channels is adjusted through 1×1 convolution layer, batch normalization layer and SiLU activation function, and the input features are used to improve the expression ability of the model through deep separable convolution. Finally, the input and features containing rich local information are added together to obtain the output of the local perception module.
[0067] like Figure 4As shown in the figure, all spatial-channel dual feature fusion modules have the same structure, which are composed of spatial attention branches and channel attention branches. The channel attention branch first normalizes the input features, and then generates queries, keys and values from the input features through linear layers and deep convolutions, calculates the similarity between queries and keys to generate attention weights, and uses the attention weights to weighted sum the values. Finally, the output is mapped back to the original space through a linear layer. The channel attention branch captures the global dependencies between different positions in the input data, enabling the model to dynamically select and focus on important information. The spatial attention branch first improves spatial relationship modeling through 3×3 deep convolution and batch normalization, and then uses 1×1 convolution, batch normalization layer and GELU activation function to adjust the number of channels. The input feature map is compressed in the channel dimension through maximum pooling and average pooling, so that the spatial features of each position are integrated, and the results of average pooling and maximum pooling are spliced in the channel dimension to form a feature map containing multiple spatial information. Then, a 7×7 convolution function is used to efficiently fuse the information of the multi-layer feature map to capture image details, and then the output of the convolution is passed through Sigmoid. Activate, limit the attention weight to the range of [0, 1], and finally add residual connection to add the weighted feature map to the input feature map, so as to alleviate the gradient vanishing problem and enhance the training stability of the model; finally, add the feature maps after the channel attention branch and the spatial attention branch, and then perform layer normalization, 1×1 point-by-point convolution, GELU activation function and 1×1 point-by-point convolution to integrate information, introduce nonlinear transformation and standardize features, which improves the feature extraction ability, computational efficiency and training stability of the model;
[0068] Step 4: Fuse the feature maps of different scales in the encoder with the high-resolution feature maps in the decoder through skip connections to maintain local details and global consistency in the depth map;
[0069] In step 5, a sub-pixel convolution upsampling module is used to gradually restore the low-resolution feature map to the original resolution of the input image.
[0070] like Figure 5 As shown in the figure, the sub-pixel convolution upsampling module consists of two 3×3 convolutions, a sub-pixel convolution operation, and a Sigmoid function. The sub-pixel convolution can reduce the number of feature channels after upsampling while retaining image details, which effectively reduces the computational burden of subsequent convolution operations and realizes the lightweight decoder.
[0071] Step 6: Each prediction head generates a depth map of the corresponding resolution. The depth map contains the depth information of the scene, and the predicted depth value is an estimate of the distance of the object in the image from the camera.
[0072] Step 7: The pose estimation network receives adjacent frame images of the input frame image, performs pose estimation to generate pose information, which is used to generate view reconstruction loss.
[0073] Specifically, the pose estimation network extracts the features of continuous image pairs and regresses the relative rotation and displacement information of the image pairs to accurately estimate the motion of the camera. The pose estimation network accepts the input frame image and the previous frame image, uses the pre-trained Resnet-18 as the encoder to extract image features, uses multiple 3×3 convolutional layers and 1x1 convolutional layers to complete the pose regression, and finally outputs the rotation axis angle and displacement vector for generating the view reconstruction loss.
[0074] Step 8, using the estimated depth map and the poses of the adjacent frames to reconstruct the current frame;
[0075] Step 9: This work uses depth estimation as an image reconstruction task. , use the deep network to predict the corresponding depth map The pose network obtains adjacent video frames and predicts the target image and the source image The relative position between Finally, the predicted depth map and relative posture As a supervisory signal, generate the reconstructed target image According to the camera parameters and the predicted pose between two adjacent views, the reconstructed target image can be obtained. The learning target modeling is to minimize the target image and the image reconstruction loss between the synthetic target image , and in the predicted depth map Upper Constrained Edge-Aware Smoothness Loss , the specific calculation formula of image reconstruction loss is as follows:
[0076] ;
[0077] Among them, according to the original image , predict pose , predicted depth and camera intrinsic properties , the reconstructed target image can be obtained using the F function , Represents the photometric error composed of the structural similarity index and L1 loss:
[0078] ;
[0079] in, It is usually set to 0.85. In addition, in order to handle out-of-view pixels and occluded objects in the source image, the minimum loss of each pixel is calculated from the previous and next frames to soften the impact of occlusion on the reprojection process:
[0080] ;
[0081] Among them, "1" and "-1" represent the forward and backward adjacent frames respectively.
[0082] Edge-aware smoothing loss,In order to smooth the generated inverse depth map, edge-aware smoothing loss is calculated.,The specific calculation method is as follows:
[0083] ;
[0084] in, is the average normalized inverse depth. In addition, an automatic mask is calculated , filtering still pixels with the same appearance between two frames of a sequence, thereby filtering objects with the same motion as the camera in both still and moving frames.
[0085] Calculate the image reconstruction loss for the output of each scale of the network and edge-aware smoothing loss , and then take the average input to the total loss calculation to train the model. The total loss calculation formula is:
[0086] ;
[0087] in Set to 1 .
[0088] Start network training, set the number of training times to 40, the weight decay to 0.01, and the initial learning rate of the pose network during training to , the initial learning rate of the depth estimation network is set to .
[0089] Embodiment: 3:
[0090] A lightweight monocular image depth estimation device, used to execute instructions to implement the lightweight monocular image depth estimation method as described in Example 1 or Example 2, such as Figure 6 As shown, including:
[0091] Image acquisition unit: used to process the input visible light image;
[0092] Image processing unit: including feature encoding module, feature fusion module and feature decoding module, used to extract feature information of the image, generate feature maps of different resolutions, and perform feature fusion and decoding;
[0093] Prediction result output unit: used to output the predicted depth map;
[0094] Pose estimation network image processing unit: used to process adjacent frame images of the input image of the depth estimation network image to generate relative pose information;
[0095] Image reconstruction unit: used to reconstruct the current frame using the predicted depth map and the pose of the adjacent frames, and output the final predicted depth image.
[0096] Furthermore, the feature encoding module includes multiple convolution blocks, the feature fusion module adopts jump connection, and the feature decoding module uses a sub-pixel convolution upsampling module for decoding.
[0097] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A lightweight monocular image depth estimation method, characterized in that: The following steps are involved: Step 1: Extract shallow features of images: prepare image datasets and extract shallow features through convolutional layers; Step 2: Multi-scale feature modeling: The extracted shallow features are converted into multi-scale features through convolutional layers and pooling layers to capture objects and scene information of different scales. Step 3, local-global feature interaction: Continuously expand the depth of convolution with large convolution kernels to build a larger receptive field, and perform spatial-channel dual feature aggregation through a lightweight attention mechanism to achieve local-global feature interaction; Step 4, multi-level feature fusion: The feature maps of different scales in the encoder are fused with the high-resolution feature maps in the decoder through skip connections to maintain local details and global consistency in the depth map; Step 5: The decoder gradually upsamples: the low-resolution feature map is gradually restored to the original resolution of the input image using a sub-pixel convolution upsampling module; Step 6, depth map prediction: Generate a depth map of corresponding resolution through the output of the decoder; Step 7, pose estimation: The pose estimation network receives adjacent frame images of the input frame image, performs pose estimation, and generates pose information; Step 8, view reconstruction: reconstruct the current frame using the estimated depth map and the poses of the adjacent frames; Step 9, model optimization: The network is optimized through view reconstruction loss and depth smoothing loss, and finally a lightweight model with high-precision depth estimation is generated.
2. The lightweight monocular image depth estimation method according to claim 1, characterized in that: In step 1, shallow feature extraction is performed through three convolution blocks, each of which consists of a 3×3 convolution, a batch normalization layer, and a GELU activation function.
3. The lightweight monocular image depth estimation method according to claim 1, characterized in that: In step 2, multi-scale feature modeling downsamples the extracted shallow features through a 3×3 convolutional layer, and performs channel splicing with the feature map after average pooling of the feature image of the previous layer to generate multi-scale features.
4. The lightweight monocular image depth estimation method according to claim 3, characterized in that: In step 3, the local-global feature interaction constructs a larger receptive field through multiple local perception modules for the spliced feature maps of different resolutions, helping the network capture long-distance spatial information while retaining local details. Then, the spatial-channel dual feature fusion module uses a lightweight attention mechanism to perform spatial-channel dual feature fusion to achieve local-global feature interaction.
5. The lightweight monocular image depth estimation method according to claim 4, characterized in that: The multiple local perception module consists of three local perception modules, each of which includes 1×1 convolution, normalization and SiLU activation function, 7×7 large convolution kernel depth expansion convolution, ECA lightweight attention mechanism and depth separable convolution.
6. The lightweight monocular image depth estimation method according to claim 4, characterized in that: The spatial-channel dual feature fusion module includes a channel attention branch and a spatial attention branch. The channel attention branch is used to capture the global dependencies between different positions in the input data, and the spatial attention branch is used to improve spatial relationship modeling.
7. The lightweight monocular image depth estimation method according to claim 1, characterized in that: In step 5, the sub-pixel convolution upsampling module consists of two 3×3 convolutions, a sub-pixel convolution operation and a Sigmoid function.
8. The lightweight monocular image depth estimation method according to claim 1, characterized in that: In step 9, the network is optimized by minimizing the image reconstruction loss based on the photometric error composed of the structural similarity index and the L1 loss, and the edge-aware smoothing loss is used to smooth the generated inverse depth map.
9. A lightweight monocular image depth estimation device, used to execute instructions to implement the lightweight monocular image depth estimation method according to any one of claims 1 to 8, characterized in that: include: Image acquisition unit: used to process the input visible light image; Image processing unit: It includes a feature encoding module, a feature fusion module and a feature decoding module, which are used to extract feature information of the image, generate feature maps of different resolutions, and perform feature fusion and decoding; Prediction result output unit: used to output the predicted depth map; Pose estimation network image processing unit: used to process adjacent frame images of the input image of the depth estimation network image to generate relative pose information; Image reconstruction unit: Use the predicted depth map and the pose of the adjacent frames to reconstruct the current frame and output the final predicted depth image.
10. The lightweight monocular image depth estimation device according to claim 9, characterized in that: The feature encoding module includes multiple convolution blocks, the feature fusion module adopts skip connection, and the feature decoding module uses a sub-pixel convolution upsampling module for decoding.
Citation Information
Patent Citations
A monocular depth estimation method based on multi-scale feature fusion
CN113870335B
Dynamic scene image deblurring method, system and equipment
CN116681622A
CNN (Convolutional Neural Network) and Transform improved lightweight monocular depth estimation method
CN117710429A
Self-supervised monocular depth estimation method for lightweight hybrid network
CN118115555A
Cited By
Photometric three-dimensional reconstruction method based on coaxial light guide
CN120147565A
A photometric stereo reconstruction method based on coaxial light guidance
CN120147565B
Monocular depth estimation-based winter wheat heading stage growth vigor monitoring method
CN120198688A
Component segmentation method and device based on unmanned aerial vehicle infrared photovoltaic image
CN120689776A
Monocular depth estimation method, device, medium and system
CN121392454A