Night depth estimation method based on self-supervision framework
By using a self-supervised framework and adaptive image enhancement module in night depth estimation, the problems of low light and high noise in night depth estimation are solved, and efficient and real-time depth estimation effects on edge devices are achieved.
Patent Information
- Application Number
- CN202510015966.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art is difficult to effectively deal with the low light and high noise problems in night depth estimation, and the existing lightweight methods are mainly trained and evaluated in daytime scenes, and cannot adapt to the complexity of night scenes.
A night depth estimation method based on a self-supervised framework is proposed, including acquiring daytime data sets and nighttime RGB images, training through depth estimation network, pose network and adversarial neural network, and processing images using adaptive image enhancement modules to realize depth estimation of night scenes.
This method can effectively perform night depth estimation on resource-constrained edge devices, provide high-quality depth maps, are suitable for night scenes, and achieve real-time and low-cost deployment.
Smart Images

Figure CN119963616A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of autonomous driving or robotics technology, and in particular relates to a nighttime depth estimation method based on a self-supervisory framework. Background Art
[0002] Monocular depth estimation is an important task in computer vision and is crucial for applications such as autonomous driving, robotic navigation, and 3D reconstruction. It is expensive to collect high-quality depth data in a wide range of environments by using expensive depth sensors such as LiDAR and TOF. Self-supervised monocular depth estimation that exploits the photometric limitations between monocular image sequences without collecting ground-truth depth maps is gaining increasing attention. In addition to research on daytime datasets such as KITTI and Make 3D. Therefore, many efforts have been made to develop self-supervised methods that train deep networks to estimate depth maps by exploring geometric cues in videos, i.e., reconstructing a target view (or frame) from another view, instead of exploiting high-quality depth data. Moreover, their performance in well-lit environments is comparable to that of supervised methods. But few methods can handle more challenging nighttime scenes. In fact, nighttime scenes include two important issues, low visibility and varying illumination, which lead to strange depth outputs from most existing self-supervised methods. Low visibility usually produces textureless areas. A dark area with indistinguishable visual texture exacerbates the model outputting a depth map with large holes, while two image patches of different brightness are cropped from the same position of two temporally adjacent frames. This brightness inconsistency leads to imperfect reconstruction of the target view, i.e., large training loss.
[0003] At the same time, the problem of lightweight models in monocular depth estimation has also attracted attention. Recent research often focuses on reducing model parameters and exploring new convolution operations through techniques such as knowledge distillation. This allows models to be deployed on edge devices to solve real-world problems. However, few people have proposed lightweight methods to address the challenges posed by more complex night scenes. The deep networks in current night methods are mainly adapted from Monodepth2, which exhibits relatively high model complexity. Existing lightweight methods are mainly trained and evaluated in daytime scenes.
[0004] In summary, the shortcomings of the prior art are as follows:
[0005] 1. It needs to rely on manual data collection, which has high labor and time costs;
[0006] 2. The method of using multiple sensors to collect data increases the hardware cost and fails to collect accurate nighttime depth information;
[0007] 3. Low light and high noise problems may occur in undesirable night scenes. The current model cannot effectively process data in the corresponding scenes.
[0008] 4. The computational cost of using the complex backbone network is too high, and real-time performance cannot be guaranteed when deployed in real scenarios.
[0009] In summary, it is of great significance to propose a lightweight deep network and integrate it into the nighttime self-supervised depth estimation architecture to deal with the low light and high noise problems in nighttime scenes. Summary of the invention
[0010] In order to solve the above technical problems, the present invention proposes a night depth estimation method based on a self-supervised framework, which is suitable for application on resource-constrained edge devices and can achieve good recognition effect on specific night scenes.
[0011] The present invention provides a nighttime depth estimation method based on a self-supervisory framework, comprising:
[0012] Obtaining a daytime data set, inputting the daytime data set into a pre-trained estimation network model, and obtaining a daytime depth map;
[0013] Obtain a target frame and a source frame of a nighttime RGB image, input the target frame into a depth estimation network model, and obtain a nighttime depth map;
[0014] Splicing the channels of the target frame and the source frame, and obtaining the posture changes of the spliced target frame and the source frame according to the posture network model;
[0015] The target frame and the source frame are enhanced by using an adaptive image enhancement module to obtain a first source frame and a first target frame;
[0016] Reconstructing the nighttime depth map, the posture change, and the first source frame to obtain a reconstructed target frame, and supervising the reconstructed target frame and the first target frame using a first loss function;
[0017] The daytime depth map and the nighttime depth map are input into an adversarial neural network model, and the estimated distribution of the self-supervised nighttime depth map is consistent with the depth distribution under normal daytime lighting scenes, wherein the adversarial neural network uses a second loss function for supervised training.
[0018] Optionally, the depth estimation network model includes: a depth encoder and a depth decoder;
[0019] The depth encoder is used to encode the target frame;
[0020] The depth decoder is used to decode the encoded target frame to obtain a nighttime depth map.
[0021] Optionally, the deep encoder includes: a convolution layer, a plurality of downsampling layers, and a plurality of feature fusion modules;
[0022] The convolution layer is used to refine feature extraction and further enhance the ability to capture local details;
[0023] The downsampling layer is used to reduce the spatial resolution of the feature map layer by layer, increase the receptive field, and retain the main scene structure information;
[0024] The feature fusion module is used to simultaneously realize the perception of local details and scene structures and integrate local and global image features.
[0025] Optionally, the depth decoder comprises: a plurality of upsampling layers and an output module;
[0026] The upsampling layer is used for upsampling by bilinear interpolation;
[0027] The output module is used to output the upsampling result.
[0028] Optionally, splicing the channels of the target frame and the source frame, and obtaining the posture changes of the spliced target frame and the source frame according to the posture network model includes:
[0029] The target frame and the source frame channels are spliced, and the spliced target frame and the source frame are input into the posture network model to obtain the relative posture between the target frame and the source frame.
[0030] Optionally, enhancing the target frame and the source frame by using an adaptive image enhancement module to obtain a first source frame and a first target frame includes:
[0031] Divide the target frame and the source frame into blocks to obtain a plurality of image blocks;
[0032] Introducing a contrast limiting parameter to obtain a histogram of the image block;
[0033] The histogram is equalized, and the image blocks after the splicing process are smoothed using bilinear difference to obtain the first source frame and the first target frame.
[0034] Optionally, the first loss function is:
[0035]
[0036] Among them, I' t is the enhanced target frame, is the reconstructed target frame, a is a parameter, and ||·||1 is the L1 norm.
[0037] Compared with the prior art, the present invention has the following advantages and technical effects:
[0038] (1) The present invention designs a lightweight architecture for monocular depth estimation under nighttime conditions, which achieves a good balance between model complexity and accuracy and is particularly suitable for application on resource-constrained edge devices.
[0039] (2) fitting depth estimation via adversarial neural networks using unpaired daytime and nighttime RGB inputs;
[0040] (3) No need to rely on complex and expensive depth cameras for data collection, training can be performed through self-supervision;
[0041] (4) By training or fine-tuning the model, depth estimation can be performed in real scenes in real time. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0043] Figure 1 is a flow chart of a nighttime depth estimation method based on a self-supervisory framework according to an embodiment of the present invention;
[0044] Figure 2 It is a structural diagram of a lightweight depth estimation network model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0045] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0046] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0047] This paper proposes a nighttime depth estimation method based on a self-supervised framework, such as Figure 1 As shown, the specific steps include:
[0048] Obtain a daytime data set, input the daytime data set into a pre-trained estimation network model, and obtain a daytime depth map;
[0049] Obtain the target frame and source frame of the nighttime RGB image, input the target frame into the depth estimation network model, and obtain the nighttime depth map;
[0050] The channels of the target frame and the source frame are spliced, and the posture changes of the spliced target frame and the source frame are obtained according to the posture network model;
[0051] The target frame and the source frame are enhanced by using an adaptive image enhancement module to obtain a first source frame and a first target frame;
[0052] Reconstructing the nighttime depth map, the posture change, and the first source frame to obtain a reconstructed target frame, and supervising the reconstructed target frame and the first target frame using a first loss function;
[0053] The daytime depth map and the nighttime depth map are input into the adversarial neural network model, and the estimated distribution of the self-supervised nighttime depth map is consistent with the depth distribution under normal daytime lighting scenes, wherein the adversarial neural network uses the second loss function for supervised training.
[0054] Specifically, the nighttime self-supervised depth estimation architecture: The system is mainly composed of a depth estimation network Φ d and posture network Φ p However, due to the particularity of night scenes, relying solely on the above method may produce a large number of abnormal depth values. In order to solve this problem, the present invention pre-trains the depth estimation network Φ′ on the daytime data set. d , and then use the adversarial neural network Patch-GANΦ A To guide night training. The present invention utilizes Φ A The discriminator in Φ d Generated night depth map D t and Φ′ d Generated daytime depth map D d .
[0055] Given an RGB image I t , we can use the trainable network Φ d :D t =Φ d (I t ) to predict the depth map D t In order to achieve self-supervision, it is necessary to use the geometric relationship T t→s =Φ p (I t , I s ) from source frame I s Reconstruct target frame I t The reconstruction process involves using the posture network to obtain I s with I t The relative posture between t→s Then, D t Point p in t Projection to I s Point p in s superior:
[0056] p s ~KT t→s D t (p t )K -1 p t
[0057] Among them, ~ represents the homogeneous equivalence relationship, and K represents the camera intrinsic parameter.
[0058] Then, I s Through differentiable bilinear sampling operation
[0059]
[0060] Subsequently, the present invention combines the L1 regularization loss and the SSIM (structural similarity) loss into the photometric loss L P , the formula is:
[0061]
[0062] Among them, ||·||1 represents the L1 norm, and the parameter α is set to 0.85 in all experiments. In addition, the present invention uses edge-aware smoothing loss to reduce the noise and discontinuity of the depth map or surface normal vector, making the depth estimation result more accurate and stable. The expression is as follows:
[0063]
[0064] in, and are the gradients of the image in the horizontal and vertical directions respectively. Finally, the present invention trains the generator Φ by minimizing the loss function of GAN d and the discriminator Φ A , the generator loss function is expressed as:
[0065]
[0066] The discriminator loss function is expressed as:
[0067]
[0068] Among them, |I d | represents the number of daytime training images, |I t | represents the number of nighttime training images. It should be noted that I t and I d does not correspond one-to-one to daytime and nighttime images, so the generated depth map D d =Φ′ d (I d ) and D t=Φ d (I t ) No pairing is required.
[0069] Furthermore, the depth estimation network model includes: a depth encoder and a depth decoder;
[0070] A deep encoder for encoding a target frame;
[0071] The depth decoder is used to decode the encoded target frame to obtain a nighttime depth map.
[0072] Furthermore, the deep encoder includes: a convolutional layer, a number of downsampling layers, and a number of feature fusion modules;
[0073] Convolutional layer, used to refine feature extraction and further enhance the ability to capture local details;
[0074] The downsampling layer is used to reduce the spatial resolution of the feature map layer by layer, increase the receptive field, and retain the main scene structure information;
[0075] The feature fusion module is used to simultaneously realize the perception of local details and scene structures and integrate local and global image features.
[0076] Further, the deep decoder includes: a number of upsampling layers and an output module;
[0077] Upsampling layer, used for upsampling by bilinear interpolation;
[0078] The output module is used to output the upsampling result.
[0079] Specifically, this lightweight depth estimation network is implemented as follows Figure 2 As shown. The main module designs are as follows:
[0080] Lightweight depth estimation network: Subsequent reasoning only needs to use DepthNet. Therefore, designing a lightweight depth network can help the present invention effectively reduce the complexity of the model and the reasoning speed. The present invention designs a lightweight depth estimation network including feature fusion blocks and cross connections:
[0081] Deep encoder: Deep encoder. Using a shallower network can effectively reduce the complexity of the model, so the present invention adopts a four-level encoder. The convolution layer consists of a 3×3 convolution with a stride of 2 and two 3×3 convolutions with a stride of 1. The downsampling module is a 3×3 convolution with a stride of 2. Since the depth estimation task has high requirements on the perception of local details and scene structures, the present invention introduces a feature fusion block that can simultaneously realize the perception of local details and scene structures and integrate local and global image features. First, the feature fusion block uses a 3×3 expanded convolution to expand the receptive field to realize the extraction of local features. Assuming a feature map x with a dimension of H×W×C, this process is as follows Figure 2 As shown:
[0082] First: a dilated convolution (Dconv) is used to expand the receptive field to perceive local features in a larger range. The output of the dilated convolution will undergo a batch normalization (BN) operation to standardize the features. Then, the feature map is adjusted point by point. Finally, the processed features are residually connected to the input features to ensure the effective integration of the newly extracted features with the original information.
[0083] In the feature fusion block of the encoder, the above operations of extracting local features are repeated multiple times to gradually enhance the feature extraction capability. Typically, these operations are repeated 3, 3, and 9 times, respectively, from top to bottom.
[0084] The process can be expressed as:
[0085] L(x)=Pω2(Pω1(BN(Dconv(x))))+x
[0086] In the formula, Dconv(·) represents dilated convolution, BN represents batch normalization, Pω1(·) and Pω2(·) represent point-by-point operations for dimensionality expansion and dimensionality reduction, respectively. In the feature fusion block of the encoder, the operation of extracting local features L(·) is repeated N times, 3, 3, and 9 from top to bottom.
[0087] The present invention replaces the original self-attention with Cross-Covariance Attention (XCA) with lower computational complexity, so as to more effectively model the global context. Specifically, it includes: global feature modeling: first, the global context features are extracted through the XCA module; then, layer normalization (LN) is used to ensure the stability of feature distribution; the features are expanded and reduced in dimension through point-by-point operations, and finally the global feature representation is obtained. Integration of local and global features: In order to enhance the integration and propagation of local features and global features, the present invention introduces a cross-connection mechanism in the downsampling module and the feature fusion block. Through cross-connection, local detail features are retained while global feature modeling, ensuring the effective transmission of context information in the entire network. The final output of the feature fusion block: The output of the feature fusion block consists of two parts: the result of multiple extractions of local features and the result of global feature modeling. The final output merges the two through the feature concatenation operation (Concat) to form a richer feature expression. Such a design effectively improves the comprehensive perception of local details and global context while reducing the computational complexity.
[0088] The present invention records this operation as G(·), and the corresponding expression is:
[0089] G(x′)=Pω2(Pω1(LN(XcA(x′))))
[0090] Where x′ is the input feature and LN represents layer normalization. In order to enhance the integration and propagation of local and global features, the present invention uses cross connections in both the downsampling module and the feature fusion block. The final output y of the feature fusion block can be expressed as:
[0091] y=Concat[N·L(x),G(N·L(x))]
[0092] Deep Decoder: Deep decoder. To reduce the complexity of the model, this paper only uses a decoder consisting of convolutional layers. The upsampling module uses bilinear interpolation for upsampling and is then connected to the disparity head of each output. These outputs are and full resolution generation.
[0093] Furthermore, the channels of the target frame and the source frame are spliced, and the posture changes of the spliced target frame and the source frame are obtained according to the posture network model, including:
[0094] The target frame and the source frame channels are spliced, and the spliced target frame and source frame are input into the posture network model to obtain the relative posture between the target frame and the source frame.
[0095] Further, reconstructing the night depth map, the posture change and the first source frame includes: given an RGB image I t, we can use the trainable network Φ d :D t =Φ d (I t ) to predict the depth map D t In order to achieve self-supervision, it is necessary to use the geometric relationship T t→s =Φ p (I t , I s ) from source frame I s Reconstruct target frame I t The reconstruction process involves using the posture network to obtain I s with I t The relative posture between t→s Then, D t Point p in t Projection to I s Point p in s superior:
[0096] p s ~KT t→s D t (p t )K -1 p t
[0097] Among them, ~ represents the homogeneous equivalence relationship, and K represents the camera intrinsic parameter.
[0098] Then, I s The reconstructed target frame is obtained by the differentiable bilinear sampling operation s(·,·)=
[0099]
[0100] Subsequently, the reconstructed target frame and the target frame are supervised by using the photometric loss, including: the present invention combines the L1 regularization loss and the SSIM (structural similarity) loss into the photometric loss L P , the formula is:
[0101]
[0102] Among them, ||·||1 represents the L1 norm, and the parameter α is set to 0.85 in all experiments. In addition, the present invention uses edge-aware smoothing loss to reduce the noise and discontinuity of the depth map or surface normal vector, making the depth estimation result more accurate and stable. The expression is as follows:
[0103]
[0104] in and are the gradients of the image in the horizontal and vertical directions, respectively.
[0105] Further, enhancing the target frame and the source frame by using an adaptive image enhancement module, and obtaining the first source frame and the first target frame includes:
[0106] Divide the target frame and the source frame into blocks to obtain a number of image blocks;
[0107] Introduce contrast limit parameter to obtain the histogram of the image block;
[0108] The histogram is equalized, and the processed image blocks are smoothed by bilinear difference to obtain a first source frame and a first target frame.
[0109] Specifically, noise-constrained adaptive image enhancement: In nighttime images, the photometric consistency between the target frame and the source frame is often not maintained, accompanied by low light and high noise levels. Therefore, inspired by adaptive histogram equalization, this paper proposes a noise-constrained adaptive image enhancement (NCAIE) module. This avoids the use of an additional image enhancement network, which is consistent with the goal of lightweight model design.
[0110] The present invention firstly converts the target frame I t and source frame I s Divide into I×J small blocks. Target frame I t Each block in the partition is denoted as t i,j , source frame I s Each block in the partition is denoted as s i,j Assuming that each block has M pixels and N gray levels, the present invention calculates the histogram h of each small block i,j (n), at the same time, in order to control the noise amplification caused by excessive contrast, the present invention introduces a contrast limit parameter β, and obtains the clipped histogram h′ i,j (n), then, the present invention calculates the distribution function CDF of the cumulative histogram as follows:
[0111]
[0112] The above expression represents the histogram equalization operation of each small block. Finally, the image blocks after smoothing and splicing are used to obtain the enhanced target frame I′. t and source frame I′ s .
[0113] In the experiment of the present invention, the present invention also introduces the method of first passing through the image enhancement module and then reconstructing. The difference between this method and the present invention includes: the size of each small block I×J is set to 8×8, and the contrast limit parameter β is set to 4. The whole enhancement process can be expressed as:
[0114] I′ t =∈(I t), I′ s =∈(I s )
[0115] Where ∈ represents the mapping function of image enhancement. Therefore, I′ s The reconstructed RGB image I′ can be obtained t :
[0116]
[0117] Finally, the overall modification definition of the photometric loss function of the present invention is as follows:
[0118]
[0119] In the normal reconstruction process, of course, it can be completed without image enhancement. The present invention also deduces the reconstruction formula based on this.
[0120] Considering that in dark light scenes, the luminance consistency between the target frame and the source frame is often not maintained, the quality of reconstruction will be poor. The present invention designs this image enhancement module, which first enhances the image and then reconstructs it, which can alleviate some of this phenomenon.
[0121] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A nighttime depth estimation method based on a self-supervised framework, characterized in that: include: Obtaining a daytime data set, inputting the daytime data set into a pre-trained estimation network model, and obtaining a daytime depth map; Obtain a target frame and a source frame of a nighttime RGB image, input the target frame into a depth estimation network model, and obtain a nighttime depth map; Splicing the channels of the target frame and the source frame, and obtaining the posture changes of the spliced target frame and the source frame according to the posture network model; The target frame and the source frame are enhanced by using an adaptive image enhancement module to obtain a first source frame and a first target frame; Reconstructing the nighttime depth map, the posture change, and the first source frame to obtain a reconstructed target frame, and supervising the reconstructed target frame and the first target frame using a first loss function; The daytime depth map and the nighttime depth map are input into an adversarial neural network model, and the estimated distribution of the self-supervised nighttime depth map is consistent with the depth distribution under normal daytime lighting scenes, wherein the adversarial neural network uses a second loss function for supervised training.
2. A nighttime depth estimation method based on a self-supervised framework according to claim 1, characterized in that: The depth estimation network model includes: a depth encoder and a depth decoder; The depth encoder is used to encode the target frame; The depth decoder is used to decode the encoded target frame to obtain a nighttime depth map.
3. A nighttime depth estimation method based on a self-supervised framework according to claim 2, characterized in that: The deep encoder includes: a convolution layer, several downsampling layers, and several feature fusion modules; The convolution layer is used to refine feature extraction and further enhance the ability to capture local details; The downsampling layer is used to reduce the spatial resolution of the feature map layer by layer, increase the receptive field, and retain the main scene structure information; The feature fusion module is used to simultaneously realize the perception of local details and scene structures and integrate local and global image features.
4. A nighttime depth estimation method based on a self-supervised framework according to claim 3, characterized in that: The deep decoder comprises: a number of upsampling layers and an output module; The upsampling layer is used for upsampling by bilinear interpolation; The output module is used to output the upsampling result.
5. The method for nighttime depth estimation based on a self-supervised framework according to claim 1, characterized in that: The channels of the target frame and the source frame are spliced, and the posture changes of the spliced target frame and the source frame are obtained according to the posture network model, including: The target frame and the source frame channels are spliced, and the spliced target frame and the source frame are input into the posture network model to obtain the relative posture between the target frame and the source frame.
6. A nighttime depth estimation method based on a self-supervised framework according to claim 1, characterized in that: The target frame and the source frame are enhanced by using an adaptive image enhancement module to obtain a first source frame and a first target frame, comprising: Divide the target frame and the source frame into blocks to obtain a plurality of image blocks; Introducing a contrast limiting parameter to obtain a histogram of the image block; The histogram is equalized, and the image blocks after the splicing process are smoothed using bilinear difference to obtain the first source frame and the first target frame.
7. The method for nighttime depth estimation based on a self-supervised framework according to claim 1, characterized in that: The first loss function is: Among them, I' t is the enhanced target frame, is the reconstructed target frame, a is a parameter, and ||·||1 is the L1 norm.
Citation Information
Patent Citations
Multi-spectral image gradient fusion model establishment method and fusion method
CN116108889A
Night scene monocular image depth estimation method and device
CN117058438A
Monocular unsupervised depth estimation method based on contextual attention mechanism
US20210390723A1