Self-Supervised Depth Estimation Method and System for Low-Light Scenes Based on Structural Regularization
By introducing a structure regularization self-supervised depth estimation method in the feature space and output space, using pseudo-labels and wavelet decomposition of normal illuminated images, combined with local multi-scale consistency losses, the domain offset problem of depth estimation in low-light scenes is solved, and more accurate depth recovery is achieved.
Patent Information
- Application Number
- CN202410538821.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-04-30
AI Technical Summary
The prior art has domain offset problems in depth estimation in low-light scenarios, resulting in inaccurate depth output. The existing domain adaptive methods are insufficient in the feature output space and the deep output space, making it difficult to effectively alleviate the impact of low visibility and lighting changes.
The self-supervised depth estimation method based on structural regularization is adopted, and the feature space adversarial identification network and the output space adversarial identification network are used to train using pseudo-feature labels and pseudo-depth labels of normal illuminated images. Combined with wavelet decomposition and global structured guidance mask, high-frequency feature weights are increased, and local multi-scale consistency losses are used to improve the structure perception and detail recovery of depth estimation.
It improves the accuracy and detail recovery ability of low-light scene depth estimation, overcomes the domain offset problem, obtains clearer boundaries and more detailed depth structures, and enhances the perception of scene structure.
Smart Images

Figure CN118314186B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of depth estimation, and in particular, to a self-supervised depth estimation method and system for low-light scenes based on structural regularization. Background Art
[0002] There is generally a large "domain shift" between low-light images and normal-light images due to different lighting conditions. On the one hand, many areas in low-light scenes are insufficiently illuminated and have low visibility. On the other hand, there are moving light sources such as car headlights and point light sources such as street lights in low-light scenes. These seriously interfere with the network's perception of the image structure. The textureless areas and light-changing areas with low visibility will cause inconsistent problems between adjacent frames, affecting the reconstruction process, and then generating non-smooth depth outputs, which limits the performance of low-light depth estimation.
[0003] To alleviate these problems, existing researchers have converted this task into a domain adaptation problem, that is, using a discriminator (patchGAN) to distinguish the features generated by low-light images from the features generated by normal-light images until the features of day and night are indistinguishable, and finally using a normal-light model for processing.
[0004] The inventor has found that although the above method can obtain acceptable results, there are still limitations: due to the large "domain shift", it is very difficult to perfectly convert low-light features to normal-light features; and when only the depth output is constrained, the influence caused by low visibility and light changes cannot be completely eliminated. In addition, the existing solution is a direct application of the domain adaptation method in the feature output space and the depth output space, and the adaptability to the depth estimation task is insufficient. That is, how to perform structured guidance during the domain adaptation process is the key to solving the problem. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a self-supervised depth estimation method and system for low-light scenes based on structural regularization, which performs structured guidance in both the domain adaptation stage and the depth estimation stage. Using the normal-light features and depth maps generated by a pre-trained normal-light depth estimation network as pseudo-labels, a discriminator is used to identify the features and depth maps obtained by the low-light depth estimation network until the discriminator cannot distinguish the features and depth maps of normal light and low light, thereby alleviating the "domain shift" between normal-light and low-light images and obtaining high-quality low-light scene depth maps.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a self-supervised depth estimation method for low-light scenes based on structural regularization, including:
[0008] Obtain the low-light image to be estimated;
[0009] Input the low-light image to be estimated into the depth estimation model to obtain a depth estimation result;
[0010] Among them, the depth estimation model includes a feature space adversarial discriminator network and an output space adversarial discriminator network, which are obtained after training based on normal-light image samples and low-light image samples; pseudo-feature labels and pseudo-depth labels are obtained based on normal-light image samples; the feature space adversarial discriminator network is used to make it difficult to distinguish between the pseudo-feature labels and the features of the low-light samples, and the output space adversarial discriminator network is used to make it difficult to distinguish between the pseudo-depth labels and the depth maps of the low-light samples.
[0011] Preferably, the step of inputting the low-light image to be estimated into the depth estimation model to obtain a depth estimation result specifically includes:
[0012] Input the target frame of the low-light image to be estimated into the depth estimation model to obtain the depth map of the target frame of the low-light image to be estimated;
[0013] Input the target frame and the source frame of the low-light image to be estimated into the pose estimation network to obtain a pose matrix;
[0014] Perform depth estimation based on the depth maps and the pose matrix of the source frame and the target frame of the low-light image to be estimated;
[0015] Complete self-supervised depth estimation using photometric loss and edge-aware smoothing loss.
[0016] Preferably, the feature space adversarial discriminator network is used to make it difficult to distinguish between the pseudo-feature labels and the features of the low-light image samples, specifically including:
[0017] Input the normal-light image samples and the low-light image samples into the encoder-decoder structure respectively to obtain normal-light sample features and low-light sample features;
[0018] Decouple the normal-light sample features and the low-light sample features into normal-light high-frequency features and low-light high-frequency features respectively; the normal-light high-frequency features are the pseudo-feature labels; the process of obtaining the low-light high-frequency features based on the low-light image samples is defined as the first generator;
[0019] Input the pseudo-feature labels and the low-light high-frequency features into the feature space discriminator, and the discrimination result is used to indicate the degree of difference between the pseudo-feature labels and the low-light high-frequency features;
[0020] Based on the discrimination result, use a loss function to alternately train the first generator and the feature space discriminator until the pseudo-feature labels and the low-light high-frequency features are indistinguishable.
[0021] Preferably, the decoupling of the normal illumination sample features and the low illumination sample features into normal illumination high-frequency features and low illumination high-frequency features specifically includes:
[0022] Using wavelet transform to decouple the normal illumination sample features into multiple normal illumination high-frequency components and normal illumination low-frequency components, and fusing the multiple normal illumination high-frequency components into normal illumination high-frequency features;
[0023] Using wavelet transform to decouple the low illumination sample features into multiple low illumination high-frequency components and low illumination low-frequency components, and fusing the multiple low illumination high-frequency components into low illumination high-frequency features.
[0024] Preferably, the output space adversarial discriminative network is used to make the pseudo-depth label and the low illumination sample depth map indistinguishable, specifically including:
[0025] Mapping the normal illumination sample features to the first normal illumination sample depth map;
[0026] Mapping the low illumination sample features to the first low illumination sample depth map;
[0027] Adding a structure-guided mask to the first normal illumination sample depth map and the first low illumination sample depth map respectively to obtain the second normal illumination sample depth map and the second low illumination sample depth map; the second normal illumination sample depth map is the pseudo-depth label; the process of obtaining the second low illumination sample depth map based on the low illumination sample features is defined as the second generator;
[0028] Inputting the pseudo-depth label and the second low illumination sample depth map into the output space discriminator, and the discrimination result is used to indicate the difference degree between the pseudo-depth label and the second low illumination sample depth map;
[0029] Based on the discrimination result, alternately training the second generator and the output space discriminator using a loss function until the pseudo-depth label and the second low illumination sample depth map are indistinguishable.
[0030] Preferably, the structure-guided mask is used to divide the first normal illumination sample depth map and the first low illumination sample depth map into multiple regions according to the scene features.
[0031] Preferably, the inputting of the to-be-estimated low illumination image into the depth estimation model to obtain a depth estimation result further includes:
[0032] Obtaining depth maps of multiple consecutive to-be-estimated low illumination images, taking two anchor points in the region with rich detailed texture of the depth map according to the scene features, and cropping the region of interest with a side length of one-eighth of the height of the depth map resolution;
[0033] Upsampling the region of interest of the previous layer to the size of the region of interest of the next layer;
[0034] The corresponding regions of interest between depth maps of different resolutions are constrained by using local multi-scale consistency loss.
[0035] In a second aspect, the present invention provides a self-supervised depth estimation system for low-light scenes based on structural regularization, including:
[0036] An acquisition module for acquiring the low-light image to be estimated;
[0037] A processing module for inputting the low-light image to be estimated into a depth estimation model to obtain a depth estimation result;
[0038] The depth estimation model includes a feature space adversarial discriminative network and an output space adversarial discriminative network, which are obtained after training based on normal-light image samples and low-light image samples; pseudo feature labels and pseudo depth labels are obtained based on normal-light image samples; the feature space adversarial discriminative network is used to make the pseudo feature labels indistinguishable from the low-light sample features, and the output space adversarial discriminative network is used to make the pseudo depth labels indistinguishable from the low-light sample depth maps.
[0039] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the self-supervised depth estimation method described in the first aspect are implemented.
[0040] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the self-supervised depth estimation method described in the first aspect are implemented.
[0041] Compared with the prior art, the beneficial effects of the present disclosure are:
[0042] (1) The present application proposes a domain adaptation module that fuses the feature space and the depth space (Feature and Output Spaces Domain Adaption, FODA), that is, adaptive constraints are adopted in both the feature space and the output space. This module simultaneously uses the adaptive constraints of the feature map space and the depth map space to better achieve the depth recovery of low-light scene images; moreover, wavelet decomposition is used to separate the features into high-frequency features and low-frequency features, and by increasing the weight of the high-frequency components containing more structured details, the obtained depth estimation result has a clearer boundary, so that the domain adaptation process pays more attention to structured information. It improves the perception of the scene structure in the depth estimation process, thus ensuring the restoration of the detailed structure of the low-light image depth map during the learning process of the depth estimation network, and overcoming the problem of inaccurate depth estimation of low-light images caused by the domain shift problem.
[0043] (2) The global Structure-Guided Mask (SGM) provided by this application rationally associates depth values with pixel positions using prior structure information to further refine the depth, promotes the smoothness of the depth within each region, and enhances the global structure information.
[0044] (3) The local multi-scale consistency loss (Crop Multi-Scale Consistency Loss, Lc) between adjacent depth outputs provided by this application focuses on regions with rich detailed textures in the scene, ensuring that the network pays more attention to the recovery of local detailed structures during the depth estimation stage.
[0045] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become apparent from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of this disclosure. The schematic embodiments and descriptions thereof of this disclosure are used to explain this disclosure and do not constitute a limitation to this disclosure.
[0047] Figure 1 It is a flowchart of a self-supervised depth estimation method for low-light scenes based on structure regularization provided by this disclosure;
[0048] Figure 2 It is a structural diagram of the depth estimation module provided by this disclosure;
[0049] Figure 3 It is an effect diagram of the depth estimation result obtained by using the self-supervised depth estimation method of this disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0051] Embodiment 1
[0052] As Figure 1 shown, this embodiment discloses a self-supervised depth estimation method for low-light scenes based on structure regularization, including the following steps:
[0053] S1: Obtain the low-light image to be estimated;
[0054] S2: Input the low-light image to be estimated into the depth estimation model to obtain the depth estimation result;
[0055] Among them, the depth estimation model includes a feature space adversarial discriminative network and an output space adversarial discriminative network, which are obtained after training based on normal illumination image samples and low-light image samples; pseudo feature labels and pseudo depth labels are obtained based on normal illumination image samples; the feature space adversarial discriminative network is used to make the pseudo feature labels indistinguishable from the low-light sample features, and the output space adversarial discriminative network is used to make the pseudo depth labels indistinguishable from the low-light sample depth maps.
[0056] In this embodiment, in order to alleviate problems such as non-uniform illumination and motion blur in the specific field of low-light images, a self-supervised monocular depth estimation network SRNSD based on structural regularization is proposed, which uses the features and depth maps of normal illumination images to obtain the depth maps of low-light images. This embodiment takes unpaired day-night images as input, which is more flexible.
[0057] Specifically, first, a depth estimation network for normal illumination images (e.g., Monodepth2) is pre-trained to serve as pseudo feature labels and pseudo depth labels. For the low-light image sequence, a depth estimation network such as Monodepth2 is used for depth estimation.
[0058] During the training process, the pre-trained normal illumination depth estimation network is used to generate normal illumination features and depth maps as pseudo labels, and a discriminator is used to identify the features and depth maps obtained by the low-light depth estimation network and the normal illumination depth estimation network until the network can no longer distinguish the features and depth maps of normal illumination and low-light, so as to alleviate the "domain shift" between normal illumination and low-light images and obtain high-quality low-light scene depth maps. During this process, three structural regularization operations are adopted to ensure the restoration of the detailed structure of the low-light image depth map during the learning process: a structured domain adaptation module based on the feature space, a structured domain adaptation module based on the output space, and a local multi-scale consistency loss. Specifically:
[0059] (1) Self-supervised monocular depth estimation module
[0060] For the low-light image sequence, a self-supervised monocular depth estimation method based on SfM is used, and the geometric consistency between two frames is used for constraint.
[0061] This module mainly consists of two parts: a depth estimation network and a pose estimation network. The depth estimation network is based on the U-Net architecture, which is an encoder and decoder with skip connections. Among them, the encoder adopts the architecture of ResNet-50, removes the fully connected layer, and replaces the max pooling with strided convolution. The decoder contains 5 convolutional layers of 3×3, and uses nearest interpolation for upsampling. Sigmoid is used for the final output, and Leaky Relu is used for other outputs. The pose estimation network adopts the ResNet-18 architecture, and each input sample corresponds to an output 6D vector.
[0062] The input of the depth estimation network is a low-light frame video sequence containing the target frame (current frame) and the source frame (previous frame or next frame). Among them, the input of the depth estimation network is the target frame used to estimate the depth map D of the target frame n ; the input of the pose estimation network is the target frame and the source frame used to estimate the camera pose T between the target frame and the source frame s→t . Through the above variables, the relationship between the target frame and the source frame can be established using the projection operation:
[0063]
[0064] where K represents the camera internal parameters, and proj() is the coordinate projection operation between the source frame and the target frame.
[0065] The model learning is based on the above projection process, that is, reconstructing the target frame from the source frame. The goal is to reduce the reconstruction error by optimizing the depth estimation network and the pose estimation network to produce more accurate outputs. The L1 loss and the SSIM loss are used as photometric errors to measure the difference between the original image and the reconstructed image :
[0066]
[0067] where α is a hyperparameter, set to 0.85.
[0068] In addition, in order to alleviate the depth blur problem caused by incorrect depth estimation, an edge-aware smoothing loss is used to enhance depth smoothing, defined as:
[0069]
[0070] where and are the gradients of the image in the horizontal and vertical directions respectively.
[0071] Reconstruct the current frame of low-light image based on this module, indirectly supervise the low-light depth map, and achieve self-supervised monocular depth estimation.
[0072] (2) Structured domain adaptation module based on feature space
[0073] Input the normal-light and low-light images into their respective depth estimation networks, and obtain the normal-light features at the decoder stage of the depth estimation network through formulas (4) and (5) and low-light features
[0074]
[0075]
[0076] Among them, and represent the depth estimation network for the normal-light scene and the depth estimation network for the low-light scene respectively. I d and I n represent the input normal-light image sequence and low-light image sequence respectively.
[0077] Structural information plays an important role in the depth estimation task. Therefore, considering the importance of structural information in the depth estimation task, this module constructs a feature space adversarial discrimination network during the feature space domain adaptation process. Among them, the ultimate goal of the discriminator is not only to make the features of normal-light images and low-light images indistinguishable, but also to make the boundary information of normal-light images and low-light images indistinguishable. Considering that high-frequency information pays more attention to boundary information and structural information, this module introduces wavelet transform (Wavelet Transform, DWT) to enhance the discriminator's perception ability of structural information.
[0078] Specifically, use wavelet transform to decouple the features of normal-light images and low-light images into high-frequency components and low-frequency components respectively. During the feature space domain adaptation adversarial discrimination process, by increasing the weight of the high-frequency components containing more structural details, the depth estimation result of the obtained low-light image has a clearer boundary.
[0079] The input of wavelet transform is the normal-light feature f d and the low-light feature f n , and the output is the high-frequency component d (or f n ) corresponding to the feature f (or ) and the low-frequency component f l d (or f l n)。After that, the high-frequency components are fused into (or ) through a merging operation. The formula for this process is defined as follows:
[0080]
[0081]
[0082]
[0083]
[0084] where Φ DWT represents the wavelet transform operation, and the operation cat(·,·) represents the concatenation operation along the channel dimension.
[0085] This module constructs an adversarial discriminative network based on the feature space in the feature space decoupled into high-frequency feature components and low-frequency feature components. The network inputs are the decoupled feature components and (or f l n and f l d ). Among them, is used as the pseudo-feature label. The depth estimation network of the weak illumination scene is used as the generator. Specifically, in the feature space, the process of inputting the weak illumination image sample into the encoder-decoder structure and obtaining is defined as the first generator. Patch-GAN is used as the feature space discriminator. The pseudo-feature label and the weak illumination high-frequency feature are input into the feature space discriminator, and the obtained discrimination result is used to indicate the difference degree between the pseudo-feature label and the weak illumination high-frequency feature. Based on this discrimination result, the first generator and the feature space discriminator are alternately trained using the generative loss and the adversarial loss until the pseudo-feature label and the weak illumination high-frequency feature are indistinguishable. The generative loss and the adversarial loss can be expressed as:
[0086]
[0087]
[0088]
[0089]
[0090] where |I d | and |I n | are the numbers of normal illumination and weak illumination training images, is a low-frequency feature component discriminator, whose goal is to make the low-frequency components f l d and f l n difficult to distinguish, while is a high-frequency feature component discriminator, whose goal is to make the high-frequency components and difficult to distinguish.
[0091] It should be noted that, in principle, the features f d and f n between corresponding layers obtained by the depth estimation network decoder can all participate in this adversarial discrimination network. Considering the comprehensive number of parameters and network performance, in this embodiment, the last two layers of features are preferably used as the input of the feature space adversarial discrimination network, that is, the input is and In addition, it should be noted that the features used in this module are the decoder features of the depth estimation network. In principle, the encoder feature part can also be used for network learning. However, considering this intensive task of depth estimation, improving the feedback at the output end where semantic expressions have been learned is better than the input end. Therefore, the decoder features are finally selected for use.
[0092] As described above, this module sets a higher weight for the high-frequency feature components, and the loss is defined as:
[0093]
[0094]
[0095] where the hyperparameter ξ is set to 0.7.
[0096] (3) Structured domain adaptation module based on the output space
[0097] Similarly, for domain adaptation in the depth map output space, attention should also be paid to adding the guidance of structured information in the adversarial discrimination process. Experiments have found that directly adopting the wavelet transform scheme proposed in the "structured domain adaptation module based on the feature space" in the domain adaptation of the depth map output space has poor effects. The reason for the analysis may be that the high-frequency components of the depth map contain too little information, and the information contained in the low-frequency components is also quite important. In order to make the discrimination process based on the depth map space pay more attention to the structural information in the depth output space, this module designs a simple and effective structured mask M p , which brings an efficient and effective prior for better weak-light depth map estimation. The specific details of M p will be described in detail below. In this module, the fusion process of the depth map and the structured mask M p and the design of the adversarial discrimination network based on the depth output space are mainly described.
[0098] First, the feature f obtained by the decoder of the depth estimation network and is mapped to obtain the first normal illumination sample depth map D d and the first low-light sample depth map D n : d n
[0099]
[0100] D d = sigmoid(f d ) (16)
[0101] D n = sigmoid(f n ) (17) where
[0102] K = 5. It should be noted that the low-light images are self-supervised trained through the self-supervised module of the "self-supervised monocular depth estimation module". Subsequently, this module builds an adversarial discriminative network based on the depth output spatial domain on the basis of the depth image merged with the structured mask. The input of the adversarial discriminative network is the second normal illumination sample depth map cat(D d , M p ) and the second low-light sample depth map cat(D n , M p ) of the depth image merged with the structured mask. Among them, cat(D d , M p ) is used as the pseudo-depth label. Similar to the adversarial discriminative network based on the feature output space, this module uses the depth estimation network of the low-light scene as the generator. Specifically, in the output space, the operation of mapping the low-light sample features to the first low-light sample depth map and adding a structure-guided mask to generate the second low-light sample depth map is defined as the second generator. Use Patch-GAN as the output space discriminator, and input the pseudo-depth label cat(D d , M p ) and the second low-light sample depth map cat(D n , M p ) into the output space discriminator, and the obtained discrimination result is used to indicate the difference degree between the pseudo-depth label and the second low-light sample depth map. Based on this discrimination result, the second generator and the output space discriminator are alternately trained using the generative loss and the adversarial loss until the pseudo-depth label and the second low-light sample depth map are indistinguishable. The generative loss and the adversarial loss of the adversarial discriminative network based on the depth output space can be expressed as:
[0103]
[0104]
[0105] Among them, |I d | and |I n | are the numbers of normal illumination and low illumination training images, is the depth output space discriminator, whose goal is to make cat(D d , M p ) and cat(D n , M p ) difficult to distinguish. It is worth mentioning that the structured mask scheme can also be applied to the adversarial discriminative network based on the feature output space at the theoretical level, but it has little effect in practice. This is mainly because the number of channels of the features is too large, and the structured mask M p that only merges a single layer of channels has little impact on the network.
[0106] (4) Global structured guidance mask
[0107] The depth of a pixel is closely related to its position in the image space, but the prior relationship used in the existing technology is not completely reasonable. It encodes the two-dimensional coordinates of each pixel into an image of the coordinates along the x-axis, that is, it represents that the depth values of the pixels from left to right in the image are from large to small or from small to large, which is inconsistent with human perception.
[0108] This module designs a simple but more effective global structure regularization mask M p to describe this relationship. From the results, it effectively alleviates the "domain shift" between low illumination images and normal illumination images, and helps to obtain better depth estimation results.
[0109] Specifically, based on the principle of linear perspective, this module roughly divides the autonomous driving scene into four regions. The structure guidance mask divides the image into four regions, that is, the buildings on the left and right sides, the sky above, and the road below. The black dots represent the "vanishing points". On this basis, a structured guidance attention map is encoded according to the two-dimensional coordinate values of the image, which is one of the input values of the discriminator, as shown in Formulas (18) and (19). M p is normalized in the range of 0 to 1 in the four regions.
[0110] In addition, the globally structured regularization mask (SGM) designed in this module can be well generalized to different autonomous driving datasets, with good generalization ability. SGM overlaps with the images of KITTI, Oxford RobotCar, and nuScenes datasets. It can not only well describe the regions of these images, that is, the left and right parts are buildings and obstacles, the upper part is the sky, and the lower part is the road, but also explain the correspondence between the depth values and pixel positions among different pixels in each region.
[0111] (5) Local multi-scale consistency loss
[0112] To further refine the depth map and recover more detailed structural information, the cropped multi-scale consistency loss is proposed by using the consistency between depth maps of different resolutions obtained at the decoder end of the depth estimation network for weakly illuminated images. Specifically, based on the observation of a large amount of autonomous driving data, instances with rich detailed textures (such as parked cars, signs, tree trunks, etc.) are mainly distributed around the line segments l1 and l2 of the structure-guided mask. Therefore, to make the network pay more attention to the regions with rich detailed textures, this module takes two anchor points on l1 and l2 of the multi-resolution depth maps obtained at the decoder end of the depth estimation network for weakly illuminated images, and cuts out a square representing the "region of interest" with a side length of one-eighth of the height of the depth map resolution, that is:
[0113]
[0114] where i and j represent the jth "region of interest" cropped from the ith depth map, i = {1, 2, 3, 4, 5}, j = {1, 2}. Then, the "region of interest" of the previous layer is upsampled to the size of the "region of interest" of the next layer to facilitate the subsequent calculation of the local multi-scale consistency loss, as shown in the following formula:
[0115]
[0116] Finally, the local multi-scale consistency loss proposed by this module is used to constrain the corresponding "regions of interest" between depth maps of different resolutions to better focus on local details. This part of the loss can be expressed as:
[0117]
[0118] where L pe is the photometric consistency loss.
[0119] (6) Loss function
[0120] The final loss consists of the photometric consistency loss of the weakly illuminated scene self-supervised monocular depth estimation module, the edge-aware smoothing loss, the generative loss and adversarial loss of the adversarial discrimination module based on the fusion of the feature output space and the depth output space. The loss of the weakly illuminated self-supervised monocular depth estimation network (SRNSD) based on structural regularization can be expressed as:
[0121]
[0122] where β, γ1, γ2, δ1 and δ2 are hyperparameters. According to experience, in this embodiment, β is set to 1.0e -3 and γ1, γ2, δ1 and δ2 are set to 2.5e -4 .
[0123] On the surface of the road with photometric changes (such as the area illuminated by vehicle headlights), the depth is smoother and not affected by photometric changes. On the large flat surfaces on both sides (such as the walls of buildings on both sides), the depth is also smoother and the boundaries are clearer. In areas with rich structured textures (such as cars by the roadside), clearer detailed textures and accurate boundaries are obtained.
[0124] In this embodiment, by designing an adversarial discrimination network to restore the texture features of the weakly illuminated image based on the normally illuminated image, the "domain shift" between the normally illuminated and weakly illuminated images is alleviated, and a high-quality depth map of the weakly illuminated scene is obtained.
[0125] Embodiment 2
[0126] This embodiment provides a weakly illuminated scene self-supervised depth estimation system based on structural regularization, including:
[0127] An acquisition module for acquiring the weakly illuminated image to be estimated;
[0128] A processing module for inputting the weakly illuminated image to be estimated into the depth estimation model to obtain a depth estimation result;
[0129] where the depth estimation model includes a feature space adversarial discrimination network and an output space adversarial discrimination network, which are obtained after training based on normally illuminated image samples and weakly illuminated image samples; pseudo feature labels and pseudo depth labels are obtained based on the normally illuminated image samples; the feature space adversarial discrimination network is used to make the pseudo feature labels and the weakly illuminated sample features indistinguishable, and the output space adversarial discrimination network is used to make the pseudo depth labels and the weakly illuminated sample depth maps indistinguishable.
[0130] Embodiment 3
[0131] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the self-supervised depth estimation method described in the above-mentioned Embodiment 1 are implemented.
[0132] Embodiment 4
[0133] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the self-supervised depth estimation method described in the above-mentioned Embodiment 1 are implemented.
[0134] The steps or modules involved in the above Embodiments 2 to 4 correspond to those in Embodiment 1. For the specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0135] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and changes can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A self-supervised depth estimation method for low-light scenes based on structural regularization, characterized in that Including: Obtain a low-light image to be estimated; Input the low-light image to be estimated into a depth estimation model to obtain a depth estimation result; Wherein, the depth estimation model includes a feature space adversarial discriminator network and an output space adversarial discriminator network, which are obtained after training based on normal light image samples and low-light image samples; pseudo-feature labels and pseudo-depth labels are obtained based on normal light image samples; the feature space adversarial discriminator network is used to make it difficult to distinguish between pseudo-feature labels and low-light sample features, and the output space adversarial discriminator network is used to make it difficult to distinguish between pseudo-depth labels and low-light sample depth maps.
2. The self-supervised depth estimation method according to claim 1, wherein The step of inputting the low-light image to be estimated into the depth estimation model to obtain a depth estimation result specifically includes: Input the target frame of the low-light image to be estimated into the depth estimation model to obtain the depth map of the target frame of the low-light image to be estimated; Input the target frame and source frame of the low-light image to be estimated into a pose estimation network to obtain a pose matrix; Perform depth estimation based on the depth maps and pose matrix of the source frame and target frame of the low-light image to be estimated; Complete self-supervised depth estimation using photometric loss and edge-aware smoothing loss.
3. The self-supervised depth estimation method according to claim 1, wherein The feature space adversarial discriminator network is used to make it difficult to distinguish between pseudo-feature labels and low-light image sample features, specifically including: Input the normal light image sample and the low-light image sample into an encoder-decoder structure respectively to obtain normal light sample features and low-light sample features; Decouple the normal light sample features and low-light sample features into normal light high-frequency features and low-light high-frequency features respectively; the normal light high-frequency features are the pseudo-feature labels; the process of obtaining low-light high-frequency features based on the low-light image sample is defined as the first generator; Input the pseudo-feature labels and low-light high-frequency features into a feature space discriminator, and the discrimination result is used to indicate the degree of difference between the pseudo-feature labels and the low-light high-frequency features; Based on the discrimination result, alternately train the first generator and the feature space discriminator using a loss function until the pseudo-feature labels and the low-light high-frequency features are indistinguishable.
4. The self-supervised depth estimation method according to claim 3, characterized in that, The step of decoupling the normal light sample features and low-light sample features into normal light high-frequency features and low-light high-frequency features respectively specifically includes: Use wavelet transform to decouple the normal light sample features into multiple normal light high-frequency components and normal light low-frequency components, and fuse the multiple normal light high-frequency components into normal light high-frequency features; Use wavelet transform to decouple the low-light sample features into multiple low-light high-frequency components and low-light low-frequency components, and fuse the multiple low-light high-frequency components into low-light high-frequency features.
5. The self-supervised depth estimation method according to claim 3, wherein The output space adversarial discriminator network is used to make it difficult to distinguish between pseudo-depth labels and low-light sample depth maps, specifically including: Map the normal light sample features to a first normal light sample depth map; Map the low-light sample features to a first low-light sample depth map; Add a structure-guided mask to the first normal illumination sample depth map and the first low illumination sample depth map respectively to obtain a second normal illumination sample depth map and a second low illumination sample depth map; the second normal illumination sample depth map is a pseudo-depth label; define the process of obtaining the second low illumination sample depth map based on the low illumination sample features as the second generator; Input the pseudo-depth label and the second low illumination sample depth map into the output space discriminator, and the discrimination result is used to indicate the degree of difference between the pseudo-depth label and the second low illumination sample depth map; Based on the discrimination result, alternately train the second generator and the output space discriminator using a loss function until the pseudo-depth label and the second low illumination sample depth map are indistinguishable.
6. The self-supervised depth estimation method according to claim 5, wherein The structure-guided mask is used to divide the first normal illumination sample depth map and the first low illumination sample depth map into multiple regions according to the scene features.
7. The self-supervised depth estimation method according to claim 2, wherein, The step of inputting the to-be-estimated low illumination image into the depth estimation model to obtain a depth estimation result further includes: Obtain depth maps of multiple consecutive to-be-estimated low illumination images, take two anchor points in the region with rich detailed texture of the depth map according to the scene features, and crop the region of interest with a side length of one-eighth of the height of the depth map resolution; Upsample the region of interest of the previous layer to the size of the region of interest of the next layer; Use the local multi-scale consistency loss to constrain the corresponding regions of interest between depth maps of different resolutions.
8. A self-supervised depth estimation system for low-light scenes based on structural regularization, characterized in that It includes: An acquisition module for acquiring the to-be-estimated low illumination image; A processing module for inputting the to-be-estimated low illumination image into the depth estimation model to obtain a depth estimation result; Among them, the depth estimation model includes a feature space adversarial discriminator network and an output space adversarial discriminator network, which are obtained after training based on normal illumination image samples and low illumination image samples; pseudo-feature labels and pseudo-depth labels are obtained based on normal illumination image samples; the feature space adversarial discriminator network is used to make the pseudo-feature labels and the low illumination sample features indistinguishable, and the output space adversarial discriminator network is used to make the pseudo-depth labels and the low illumination sample depth maps indistinguishable.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the self-supervised depth estimation method according to any one of claims 1-7.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the self-supervised depth estimation method according to any one of claims 1-7.
Citation Information
Patent Citations
Underwater binocular depth estimation method based on unsupervised adaptive network
CN114299130A
Night scene monocular image depth estimation method and device
CN117058438A