Stable adversarial training method for generalization self-supervision monocular depth estimation

Through the stable adversarial training method, the noise generator and scalable deep network are used to solve the problems of insufficient generalization and training instability of the monocular depth estimation model in unseen scenes, achieving higher generalization ability and stability, and is suitable for depth estimation in complex environments.

CN120409611APending Publication Date: 2025-08-01HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510344771.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing monocular depth estimation model lacks generalization ability when facing unseen scenarios, and the training instability and gradient conflict caused by adversarial data enhancement are serious, affecting its application in the real world.

Method used

The stable adversarial training method is adopted to adaptively generate image noise through the noise generator network, adjust the long-hop connection coefficient, design a scalable depth estimation network, and propose a gradient conflict repair strategy to alleviate the double optimization conflict during the training process.

Benefits of technology

It improves the generalization performance of the model in unknown scenarios, realizes a more stable training process and higher depth estimation accuracy, and is suitable for a variety of complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409611A_ABST
    Figure CN120409611A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of depth estimation, in particular to a stable adversarial training method for generalization self-supervision monocular depth estimation. The method comprises the following steps: step 1, based on an input clean image, carrying out noise addition through a noise generator network; step 2, designing a scalable depth estimation network for a mixed clean image and a noisy image, and reducing a reconstruction error by adjusting a long-hop connection coefficient; and 3, proposing a gradient conflict repair strategy, and incrementally applying antagonism enhancement from a plurality of iterations. According to the technical scheme, the existing multiple models can have the out-of-distribution generalization ability, stable training of the models can be guaranteed, and the generalization performance of depth estimation and the practicability facing the ground are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of depth estimation, and particularly relates to a stable adversarial training method for generalizable self-supervised monocular depth estimation. Background Art

[0002] Monocular depth estimation plays an important role in various 3D perception fields (such as robot navigation, autonomous driving, and 3D reconstruction). However, due to the dynamic nature of the real world, even a tiny perturbation in the environment may cause a significant domain shift in visual observations, making it difficult for a trained model to generalize to unseen scenarios, thus limiting its application in the physical world. To improve the generalization ability, some studies have utilized data augmentation methods to generate synthetic data and diversify the training environment, achieving a considerable performance improvement.

[0003] Currently, a series of studies have explored how to improve the domain generalization ability of monocular depth estimation (abbreviated as MDE) by using data augmentation or synthetic data generation. For example, ADDS-DepthNet uses a generative adversarial network (abbreviated as GAN) to generate night-time image pairs from daytime images, training the network to be effective in both daytime and night-time domains. This algorithm uses day-depth estimation as pseudo-supervision for night scenes, restricting the night-depth estimation to the accuracy of day-depth estimation. In addition, ADDS-DepthNet focuses on reconstructing the night-time images generated by the GAN, which produces a reconstruction target with high variance and is not conducive to the self-supervised MDE algorithm. ITDFA adopts a fixed-depth decoder and adjusts the encoder for each domain to enforce cross-domain feature consistency, but this also limits the performance. EPC-depth proposes an offline data augmentation method. By exploiting the correspondence between un-augmented and augmented data, a pseudo-supervised loss is introduced for depth and pose estimation. However, previous methods limit the ability to handle environments with predefined variations and require significant modifications to the network architecture to obtain accurate depth estimation, resulting in severe domain bias when applied to real-world scenarios.

[0004] However, most existing methods select some specific data augmentation schemes for specific scenarios and show poor generalization performance in environments that are quite different from the augmented images. Compared with the above offline data augmentation, adversarial data augmentation (ADA) makes no assumptions about the target distribution and synchronously optimizes the adversarial generator during the training phase, providing a promising solution. However, self-supervised algorithms are quite sensitive to this type of data augmentation and cannot directly benefit from adversarial data augmentation. There are mainly two factors contributing to this phenomenon: (1) the inherent sensitivity of long skip connections (LSCs) in UNet-like deep networks. The existence of these shortcut connections will amplify the adversarial gradients, which will lead to serious training instability and collapse when combined with pixel-level adversarial data augmentation; (2) the dual optimization conflict caused by over-regularization. Since adversarial augmented data are usually the worst-case training examples, they tend to cause over-regularization to the model training, resulting in the optimization gradient being opposite to that of the original data, thus causing serious gradient conflicts and further reducing the convergence performance. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a stable adversarial training method for generalizable self-supervised monocular depth estimation. Without the need for offline data augmentation of images, it uses adversarial training to adaptively generate image noise during the training process to improve the generalization performance of self-supervised monocular depth estimation.

[0006] To achieve the above objective, the following technical solutions are adopted:

[0007] A stable adversarial training method for generalizable self-supervised monocular depth estimation includes the following steps:

[0008] Step 1: Based on the input clean image, add noise through a noise generator network;

[0009] Step 2: For the mixed clean images and noisy images, design a scalable depth estimation network to reduce the reconstruction error by adjusting the long skip connection coefficient;

[0010] Step 3: Propose a gradient conflict repair strategy to incrementally apply adversarial augmentation from multiple iterations to alleviate the dual optimization conflict faced when directly applying adversarial data augmentation.

[0011] Further, in Step 1, the adding noise based on the input clean image through a noise generator network is specifically as follows:

[0012] Given the clean image I at time t t , use the noise generation network f φ to add noise to the input image:

[0013]

[0014] In the formula represents the input of the noisy image, where i t represents the clean image at time t, and δ represents the image noise added by the noise generator network.

[0015] Furthermore, in step 2, for the mixed clean image and the noisy image, a scalable depth estimation network is designed to reduce the reconstruction error by adjusting the long skip connection coefficient, which specifically includes the following steps:[[]]

[0016] Step 2.1: Use the mixed clean image and the noisy image as the input image, and obtain the corresponding depth map D t through the depth prediction network; combine the camera internal parameter K, and reconstruct the current image through the images at the previous and next moments; calculate the photometric error function pe;

[0017] Step 2.2: Define the optimization objective L p of the depth prediction network and the objective function L AD of the noise generation network;

[0018] Step 2.3: Propose a scaled depth network to reduce the reconstruction error by adjusting the long skip connection coefficient κi.

[0019] Furthermore, in step 2.1, using the mixed clean image and the noisy image as the input image, the corresponding depth map D t is obtained through the depth prediction network; combine the camera internal parameter K, and reconstruct the current image through the images at the previous and next moments; calculate the photometric error function pe, specifically as follows:[[]]

[0020] Define I′ t ∈{I t-1 ,I t+1} as the video frame images at times t + 1 and t - 1. Given I t and I′ t as the input, the camera pose T t→t′ is obtained through the pose prediction network;

[0021] Combine the camera internal parameter K, and reconstruct the current image through the images at the previous and next moments:[[]]

[0022] I t ' →t =I' t <proj(D t ,T t→t ',K)> (2)[[]]

[0023] In the formula, I t′→t is the reconstructed image when the clean image is used as the input, and I' t∈{I t-1 ,I t+1} is the video frame image at time t+1 and t-1, proj(·) represents the 2D coordinate corresponding to the generated depth map, and <> represents the sampling operation, D t represents the depth map, T t→t′ represents the camera pose, K represents the camera intrinsic parameter;

[0024] The photometric error function pe is calculated as follows:

[0025]

[0026] Where pe is the photometric error function, α is the weighting coefficient, I a , I b Represents two random input images, SSIM represents structural similarity; ||I a -I b ||1 means calculating the L1 loss between two random input images, and the L1 loss is the mean absolute error.

[0027] Furthermore, in step 2.2, the optimization target L of the depth prediction network is defined p and the objective function L of the noise generation network AD , specifically:

[0028]

[0029] Where, L p is the optimization target of the depth prediction network, I t′→t is the reconstructed image when the clean image is used as input, and is the reconstructed image when the data-enhanced image is used as input, and pe is the photometric error function;

[0030]

[0031] Where, L AD represents the objective function of the noise generation network, The reconstructed image when the data augmented image is used as input.

[0032] Furthermore, in step 2.3, the proposed scaling depth network reduces the reconstruction error by adjusting the long jump connection coefficient κi, specifically:

[0033] DepthNet(x)=f0(x) (6)

[0034]

[0035] Where x is the input, and the depth estimation network DepthNet represents f i (x)(i≥1), ai and b i (i≥1) are the trainable parameters of the i-th block.

[0036] Furthermore, in step 3, the proposed gradient conflict repair strategy incrementally applies adversarial enhancements from multiple iterations, and the specific process is as follows:

[0037] Step 3.1: In each training cycle, the system first randomly selects a saved noise generator from the historical buffer, uses it to generate adversarial noise, and superimposes it on the clean image to obtain an adversarial image.

[0038] Step 3.2: Input the clean image and the adversarial image into the depth prediction network and the pose network respectively to obtain their respective depth maps and pose information.

[0039] Step 3.3: Use the same target to perform reprojection reconstruction on the clean image and the adversarial image.

[0040] Step 3.4: Optimize the network by minimizing the loss L of the depth estimation network p while optimizing the noise generator by maximizing the objective function L of the noise generation network AD At the end of each training cycle, the system saves the weight parameters of the noise generator and adds them to the historical buffer for use in the next round of training.

[0041] Based on the same inventive concept, the present invention also provides a stable adversarial training system for generalizable self-supervised monocular depth estimation, which is used to implement the above-mentioned stable adversarial training method for generalizable self-supervised monocular depth estimation.

[0042] The present invention also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-mentioned stable adversarial training method for generalizable self-supervised monocular depth estimation.

[0043] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned stable adversarial training method for generalizable self-supervised monocular depth estimation.

[0044] Compared with the prior art, the present invention has the following beneficial technical effects:

[0045] This invention not only enables multiple existing models to generalize outside of the distribution, but also ensures stable model training, significantly improving the generalization performance of depth estimation and its practicality in the field. The present invention conducted extensive experiments in various scenarios under five benchmarks, achieving optimal performance on multiple datasets, validating the advantages of the present invention's stable adversarial enhancement framework, enabling existing models to generalize outside of the distribution and a more stable training process, thus realizing a generalizable self-supervised monocular depth estimation solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is the overall framework diagram of the SCAT method of the present invention;

[0047] Figure 2 This is a schematic diagram of an image after adding anti-noise in an embodiment;

[0048] Figure 3 The gradient cosine similarity between the perturbation data and the original data and the training oscillation problem diagram in the embodiment are shown.

[0049] Figure 3 (a) is the gradient cosine similarity statistics, Figure 3 (b) The training oscillation problem caused by gradient conflict;

[0050] Figure 4 A visualization of the comparison of various domain generalization methods in experimental results;

[0051] Figure 4 (a) is a comparison diagram of the SCAT method and the offline scene-specific data enhancement method.

[0052] Figure 4 (b) is a comparison diagram of the SCAT method with the original Gaussian noise and the original adversarial data enhancement.

[0053] Figure 4 (c) Comparison of the scaling factors of LSCs with different length jump connections using the SCAT method;

[0054] Figure 5 This is a qualitative analysis diagram of the KITTI-C dataset in the experimental results;

[0055] Figure 6 This is a qualitative analysis diagram of the Foggy Cityscapes dataset in the experimental results;

[0056] Figure 7 This is a qualitative analysis diagram of the DrivingStereo and NuScenes datasets in the experimental results. DETAILED DESCRIPTION

[0057] To make the objectives, technical solutions, and advantages of the present invention clearer, the following provides a detailed description of the embodiments of the present invention with reference to the accompanying drawings.

[0058] To address the training collapse under adversarial data augmentation, the present invention proposes a general adversarial training framework called Stabilized Conflict-optimization Adversarial Training (abbreviated as SCAT). SCAT is a simple and effective model-agnostic framework that uses stable adversarial training to enhance the generalization ability of self-supervised MDE models in unknown scenarios. The present invention proposes a scalable depth estimation network, adjusts the coefficients of the long skip connections in the UNet architecture, and theoretically ensures a more stable training process; it also proposes a gradient conflict repair method to gradually fuse adversarial gradients and guide the model to optimize in a conflict-free direction.

[0059] The present invention provides a stable adversarial training method for generalizable self-supervised monocular depth estimation, as Figure 1 shown, including the following steps:

[0060] Step 1: Based on the input clean image, add noise through a noise generator network.

[0061] Given the clean image I t at time t, the present invention uses the noise generation network f φ to add noise to the input image:

[0062]

[0063] where represents the noisy image input, I t represents the clean image at time t, and δ represents the image noise added by the noise generator network.

[0064] Step 2: For the mixed clean images and noisy images, design a scalable depth estimation network to reduce the reconstruction error by adjusting the coefficients of the long skip connections (abbreviated as LSCs), specifically including the following steps:

[0065] Step 2.1: Use the mixed clean images and noisy images as input images, and obtain the corresponding depth map D t through the depth prediction network.

[0066] Define I′ t ∈ {I t-1 , I t+1} as the video frame images at times t + 1 and t - 1. Given I t and I′ t as inputs, obtain the camera pose T t→t′ through the pose prediction network.。

[0067] Reconstruct the current image through the images at the previous and current moments in combination with the camera internal parameter K:

[0068] I t'→t = I' t <proj(D t , T t→t' , K)> (2)

[0069] In the formula, I t′→t is the reconstructed image when the clean image is used as the input, and I' t ∈ {I t-1 , I t+1} is the video frame image at times t + 1 and t - 1. proj(·) represents generating the 2D coordinates corresponding to the depth map, and <> represents the sampling operation. D t represents the depth map, T t→t′ represents the camera pose, and K represents the camera internal parameter.

[0070] The photometric error function pe is calculated as follows:

[0071]

[0072] In the formula, pe is the photometric error function, α is the weighting coefficient, I a , I b represent two randomly input images. SSIM represents the structural similarity, which is an index for measuring the similarity between two images; ||I a - I b ||1 represents calculating the L1 loss between the two randomly input images. The L1 loss is also called the Mean Absolute Error (MAE), which is a loss function. The weighted combination of the L1 loss and SSIM is used to represent the photometric error function pe.

[0073] Step 2.2, Define the optimization objective L p of the depth prediction network and the objective function L AD of the noise generation network.

[0074] Define the optimization objective of the depth prediction network as L p :

[0075]

[0076] In the formula, L p is the optimization objective of the depth prediction network, I t′→t is the reconstructed image when the clean image is used as the input, and is the reconstructed image when the data-augmented image is used as the input. pe is the photometric error function.

[0077] Contrary to the optimization objective of the depth prediction network, the adversarial noise generator network needs to maximize the reconstruction loss of the noisy image to interfere with the depth prediction network as much as possible:

[0078]

[0079] In the formula, L AD represents the objective function of the noise generation network, is the reconstructed image when the data-augmented image is used as the input.

[0080] Step 2.3: To mitigate the impact of additional noise, the present invention proposes a Scaling Depth Network (SDN), which reduces the reconstruction error by adjusting the coefficient κi of the long skip connections (LSCs for short), specifically:

[0081] DepthNet(x) = f0(x) (6)

[0082]

[0083] where x is the input, and the depth estimation network DepthNet represents f i (x) (i ≥ 1), a i and b i (i ≥ 1) are the trainable parameters of the i-th block. The scaling coefficient κi in the original UNet depth network is fixed at 1. In the scaling depth network, we adjust the coefficient κi of each layer to cope with different degrees of perturbation. By setting different values of κi, we can balance the stability and generalization ability of the network. Verified by the heuristic algorithm, when κi is set to 0.7, the network shows strong generalization ability when dealing with unknown scenarios, while maintaining good performance on the original dataset.

[0084] Step 3: Propose a gradient conflict repair strategy to incrementally apply adversarial enhancements from multiple iterations to alleviate the dual optimization conflict faced when directly applying adversarial data augmentation.

[0085] When directly applying Adversarial Data Augmentation (ADA for short) in self-supervision, the problem of dual optimization conflict will be faced. This is because the gradient optimization directions between the adversarial perturbation data and the original data are opposite, resulting in over-regularization. To alleviate this problem, we propose a gradient conflict repair strategy to gradually introduce noisy data with different degrees, and the specific process is as follows:

[0086] Step 3.1: In each training cycle, the system first randomly selects a saved noise generator from the historical buffer, uses it to generate adversarial noise and superimposes it on the clean image to obtain an adversarial image.

[0087] Step 3.2: Input the clean image and the adversarial image into the depth prediction network and the pose network respectively to obtain their respective depth maps and pose information.

[0088] Step 3.3: Use the same target to perform reprojection reconstruction on the clean image and the adversarial image.

[0089] Step 3.4: Optimize the network by minimizing the loss L of the depth estimation network p while optimizing the noise generator by maximizing the objective function L of the noise generation network AD At the end of each training cycle, the system saves the weight parameters of the noise generator and adds them to the historical buffer for use in the next round of training.

[0090] The present invention also provides a stable adversarial training system for generalizable self-supervised monocular depth estimation, which is used to implement the above-mentioned stable adversarial training method for generalizable self-supervised monocular depth estimation.

[0091] The present invention also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-mentioned stable adversarial training method for generalizable self-supervised monocular depth estimation.

[0092] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned stable adversarial training method for generalizable self-supervised monocular depth estimation.

[0093] Embodiment:

[0094] The following further details with a specific embodiment and verifies the beneficial effects of the present invention through experimental verification.

[0095] Step 1: Based on the input clean image, add noise through the noise generator network.

[0096] An adversarial noise generator composed of a 4-layer 1x1 convolutional neural network is used to add noise to the input clean image. The image before adding noise is I mentioned in the present invention t , and the image after adding noise is as shown in Figure 2 , which also corresponds to

[0097] Step 2: For the mixed clean images and noisy images, design a scalable depth estimation network to reduce the reconstruction error by adjusting the coefficients of the long skip connections.

[0098] As described in the specific implementation, adjust the scaling coefficient κi in each layer of the UNet depth network, and finally adopt κi = 0.7 as the scaling coefficient.

[0099] Step 3: Propose a gradient conflict repair strategy to incrementally apply adversarial enhancements from multiple iterations to alleviate the dual optimization conflict faced when directly applying adversarial data augmentation.

[0100] As Figure 3 shown, the present invention extracts 1,000 images from the KITTI dataset and records the gradient cosine similarity between the adversarial perturbation data and the original data. Without using the conflict gradient repair strategy, as the adversarial training iterates, gradient conflict becomes a common problem, leading to slower convergence and performance degradation. By applying the gradient conflict repair strategy, we transform the distribution of the cosine similarity from negative skewness to positive skewness, alleviating the previously prevalent gradient conflict problem. Specifically, given a set of gradient vectors, construct adversarial perturbation data while using multiple adversarial generators from previous iterations, and suppress the over-regularization caused by the conflict gradients through conflict gradient surgery.

[0101] Experimental verification:

[0102] The present invention uses the KITTI dataset as the training dataset and verifies the generalization performance of the model of the present invention on multiple challenging cross-domain datasets, namely KITTI-C, Cityscapes, and NuScenes. In addition, the stability enhancement effect of the method of the present invention on multiple state-of-the-art models is also verified. We analyzed the disadvantages of existing offline data augmentation methods in terms of training efficiency, while the method of the present invention does not require additional data preparation and can converge quickly.

[0103] (I). Dataset:

[0104] 1. KITTI: The KITTI dataset is a large road dataset collected for mobile robots and autonomous driving, which contains hours of traffic scenes recorded using multiple sensor modalities, including high-resolution RGB, grayscale stereo cameras, and 3D lidar scanners. We used 39,810 images for training and 4,424 images for validation. Subsequently, we strictly evaluated the proposed method and other comparison methods on the KITTI eigen test dataset.

[0105] 2. KITTI-C: To evaluate the robustness and safety of our method in out-of-distribution (OoD) scenarios, we used the KITTI-C dataset as a benchmark. KITTI-C contains 18 common corruption patterns, covering changes in weather and lighting conditions, sensor failures, motion, and noise during data processing. These diverse corruption patterns make it possible to simulate the reliability of the perturbation distribution that may occur in real-world scenarios.

[0106] 3. Cityscapes: The Cityscapes dataset focuses on the semantic understanding of road scenes and contains high-resolution images taken in 50 cities under different seasons and weather conditions. It is used for the evaluation of cross-domain generalization ability, especially the performance test under complex weather conditions such as haze.

[0107] 4. NuScenes: The NuScenes dataset is a large-scale autonomous driving dataset that contains high-resolution images and lidar data collected under different weather and time periods, covering various traffic scenes in urban environments. It is used to evaluate the performance of the model in complex urban environments and different lighting conditions.

[0108] (II) Objective Results:

[0109] 1. Performance of the ideal dataset and the generalized out-of-distribution dataset

[0110] To verify the improvement of the generalization ability of the present invention, we compared the generalization abilities of various baseline methods in common data augmentation and scene adversarial augmentation.

[0111] Table 1 Quantitative analysis of the KITTI-C dataset

[0112]

[0113]

[0114] Table 1 is the quantitative analysis table of the KITTI-C dataset. The bold fonts in the table are the variants with the strongest generalization ability under each method. As shown in Table 1, it can be seen that various methods achieved their best generalization ability after combining the stability data augmentation framework SCAT.

[0115] Figure 4 It is a visualization graph of the comparison of various methods for domain generalization, where Figure 4 (a) is the comparison graph of the SCAT method and the offline scene-specific data augmentation method, Figure 4 (b) is the comparison graph of the SCAT method and the original Gaussian noise and the original adversarial data augmentation, Figure 4 (c) is the comparison graph of different long skip connection (LSC) scale factors of the SCAT method. Combining Table 1 and Figure 4It can be seen that the quantitative results on the out-of-distribution dataset KITTI-C show that SCAT has significant advantages in terms of the mean Corruption Error (mCE) and mean Resilience Rate (mRR) scores. Compared with the baseline models, SCAT outperforms its competitors on all baseline models, providing a general framework to enhance the cross-domain generalization ability of existing self-supervised monocular depth estimation methods.

[0116] Figure 5 is the qualitative analysis of the KITTI-C dataset. Combining Figure 5 with Table 1, it can be seen that when combined with SCAT, various baseline models can infer clearer and more edge-precise depth images in multiple out-of-distribution scenarios.

[0117] 2. Generalization ability on multiple out-of-distribution cross-domain datasets

[0118] The experimental results on the Foggy CityScapes, DrivingStereo, and NuScenes datasets show that the depth estimation results generated by SCAT under all weather conditions have more reliable and clear contours. On these synthetic and real-world datasets, SCAT effectively recovers distant objects occluded by fog or complex scenes.

[0119] Figure 6 is the qualitative analysis of the Foggy CityScapes dataset. From Figure 6 it can be seen that compared with the original baseline model, the model of the present invention based on SCAT shows higher accuracy in predicting the depth of distant objects occluded by fog.

[0120] Figure 7 is the qualitative analysis of the DrivingStereo and NuScenes datasets. From Figure 7 it can be seen that the SCAT method of the present invention can effectively recover challenging distant objects, such as vehicles or streetlights, when dealing with complex and unknown images in real scenes.

[0121] It should be noted that: the embodiments of the present invention are preferred embodiments, not limitations thereof. It should be pointed out that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, modifications can be made to the specific implementation manners or some technical features can be equivalently replaced, and all of them should be regarded as falling within the scope of the present invention.

Claims

1. A stable adversarial training method for generalizable self-supervised monocular depth estimation, characterized in that, It includes the following steps: Step 1: Based on the input clean image, add noise through a noise generator network; Step 2: For the mixed clean image and noisy image, design a scalable depth estimation network to reduce the reconstruction error by adjusting the long skip connection coefficient; Step 3: Propose a gradient conflict repair strategy and incrementally apply adversarial enhancements from multiple iterations to alleviate the dual optimization conflict faced when directly applying adversarial data augmentation.

2. The stable adversarial training method for generalizable self-supervised monocular depth estimation according to claim 1, wherein In Step 1, adding noise to the input clean image through a noise generator network is specifically as follows: Given a clean image \(I\) at time \(t\) t , use the noise generation network \(f\) φ to add noise to the input image: where represents the noisy image input, I t represents the clean image at time t, and δ represents the image noise added by the noise generator network.

3. The stable adversarial training method for generalizable self-supervised monocular depth estimation according to claim 1, wherein In Step 2, for the mixed clean image and noisy image, designing a scalable depth estimation network to reduce the reconstruction error by adjusting the long skip connection coefficient specifically includes the following steps: Step 2.1: Mix the clean image and the noisy image as the input image, and obtain the corresponding depth map D through the depth prediction network t ; Combine the camera intrinsic parameter K, and reconstruct the current image through the images at the previous and current moments; Calculate the photometric error function pe; Step 2.2: Define the optimization objective L of the depth prediction network p and the objective function L of the noise generation network AD ; Step 2.3: Propose a scaled depth network to reduce the reconstruction error by adjusting the long skip connection coefficient κi.

4. The stable adversarial training method for generalizable self-supervised monocular depth estimation according to claim 3, wherein In step 2.1, the mixed clean image and the noisy image are used as the input image, and the corresponding depth map D is obtained through the depth prediction network t ; combined with the camera internal parameter K, the image at the current moment is reconstructed through the images at the previous and current moments; the photometric error function pe is calculated, specifically as follows: Define I′ t ∈{I t-1 ,I t+1} are the video frame images at times t + 1 and t - 1. Given I t and I′ t as inputs, the camera pose T t→t′ is obtained through the pose prediction network; Combined with the camera internal parameter K, reconstruct the current image through the images at the previous and current moments: I t ' →t = I' t <proj(D t ,T t→t' ,K)> (2) where I t′→t is the reconstructed image when the clean image is used as the input, I' t ∈ {I t-1 , I t+1} are the video frame images at times t+1 and t-1, proj(·) represents generating the 2d coordinates corresponding to the depth map, and <> represents the sampling operation, D t represents the depth map, T t→t′ represents the camera pose, and K represents the camera intrinsics; The photometric error function pe is calculated as follows: where pe is the photometric error function, α is the weighting coefficient, and I a and I b represent two randomly input images, and SSIM represents structural similarity; ||I a - I b ||1 represents calculating the L1 loss between the two randomly input images, and the L1 loss is the mean absolute error.

5. The stable adversarial training method for generalizable self-supervised monocular depth estimation according to claim 3, characterized in that, In Step 2.2, the optimization objective L of the defined depth prediction network p and the objective function L of the noise generation network AD are specifically as follows: Where L p is the optimization objective of the depth prediction network, I t′→t is the reconstructed image when the clean image is used as the input, while is the reconstructed image when the data-augmented image is used as the input, and pe is the photometric error function; where L AD represents the objective function of the noise generation network, is the reconstructed image when the data-augmented image is used as the input.

6. The stable adversarial training method for generalizable self-supervised monocular depth estimation according to claim 3, wherein In Step 2.3, proposing a scaled depth network to reduce the reconstruction error by adjusting the long skip connection coefficient κi is specifically as follows: DepthNet(x) = f0(x) (6) where x is the input, and the depth estimation network DepthNet is denoted as f i (x) (i ≥ 1), a i and b i (i ≥ 1) are the trainable parameters of the i-th block.

7. The stable adversarial training method for generalizable self-supervised monocular depth estimation according to claim 1, characterized in that In Step 3, proposing a gradient conflict repair strategy and incrementally applying adversarial enhancements from multiple iterations, the specific process is as follows: Step 3.1: In each training cycle, the system first randomly selects a saved noise generator from the historical buffer, uses it to generate adversarial noise and superimposes it on the clean image to obtain an adversarial image; Step 3.2: Input the clean image and the adversarial image into the depth prediction network and the pose network respectively to obtain their respective depth maps and pose information; Step 3.3: Use the same target to reproject and reconstruct the clean image and the adversarial image; Step 3.4: Optimize the network by minimizing the loss L of the depth estimation network, and optimize the noise generator by maximizing the objective function L of the noise generation network; at the end of each training cycle, the system saves the weight parameters of the noise generator and adds them to the historical buffer for use in the next round of training. p Optimize the network while optimizing the noise generator by maximizing the objective function L of the noise generation network. AD At the end of each training cycle, the system saves the weight parameters of the noise generator and adds them to the historical buffer for use in the next round of training.

8. A stable adversarial training system for generalizable self-supervised monocular depth estimation, characterized in that, The system is used to implement the stable adversarial training method for generalizable self-supervised monocular depth estimation as described in any one of claims 1 to 7.

9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the stable adversarial training method for generalizable self-supervised monocular depth estimation as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the stable adversarial training method for generalizable self-supervised monocular depth estimation as described in any one of claims 1 to 7.