Self-Supervised Monocular Depth Estimation Method Based on Swin-Transformer and CNN Parallel Network

Through feature fusion and self-supervised training of Swin-Transformer and CNN parallel network structures, the problem of loss of image structure information and insufficient long-range correlation in self-supervised monocular depth estimation is solved, and the accuracy of depth estimation is improved.

CN115731280BActive Publication Date: 2025-07-11HARBIN ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211467771.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-07-11
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

The existing self-supervised monocular depth estimation method loses image structure information when using Transformer, resulting in low depth estimation accuracy. Although CNN retains structural information, the long-range correlation is insufficient, making it difficult to achieve good results in the depth estimation task.

Method used

Swin-Transformer and CNN parallel network structures are adopted, feature fusion is performed through the SCFuse module, and self-supervised training is carried out for scale-by-scale self-distillation loss, single-scale image reconstruction loss and edge smoothing loss to build a deep network and a pose network to realize feature extraction and reconstruction.

Benefits of technology

It improves the accuracy of depth estimation, balances the long-range correlation of Transformer and the structural information retention ability of CNN, and improves the accuracy of self-supervised monocular depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731280B_ABST
    Figure CN115731280B_ABST
Patent Text Reader

Abstract

The present invention provides a self-supervised monocular depth estimation method based on a parallel network of Swin-Transformer and CNN, aiming to propose a self-supervised monocular depth estimation method based on a parallel network of Swin-Transformer and convolutional neural network (CNN). The present invention simultaneously uses Swin-Transformer and CNN for feature extraction and fuses the extracted features, which can balance the network between establishing long-range correlations and retaining spatial structure information, strengthen the network's ability to learn features, and combine the scale-by-scale self-distillation loss proposed by the present invention for self-supervised training of the network, thereby improving the accuracy of self-supervised monocular depth estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and relates to a self-supervised depth estimation method based on a parallel network of Swin-Transformer and CNN. Background Art

[0002] Depth estimation has always been one of the important issues in the field of computer vision. In recent years, fields such as autonomous driving, human-computer interaction, virtual reality, and robotics have developed extremely rapidly. Especially, visual solutions have achieved amazing results in autonomous driving. In these application scenarios, how to obtain depth information in the scene is very crucial.

[0003] At the same time, it is easy to distinguish the boundaries of objects by obtaining scene depth information from the depth map, and applying it to other tasks in computer vision, such as 3D object detection and segmentation, scene understanding, etc., can simplify the original algorithms.

[0004] Compared with the supervised method, self-supervised monocular depth estimation can achieve monocular depth estimation without relying on real depth labels, greatly saving the cost of collecting real depth values.

[0005] The Transformer structure has great advantages in establishing long-range correlations and has achieved good results in many sub-tasks in vision. However, Transformer modeling will lose the original structural information of the image. For the depth estimation task, the structural information of the image will have a certain impact on the accuracy of depth estimation. Although the convolutional neural network (CNN) has insufficient ability to establish long-range correlations, it can well preserve the structural information of the image. Among the many evolved structures of Transformer, Swin-Transformer can provide multi-scale feature information. The present invention proposes to use Swin-Transformer and CNN in a parallel combination for the feature extraction part in the depth estimation task to achieve the complementary advantages of the two. At the same time, use Swin-Transformer and CNN to provide multi-scale hierarchical features and fuse them to enhance the feature extraction ability of the network to obtain higher-precision depth prediction. Summary of the Invention

[0006] The object of the present invention is to propose a self-supervised monocular depth estimation method based on a parallel network of Swin-Transformer and CNN, which can perform self-supervised training on unlabeled monocular video sequences. An encoder composed of parallel branches of Swin-Transformer and CNN is constructed to extract features respectively, and the extracted features are fused through a Swin-Transformer and CNN information fusion (SCFuse) module. And a scale-by-scale self-distillation loss is designed, which is combined with a single-scale image reconstruction loss and an edge smoothing loss to jointly constitute the overall loss of the network.

[0007] The object of the present invention is achieved as follows: The steps are as follows:

[0008] Step 1: Use a monocular camera to take pictures and process them to obtain a series of image sequences with a resolution of H*W and a length of N;

[0009] Step 2: Select a frame of image I t from the image sequence in Step 1 as the input of the depth network with a parallel structure of Swin-Transformer and convolutional neural network (CNN), and the output is depth maps D of different scales i . Concatenate I t and the adjacent frame image I t-1 in the channel dimension and use it as the input of the pose network with a pure convolutional neural network structure, and the output is the relative pose T t→t-1 of the two frames of images;

[0010] Step 3: Based on the depth map D0 finally output by the depth network and the relative pose T t→t-1 output by the pose network in Step 2, perform view reconstruction on the input image I t to obtain the reconstructed image I t ', and calculate the single-scale image reconstruction loss L rc . Calculate the scale-by-scale self-distillation loss L i and the edge smoothing loss L esd based on the depth maps D s of different resolutions output by the depth network in Step 2;

[0011] Step 4: Based on the single-scale image reconstruction loss L rc , the scale-by-scale self-distillation loss L esd and the edge smoothing loss L s , construct the overall loss function L total of the depth network and the pose network, and use the monocular video to perform self-supervised training of the network until the overall loss function L total converges; obtain the trained depth network;

[0012] Step 5: Input a single image into the trained deep network. The network outputs a depth map D0 with the same resolution as the input image, and use the depth map D0 as the monocular depth estimation result of the input image.

[0013] The present invention also includes the following structural features:

[0014] 1. The deep network constructed in Step 2 consists of an encoder and a decoder, with cross-layer skip connections between the encoder and the decoder. The encoder is composed of a Swin-Transformer branch and a CNN branch in parallel. The Swin-Transformer branch and the CNN branch are used to extract image features respectively to obtain feature maps of different scales. The Swin-Transformer branch contains n Swin-Transformer modules, and the input image passes through the Swin-Transformer branch to obtain a total of n different scales of feature maps X i . The CNN branch is composed of CNN modules, and the input image passes through the CNN branch to obtain a total of n different scales of feature maps Y i , where the value of n can be selected according to the resolution of the input image to adapt to different resolution inputs.

[0015] 2. The decoder of the deep network constructed in Step 2 is composed of Swin-Transformer modules, which can output n + 1 depth maps D0, D1, D2, …, D n with gradually decreasing resolutions, where D0 has the same resolution as the input image I t .

[0016] 3. The encoder part of the deep network in Step 2 fuses the feature maps X i and Y i of different scales output by the Swin-Transformer branch and the CNN branch of the encoder through the Swin-Transformer and CNN Information Fusion (SCFuse) module to obtain n different scales of fused feature maps Z i . The operation of the SCFuse module is shown in Equation (1):

[0017]

[0018] where X i and Y i respectively represent two input feature maps, Z i is the output of the i-th SCFuse module, represents concatenating the two feature maps in the channel dimension, and CONV 1×1Represents a convolution step with a step size of 1 and a convolution kernel size of 1*1.

[0019] 4. Step three proposes a per-scale self-distillation loss for self-supervised training of the network. The per-scale self-distillation loss L esd is defined as shown in Equation (2):

[0020]

[0021] where D i represents the i-th depth map output by the decoder part of the deep network, upsample(·) represents the upsampling operation, and ||·||2 represents the image similarity function, which is defined as shown in Equation (3):

[0022]

[0023] where and represent the pixel values of two images, and n represents the total number of pixel points in the image.

[0024] 5. Step four constructs the overall loss function L rc of the deep network and the pose network based on the single-scale image reconstruction loss L esd , the per-scale self-distillation loss L s and the edge smoothing loss L total . The single-scale image reconstruction loss L rc is defined as shown in Equation (4):

[0025] L rc = α(1 - SSIM(I t , I t ')) + β||I t - I' t ||1 (4)

[0026] where I t and I' t represent the input image and the reconstructed image respectively, SSIM represents the structural similarity function, and α and β represent the constraint balance factors.

[0027] The edge smoothing loss L s is defined as shown in Equation (5):

[0028]

[0029] where and represent the horizontal and vertical gradients of the input image I t respectively, and p t represents the pixel coordinates of a certain point in the input image I t , Represents the average depth value.

[0030] The overall loss of the network is shown in Equation (6):

[0031] L total = λ1L rc + λ2L s + λ3L esd (6)

[0032] Where L total represents the overall loss function of the network, and λ1, λ2, and λ3 are constraint balance factors.

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a method of parallel combination of Swin-Transformer and CNN for depth estimation tasks. By separately extracting multi-scale features and then fusing them, the advantages of the long-range correlation of Transformer and the advantage of CNN in effectively maintaining the spatial structure information of images are combined, and the entire network is trained in a self-supervised manner. The per-scale self-distillation loss proposed by the present invention can reduce the network's learning from the repeated weak supervision signals caused by the image reconstruction loss, and can enable the network to learn better intermediate feature representations to improve the accuracy of depth estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a schematic diagram of the model structure of the present invention;

[0035] Figure 2 is a structural diagram of the depth network of the present invention;

[0036] Figure 3 is a structural diagram of the Swin-Transformer module of the present invention;

[0037] Figure 4 is a schematic diagram of the merging operation of the present invention;

[0038] Figure 5 is a schematic diagram of the expansion operation of the present invention;

[0039] Figure 6 is a structural diagram of CNN Module 1 of the present invention;

[0040] Figure 7 is a structural diagram of CNN Module 2 of the present invention;

[0041] Figure 8 is a structural diagram of the SCFuse module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0043] A self-supervised monocular depth estimation method based on a parallel network of Swin-Transformer and convolutional neural network (CNN) according to the present invention, the model of which is as shown in the accompanying Figure 1 figures, and is composed of two parts: a depth network and a pose network. The depth network outputs single-channel depth maps D of different scales i , where the depth map D0 with the same resolution as the input image is used as the depth estimation result. Each pixel value of the depth map D0 represents the depth value. The pose network outputs the relative pose between two frames of images. The specific steps of this method are described as follows:

[0044] Step 1: Use a monocular camera with known camera intrinsics to capture a video, and after processing, obtain a series of image sequences with a resolution of H*W and a length N of 2.

[0045] Step 2: Construct a depth network with a parallel structure of Swin-Transformer and convolutional neural network (CNN). Take a frame of image I t in the image sequence as the input of the depth network, and output corresponding depth maps D0, D1, D2,..., D n with decreasing resolutions in turn, where D0 has the same resolution as the input image I t . Construct a pose network with a pure convolutional network structure. Concatenate the image I t and the adjacent frame image I t-1 on the channel dimension as the input of the pose network, and the output is the relative pose T t→t-1 between the two frames of images.

[0046] Step 2(1) The depth network is mainly composed of two parts: an encoder and a decoder.

[0047] The specific construction steps of the encoder are as follows:

[0048] As shown in the accompanying Figure 2 figures, first construct a Swin-Transformer branch network. The Swin-Transformer branch network is composed of repeated stacking of block operations, linear mappings, Swin-Transformer modules, and merging operations n times. The input image I with a channel number of C and a size of H×W tFirst, a feature map with a channel number of C and a size of H / 4×W / 4 is obtained through block operation and linear mapping. Among them, the implementation method of the block operation is to take every adjacent 4*4 size as a block, flatten it in the channel dimension, and obtain a feature map with a channel number of 16*C and a size of H / 4×W / 4. The linear mapping reduces the channel dimension from 4*C to C through a 1*1 convolution, and the size of the feature map remains unchanged. After that, every time a Swin-Transformer module and a merging operation are passed through, the channel number will double, and the size of the feature map will be halved. Finally, the feature map output by the Swin-Transformer branch has a channel number of n*C and a size of H / 2 n +1 ×W / 2 n+1 The input image I t A total of n feature maps X with different scales will be obtained after passing through the Swin-Transformer branch i 。The specific operation of the merging operation is as attached Figure 4 As shown, first extract the information at the same position of the input feature map by blocks and splice it in the channel dimension, and then through layer normalization operation and linear mapping operation, the downsampling operation of the input feature map is realized.

[0049] Then construct a CNN branch network. The CNN branch network consists of CNN module 1 and CNN module 2 stacked repeatedly n times. The specific structure of CNN module 1 is as attached Figure 6 As shown, CNN module 1 consists of two convolutional layers and one max pooling layer. The convolutional kernels of the two convolutional layers are both 3*3, the strides are both 1, and each convolutional layer is followed by a ReLU layer. The pooling window size of the max pooling layer is 2*2, and the stride is 2. The input image I t Outputs a feature map with a channel number of C / 2 and a size of H / 2×W / 2 through CNN module 1. The specific structure of CNN module 2 is as Figure 7 As shown, CNN module 2 also consists of two convolutional layers and one max pooling layer. The convolutional kernels of the two convolutional layers are both 3*3, the strides are both 1, and each convolutional layer is followed by a ReLU layer. The pooling window size of the max pooling layer is 2*2, and the stride is 2. Different from CNN structure 1, CNN structure 2 doubles the channel number of the input feature map and halves the size. Finally, the feature map output by the CNN branch has a channel number of n*C and a size of H / 2 n+1 ×W / 2 n+1 The input image I t A total of n feature maps Y with different scales will be obtained after passing through the CNN branch i 。

[0050] Finally, construct an SCFuse module to fuse the feature maps of different scales output by the Swin-Transformer branch and the CNN branch in the encoder. As attachedFigure 8 As shown in the figure, the SCFuse module can deeply fuse two feature maps of the same scale and output a fused feature map with the same size as the input. The specific operation is shown in Equation (1):

[0051]

[0052] Among them, X i and Y i respectively represent two input feature maps, and Z i is the output of the i-th SCFuse module. represents concatenating the two feature maps in the channel dimension, and CONV 1×1 represents a convolution step with a stride of 1 and a kernel size of 1*1. Through this convolution step, the network can learn by itself the weights assigned to the feature information from the Swin-Transformer branch and the CNN branch, thus achieving more flexible feature fusion. Through n SCFuse modules, n fused feature maps Z i with different resolutions are output.

[0053] The specific construction method of the decoder of the deep network is as follows:

[0054] As shown in the appendix Figure 2 The decoder part consists of an expansion operation and repeated stacking of Swin-Transformer modules. The role of the expansion operation is to upsample the input feature map. The specific steps are shown in the appendix Figure 5 As shown, first, it passes through a linear layer to double the channel dimension, and then the size of the input feature map is doubled through a rearrangement operation opposite to that in the merging operation, thus realizing the upsampling operation of the input feature map. The input of the decoder comes from the output of the Swin-Transformer branch of the encoder and then passes through a Swin-Transformer module. The feature map output by the decoder undergoes a linear transformation to obtain n + 1 depth maps D0, D1, D2,..., D n+1 with different resolutions.

[0055] The encoder part and the decoder part of the deep network are connected by a Swin-Transformer structure. The output of the Swin-Transformer branch of the encoder and the n fused feature maps Z i output after the fusion of the CNN branch through the SCFuse module are input to each layer of the encoder through skip connections. The specific method is to add them pixel by pixel to the output of each expansion operation of the decoder of the deep network and then input them into each Swin-Transformer module of the decoder, so as to fuse the underlying information into the deep network and enable the deep network to obtain the underlying key information.

[0056] The internal structures of all Swin-Transformer modules in the deep network are the same, and the specific structure is as shown in the appendix. Figure 3 For the Swin-Transformer modules in the encoder of the deep network, except for the first one, the inputs of the rest come from the output of the previous merging operation, and the outputs are given to the next merging operation and the corresponding SCFuse module for feature fusion. For each Swin-Transformer module in the decoder part of the deep network, the input comes from the feature map after fusing the output of the previous expansion operation and the output of the decoder SCFuse module, and the output is given to the next expansion operation.

[0057] Step (2): The pose network is composed of a pure convolutional structure.

[0058] The pose network mainly includes seven convolutional layers. The number of convolutional kernels in each layer is 16, 32, 64, 128, 256, 256, 256 respectively, and the stride is 2 for all. The size of the convolutional kernel in the first layer is 7*7, the second layer is 5*5, and the rest are 3*3. After each convolution, it is activated by a ReLU activation layer. Finally, through a 1*1 convolution, the relative pose between two frames is output, which includes three Euler angles and three translation amounts, describing the relative movement of the camera when shooting two frames. The adjacent two-frame images I t and I t-1 are concatenated in the channel dimension and used as the input of the pose network, and the output is the relative pose T t between the two-frame images I t-1 and I t→t-1 .

[0059] Step three: Based on the depth map D0 with the maximum resolution output by the deep network in step two and the relative pose T t→t-1 output by the pose network, the view reconstruction of the input image I t is performed to obtain the reconstructed image I t '. The specific reconstruction steps are as follows. Let p t be a pixel point in the input image I t . Using the reprojection formula (2), the projection point p t of p t-1 on the adjacent frame image I s can be obtained.

[0060] p s = KT t→t-1 D0(p t )K -1 p t (2)

[0061] Then, by interpolating the value of p t-1 in the adjacent frame image I sBilinear sampling is performed on the four nearest pixel points to obtain the pixel value of p t which, together with p t forms the reconstructed image I′ t . The single-scale image reconstruction loss is constructed using the reconstructed image and the input image. The single-scale image reconstruction loss L rc is shown in Equation (3).

[0062] L rc = α(1 - SSIM(I t , I t ′)) + β||I t - I′ t ||1 (3)

[0063] where SSIM represents the structural similarity function, and α and β represent the constraint balance factors. Theoretically, if the reconstructed image I t ′ is exactly the same as the input image I t , the single-scale reconstruction loss is zero.

[0064] Based on the depth maps D i of different scales output by the decoder part of the deep network in step 2, the per-scale self-distillation loss is calculated. The per-scale self-distillation loss L esd is shown in Equation (4):

[0065]

[0066] where D i represents the i-th depth map output by the decoder, and upsample(·) represents the upsampling step, which means sampling the depth map D i+1 with a smaller resolution to the same size as D i . ||·||2 represents the image similarity function, and its definition is shown in Equation (5):

[0067]

[0068] where I a and I b represent two images of the same size, and represent the pixel values of the two images, and n represents the total number of pixel points in the image.

[0069] For the depth map D0 finally output by the decoder of the deep network, the edge smoothing loss L s is shown in Equation (6):

[0070]

[0071] where and respectively represent the input image I t horizontal and vertical gradients, p t represent the input image I t the pixel coordinates of a certain point represent the average depth value, the edge smoothing loss L s make the depth change at the object edge sharper and the depth change in the non-object edge area smoother.

[0072] Step Four: Based on the single-scale image reconstruction loss L rc and the per-scale self-distillation loss L esd and the edge smoothing loss L s construct the overall loss function L of the depth network and the pose network total The overall loss function L of the network total is defined as shown in Equation (7):

[0073] L total = λ1L rc + λ2L s + λ3L esd (7)

[0074] where λ1, λ2, and λ3 are constraint balance factors.

[0075] Use the monocular video for self-supervised training of the network until the overall loss function L total converges. Obtain the trained depth network.

[0076] Step Five: Input a single image into the trained depth network. The network outputs a depth map D0 with the same resolution as the input image, and use the depth map D0 as the monocular depth estimation result of the input image.

[0077] In summary, the purpose of the present invention is to propose a self-supervised monocular depth estimation method based on a parallel network of Swin-Transformer and convolutional neural network (CNN). The present invention uses Swin-Transformer and CNN for feature extraction at the same time, and fuses the extracted features, which can balance the network between establishing long-range correlations and retaining spatial structure information, strengthen the network's ability to learn features, and combine the per-scale self-distillation loss proposed by the present invention for self-supervised training of the network, thereby improving the accuracy of self-supervised monocular depth estimation.

Claims

1. A self-supervised monocular depth estimation method based on a parallel network of Swin-Transformer and CNN, characterized in that, The steps are as follows: Step 1: Use a monocular camera to take pictures and process them to obtain a series of image sequences with a resolution of H*W and a length of N; Step 2: Select a frame image I from the image sequence in Step 1 t As the input of the deep network with the parallel structure of Swin-Transformer and convolutional neural network, the output is depth maps D of different scales i , take I t and the adjacent frame image I t-1 Concatenate them along the channel dimension and use them as the input of the pose network of the pure convolutional neural network structure, and the output is the relative pose T of the two frames of images t→t-1 ; Step 3: Based on the depth map D0 finally output by the depth network and the relative pose T output by the pose network t→t-1 perform view reconstruction on the input image I t to obtain the reconstructed image I′ t , and calculate the single-scale image reconstruction loss L rc ; Based on the depth maps D of different resolutions output by the depth network in Step 2 i calculate the per-scale self-distillation loss L esd and the edge smoothing loss L s ; Step 4: Based on the single-scale image reconstruction loss L rc , the per-scale self-distillation loss L esd and the edge smoothing loss L s construct the overall loss function L of the depth network and the pose network total , and use the monocular video for self-supervised training of the network until the overall loss function L total converges; obtain the trained depth network; Step 5: Input a single image into the trained deep network. The network outputs a depth map D0 with the same resolution as the input image, and use the depth map D0 as the monocular depth estimation result of the input image.

2. The self-supervised monocular depth estimation method based on the Swin-Transformer and CNN parallel network according to claim 1, wherein: The deep network constructed in Step 2 consists of an encoder and a decoder, with cross-layer skip connections between the encoder and the decoder; the encoder is composed of a Swin-Transformer branch and a CNN branch in parallel, and the Swin-Transformer branch and the CNN branch are used to extract image features respectively to obtain feature maps of different scales; there are n Swin-Transformer modules in the Swin-Transformer branch, and the input image passes through the Swin-Transformer branch to obtain a total of n feature maps X of different scales i ; the CNN branch is composed of CNN modules, and the input image passes through the CNN branch to obtain a total of n feature maps Y of different scales i , where the value of n is selected according to the resolution of the input image to achieve the purpose of adapting to inputs of different resolutions.

3. The self-supervised monocular depth estimation method based on the Swin-Transformer and CNN parallel network according to claim 2, characterized in that: The decoder consists of Swin-Transformer modules, which can output depth maps D0, D1, D2, …, D of n+1 different resolutions, and the resolutions decrease in sequence, where D0 has the same resolution size as the input image I n , t ​ 4. The self-supervised monocular depth estimation method based on the Swin-Transformer and CNN parallel network according to claim 2, characterized in that: The encoder part fuses the feature maps X i and Y i output by the Swin-Transformer branch and the CNN branch of the encoder through the Swin-Transformer and CNN information fusion module to obtain n fused feature maps Z i ; The operation of the SCFuse module is as follows: Among them, X i and Y i respectively represent the feature maps of two inputs, and Z i is the output of the i-th SCFuse module. denotes concatenating the two feature maps in the channel dimension, and CONV 1×1 represents a convolution step with a stride of 1 and a convolution kernel size of 1*1.

5. The self-supervised monocular depth estimation method based on the Swin-Transformer and CNN parallel network according to claim 1, characterized in that: The per-scale self-distillation loss L in Step 3 esd is as follows: Among them, D i represents the i-th depth map output by the deep network decoder part, upsample(·) represents the upsampling operation, and ||·||2 represents the image similarity function, which is defined as follows: wherein, and represent the pixel values of two images, and n represents the total number of pixel points of the image.

6. The self-supervised monocular depth estimation method based on the Swin-Transformer and CNN parallel network according to claim 1, characterized in that: Step 4: Construct the overall loss function \(L\) of the depth network and the pose network based on the single-scale image reconstruction loss \(L\) rc , the per-scale self-distillation loss \(L\) esd and the edge smoothing loss \(L\) s ; The definition of the single-scale image reconstruction loss \(L\) total is as follows: rc ​ L rc = α(1 - SSIM(I t , I′ t )) + β ||I t - I′ t ||1 Among them, I t and I' t represent the input image and the reconstructed image respectively, SSIM represents the structural similarity function, and α and β represent the constraint balance factors; Edge smoothing loss L s is as follows: Among them, and respectively represent the horizontal and vertical gradients of the input image I t and p t represents the pixel coordinates of a certain point in the input image I t , represents the average depth value; The overall loss of the network is: L total = λ1L rc + λ2L s + λ3L esd Among them, L total represents the overall network loss function, and λ1, λ2, and λ3 are constraint balance factors.

Citation Information

Patent Citations

  • Self-lifting learning method and device for self-supervised monocular depth estimation and equipment

    CN113724155A

  • Monocular unsupervised depth estimation method based on contextual attention mechanism

    US20210390723A1