Lightweight monocular depth estimation model training method and device based on self-supervised learning

By constructing a lightweight depth estimation model and combining self-supervised learning with feature extraction and cleansing modules, the problems of high computational complexity and insufficient performance of self-supervised depth estimation methods are solved, and efficient depth estimation is achieved on edge devices.

CN120930712APending Publication Date: 2025-11-11HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510808712.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing self-supervised depth estimation methods suffer from high computational complexity and slow inference speed due to the use of large backbone networks and depth, making them difficult to deploy on edge devices. Furthermore, lightweight methods exhibit blurring in detail-sensitive regions, impacting model performance.

Method used

A lightweight monocular depth estimation model training method based on self-supervised learning is adopted. By constructing a lightweight depth prediction network and a self-supervised pose network, and combining a hole shuffling module, an adaptive rotation kernel attention module, and a depth frequency domain cleanup module, the efficient self-supervised training method is used to reduce the number of model parameters and enhance generalization performance.

Benefits of technology

By avoiding expensive deep annotation, we learn scene-specific features, enhance model generalization performance, reduce the number of training parameters, improve computational efficiency and inference speed, and achieve optimal performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930712A_ABST
    Figure CN120930712A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight monocular depth estimation model training method based on self-supervised learning. The method comprises the following steps: acquiring a high-resolution driving scene video; sampling and preprocessing are carried out, and a training set and a verification set are constructed; constructing a self-supervised lightweight monocular depth estimation model; and inputting the images in the verification set into the trained self-supervised lightweight monocular depth estimation model to predict the pixel-by-pixel depth. According to the method, while expensive depth labeling is avoided, general characteristics of a scene are learned, and the generalization performance of a model is enhanced; the built model has the characteristics of efficient feature extraction, function enhancement and frequency domain signal purification, the performance is excellent in multiple real scenes, the training parameter quantity of the model is greatly reduced, and meanwhile, the reasoning speed on edge equipment is excellent; compared with a traditional self-supervised lightweight monocular depth estimation model, the method has the advantages that optimal performance including accuracy, error and calculation efficiency is achieved with the lowest model parameter quantity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of monocular depth estimation technology in computer vision, and in particular to a lightweight monocular depth estimation model training method and device based on self-supervised learning. Background Technology

[0002] Numerous applications in fields such as robotics, drones, and virtual reality rely on accurate depth perception. Monocular depth estimation, the task of inferring depth information from a single image, has become a fundamental technique for reconstructing the geometry of 3D scenes from single images. Since many real-world scenes lack the large-scale, dense real-world depth required for supervised learning, self-supervised methods offer a significant advantage by utilizing stereo image pairs or monocular video streams, which do not require real-world depth.

[0003] While recent advances in self-supervised depth estimation have reduced the reliance on true depth, many existing mainstream methods are limited by using large backbone networks and deeper networks to achieve significant performance gains. Due to their extremely high training parameters and computational complexity, these methods exhibit slow inference speeds and place high demands on computational resources, thus limiting their deployment on edge devices.

[0004] As research progressed, these limitations spurred innovations in several lightweight solutions. While achieving the lightweight requirement of the model, these lightweight solutions exhibited a low receptive field or poor performance in detail-sensitive regions, such as significant blurring of image contours and structural details, which is crucial for depth estimation performance and thus limits the model's performance.

[0005] Therefore, a lightweight monocular depth estimation model training method based on self-supervised learning is needed. Summary of the Invention

[0006] To address the issues of bulky architectures that sacrifice practicality and structural accuracy in monocular depth estimation, the primary objective of this invention is to provide a lightweight monocular depth estimation model training method based on self-supervised learning that enhances the generalization performance of the model, significantly reduces the number of training parameters, and achieves optimal performance with the lowest possible number of model parameters.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: a lightweight monocular depth estimation model training method based on self-supervised learning, the method comprising the following sequential steps:

[0008] (1) High-resolution driving scene video is acquired using a visible light camera mounted on a car;

[0009] (2) Sample and preprocess the acquired video data to construct training and validation sets;

[0010] (3) Construct a self-supervised lightweight monocular depth estimation model, optimize the images in the training set and input them into the self-supervised lightweight monocular depth estimation model for training, and save the weights obtained during training.

[0011] (4) Input the images in the validation set into the trained self-supervised lightweight monocular depth estimation model to predict pixel-by-pixel depth.

[0012] In step (1), the visible light camera has an effective pixel count of more than 5 million and uses a fixed-focus lens; the driving scene includes more than 60 different areas.

[0013] Step (2) specifically includes the following steps:

[0014] (2a) Sample the video captured by the visible light camera at 10 frames per second;

[0015] (2b) Edge cropping is performed on the sampled images to construct training and validation sets from video streams from different regions;

[0016] (2c) Extract neighboring frames from the images in the training set to construct the optimized training set.

[0017] Step (3) specifically includes the following steps:

[0018] (3a) Construct a self-supervised lightweight monocular depth estimation network, which includes a lightweight depth prediction network and a self-supervised pose network; the lightweight depth prediction network includes a hole shuffling module, an adaptive rotation kernel attention module, and a depth frequency domain cleanup module, and the hole shuffling module outputs the fused feature tensor. The adaptive rotation kernel attention module outputs a tensor X. * The deep frequency domain cleanup module outputs the final tensor X. out ;

[0019] (3b) Select an image from the training set As input to a lightweight depth prediction network, the image has a size of H×W, where H and W represent the height and width of the image, respectively, and the channel dimension is C. The output is the depth map D of this image. t Then, the optimized training set and image X are selected. t Corresponding adjacent frame X t' As input to the self-supervised pose network, the self-supervised pose network consists of four convolutional layers used to estimate X. t With X t' The relative poses between the two positions are sampled to obtain image T. t'→tThen, the SSIM loss and L1 loss are used to calculate the reprojection error L. p :

[0020]

[0021] Wherein, α was experimentally set to 0.85;

[0022] Applying the minimum photometric loss calculation method:

[0023] L p (X s ,X t )=minL p (X t' ,X t )

[0024] Among them, X s Represents relative to image X t The previous and next frames are compared, and then the edge-aware smoothing loss L is applied. s Apply auxiliary constraints to obtain a smoother depth map:

[0025]

[0026] in, This represents the disparity after mean normalization;

[0027] The final loss of the self-supervised lightweight monocular depth estimation network is:

[0028] L final =μL p +τL s

[0029] Where μ is obtained during training, and τ is set to 0.001;

[0030] Save the weights obtained during training.

[0031] Step (4) specifically refers to: inputting the high-resolution images in the validation set into the trained self-supervised lightweight monocular depth estimation model, and using the weights obtained during training to finally obtain the pixel-by-pixel depth and output a dense depth map.

[0032] In step (3a), the dilated shuffling module operates as follows: First, the input tensor is divided into two parts along the channel dimension. One part, Shu1(X), is directly retained for subsequent concatenation, while the other part, Shu2(X), enters the main branch for feature extraction. The main branch consists of three concatenated modules, each consisting of a dilated convolution and two fully connected layers. Normalization and LR operations are performed after each concatenated module. Then, the output of the entire main branch is added to the residual of Shu2(X) to form the main branch result X'. Finally, the main branch result X' is concatenated with Shu1(X) through channels, and the information interaction between different channels is enhanced through a global channel shuffling operation, outputting the fused feature tensor.

[0033]

[0034] F r (X)=Linear(LR(Norm(DConv r (X))))

[0035]

[0036] In the formula, Shu G For global channel shuffling operation, F r (X) is a combined module with dilated convolution rate r in the main branch, Linear is two fully connected layers cascaded, LR is the Leaky ReLU activation function, Norm is a normalization layer, and DConv... r (X) represents dilated convolution. This refers to composite function operations.

[0037] In step (3a), the adaptive rotating kernel attention module operates as follows: the output tensor of the hole shuffling module is... Three types of rotations are performed: channel-width, height-channel, and height-width. Channel-width involves swapping the channel and width dimensions; height-channel involves swapping the height and channel dimensions; and height-width involves swapping the height and width dimensions. Different rotation tensors x are constructed. i i = 1, 2, 3, and then rotate the tensor x i Perform max pooling and average pooling operations separately, and then concatenate them along the channel dimension; the resulting tensor is denoted as x. pool , then x pool The array is passed sequentially through an adaptive kernel convolutional layer and a sigmoid activation function, and then through a rotation tensor x. iAfter performing the Hadamard product, an inverse rotation is performed according to the corresponding rotation mode; finally, the attention-weighted result is weighted by λ for different rotation modes. i Combine:

[0038]

[0039] x pool =[MaxPool(x i ), AvgPool(x i )]

[0040]

[0041] In the formula, Rotation i For different rotation methods, MaxPool is max pooling, AvgPool is average pooling, and X... * Let σ be the output tensor of the adaptive rotating kernel attention module, and σ be the sigmoid activation function. k For adaptive kernel convolutional layers, For different reverse rotation methods.

[0042] In step (3a), the operation of the depth frequency domain purification module is as follows: First, it receives the output tensor X of the adaptive rotating kernel attention module. * As input, it is divided into Shu1(X) by channel shuffling operation. * ) and Shu2(X * The two parts, after processing Shu2(X) * Applying Fourier transform Its frequency domain representation is obtained Then it is decomposed into real and imaginary parts, and then purified by the purifying operator. Extract useful frequency domain features; purification operator The real and imaginary parts of the input complex features are processed separately. First, the Hadamard product is multiplied by a low-pass filter mask M, then passed through a GELU activation function and a batch normalization layer. Finally, the purified real and imaginary parts are obtained through a Conv1 convolution with a kernel of 1, and combined to form the complex features. Perform the Hadamard product, then apply the inverse Fourier transform. Reconstructing the temporal representation Then, it is concatenated with the input Shu1(X) in the pointwise convolutional layer along the channel dimension, and finally passed through a Conv convolution with a kernel size of p. p The final tensor X is obtained. out :

[0043]

[0044] In the formula, Conv is a pointwise convolutional layer. For the reconstructed time-domain representation, BN real and BN imag These represent batch normalization operations performed on the real and imaginary parts, respectively. GELU is the GELU activation function, and Re(x) and Im(x) represent the real and imaginary parts of the tensor x, respectively.

[0045] Another object of the present invention is to provide an electronic device comprising:

[0046] Processor; and

[0047] The memory stores computer program instructions that, when executed by the processor, cause the processor to perform the lightweight monocular depth estimation model training method based on self-supervised learning as described above.

[0048] The present invention also provides a computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the lightweight monocular depth estimation model training method based on self-supervised learning as described above.

[0049] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows: First, the present invention uses a fixed-focus camera to capture high-resolution driving scene videos, and then obtains model training datasets and validation datasets through sampling, preprocessing, etc. Finally, it uses a self-supervised learning method to train a lightweight monocular depth estimation network. This method avoids expensive depth annotation while learning the general features of the scene, enhancing the generalization performance of the model. Second, the model constructed by the present invention has the characteristics of efficient feature extraction and enhancement functions and frequency domain signal cleansing, which performs well in multiple real-world scenarios. It avoids the use of a bulky backbone network, greatly reduces the number of training parameters of the model, and has excellent inference speed on edge devices. Compared with traditional self-supervised lightweight monocular depth estimation models, the present invention achieves optimal performance, including accuracy, error, and computational efficiency, with the lowest number of model parameters. Attached Figure Description

[0050] Figure 1 This is a flowchart of the method of the present invention;

[0051] Figure 2 , 3 4 and 5 are schematic diagrams of the structure of the cavity mixing module, the adaptive rotating kernel attention module, and the deep frequency domain purification module in this invention, respectively.

[0052] Figure 5 This is a real-world depth map predicted by the lightweight monocular depth estimation model used in this invention. Detailed Implementation

[0053] like Figure 1 As shown, a lightweight monocular depth estimation model training method based on self-supervised learning is presented, which includes the following sequential steps:

[0054] (1) High-resolution driving scene video is acquired using a visible light camera mounted on a car;

[0055] (2) Sample and preprocess the acquired video data to construct training and validation sets;

[0056] (3) Construct a self-supervised lightweight monocular depth estimation model. Optimize the images in the training set and input them into the self-supervised lightweight monocular depth estimation model for training. Save the weights obtained during training. Utilize the unique feature extraction and enhancement functions and frequency domain signal purification mechanism of the self-supervised lightweight monocular depth estimation model to retain key features and remove useless information such as noise.

[0057] (4) Input the images in the validation set into the trained self-supervised lightweight monocular depth estimation model to predict pixel-by-pixel depth.

[0058] In step (1), the visible light camera has an effective pixel count of more than 5 million and uses a fixed-focus lens; the driving scene includes more than 60 different areas.

[0059] Step (2) specifically includes the following steps:

[0060] (2a) Sample the video captured by the visible light camera at 10 frames per second;

[0061] (2b) Edge cropping is performed on the sampled images to construct training and validation sets from video streams from different regions;

[0062] (2c) Extract neighboring frames from the images in the training set to construct the optimized training set.

[0063] Step (3) specifically includes the following steps:

[0064] (3a) Construct a self-supervised lightweight monocular depth estimation network, which includes a lightweight depth prediction network and a self-supervised pose network; the lightweight depth prediction network includes a hole shuffling module, an adaptive rotation kernel attention module, and a depth frequency domain cleanup module, and the hole shuffling module outputs the fused feature tensor. The adaptive rotation kernel attention module outputs a tensor X. * The deep frequency domain cleanup module outputs the final tensor X. out ;

[0065] (3b) Select an image from the training set As input to a lightweight depth prediction network, the image has a size of H×W, where H and W represent the height and width of the image, respectively, and the channel dimension is C. The output is the depth map D of this image. t Then, the optimized training set and image X are selected. t Corresponding adjacent frame X t' As input to the self-supervised pose network, the self-supervised pose network consists of four convolutional layers used to estimate X. t With X t' The relative poses between the two positions are sampled to obtain image T. t'→t Then, the SSIM loss and L1 loss are used to calculate the reprojection error L. p :

[0066]

[0067] Wherein, α was experimentally set to 0.85;

[0068] Applying the minimum photometric loss calculation method:

[0069] L p (X s ,X t )=minL p (X t' ,X t )

[0070] Among them, X s Represents relative to image X t The previous and next frames are compared, and then the edge-aware smoothing loss L is applied. s Apply auxiliary constraints to obtain a smoother depth map:

[0071]

[0072] in, This represents the disparity after mean normalization;

[0073] The final loss of the self-supervised lightweight monocular depth estimation network is:

[0074] L final =μL p +τL s

[0075] Where μ is obtained during training, and τ is set to 0.001;

[0076] Save the weights obtained during training.

[0077] Step (4) specifically refers to: inputting the high-resolution images in the validation set into the trained self-supervised lightweight monocular depth estimation model, and using the weights obtained during training to finally obtain the pixel-by-pixel depth and output a dense depth map.

[0078] like Figure 2 As shown, in step (3a), the dilated shuffling module operates as follows: First, the input tensor is divided into two parts along the channel dimension. One part, Shu1(X), is directly retained for subsequent concatenation, while the other part, Shu2(X), enters the main branch for feature extraction. The main branch consists of three concatenated modules, each consisting of a dilated convolution and two fully connected layers. Normalization and LR operations are performed after each concatenation module. Then, the output of the entire main branch is added to the residual of Shu2(X) to form the main branch result X'. Finally, the main branch result X' is concatenated with Shu1(X) through channels, and the information interaction between different channels is enhanced through a global channel shuffling operation, outputting the fused feature tensor.

[0079]

[0080] F r (X)=Linear(LR(Norm(DConv r (X))))

[0081]

[0082] In the formula, Shu G For global channel shuffling operation, F r (X) is a combined module with dilated convolution rate r in the main branch, Linear is two fully connected layers cascaded, LR is the Leaky ReLU activation function, Norm is a normalization layer, and DConv... r (X) represents dilated convolution. This refers to composite function operations.

[0083] By efficiently fusing channel shuffling and dilated convolution operations, the receptive field is increased without introducing additional training parameters, and information exchange between channels is fully realized.

[0084] like Figure 3 As shown, in step (3a), the operation of the adaptive rotating kernel attention module is as follows: the output tensor of the hole shuffling module is... Three types of rotations are performed: channel-width, height-channel, and height-width. Channel-width involves swapping the channel and width dimensions; height-channel involves swapping the height and channel dimensions; and height-width involves swapping the height and width dimensions. Different rotation tensors x are constructed. i i = 1, 2, 3, and then rotate the tensor x i Perform max pooling and average pooling operations separately, and then concatenate them along the channel dimension; the resulting tensor is denoted as x. pool , then x pool The array is passed sequentially through an adaptive kernel convolutional layer and a sigmoid activation function, and then through a rotation tensor x. i After performing the Hadamard product, an inverse rotation is performed according to the corresponding rotation mode; finally, the attention-weighted result is weighted by λ for different rotation modes. i Combine:

[0085]

[0086] x pool =[MaxPool(x i ), AvgPool(x i )]

[0087]

[0088] In the formula, Rotation i For different rotation methods, MaxPool is max pooling, AvgPool is average pooling, and X... * Let σ be the output tensor of the adaptive rotating kernel attention module, and σ be the sigmoid activation function. k For adaptive kernel convolutional layers, For different reverse rotation methods.

[0089] The dependencies between dimensions are established by combining rotation operations and residual connections. Attention gates are used to dynamically compute attention weights for specific dimensions in the spatial domain (channels, width, height). An adaptive convolutional kernel scaling mechanism is adopted to adapt to multi-scale features and enhance the features.

[0090] like Figure 4 As shown, in step (3a), the operation of the deep frequency domain purification module is as follows: First, it receives the output tensor X of the adaptive rotating kernel attention module. * As input, it is divided into Shu1(X) by channel shuffling operation. * ) and Shu2(X * The two parts, after processing Shu2(X) * Applying Fourier transform Its frequency domain representation is obtained Then it is decomposed into real and imaginary parts, and then purified by the purifying operator. Extract useful frequency domain features; purification operator The real and imaginary parts of the input complex features are processed separately. First, the Hadamard product is multiplied by a low-pass filter mask M, then passed through a GELU activation function and a batch normalization layer. Finally, the purified real and imaginary parts are obtained through a Conv1 convolution with a kernel of 1, and combined to form the complex features. Perform the Hadamard product, then apply the inverse Fourier transform. Reconstructing the temporal representation Then, it is concatenated with the input Shu1(X) in the pointwise convolutional layer along the channel dimension, and finally passed through a Conv convolution with a kernel size of p. p The final tensor X is obtained. out :

[0091]

[0092] In the formula, Conv is a pointwise convolutional layer. For the reconstructed time-domain representation, BN real and Bn imag These represent batch normalization operations performed on the real and imaginary parts, respectively. GELU is the GELU activation function, and Re(x) and Im(x) represent the real and imaginary parts of the tensor x, respectively.

[0093] Fourier transform is used to clean up the global signal in the frequency domain, filtering out useless features such as noise while preserving key features such as structural details.

[0094] like Figure 5 As shown, given a displayed driving scene image, the image is input into the lightweight depth prediction network trained by this invention, and the corresponding dense depth map is output.

[0095] In summary, this invention uses a fixed-focus camera to capture high-resolution driving scene videos, then obtains model training and validation datasets through sampling and preprocessing. Finally, it trains a lightweight monocular depth estimation network using self-supervised learning. This method avoids expensive depth annotation while learning general features of the scene, enhancing the model's generalization performance. The model constructed in this invention features efficient feature extraction and enhancement, as well as frequency domain signal cleansing, demonstrating excellent performance in multiple real-world scenarios. It avoids using a cumbersome backbone network, significantly reducing the number of training parameters, and exhibits excellent inference speed on edge devices. Compared to traditional self-supervised lightweight monocular depth estimation models, this invention achieves optimal performance, including accuracy, error, and computational efficiency, with the lowest number of model parameters. This invention does not require real-world scene depth; training can be performed using a given monocular video stream. Utilizing the model's efficient feature extraction and enhancement, along with frequency domain signal cleansing, it demonstrates excellent generalization ability across multiple real-world scenarios.

[0096] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A lightweight monocular depth estimation model training method based on self-supervised learning, characterized in that: The method includes the following steps in sequence: (1) High-resolution driving scene video is acquired using a visible light camera mounted on a car; (2) Sample and preprocess the acquired video data to construct training and validation sets; (3) Construct a self-supervised lightweight monocular depth estimation model, optimize the images in the training set and input them into the self-supervised lightweight monocular depth estimation model for training, and save the weights obtained during training. (4) Input the images in the validation set into the trained self-supervised lightweight monocular depth estimation model to predict pixel-by-pixel depth.

2. The lightweight monocular depth estimation model training method based on self-supervised learning according to claim 1, characterized in that: In step (1), the visible light camera has an effective pixel count of more than 5 million and uses a fixed-focus lens; the driving scene includes more than 60 different areas.

3. The lightweight monocular depth estimation model training method based on self-supervised learning according to claim 1, characterized in that: Step (2) specifically includes the following steps: (2a) Sample the video captured by the visible light camera at 10 frames per second; (2b) Edge cropping is performed on the sampled images to construct training and validation sets from video streams from different regions; (2c) Extract neighboring frames from the images in the training set to construct the optimized training set.

4. The lightweight monocular depth estimation model training method based on self-supervised learning according to claim 1, characterized in that: Step (3) specifically includes the following steps: (3a) Construct a self-supervised lightweight monocular depth estimation network, which includes a lightweight depth prediction network and a self-supervised pose network; the lightweight depth prediction network includes a hole shuffling module, an adaptive rotation kernel attention module, and a depth frequency domain cleanup module, and the hole shuffling module outputs the fused feature tensor. The adaptive rotation kernel attention module outputs a tensor X. * The deep frequency domain cleanup module outputs the final tensor X. out ; (3b) Select an image from the training set As input to a lightweight depth prediction network, the image has a size of H×W, where H and W represent the height and width of the image, respectively, and the channel dimension is C. The output is the depth map D of this image. t Then, the optimized training set and image X are selected. t Corresponding adjacent frame X t' As input to the self-supervised pose network, the self-supervised pose network consists of four convolutional layers used to estimate X. t With X t' The relative poses between the two positions are sampled to obtain image T. t'→t Then, the SSIM loss and L1 loss are used to calculate the reprojection error L. p : Wherein, α was experimentally set to 0.85; Applying the minimum photometric loss calculation method: L p (X s ,X t )=min L p (X t' ,X t ) Among them, X s Represents relative to image X t The previous and next frames are compared, and then the edge-aware smoothing loss L is applied. s Apply auxiliary constraints to obtain a smoother depth map: in, This represents the disparity after mean normalization; The final loss of the self-supervised lightweight monocular depth estimation network is: L final =μL p +τL s Where μ is obtained during training, and τ is set to 0.001; Save the weights obtained during training.

5. The lightweight monocular depth estimation model training method based on self-supervised learning according to claim 1, characterized in that: Step (4) specifically refers to: inputting the high-resolution images in the validation set into the trained self-supervised lightweight monocular depth estimation model, and using the weights obtained during training to finally obtain the pixel-by-pixel depth and output a dense depth map.

6. The lightweight monocular depth estimation model training method based on self-supervised learning according to claim 4, characterized in that: In step (3a), the dilated shuffling module operates as follows: First, the input tensor is divided into two parts along the channel dimension. One part, Shu1(X), is directly retained for subsequent concatenation, while the other part, Shu2(X), enters the main branch for feature extraction. The main branch consists of three concatenated modules, each consisting of a dilated convolution and two fully connected layers. Normalization and LR operations are performed after each concatenated module. Then, the output of the entire main branch is added to the residual of Shu2(X) to form the main branch result X'. Finally, the main branch result X' is concatenated with Shu1(X) through channels, and the information interaction between different channels is enhanced through a global channel shuffling operation, outputting the fused feature tensor. F r (X)=Linear(LR(Norm(DConv r (X)))) In the formula, Shu G For global channel shuffling operation, F r (X) is a combined module with dilated convolution rate r in the main branch, Linear is two fully connected layers cascaded, LR is the Leaky ReLU activation function, Norm is a normalization layer, and DConv... r (X) represents dilated convolution. This refers to composite function operations.

7. The lightweight monocular depth estimation model training method based on self-supervised learning according to claim 4, characterized in that: In step (3a), the adaptive rotating kernel attention module operates as follows: the output tensor of the hole shuffling module is... Three types of rotations are performed: channel-width, height-channel, and height-width. Channel-width involves swapping the channel and width dimensions; height-channel involves swapping the height and channel dimensions; and height-width involves swapping the height and width dimensions. Different rotation tensors x are constructed. i i = 1, 2, 3, and then rotate the tensor x i Perform max pooling and average pooling operations separately, and then concatenate them along the channel dimension; the resulting tensor is denoted as x. pool , then x pool The array is passed sequentially through an adaptive kernel convolutional layer and a sigmoid activation function, and then through a rotation tensor x. i After performing the Hadamard product, an inverse rotation is performed according to the corresponding rotation mode; finally, the attention-weighted result is weighted by λ for different rotation modes. i Combine: x pool =[MaxPool(x i ),AvgPool(x i )] In the formula, Rotation i For different rotation methods, MaxPool is max pooling, AvgPool is average pooling, and X... * Let σ be the output tensor of the adaptive rotating kernel attention module, and σ be the sigmoid activation function. k For adaptive kernel convolutional layers, For different reverse rotation methods.

8. The lightweight monocular depth estimation model training method based on self-supervised learning according to claim 4, characterized in that: In step (3a), the operation of the depth frequency domain purification module is as follows: First, it receives the output tensor X of the adaptive rotating kernel attention module. * As input, it is divided into Shu1(X) by channel shuffling operation. * ) and Shu2(X * The two parts, after processing Shu2(X) * Applying Fourier transform Its frequency domain representation is obtained Then it is decomposed into real and imaginary parts, and then purified by the purifying operator. Extract useful frequency domain features; Purification Operator The real and imaginary parts of the input complex features are processed separately. First, the Hadamard product is multiplied by a low-pass filter mask M, then passed through a GELU activation function and a batch normalization layer. Finally, the purified real and imaginary parts are obtained through a Conv1 convolution with a kernel of 1, and combined to form the complex features. Perform the Hadamard product, then apply the inverse Fourier transform. Reconstructing the temporal representation Then, it is concatenated with the input Shu1(X) in the pointwise convolutional layer along the channel dimension, and finally passed through a Conv convolution with a kernel size of p. p The final tensor X is obtained. out : In the formula, Conv is a pointwise convolutional layer. For the reconstructed time-domain representation, BN real and Bn imag These represent batch normalization operations performed on the real and imaginary parts, respectively. GELU is the GELU activation function, and Re(x) and Im(x) represent the real and imaginary parts of the tensor x, respectively.

9. An electronic device, comprising: processor; as well as A memory storing computer program instructions that, when executed by the processor, cause the processor to perform the lightweight monocular depth estimation model training method based on self-supervised learning as described in any one of claims 1-8.

10. A computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the lightweight monocular depth estimation model training method based on self-supervised learning as described in any one of claims 1-8.