Lightweight monocular depth estimation method and system based on large model fine tuning
By combining the lightweight monocular depth estimation method with CNN and dual-stream Transformer structures, the problem of high computational complexity of monocular depth estimation in resource-constrained scenarios is solved, and efficient and accurate depth estimation is achieved, which is suitable for edge computing devices.
Patent Information
- Application Number
- CN202510521250.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-25
AI Technical Summary
The existing monocular depth estimation methods have high computational complexity and high resource requirements, making it difficult to maintain high accuracy and low computing overhead in resource-constrained application scenarios.
A lightweight monocular depth estimation method based on large model fine-tuning is adopted, combining convolutional neural network (CNN) and dual-stream Transformer structures, features are enhanced through feature extraction and dual-stream Transformer modules, and depth estimation and pose estimation are used to optimize network parameters.
While maintaining high estimation accuracy, it significantly reduces the computational overhead and parameter amount of the model. It is suitable for resource-constrained edge computing scenarios, improving the generalization ability and adaptability of the model.
Smart Images

Figure CN120374732A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and deep learning, and in particular, to a lightweight monocular depth estimation method and system based on large model fine-tuning. Background Art
[0002] With the rapid development of computer vision technology, depth estimation, as one of the important research directions, is widely used in multiple fields such as autonomous driving, robotics, virtual reality, etc. (Monodepth2, Godard et al., 2019). Traditional depth estimation methods rely on multi-sensor systems, and usually require high hardware costs and computational burdens when obtaining depth information. With the rise of deep learning technology, monocular depth estimation methods based on convolutional neural networks (CNNs) have become a research hotspot (Lite-Mono, Ning Zhang et al., 2023). Such methods can achieve effective depth prediction only relying on a single camera image, and have great application potential.
[0003] However, the main problems existing in the prior art are that most of the existing monocular depth estimation methods rely on complex and large network structures, which lead to high demands for computing resources and bottlenecks in inference speed (Depth-Anything, LiheYang et al., 2024). Especially in resource-constrained application scenarios, how to design a depth estimation model that can not only ensure high accuracy but also meet the low computing resource requirements has become an urgent challenge to be solved.
[0004] The difficulty in solving the above problems is that in recent years, although some depth estimation models such as Monodepth2 and Mono-Depth2 have been proposed, these methods often have problems of high computational cost or unstable performance. In order to further improve the parameter efficiency and computational efficiency of the model, researchers have explored in many aspects in the design of lightweight network architectures, but how to achieve low computational overhead while ensuring high accuracy is still a difficulty in the current technology.
[0005] The significance of solving the above problems is to provide a more efficient and accurate depth estimation solution by providing a new type of lightweight monocular depth estimation model. Summary of the Invention
[0006] The present invention provides a lightweight monocular depth estimation method and system based on large model fine-tuning. By combining a convolutional neural network (CNN) and an innovative two-stream Transformer structure, the framework of the present invention significantly reduces the computational overhead and the number of parameters of the model while maintaining high estimation accuracy, and is particularly suitable for resource-constrained edge computing scenarios, solving the deficiencies of existing models in terms of computational complexity and parameter efficiency.
[0007] The technical solution of the present invention is as follows: According to one aspect of the present invention, a lightweight monocular depth estimation method based on large model fine-tuning is provided, including the following steps: S1. Image feature extraction: The input monocular image is subjected to feature extraction operations through a feature extraction module to extract preliminary image features; S2. The preliminary image features are further enhanced through a dual-stream Transformer module to obtain enhanced features; S3. Depth estimation and self-supervised training: The self-supervised training module performs depth estimation and model optimization based on the enhanced features; S4. Pose estimation and joint optimization: The camera pose is estimated and jointly optimized through a pose estimation network in combination with the self-supervised training module.
[0008] Optionally, in the above-mentioned lightweight monocular depth estimation method based on large model fine-tuning, in step S1, the feature extraction module preprocesses the image, including pixel value normalization, image size adjustment, and data augmentation operations, so as to extract preliminary image features.
[0009] Optionally, in the above-mentioned lightweight monocular depth estimation method based on large model fine-tuning, in step S2, the dual-stream Transformer module includes a channel attention stream and a spatial attention stream. First, the input is prepared through a preprocessing stage; then in the channel attention stream, channel stream features are generated through feature dimensionality reduction, attention map calculation, and feature weighted aggregation; in the spatial attention stream, the spatial dimension is flattened through feature rearrangement, spatial position correlation is calculated, and attention weights are applied to generate spatial stream features; finally, the channel stream features and the spatial stream features are fused through an adaptive weighting method to obtain an enhanced feature representation.
[0010] Optionally, in the above-mentioned lightweight monocular depth estimation method based on large model fine-tuning, in step S3, the self-supervised training module receives the fused features output by the dual-stream Transformer module, predicts the depth map of the monocular image through a depth estimation network, and at the same time trains the depth estimation network in combination with a self-supervised learning strategy. During the training process, the self-supervised training module does not require real depth labels, but automatically generates supervision signals through the photometric difference and geometric constraints between consecutive frames, so as to optimize the parameters of the depth estimation network and enable it to accurately predict the depth map.
[0011] Optionally, in the above-mentioned lightweight monocular depth estimation method based on large model fine-tuning, in step S4, the pose estimation network receives the image features of consecutive frames, predicts the relative pose of the camera between different frames, and feeds the pose estimation result back to the self-supervised training module. The self-supervised training module uses the output of the pose estimation network, combines the depth estimation result, and further optimizes the parameters of the depth estimation network and the pose estimation network through the reprojection error.
[0012] According to another aspect of the present invention, a lightweight monocular depth estimation system based on large model fine-tuning is provided, including: a feature extraction module for performing feature extraction operations on the input monocular image to extract preliminary image features; a dual-stream Transformer module for further enhancing the preliminary image features to obtain an enhanced feature representation; and a pose estimation network and a self-supervised training module for performing camera pose estimation and joint optimization through the pose estimation network in combination with the self-supervised training module.
[0013] Optionally, in the above lightweight monocular depth estimation system based on large model fine-tuning, the dual-stream Transformer module includes a channel attention stream and a spatial attention stream. First, through a preprocessing stage (preparing the input; then in the channel attention stream, channel stream features are generated through feature dimensionality reduction, attention map calculation, and feature weighted aggregation; in the spatial attention stream, spatial stream features are generated through feature rearrangement, calculation of spatial position correlation, and application of attention weights; finally, the channel stream features and the spatial stream features are fused through an adaptive weighting method to obtain an enhanced feature representation.
[0014] According to the technical solution of the present invention, the beneficial effects are as follows: By improving the network architecture and optimizing the training strategy, problems such as high computational complexity, slow inference speed, and unstable model accuracy in the prior art are solved, thereby providing a more efficient and accurate depth estimation solution.
[0015] The present invention proposes a lightweight monocular depth estimation system based on fine-tuning of a pre-trained model, which achieves an optimal balance between computational efficiency and estimation accuracy by combining a convolutional neural network (CNN) and a two-stream Transformer architecture. Monocular depth estimation, as a fundamental task in computer vision, is of great significance for numerous application scenarios such as autonomous driving, robot perception, and mixed reality systems. Traditional methods often rely on geometric constraints and sensor measurements provided by LiDAR or stereo vision systems, while the present invention can effectively recover the three-dimensional scene geometric information from a single two-dimensional image. The present invention makes full use of the features of the base model through an efficient fine-tuning strategy, improving the generalization ability of the model while reducing the computational overhead. Specifically, the present invention uses a pre-trained model as the basis for feature extraction and designs a specific fine-tuning strategy to enable the model to quickly adapt to the monocular depth estimation task. This pre-training-based method not only reduces the dependence on large-scale labeled data but also improves the generalization performance of the model in different scenarios through transfer learning. With the rich feature representations extracted by the pre-trained model and the task-specific optimization strategy, it demonstrates superior performance in the monocular depth estimation task. The streamlined architecture proposed by the present invention achieves an optimal balance between performance and computational efficiency, significantly reducing the computational complexity and the number of model parameters while maintaining competitive accuracy; this lightweight design particularly takes into account the resource limitations of edge computing devices, and by optimizing the network structure and computational process, it greatly reduces the storage and computational requirements of the model. The framework of the present invention is not only applicable to high-performance computing platforms but can also be efficiently deployed on resource-constrained edge devices, providing more possibilities for practical application scenarios.
[0016] The novel CNN-Transformer hybrid framework introduced by the present invention can effectively capture local and global features while maintaining computational efficiency. This innovative two-stream structure design enables the model to simultaneously utilize the local feature extraction ability of CNN and the global context modeling ability of Transformer. The CNN stream is responsible for extracting local texture, edges, and other detailed features of the image, while the Transformer stream focuses on the long-range dependencies and global semantic information of the scene. The feature fusion mechanism between the two streams ensures that the model can comprehensively understand the depth information of the scene, thereby producing more accurate depth estimation results.
[0017] A large number of experiments have shown that the framework of the present invention significantly reduces the computational complexity and model parameters while still maintaining competitive accuracy. Compared with existing methods, the present invention has demonstrated superior performance on multiple benchmark data sets, especially in terms of the balance between computational efficiency and accuracy. This efficient design makes the present invention particularly suitable for deployment in resource-constrained edge computing scenarios, providing a practical solution for monocular depth estimation tasks. The innovative architecture design and optimization strategy of the present invention not only improve the deployment efficiency of the model, but also provide new ideas for solving the inherent challenges in monocular depth estimation. By dealing with complex situations such as occlusion and dynamic scene elements, the framework of the present invention exhibits strong environmental adaptability. This comprehensive performance improvement provides important technical support for practical applications in the field of computer vision and promotes the development and innovation of related technologies.
[0018] In order to better understand and illustrate the concept, working principle and effect of the present invention, the present invention is described in detail below through specific embodiments in conjunction with the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific implementation of the present invention or the technical solution in the prior art, the drawings required for use in the specific implementation or the description of the prior art are briefly introduced below.
[0020] Figure 1 It is a flowchart of the steps of the lightweight monocular depth estimation method based on large model fine-tuning of the present invention; Figure 2 It is a flowchart of a lightweight monocular depth estimation method based on large model fine-tuning of the present invention; Figure 3 This is the architecture diagram of the two-stream Transformer; Figure 4 It is the comparison of depth estimation results of two input scenes. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical method and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific examples. These examples are only illustrative and not limiting of the present invention.
[0022] The lightweight monocular depth estimation method and system based on large model fine-tuning of the present invention achieve an optimal balance between computational efficiency and accuracy through a two-stream Transformer architecture design and a self-supervised training strategy. The lightweight monocular depth estimation method and system based on large model fine-tuning of the present invention adopt the method of fine-tuning a pre-trained model, significantly reducing the model complexity while ensuring the depth estimation accuracy. It not only solves the problem of high computational complexity of traditional depth estimation methods but also improves the adaptability to edge scenarios through an innovative architecture design.
[0023] As Figure 1 and Figure 2 shown, the lightweight monocular depth estimation method based on large model fine-tuning of the present invention initializes parameters using pre-trained weights. The input image is processed by a feature fusion module, which consists of two parts: position encoding and a two-stream transformer structure. The result after feature fusion is fed into a pose network, and finally, the predicted depth map is output by the pose network. Throughout the process, the pre-trained weights provide basic parameter support for the feature fusion module. The image data is sequentially processed by the feature fusion and the pose network, and finally, a depth map is generated and output. Specifically, the method includes the following steps: S1. Image feature extraction: The system first performs feature extraction operations on the input monocular image (input picture) through a feature extraction module. The shape of the input image is B×C×H×W, where B represents the batch size, C represents the number of channels (3 for RGB images), and H and W represent the height and width of the image, respectively. The feature extraction module preprocesses the image, including normalizing pixel values (normalizing pixel values to the range [-1, 1]), adjusting the image size (adjusting to a standard size, such as 640×192), and data augmentation operations (including random cropping, horizontal flipping, etc.), so as to extract preliminary image features and provide a basic feature representation for subsequent depth estimation and pose estimation.
[0024] S2. Further enhance the preliminary image features through the dual-stream Transformer module to obtain enhanced features: After the feature extraction module extracts the preliminary image features, the system further enhances the features through the dual-stream Transformer module. The dual-stream Transformer module includes a channel attention stream and a spatial attention stream. First, prepare the input through a preprocessing stage (including generating query Q, key K, and value V feature matrices through linear projection, adding position encoding to enhance the expression of spatial information, and applying LayerNorm for feature normalization); then in the channel attention stream, generate channel feature Yc through feature dimensionality reduction, attention map calculation, and feature weighted aggregation; in the spatial attention stream, generate spatial feature Ys through feature rearrangement (flattening the spatial dimension), calculating spatial position correlation, and applying attention weights; finally, fuse the channel stream feature Yc and the spatial stream feature Ys through an adaptive weighting method (the fusion result is Yfusion, where Yfusion is the fused feature, and γ1, γ2 are learnable fusion parameters) to obtain an enhanced feature representation, providing richer feature information for subsequent depth estimation.
[0025] S3. Depth estimation and self-supervised training: Use the self-supervised training module to perform depth estimation and model optimization based on the enhanced features. The self-supervised training module receives the fused feature Yfusion output by the dual-stream Transformer module, predicts the depth map of the monocular image through the depth estimation network, and simultaneously trains the depth estimation network in combination with self-supervised learning strategies (such as based on photometric consistency loss and geometric consistency loss). During the training process, the self-supervised training module does not require real depth labels, but automatically generates supervision signals through the photometric differences and geometric constraints between consecutive frames, thereby optimizing the parameters of the depth estimation network to enable it to accurately predict the depth map.
[0026] S4. Pose estimation and joint optimization: Through the pose estimation network ( Figure 2 the pose network in
[0027] ), combine with the self-supervised training module to perform camera pose estimation and joint optimization. The pose estimation network receives the image features of consecutive frames (extracted by the feature extraction module), predicts the relative pose (including rotation and translation) of the camera between different frames, and feeds the pose estimation result back to the self-supervised training module. The self-supervised training module uses the output of the pose estimation network, combines with the depth estimation result, and further optimizes the parameters of the depth estimation network and the pose estimation network through the reprojection error (the difference between the reprojection based on depth and pose and the actual image), and outputs the depth map, thereby realizing the joint optimization of depth estimation and pose estimation, and finally improving the overall performance of the system. The lightweight monocular depth estimation system based on large model fine-tuning of the present invention includes: A feature extraction module for performing feature extraction operations on the input monocular image, extracting preliminary image features, and providing a basic feature representation for subsequent depth estimation and pose estimation; A dual-stream Transformer module for further enhancing the features to obtain an enhanced feature representation; wherein, the dual-stream Transformer module includes a channel attention stream and a spatial attention stream. First, in the preprocessing stage (including generating query Q, key K, and value V feature matrices through linear projection, adding positional encoding to enhance the expression of spatial information, and applying LayerNorm for feature normalization) to prepare the input; then in the channel attention stream, generating channel stream feature Yc through feature dimensionality reduction, attention map calculation, and feature weighted aggregation; in the spatial attention stream, generating spatial stream feature Ys through feature rearrangement (flattening the spatial dimension), calculating spatial position correlation, and applying attention weights; finally, fusing the channel stream feature Yc and the spatial stream feature Ys through an adaptive weighting method (the fusion result is Yfusion, where Yfusion is the fused feature, and γ1, γ2 are learnable fusion parameters) to obtain an enhanced feature representation, providing richer feature information for subsequent depth estimation; and A pose estimation network and a self-supervised training module for performing camera pose estimation and joint optimization through the pose estimation network combined with the self-supervised training module.
[0028] Among them, the dual-stream Transformer module, as Figure 3 shown, the specific implementation of the dual-stream Transformer module includes the following key parts: The dual-stream Transformer module receives an input tensor with the shape of B×C×H×W, where: B represents the batch size for parallel processing of multiple samples; C represents the number of channels (3 for RGB images), representing the color information of the image; H and W respectively represent the height and width of the image, determining the spatial resolution of the feature map.
[0029] 1. Input layer normalization processing The input layer of the Transformer module first normalizes the image, including pixel value normalization (normalizing pixel values to the range [-1, 1]), image size adjustment (adjusting to a standard size, such as 640×192), data augmentation operations (including random cropping, horizontal flipping, etc.), and a feature extraction process for preparing subsequent feature extraction.
[0030] 2. Feature extraction process The feature extraction process is divided into the following steps: (1) In the preprocessing stage, three feature matrices of query (Q), key (K), and value (V) are obtained through a linear projection layer. The feature map is position-encoded to enhance the expression of spatial information, and LayerNorm is applied for feature normalization; (2) Channel attention flow: The input features first undergo split feature processing, and then are transformed into a representation space suitable for attention calculation through a channel embedding layer and input into a channel transformer module for channel dimension feature interaction. Here, Qc and Kc represent the query and key matrices in the channel dimension respectively, and normalization processing is achieved by introducing a scaling factor dk and the softmax function. (3) The spatial attention branch adopts a symmetric structure design and also undergoes channel embedding and channel transformer processing. Here, Qs and Ks represent the query and key matrices in the spatial dimension, and the HW parameter characterizes the spatial position of the flattened feature map. The B×HW×C dimensional channel features output by the two branches are integrated through a cross-flow fusion module.
[0031] 3. Feature fusion mechanism The fusion of two-stream features adopts an adaptive weighting method, and the fusion result is Yfusion, where Yfusion is the fused feature with a shape of B×C×H×W, γ1 and γ2 are learnable fusion parameters, Yc is the channel flow feature ( Figure 3 the channel feature in Figure 3 ), and Ys is the spatial flow feature (
[0032] Self-supervised training module. The present invention adopts an innovative self-supervised training method, which does not require depth ground truth supervision and significantly reduces the dependence on labeled data. The core idea of this training is to learn depth estimation by solving the image reconstruction problem between temporally adjacent frames. The system calculates the difference between the target image and the synthesized image to guide the network to learn accurate depth information. During the training process, the system first constructs temporally adjacent image pairs, including the target frame and its adjacent frames. Through the predicted depth information and camera motion parameters, the system can project the adjacent frames into the perspective of the target frame, thereby generating a synthesized image. This view synthesis process fully utilizes multi-view geometric constraints, enabling the network to learn the three-dimensional structure of the scene without supervision. To improve the robustness of the training, the present invention introduces multiple innovative mechanisms. First, the structural similarity metric and pixel-level differences are used as the evaluation criteria for reconstruction quality. This combination takes into account both the structural information of the image and ensures the accuracy of details. Second, an automatic masking mechanism is designed to handle complex situations such as dynamic objects and occluded regions, ensuring that the network only learns from reliable image regions. Finally, an edge-aware smoothing constraint is introduced to suppress noise in the depth map while keeping the depth boundaries sharp. This self-supervised training strategy not only solves the problem of difficult depth data acquisition but also improves the generalization ability of the model in real-world scenarios. Through the integration of multi-frame information and the application of geometric constraints, the system can learn a more stable and accurate depth estimation ability.
[0033] Pose estimation network, which adopts a lightweight design and is used to predict the motion state of the camera between consecutive frames. This network can accurately estimate the six-degree-of-freedom camera motion parameters, including rotation and translation information in three-dimensional space, providing the necessary geometric constraints for depth estimation. In terms of network design, an efficient feature extraction structure is adopted. Through multi-level feature extraction and fusion, the motion information in the image is captured. The network output undergoes special normalization processing to ensure the numerical stability of the prediction results. At the same time, the axis-angle representation is used to describe the rotational motion, which is both intuitive and convenient for optimization. The pose estimation network works in cooperation with the depth estimation network to form a complete self-supervised learning system. Through accurate pose estimation, the system can better understand the geometric structure of the scene, thereby improving the accuracy of depth estimation. This joint learning method fully utilizes temporal information, making the entire system more robust.
[0034] Figure 4The results show a comparison of depth estimation results for two input scenarios. The first row is the input image, the second and third rows are the depth estimation results of the Monodepth2 and Lite-Mono models respectively, and the last row is the model prediction result of this study. The input image shows a road scene, including vehicles, roads, and the surrounding environment. The depth estimation results represent depth information through colors, where warm colors (such as yellow) indicate closer distances and cold colors (such as purple) indicate farther distances. In the first scenario, the depth estimation results of Monodepth2 and Lite-Mono appear blurred at the boundaries of the road and vehicles, and the depth transition is not smooth enough. In particular, the depth prediction for distant vehicles and the road edges is not accurate enough, showing obvious depth discontinuity. In contrast, the model of this study shows clearer depth estimation results in the last row. The boundaries between the road and vehicles are more distinct, the depth transition is smooth, and it can more accurately reflect the relationship between near and far. Especially at the edges of distant vehicles and roads, the depth information is more continuous and reasonable. In the second scenario, the depth estimation results of Monodepth2 and Lite-Mono also have problems. For example, the depth prediction of vehicles in the middle of the road is inaccurate, and the depth distribution in the distant road area is uneven. However, the model of this study shows better results in the last row. The depth information of vehicles and roads in the depth map is more accurate, the overall depth distribution is more consistent, and the expression of the relationship between near and far is more natural. In summary, the model of this study has better results in depth estimation prediction, specifically manifested as clearer boundaries, smoother depth transitions, and more accurate expression of the relationship between near and far. Compared with the Monodepth2 and Lite-Mono models, it can more effectively capture the depth information in the scene and improve the performance of monocular depth estimation.
[0035] The present invention has the following advantages: Pre-training and fine-tuning strategy: The present invention improves the generalization ability of the depth estimation model and reduces the computational overhead by efficiently fine-tuning using the features of the pre-trained model. During the fine-tuning process, through targeted optimization, the model can adapt to different tasks while retaining the rich feature representations of the pre-trained model.
[0036] Lightweight design: The present invention proposes a streamlined network architecture, which achieves efficient deployment in resource-constrained environments by optimizing the balance between performance and computational efficiency. This design significantly reduces the number of model parameters and computational complexity, enabling the framework to achieve efficient inference on low-computing-resource platforms such as mobile devices.
[0037] Dual-Stream Architecture: The present invention introduces an innovative CNN-Transformer hybrid framework, which improves the accuracy and robustness of depth estimation by effectively capturing local features and global context information. While maintaining computational efficiency, this dual-stream architecture enhances the adaptability to complex scenes and diverse environmental conditions.
[0038] The above description is the best embodiment based on the concept and working principle of the invention. The above embodiments should not be construed as limiting the protection scope of the present claims. Combinations of other implementation manners and realization manners in accordance with the concept of the present invention all fall within the protection scope of the present invention.
Claims
1. A lightweight monocular depth estimation method based on large model fine-tuning, characterized in that, It includes the following steps: S1. Image feature extraction: The feature extraction module performs feature extraction operations on the input monocular image to extract preliminary image features; S2. Further enhance the preliminary image features through the dual-stream Transformer module to obtain enhanced features; S3. Depth estimation and self-supervised training: The self-supervised training module performs depth estimation and model optimization based on the enhanced features; S4. Pose estimation and joint optimization: The pose estimation network combines with the self-supervised training module to perform camera pose estimation and joint optimization.
2. The lightweight monocular depth estimation method based on large model fine-tuning according to claim 1, wherein In step S1, the feature extraction module preprocesses the image, including pixel value normalization, image size adjustment, and data augmentation operations, so as to extract the preliminary image features.
3. The lightweight monocular depth estimation method based on large model fine-tuning according to claim 1, characterized in that, In step S2, the dual-stream Transformer module includes a channel attention stream and a spatial attention stream. First, the input is prepared through a preprocessing stage; then in the channel attention stream, channel stream features are generated through feature dimensionality reduction, attention map calculation, and feature weighted aggregation; In the spatial attention stream, the spatial dimension is flattened through feature rearrangement, spatial position correlation is calculated, and attention weights are applied to generate spatial stream features; Finally, the channel stream features and the spatial stream features are fused through an adaptive weighting method to obtain the enhanced feature representation.
4. The lightweight monocular depth estimation method based on large model fine-tuning according to claim 1, wherein In step S3, the self-supervised training module receives the fused features output by the dual-stream Transformer module, predicts the depth map of the monocular image through the depth estimation network, and at the same time trains the depth estimation network in combination with the self-supervised learning strategy. During the training process, the self-supervised training module does not require real depth labels, but automatically generates supervision signals through the photometric difference and geometric constraints between consecutive frames, so as to optimize the parameters of the depth estimation network and enable it to accurately predict the depth map.
5. The lightweight monocular depth estimation method based on large model fine-tuning according to claim 1, characterized in that In step S4, the pose estimation network receives the image features of consecutive frames, predicts the relative pose of the camera between different frames, and feeds the pose estimation result back to the self-supervised training module. The self-supervised training module uses the output of the pose estimation network, combines with the depth estimation result, and further optimizes the parameters of the depth estimation network and the pose estimation network through the reprojection error.
6. A lightweight monocular depth estimation system based on large model fine-tuning, characterized in that It includes: A feature extraction module for performing feature extraction operations on the input monocular image to extract preliminary image features; A dual-stream Transformer module for further enhancing the preliminary image features to obtain an enhanced feature representation; And A pose estimation network and a self-supervised training module, which perform camera pose estimation and joint optimization through the pose estimation network in combination with the self-supervised training module.
7. The lightweight monocular depth estimation system based on large model fine-tuning according to claim 6, wherein, The dual-stream Transformer module includes a channel attention stream and a spatial attention stream. First, the input is prepared through a preprocessing stage; then in the channel attention stream, channel stream features are generated through feature dimensionality reduction, attention map calculation, and feature weighted aggregation; In the spatial attention flow, spatial flow features are generated by feature rearrangement, calculating spatial position correlations, and applying attention weights. Finally, the channel flow features and the spatial flow features are fused through an adaptive weighting method to obtain the enhanced feature representation.