Image depth estimation method based on geometric reconstruction up-sampling and dynamic offset constraint
Through the method of geometric reconstruction upsampling and dynamic offset constraints, the problems of chaotic sampling point distribution and high model complexity in monocular depth estimation are solved, and high-precision and lightweight depth map recovery is achieved, which is suitable for scenarios such as autonomous driving and augmented reality.
Patent Information
- Application Number
- CN202510794788.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-14
- Publication Date
- 2025-09-12
AI Technical Summary
Existing monocular depth estimation methods suffer from problems such as blurred depth edges, loss of small-scale structures, and artifact generation in complex scenes. Unconstrained offset learning also leads to chaotic distribution of sampling points, high model complexity, and difficulty in deployment in resource-constrained scenarios.
The method of geometric reconstruction upsampling and dynamic offset constraint is adopted. Through the self-supervised training neural network model, the geometric reconstruction decoder and dynamic offset generation and feature fusion are combined. The pose estimation network is used to constrain the photometric consistency loss and optimize the network parameters to achieve high-precision depth map recovery.
It significantly improves the geometric fidelity and scene adaptability of depth estimation, takes into account both high-resolution detail reconstruction and low computing resource consumption, and achieves stable and reliable depth reconstruction in complex scenes. It is suitable for fields such as autonomous driving and augmented reality.
Smart Images

Figure CN120635167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to an image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints. Background Art
[0002] Monocular depth estimation is a core task in computer vision, aiming to predict three-dimensional depth information from a single two-dimensional image. Unlike methods based on stereo vision, multi-viewing, or lidar, monocular depth estimation relies solely on a single camera, offering the advantages of low hardware cost and lightweight equipment. It is widely used in fields such as autonomous driving, augmented reality (AR), and robotic navigation. Existing methods often employ an encoder-decoder architecture, where the encoder extracts multi-scale features via a convolutional network, while the decoder relies on upsampling to gradually restore a high-resolution depth map. However, the fixed interpolation methods used in traditional decoders, such as bilinear interpolation, suffer from significant drawbacks: their pre-defined regularized sampling patterns struggle to adapt to complex scene geometry, leading to blurred depth edges, loss of small-scale structure, and artifact generation, severely limiting depth estimation accuracy. This problem is particularly acute in real-world scenes with frequent occlusion, sparse textures, or variable illumination.
[0003] In recent years, researchers have tried to improve the upsampling process through dynamic methods such as deformable convolution and attention mechanisms, but such schemes still face key challenges: on the one hand, unconstrained offset learning can easily lead to chaotic distribution of sampling points, causing feature dislocation and computational instability; on the other hand, the surge in the number of parameters in multi-stage upsampling increases the complexity of the model, making it difficult to deploy in resource-constrained scenarios. In addition, existing methods mostly rely on fixed initialization strategies (such as bilinear grids), lack an adaptive balance between geometric priors and data-driven, and limit the robustness of cross-scale feature reconstruction. In response to the above problems, there is an urgent need for an upsampling mechanism that takes into account efficiency, geometric fidelity, and generalization capabilities to enhance the depth estimation model's ability to parse complex scenes while meeting high precision and lightweight requirements. The geometric reconstruction upsampling and dynamic offset constraint method proposed in this patent is designed to solve this technical bottleneck, providing an efficient and reliable feature recovery solution for the deep decoder. Summary of the Invention
[0004] The present invention proposes an image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints, which can accurately and effectively perform depth estimation on low-resolution images.
[0005] The present invention adopts the following technical solutions.
[0006] An image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints first collects images to form a training dataset and inputs them into a depth encoder to extract multi-scale features. Subsequently, a geometric reconstruction decoder achieves high-precision geometric structure recovery through dynamic offset generation and feature fusion, gradually upsampling to restore scene depth and generate a high-precision depth map. The estimation model used in the method is a neural network model based on a self-supervised training method. The pose estimation network is used to predict the pose of the camera used for image capture and a photometric consistency loss constraint is constructed to optimize the network parameters for model training.
[0007] The method comprises the following steps:
[0008] Step S1: Collect the original image dataset and preprocess the data to remove images with poor visual effects; input the image dataset to the model's deep encoder, use residual connections and dilated convolution structures to extract features layer by layer, and output a multi-scale feature map;
[0009] Step S2: In the depth decoder, a dynamic offset parameter is generated through linear transformation. The offset range is adaptively controlled by combining a preset constraint factor to avoid overlapping or over-dispersion of the sampling points. Subsequently, the offset parameter is superimposed on the basic sampling grid to construct an adaptive sampling set, which provides dynamic spatial guidance for subsequent upsampling.
[0010] Step S3: Use the dual-mode geometric reconstruction upsampling module to restore the high-resolution depth map, select the high-precision mode (LPR) or the lightweight mode (PRL) to fuse multi-scale features, upsample step by step, and output a high-resolution depth map;
[0011] Step S4: Predict the camera poses of adjacent frames through a lightweight pose estimation network, calculate the relevant losses for geometric constraints, iteratively optimize the network parameters, and save the optimal model based on the verified geometric error and prediction accuracy.
[0012] Step S1 specifically includes the following steps:
[0013] Step S11: Using the public image dataset KITTI dataset as a training dataset, perform data preprocessing and remove images with poor visual effects;
[0014] Step S12: Build a deep encoder architecture with residual connections as the core and design a multi-layer convolution module combination: each residual unit contains a convolution layer, a normalization layer, and a nonlinear activation function, and cross-layer feature reuse is achieved through skip connections; at the same time, multiple groups of dilated convolution layers are embedded, and different void rates are used to expand the receptive field;
[0015] Step S13: Read the original input frame image I t ∈R C×H×WInput the depth encoder, where C represents the number of feature channels, H represents the height of the image, and w represents the height of the image. Through multi-stage progressive downsampling and hierarchical feature extraction, it outputs a multi-scale feature map sequence from high resolution to low resolution;
[0016] In step S13, three stages of progressive downsampling and hierarchical feature extraction are performed.
[0017] Step S2 specifically includes the following steps:
[0018] Step S21: Multi-scale feature map X∈R output by the encoder C×H×W Perform channel normalization to eliminate feature distribution differences. 2 , r is the linear projection layer of the upsampling scale factor to generate the initial offset O init , which is specifically expressed as:
[0019] O init =W linear ·X+b linear
[0020] Where W linear is the weight matrix, b linear is the bias term, and the output tensor size is 2r 2 ×H×W, representing the r of each spatial position 2 The initial offset coordinates of the sampling points;
[0021] Step S22: Introduce the reconstruction constraint factor γ and set it through theoretical derivation Limit the maximum displacement range of the offset to avoid overlapping or excessive dispersion of adjacent sampling points; after the constraint, each coordinate component of the offset O is compressed to interval, ensuring that the sampling points in the local neighborhood are evenly covered and have no overlap, which can be specifically expressed as:
[0022] O=γ·O init ;
[0023] Step S23: Construct a bilinear initialization grid G of size 2×rH×rW, where each output position coordinate (x, y) corresponds to the input feature map. Neighborhood, the initial sampling points are distributed according to the bilinear interpolation weights; Step S24: reshape the constrained offset O into a size of 2×rH×rW through a pixel reconstruction operation, denoted as O′, and add it element-by-element to the basic grid G to generate a dynamic sampling set S.
[0024] S=G+O′.
[0025] Step S3 specifically includes the following steps:
[0026] Step S31: Select an operation mode based on computing resources and accuracy requirements:
[0027] LPR mode: Keep the original size C×H×W of the input feature map X and generate 2r directly through the linear layer 2 ×H×W offset, and then reshaped to 2×rH×rW, preserving the full spatial correlation;
[0028] PRL mode: reshape X into Then, a 2×rH×rW offset is generated through a lightweight linear layer, reducing the number of GFLOPs parameters by about r 2 times;
[0029] Step S32: Based on the dynamic sampling set S, perform geometric-aware feature resampling on the input feature map X to output upsampled features X′ with a size of c×rH×rW, which is specifically expressed as follows:
[0030]
[0031] N(s) is the neighborhood pixel index of the sampling point s in the dynamic sampling set S, and ω(n,s) is the bilinear interpolation weight;
[0032] Step S33: The deep decoder combines the skip connection to gradually fuse the low-scale sampling features while maintaining the high-resolution feature representation and gradually restore the feature resolution;
[0033] Step S34: Perform depth regression and range normalization on the final fusion feature map, map the pixel values to the preset depth range, and obtain the final inverse depth map D t .
[0034] In step S31 , the LPR mode is used for high-precision sampling requirements, and when the computing resources are lightweight computing resources, the PRL mode is used.
[0035] Step S4 specifically includes the following steps:
[0036] Step S41: Input the posture network encoder with continuous frames to extract spatiotemporal features, construct motion perception representation through 1×1 convolution dimensionality reduction and feature splicing, and output the relative posture matrix T through 3×3 convolution layer and ReLU activation. t→s ;
[0037] Step S42: Based on the inverse depth map D generated in step S34 t and the relative pose matrix T output in step S41 t→s , reconstruct the target synthetic image using projection mapping and bilinear interpolation
[0038] Step S43: synthesize the target image With L tPerform geometric consistency constraints and calculate photometric reprojection loss The formula is as follows:
[0039]
[0040] Among them, SIIM(·) represents the sum of pixel similarities, α is set to a constant, and the minimum reprojection loss is used to reduce occlusion and artifacts. The specific formula is as follows:
[0041]
[0042] Among them, I s Indicates the previous or next frame of the target image;
[0043] In order to make the depth map consistent in smooth areas while retaining clear boundaries in edge areas, an edge-aware smoothness loss is calculated. The specific formula is as follows:
[0044]
[0045] where d * Indicates that the inverse depth is normalized. is the set parameter;
[0046] Final loss The calculation formula is as follows:
[0047]
[0048] γ is set to a constant, and λ is set to a constant.
[0049] Step S44: Model training iteratively optimizes parameters to minimize the loss function, learns the mapping relationship between the input image and the corresponding depth map, and trains until the preset maximum number of iterations is reached. At this time, the model performance is evaluated and the weight with the best prediction accuracy is selected as the final model.
[0050] In step S43, α is set to 0.85; γ is set to 1.2; and λ is set to 1e -3 .
[0051] The estimation model is used to perform geometric reconstruction upsampling and dynamic offset constraint mechanisms in computer vision tasks for monocular depth estimation. The geometric reconstruction upsampling and dynamic offset constraint mechanisms are seamlessly integrated with the self-supervised depth estimation framework, and do not need to rely on real depth labels when modeling high-fidelity depth.
[0052] Compared with the existing technology, the present invention has the following beneficial effects:
[0053] 1. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints constructed in the present invention can improve the sampling point overlap and geometric distortion problems caused by unconstrained offset in upsampling in traditional depth estimation tasks, and significantly improve the integrity of depth edges and the fidelity of scene structure.
[0054] 2. The innovative dual-mode adaptive upsampling architecture of this invention takes into account both high-resolution detail reconstruction and low computing resource consumption requirements, enabling flexible adaptation of the algorithm between complex scenarios and edge devices.
[0055] 3. The collaborative architecture of geometric prior guidance and dynamic offset learning adopted in this invention combines the regularization characteristics of bilinear interpolation grid with data-driven adaptive sampling to achieve stable and reliable depth reconstruction in complex occlusion and weak texture areas.
[0056] 4. This invention combines the self-supervised multi-frame joint optimization paradigm to achieve joint optimization of geometric fidelity and computational efficiency through end-to-end training, without relying on complex post-processing or additional annotation information, significantly improving the deployment convenience and real-time performance of the model in industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0058] Attachment Figure 1 It is a schematic diagram of the principle of the present invention. DETAILED DESCRIPTION
[0059] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0060] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.
[0061] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0062] like Figure 1As shown in the figure, an image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints first collects images to form a training data set and inputs them into a depth encoder to extract multi-scale features; then, a geometric reconstruction decoder generates dynamic offsets and fuses features to achieve high-precision geometric structure recovery, gradually upsamples to restore scene depth, and generates a high-precision depth map; the estimation model used in the method is a neural network model based on a self-supervised training method, which predicts the pose of the camera used for image capture through a posture estimation network and constructs a photometric consistency loss constraint to optimize network parameters for model training.
[0063] The method comprises the following steps:
[0064] Step S1: Collect the original image dataset and preprocess the data to remove images with poor visual effects; input the image dataset to the model's deep encoder, use residual connections and dilated convolution structures to extract features layer by layer, and output a multi-scale feature map;
[0065] Step S2: In the depth decoder, a dynamic offset parameter is generated through linear transformation. The offset range is adaptively controlled by combining a preset constraint factor to avoid overlapping or over-dispersion of the sampling points. Subsequently, the offset parameter is superimposed on the basic sampling grid to construct an adaptive sampling set, which provides dynamic spatial guidance for subsequent upsampling.
[0066] Step S3: Use the dual-mode geometric reconstruction upsampling module to restore the high-resolution depth map, select the high-precision mode (LPR) or the lightweight mode (PRL) to fuse multi-scale features, upsample step by step, and output a high-resolution depth map;
[0067] Step S4: Predict the camera poses of adjacent frames through a lightweight pose estimation network, calculate the relevant losses for geometric constraints, iteratively optimize the network parameters, and save the optimal model based on the verified geometric error and prediction accuracy.
[0068] Step S1 specifically includes the following steps:
[0069] Step S11: Using the public image dataset KITTI dataset as a training dataset, perform data preprocessing and remove images with poor visual effects;
[0070] Step S12: Build a deep encoder architecture with residual connections as the core and design a multi-layer convolution module combination: each residual unit contains a convolution layer, a normalization layer, and a nonlinear activation function, and cross-layer feature reuse is achieved through skip connections; at the same time, multiple groups of dilated convolution layers are embedded, and different void rates are used to expand the receptive field;
[0071] Step S13: Read the original input frame image I t ∈R C×H×WInput the depth encoder, where C represents the number of feature channels, H represents the height of the image, and w represents the height of the image. Through multi-stage progressive downsampling and hierarchical feature extraction, it outputs a multi-scale feature map sequence from high resolution to low resolution;
[0072] In step S13, three stages of progressive downsampling and hierarchical feature extraction are performed.
[0073] Step S2 specifically includes the following steps:
[0074] Step S21: Multi-scale feature map X∈R output by the encoder C×H×W Perform channel normalization to eliminate feature distribution differences. 2 , r is the linear projection layer of the upsampling scale factor to generate the initial offset O init , which is specifically expressed as:
[0075] O init =W linear ·X+b linear
[0076] Where W linear is the weight matrix, b linear is the bias term, and the output tensor size is 2r 2 ×H×W, representing the r of each spatial position 2 The initial offset coordinates of the sampling points;
[0077] Step S22: Introduce the reconstruction constraint factor γ and set it through theoretical derivation Limit the maximum displacement range of the offset to avoid overlapping or excessive dispersion of adjacent sampling points; after the constraint, each coordinate component of the offset O is compressed to interval, ensuring that the sampling points in the local neighborhood are evenly covered and have no overlap, which can be specifically expressed as:
[0078] O=γ·O init ;
[0079] Step S23: Construct a bilinear initialization grid G of size 2×rH×rW, where each output position coordinate (x, y) corresponds to the input feature map. Neighborhood, the initial sampling points are distributed according to the bilinear interpolation weights; Step S24: reshape the constrained offset O into a size of 2×rH×rW through a pixel reconstruction operation, denoted as O′, and add it element-by-element to the basic grid G to generate a dynamic sampling set S.
[0080] S=G+O′.
[0081] Step S3 specifically includes the following steps:
[0082] Step S31: Select an operation mode based on computing resources and accuracy requirements:
[0083] LPR mode: Keep the original size C×H×W of the input feature map X and generate 2r directly through the linear layer 2 ×H×W offset, and then reshaped to 2×rH×rW, preserving the full spatial correlation;
[0084] PRL mode: reshape X into Then, a 2×rH×rW offset is generated through a lightweight linear layer, reducing the number of GFLOPs parameters by about r 2 times;
[0085] Step S32: Based on the dynamic sampling set S, perform geometric-aware feature resampling on the input feature map X to output upsampled features X′ with a size of c×rH×rW, which is specifically expressed as follows:
[0086]
[0087] N(s) is the neighborhood pixel index of the sampling point s in the dynamic sampling set S, and ω(n,s) is the bilinear interpolation weight;
[0088] Step S33: The deep decoder combines the skip connection to gradually fuse the low-scale sampling features while maintaining the high-resolution feature representation and gradually restore the feature resolution;
[0089] Step S34: Perform depth regression and range normalization on the final fusion feature map, map the pixel values to the preset depth range, and obtain the final inverse depth map D t .
[0090] In step S31 , the LPR mode is used for high-precision sampling requirements, and when the computing resources are lightweight computing resources, the PRL mode is used.
[0091] Step S4 specifically includes the following steps:
[0092] Step S41: Input the posture network encoder with continuous frames to extract spatiotemporal features, construct motion perception representation through 1×1 convolution dimensionality reduction and feature splicing, and output the relative posture matrix T through 3×3 convolution layer and ReLU activation. t→s ;
[0093] Step S42: Based on the inverse depth map D generated in step S34 t and the relative pose matrix T output in step S41 t→s , reconstruct the target synthetic image using projection mapping and bilinear interpolation
[0094] Step S43: synthesize the target image With L tPerform geometric consistency constraints and calculate photometric reprojection loss The formula is as follows:
[0095]
[0096] Among them, SIIM(·) represents the sum of pixel similarities, α is set to a constant, and the minimum reprojection loss is used to reduce occlusion and artifacts. The specific formula is as follows:
[0097]
[0098] Among them, I s Indicates the previous or next frame of the target image;
[0099] In order to make the depth map consistent in smooth areas while retaining clear boundaries in edge areas, an edge-aware smoothness loss is calculated. The specific formula is as follows:
[0100]
[0101] where d * Indicates that the inverse depth is normalized. is the set parameter;
[0102] Final loss The calculation formula is as follows:
[0103]
[0104] γ is set to a constant, and λ is set to a constant.
[0105] Step S44: Model training iteratively optimizes parameters to minimize the loss function, learns the mapping relationship between the input image and the corresponding depth map, and trains until the preset maximum number of iterations is reached. At this time, the model performance is evaluated and the weight with the best prediction accuracy is selected as the final model.
[0106] In step S43, α is set to 0.85; γ is set to 1.2; and λ is set to 1e -3 .
[0107] The estimation model is used to perform geometric reconstruction upsampling and dynamic offset constraint mechanisms in computer vision tasks for monocular depth estimation. The geometric reconstruction upsampling and dynamic offset constraint mechanisms are seamlessly integrated with the self-supervised depth estimation framework, and do not need to rely on real depth labels when modeling high-fidelity depth.
[0108] This example proposes an image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints, which can effectively address the problems of geometric detail loss, artifact generation, and structural distortion commonly encountered in the upsampling operation of traditional depth decoders. This method proposes a dynamic sampling point distribution optimization and geometric reconstruction upsampling mechanism, significantly improving the geometric fidelity and scene adaptability of depth estimation. First, images are collected to form a training dataset and input into a depth encoder to extract multi-scale features. Subsequently, the geometric reconstruction decoder achieves high-precision geometric structure recovery through dynamic offset generation and feature fusion, and gradually upsamples to restore the scene depth and generate a high-precision depth map. The entire network adopts a self-supervised training method, using a pose estimation network to predict the camera pose and construct a photometric consistency loss constraint to optimize the network parameters. End-to-end training can be achieved without the need for real depth labels.
[0109] The geometric reconstruction upsampling and dynamic offset constraint mechanism proposed in this example can be seamlessly integrated with the self-supervised depth estimation framework, achieving high-fidelity depth modeling for any scene without relying on real depth labels: through the adaptive geometric reconstruction capability of the dynamic sampling set, the reprojection error of multi-scale features is optimized in the decoder upsampling stage, and the offset constraint factor is used to ensure the geometric consistency of cross-resolution depth prediction. At the same time, combined with the dual-mode upsampling strategy, the texture prior knowledge implicit in the high-resolution image is transferred to the low-resolution depth reconstruction process, effectively overcoming the edge distortion and structural fracture problems caused by interpolation blur in traditional self-supervised methods. Compared with existing solutions that rely on fixed interpolation or complex annotation, this method significantly enhances the model's robustness to image degradation (such as blur and noise) through the coordinated optimization of data-driven offset learning and theoretical constraints. In depth estimation scenarios such as mobile 3D perception and wide-area monitoring equipment, it has the dual advantages of high-precision reconstruction and lightweight deployment.
[0110] The above description is only a preferred embodiment of the present invention, and all equivalent changes and modifications made within the scope of the patent application of the present invention should be covered by the present invention.
Claims
1. An image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints, characterized by: First, collect images to form a training dataset and input them into a deep encoder to extract multi-scale features; Subsequently, the geometric reconstruction decoder achieves high-precision geometric structure recovery through dynamic offset generation and feature fusion, gradually upsampling to restore the scene depth and generate a high-precision depth map; the estimation model used in the method is a neural network model based on a self-supervised training method, which predicts the pose of the camera used for image capture through a posture estimation network and constructs a photometric consistency loss constraint to optimize the network parameters for model training.
2. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 1, characterized in that: The method comprises the following steps: Step S1: Collect the original image dataset and preprocess the data to remove images with poor visual effects; input the image dataset to the model's deep encoder, use residual connections and dilated convolution structures to extract features layer by layer, and output a multi-scale feature map; Step S2: In the depth decoder, a dynamic offset parameter is generated through linear transformation, and the offset range is adaptively controlled in combination with a preset constraint factor to avoid overlapping or over-dispersion of the sampling point distribution; Subsequently, the offset parameters are superimposed on the basic sampling grid to construct an adaptive sampling set, which provides dynamic spatial guidance for subsequent upsampling; Step S3: Use the dual-mode geometric reconstruction upsampling module to restore the high-resolution depth map, select the high-precision mode LPR or the lightweight mode PRL to fuse multi-scale features, upsample step by step, and output a high-resolution depth map; Step S4: Predict the camera poses of adjacent frames through a lightweight pose estimation network, calculate the relevant losses for geometric constraints, iteratively optimize the network parameters, and save the optimal model based on the verified geometric error and prediction accuracy.
3. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 2, characterized in that: Step S1 specifically includes the following steps: Step S11: Using the public image dataset KITTI dataset as a training dataset, perform data preprocessing and remove images with poor visual effects; Step S12: Build a deep encoder architecture with residual connections as the core and design a multi-layer convolution module combination: each residual unit contains a convolution layer, a normalization layer, and a nonlinear activation function, and cross-layer feature reuse is achieved through skip connections; at the same time, multiple groups of dilated convolution layers are embedded, and different void rates are used to expand the receptive field; Step S13: Read the original input frame image I t ∈R C×H×W Input the depth encoder, where C represents the number of feature channels, H represents the height of the image, and W represents the height of the image. Through multi-stage progressive downsampling and hierarchical feature extraction, a multi-scale feature map sequence from high resolution to low resolution is output.
4. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 2, characterized in that: In step S13, three stages of progressive downsampling and hierarchical feature extraction are performed.
5. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 2, characterized in that: Step S2 specifically includes the following steps: Step S21: Multi-scale feature map X∈R output by the encoder C×H×W Perform channel normalization to eliminate feature distribution differences. 2 , r is the linear projection layer of the upsampling scale factor to generate the initial offset O init , which is specifically expressed as: O init =W linear ·X+b linear Where W linear is the weight matrix, b linear is the bias term, and the output tensor size is 2r 2 ×H×W, representing the r of each spatial position 2 The initial offset coordinates of the sampling points; Step S22: Introduce the reconstruction constraint factor γ and set it through theoretical derivation Limit the maximum displacement range of the offset to avoid overlapping or excessive dispersion of adjacent sampling points; after the constraint, each coordinate component of the offset O is compressed to interval, ensuring that the sampling points in the local neighborhood are evenly covered and have no overlap, which can be specifically expressed as: O=γ·O init ; Step S23: Construct a bilinear initialization grid G of size 2×rH×rW, where each output position coordinate (x, y) corresponds to the input feature map. Neighborhood, initial sampling points are distributed according to bilinear interpolation weights; Step S24: reshape the constrained offset O into a size of 2×rH×rW through a pixel reconstruction operation, denoted as O′, and add it element-by-element to the base grid G to generate a dynamic sampling set S. S=G+O′.
6. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 2, characterized in that: Step S3 specifically includes the following steps: Step S31: Select an operation mode based on computing resources and accuracy requirements: LPR mode: Keep the original size C×H×W of the input feature map X and generate 2r directly through the linear layer 2 ×H×W offset, and then reshaped to 2×rH×rW, preserving the full spatial correlation; PRL mode: reshape X into Then generate 2×rH×rW offset through the lightweight linear layer, reducing the number of GFLOPs parameters by about r 2 times; Step S32: Based on the dynamic sampling set S, perform geometric-aware feature resampling on the input feature map X to output upsampled features X′ with a size of c×rH×rW, which is specifically expressed as follows: N(s) is the neighborhood pixel index of the sampling point s in the dynamic sampling set S, and ω(n, s) is the bilinear interpolation weight; Step S33: The deep decoder combines the skip connection to gradually fuse the low-scale sampling features while maintaining the high-resolution feature representation and gradually restore the feature resolution; Step S34: Perform depth regression and range normalization on the final fusion feature map, map the pixel values to the preset depth range, and obtain the final inverse depth map D t .
7. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 6, characterized in that: In step S31 , the LPR mode is used for high-precision sampling requirements, and when the computing resources are lightweight computing resources, the PRL mode is used.
8. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 2, characterized in that: Step S4 specifically includes the following steps: Step S41: Input the posture network encoder with continuous frames to extract spatiotemporal features, construct motion perception representation through 1×1 convolution dimensionality reduction and feature splicing, and output the relative posture matrix T through 3×3 convolution layer and ReLU activation. t→s ; Step S42: Based on the inverse depth map D generated in step S34 t and the relative pose matrix T output in step S41 t→s , reconstruct the target synthetic image using projection mapping and bilinear interpolation Step S43: synthesize the target image with I t Perform geometric consistency constraints and calculate photometric reprojection loss The formula is as follows: Where SSIM(·) represents the sum of pixel similarities, α is set to a constant, and the minimum reprojection loss is used to reduce occlusion and artifacts. The specific formula is as follows: Among them, I s Indicates the previous or next frame of the target image; In order to make the depth map consistent in smooth areas while retaining clear boundaries in edge areas, an edge-aware smoothness loss is calculated. The specific formula is as follows: where d * Indicates that the inverse depth is normalized. is the set parameter; Final loss The calculation formula is as follows: γ is set to a constant, and λ is set to a constant. Step S44: Model training iteratively optimizes parameters to minimize the loss function, learns the mapping relationship between the input image and the corresponding depth map, and trains until the preset maximum number of iterations is reached. At this time, the model performance is evaluated and the weight with the best prediction accuracy is selected as the final model.
9. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 8, characterized in that: In step S43, α is set to 0.85; γ is set to 1.2; and λ is set to 1e -3 .
10. The image depth estimation method based on geometric reconstruction upsampling and dynamic offset constraints according to claim 1, characterized in that: The estimation model is used to perform geometric reconstruction upsampling and dynamic offset constraint mechanisms in computer vision tasks for monocular depth estimation. The geometric reconstruction upsampling and dynamic offset constraint mechanisms are seamlessly integrated with the self-supervised depth estimation framework, and do not need to rely on real depth labels when modeling high-fidelity depth.