Low-resolution monocular depth estimation method based on wavelet domain feature enhancement

By combining wavelet domain feature enhancement and self-supervised training, the problems of detail loss and misjudgment of blurred areas in depth estimation of low-resolution images are solved, and high-precision depth reconstruction is achieved.

CN120655692APending Publication Date: 2025-09-16FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510745691.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively restoring high-frequency details and correcting misjudgments of low-frequency blurred areas in depth estimation of low-resolution images, resulting in insufficient prediction accuracy.

Method used

The wavelet domain feature enhancement method is used to decompose image features into high-frequency sub-bands and low-frequency sub-bands, and edge enhancement and fuzzy mapping processing are performed respectively. Combined with cross-scale consistency constraints, the model is optimized through self-supervised training.

Benefits of technology

It significantly improves the depth estimation accuracy of low-resolution images, enhances the model's robustness to hardware downsampling and motion blur, reduces computational overhead, and is suitable for mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655692A_ABST
    Figure CN120655692A_ABST
Patent Text Reader

Abstract

The invention provides a low-resolution monocular depth estimation method based on wavelet domain feature enhancement, and the method comprises the steps: carrying out the multi-scale enhancement processing of an input image, and generating image groups with different resolutions; extracting multi-scale features of the image group, and decomposing the multi-scale features into a high-frequency sub-band containing edge details and a low-frequency sub-band containing a smooth region through wavelet transform; carrying out edge enhancement processing on the high-frequency sub-band to enhance detail information; gradient symbol consistency and a cross-frame reprojection error are calculated for the low-frequency sub-band, and a fuzzy mapping value is generated to correct a fuzzy region; inputting the processed features into a depth decoder, and generating an inverse depth map of the scene through progressive up-sampling; predicting a camera pose by combining a pose estimation network, and reconstructing a target composite image; network parameters are optimized through a loss function, and end-to-end monocular depth estimation is completed; the loss function at least comprises cross-scale depth consistency loss and is used for constraining structural similarity and numerical value consistency of depth maps with different resolutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a low-resolution monocular depth estimation method based on wavelet domain feature enhancement. Background Art

[0002] Monocular depth estimation, a core task in computer vision, aims to recover the three-dimensional geometric information of a scene from a single two-dimensional image. Compared to hardware solutions that rely on multiple cameras (such as stereo vision) or lidar, monocular methods require only a single standard camera, significantly reducing hardware cost and device complexity. Therefore, they have unique advantages in mobile applications (such as smartphones and drones) and lightweight embedded systems. In recent years, with breakthroughs in deep learning technology, monocular depth estimation models based on convolutional neural networks (CNNs) have made significant progress, demonstrating high prediction accuracy in outdoor scenes (such as the KITTI dataset) and indoor environments (such as the NYU Depth V2 dataset).

[0003] However, existing research primarily focuses on depth estimation from high-resolution input images. These methods typically assume that the input image possesses rich texture details and clear edge structures, providing sufficient visual cues for the model. For example, they employ encoder-decoder architectures (such as U-Net and HRNet) to extract multi-scale features, combined with skip connections to preserve high-frequency details, or employ attention mechanisms to enhance feature responses in key regions. While these methods perform well in high-resolution scenarios, their generalization capabilities to low-resolution images are severely limited.

[0004] Low-resolution images are widely used in practical applications, mainly due to the following factors:

[0005] Hardware constraints: Security cameras, vehicle recorders, and other devices are often limited by cost and power consumption and often use low-resolution sensors.

[0006] Transmission and storage bottlenecks: Mobile network bandwidth limitations (such as 4G / 5G real-time video streaming) and edge device storage capacity force image downsampling;

[0007] Environmental interference: Severe weather conditions such as fog, haze, rain, snow, or motion blur further reduce image recognizability.

[0008] Under low-resolution conditions, existing depth estimation models face two core challenges:

[0009] Detail loss problem: Image downsampling causes severe attenuation of high-frequency information (such as edges and textures), making it difficult for traditional upsampling methods (such as bilinear interpolation) to restore the true geometric structure;

[0010] Misjudgment of fuzzy areas: Low-frequency areas (such as smooth walls and the sky) lack effective features and are easily mispredicted by the model as planes with no depth change, thus destroying the continuity of the scene.

[0011] Current attempts to improve low-resolution depth estimation mainly include:

[0012] Multi-scale feature fusion: Contextual information of different scales is fused through the input image pyramid or feature pyramid network (FPN), but the recovery effect of high-frequency details is limited;

[0013] Super-resolution assistance: Jointly train the depth estimation and super-resolution branches, but the additional computational overhead makes it difficult to meet real-time requirements;

[0014] Self-supervised optimization: It uses the photometric consistency constraint of adjacent frames to reduce the dependence on the real depth label, but it is easily disturbed by blurred areas at low resolution, resulting in failure of reprojection error.

[0015] While these methods partially alleviate the low-resolution problem, they still lack a systematic approach to modeling the imbalance in time-frequency domain features. Simultaneously enhancing the robust representation of high-frequency details while correcting for blurred predictions in low-frequency regions is a key bottleneck in improving low-resolution depth estimation accuracy. Summary of the Invention

[0016] To address the common problems of edge detail loss and inaccurate prediction of blurred regions in depth estimation from low-resolution images, this paper proposes an innovative method that combines wavelet-domain feature enhancement with self-supervised optimization. The core of this method is to leverage the time-frequency analysis capabilities of the wavelet transform to decompose image features into high-frequency edge details and low-frequency smooth regions, designing enhancement strategies for each, and incorporating cross-scale consistency constraints to improve the model's robustness to resolution variations.

[0017] First, the multi-scale image features are decomposed into high-frequency subbands (containing details such as contours and textures) and low-frequency subbands (containing smooth regions) using a discrete wavelet transform. For the high-frequency subbands, thresholding is used to suppress noise interference, and coefficient scaling is used to enhance edge response strength, significantly improving the attenuation of key geometric information caused by downsampling. For the low-frequency subbands, a gradient sign consistency detection mechanism is introduced: when the product of the horizontal and vertical gradient signs of a pixel is negative, it is determined to belong to a blurred region. Furthermore, combined with cross-frame reprojection errors, the optimal reference frame is dynamically selected and the blur map value is calculated to quantify the depth uncertainty of the low-frequency region, thereby correcting the misjudgment of smooth regions by traditional methods.

[0018] To address the challenge of multi-scale feature alignment at low resolution, this solution designs a cross-scale depth consistency loss. By constraining the structural similarity and numerical differences between depth maps of different resolutions, the model is forced to maintain the consistency of the scene structure when generating high, medium, and low resolution depth maps. At the same time, the joint pose estimation network predicts the camera motion pose, and uses the reprojection error to construct a dynamic reference frame selection mechanism to effectively alleviate motion blur interference. The entire training process is implemented in a self-supervised manner: combining the photometric reprojection loss to ensure the authenticity of the synthesized image, supplemented by the edge-aware smoothness loss to optimize the smooth area while retaining clear boundaries, and finally balancing the various optimization objectives through a weighted loss function.

[0019] The technical solution specifically adopted by the present invention to solve the technical problem is:

[0020] A low-resolution monocular depth estimation method based on wavelet domain feature enhancement, comprising:

[0021] Perform multi-scale enhancement on the input image to generate image groups with different resolutions, including:

[0022] A high-resolution image is generated by enlarging the original image at a set scaling ratio and then cropping it;

[0023] Keep the original size and apply perturbations to generate medium-resolution images;

[0024] Generate a low-resolution image by downsampling and padding at a set scaling ratio;

[0025] Extracting multi-scale features of the image group and decomposing them into high-frequency sub-bands containing edge details and low-frequency sub-bands containing smooth areas through wavelet transform;

[0026] Perform edge enhancement processing on high-frequency sub-bands to enhance detail information;

[0027] Calculate the gradient sign consistency and cross-frame reprojection error for the low-frequency subband to generate a blur map value to correct the blurred area;

[0028] The processed features are input into the depth decoder to generate the inverse depth map of the scene through progressive upsampling;

[0029] The joint pose estimation network predicts the camera pose and reconstructs the target composite image;

[0030] End-to-end monocular depth estimation is completed by optimizing network parameters through a loss function; the loss function at least includes a cross-scale depth consistency loss, which is used to constrain the structural similarity and numerical consistency of depth maps of different resolutions.

[0031] Furthermore, the wavelet transform adopts discrete wavelet transform to decompose the features into a low-frequency sub-band and three high-frequency sub-bands: horizontal, vertical, and diagonal.

[0032] Furthermore, the edge enhancement of the high frequency sub-band includes: performing threshold processing on the high frequency coefficients, and enhancing the edge response strength by coefficient scaling.

[0033] Furthermore, the generation of the fuzzy mapping value includes:

[0034] If the product of the gradient signs of a pixel in the horizontal and vertical directions is negative, it is determined to be a fuzzy area;

[0035] The blur map value is calculated by dynamically selecting the reference frame with the smallest reprojection error and combining the blur information of adjacent frames.

[0036] Furthermore, the cross-scale depth consistency loss constrains the consistency of depth maps of different resolutions through the structural similarity index and L1 norm.

[0037] Furthermore, the low-resolution image is generated by downsampling and then symmetrically filling the four quadrants.

[0038] Furthermore, the loss function is a weighted loss function, further comprising:

[0039] Photometric reprojection loss, used to constrain the similarity between the synthesized image and the original image;

[0040] Edge-aware smoothness loss for smoothness constraints on weighted depth gradients.

[0041] And, a low-resolution monocular depth estimation system based on wavelet domain feature enhancement, comprising:

[0042] The image processing unit is used to perform multi-scale enhancement processing on the input image to generate image groups with different resolutions, including:

[0043] A high-resolution image is generated by enlarging the original image at a set scaling ratio and then cropping it;

[0044] Keep the original size and apply perturbations to generate medium-resolution images;

[0045] Generate a low-resolution image by downsampling and padding at a set scaling ratio;

[0046] a feature enhancement unit configured to extract multi-scale features of the image group and decompose them into high-frequency sub-bands containing edge details and low-frequency sub-bands containing smooth areas through wavelet transform; perform edge enhancement processing on the high-frequency sub-bands to enhance detail information; calculate gradient sign consistency and cross-frame reprojection error on the low-frequency sub-bands to generate blur map values ​​to correct blurry areas;

[0047] The depth estimation unit is used to input the processed features into the depth decoder and generate the inverse depth map of the scene through progressive upsampling;

[0048] The pose estimation unit is used to predict the camera pose in conjunction with the pose estimation network and reconstruct the target synthetic image;

[0049] An optimization unit is used to optimize network parameters through a loss function to complete end-to-end monocular depth estimation; the loss function at least includes a cross-scale depth consistency loss, which is used to constrain the structural similarity and numerical consistency of depth maps of different resolutions.

[0050] And, a computer device includes a memory, a processor and a computer program stored in the memory, and the processor implements the above method when executing the computer program.

[0051] A non-transitory computer-readable storage medium stores a computer program, which implements the method described above when executed by a processor.

[0052] Compared with the prior art, the present invention and its preferred embodiments have at least the following beneficial effects:

[0053] 1. Breakthrough the bottleneck of time-frequency domain modeling for low-resolution depth estimation

[0054] Through the wavelet domain dual-module coupling mechanism, the coordinated optimization of high-frequency edge details and low-frequency blurred areas is achieved for the first time:

[0055] High-frequency sub-band enhancement significantly improves edge contour restoration capabilities, effectively alleviating the texture loss problem caused by downsampling;

[0056] The low-frequency fuzzy mapping value dynamically corrects the depth prediction error in the smooth area, avoiding the misjudgment of areas such as the sky and wall by traditional methods.

[0057] 2. Enhance cross-scale scenario adaptability

[0058] The cross-scale depth consistency loss forces depth maps of different resolutions to maintain structural uniformity:

[0059] Solve the problem of multi-scale feature alignment distortion and improve the model's robustness to hardware downsampling;

[0060] Combined with the dynamic reference frame selection mechanism, the interference of motion blur on reprojection error is effectively suppressed.

[0061] 3. Reduce deployment threshold and computing overhead

[0062] The self-supervised training paradigm eliminates the reliance on real depth labels and supports training on massive unlabeled data.

[0063] The progressive upsampling decoder avoids the information loss of traditional interpolation and takes into account both accuracy and real-time requirements.

[0064] 4. Expanding industrial application potential

[0065] Lightweight architecture adapted to mobile devices (such as surveillance cameras and in-vehicle systems);

[0066] Achieve high-precision depth reconstruction in low-resolution scenes such as video surveillance and satellite remote sensing. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0068] Figure 1 This is a schematic diagram of the process and principle of an embodiment of the present invention. DETAILED DESCRIPTION

[0069] In order to make the features and advantages of the present invention more clearly understood, the following embodiments are given for detailed description:

[0070] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs.

[0071] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0072] like Figure 1 As shown, this embodiment provides a low-resolution monocular depth estimation method based on wavelet domain feature enhancement, which can effectively solve the problems of edge detail loss and inaccurate prediction of blurred areas in depth estimation in low-resolution scenes. This method uses the wavelet time-frequency domain perception optimization framework to achieve high-precision depth prediction for low-resolution images. First, an image dataset is collected to form a training dataset as input. After the multi-scale features are extracted by the depth encoder, the discrete wavelet transform is used to decompose the features into high-frequency and low-frequency sub-bands. The multi-scale wavelet edge enhancement and wavelet blur mapping modules are used to improve the capture of edge details and reduce the depth estimation error caused by feature blur. Subsequently, the multi-scale features are progressively upsampled by the depth decoder to gradually restore the depth information of the scene, and finally a high-precision depth map is output. The entire network adopts a self-supervised training method. The pose estimation network predicts the camera pose and constructs a photometric consistency loss constraint to optimize the network parameters. End-to-end training can be achieved without the need for real depth labels.

[0073] The implementation process specifically includes the following steps:

[0074] Step S1: Collect the original image dataset and generate high, medium and low resolution image groups through data enhancement to input into the deep decoder.

[0075] Step S2: Multi-scale feature extraction is performed in the deep encoder, and time-frequency information is enhanced through wavelet domain analysis. The multi-scale wavelet edge enhancement module progressively enhances high-frequency details in layers 2 to 4 of the encoding phase, optimizing depth estimation. The wavelet blur mapping module mitigates depth estimation errors caused by feature blur through supervised learning of low-frequency blur features.

[0076] Step S3: The multi-scale features are progressively upsampled through the depth decoder, and the joint pose estimation network predicts the camera pose to gradually restore the depth information of the scene and output a high-precision depth map.

[0077] Step S4: Perform iterative training according to the specified training parameters, update the model parameters by calculating the relevant losses, continuously save the optimal model according to the verification error rate and accuracy, and finally save the model weight with the highest prediction accuracy.

[0078] In this embodiment, step S1 specifically includes the following steps:

[0079] Step S11: Use the public image dataset KITTI dataset as a training dataset, perform data preprocessing, and remove images with poor visual effects.

[0080] Step S12: Read the original input frame image ,in represents the number of feature channels, Represents the height of the image, Represents the height of the image and performs multi-scale enhancement: by sampling scaling ratio Enlarge the image and randomly crop it to generate a high-resolution image ;Retain the original size and apply mixed perturbations to generate medium-resolution images ; By sampling scaling After downsampling, it is expanded to a low-resolution image by four-quadrant symmetric padding .

[0081] In this embodiment, step S2 specifically includes the following steps:

[0082] Step S21: In the first stage of the deep encoder, discrete wavelet transform (DWT) is used to transform the multi-scale feature maps 、 、 Decomposed into multiple sub-bands, each scale feature map includes a low-frequency sub-band and three high-frequency sub-bands (horizontally, vertically, and diagonally). The edge information in the high-frequency sub-band is enhanced through threshold processing and coefficient enhancement, and then the enhanced sub-band is reconstructed into a time domain feature map through inverse wavelet transform (IDWT) to obtain a fused enhanced feature map. .

[0083] Step S22: The obtained fusion features are input into the 2-4 stages of the deep encoder, and sequentially pass through two convolution layers with a convolution kernel size of 6×6 and one with a convolution kernel size of 8×8 to capture a wide range of context information while keeping the size of the output feature map unchanged. After processing by the dilated convolution module, the output feature map It can be expressed as:

[0084]

[0085] in, Represents point-by-point convolution operation, followed by activation. is the batch normalization layer.

[0086] Step S23: For the information of the low-frequency area, calculate the gradient sign consistency of each pixel with its adjacent pixels in the horizontal and vertical directions. If the product sign is negative, it is determined to be a blurred area and step S24 is executed. Otherwise, it is not executed.

[0087] Step S24: Calculate the reprojection error between different frames for each pixel in the blurred area, and measure the Frame pixels in the reference frame ( ) in the matching degree. Pixel depth information in the frame , the projection function reprojects the pixel to the corresponding position at different time points, expressed as:

[0088]

[0089] in, Indicates the The projection function of the frame, It is Depth through the frame The calculated 3D coordinates.

[0090] Then, the reprojection error is calculated is the current pixel With the Frame reprojected pixels The distance between is defined as:

[0091]

[0092] Select the frame with the smallest error As the reference frame of the current pixel, it is represented as follows:

[0093]

[0094] Select target frame blur information and the blurred information obtained by reprojection of adjacent frames The maximum value in the pixel The fuzzy mapping value of , realizing the mapping from pixel space to fuzzy value space.

[0095]

[0096] In this embodiment, step S3 specifically includes the following steps:

[0097] Step S31: The depth decoder inherits the multi-scale features of the depth encoder, combines the jump connection to gradually fuse the low-scale features, while maintaining the high-resolution feature representation, and restores the feature resolution step by step through upsampling operations to generate an inverse depth map .

[0098] Step S32: Take the continuous frames as input to the posture network encoder to extract spatiotemporal features, construct motion perception representation through 1×1 convolution dimensionality reduction and feature splicing, and output the relative posture matrix through 3×3 convolution layer and ReLU activation. .

[0099] Step S33: Based on the inverse depth map generated in step S31 and the relative pose matrix output in step S32 , reconstruct the target synthetic image using projection mapping and bilinear interpolation .

[0100] Step S34: Jointly optimize the depth and pose networks to obtain the target composite image and Perform geometric consistency constraints.

[0101] In this embodiment, step S4 specifically includes the following steps:

[0102] Step S41: Calculate photometric reprojection loss , the formula is as follows:

[0103]

[0104] in, represents the sum of pixel similarities, Set to 0.85. In order to reduce occlusion and artifacts, it is necessary to use the minimum reprojection loss. The specific formula is as follows:

[0105]

[0106] in, Indicates the previous or next frame of the target image.

[0107] In addition, the depth information of the same scene should be consistent at different resolutions. , calculate cross-scale depth consistency loss , the specific formula is as follows:

[0108]

[0109] Step S42: To make the depth map consistent within the smooth area while retaining clear boundaries in the edge area, calculate the edge-aware smoothness loss , the specific formula is as follows:

[0110]

[0111] in Indicates that the inverse depth is normalized. 、 The parameters to be set.

[0112] Step S43: Final loss The calculation formula is as follows:

[0113]

[0114] Set to 1.0, Set to , Set to 1.0.

[0115] Step S44: The model is trained through continuous iteration to find the optimal model parameters to minimize the loss function and learn the mapping relationship between the input image features and the corresponding depth map, thereby obtaining more accurate depth estimation results. When the number of iterations reaches the preset maximum number of iterations, the training process ends and an evaluation is performed, returning the model weight with the highest prediction accuracy.

[0116] In particular, this embodiment employs a self-supervised depth estimation framework, achieving high-precision modeling of low-resolution scenes without relying on real-world depth labels. Through a multi-scale wavelet decomposition mechanism, the self-supervised learning objective is decoupled into a dual path: high-frequency edge feature enhancement and low-frequency blur feature correction. High-frequency subband reconstruction is used to enforce geometric consistency constraints on image edges, while low-frequency subband blur mapping guides scene structure continuity learning, effectively overcoming the shortcomings of traditional self-supervised methods in low-resolution environments, such as detail loss and misjudgment of blurred areas. Combined with a cross-resolution feature alignment strategy, the implicit texture prior knowledge of high-resolution images can be transferred to the low-resolution depth estimation process, significantly improving the robustness of self-supervised training to image degradation. Compared to supervised learning methods that rely on manual annotation, this solution not only expands the training scale by leveraging massive amounts of unlabeled data, but also enhances the model's generalization ability for complex scenes through a wavelet time-frequency domain joint optimization mechanism. This demonstrates unique technical advantages in low-resolution depth reconstruction scenarios such as mobile visual perception and wide-area surveillance equipment.

[0117] The advantages of the above solution provided in this embodiment include at least:

[0118] 1. It can accurately and effectively perform monocular depth estimation on low-resolution images, achieve high accuracy, and show strong generalization ability and adaptability.

[0119] 2. The innovatively designed wavelet domain dual-module coupling mechanism realizes the joint optimization of high-frequency edge detail enhancement and low-frequency fuzzy feature supervised learning for the first time, breaking through the limitation of traditional methods that are difficult to take into account both time- and frequency-domain feature modeling.

[0120] 3. The wavelet decomposition and reconstruction strategy adopted separates and specifically processes high-frequency / low-frequency sub-band features through discrete wavelet transform, effectively solving the problem of missing details and visual blur caused by information loss in low-resolution images.

[0121] 4. The constructed self-supervised multi-frame joint optimization paradigm combines reprojection error to select dynamic reference frames and performs cross-scale depth consistency constraints, achieving robust training without the need for high-precision labeled data, significantly improving the model's practicality and feasibility of deployment in industrial scenarios.

[0122] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0123] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, performs the above-described method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0124] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0125] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.

[0126] The present invention is not limited to the above-mentioned optimal implementation mode. Anyone can derive various other forms of low-resolution monocular depth estimation methods based on wavelet domain feature enhancement under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.

Claims

1. A low-resolution monocular depth estimation method based on wavelet domain feature enhancement, characterized in that: include: Perform multi-scale enhancement on the input image to generate image groups with different resolutions, including: A high-resolution image is generated by enlarging the original image at a set scaling ratio and then cropping it; Keep the original size and apply perturbations to generate medium-resolution images; Generate a low-resolution image by downsampling and padding at a set scaling ratio; Extracting multi-scale features of the image group and decomposing them into high-frequency sub-bands containing edge details and low-frequency sub-bands containing smooth areas through wavelet transform; Perform edge enhancement processing on high-frequency sub-bands to enhance detail information; Calculate the gradient sign consistency and cross-frame reprojection error for the low-frequency subband to generate a blur map value to correct the blurred area; The processed features are input into the depth decoder to generate the inverse depth map of the scene through progressive upsampling; The joint pose estimation network predicts the camera pose and reconstructs the target composite image; End-to-end monocular depth estimation is completed by optimizing network parameters through a loss function; the loss function at least includes a cross-scale depth consistency loss, which is used to constrain the structural similarity and numerical consistency of depth maps of different resolutions.

2. The low-resolution monocular depth estimation method based on wavelet domain feature enhancement according to claim 1, characterized in that: The wavelet transform adopts discrete wavelet transform to decompose the feature into a low-frequency sub-band and three high-frequency sub-bands: horizontal, vertical, and diagonal.

3. The low-resolution monocular depth estimation method based on wavelet domain feature enhancement according to claim 2, characterized in that: The edge enhancement of the high frequency sub-band includes: performing threshold processing on the high frequency coefficients, and enhancing the edge response strength by coefficient scaling.

4. The low-resolution monocular depth estimation method based on wavelet domain feature enhancement according to claim 1, characterized in that: The generation of the fuzzy mapping value includes: If the product of the gradient signs of a pixel in the horizontal and vertical directions is negative, it is determined to be a fuzzy area; The blur map value is calculated by dynamically selecting the reference frame with the smallest reprojection error and combining the blur information of adjacent frames.

5. The low-resolution monocular depth estimation method based on wavelet domain feature enhancement according to claim 1, characterized in that: The cross-scale depth consistency loss constrains the consistency of depth maps of different resolutions through the structural similarity index and L1 norm.

6. The low-resolution monocular depth estimation method based on wavelet domain feature enhancement according to claim 1, characterized in that: The low-resolution image is generated by downsampling and then symmetrically filling the four quadrants.

7. The low-resolution monocular depth estimation method based on wavelet domain feature enhancement according to claim 1, characterized in that: The loss function is a weighted loss function, and further includes: Photometric reprojection loss, used to constrain the similarity between the synthesized image and the original image; Edge-aware smoothness loss for smoothness constraints on weighted depth gradients.

8. A low-resolution monocular depth estimation system based on wavelet domain feature enhancement, characterized in that: include: The image processing unit is used to perform multi-scale enhancement processing on the input image to generate image groups with different resolutions, including: A high-resolution image is generated by enlarging the original image at a set scaling ratio and then cropping it; Keep the original size and apply perturbations to generate medium-resolution images; Generate a low-resolution image by downsampling and padding at a set scaling ratio; a feature enhancement unit configured to extract multi-scale features of the image group and decompose them into high-frequency sub-bands containing edge details and low-frequency sub-bands containing smooth areas through wavelet transform; perform edge enhancement processing on the high-frequency sub-bands to enhance detail information; calculate gradient sign consistency and cross-frame reprojection error on the low-frequency sub-bands to generate blur map values ​​to correct blurry areas; The depth estimation unit is used to input the processed features into the depth decoder and generate the inverse depth map of the scene through progressive upsampling; The pose estimation unit is used to predict the camera pose in conjunction with the pose estimation network and reconstruct the target synthetic image; An optimization unit is used to optimize network parameters through a loss function to complete end-to-end monocular depth estimation; the loss function at least includes a cross-scale depth consistency loss, which is used to constrain the structural similarity and numerical consistency of depth maps of different resolutions.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.