Multi-view three-dimensional depth reasoning method for aerial image of unmanned aerial vehicle

By employing a multi-view stereo depth inference network with frequency-aware feature extraction and dual regularization processing, the reconstruction error problem of UAV aerial images in complex scenes is solved, achieving higher accuracy in depth estimation and 3D reconstruction.

CN121053292APending Publication Date: 2025-12-02NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511140508.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing multi-view stereo depth reasoning methods for UAV aerial images suffer from large errors and low reconstruction accuracy when dealing with complex scenes, especially in areas with overlapping building structures, road networks, and vegetation.

Method used

A multi-view stereo depth inference network employing frequency-aware feature extraction and dual regularization mechanisms, including multi-scale frequency aliasing suppression, BiFormer visual Transformer, frequency adaptive feature aggregation, and dual regularization processing, improves the accuracy of feature extraction and depth estimation.

Benefits of technology

It significantly improves the depth inference accuracy of building edge areas in UAV aerial images, reduces reconstruction errors in complex scenes, and enhances the overall accuracy and efficiency of 3D reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053292A_ABST
    Figure CN121053292A_ABST
Patent Text Reader

Abstract

The invention provides a multi-view three-dimensional depth reasoning method for unmanned aerial vehicle aerial images, which comprises the steps of data preprocessing and input, frequency perception feature extraction, cost body construction, double regularization and depth map generation and loss calculation, and is used for a multi-view depth estimation task in an unmanned aerial vehicle aerial scene. A traditional method mainly depends on commercial software such as SURE, while a current deep learning method is insufficient in generalization ability and often generates more errors in a boundary region and a weak texture region when processing deep reasoning and reconstruction of multi-view aerial images of an unmanned aerial vehicle. According to the method, spatial features are extracted by adopting frequency adaptive feature extraction and BiFormer global feature modeling, and a depth map is generated and optimized by adopting a double regularization module, so that the problem of reconstruction errors caused by depth inference errors of a boundary region and a weak texture region is solved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] [Technical Field] This invention belongs to the field of computer vision and 3D reconstruction technology, specifically involving a multi-view stereo depth reasoning method based on frequency perception and dual regularization, which is particularly suitable for 3D reconstruction of long-range UAV aerial images.

[0002] [Background Technology] In recent years, 3D reconstruction technology based on unmanned aerial vehicle (UAV) aerial imagery, especially multi-view stereo (MVS) technology, has demonstrated significant value in multiple fields. MVS recovers the 3D geometric information of a scene from 2D images captured from different perspectives, providing efficient data support for applications such as digital airspace construction, urban planning, geological disaster monitoring, and cultural relic protection. Combining the advantages of UAV aerial photography—flexible operation, low cost, and high resolution—it can quickly acquire image data over large areas and use MVS technology for 3D reconstruction.

[0003] Traditional methods primarily rely on commercial software packages such as ContextCaptur, SURE, and Pix4D. However, these traditional solutions have significant limitations: they are prone to matching errors in complex scenes, and post-processing requires extensive manual correction, resulting in time-consuming and labor-intensive workflows. Meanwhile, deep learning has brought transformative breakthroughs to MVS applications. Algorithms represented by MVSNet have performed exceptionally well on standard benchmarks such as the DTU dataset, demonstrating not only the ability to autonomously learn scene features but also significantly improving reconstruction accuracy and computational efficiency. This data-driven approach significantly reduces reliance on manual intervention, providing a more intelligent solution for 3D reconstruction. Deep learning-based MVS methods are gradually becoming a research frontier with broad application prospects. Despite these advancements, current deep learning methods still have inherent limitations when processing deep inference and reconstruction of UAV multi-view aerial imagery. First, most existing MVS algorithms are trained on images acquired in laboratories, resulting in insufficient model generalization ability when transferred to long-range aerial imagery captured by UAVs. Furthermore, compared to near-range deep reasoning and reconstruction, deep reasoning from the perspective of drone aerial photography, with its vast spatial coverage, faces unique challenges: the complex interactions of building structures, road networks, terrain features, and vegetation create intricate and interwoven boundaries. Existing methods often generate more errors in these boundary areas, especially when dealing with the complex spatial relationships between various geographic elements in large-scale aerial scenes.

[0004] In summary, the research and development of a multi-view stereo depth inference method for UAV aerial images, which reduces reconstruction errors caused by errors at boundaries and weak texture regions during depth estimation, has strong practical significance. [Summary of the Invention]

[0005] This invention provides an innovative multi-view stereo depth inference network architecture. This network significantly improves depth estimation accuracy in complex scenes through frequency-aware feature extraction and a dual regularization mechanism. It addresses reconstruction errors caused by intricate and overlapping boundaries formed by complex interactions, and is used for 3D reconstruction of UAV aerial images. The specific technical solution includes the following steps:

[0006] Step 1. Data Preprocessing and Input

[0007] The drone aerial images from multiple perspectives are normalized and resized to a uniform resolution. A copy is then made and downsampled to half the uniform resolution.

[0008] Step 2. Multi-scale frequency-aware feature extraction

[0009] 2.1 Multi-scale frequency aliasing suppression feature extraction

[0010] To suppress frequency aliasing during feature extraction, a UNet network is preferably used as the basic feature extraction architecture. The following components are integrated sequentially into each layer of the UNet encoder: 1) an anti-aliasing filter: deployed before each convolutional downsampling operation to effectively suppress frequency aliasing during feature extraction; 2) a standard convolutional layer: performing feature extraction; and 3) a frequency fusion module: placed after each convolutional downsampling operation to perform frequency-adaptive weighted fusion of the feature maps. These components together constitute the Frequency Alising Reduced Convolution module (FARC). The FARC is then used to perform frequency aliasing-suppressed feature extraction on UAV aerial images.

[0011] 2.2 Global Feature Modeling

[0012] To model the global features of multi-view aerial images, the preferred approach is to use the BiFormer visual Transformer to capture and replicate the long-range dependencies in the downsampled UAV aerial multi-view images. The global features extracted by BiFormer are progressively fused through multi-level convolutional layers. Finally, the global features output by the BiFormer branch are deeply fused with the local features extracted by the FARC module to form a comprehensive feature representation that combines details and contextual information.

[0013] 2.3 Frequency-Adaptive Feature Aggregation

[0014] Multiscale features from the encoder (Where F3 represents shallow high-resolution features and F1 represents deep semantic features) The features are processed sequentially by standard convolution, three frequency-adaptive dilated convolutions, and then another standard convolution. The standard convolution uses a 3×3 kernel to extract local spatial features; the frequency-adaptive dilated convolution uses a 3×3 kernel and dynamically adjusts the receptive field to extract features based on the frequency characteristics of the region; the initial features are fused with the processed features through skip connections, specifically including: 1) Main path skip connection: directly connected from the input of the first standard convolution to the output of the last standard convolution; 2) Frequency path skip connection: directly connected from the input of the first dilated convolution to the input of the last dilated convolution.

[0015] Step 3. Feature Construction

[0016] The lowest resolution multi-scale features processed by the frequency adaptive feature aggregation module are transformed into the frustum space of the reference view using a differentiable homography transformation; an initial cost volume pair is constructed between the feature maps of each source view and the reference view by calculating the feature dot product; and a cascade strategy is used to construct multi-scale cost volumes from coarse to fine.

[0017] Step 4. Double Regularization

[0018] To address the problem of deteriorated matching costs caused by early cost aggregation in existing technologies, this invention proposes a cost volume optimization method based on dual regularization, comprising two stages: pre-regularization and post-regularization. The specific implementation scheme is as follows:

[0019] 4.1. Pre-regularization stage

[0020] 1) Cost volume compression: After performing correlation calculation on paired cost volumes, the preferred method is to use 3D convolution for channel compression to reduce computational complexity.

[0021] 2) Single-view regularization: In order to perform regularization between a single source view and a reference view, preferably, an improved 2D U-Net network (integrating dealiasing filter and frequency mixing module) is used to perform preliminary regularization on the single-view cost volume, generate a probability volume and regress it into a depth map;

[0022] 3) Supervision mechanism: In order to supervise the difference between the predicted depth map and the ground truth to guide the learning, preferably, Smooth L1 Loss is used to supervise the depth map prediction in the pre-regularization stage to ensure the accuracy of single-view optimization.

[0023] 4.2. Post-regularization stage

[0024] 1) Multi-view fusion: In order to fuse the pre-regularization results of multiple views, preferably, N-1 pairs of pre-regularized cost volumes are spliced ​​along the view dimension to form the cost volume of multi-view fusion;

[0025] 2) Cross-view optimization: To further perform cross-view regularization on the pre-regularization results of merging multiple views, preferably, ensemble product optimization is used.

[0026] Lightweight recurrent convolutional networks with gated recurrent units are used for cross-view regularization to enhance multi-view consistency.

[0027] Step 5. Depth Map Generation and Loss Calculation

[0028] To obtain the depth map and calculate the loss, preferably, the depth map output by the regularized network after the third stage is selected as the final depth prediction, and a cascaded weighted loss is used, L. i Let λ be the depth prediction loss for the i-th cascade stage; i λ1 = 0.5, λ2 = 1.0, λ3 = 2.0; V is the number of input views; D pre For the predicted depth map, D gt For the true depth map

[0029] [Attached Image Description]

[0030] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention, but do not constitute a limitation thereof. Wherein:

[0031] Figure 1 The flowchart illustrates the overall process steps of a multi-view stereo depth reasoning method for UAV aerial images.

[0032] Figure 2 The diagram illustrates the overall architecture of a multi-view stereo depth inference method for UAV aerial images.

[0033] Figure 3 The diagram shows the structure of the frequency adaptive feature aggregation module and its workflow.

[0034] Figure 4 The flowchart of the double regularization module illustrates the specific implementation of the double regularization module.

[0035] Figure 5 The visualization comparison chart of depth prediction results from UAV aerial images shows the comparison of visualization results of multi-view stereo depth estimation methods based on depth maps on the WHU-MVS dataset.

[0036] This method performs excellently in both three-view and five-view scenarios. In the three-view scenario, it achieves the smallest mean absolute error (MAE). Extending to five-view scenarios, its MAE is improved by 12.6% compared to the best mainstream methods. Furthermore, although previous methods achieved 98.1% on the <0.6m metric, this method further improves this metric to 98.5%. Compared to the cascaded baseline models Ada-MVS and MSRED-Net, this method maintains the best performance on the 3-interval metric. Figure 5 This method demonstrates that it maintains minimal error at both the building's edges and the overall area.

[0037] In summary, this invention proposes an innovative multi-view stereo vision depth estimation method. Compared with existing technologies, the beneficial effects of this invention are as follows:

[0038] 1) This invention proposes a frequency-adaptive feature aggregation module, which, together with the dealiasing filter and the frequency mixing module, constitutes a frequency-aware feature extraction network, significantly improving the depth inference accuracy of building edge regions in aerial images. Specifically, through a frequency-adaptive mechanism, optimized extraction and fusion of features from different frequency bands are achieved.

[0039] 2) This invention introduces a parallel ViT branch in the feature extraction stage, making full use of its global feature processing capabilities to significantly improve the accuracy of depth map estimation.

[0040] 3) This invention proposes a dual regularization module that uses a delayed cost aggregation strategy to replace the traditional early variance and cost weighting method. This module includes: pre-regularization processing of cost for single views and post-regularization optimization of cost for multiple views. Through this dual optimization mechanism, the overall depth estimation error is effectively reduced.

Detailed Implementation Methods

[0041] This invention aims to illustrate its specific implementation based on a case study of long-tail knowledge graph entity alignment in the industrial field. Figure 1 This is a flowchart of a knowledge graph entity alignment method provided in Embodiment 1 of the present invention.

[0042] Example 1:

[0043] The following is a flowchart of the multi-view depth estimation method for UAV aerial images provided in Embodiment 1 of the present invention, including the following steps:

[0044] Step 1: Data Preprocessing and Input

[0045] - Acquire multi-view image data from drone aerial photography, including 1 reference image and N-1 source images;

[0046] - Normalize the input image and resize it to a uniform resolution (e.g., 768×384 pixels);

[0047] - Copy the multi-view image data and downsample it.

[0048] Step 2: Frequency-aware feature extraction

[0049] 2.1 Multi-scale feature extraction using a multi-scale frequency-aware feature extraction network:

[0050] - Run an anti-aliasing filter before downsampling each layer of the encoder to suppress frequency aliasing effects;

[0051] - Run the frequency mixing module after downsampling at each layer of the encoder to achieve frequency domain adaptive feature fusion;

[0052] 2.2 Running BiFormer branches in parallel:

[0053] - Use BiFormer to capture global long-distance dependencies;

[0054] - Feature dimensions are aligned using 3×3 convolutional layers;

[0055] 2.3 The features of UNet and BiFormer branches are fused at the highest level.

[0056] Step 3: Cost Body Construction

[0057] 3.1 Differentiable homography transformation is used to transform the features of the source view to the view cone space of the reference view;

[0058] 3.2 Construct the initial cost volume by calculating the feature dot product;

[0059] 3.3 A cascaded strategy is adopted to construct a multi-scale cost volume from coarse to fine:

[0060] - Phase 1: 48 depth hypotheses, interval (d) max -d min ) / 48;

[0061] - Phase Two: 32 assumptions, with fixed intervals of 0.2m;

[0062] - Phase 3: 8 assumptions, with a fixed interval of 0.1m.

[0063] Step 4: Double Regularization

[0064] 4.1 Pre-regularization stage:

[0065] - Compress cost volume channels using 3D convolution;

[0066] - Single-view cost volume regularization using a 2D U-Net with integrated dealiasing filters;

[0067] - Output the probability volume and regress it into a depth map;

[0068] 4.2 Post-regularization stage:

[0069] - Consolidate N-1 pre-regularized cost bodies along the view dimension;

[0070] - Use a lightweight recurrent convolutional network (including the RED module) for cross-view regularization;

[0071] 4.3 Loss Calculation:

[0072] - Use cascaded weighted smoothing L1 loss (weight coefficients λ1 = 0.5, λ2 = 1.0, λ3 = 2.0);

[0073] - Jointly optimize the pre-regularization and post-regularization outputs at each stage.

[0074] Step 5: Depth Map Generation and Loss Calculation

[0075] 5.1 The depth map output by the regularized network after the third stage is selected as the final result;

[0076] Example Parameter Configuration

[0077] - Training settings:

[0078] - Optimizer: RMSProp (initial learning rate 0.001)

[0079] - Batch size: 1

[0080] - Training cycle: 30 epochs

[0081] - Reasoning settings:

[0082] - Number of input views: 5

[0083] -Processing resolution: Original image size

[0084] - Computing device: NVIDIA 3090 GPU

[0085] This method achieves an MAE of 7.6m on the WHU-MVS dataset, a 13.1% improvement over the baseline; in this embodiment, Figure 4 This image shows the depth prediction result obtained by the multi-view depth estimation method for UAV aerial images proposed in this invention. An example illustrates that the method of this invention can effectively estimate the depth of UAV aerial images.

[0086] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-view stereo depth reasoning method for UAV aerial images, characterized in that... Includes the following steps: Data preprocessing and input, multi-scale frequency-aware feature extraction, overall feature construction, depth map generation and loss calculation.

2. The multi-view stereo depth reasoning method for UAV aerial images according to claim 1, characterized in that... In step 1, the multi-view aerial images taken by the drone are normalized and resized to a uniform resolution. A copy is then made and downsampled to half the uniform resolution.

3. The multi-view stereo depth reasoning method for UAV aerial images according to claim 1, characterized in that... In step 2, the frequency aliasing suppression convolution module FARC is first used to extract multi-scale frequency aliasing suppression features from the multi-view UAV aerial images, and the BiFormer visual Transformer is used to capture the long-distance dependencies in the downsampled UAV aerial multi-view images for global feature modeling. Finally, the global features output by the BiFormer branch are fused with the local features extracted by the FARC module, and the fused features are processed by the frequency adaptive feature aggregation module FAA.

4. The multi-view stereo depth reasoning method for UAV aerial images according to claims 1 and 3, characterized in that... The Frequency Altering Convolutional Module (FARC) in step 2 for multi-scale frequency-aware feature extraction includes:

1. An anti-aliasing filter: deployed before each convolutional downsampling operation to effectively suppress frequency aliasing during feature extraction; 2. A standard convolutional layer: performing feature extraction; 3. A frequency fusion module: set after each convolutional downsampling operation to perform frequency-adaptive weighted fusion of the feature maps; the above components together constitute the Frequency Altering Convolutional Module.

5. The multi-view stereo depth reasoning method for UAV aerial images according to claims 1 and 3, characterized in that... In step 2, the BiFormer visual Transformer is used to capture long-distance dependencies in the multi-view images of drone aerial photography after copying and downsampling; the global features extracted by BiFormer are progressively fused through multi-level convolutional layers; finally, the global features output by the BiFormer branch are deeply fused with the local features extracted by the FARC module to form a comprehensive feature representation that combines details and contextual information.

6. The multi-view stereo depth reasoning method for UAV aerial images according to claims 1 and 3, characterized in that... In step 2, the frequency-adaptive feature aggregation module (FAA) consists of a standard convolution, three frequency-adaptive dilated convolutions, and a standard convolution. The standard convolution uses a 3×3 kernel to extract local spatial features; the frequency-adaptive dilated convolution uses a 3×3 kernel and dynamically adjusts the receptive field to extract features based on the frequency characteristics of the region. Skip connections are used to fuse initial features with processed features. Specifically, they include:

1. Main path skip connections: directly connecting from the first standard convolution input to the last standard convolution output; 2. Frequency path skip connections: directly connecting from the first dilated convolution input to the last dilated convolution input.

7. The multi-view stereo depth reasoning method for UAV aerial images according to claim 1, characterized in that... In step 3, the lowest resolution multi-scale features processed by the frequency adaptive feature aggregation module are transformed into the frustum space of the reference view using differentiable homography transformation; an initial cost volume pair is constructed between the feature maps of each source view and the reference view by calculating the feature dot product; and a cascade strategy is used to construct the multi-scale cost volume from coarse to fine.

8. The multi-view stereo depth reasoning method for UAV aerial images according to claim 1, characterized in that... In step 4, after calculating the correlation of the cost pairs, a pre-regularization module is used to regularize the cost pairs between a single source view and a reference view: 3D convolution is used to compress channels to reduce computational complexity, and a 2D U-Net network with integrated dealiasing filter and frequency mixing module is used to pre-regularize the cost pairs of a single view, generating a probability volume and regressing it into a depth map; a post-regularization module is used to optimize the cross-view of multiple pre-regularized view cost pairs: multiple pre-regularized view cost pairs are concatenated along the view dimension, and 2DU-Net and convolutional gated recurrent units are used to perform cross-view regularization on the concatenated view cost pairs.

9. The multi-view stereo depth reasoning method for UAV aerial images according to claim 1, characterized in that... In step 5, the depth map output by the regularized network after the third stage is selected as the final depth prediction. A cascaded weighted loss is used, and the error between the predicted depth map and the true depth map is calculated using smoothed L1 loss for each stage. i Let λ be the depth prediction loss for the i-th cascade stage; i λ1 = 0.5, λ2 = 1.0, λ3 = 2.0; V is the number of input views; D pre For the predicted depth map, D gt For the true depth map, the loss function is expressed as: