Unmanned aerial vehicle image three-dimensional reconstruction method based on feature point extraction and reinforcement

By introducing multi-scale feature fusion and feature descriptor enhancement modules, the problem of insufficient robustness of deep learning feature point extraction algorithms in complex environments is solved, and the accuracy and robustness of UAV image 3D reconstruction are improved.

CN121053286APending Publication Date: 2025-12-02CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510914579.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing deep learning-based feature point extraction algorithms lack robustness in complex environments and have high matching error rates, leading to a decrease in the accuracy and robustness of 3D reconstruction of UAV images, especially in areas with large scale and large viewpoint changes, lighting changes, and repetitive textures.

Method used

A multi-scale feature fusion module and a feature descriptor enhancement module are introduced. The stability of feature point detection is improved by a multi-scale joint feature point detection module, and the spatial context information of the feature descriptor is enhanced by a Transformer network. The consistency and matching accuracy of feature point detection are improved by a learnable weight mechanism and relative position encoding.

Benefits of technology

It significantly improves the robustness and matching accuracy of feature point extraction in complex environments, reduces noise and distortion in 3D reconstruction, and enhances the accuracy and robustness of the reconstruction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053286A_ABST
    Figure CN121053286A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle image three-dimensional reconstruction method based on feature point extraction and enhancement, and relates to the field of image processing, and the method comprises the steps: obtaining an unmanned aerial vehicle image; constructing a feature point extraction and enhancement network based on a SuperPoint network and multi-scale feature fusion; the feature point extraction and enhancement network comprises a multi-scale joint feature point detection module and a feature descriptor enhancement module; training the feature point extraction and enhancement network through the unmanned aerial vehicle picture; obtaining a to-be-extracted unmanned aerial vehicle picture; feature point extraction is carried out on a to-be-extracted unmanned aerial vehicle picture through the trained feature point extraction and enhancement network; and completing three-dimensional reconstruction of the unmanned aerial vehicle image through the extracted feature points. According to the technical scheme, the three-dimensional reconstruction precision of the unmanned aerial vehicle image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular to a method for 3D reconstruction of UAV images based on feature point extraction and enhancement. Background Technology

[0002] Drones are typically equipped with high-resolution visible light cameras that fly along predetermined routes to capture images of target areas. Using highly overlapping image sequences as input, and employing algorithms such as Structured Bundle Method (SfM) and Multi-View Stereo Matching (MVS), they ultimately generate dense point clouds, mesh models, and texture maps. By rationally planning the flight path and adjusting the overlap, centimeter-level mapping requirements can be met.

[0003] In the 3D reconstruction process based on UAV imagery, feature point extraction and matching are crucial steps in ensuring the accuracy of the 3D model. Feature point extraction aims to find a stable and repeatable set of key points from multiple overlapping images and assign a descriptor to each key point so that they can be matched with each other from different viewpoints, thereby completing the camera motion estimation. Traditional feature point extraction algorithms include SIFT, SURF, and ORB. These methods rely on manually designed feature point extraction methods and can achieve certain results in terms of scale invariance and rotation invariance.

[0004] In recent years, the rapid development of deep learning technology has brought new ideas and methods to feature point extraction. By using convolutional neural networks for end-to-end feature learning of images, researchers have proposed a number of deep learning-based feature point detection and description algorithms, such as LIFT, SuperPoint, D2-Net, R2D2, and ContextDesc. These methods, through a self-training phase, input a large number of image pairs from different scenes into the network, automatically learning feature point detectors and descriptors under contrast loss or reprojection loss, achieving stronger generalization capabilities. Among them, the SuperPoint algorithm, as a typical deep network based on self-supervised learning, achieves a good balance between accuracy and speed in feature point detection tasks.

[0005] However, as drone aerial photography increasingly exposes complex situations such as varying lighting, changing viewpoints, scale differences, and large areas of repetitive textures, the matching stability and accuracy of traditional feature point extraction algorithms significantly decrease under extreme conditions. While deep learning methods have significantly improved feature point extraction performance in most conventional scenarios, they still face limitations in more complex environments and may lead to feature point matching errors, resulting in point cloud misalignment and inaccurate 3D reconstruction results. Specifically, let's take the relatively mature SuperPoint algorithm as an example: In scenarios with significant scale and viewpoint variations, while SuperPoint exhibits good robustness against rotation and local deformation, it may fall short in terms of feature point repeatability and consistency for large-scale and large-viewpoint changes. When objects in an image undergo significant deformation at different scales and viewpoints, the detected keypoints may fail to maintain a stable correspondence across different views, leading to matching errors. This affects the accuracy of camera pose estimation and triangulation, ultimately resulting in structural distortion or loss of detail in the 3D reconstruction model.

[0006] In environments with significant lighting variations, while SuperPoint can adapt to lighting changes to some extent, the local gradient information of the image changes significantly under extreme lighting conditions (such as strong shadows, highlights, or drastic changes in ambient light), thus affecting the detection and description of key points. Reduced matching accuracy leads to registration errors between images, affecting depth estimation and sparse point cloud generation during reconstruction, ultimately causing noise and blurred details in the reconstruction model in areas of transitional lighting.

[0007] In areas with large-scale repetitive imagery, SuperPoint may detect a large number of similar feature points due to the high similarity of textures and structures in the scene, thus causing ambiguity during the matching stage. Such repetition can easily lead to incorrect matching, causing deviations in camera pose estimation and triangulation. Ultimately, this may result in model distortion, increased noise, or local structural errors during 3D reconstruction, reducing the overall accuracy and robustness of the reconstruction results. Summary of the Invention

[0008] The purpose of this invention is to provide a method for 3D reconstruction of UAV images based on feature point extraction and enhancement, in order to solve the problems of insufficient robustness, high matching error rate and uneven distribution of key points in existing deep learning-based feature point extraction algorithms.

[0009] The above-mentioned objective of this application is achieved through the following technical solution: S1: Acquire drone images; S2: Based on the SuperPoint network and multi-scale feature fusion, a feature point extraction and enhancement network is constructed; the feature point extraction and enhancement network includes: a multi-scale joint feature point detection module and a feature descriptor enhancement module; S3: The feature point extraction and enhancement network is trained using drone images; S4: Obtain the drone image to be extracted; extract feature points from the drone image using the trained feature point extraction and enhancement network; complete the 3D reconstruction of the drone image using the extracted feature points.

[0010] Optionally, the multi-scale joint feature point detection module includes: a multi-scale feature extraction unit, a Softmax-Reshape module, and a fusion unit; the multi-scale feature extraction unit, the Softmax-Reshape module, and the fusion unit are connected sequentially. The multi-scale feature extraction unit is used to extract the original feature map of drone images. Feature map with dimensions W / 4 × H / 4 and a feature map with dimensions W / 2×H / 2 W and H represent the width and height of the feature map; The Softmax-Reshape module is used to reshape the original feature maps of drone images. Feature map and feature map Reconstructed into a thermal detection map , and ; The fusion unit is used for thermal detection maps , and By combining the results, a fusion heatmap is obtained.

[0011] Optionally, the Softmax-Reshape module is used to reshape the original feature map of the drone image. Feature map and feature map Reconstructed into a thermal detection map , and Specifically, it includes: The Softmax-Reshape module includes: Softmax units, Reshape units, and the ReLU activation function; the Softmax units are connected to the Reshape units; The Softmax unit is used to process the original feature map. Feature map and feature map Perform channel-level softmax operation; The Reshape unit is used to restore the feature map after the softmax operation to the original image size, resulting in three thermal detection maps.

[0012] Optionally, the feature descriptor enhancement module includes: a descriptor self-enhancement unit and a descriptor mutual enhancement unit; the descriptor self-enhancement unit is connected to the descriptor mutual enhancement unit.

[0013] Optionally, the descriptor self-enhancement unit specifically includes: Extract a D-dimensional descriptor for each pixel in the fused heatmap. ; Using the self-enhancing unit of the MLP network descriptor as the mapping function, the descriptor... Perform the transformation to obtain the descriptor. ; The MLP network is trained using a loss function based on Euclidean or Hamming distance constraints.

[0014] Optionally, the descriptor mutual enhancement unit specifically includes: Descriptor of N feature points Input a Transformer network and obtain the descriptor after cross-amplification, as shown in the following formula:

[0015] The Transformer network uses the AFT-Simple attention mechanism for computation; RoPE relative position encoding is introduced into the AFT-Simple mechanism.

[0016] An electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to enable the electronic device to perform a method for three-dimensional reconstruction of UAV images based on feature point extraction and enhancement.

[0017] A computer-readable storage medium storing instructions that, when executed, perform a method for 3D reconstruction of UAV imagery based on feature point extraction and enhancement.

[0018] The beneficial effects of the technical solution provided in this application are: Building upon the existing SuperPoint network structure, a multi-scale feature fusion module is introduced to extract multi-feature information at different scales. A learnable weighting mechanism adaptively adjusts the weights of features at each scale, enhancing the consistency of keypoint detection across large and small scales and varying scale scenarios. Simultaneously, to improve the representational power of feature points, relative position encoding is used to enhance the descriptors of feature points. Furthermore, a Transformer network is employed to understand the potential positional relationships of feature points within the same image, extracting spatial contextual information and thus improving the discriminative power of the feature descriptors. Attached Figure Description

[0019] The present application will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of an embodiment of this application; Figure 2 This is a schematic diagram of a multi-scale joint feature point detection network in an embodiment of this application; Figure 3 This is a schematic diagram of the descriptor enhancement module in an embodiment of this application; Figure 4 This is a schematic diagram of the electronic device structure in the embodiments of this application; Figure 5 This is a feature point detection and matching map of the algorithm in the embodiments of this application in a scene with repeated regions and large lighting changes; Figure 6 These are experimental images of feature point detection and matching using the algorithm in this application embodiment under scenarios with changing viewpoints and scales; Figure 7 This is a statistical chart of data from the algorithm in this application embodiment under scenarios with repetitive regions and significant lighting variations; Figure 8 This is a statistical chart of data from the algorithm in the embodiments of this application under scenarios with changes in viewpoint and scale. Detailed Implementation

[0020] To provide a clearer understanding of the technical features, objectives, and effects of this application, the specific embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0021] The embodiments of this application provide a method for 3D reconstruction of UAV images based on feature point extraction and enhancement.

[0022] Please refer to Figure 1 , Figure 1 This is a flowchart of a UAV image 3D reconstruction method based on feature point extraction and enhancement, as described in an embodiment of this application, including: As one embodiment, the feature point extraction and enhancement method mainly consists of two parts. A multi-scale joint feature point detection module decodes feature maps at different scales to obtain feature point detection heatmaps at the corresponding scales. Each heatmap is treated as a different channel, and learnable weights are used to weight and fuse the detection results. This method enables the network to adaptively allocate the contribution of different scales in the fusion, avoiding the problem of indistinct features caused by simple addition or taking the maximum value. The location of feature points can be obtained through the fused feature point heatmap, and then the feature descriptors are passed to the description enhancement module.

[0023] As one embodiment, in the feature descriptor enhancement module, relative position encoding is performed based on the position of feature points in the image and embedded into the original descriptor to achieve self-enhancement of the second number. Subsequently, the rotation position encoding is integrated into the attention score calculation of the self-attention layer. The purpose is to mine the potential positional relationship of feature points in the same image and extract spatial context information to improve the distinguishability of the feature descriptor.

[0024] The multi-scale joint feature point detection module includes: a multi-scale feature extraction unit, a Softmax-Reshape module, and a fusion unit; the multi-scale feature extraction unit, the Softmax-Reshape module, and the fusion unit are connected sequentially. The multi-scale feature extraction unit is used to extract the original feature map of drone images. Feature map with dimensions W / 4 × H / 4 and a feature map with dimensions W / 2×H / 2 W and H represent the width and height of the feature map; As one embodiment, the original SuperPoint network consists of a shared encoder, a feature point detection decoder, and a feature vector generation decoder. The shared decoder uses a VGG-like network structure. To improve feature detection speed, three convolutional-max pooling layers are used to obtain a feature map I1 of size W / 8×H / 8. However, scaling down the image results in the loss of some shallow features, causing feature point detection to perform poorly in scenes with scale variations. Therefore, we consider retaining a feature map of size W / 4×H / 4 after two and one max pooling layers. Feature map with dimensions W / 2×H / 2 Original feature map of drone images By combining these features, a multi-scale joint feature detection network with different resolutions and channel depths can be formed.

[0025] The Softmax-Reshape module is used to reshape the original feature maps of drone images. Feature map and feature map Reconstructed into a thermal detection map , and ; The fusion unit is used for thermal detection maps , and By combining the results, a fusion heatmap is obtained.

[0026] As an example, the original SuperPoint feature point detection algorithm relies on a neural network, possessing strong feature representation capabilities, and its reliability is improved through self-supervised training. However, when faced with large-scale changes in viewpoint and scale in images captured from a drone's perspective, the SuperPoint algorithm still suffers from unstable feature point detection and poor matching results, which limits its widespread application in the SfM algorithm. When performing SfM calculations on large areas, the viewpoint and scale changes of the images are more drastic, placing greater emphasis on the repeatability and distinguishability of feature points in terms of viewpoint and scale. Poor detection performance can easily lead to poor matching results, resulting in incomplete reconstruction results and even point cloud misalignment. To address these issues, a multi-scale joint detection module is added to improve the multi-scale detection capability of SuperPoint. Based on the SuperPoint algorithm, the original network is improved, and the improved network structure is as follows: Figure 2 As shown.

[0027] The Softmax-Reshape module is used to reshape the original feature map of the drone image. Feature map and feature map Reconstructed into a thermal detection map , and Specifically, it includes: The Softmax-Reshape module includes: Softmax units, Reshape units, and the ReLU activation function; the Softmax units are connected to the Reshape units; The Softmax unit is used to process the original feature map. Feature map and feature map Perform channel-level softmax operation; The Reshape unit is used to restore the feature map after the softmax operation to the original image size, resulting in three thermal detection maps; In one embodiment, in the feature point decoder section, feature maps of three different scales are fed into the decoder for decoding, and the feature point location probability tensor is obtained through sub-pixel convolution. , and Then, channel-level softmax operations are performed on the three tensor maps; the "no interest point" dimension is removed, and the matrix is ​​restored to the original image size to obtain three heat maps.

[0028] In one embodiment, the heatmap represents the confidence level that a pixel is a special point. Previous methods assumed all input features contributed equally to the output, using either the maximum value or summation. However, experiments showed that features at different resolutions often require different levels of importance. Therefore, learnable weights are introduced for each input branch, allowing the network to learn the importance of each input feature.

[0029] As one example, a fusion method based on Softmax is proposed, which normalizes the weights to a probability distribution in the [0,1] interval through softmax to reflect the importance of each input feature. However, it was observed that the additional softmax calculation significantly increases the GPU inference latency, so the fusion method is further improved.

[0030]

[0031] The improved method ensures this by adding a ReLU activation function after the weights. and with a minimal constant To avoid numerical instability, the normalized weights are also set to fall within the [0,1] range. This method abandons softmax, is more efficient, and can improve inference speed while maintaining accuracy similar to softmax.

[0032]

[0033] in O The final output feature point heatmap is the result of weighted fusion of all input feature maps according to the weights in the above formula. Represents the learnable scalar weights corresponding to the i-th input feature map; This represents the summation of all input weights, used for each... Perform normalization; This represents the i-th input feature map, which comes from feature layers of different scales; It is a very small constant used to avoid the denominator being zero or the value being unstable.

[0034] Through the above improvements, the network has the ability to detect feature points at multiple scales, thus improving the scale repeatability of feature detection.

[0035] The feature descriptor enhancement module includes: a descriptor self-enhancement unit and a descriptor mutual enhancement unit; the descriptor self-enhancement unit is connected to the descriptor mutual enhancement unit.

[0036] As one example, descriptors are used to describe the similarity between feature points. Correspondence between feature points can be established by comparing the magnitude of similarity, i.e., feature point matching. Hand-designed descriptors are widely used in computer vision tasks, but as seen in the well-known SIFT and ORB algorithms, they are limited to local regions around feature points and lack the ability to represent higher-level features. After obtaining feature descriptors, nearest neighbor search is usually used to determine the correspondence between images; however, this ignores the spatial and visual relationships between features. The locations of features and the distribution of descriptors in the entire image constitute a global context, which plays an important role in feature matching. The method proposed in this paper integrates global context information into the original descriptors to improve their discriminative ability. The feature descriptor enhancement module design mainly consists of two stages: descriptor self-enhancement and mutual enhancement. The enhancement process is as follows: Figure 3 As shown.

[0037] The descriptor self-enhancing unit specifically includes: As one example, in the SuperPoint network, a D-dimensional descriptor vector is extracted from each pixel of the fused heatmap, and Euclidean distance is typically used to measure the similarity between descriptors. Previous studies have shown that using Hellinger distance to measure similarity yields better matching performance than using Euclidean distance. A significant difference compared to Euclidean distance is that each element in the descriptor is taken as the square root; that is, different similarity measures can be achieved by transforming the original descriptor. This transformation can be viewed as mapping the descriptor to a new feature space, where using Euclidean distance in the new feature space still yields better results.

[0038] Extract a D-dimensional descriptor for each pixel in the fused heatmap. ; Using the self-enhancing unit of the MLP network descriptor as the mapping function, the descriptor... Perform the transformation to obtain the descriptor. ; The MLP network is trained using a loss function based on Euclidean or Hamming distance constraints.

[0039] As one embodiment, an MLP is used as the mapping function. The MLP learns complex patterns in the input descriptors through nonlinear transformations, learning how to map the original descriptors to a more suitable matching space. The training process uses a loss function based on Euclidean or Hamming distance constraints, and the mapped descriptors are more suitable for measuring similarity in the corresponding distance space. The transformed descriptors are recorded as follows: It is the result of the original descriptor passing through an MLP network.

[0040] The geometric information is fed into another MLP network, and the high-dimensional embedded geometric information is added to the descriptor. The self-enhanced descriptor is obtained from ,as follows:

[0041] in Indicates the first The geometric information of each pixel includes its 2D position, scale, orientation, and detection confidence level on the image.

[0042] The descriptor mutual enhancement unit specifically includes: As one example, in drone imagery with large, repetitive areas or weak textures (such as extensive forests), the low discriminative power of objects means that feature points at different locations may have similar feature vectors, making accurate matching difficult using only the set of descriptors. To address this issue, these descriptors are further processed in a subsequent stage, employing cross-enhancement. A Transformer network is used to understand the potential positional relationships of feature points within the same image, extracting spatial context information to improve the discriminative power of the feature descriptors. Through this processing, the receptive field of local feature descriptors is expanded, and they can adaptively adjust based on neighboring features, significantly enhancing their discriminative ability.

[0043] Descriptor of N feature points Input a Transformer network and obtain the descriptor after cross-amplification, as shown in the following formula:

[0044] As one example, the computational complexity of traditional scaled dot product attention is proportional to the square of the number of feature points N2. However, in large-scale images captured by drones, an image may have hundreds or thousands of feature points (pixels), which greatly increases the computational complexity and requires high memory and computational overhead. Therefore, the AFT-Simple attention calculation formula is used.

[0045] The Transformer network uses the AFT-Simple attention mechanism for computation; AFT-Simple does not use matrix multiplication to approximate dot product attention. Instead, it directly performs element-wise multiplication on K and V, avoiding the expensive matrix multiplication operations in traditional attention computation and thus greatly reducing computational complexity. Since matrix multiplication is no longer performed, the computational complexity of attention is linearly related to the number and dimensionality of features, making it very suitable for processing images containing a large number of local features.

[0046] In drone aerial images, feature points exhibit regional distribution. When calculating the correlation between feature point contexts, relative position is more valuable than absolute position; therefore, RoPE relative position encoding is introduced in the AFT-Simple mechanism.

[0047] RoPE relative position encoding is introduced into the AFT-Simple mechanism.

[0048] As one example, RoPE relative position encoding is introduced into the AFT-Simple mechanism, as follows: Let the feature dimension of the descriptor be D. Divide the descriptor into two equal parts, each with a dimension of D / 2. Introduce frequencies to the p-th pair of descriptor subvectors, that is, introduce frequencies to every two dimensions:

[0049] For a feature point located at coordinates (x, y) on the image, the rotation angle on the p-th pair of sub-vectors is defined as:

[0050] That is, first add the rotation amount corresponding to x to the rotation amount corresponding to y to obtain a composite rotation angle. .

[0051] This application, through the above technical solution, first adds the rotation amount corresponding to x to the rotation amount corresponding to y to obtain a composite rotation angle. The angle calculation formula of traditional rotation encoding is improved by taking into account both x and y position variables, which can better reflect the positional relationship between different feature points.

[0052] The calculated introductory frequencies are applied to the query Q and key value K, and the descriptor vector is cut into D / 2 pairs of two-dimensional subvectors.

[0053]

[0054] in The i-th row of the query matrix Q without rotation encoding can be understood as the specific representation of the eigenvector of the i-th feature point in the query matrix Q. Using the rotation angle, for the p-th pair Rotate the vector:

[0055] in and Indicates the rotated vector; It is the specific representation of the eigenvector of the j-th feature point in the key-value matrix K. Similarly, for... Use it Perform rotation; By concatenating the rotated vectors in pairs, we get:

[0056] Will and Applying this to the original AFT-Simple attention mechanism, we get:

[0057] in, This represents the feature descriptor enhanced by the attention mechanism; After introducing RoPE, the gate vector By directly relying on the coordinates of feature points, feature points at different locations will receive different gating coefficients. Location-based cosine and sinine values ​​are compressed. Meanwhile, attention weights... Depend on The weight distribution for each dimension d is generated as follows:

[0058] in It is already related to the position of each feature point. That is to say, for a certain dimension d, the coordinates of the feature points determine the magnitude of the weight, which in turn affects which type of feature points in that dimension are given higher weights. This makes the aggregation context bias towards points in different positions.

[0059] For different feature points and feature points The coordinates, and the output difference of the attention mechanism are:

[0060] The vectors are caused by gating differences, which in turn originate from coordinate differences. Different feature points produce different outputs due to their different relative positions, thus implicitly preserving two-dimensional relative position information, which is more helpful for matching.

[0061] As one example, in image matching tasks, the relative positions encoded by RoPE help the model distinguish true matching pairs from false matching points by spatial layout information, even under conditions of repetitive textures or changes in lighting. The gating mechanism and weighted aggregation in AFT-Simple inject this positional information into the final descriptor through element-wise multiplication, thereby improving the accuracy and robustness of matching.

[0062] This application also discloses feature point extraction and matching experiments on images with repetitive regions and significant lighting variations. The RANSAC algorithm was used for testing, with a preliminary reprojection error threshold of 8 pixels. For comparison, experiments were conducted using the SIFT algorithm, the original SuperPoint algorithm, and the improved SuperPoint algorithm described in this paper. As shown in Figure 1, a pair of building images under strong and weak light conditions were selected, with the buildings containing many repetitive window patterns. Feature point detection and matching were performed using SIFT, SuperPoint, and the improved SuperPoint algorithm, respectively. Correct and incorrect matching results were marked with red and green lines, respectively. The experimental results are shown in Figure 2. Figure 5 As shown, the data statistics are as follows: Figure 7 As shown.

[0063] Compared to the SIFT and SuperPoint algorithms, the improved SuperPoint algorithm can detect more feature points and achieve a higher number of matches. Traditional SIFT is prone to mismatches in regions with repetitive textures. While the SuperPoint algorithm has improved upon this, mismatches are still relatively common. The improved SuperPoint algorithm proposed in this paper, after descriptor self-enhancement, provides better representation of lighting conditions. Furthermore, after descriptor cross-enhancement, it captures spatial context information, enhancing its performance in regions with many repetitive textures, effectively improving matching accuracy and ensuring the precision of 3D reconstruction. Overall, the improved SuperPoint algorithm outperforms the original algorithm in both feature point extraction and matching performance.

[0064] To verify the feature point detection and matching performance of the improved SuperPoint algorithm under changes in viewpoint and scale, another set of images was selected for experimentation. The SIFT algorithm, the original SuperPoint algorithm, and the improved SuperPoint algorithm were used in the experiments, with a reprojection error threshold of 8 pixels. The experimental results are shown below. Figure 6 As shown, the data statistics are as follows: Figure 8 As shown, the improved SuperPoint can detect more feature points with the help of a multi-scale detection module, and the matching success rate is also improved with the help of descriptor enhancement.

[0065] An electronic device. (Refer to...) Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.

[0066] The communication bus 502 is used to enable communication between these components.

[0067] The user interface 503 may include a display screen, and optionally, the user interface 503 may also include a standard wired interface or a wireless interface.

[0068] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0069] This application also discloses a computer-readable storage medium storing multiple instructions adapted for loading by a processor to execute the above-described method for 3D reconstruction of UAV images based on feature point extraction and enhancement.

[0070] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure.

[0071] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for 3D reconstruction of UAV images based on feature point extraction and enhancement, characterized in that, The method includes the following steps: S1: Acquire drone images; S2: Based on the SuperPoint network and multi-scale feature fusion, a feature point extraction and enhancement network is constructed; the feature point extraction and enhancement network includes: a multi-scale joint feature point detection module and a feature descriptor enhancement module; S3: The feature point extraction and enhancement network is trained using drone images; S4: Obtain the drone image to be extracted; extract feature points from the drone image using the trained feature point extraction and enhancement network; complete the 3D reconstruction of the drone image using the extracted feature points.

2. The method for 3D reconstruction of UAV images based on feature point extraction and enhancement as described in claim 1, characterized in that, The multi-scale joint feature point detection module includes: a multi-scale feature extraction unit, a Softmax-Reshape module, and a fusion unit; the multi-scale feature extraction unit, the Softmax-Reshape module, and the fusion unit are connected sequentially. The multi-scale feature extraction unit is used to extract the original feature map of drone images. Feature map with dimensions W / 4 × H / 4 and a feature map with dimensions W / 2×H / 2 W and H represent the width and height of the feature map; The Softmax-Reshape module is used to reshape the original feature maps of drone images. Feature map and feature map Reconstructed into a thermal detection map , and ; The fusion unit is used for thermal detection maps , and By combining the results, a fusion heatmap is obtained.

3. The method for 3D reconstruction of UAV images based on feature point extraction and enhancement as described in claim 2, characterized in that, The Softmax-Reshape module is used to reshape the original feature map of the drone image. Feature map and feature map Reconstructed into a thermal detection map , and Specifically, it includes: The Softmax-Reshape module includes: Softmax units, Reshape units, and the ReLU activation function; the Softmax units are connected to the Reshape units; The Softmax unit is used to process the original feature map. Feature map and feature map Perform channel-level softmax operation; The Reshape unit is used to restore the feature map after the softmax operation to the original image size, resulting in three thermal detection maps.

4. The method for 3D reconstruction of UAV images based on feature point extraction and enhancement as described in claim 2, characterized in that, The feature descriptor enhancement module includes: a descriptor self-enhancement unit and a descriptor mutual enhancement unit; the descriptor self-enhancement unit is connected to the descriptor mutual enhancement unit.

5. The method for 3D reconstruction of UAV images based on feature point extraction and enhancement as described in claim 4, characterized in that, The descriptor self-enhancing unit specifically includes: Extract a D-dimensional descriptor for each pixel in the fused heatmap. ; Using the self-enhancing unit of the MLP network descriptor as the mapping function, the descriptor... Perform the transformation to obtain the descriptor. ; The MLP network is trained using a loss function based on Euclidean or Hamming distance constraints.

6. The method for 3D reconstruction of UAV images based on feature point extraction and enhancement as described in claim 5, characterized in that, The descriptor mutual enhancement unit specifically includes: Descriptor of N feature points Input a Transformer network and obtain the descriptor after cross-amplification, as shown in the following formula: The Transformer network uses the AFT-Simple attention mechanism for computation; RoPE relative position encoding is introduced into the AFT-Simple mechanism.

7. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a computer, perform the method as described in any one of claims 1-6.