Unmanned aerial vehicle remote sensing image splicing method fusing depth features of spatial domain and frequency domain

The integration of spatial and frequency domain features in a deep learning-based image stitching method for UAV remote sensing images addresses alignment challenges, enhancing precision and quality in complex environments.

CN120318063APending Publication Date: 2025-07-15SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317463.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing drone remote sensing image stitching algorithm is difficult to achieve high-quality image registration when processing complex backgrounds and low feature areas. The traditional spatial domain feature method is not effective when there is insufficient parallax and key points, while the frequency domain feature method leads to unclear information and loss of spatial information.

Method used

The drone remote sensing image stitching method that integrates the spatial domain frequency domain depth features is adopted. By constructing an image stitching network model, combining the spatial feature extraction module, the spatial domain frequency domain depth feature fusion module and the image spatial transformation parameter prediction module, the Haar wavelet transform and the Manhattan distance loss function are used to optimize feature matching to achieve more accurate image stitching.

Benefits of technology

It significantly improves the accuracy of feature point matching and sensitivity of image grayscale changes, optimizes the stitching effect, especially in complex backgrounds and low feature areas, and improves the stitching accuracy and panoramic image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318063A_ABST
    Figure CN120318063A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle remote sensing image splicing method fusing depth features of a space domain and a frequency domain, and the method comprises the steps: dividing an obtained unmanned aerial vehicle remote sensing image data set into a training data set and a test data set according to a proportion; an image splicing network model fusing the spatial domain and frequency domain depth features is constructed, the image splicing network model comprises a spatial feature extraction module, a spatial domain and frequency domain depth feature fusion module and an image spatial transformation parameter prediction module, and an image transformation matrix is obtained; gaussian distribution with the mean value being zero is utilized to randomly initialize the weight and bias of the convolution kernel, and the training data set is adopted to carry out learning training on the image splicing network model to obtain an optimal image splicing network model; and inputting the test data set into the optimal image splicing network model to complete the splicing task of the remote sensing image of the test unmanned aerial vehicle. Compared with an existing method, the method has the advantages that the splicing precision, the feature recognition capability and the panoramic image quality are remarkably improved, and meanwhile, the automation degree and the efficiency are higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) remote sensing image processing, and specifically relates to a method for stitching UAV remote sensing images by fusing spatial domain, frequency domain and depth features. Background Art

[0002] In recent years, with the continuous improvement of the automation level of UAV platforms, their advantages of convenient and fast deployment have become increasingly prominent. The remote sensing images obtained by UAVs have many significant advantages. The coverage range is extremely wide, and it can easily cover large geographical areas; the observation perspectives are rich and diverse, and whether it is a vertical overlook or an oblique shot, it can accurately capture surface features; the data update cycle is short, and it can dynamically monitor surface changes almost in real time. However, limited by factors such as aperture size, sensor sensitivity and flight altitude, it is difficult for UAVs to cover the entire target area in one shot. In order to obtain images of large areas, it is necessary to stitch multiple partially overlapping UAV images.

[0003] When image stitching technology is applied to UAV remote sensing images, due to problems such as complex object features in UAV remote sensing images, such as small size, high similarity and mutual occlusion, traditional image stitching algorithms often fail to achieve ideal accuracy; however, the current requirements for the quality of the stitched panoramic images are extremely high. Therefore, it is particularly important to improve the stitching algorithm of UAV remote sensing images and optimize the stitching effect. The UAV remote sensing image stitching technology mainly includes four steps: image acquisition, image preprocessing, image registration and image fusion. Image registration is the most critical step among them. Image registration aligns multiple images in geometric positions according to the feature information between the images to be stitched, and realizes the image position registration in space.

[0004] The existing deep image stitching methods mainly rely on spatial domain features. Although they can capture the detailed information of images, they often appear powerless when facing problems such as parallax and insufficient key points in UAV remote sensing images. These methods have poor stitching effects when dealing with complex backgrounds and low-feature areas, resulting in a decline in the registration quality of the overall image. A few frequency domain feature-based methods directly replace the spatial domain features with the frequency domain features of wavelet transform, which improves the model performance and generalization ability to a certain extent. However, due to the particularity of remote sensing images, using only frequency domain features will lead to unclear information and loss of spatial information. In contrast, spatial domain features can better capture the semantic information of different categories in images, and the spatial information is more accurate. Summary of the Invention

[0005] In order to overcome the disadvantages and deficiencies of the existing technology, the present invention provides a method for stitching UAV remote sensing images by fusing spatial domain, frequency domain and depth features.

[0006] The present invention proposes a method for stitching UAV remote sensing images that fuses spatial domain, frequency domain, and depth features. This method not only retains the fine details of spatial domain features but also ingeniously incorporates frequency domain information to enhance the sensitivity to image gray-scale changes and the recognition ability of feature points. Through this fusion, feature point matching can be performed more precisely, thereby calculating a more accurate transformation matrix to achieve higher-quality image stitching. The present invention can obtain a better stitching effect while ensuring image accuracy, meeting the requirements of actual industrial applications.

[0007] To achieve the above objectives, the present invention adopts the following technical solutions:

[0008] A method for stitching UAV remote sensing images that fuses spatial domain, frequency domain, and depth features, comprising:

[0009] Step 1: Divide the obtained UAV remote sensing image dataset into a training dataset and a test dataset according to a ratio;

[0010] Step 2: Construct an image stitching network model that fuses spatial domain, frequency domain, and depth features. The image stitching network model includes a spatial feature extraction module, a spatial domain-frequency domain-depth feature fusion module, and an image spatial transformation parameter prediction module to obtain an image transformation matrix;

[0011] Step 3: Randomly initialize the weights and biases of the convolution kernels using a Gaussian distribution with a mean of zero, and use the training dataset to learn and train the image stitching network model to obtain an optimal image stitching network model;

[0012] Step 4: Input the test dataset into the optimal image stitching network model to complete the stitching task of the test UAV remote sensing images.

[0013] Further, the learning and training is specifically to optimize the image stitching network model with the goal of improving the similarity of the overlapping regions of the image after homography transformation. The Manhattan distance loss function is selected to calculate the error of the predicted homography transformation matrix for image registration, and this error is used as the prediction loss of the network model to train the network model by the method of backpropagation.

[0014] Further, the step 1 further includes a preprocessing step, and the preprocessing step is to adjust the size of the input image to a fixed value of 512*512, and every two UAV remote sensing images with an overlapping region are input as a pair.

[0015] Further, the spatial feature extraction module includes a VGG16 network and an L2-Normalization module connected in sequence. The input picture I∈R 512×512×3 , and the features obtained through the VGG16 network are transformed to the range of 0 to 1 to obtain spatial features I1∈R 128×128×256 .

[0016] Furthermore, the VGG16 network includes three convolutional blocks connected in sequence. The output of each convolutional block is non-linearly processed through the ReLU activation function, and max pooling layers are used for downsampling.

[0017] Furthermore, the spatial-domain frequency-domain depth feature fusion module is specifically as follows:

[0018] First, the spatial feature I1 is decomposed by Haar wavelet transform into four frequency-domain components: a low-frequency component A, a horizontal high-frequency component H, a vertical high-frequency component V, and a diagonal high-frequency component D. The sizes of the above frequency-domain features are all I f ∈R 64×64×256 ;

[0019] Second, the low-frequency component A is subjected to a convolution operation to obtain a low-frequency feature I lf ∈R 64×64×64 , and the three high-frequency components H, V, and D are concatenated into a feature map and then convolved to obtain a high-frequency feature I hf ∈R 64×64×64 ;

[0020] Finally, the spatial feature I1 and the frequency-domain feature are fused to complete the operation of extracting the features of the UAV remote sensing image by fusing the spatial-domain and frequency-domain features, and the remote sensing image features are obtained.

[0021] Furthermore, after downsampling I1 using a 2*2 max pooling layer with a stride of 2, it is convolved with 64 1*1 convolutional kernels with a stride of 1 to obtain a spatial feature I2 ∈ R 64×64×64 that is the same size as the frequency-domain feature. Then, I2, I lf , and I hf are concatenated to complete the operation of extracting the features of the UAV remote sensing image by fusing the spatial-domain and frequency-domain features.

[0022] Furthermore, the specific process of the image spatial transformation parameter prediction module is as follows:

[0023] First, calculate the correlation between the features of image A and the features of image B:

[0024]

[0025] where is the distance of the feature in the two-dimensional space, |x| is the modulus of the calculated vector, represents the inner product between vectors, and image A and image B are two images to be stitched;

[0026] Second, perform convolution through three convolutional layers, and perform rectified linear unit activation processing after each convolution. Then, flatten the activated feature vectors;

[0027] Finally, a two-dimensional image homography matrix is output through two fully-connected layers.

[0028] Furthermore, the three convolutional blocks are respectively a first convolutional block, a second convolutional block, and a third convolutional block;

[0029] The first convolutional block includes two convolutional layers of equal size and one pooling layer;

[0030] The second convolutional block includes two convolutional layers of equal size and one pooling layer;

[0031] The third convolutional block includes three convolutional layers of equal size.

[0032] Furthermore, after the image transformation matrix, image registration is performed, and finally, through the image fusion step, the final stitched image is obtained using the average value fusion method.

[0033] Advantages of the present invention:

[0034] The method for stitching UAV remote sensing images that fuses spatial domain, frequency domain, and depth features proposed by the present invention significantly improves the accuracy of feature point matching and the sensitivity of the model to image gray-scale changes by combining spatial domain and frequency domain features, and optimizes the stitching effect;

[0035] The present invention is particularly suitable for UAV remote sensing images with complex backgrounds and low feature areas, and effectively overcomes the deficiencies of traditional methods in dealing with small-sized, highly similar, and mutually occluding objects;

[0036] Compared with existing methods, the present invention has significantly improved stitching accuracy, feature recognition ability, and panoramic image quality, and at the same time has higher automation and efficiency. Description of the Drawings

[0037] Figure 1 is a schematic diagram of the method for stitching UAV remote sensing images that fuses spatial domain, frequency domain, and depth features of the present invention;

[0038] Figure 2 is a hierarchical structure diagram of the spatial domain, frequency domain, and depth feature fusion module of the present invention;

[0039] Figure 3(a) is the original UAV remote sensing image to be stitched, Figure 3(b) is the stitching effect diagram of the method based on spatial features, and Figure 3(c) is the stitching effect diagram of the method described in this embodiment;

[0040] Figure 4 is the workflow diagram of the present invention. Specific Embodiments

[0041] The present invention will be further described in detail below in conjunction with embodiments and the drawings, but the embodiments of the present invention are not limited thereto.

[0042] This embodiment provides a method for stitching UAV remote sensing images that fuses spatial domain, frequency domain, and depth features. The flow chart is as shown in Figure 1 and Figure 4 shown. Combining specific numerical examples, it includes the following steps:

[0043] Step 1: Preprocess the acquired UAV remote sensing image dataset. Divide the dataset into a training dataset and a test dataset in a ratio of 8:2. The preprocessing includes adjusting the size of the input image to a fixed value of 512*512, and taking every two UAV remote sensing images with an overlapping area as a pair of inputs;

[0044] Step 2: Construct an image stitching network model that fuses spatial domain, frequency domain, and depth features. Randomly initialize the weights and biases of the convolutional kernels using a Gaussian distribution with a mean of zero, and input the divided training dataset into the image stitching network that fuses spatial domain, frequency domain, and depth features for training and learning;

[0045] Step 3: Adopt the idea of unsupervised learning. Take improving the similarity of the overlapping area after the homography transformation of the image pair as the optimization goal of the entire network model. Select the Manhattan distance loss function to calculate the error size after image registration using the predicted homography transformation matrix, and use it as the prediction loss of the entire network model. The specific calculation method of this loss function is:

[0046]

[0047] where I A , I b represent the pixel matrices of the images A and B to be stitched, ⊙ represents the multiplication of pixel matrices, and the size of E is the same as that of I a same. represents the homography transformation matrix. This formula measures the effect of image registration by calculating the difference degree of pixel points between the registration of images A and B. The smaller the Loss, the better the effect of image registration;

[0048] Step 4: Use the Adam optimizer to perform backpropagation on the error, and iteratively update the weights and biases of the convolutional kernels. When the loss function reaches the minimum value or reaches the set maximum number of iteration steps, it is regarded as obtaining the current optimal image stitching network model;

[0049] Step 5: Input the test dataset into the optimal image stitching network model obtained in Step 4, and the stitching task of the test UAV remote sensing images can be completed.

[0050] Furthermore, the image stitching network model in Step 2 includes a spatial feature extraction module, a spatial domain, frequency domain, and depth feature fusion module, and an image spatial transformation parameter prediction module.

[0051] Step 2.1. Specific process of the spatial feature extraction module:

[0052] For the input image I ∈ R 512×512×3 , spatial features are extracted. The VGG16 network model is adopted. The feature extraction structure of the VGG16 network before its third pooling layer is intercepted for this module, and then an L2-Normalization module is connected to transform the extracted features into the range of 0 to 1, corresponding to the operation of normalizing the feature descriptor in the traditional feature extraction algorithm, weakening the influence brought by the illumination difference. The two feature extraction modules in the overall network structure share parameters to ensure that the same features are extracted from the two input images. Finally, the output feature I1 ∈ R 128 ×128×256 .

[0053] The specific part of the VGG16 network used therein includes:

[0054] The first convolutional block:

[0055] Convolutional layer 1: 64 convolutional kernels, with a size of 3*3, a stride of 1. The input image is convolved through 64 3*3 convolutional kernels, and the ReLU activation function is used.

[0056] Convolutional layer 2: 64 convolutional kernels, with a size of 3*3, a stride of 1. The output of the previous layer is convolved through 64 3*3 convolutional kernels, and the ReLU activation function is used.

[0057] Pooling layer 1: Use a 2*2 max-pooling layer with a stride of 2 to downsample the output of convolutional layer 2.

[0058] The second convolutional block:

[0059] Convolutional layer 3: 128 convolutional kernels, with a size of 3*3, a stride of 1. The output of the previous layer is convolved through 128 3*3 convolutional kernels, and the ReLU activation function is used.

[0060] Convolutional layer 4: 128 convolutional kernels, with a size of 3*3, a stride of 1. The output of the previous layer is convolved through 128 3*3 convolutional kernels, and the ReLU activation function is used.

[0061] Pooling layer 2: Use a 2*2 max-pooling layer with a stride of 2 to downsample the output of convolutional layer 4.

[0062] The third convolutional block:

[0063] Convolutional layer 5: 256 convolutional kernels, with a size of 3*3, a stride of 1. The output of the previous layer is convolved through 256 3*3 convolutional kernels, and the ReLU activation function is used.

[0064] Convolutional layer 6: 256 convolution kernels of size 3*3, stride 1. The output of the previous layer is convolved with 256 3*3 convolution kernels, using the ReLU activation function.

[0065] Convolutional layer 7: 256 convolution kernels of size 3*3, stride 1. The output of the previous layer is convolved with 256 3*3 convolution kernels, using the ReLU activation function.

[0066] The output of each convolutional layer is processed nonlinearly through the ReLU activation function. Within each convolutional block, the output of the convolutional layer is continuous, that is, the output of one convolutional layer is directly used as the input of the next convolutional layer. After each convolutional block, a maximum pooling layer is used for downsampling to reduce the size of the feature map and extract the most important features to obtain the spatial features I1∈R 128×128×256 .

[0067] Step 2.2: spatial domain and frequency domain deep feature fusion module, specifically:

[0068] For the spatial feature I1∈R obtained in step 2.1 128×128×256 Perform deep feature fusion in spatial and frequency domains. The module structure is shown in the following figure. Figure 2 As shown, the specific process is:

[0069] First, Haar wavelet transform is used to decompose feature I1, and the spatial feature is decomposed into four frequency domain components: low frequency component A, horizontal high frequency component H, vertical high frequency component V, and diagonal high frequency component D. The size of these frequency domain features is I f ∈R 64 ×64×256 These frequency domain components contain the information of the image in different frequency ranges, thus enriching the feature representation ability of the model. Then the low-frequency component A is convolved with 64 1*1 convolution kernels with a step size of 1 to obtain the low-frequency feature I lf ∈R 64×64×64 ; Connect the three high-frequency components H, V, and D into a feature map, and then use 64 1*1 convolution kernels with a step size of 1 to perform convolution to obtain the high-frequency feature I hf ∈R 64×64×64 .

[0070] Finally, the spatial features and frequency domain features are fused: I1 is downsampled using a 2*2, step-size 2 maximum pooling layer, and then convolved with 64 1*1, step-size 1 convolution kernels to obtain the spatial feature I2∈R that is consistent with the frequency domain feature size. 64×64×64 , I2, I lf ,I hfConnect them to complete the feature extraction operation of UAV remote sensing images that fuse spatial and frequency domain features. And this feature extraction module shares weights between two input images;

[0071] Step 2.3, Image Spatial Transformation Parameter Prediction Module, specifically:

[0072] Predict the spatial transformation matrix of the image using the feature map obtained by fusing spatial and frequency domain features in Step 2.2, which is completed through a spatial transformation parameter prediction module. The specific structure of the module is as follows:

[0073] First, calculate the correlation between the features of Image A and the features of Image B:

[0074]

[0075] In the formula is the distance of the feature in the two-dimensional space. |x| is the modulus length of the calculated vector, represents the inner product between vectors.

[0076] Then, it is processed through three convolutional layers. Each layer is convolved with 512 3*3 convolutional kernels with a stride of 1, and ReLU activation processing is performed after each convolution.

[0077] Furthermore, the feature vector is flattened, and then two fully connected layers are connected. The output lengths are set to 1024 and 8 respectively, and ReLU activation processing is performed between the two fully connected layers.

[0078] Furthermore, the final output of the 8 parameters corresponds to the two-dimensional image homography matrix

[0079]

[0080] where, (x, y) and (x ′ , y ′ ) represent the coordinates of a certain point in the image before and after transformation respectively, that is, the homography matrix. In the two-dimensional space, four points determine a plane, with a total of 8 parameters;

[0081] Step 2.4, After obtaining the image transformation matrix, perform image registration, and finally through the image fusion step, use the average value fusion method to obtain the final stitched image.

[0082] Figure 3(a) is the original remotely sensed UAV image to be stitched, Figure 3(b) is the stitching effect diagram of the method based on spatial features, the right figure is the enlarged view of the left square, Figure 3(c) is the stitching effect diagram of the method described in this embodiment, and the right figure is the enlarged view of the left square; in the stitching area of the method based on spatial features, obvious registration errors occurred, and the images in some areas were misaligned, resulting in poor coherence of the entire picture. Moreover, there are also ghosting and blurring situations in the picture, causing some detailed information of the image to be lost, and the overall visual effect is greatly reduced, obviously it is difficult to meet the actual industrial application scenarios with high requirements for stitching quality; while the method of fusing spatial domain, frequency domain and depth features proposed in this embodiment can obtain a higher quality stitching effect, which better meets the strict requirements for the accuracy and integrity of image stitching in practical applications, and has good practical value and application prospects.

[0083] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the described embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for stitching UAV remote sensing images that integrates spatial domain, frequency domain, and depth features, characterized in that Including: Step 1: Divide the acquired UAV remote sensing image dataset into a training dataset and a test dataset according to a ratio. Step 2: Construct an image stitching network model that fuses spatial domain and frequency domain depth features. The image stitching network model includes a spatial feature extraction module, a spatial domain and frequency domain depth feature fusion module, and an image spatial transformation parameter prediction module to obtain an image transformation matrix. Step 3: Randomly initialize the weights and biases of the convolution kernels with a Gaussian distribution having a mean of zero, and use the training dataset to learn and train the image stitching network model to obtain an optimal image stitching network model. Step 4: Input the test dataset into the optimal image stitching network model to complete the stitching task of the test UAV remote sensing images.

2. The method for stitching drone remote sensing images according to claim 1, characterized in that, The learning and training is specifically to optimize the image stitching network model with the goal of improving the similarity of the overlapping regions of the images after homography transformation. The Manhattan distance loss function is selected to calculate the error of the predicted image transformation matrix for image registration, and this error is used as the prediction loss of the image stitching network model, and the image stitching network model is trained by the backpropagation method.

3. The method for stitching UAV remote sensing images according to claim 1, wherein The said Step 1 further includes a preprocessing step, and the preprocessing step is to adjust the size of the input image to a fixed value of 512*512, and every two UAV remote sensing images with overlapping regions are input as a pair.

4. The method for stitching drone remote sensing images according to claim 1, wherein, The said Spatial feature extraction module, including a VGG16 network and an L2-Normalization module connected in sequence, with the input image \(I\in\mathbb{R}\) 512×512×3 , transforms the features obtained through the VGG16 network to the range of 0 to 1, obtaining spatial features \(I_1\in\mathbb{R}\) 128×128×256 .

5. The method for stitching UAV remote sensing images according to claim 4, wherein The VGG16 network includes three convolutional blocks connected in sequence. The output of each convolutional block is non-linearly processed through a ReLU activation function, and max pooling layers are used for downsampling.

6. The method for stitching UAV remote sensing images according to claim 1, wherein The spatial domain and frequency domain depth feature fusion module is specifically: First, perform the spatial feature I1 decomposition using the Haar wavelet transform, and decompose the spatial feature into four frequency-domain components: the low-frequency component A, the horizontal high-frequency component H, the vertical high-frequency component V, and the diagonal high-frequency component D. The sizes of the above frequency-domain features are all I f ∈R 64 ×64×256 ; Secondly, perform a convolution operation on the low-frequency component A to obtain the low-frequency feature I lf ∈R 64×64×64 , connect the three high-frequency components H, V, and D into a feature map, perform convolution, and obtain the high-frequency feature I hf ∈R 64×64×64 ; Finally, fuse the spatial feature I1 and the frequency domain feature to complete the feature extraction operation of the UAV remote sensing image that fuses spatial domain and frequency domain features, and obtain the remote sensing image features.

7. The method for stitching drone remote sensing images according to claim 6, wherein After downsampling I1 using a 2*2 max pooling layer with a stride of 2, it is convolved with 64 1*1 convolutional kernels with a stride of 1 to obtain a spatial feature I2 ∈ R that is the same size as the frequency domain feature. 64×64×64 , combine I2, I lf , and I hf to complete the feature extraction operation of the UAV remote sensing image that fuses spatial and frequency domain features.

8. The method for stitching UAV remote sensing images according to claim 1, wherein The process of the image spatial transformation parameter prediction module is specifically: First, calculate the correlation between the features of image A and the features of image B. where is the distance of the feature in the two-dimensional space, and |x| is the norm of the calculated vector, represents the inner product between vectors, and Image A and Image B are two images to be stitched; Secondly, perform convolution through three convolutional layers, and perform rectified linear activation processing after each convolution, and flatten the activated feature vectors. Finally, output the image transformation matrix through two fully connected layers.

9. The method for stitching UAV remote sensing images according to claim 5, wherein The said three convolutional blocks are respectively the first convolutional block, the second convolutional block and the third convolutional block. The first convolutional block includes two convolutional layers of equal size and one pooling layer. The second convolutional block includes two convolutional layers of equal size and one pooling layer. The third convolutional block includes three convolutional layers of equal size.

10. The method for stitching drone remote sensing images according to claim 1, wherein After the image transformation matrix, image registration is performed, and finally through the image fusion step, the final stitched image is obtained by using the average value fusion method.