Unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention

By using the HTMA-RS model, multi-scale Transformer modules and hybrid attention mechanisms, the problems of simple image degradation methods and high training data costs in UAV image super-resolution are solved. This model achieves detail restoration and sharpness enhancement of high-resolution images, and is suitable for target object recognition in complex scenes.

CN120655511BActive Publication Date: 2025-11-04HUANTIAN SMART TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511149616.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-04
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing UAV image super-resolution methods have limitations such as simple image degradation methods, high training data acquisition costs, and the inability to capture global contextual information.

Method used

The HTMA-RS model based on multi-scale hybrid attention is adopted. Through multi-scale Transformer module, multi-channel-multi-space hybrid attention module and feature fusion reconstruction module, combined with random kernel degradation and JPEG compression, high-resolution images are generated.

Benefits of technology

It restores detailed features in drone imagery, enhances image clarity, reduces blur and jagged edges, and improves image quality and accuracy, making it suitable for target object recognition in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655511B_ABST
    Figure CN120655511B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-scale mixed attention unmanned aerial vehicle image super-resolution method, steps are as follows: first cutting the data set of unmanned aerial vehicle image obtained;Image degradation is then obtained respectively low, high-resolution image data set;Then create and train HTMA-RS model;Finally, the high-resolution image output by model is spliced, and complete high-resolution aerial image is obtained.Super-resolution technology can reconstruct high-resolution image from low-resolution image, restore building texture and other details, increase the number of pixels, make the image delicate, enlarge not easy to blur.Multi-scale mixed attention mechanism extracts features from different scales, and fine texture and overall structure are grasped in different sizes.Scale attention mechanism lets model focus on important areas and suppress interference.The trained HTMA-RS model accurately learns image mapping relationship, ensures reconstruction quality, and provides high-quality image support for unmanned aerial vehicle image in multi-field application.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image processing, and specifically relates to a method for super-resolution of unmanned aerial vehicle (UAV) images based on multi-scale hybrid attention. BACKGROUND

[0002] With the rapid development and popularization of remote sensing technology, especially with the improvement of image spatial resolution, people's demand for the application of remote sensing data is increasing, and the requirement for the resolution of remote sensing images is also increasing. At present, there are few remote sensing satellites with a spatial resolution better than 1 meter worldwide, which is difficult to meet the needs of users. Therefore, unmanned aerial vehicle technology has gradually become an important complementary means. Unmanned aerial vehicles not only provide high-resolution image data, but also have the advantages of high flexibility, low cost, and simple operation, and can obtain large-area high-precision remote sensing data in a short time. Therefore, unmanned aerial vehicles have a very broad application prospect in the fields of agriculture, environmental monitoring, urban planning, disaster emergency, etc.

[0003] However, although the unmanned aerial vehicle image has high resolution, in actual application, due to factors such as flight height and sensor performance, the image data obtained still has the problem of insufficient resolution. In order to further improve the application value of the unmanned aerial vehicle image, it is necessary to apply the super-resolution reconstruction technology to the unmanned aerial vehicle image. The super-resolution reconstruction technology generates a high-resolution image by processing a low-resolution image, thereby making up for the defects of insufficient resolution of the unmanned aerial vehicle image.

[0004] At present, there are some methods and systems for super-resolution reconstruction of unmanned aerial vehicle images.

[0005] In patent publication CN111667409A, high-resolution high-voltage cable insulator images collected by an unmanned aerial vehicle are down-sampled to simulate low-resolution images, and a deep learning model is used to learn the deep correspondence between low-resolution and high-resolution images. Finally, through the super-resolution model obtained by training, the low-resolution insulator image can be enhanced to a high-resolution image.

[0006] In patent publication CN113469881A, the SRDenseNet network is improved to reduce the amount of calculation and improve the real-time performance, so as to realize the super-resolution reconstruction of unmanned aerial vehicle aerial images and improve the image quality.

[0007] In patent publication CN117893410A, a method for super-resolution reconstruction of unmanned aerial vehicle aerial images is introduced. By obtaining low-resolution and high-resolution images at different heights, image registration, color alignment, and deep convolutional neural network training are performed to realize the detail reconstruction of the low-resolution image and improve the image quality and accuracy. SUMMARY

[0008] The unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention aims to solve the problems of the existing solutions in the background art, such as simple image degradation method, high training data acquisition cost, and certain limitations in capturing global context information.

[0009] To solve the above technical problems, the technical solution adopted by the present application is:

[0010] An unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention, comprising the following steps:

[0011] Step S1, preparing the unmanned aerial vehicle image that has been acquired, cutting it to make a data set;

[0012] Step S2, performing degradation processing on the acquired unmanned aerial vehicle image data, the degraded image serving as a low-resolution image data set, and the corresponding high-resolution image serving as a high-resolution image data set;

[0013] Step S3, creating and training an HTMA-RS model;

[0014] Step S4, splicing the high-resolution image output by the model to obtain a complete high-resolution real-time aerial image.

[0015] According to the above technical solution, in step S2, the degradation processing on the unmanned aerial vehicle image data specifically comprises the following steps:

[0016] Step S201, cutting the unmanned aerial vehicle image data to obtain a high-resolution image;

[0017] Step S202, performing first degradation processing on the cut unmanned aerial vehicle image data, including random kernel, JPEG compression, and random noise processing;

[0018] Step S203, performing second degradation processing on the unmanned aerial vehicle image data, including random kernel, sinc kernel, JPEG compression, and random noise processing;

[0019] Step S204, obtaining a low-resolution data set after the processing is completed.

[0020] According to the above technical solution, in the first degradation processing, the random kernel includes iso kernel, aniso kernel, generalized_iso kernel, generalized_aniso kernel, plateau_iso kernel, and plateau_aniso kernel; the size of the kernel is randomly selected between 7 and 21, and image edge padding is performed to adapt to the processing requirements.

[0021] According to the above technical scheme, after the first and second quality degradation processing, a JPEG compression is added to the image to generate a sinc low-pass filter kernel, and the formula is as follows:

[0022]

[0023] In the formula, sinc(x) represents the low-pass filter kernel, and sin(πx) represents the function value of the sine function when the independent variable is πx.

[0024] The low-pass filter kernel is used for filtering after JPEG compression to reduce the artifacts caused by compression.

[0025] According to the above technical scheme, in step S3, the HTMA-RS model is created as follows: the HTMA-RS model includes a multi-scale Transformer module, a multi-channel-multi-space hybrid attention module, and a feature fusion reconstruction module.

[0026] The multi-scale Transformer module is used for extracting multi-scale features of the input image, and specifically, the input image is first divided into fixed-size non-overlapping blocks. Assuming that the size of the input image is HxW, and the size of each block is PxP, the image is divided into blocks; linear embedding is performed on each block to map each block to a high-dimensional feature space; assuming that the channel number of the input block is , and the feature dimension after embedding is D, the linear embedding process can be represented as:

[0027]

[0028] wherein, and represent the weights and biases of linear embedding, respectively; SwinT uses window-based multi-head self-attention (W-MSA) and cross-window multi-head self-attention (SW-MSA) to alternately extract features.

[0029] W-MSA performs self-attention calculation within a fixed-size window. Assuming that the window size is MxM, in each window, multi-head self-attention is calculated:

[0030]

[0031] In the formula, , , represent the query, key, and value matrices, respectively, represents the dimension of the key, represents the weighted feature matrix.

[0032] SW-MSA realizes cross-window feature interaction through a sliding window mechanism, and the sliding window size is .

[0033] According to the technical solution, the multi-channel-multi-space mixed attention module introduces a mixed attention mechanism of channel and space dimensions, enhances the attention to different regions and features in the remote sensing image, and the specific calculation is as follows:

[0034] Suppose that the multi-scale features extracted from the Swin Transformer are wherein represents the feature map of the i-th layer, , , respectively represent the channel number, height and width of the i-th layer feature map; the channel feature vector is obtained by using global average pooling (GAP) and global maximum pooling (GMP), wherein the GAP extracts the average value of each channel, and pays attention to the global context; the GMP extracts the maximum value of each channel, and pays attention to the most significant feature:

[0035]

[0036] In the formula, is the feature map of the i-th stage, is the channel index, is the height index, is the width index, represents the value of the feature map at the channel c, height h, and width w position, is the result of applying global maximum pooling to the channel c of the feature map;

[0037] The two pooling results are spliced and passed through a shared fully connected layer to generate the channel weight:

[0038]

[0039] wherein, represents an activation function, is the weight of the fully connected layer, represents the channel attention weight vector; the feature map is weighted by the channel attention weight:

[0040]

[0041] In the formula, represents the i-th stage feature map after weighting;

[0042] The spatial attention mechanism selectively enhances the features by learning the importance weight of each spatial position; the global average pooling and the global maximum pooling are used to obtain the spatial feature map:

[0043]

[0044] wherein, the global average pooling of the spatial dimension; denotes the spatial feature map of the feature map in the channel c, and the dimension is that is, a two-dimensional matrix;

[0045]

[0046] wherein, the global maximum pooling of the spatial dimension, denotes the maximum value of all channels in each spatial position (h, w);

[0047] The two pooling results are spliced and a shared convolution layer is used to generate the spatial weight; the feature map is weighted through the spatial weight; the processed multi-scale features are up-sampled and spliced to obtain the fusion feature map.

[0048] According to the above technical scheme, the two pooling results are spliced and a shared convolution layer is used to generate the spatial weight, and specifically:

[0049]

[0050] wherein, denotes the spatial attention weight, which is derived from the splicing of the two pooling results and the generation of the shared convolution layer;

[0051] The feature map is weighted through the spatial weight:

[0052]

[0053] In the formula, denotes the feature map of the i-th stage after being weighted by the channel attention and the spatial weight; denotes the spatial attention weight.

[0054] According to the above technical scheme, the processed multi-scale features are up-sampled and spliced to obtain the fusion feature map, and specifically:

[0055]

[0056] wherein, denotes the feature map of the i-th stage after being weighted by the channel attention and the spatial attention; denotes the up-sampling operation; represents a splicing operation; represents a fused feature map obtained by splicing.

[0057] The feature fusion reconstruction module uses a 3x3 convolution layer and a pixel shuffle layer to convert the feature map into a high-resolution image;

[0058] First, the spliced feature map is fused into a single feature map using a convolution layer:

[0059]

[0060] In the formula, conv(·) represents a 3x3 convolution layer, represents a fused feature map, the number of channels is , and the size is ;

[0061] The fused feature map is transformed into a three-channel feature map through a 1x1 convolution layer:

[0062]

[0063] In the formula, represents a 1x1 convolution layer, represents the output three-channel feature map.

[0064] According to the above technical solution, converting the feature map into a high-resolution image is specifically:

[0065] The three-channel feature map is upsampled to the target high-resolution image using sub-pixel convolution (Pixel Shuffle), and the sub-pixel convolution is divided into convolution to increase the number of channels and Pixel Shuffle rearrangement:

[0066]

[0067] In the formula, k represents the size of the convolution kernel, represents the feature map after increasing the number of channels, and r is the upsample factor; Pixel Shuffle rearranges channels to the spatial dimension, and the resolution is enlarged by r times;

[0068]

[0069] In the formula, represents the obtained high-resolution reconstructed 3-channel high-quality image.

[0070] Compared with the prior art, the present application has the following beneficial effects:

[0071] In the present application, the core of super-resolution technology is to reconstruct high-resolution images from low-resolution images, which can restore more detailed features, making the building texture, road signs, vegetation details, etc. in the unmanned aerial photography clearer and more discernible; increasing the number of pixels per unit area, the image is more delicate, and when the image is enlarged, it is not easy to appear blurred, jagged, etc.

[0072] The multi-scale hybrid attention mechanism can extract image features from different scales and comprehensively capture details, such as focusing on fine textures at small scales and grasping overall structures at large scales.

[0073] The attention mechanism allows the model to focus on important areas and suppress irrelevant information interference, such as highlighting target objects in complex scenes.

[0074] The trained HTMA-RS model can accurately learn the mapping relationship from low-resolution to high-resolution images, ensuring the reconstruction quality. BRIEF DESCRIPTION OF DRAWINGS

[0075] Figure 1 The flowchart of the present application;

[0076] Figure 2 The HTMA-RS model architecture diagram of the present application;

[0077] Figure 3 The flowchart of the present application for making data sets;

[0078] Figure 4 The flowchart of the present application for processing real-time aerial images of unmanned aerial vehicles. DETAILED DESCRIPTION

[0079] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0080] Embodiment one

[0081] As shown in Figure 1 , a multi-scale hybrid attention-based unmanned aerial image super-resolution method includes the following steps:

[0082] Step S1, prepare the unmanned aerial images that have been acquired, and cut them to make data sets;

[0083] Step S2, perform degradation processing on the acquired unmanned aerial image data, and the degraded image is used as a low-resolution image data set, while the corresponding high-resolution image is used as a high-resolution image data set;

[0084] Step S3, creating and training the HTMA-RS model;

[0085] Step S4, stitching the high-resolution image output by the model; obtaining a complete high-resolution real-time aerial image.

[0086] In the present application, the core of super-resolution technology is to reconstruct high-resolution images from low-resolution images, which can restore more detailed features, making the building textures, road signs, and vegetation details in the aerial images of the UAV clearer and more distinguishable; increasing the number of pixels per unit area, making the image more delicate, and reducing the occurrence of blurring and jaggedness when zooming in on the image.

[0087] The multi-scale hybrid attention mechanism can extract image features from different scales and capture details comprehensively, such as focusing on fine textures at small scales and grasping overall structures at large scales.

[0088] The attention mechanism allows the model to focus on important areas and suppress irrelevant information interference, highlighting target objects in complex scenes.

[0089] The trained HTMA-RS model can accurately learn the mapping relationship from low-resolution to high-resolution images, ensuring the quality of the reconstruction.

[0090] Embodiment Two

[0091] The present embodiment provides a specific implementation.

[0092] The present application provides a UAV image super-resolution method based on multi-scale hybrid attention, called Hybrid Transformer and Multi-scale Attention for Remote Sensing Super-Resolution (HTMA-RS), and a high-fidelity image degradation method. The model combines self-attention mechanism and multi-scale feature extraction, which can better capture complex textures and spatial information in remote sensing images. The high-fidelity image degradation method can better simulate low-quality images in the real world. The details of the low-quality UAV image are reconstructed to improve the image quality and accuracy.

[0093] First, build the HTMA-RS model, as shown in the scheme Figure 2

[0094] The HTMA-RS model includes three key components: a multi-scale Transformer module, a multi-channel-multi-space hybrid attention module, and a feature fusion reconstruction module.

[0095] ​1. The multi-scale Transformer module uses Swin Transformer (SwinT) designed for processing multi-scale features, which realizes the extraction of multi-scale features of the input image through a hierarchical block structure and a sliding window mechanism. The specific process of extracting multi-scale features first divides the input image into non-overlapping blocks of a fixed size. Assuming that the size of the input image is HxW, and the size of each block is PxP, the image is divided into blocks. Linear embedding is performed on each block to map each block to a high-dimensional feature space. Assuming that the channel number of the input block is C, and the feature dimension after embedding is D, the linear embedding process can be represented as:

[0096]

[0097] wherein and represent the weights and biases of linear embedding, respectively; SwinT adopts window multi-head self-attention (W-MSA) and cross-window multi-head self-attention (SW-MSA) to alternately extract features;

[0098] W-MSA performs self-attention calculation within a fixed-size window. Assuming that the window size is MxM, in each window, multi-head self-attention is calculated:

[0099]

[0100] wherein , , represent the query, key and value matrices, respectively, represents the dimension of the key, represents the weighted feature matrix; SW-MSA realizes cross-window feature interaction through the sliding mechanism of the window, and the sliding window size is Similar to W-MSA, SW-MSA reduces the spatial resolution and increases the feature dimension through block merging operation layer by layer. Assuming that the block size of the th layer is , and the channel number is , the block size of the th layer is , and the channel number is . After processing by multiple SwinTransformer blocks, feature maps of different scales can be obtained. The advantage of Swin Transformer as a feature extractor lies in its hierarchical structure and sliding window mechanism, which can continuously optimize the weighted feature matrix , the model generates a weighted feature matrix by selectively focusing on important information in the input features, thus better adapting to downstream tasks. This design enables the model to effectively capture key information in the features and perceive scale changes.

[0101] SW-MSA achieves cross-window feature interaction through the sliding mechanism of the window, and the size of the sliding window is , that is, the window boundary of the feature map moves right and down by pixels. To achieve the offset, SwinT uses a cyclic shift operation, specifically, the pixels of the feature map are cyclically shifted by units along the horizontal and vertical directions. The feature map after cyclic shift is re-divided into MxM windows, and these windows contain cross-window features compared with the regular windows. For the incomplete windows after offset, attention masks are added to ensure that attention is only calculated in the valid area.

[0102] In the window after offset, multi-head self-attention is performed. Since the new window contains features from different original windows, the attention mechanism naturally achieves cross-window information interaction. The calculation formula of multi-head self-attention is consistent with the formula in the above W-MSA.

[0103] 2. Multi-channel and multi-space hybrid attention module combines the latest natural image super-resolution research model HAT, introduces a hybrid attention mechanism in the channel and spatial dimensions, and enhances the attention to different regions and features in remote sensing images. Assume that the multi-scale features extracted from Swin Transformer are , where represents the feature map of the i-th layer, , , respectively represent the channel number, height and width of the i-th layer feature map; the channel attention mechanism selectively enhances the features by learning the importance weight of each channel. Global Average Pooling (GAP) and GlobalMax Pooling (GMP) are used to obtain the channel feature vector:

[0104]

[0105] The two pooling results are spliced and passed through a shared fully connected layer to generate the channel weight:

[0106]

[0107] where is an activation function, is the weight of the fully connected layer, denotes a channel attention weight vector; the feature map is weighted by the channel weight:

[0108]

[0109] wherein, denotes the feature map after the i-th stage of weighting;

[0110] The spatial attention mechanism selectively enhances the features by learning the importance weight of each spatial position; the spatial feature map is obtained using global average pooling and global maximum pooling:

[0111]

[0112] wherein, denotes the global average pooling of the spatial dimension; denotes the spatial feature map of the feature map in the channel c, with the dimension , i.e. a two-dimensional matrix;

[0113]

[0114] wherein, denotes the global maximum pooling of the spatial dimension, denotes the maximum value of all channels at each spatial position (h, w);

[0115] The two pooling results are spliced and a shared convolution layer is used to generate spatial weights, specifically:

[0116]

[0117] wherein, denotes the spatial attention weight, which is derived from splicing the two pooling results and generating through a shared convolution layer;

[0118] The feature map is weighted by the spatial weight:

[0119]

[0120] wherein, denotes the feature map after the i-th stage of weighting by the channel attention and the spatial weight; denotes the spatial attention weight.

[0121] The processed multi-scale features are upsampled and spliced to obtain a fused feature map, specifically:

[0122]

[0123] wherein, ​denotes the feature map after passing through channel attention and spatial attention weighting in the i-th stage; denotes an up-sampling operation; denotes a concatenation operation; denotes the fused feature map obtained by concatenation.

[0124] 3. The feature fusion reconstruction module uses a 3x3 convolution layer and a pixel shuffle layer to convert the feature map into a high-resolution image. The fused feature map after concatenation is fused into a single feature map using a convolution layer:

[0125]

[0126] wherein, conv(·) denotes a 3x3 convolution layer, denotes the fused feature map, the number of channels is , and the size is ;

[0127] The fused feature map is transformed into a three-channel feature map through a 1x1 convolution layer:

[0128]

[0129] wherein, denotes a 1x1 convolution layer, denotes the output three-channel feature map.

[0130] The three-channel feature map is up-sampled to the target high-resolution image using sub-pixel convolution (Pixel Shuffle), which is divided into convolution to increase the number of channels and Pixel Shuffle rearrangement:

[0131]

[0132] wherein, k denotes the size of the convolution kernel, denotes the feature map after increasing the number of channels, and r is the up-sampling factor;

[0133] Pixel Shuffle rearranges channels to the spatial dimension, and the resolution is enlarged by r times;

[0134]

[0135] wherein, denotes the obtained high-resolution reconstructed 3-channel high-quality image.

[0136] The high-fidelity image degradation method is redesigned, and the scheme is shown in Figure 3 :

[0137] The high-fidelity image degradation method uses three different types of kernels for image processing and adds two different types of noise with random probability. The kernel type for the first degradation is randomly generated, and the kernel types include: iso (isotropic), aniso (anisotropic), generalized_iso (generalized isotropic), generalized_aniso (generalized anisotropic), plateau_iso (plateau isotropic), and plateau_aniso (plateau anisotropic). The kernel size is randomly selected between 7-21, and image edge padding is performed to meet processing requirements. Finally, a JPEG compression is added. The kernel type and generation method for the second degradation are similar to the first degradation kernel, but the parameters may be different. The final sinc kernel (sinc_kernel) generates a sinc low-pass filter kernel for the final processing step. This kernel is used for filtering after JPEG compression to reduce compression artifacts.

[0138] 1. Isotropic kernel refers to a kernel with the same characteristics in all directions. The most common is the Gaussian kernel, whose formula is as follows:

[0139]

[0140] where, is the standard deviation, which controls the spread of the kernel.

[0141] 2. Anisotropic kernel has different characteristics in different directions. An elliptical Gaussian kernel can be used to achieve this, whose formula is as follows:

[0142]

[0143] where, and control the spread in the horizontal and vertical directions, respectively.

[0144] 3. Generalized isotropic kernel is an extension of the isotropic kernel, which can be achieved by adjusting the shape parameter of the kernel. For example, a generalized Gaussian kernel can be used, whose formula is as follows:

[0145]

[0146] where, is the shape parameter, which controls the shape of the kernel.

[0147] 4. Generalized anisotropic kernel is an extension of the anisotropic kernel, which is achieved by adjusting the shape parameter of the kernel. For example, a generalized elliptical Gaussian kernel can be used, whose formula is as follows:

[0148]

[0149] where, and are shape parameters, controlling the shape in horizontal and vertical directions, respectively.

[0150] 5. The platform isotropic kernel is an isotropic kernel with a flat top, which can be achieved by adjusting the parameters of the kernel. For example, a flat-top Gaussian kernel can be used, whose formula is as follows:

[0151]

[0152] where r is the radius of the platform.

[0153] 6. The platform anisotropic kernel is an anisotropic kernel with a flat top, which can be achieved by adjusting the parameters of the kernel. For example, a flat-top elliptical Gaussian kernel can be used, whose formula is as follows:

[0154]

[0155] where r is the radius of the platform.

[0156] The size of the kernel is randomly selected between 7 and 21, and image edge padding is performed to adapt to the processing requirements. The specific steps are as follows: 1. Randomly select the size k of the kernel between 7 and 21; 2. Generate the kernel according to the selected kernel type and parameters; 3. Use the generated kernel to perform convolution operation on the image to process the image.

[0157] After the first and second quality reduction processing, a JPEG compression is added to the image. In the final processing step, a sinc low-pass filter kernel is generated, whose formula is as follows:

[0158]

[0159] In the formula, sinc(x) represents the low-pass filter kernel, and sin(πx) represents the function value of the sine function when the independent variable is πx; this kernel is used for filtering after JPEG compression to reduce artifacts caused by compression.

[0160] In image degradation processing, adding noise is an important step. This scheme involves two main types of noise: Gaussian noise and Poisson noise. The addition probability of the two types of noise is 50%, and the standard deviation range of Gaussian noise is set to 1-15, and the scale range of Poisson noise is set to 0.05-1. In addition, there is a 40% probability of continuing to add gray noise.

[0161] Gaussian noise refers to noise that follows a Gaussian (normal) distribution, typically represented by a Gaussian distribution with a mean of zero. The formula for adding Gaussian noise is as follows:

[0162]

[0163] where, is the pixel value of the original image at position , is the pixel value of the image after adding noise at position . is the Gaussian noise with a mean of 0 and a variance of .

[0164] Poisson noise refers to noise that follows a Poisson distribution, which is usually proportional to the intensity of the image. The formula for adding Poisson noise is as follows:

[0165]

[0166] where, is the pixel value of the original image at position , is the pixel value of the image after adding noise at position . is a random variable following a Poisson distribution with parameter .

[0167] Salt-and-Pepper noise is a common type of noise that randomly sets some pixel values to the minimum or maximum value in the image. The formula for adding salt-and-pepper noise is as follows:

[0168]

[0169] where, is the pixel value of the original image at position , is the pixel value of the image after adding noise at position . P is the proportion of noise, indicating how many proportions of pixels will be affected by noise.

[0170] After the construction is completed, the constructed network model is trained:

[0171] The first step is to prepare the unmanned aerial vehicle image that has been obtained, and cut it to make a data set. After two times of quality reduction, the reduced image is used as a low-resolution image data set, and its corresponding high-resolution image is used as a high-resolution image data set.

[0172] The second step is to input the processed data set into the network for training until fitting or reaching the maximum number of training times.

[0173] Third step, save the trained network model.

[0174] Real-time processing of UAV aerial images, as shown in the scheme Figure 4

[0175] First, obtain the UAV aerial images of a certain area;

[0176] Second, cut the obtained images into small blocks;

[0177] Third, input the cut images into the trained network model to get high-resolution output results;

[0178] Fourth, splice the output high-resolution images according to the original images to get complete high-resolution images.

[0179] The following are the key formulas of the image segmentation module and the splicing module and their simple introduction.

[0180] First, determine the top-left corner coordinates of the small block , where is the input size of the model. Since it does not include overlapping areas, the step is directly equal to the size of the small block, i.e. . Then the number of small blocks in the height and width directions can be represented as:

[0181]

[0182] The total number of small blocks is the original size of the remote sensing image. Then the cutting formula can be expressed as:

[0183]

[0184] where is the original input remote sensing image. The function of this module is to cut the large size remote sensing image into n non-overlapping small blocks. i, j are the indices of the small block, the index starts from 0, representing the i-th row and j-th column small block, which is a rectangular region extracted from the original image, corresponding to a small block with a certain index position (i, j), and each small block size is to adapt to the model input. The cutting formula is to take a WxH small block with a width interval W and a height interval H. For example, assuming that the size of the original remote sensing image is 100x100, and H and W are set to 10 and 10 respectively. Then a 10x10 sub-block can be taken out each time, and 100 non-overlapping sub-blocks can be taken out.

[0185] First, determine the top-left corner coordinates of the small block after super-resolution , where ​and is the post-super-resolution patch size, r is the up-sampling factor, i.e. where H and W are the original patch size.

[0186] Since the non-overlapping cropping is used when cropping, the stride after super-resolution is directly equal to the post-super-resolution patch size, i.e. Then the stitching formula can be expressed as:

[0187]

[0188] where, is the complete super-resolution image after stitching, is the (i, j)th post-super-resolution patch. and The position range of the patch in the reconstructed image is defined. The role of this module is to stitch the n post-super-resolution non-overlapping patches into a complete remote sensing image The post-super-resolution stride is used to align the patch positions, ensuring accurate reconstruction of the entire image, which is suitable for output reconstruction of the super-resolution task.

[0189] Through the method in the present application, the analysis accuracy can be effectively improved: in the fields of surveying and mapping, inspection, etc., high-resolution images provide more accurate data for terrain analysis, target detection and identification, etc., such as clearly distinguishing the state of line components during power inspection. Make the UAV image applied in more scenes with high requirements for image quality, such as high-definition image map making, film shooting, etc.

[0190] It should be noted that in this text, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device.

[0191] Finally, it should be noted that: the above only describes the preferred embodiments of the present application, and is not used to limit the present application, although the present application has been described in detail with reference to the foregoing embodiments, and those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for unmanned aerial vehicle (UAV) image super-resolution based on multi-scale hybrid attention, characterized in that: The method comprises the following steps: Step S1, preparing the acquired unmanned aerial vehicle image, cutting it to make a data set; Step S2, performing quality reduction processing on the acquired unmanned aerial vehicle image data, the image after quality reduction being used as a low-resolution image data set, and the corresponding high-resolution image being used as a high-resolution image data set; Step S3, creating and training an HTMA-RS model; the HTMA-RS model comprises a multi-scale Transformer module, a multi-channel-multi-space hybrid attention module, and a feature fusion reconstruction module; The multi-scale Transformer module is used for extracting multi-scale features of the input image, specifically: first, the input image is divided into fixed-size non-overlapping blocks, assuming that the size of the input image is HxW, and the size of each block is PxP, then the image is divided into blocks; performing linear embedding on each block, mapping each block to a high-dimensional feature space; assuming that the channel number of the input block is , and the dimension of the embedded feature is D, then the linear embedding process can be represented as: wherein, and denote the weights and biases of linear embedding, respectively; SwinT adopts windowed multi-head self-attention (W-MSA) and cross-window multi-head self-attention (SW-MSA) to alternate feature extraction; W-MSA performs self-attention calculation in a fixed-size window; assuming that the window size is MxM, in each window, multi-head self-attention is calculated: wherein, , , denote the query, key and value matrices, respectively, denotes the dimension of the key, denotes the weighted feature matrix; SW-MSA realizes the cross-window feature interaction through the sliding mechanism of the window, and the sliding window size is ; The multi-channel-multi-space hybrid attention module introduces a hybrid attention mechanism of channel and space dimensions, enhances the attention to different regions and features in the remote sensing image, and the specific calculation is as follows: Assume that the multi-scale features extracted from the Swin Transformer are wherein represents the feature map of the i-th layer, , , respectively represent the channel number, height and width of the i-th layer feature map; the channel feature vector is obtained using global average pooling (GAP) and global maximum pooling (GMP), wherein the GAP extracts the average value of each channel, focusing on the global context; the GMP extracts the maximum value of each channel, focusing on the most significant features: wherein is a feature map of the i-th stage, is a channel index, is a height index, is a width index, denotes the value of the feature map at channel c, height h, width w position, is the result after applying global max pooling to the channel c of the feature map; The two pooling results are spliced and passed through a shared fully connected layer to generate channel weights: wherein, denotes an activation function, is a fully connected layer weight, denotes a channel attention weight vector; the feature map is weighted by the channel attention weight: In the formula, denotes the elapsed the weighted feature map of the i-th stage; The spatial attention mechanism learns the importance weights of each spatial position to selectively enhance the features; global average pooling and global maximum pooling are used to obtain spatial feature maps: wherein, global average pooling representing the spatial dimension; feature map spatial feature map on channel c, dimension i.e. a two-dimensional matrix; wherein, global max pooling over spatial dimensions, taking the maximum value over all channels at each spatial position (h, w); The two pooling results are spliced and passed through a shared convolution layer to generate spatial weights: the feature maps are weighted through the spatial weights: the processed multi-scale features are upsampled and spliced to obtain a fusion feature map; The feature fusion reconstruction module uses a 3x3 convolution layer and a pixel shuffle layer to convert the feature map into a high-resolution image; First, a convolution layer is used to fuse the spliced feature maps into a single feature map: In the formula, conv(·) represents a 3x3 convolution layer, represents a fusion feature map, the number of channels is , and the size is ; A 1x1 convolution layer is used to transform the fusion feature map into a three-channel feature map: In the formula, denotes a 1x1 convolutional layer, denotes the output three-channel feature map; Step S4, splicing the high-resolution image output by the model; obtaining a complete high-resolution real-time aerial image.

2. The unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention according to claim 1, characterized in that: In step S2, the quality reduction processing on the unmanned aerial vehicle image data comprises the following steps: Step S201, cropping the unmanned aerial vehicle image data to obtain a high-resolution image; Step S202, performing a quality reduction processing on the cropped unmanned aerial vehicle image data, including random kernel, JPEG compression, and random noise processing; Step S203, performing a secondary quality reduction processing on the unmanned aerial vehicle image data, including random kernel, sinc kernel, JPEG compression, and random noise processing; Step S204, obtaining a low-resolution data set after the processing.

3. The unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention according to claim 2, characterized in that: In the primary quality reduction processing, the random kernel includes iso kernel, aniso kernel, generalized_iso kernel, generalized_aniso kernel, plateau_iso kernel, and plateau_aniso kernel; the size of the kernel is randomly selected between 7 and 21, and image edge padding is performed to meet the processing requirements.

4. The unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention according to claim 3, characterized in that: After the primary and secondary quality reduction processing, a JPEG compression is added to the image to generate a sinc low-pass filter kernel, and the formula is as follows: In the formula, sinc(x) represents the low-pass filter kernel, and sin(πx) represents the function value of the sine function when the independent variable is πx. The low-pass filter kernel is used for filtering after JPEG compression to reduce the artifacts caused by compression.

5. The unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention according to claim 1, characterized in that: The two pooling results are spliced and a shared convolution layer is used to generate spatial weights, which are specifically: wherein, denotes spatial attention weights, derived from concatenating the two pooling results and passing through a shared convolutional layer; The feature map is weighted by the spatial weight: In the formula, denotes the feature map of the i-th stage weighted by the channel attention and spatial weight; denotes the spatial attention weight.

6. The unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention according to claim 5, characterized in that: The processed multi-scale features are up-sampled and spliced to obtain the fused feature map, which is specifically: wherein, represents the feature map of the i-th stage after passing through channel attention and spatial attention weighting; represents an up-sampling operation; represents a concatenation operation; represents the fusion feature map obtained by concatenation.

7. The unmanned aerial vehicle image super-resolution method based on multi-scale hybrid attention according to claim 1, characterized in that: The feature map is converted to a high-resolution image, which is specifically: The three-channel feature map is up-sampled to the target high-resolution image using sub-pixel convolution (Pixel Shuffle), which includes convolution to increase the number of channels and Pixel Shuffle to rearrange: where k represents the convolution kernel size, represents the feature map after increasing the number of channels, r is the up-sampling factor; Pixel Shuffle rearranges the channels to the spatial dimension, and the resolution is enlarged r times. wherein, denotes the resulting high-resolution reconstructed 3-channel high-quality image.

Citation Information

Patent Citations

  • Insulator image resolution enhancement method based on super-resolution algorithm

    CN111667409A

  • Method for improving unmanned aerial vehicle aerial image resolution based on deep learning

    CN113469881A

  • Unmanned aerial vehicle image super-resolution reconstruction method, system and equipment

    CN117893410A

  • Image super-resolution reconstruction method, device and equipment

    CN115564649A

  • Enhancement processing method and system for aerial image of unmanned aerial vehicle

    CN117911908A