Real-time multi-source remote sensing image small target detection method using super-resolution assisted reasoning
Patent Information
- Application Number
- CN202510082463.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-17
- Filing Date
- 2025-01-20
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-01-20
AI Technical Summary
[0004]同时,将超分辨率(Super Resolution,SR)领域的重建网络直接应用在遥感领域,由于影像背景复杂,会出现特征重建不明显的现象,同时SR重建更多关注全局特征重建,对于影像中小目标特征不能有效增强
[0020]本方法先进、科学,通过结合超分辨率技术能有效重建出更多小目标特征,提升多源遥感影像小目标检测的精度。结果表明,该方法具有较高的检测准确率,且降低了模型的计算开销,具有很好的应用价值。
Smart Images

Figure CN119888195B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent interpretation of remote sensing images, and in particular to super-resolution technology and methods for small target detection in multi-source remote sensing images. Background Technology
[0002] Remote sensing images contain rich information about ground features. By extracting different target features from these images, various targets can be accurately identified and located, playing a crucial role in the intelligent processing and analysis of remote sensing images. Currently, remote sensing technology is widely used in urban planning and construction, resource and environmental monitoring, smart agriculture, and military reconnaissance, possessing significant civilian and military application value. With the continuous development of sensor technology, more and more remote sensing satellites are being launched, bringing massive amounts of remote sensing image data. This rich image resource presents both new opportunities and challenges for the intelligent interpretation of remote sensing images. Target detection in remote sensing images, as one of the important research areas in the field, accurately identifies and detects targets by extracting image features. Therefore, researching target detection methods in remote sensing images is of great significance for broadening the application scope of remote sensing images.
[0003] The fundamental reasons why small targets are difficult to detect in remote sensing images can be divided into two points: First, as can be seen from the image imaging principle, the Ground Sample Distance (GSD) varies between different images, resulting in large differences in target size within the same category. Under the same GSD conditions, small targets occupy a smaller proportion of image pixels and lack the texture and detail features required for target detection, making them easily confused with the complex background of the remote sensing scene. Therefore, they are more difficult to detect compared to medium and large targets. Second, in current target detection frameworks based on deep learning technology, the model extracts features through multiple convolutions. However, the features of small targets are mostly present in shallow feature maps, and these features are easily partially or completely lost in the deep feature maps obtained from multiple convolutions. Therefore, how to solve the problem of small target feature loss during convolution is another challenge.
[0004] Furthermore, directly applying reconstruction networks from the Super Resolution (SR) domain to remote sensing often results in indistinct feature reconstruction due to complex image backgrounds. Additionally, SR reconstruction focuses primarily on global feature reconstruction, failing to effectively enhance the features of small targets within the image. Moreover, methods that fuse different image features to improve detection accuracy are complex, computationally expensive, and yield unsatisfactory results. Therefore, only by rationally utilizing SR technology can the detection accuracy of small targets in remote sensing images be effectively improved.
[0005] Therefore, this invention discloses a real-time multi-source remote sensing image small target detection method using super-resolution assisted inference. Its main contents are as follows: First, a pixel-level fusion method combining channel attention mechanism is proposed. By capturing global channel features, it focuses on more features beneficial to the detection task in visible light and infrared images, reducing the fusion of redundant features. Second, a super-resolution assisted inference branch combining boundary enhancement is constructed to improve the backbone network's ability to extract and represent high-resolution features of small targets in multi-source remote sensing images. Finally, to reduce computation and accelerate model inference time, medium and large-sized anchor boxes in the YOLO architecture are removed, and anchor boxes for small targets are added, reducing the number of feature layer fusion layers. Under the condition of reducing the number of parameters, the IR and RGB image features are fully fused, improving the small target detection accuracy. Summary of the Invention:
[0006] In view of this, the present invention proposes a real-time multi-source remote sensing image small target detection method using super-resolution assisted inference, including three aspects: multi-source feature fusion, feature extraction and super-resolution reconstruction of feature maps, and target detection.
[0007] Firstly, this invention proposes a multi-source feature fusion method, the steps of which are as follows:
[0008] S1: Extract features from the input image using 1×1 convolution to generate a feature matrix of fixed size.
[0009] S2: Use channel attention mechanism to capture global channel information, pay more attention to features that have a gain on the target region among different modal channel features, and reduce the fusion of redundant features.
[0010] S3: The extracted features from different modalities are integrated channel by channel, and the features are fused using an attention mechanism and a splicing method.
[0011] The second aspect involves feature extraction and super-resolution reconstruction of feature maps, with the following steps:
[0012] S4: Use the YOLOv8s method to extract model features by progressively extracting features through stacked Fusion, C2f, and SPPF modules.
[0013] S5: The backbone network uses the FPN+PANET method to achieve multi-scale feature fusion of the extracted features, enabling the model to focus on more targets with large scale differences.
[0014] S6: The features extracted by the backbone network are fed into the constructed ES-WDSR branch to reconstruct features. This branch combines the Sobel edge feature extraction operator to reconstruct the edge features of the target, reconstructs a high-resolution feature map, and then merges it with the backbone network features again to improve the extraction and representation capabilities of the backbone network and high-resolution features.
[0015] Thirdly, target detection, the steps are as follows:
[0016] S7: Initially detect the target and calculate the loss value using the loss function.
[0017] S8: Backpropagation, gradient update, further training, until the loss function converges.
[0018] S9: Obtain the optimal weights and classify and locate the targets.
[0019] S10: End.
[0020] This method is advanced and scientific. By combining super-resolution technology, it can effectively reconstruct more features of small targets, thereby improving the accuracy of small target detection in multi-source remote sensing images. Results show that this method has high detection accuracy and reduces the computational cost of the model, demonstrating significant application value. Attached image description:
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only schematic diagrams of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0022] Figure 1 Flowchart
[0023] Figure 2 Framework diagram of target detection network for multi-source remote sensing images
[0024] Figure 3 For the fusion module diagram
[0025] Figure 4 ES-WDSR decoder diagram
[0026] Figure 5 Vehicle target detection results from multi-source remote sensing imagery
[0027] Figure 6 pedestrian target detection results Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0029] The following is the multi-source feature fusion section:
[0030] Step 1: Extract features from the input image using a 1×1 convolution to generate a fixed-size feature matrix. Specifically, if the input is IR and RGB remote sensing images, first extract image features using a 1×1 convolution, with the following formula:
[0031] F RGB =Conv(I RGB ), F IR =Conv(I IR (1)
[0032] In the formula, I RGB and I IR The inputs are the visible light image (RGB) and the near-infrared image (IR), respectively. Conv is a regular 1×1 convolution, and F... RGB and F IR These are the extracted visible light and near-infrared features, respectively.
[0033] Step 2: Use the channel attention mechanism to capture global channel information, pay more attention to the features that have a gain on the target region among the features of different modal channels, and reduce the fusion of redundant features.
[0034] A RGB =CA(F RGB ), A IR =CA(F IR (2)
[0035] In the formula, CA represents the channel attention mechanism, and A RGB and A IR The input features are the features after channel attention processing.
[0036] Step 3: Integrate the extracted features from different modalities channel by channel, using an attention mechanism and a concatenation method to fuse the features. The calculation method is shown in formulas (3) and (4):
[0037] F fus1 =F RGB +A RGB F fus2 =F IR +A IR (3)
[0038] F ful =CA(Concat(F) fus1 , F fus2 (4)
[0039] In the formula, Concat is the concatenation method, and F ful The fused features are obtained by processing the spliced features using the ECA method.
[0040] The following is the feature extraction and super-resolution reconstruction feature map section:
[0041] Step 4: Use the YOLOv8s method to extract model features by gradually extracting features F through stacked Fusion, C2f, and SPPF modules.
[0042] Step 5: The backbone network uses the FPN+PANET method to achieve multi-scale feature fusion of the extracted features, enabling the model to focus on more targets with large scale differences.
[0043] Step 6: Feed the features F extracted by the backbone network into the constructed ES-WDSR branch to reconstruct features. This branch combines the Sobel edge feature extraction operator to reconstruct the edge features of the target, reconstructs a high-resolution feature map, and then merges it with the backbone network features again to improve the backbone network's ability to extract and represent high-resolution features.
[0044] The soft threshold calculation method is shown in formula (5).
[0045]
[0046] In the formula, F Soft F is the output feature, τ is the input feature, and τ is the threshold. Features in the interval [-τ, τ] are set to 0. Features greater than τ are subtracted by τ, and features less than -τ are added by τ. That is, features with absolute values below a certain threshold are set to zero, and other features are also adjusted towards zero, thus achieving a shrinkage effect and removing redundant features. Specifically, after convolution in the super-resolution branch residual module, the features are input into the soft thresholding module, and the soft thresholds for each channel are obtained using formulas (6), (7), and (8). First, the absolute values of the input features are taken in the sub-network, and after global average pooling, the feature map is adjusted to... Secondly, the processed features are passed through a fully connected layer and an RBL layer to obtain Z. The result is then activated by a Sigmoid nonlinear activation to obtain α. Finally, the result of global average pooling is averaged and multiplied by α to obtain a soft threshold τ. The original input is then thresholded using formula (1).
[0047] Z = FC(RBL(GAP(|F|)) (6)
[0048]
[0049] τ=α·average(GAP(F)) (8)
[0050] In the formula, GAP is the global average pooling method, |F| is the absolute value of the input feature F, RBL is a set of three operations: LeakyReLU activation function, BathNormLize batch normalization, and Linear layer, respectively, and average is the averaging method. The improved super-resolution branch ensures that the thresholds of all channels are positive and avoids the case where all outputs are zero. In addition, the soft thresholding module stacked in the neck network can adaptively generate an independent set of thresholds for each group of samples, reducing the influence of inconsistent noise caused by different backgrounds in the samples, thereby effectively reducing the weights of each channel.
[0051] The edge feature extraction method employs the classic first-order edge detection operator, the Sobel operator, which utilizes the gradient information of the image to detect edges. The Sobel operator uses two convolutional kernels to calculate the gradients of the image in the horizontal and vertical directions, respectively. The forms of these two convolutional kernels are as follows:
[0052] Horizontal gradient direction
[0053] Vertical gradient direction In the formula, M x M is the convolution kernel along the horizontal gradient direction. y It is a vertical convolution kernel.
[0054] First, a Sobel horizontal convolution kernel is applied to the processed grayscale image to obtain the horizontal gradient. Then, a Sobel vertical convolution kernel is used to extract the vertical gradient information. Finally, by combining the obtained horizontal and vertical gradient values, the edge strength and direction of each pixel in the image are calculated to extract edge features. These edge features are then fused with the features processed by the soft-thresholding network to enhance the edge features of the target in the reconstruction network, as shown in formulas (11)(12)(13)(14).
[0055] G Sobel_x =F input ×M x (11)
[0056] G Sobel_y =F input ×M y (12)
[0057]
[0058] F Out =F Soft +F Sobel (14)
[0059] In the formula, G Sobel_x G Sobel_yF represents the gradient values on the horizontal and vertical gradients, respectively. input F represents the features extracted by the encoder. Sobel For the extracted gradient features, F Soft F represents the features obtained by the soft thresholding method. Out For F Soft Output features after fusing edge features.
[0060] The following is the object detection section:
[0061] Step 7: Initially detect the target and calculate the loss value using a loss function.
[0062] Step 8: Backpropagation, gradient update, further training, until the loss converges.
[0063] In this invention, the overall network loss L is detected. all It consists of two parts: target detection loss L o and super-resolution branch loss L sr ,Right now:
[0064] L all =a1L o +a2L sr (15)
[0065] In the formula, a1 and a2 are the modulation coefficients of the two tasks, used to balance the two training tasks.
[0066] L sr =||SR-X||1 (16)
[0067] In the formula, SR is the reconstructed feature image, and X is the input image.
[0068] Object detection loss function L o It consists of three parts: localization loss L loc Confidence loss L obj And classification loss L cls ,Right now:
[0069]
[0070] In the formula, l represents two scales of the detection head, and a l b l c l These are the weights of the detection head, and λ represents the weights of the detection head. loc , λ obj , λ cls These are the modulation coefficients for the three parts, used to adjust the weights of localization loss, confidence loss, and classification loss on the target detection task.
[0071] Step 9: Obtain the optimal weights and classify and locate the targets.
[0072] Step 10: End.
[0073] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. It should be noted that those skilled in the art can make several improvements and modifications without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A real-time multi-source remote sensing image small target detection method using super-resolution assisted inference, comprising a multi-source feature fusion step, a feature extraction and super-resolution feature map reconstruction step, and a target detection step: The steps of multi-source feature fusion are as follows: S1: Extract features from the input image using 1×1 convolution to generate a feature matrix of fixed size; S2: Use channel attention mechanism to capture global channel information, pay more attention to features that have a gain on the target region among different modal channel features, and reduce the fusion of redundant features; S3: The extracted features from different modalities are integrated channel by channel, and the features are fused using an attention mechanism and a splicing method; The steps for feature extraction and super-resolution feature map reconstruction are as follows: S4: Use the YOLOv8s algorithm to extract model features, and extract features step by step through stacked Fusion, C2f, and SPPF modules; S5: The backbone network uses the FPN+PANET method to achieve multi-scale feature fusion of the extracted features, enabling the model to focus on more targets with large scale differences; S6: The features extracted by the backbone network are fed into the constructed ES-WDSR branch to reconstruct features. This branch combines the Sobel edge feature extraction operator to reconstruct the edge features of the target, reconstructs a high-resolution feature map, and then merges it with the backbone network features to improve the extraction and representation capabilities of the backbone network and high-resolution features. The target detection steps are as follows: S7: Initial target detection, using a loss function to calculate the loss value; S8: Backpropagation, gradient update, further training, until the loss converges; S9: Obtain the optimal weights and classify and locate the targets; S10: End.
2. The method for real-time multi-source remote sensing image small target detection using super-resolution assisted inference as described in claim 1, characterized in that, In step S7, the target detection model weights generated in steps S1 to S6 are used.
Citation Information
Patent Citations
Super-resolution reconstruction method, system and device for underwater fish image
CN111080531A
Infrared image multi-target pedestrian identification method
CN111597967A