Remote sensing target detection method based on serial pyramid convolution
By adopting a serial pyramid convolutional network structure in remote sensing object detection, the problems of large amount of parameters and high computational complexity in the prior art are solved, and the object detection method for lightweight design and efficient calculation are realized.
Patent Information
- Application Number
- CN202510090052.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing remote sensing object detection methods have problems such as large amount of parameters and high calculation complexity, making it difficult to achieve a lightweight design.
A network structure based on serial pyramid convolution is adopted, and multi-scale convolution is achieved by connecting a 3×3 convolution kernel in series, reducing the amount of parameters and improving the computing efficiency.
It effectively reduces the amount of parameters and calculation complexity, improves the lightweight design performance of remote sensing target detection, and ensures the accuracy of output results.
Smart Images

Figure CN120125958A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a remote sensing target detection method based on serial pyramid convolution. Background Art
[0002] Remote sensing technology first emerged in the early 20th century and was used for military reconnaissance and geographical survey. With the development of aviation and satellite technology, especially the maturity of satellite remote sensing technology in the 1960s, the acquisition of remote sensing data has become more popular. Initially, the processing of remote sensing data mainly relied on manual annotation and simple image processing algorithms. However, with the enhancement of computing power and the rise of artificial intelligence (AI) technology, automatic target detection technology based on deep learning has been widely applied in the field of remote sensing.
[0003] The definition of remote sensing target detection is to use remote sensing image data to automatically identify and locate the objects of interest in the target area. Usually, these objects include natural resources (such as forests, oceans), man-made objects (such as buildings, bridges, roads), military facilities (such as tanks, airplanes), crops, etc.
[0004] Remote sensing target detection technology is widely applied in multiple fields, mainly including:
[0005] Land use and land cover classification: By detecting objects such as buildings, roads, and farmland, the land use situation and changes can be understood.
[0006] Environmental monitoring and natural disaster assessment: Timely monitoring and assessment of natural disasters such as floods, fires, and earthquakes help formulate emergency response plans.
[0007] Urban planning and infrastructure monitoring: Remote sensing images are used for monitoring and planning urban expansion, and identifying changes in infrastructure such as roads and bridges.
[0008] Agricultural monitoring: By detecting the planting area and growth status of crops, remote sensing can help evaluate agricultural yields and provide decision-making support.
[0009] Military reconnaissance and national defense security: In the military field, remote sensing is used to identify and track military targets such as vehicles, ships, and airplanes.
[0010] Ocean and climate research: Monitoring environmental changes such as ocean vessels, glacier melting, and ocean pollution.
[0011] Traditional remote sensing target detection methods mainly include:
[0012] Threshold-based detection methods: According to the gray-scale or color feature differences of different targets in the image, the targets are separated from the background by setting appropriate thresholds.
[0013] Edge detection-based detection method: Utilize the gray-scale changes at the edges between the target and the background in the image for detection.
[0014] Texture feature-based detection method: Based on the statistical or frequency features of the image, extract the texture information of the target area and distinguish it from the background.
[0015] Shape feature-based detection method: Utilize the geometric features of the target object for the recognition and localization of the target.
[0016] In recent years, relying on their powerful feature representation capabilities, deep learning models have proposed a large number of high-quality solutions for remote sensing target detection tasks.
[0017] A large number of studies have focused on methods such as improving the oriented bounding box detector, improving the loss function, and improving the backbone network to enhance the target detection performance, and have achieved remarkable success.
[0018] In terms of the oriented bounding box detector, studies represented by O-RCNN have effectively solved the target orientation problem by introducing neck refinement techniques, extracting rotated regions of interest, designing specific detections, and developing new target representation methods.
[0019] In terms of image feature extraction, studies such as ARC have effectively improved the network performance through designs that adapt to the target orientation, but have introduced a large amount of additional computational complexity. Studies such as LSKNet and PKINet focus on achieving multi-level feature extraction through diverse convolutional kernel sizes and receptive fields. However, the use of large kernels increases the model's parameter quantity and computational complexity.
[0020] Therefore, designing a lightweight remote sensing target detection network is of great research significance and practical value. Summary of the Invention
[0021] The purpose of the present invention is to overcome the above-mentioned shortcomings and deficiencies of the prior art, and provide a remote sensing target detection method based on serial pyramid convolution. The present invention replaces large kernel convolutions with serial 3×3 convolution kernels, maximally shares parameters, and adopts a pyramid structure to organically combine small-scale features with large-scale features, effectively reducing the parameter quantity and improving the operation efficiency.
[0022] The present invention is realized through the following technical solutions:
[0023] A remote sensing target detection method based on serial pyramid convolution, comprising the following steps:
[0024] S1 Preprocess the input remote sensing image to obtain the preprocessed image, that is, perform operations such as rotating, flipping, cropping, scaling, and normalizing the image. Image preprocessing can enhance image features, reduce noise, and improve the processing speed and algorithm accuracy.
[0025] S2-1 Input the preprocessed remote sensing image into the Stem Layer of the encoder based on our proposed Serial Pyramid Convolutional Network (SPCNet). This layer uses multiple convolutional layers to extract low-level features of the input image, reducing the spatial dimension of the feature map through convolutional operations, thereby reducing computational complexity and increasing the receptive field. Output the feature map at this stage.
[0026] S2-2 First, perform resolution adjustment and channel adjustment on the obtained feature map through downsampling. Subsequently, change the number of channels to twice the number of hidden layer channels through 1×1 convolution for channel splitting operations. Then, through channel splitting, the feature is split into two features with shapes both being [C hid ,H out ,W out . These two features respectively enter the feed-forward neural network and the SPCLayer for feature extraction.
[0027] S2-3 The feed-forward neural network first performs 1×1 convolution on the feature map, then performs depthwise separable convolution with n×n (n = 5, 7, 9, 11), and finally performs 1×1 convolution and outputs the feature map
[0028] S2-4 The serial pyramid convolution module realizes the effect of multi-scale convolution by concatenating 3×3 convolutional kernels. Specifically, it extracts features with receptive field ranges of 3×3, 7×7, 11×11, and 17×17 through 4 serial 3×3 dilated convolutions. To expand the receptive field while reducing the number of parameters, these 4 convolutions are all depthwise separable convolutions with dilation rates of 1, 2, 2, and 3 respectively.
[0029] S3 Obtain image features at different levels from texture to semantics respectively.
[0030] S4 Use the feature fusion module to fuse image features at different levels; perform channel concatenation operations on multi-scale features and use 1×1 convolution for feature aggregation.
[0031] S5 Input the fused features into the decoder for decoding and output the predicted saliency map.
[0032] Through the above process, the predicted output map for remote sensing object detection can be obtained.
[0033] Compared with the prior art, the present invention has the following advantages and effects:
[0034] The present invention uses a deep learning-based algorithm, avoiding the disadvantage of the traditional algorithm that requires a large number of manually designed features. In addition, the method based on the traditional algorithm often can only mark out a rough area, while the deep learning method only needs to input the data into the network for multiple rounds of training to obtain a high-quality saliency detection result.
[0035] The present invention converts the parallel large-kernel convolution into a shared serial small-kernel convolution by introducing stacked dilated convolutions, greatly compressing the computational complexity while locking the receptive field, and achieving a lightweight design.
[0036] The present invention combines the mainstream feature fusion and feature decoding technologies, effectively ensuring the accuracy of the output result. Brief Description of the Drawings
[0037] Figure 1 It is a flowchart of remote sensing target detection based on serial pyramid convolution of the present invention.
[0038] Figure 2 It is the operation process of the SPC layer.
[0039] Figure 3 It is the operation process of the SPC block.
[0040] Figure 4 It is the original image on DOTA and the remote sensing target detection result of SPCNet, and the edge of the segmentation area is marked with a red line on the blank image. Detailed Embodiments
[0041] The following further describes the present invention in detail with specific embodiments.
[0042] The present invention discloses a remote sensing target detection method based on serial pyramid convolution, which can be realized through the following steps:
[0043] S1 Preprocess the input remote sensing image to obtain the preprocessed image; that is, perform operations such as rotation, flipping, cropping, scaling, and normalization on the image. Image preprocessing can enhance image features, reduce noise, and improve the processing speed and algorithm accuracy.
[0044] S2-1 Input the preprocessed remote sensing image into the Stem Layer of the encoder of the Serial Pyramid Convolutional Network (SPCNet) proposed by us. This layer uses multiple convolutional layers to extract the low-level features of the input image, reducing the spatial dimension of the feature map through convolutional operations, thereby reducing the computational complexity and increasing the receptive field. Output the feature map at this stage.
[0045] S2-2 First, perform resolution adjustment and channel adjustment on the obtained feature map through downsampling, and then change the number of channels to twice the number of hidden layer channels through 1×1 convolution for channel splitting operations. Then, through channel splitting, the feature is split into two features with shapes both being [C hid ,H out ,W out . These two features respectively enter the feedforward neural network and the SPCLayer for feature extraction.
[0046] S2-3 The feedforward neural network first performs 1×1 convolution on the feature map, then performs depthwise separable convolution with n×n (n = 5, 7, 9, 11), and finally performs 1×1 convolution and outputs the feature map
[0047] S2-4 The serial pyramid convolution module realizes the effect of multi-scale convolution by concatenating 3×3 convolution kernels. Specifically, it extracts features with receptive field ranges of 3×3, 7×7, 11×11, and 17×17 through 4 serial 3×3 dilated convolutions. To expand the receptive field while reducing the number of parameters, these 4 convolutions are all depthwise separable convolutions with dilation rates of 1, 2, 2, and 3 respectively.
[0048] S3 Obtain image features at different levels from texture to semantics respectively.
[0049] S4 Use the feature fusion module to fuse image features at different levels. Perform channel concatenation operations on the multi-scale features and use 1×1 convolution for feature aggregation.
[0050] S5 Input the fused features into the decoder for decoding and output the predicted saliency map.
[0051] As described above, the present invention can be preferably implemented.
[0052] The implementation manners of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.
Claims
1. A remote sensing target detection method based on serial pyramid convolution, characterized in that The steps include: S1 preprocesses the input remote sensing image to obtain a preprocessed image; S2 inputs the preprocessed remote sensing image into the Stem Layer of the serial pyramid convolutional network encoder, which uses multiple convolutional layers to extract low-level features of the input image and reduces the spatial dimension of the feature map through convolution operations to reduce computational complexity and increase the receptive field; outputs the feature map of this stage; S3 obtains image features at different levels from texture to semantics; S4 uses the feature fusion module to fuse image features at different levels; performs channel splicing operations on multi-scale features and uses 1×1 convolution for feature aggregation; S5 inputs the fused features into the decoder for decoding and outputs the predicted saliency map.
2. The remote sensing target detection method based on serial pyramid convolution according to claim 1, characterized in that: In step S1, the preprocessed image is obtained by rotating, flipping, cropping, scaling, and normalizing the image.
3. The remote sensing target detection method based on serial pyramid convolution according to claim 1, characterized in that: The next step of outputting the feature map of this stage in step S2 is: S21 first adjusts the resolution and channels of the obtained feature map by downsampling, and then changes the number of channels to twice the number of hidden layer channels by 1×1 convolution to perform channel splitting operation; Then, by channel splitting, the feature is split into two shapes [C hid , H out , W out ] features, these two features enter the feedforward neural network and SPCLayer for feature extraction respectively; The S22 feedforward neural network first performs a 1×1 convolution on the feature map, then performs an n×n (n=5, 7, 9, 11) depthwise separable convolution, and finally performs a 1×1 convolution and outputs the feature map; The S23 serial pyramid convolution module achieves the effect of multi-scale convolution by connecting 3×3 convolution kernels in series.
4. The remote sensing target detection method based on serial pyramid convolution according to claim 3 is characterized in that: The effect of implementing multi-scale convolution in step S23 specifically refers to extracting features with receptive fields within the range of 3×3, 7×7, 11×11 and 17×17 through four serial 3×3 dilated convolutions.
5. The remote sensing target detection method based on serial pyramid convolution according to claim 4, characterized in that: In order to expand the receptive field while reducing the number of parameters, these four convolutions are all depth-wise separable convolutions, and the expansion rates are 1, 2, 2, and 3 respectively.