Lightweight remote sensing target detection method based on spatial perception dynamic feature selection
By adopting a spatially perceived dynamic feature selection method based on the RetinaNet framework in remote sensing object detection, using a deep expansion convolutional cascade sequence and an adaptive rotation attention mechanism, the problem of insufficient remote sensing image detection accuracy is solved, and the accuracy improvement and calculation complexity of remote sensing image object detection is achieved, which is suitable for remote sensing hardware platforms.
Patent Information
- Application Number
- CN202510747249.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-25
AI Technical Summary
The existing remote sensing object detection methods are difficult to effectively utilize the prior knowledge of remote sensing images, especially extensive and dynamically changing context information, which leads to insufficient detection accuracy and high computational complexity, making it difficult to effectively apply on remote sensing hardware platforms.
The spatially perceived dynamic feature selection method based on the RetinaNet framework is adopted. By constructing a deep-expanded convolutional cascade sequence with different scales of receptive fields, combining the adaptive rotation attention mechanism and dynamic spatial selection mechanism, the remote sensing image is extracted and fused, and the pyramid hierarchical structure is used for step-by-step feature extraction, and the training process is optimized through the equilibrium loss function.
It improves the accuracy of remote sensing image object detection, achieves a balance of speed and accuracy, adapts to the context information needs of different objects, reduces the computational complexity, and is suitable for the application of remote sensing hardware platforms.
Smart Images

Figure CN120374957A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning, and particularly to a lightweight remote sensing object detection method based on spatial perception dynamic feature selection. Background Art
[0002] Remote sensing object detection is a field of computer vision that aims to automatically identify and locate valuable objects (such as airplanes, ships, bridges, etc.) in aerial images, and has been widely applied in multiple fields such as environmental monitoring, resource exploration, and smart cities. In recent years, with the rapid development of deep learning technologies such as the wide application of Convolutional Neural Network (CNN) and the support of a large number of high-resolution remote sensing image datasets, the remote sensing object detection algorithms based on deep learning have developed rapidly, significantly improving the performance of remote sensing object detection tasks.
[0003] Remote sensing images are usually taken from an aerial view with high resolution, and most of the objects are small in size, making it difficult to accurately identify these objects based solely on appearance information. Instead, the accurate identification of these objects relies more on extensive context information, as the surrounding environment is closely related to their shape, orientation, and other features. In the case of insufficient context information, it is very likely to misdetect small remote sensing objects; in addition, there are significant differences in the requirements for the context information range for different types of objects. For example, some objects may only require local information, while the identification of other objects requires a larger range of context information. This dependence on extensive and dynamically changing context information constitutes unique and important prior knowledge in remote sensing images.
[0004] In recent years, the research on remote sensing object detection based on CNN has mainly focused on the optimization of rotated bounding box representation and the extraction of key features. Although these methods have made significant progress in solving the rotation invariance problem, they often ignore the valuable prior knowledge contained in remote sensing images. The Transformer-based models have shown excellent performance superior to the CNN architecture in many vision tasks, but they also have some limitations, such as higher computational complexity, greater training and inference costs, and the need for a large amount of training data. Due to the limited memory and computational resources of remote sensing hardware platforms, it is difficult to directly apply the Transformer architecture to remote sensing object detection tasks. Some studies have shown that a large receptive field is a key factor for the success of the Transformer architecture. Therefore, the careful design with a large receptive field, such as the application of large-size convolutional kernels, has become an ideal optimization scheme to improve the performance of convolutional neural networks. However, an overly large convolutional kernel size is likely to exceed the optimization range of conventional operators and will significantly increase the computational complexity. In addition, due to the characteristics of remote sensing images, models with different scale receptive fields are very suitable for remote sensing object detection tasks.
[0005] Therefore, how to make full use of the prior knowledge of remote sensing images: it requires extensive and dynamically changing contextual information, design a lightweight remote sensing target detection method, and achieve a significant improvement in detection accuracy, which is an urgent problem to be solved by technicians in this field. Summary of the invention
[0006] The purpose of this application is to provide a lightweight remote sensing target detection method based on spatial perception dynamic feature selection to improve the accuracy of remote sensing image target detection and achieve an effective balance between speed and accuracy.
[0007] In order to solve the above technical problems, the present application provides a lightweight remote sensing target detection method based on spatial perception dynamic feature selection, including: Step 1: Build a dataset; Preprocess the original remote sensing images; construct a data set using the preprocessed remote sensing images; Step 2: construct a network model based on spatial perception dynamic feature selection; use the data of the data set to train the network model based on spatial perception dynamic feature selection to obtain the optimal network model based on spatial perception dynamic feature selection; the input of the network model based on spatial perception dynamic feature selection is the remote sensing image in the data set; the output of the network model based on spatial perception dynamic feature selection is a remote sensing image with a detection frame; the detection frame includes the category and confidence of the image; The network model based on spatial perception dynamic feature selection is based on the RetinaNet framework; the backbone network backbone of the network model based on spatial perception dynamic feature selection decomposes the large kernel convolution to construct a set of deep dilated convolution cascade sequences with different scale receptive fields, and extracts features from the remote sensing image to be tested according to the deep dilated convolution cascade sequence to obtain a multi-scale spatial feature map; the multi-scale spatial feature map includes a first multi-scale spatial feature map, a second multi-scale spatial feature map and a third multi-scale spatial feature map; the fine-grained directional information of the multi-scale spatial feature map is captured through an adaptive rotation attention mechanism to enhance the feature representation capability of objects in any direction; the multi-scale spatial feature map and the corresponding directional information are adaptively weighted fused through a dynamic spatial selection mechanism to obtain feature maps of different scales; the feature maps of different scales include a first feature map F1, a second feature map F2, a third feature map F3 and a fourth feature map F4; the feature maps of different scales are input into a pyramid feature extraction module, and feature extraction is performed step by step on the feature maps of different scales to obtain detection results; Step 3: Use the optimal network model based on spatial perception dynamic feature selection to identify and detect remote sensing images.
[0008] Preferably, the backbone of the network model based on spatial perception dynamic feature selection includes a Stage I module, a Stage II module, a Stage III module, and a Stage IV module connected in sequence; the input of the backbone is the remote sensing image in the dataset; the structures of the Stage I module, the Stage II module, the Stage III module, and the Stage IV module are the same; The input of the Stage I module is the remote sensing image in the dataset; the output of the Stage I module is the first feature map F1; the output of the Stage II module is the second feature map F2; the output of the Stage III module is the third feature map F3; the output of the Stage IV module is the fourth feature map F4; the pyramid feature extraction module of the network model based on spatial perception dynamic feature selection includes a top layer, a middle layer, and a bottom layer connected in sequence; the input of the top layer is the fourth feature map F4; the input of the middle layer includes the top layer feature map output by the top layer and the third feature map F3; the input of the bottom layer includes the middle layer feature map output by the middle layer and the second feature map F2.
[0009] Preferably, the Stage I module includes a first block module and a second block module connected in sequence; the first block module and the second block module have the same structure; the input of the second block module is the second output feature map; the output of the second block module is the first feature map F1; The first block module includes a first batch normalization module, a feature selection sub-block, a second batch normalization module, and a feed-forward neural network sub-block connected in sequence; The input of the first batch normalization module is the remote sensing image in the dataset; the output of the first batch normalization module is the first normalized image; a residual connection is established between the input of the first batch normalization module and the output of the feature selection sub-block to obtain the first output feature map; the input of the second batch normalization module is the first output feature map; a residual connection is established between the input of the second batch normalization module and the output of the feed-forward neural network sub-block to obtain the second output feature map; the output of the second batch normalization module is the second normalized image; The first batch normalization module and the second batch normalization module have the same structure; the first batch normalization module and the second batch normalization module normalize the input image features to the same scale on each channel of the image.
[0010] Preferably, the feature selection sub-block includes a first convolutional layer, a first GELU activation function, a spatial perception dynamic feature selection module, and a second convolutional layer connected in sequence; the input of the feature selection sub-block is the first normalized image; both the first convolutional layer and the second convolutional layer have a 1×1 convolutional kernel.
[0011] Preferably, the spatial perception dynamic feature selection module includes a depth dilated convolution cascade sequence, a dynamic spatial selection mechanism, a first adaptive rotation attention mechanism, a second adaptive rotation attention mechanism, and a third adaptive rotation attention mechanism; The structures of the second adaptive rotation attention mechanism and the third adaptive rotation attention mechanism are the same as the structure of the first adaptive rotation attention mechanism; The output of the first adaptive rotation attention feature map is the first adaptive rotation attention feature map; The output of the second adaptive rotation attention mechanism is the second adaptive rotation attention feature map; The output of the third adaptive rotation attention mechanism is the third adaptive rotation attention feature map; The depth dilated convolution cascade sequence includes a first convolution sequence, a second convolution sequence, a third convolution sequence, a third convolution layer, a fourth convolution layer, and a fifth convolution layer connected in sequence; The first convolution sequence, the second convolution sequence, and the third convolution sequence all have 3×3 convolution kernels; The output of the first convolution sequence is the first convolution sequence feature map U1; the output of the second convolution sequence is the second convolution sequence feature map U2; the output of the third convolution sequence is the third convolution sequence feature map U3; the input of the third convolution layer is the first convolution sequence feature map U1; the input of the fourth convolution layer is the second convolution sequence feature map U2; the input of the fifth convolution layer is the first convolution sequence feature map U3; The output of the third convolution layer is the first multi-scale spatial feature map; the output of the fourth convolution layer is the second multi-scale spatial feature map; the output of the fifth convolution layer is the third multi-scale spatial feature map; The third convolution layer, the fourth convolution layer, and the fifth convolution layer all have 1×1 convolution kernels; The pointwise convolution layer includes a number of convolution kernels; the number of convolution kernels in the pointwise convolution layer is the same as the number of convolution sequences in the depth dilated convolution cascade sequence; The output of the depth dilated convolution cascade sequence includes the first multi-scale spatial feature map, the second multi-scale spatial feature map, and the third multi-scale spatial feature map; The input of the first adaptive rotation attention mechanism is the first multi-scale spatial feature map; the output of the first adaptive rotation attention mechanism is the first adaptive rotation attention feature map; The input of the second adaptive selection attention mechanism is the second multi-scale spatial feature map; the output of the second adaptive selection attention mechanism is the second adaptive rotation attention feature map; The input of the third adaptive selection attention mechanism is the third multi-scale spatial feature map; the output of the third adaptive selection attention mechanism is the third adaptive rotation attention feature map; The dynamic spatial selection mechanism includes a first splicing layer, an average pooling processing layer, a max pooling processing layer, a second splicing layer, a pointwise convolution layer, a Sigmoid activation function, and a selection fusion processing layer; The inputs of the dynamic spatial selection mechanism are a first multi-scale spatial feature map, a second multi-scale spatial feature map, and a third multi-scale spatial feature map; The first splicing layer splices the input features to obtain a spliced feature map; The spliced feature map is input into the average pooling processing layer to obtain a first pooled feature map; the spliced feature map is input into the max pooling processing layer to obtain a second pooled feature map; the first pooled feature map and the second pooled feature map are respectively input into the second splicing layer to obtain a fused feature map; The input of the pointwise convolution layer is the fused feature map; the output of the pointwise convolution layer is a spatial attention feature map; The input of the Sigmoid activation function is the spatial attention feature map; the output of the Sigmoid activation function is a spatial selection mask; The spatial selection mask is multiplied pointwise with the first adaptive rotation attention feature map, the second adaptive rotation attention feature map, and the third adaptive rotation attention feature map respectively, and then the results of the pointwise multiplications are respectively input into the selection fusion processing layer; the selection fusion processing layer adds the input feature maps pointwise and outputs a first target feature map; the first target feature map and the input of the first batch normalization module are processed through a residual connection to obtain a first output feature map.
[0012] Preferably, the feed-forward neural network sub-block includes a sixth convolution layer, a second depth convolution layer, a second GELU activation function, and a seventh convolution layer connected in sequence; both the sixth convolution layer and the seventh convolution layer have a 1×1 convolution kernel; the input of the feed-forward neural network sub-block is a second normalized image; the output of the feed-forward neural network sub-block is a second target feature map; the second target feature map and the input of the second batch normalization module are processed through a residual connection to obtain a second output feature map.
[0013] Preferably, the first adaptive rotation attention mechanism includes a first depth convolution layer, a Gelu & layer normalization layer, a global average pooling layer, a first linear layer & Softsign activation function, a second linear layer & Sigmoid activation function, and several rotation convolution kernels; the first depth convolution layer, the Gelu & layer normalization layer, and the global average pooling layer are connected in sequence; the first depth convolution layer, the Gelu & layer normalization layer, and the global average pooling layer process the input feature map in sequence to obtain a feature vector; the feature vector is input into the first linear layer & Softsign activation function to obtain the predicted i-th rotation angle; the feature vector is input into the second linear layer & Sigmoid activation function to obtain the scaling factor corresponding to the i-th rotation angle; After rotating the i-th rotation convolution kernel by the i-th rotation angle, it is multiplied by the scaling factor corresponding to the i-th rotation angle to obtain a direction feature vector; all the direction feature vectors are respectively convolved with the input of the first adaptive rotation attention mechanism to obtain a first adaptive rotation attention feature map; Preferably, the multi-task loss function of the network model based on spatial perception dynamic feature selection is:
[0014] wherein, is the target recognition function; is the localization target function; The predictions and targets in and are respectively represented as and ; is the regression result corresponding to the th class, is the regression target, is used to adjust the loss weights under multi-task; Using the balanced loss function to replace , that is:
[0015] The balanced loss function is: ; is the deviation amount of the predicted box and the true box of the target in geometric parameters; is a factor for controlling the increase of the normal value gradient. A smaller will cause a smaller gradient to be enhanced more, while the gradient of the outlier is not affected; is a constant for controlling the gradient enhancement magnification, which adjusts the upper limit of the regression gradient; k is a control parameter used to ensure that whenL b =( x = 1) when ; is a constant; When using the training set in the dataset for training, when the network model based on spatial perception dynamic feature selection reaches the highest mean average precision mAP on the validation set, and remains stable without significant decline for multiple consecutive epochs, and at the same time the multi-task loss function tends to be stable without obvious oscillation, that is, the ratio of the standard deviation to the mean of the multi-task loss function value does not exceed 10%, and the average fluctuation amplitude of the multi-task loss function value between adjacent epochs does not exceed 0.005, then the network model at this time is considered the optimal network model based on spatial perception dynamic feature selection.
[0016] Preferably, the preprocessing makes the original remote sensing image have a unified size; the size is 1024×1024.
[0017] A lightweight remote sensing object detection method based on spatial perception dynamic feature selection provided by the present application includes: obtaining a to-be-detected remote sensing image captured by a photographing device; decomposing large kernel convolution to construct a series of depth dilated convolution cascades with different scale receptive fields in a remote sensing object detection network, and extracting features from the to-be-detected remote sensing image according to the depth dilated convolution cascade sequence to obtain a multi-scale spatial feature map; capturing the fine-grained direction information of the multi-scale spatial feature map through an adaptive rotation attention mechanism in the remote sensing object detection network to enhance the feature representation ability for objects in any direction; performing adaptive weighted fusion processing on the multi-scale spatial feature map and the corresponding direction information through a dynamic spatial selection mechanism in the remote sensing object detection network to obtain an object feature map; performing hierarchical feature extraction on the object feature map according to a pyramid hierarchical structure processing mechanism to obtain a detection result. It can be seen that the present application uses a series of depth dilated convolution cascades with different scale receptive fields to extract features from the to-be-detected remote sensing image, thereby realizing feature extraction of the to-be-detected remote sensing image under large receptive fields of different scales and effectively modeling long-range context information in multiple ranges. And the present application further performs weighting and fusion on a series of generated multi-scale spatial feature maps in the spatial dimension through a dynamic spatial selection mechanism to achieve efficient and accurate feature selection, so that the most suitable receptive field can be dynamically selected according to different objects, enhancing the ability of the detection network to focus on its most relevant spatial context area. In addition, by balancing the loss function to balance the gradients generated by different samples, the training of the network is made more stable. The present application effectively improves the accuracy of remote sensing image object detection and realizes the balance between speed and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of a lightweight remote sensing target detection method based on spatial perception dynamic feature selection provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the change of the receptive field provided by an embodiment of the present application; Figure 3 It is a flow structure diagram of an adaptive rotation attention mechanism provided by an embodiment of the present application; Figure 4 It is a flow structure diagram provided by an embodiment of the present application; Figure 5 It is a structure diagram of a lightweight backbone network provided by an embodiment of the present application; Figure 6 It is a schematic diagram of the structure of a block provided by an embodiment of the present application; Figure 7 It is an overall flow structure diagram provided by an embodiment of the present application. Specific embodiments
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0021] The lightweight remote sensing target detection method based on spatial perception dynamic feature selection includes the following steps: Step 1: Construct a dataset; Preprocess the original remote sensing image; use the preprocessed remote sensing image to construct a dataset; The preprocessing makes the original remote sensing image have a unified size; the size is 1024×1024; Step 2: Construct a network model based on spatial perception dynamic feature selection; use the data in the dataset to train the network model based on spatial perception dynamic feature selection to obtain the optimal network model based on spatial perception dynamic feature selection; the input of the network model based on spatial perception dynamic feature selection is the remote sensing image in the dataset; the output of the network model based on spatial perception dynamic feature selection is the remote sensing image with detection frames; the detection frames include the category and confidence of the image; The network model based on spatial perception dynamic feature selection is based on the RetinaNet framework; the backbone network backbone of the network model based on spatial perception dynamic feature selection includes a Stage I module, a Stage II module, a Stage III module and a Stage IV module connected in sequence; the input of the backbone network backbone is a remote sensing image in a data set; the structures of the Stage I module, the Stage II module, the Stage III module and the Stage IV module are the same; The input of the Stage I module is a remote sensing image in a data set; the output of the Stage I module is a first feature map F1; the output of the Stage II module is a second feature map F2; the output of the Stage III module is a third feature map F3; the output of the Stage IV module is a fourth feature map F4; the pyramid feature extraction module of the network model based on spatial perception dynamic feature selection includes sequentially connecting the top layer, the middle layer and the bottom layer; the input of the top layer is the fourth feature map F4; the input of the middle layer includes the top feature map output by the top layer and the third feature map F3; the input of the bottom layer includes the middle layer feature map output by the middle layer and the second feature map F2; The backbone network backbone decomposes the large kernel convolution to construct a set of deep dilated convolution cascade sequences with receptive fields of different scales, and extracts features of the remote sensing image to be tested according to the deep dilated convolution cascade sequence to obtain a multi-scale spatial feature map; the fine-grained directional information of the multi-scale spatial feature map is captured through an adaptive rotational attention mechanism to enhance the feature representation capability of objects in any direction; the multi-scale spatial feature map and the corresponding directional information are adaptively weighted fused through a dynamic spatial selection mechanism to obtain feature maps of different scales; the feature maps of different scales are input into a pyramid feature extraction module, and feature extraction of the feature maps of different scales is performed step by step to obtain detection results.
[0022] The Stage I module includes a first module and a second module connected in sequence; the first module and the second module have the same structure; The first block module includes a first batch normalization module, a feature selection sub-block, a second batch normalization module and a feedforward neural network sub-block connected in sequence; the input of the first batch normalization module is a remote sensing image in a data set; the output of the first batch normalization module is a first normalized image; a residual connection is established between the input of the first batch normalization module and the output of the feature selection sub-block to obtain a first output feature map; the input of the second batch normalization module is the first output feature map; a residual connection is established between the input of the second batch normalization module and the output of the feedforward neural network sub-block to obtain a second output feature map; the output of the second batch normalization module is a second normalized image; The first batch normalization module and the second batch normalization module have the same structure; the first batch normalization module and the second batch normalization module normalize the input image features to the same scale on each channel of the figure; The feature selection sub-block includes a first convolutional layer, a first GELU activation function, a spatial perception dynamic feature selection module, and a second convolutional layer connected in sequence; the input of the feature selection sub-block is the first normalized image; Both the first convolutional layer and the second convolutional layer have a 1×1 convolutional kernel; The spatial perception dynamic feature selection module includes a depth dilated convolution cascade sequence, a dynamic spatial selection mechanism, a first adaptive rotation attention mechanism, a second adaptive rotation attention mechanism, and a third adaptive rotation attention mechanism; The depth dilated convolution cascade sequence includes a first convolutional sequence, a second convolutional sequence, a third convolutional sequence, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer connected in sequence; The output of the depth dilated convolution cascade sequence includes a first multi-scale spatial feature map, a second multi-scale spatial feature map, and a third multi-scale spatial feature map; The input of the first adaptive rotation attention mechanism is the first multi-scale spatial feature map; the output of the first adaptive rotation attention mechanism is the first adaptive rotation attention feature map; The input of the second adaptive selection attention mechanism is the second multi-scale spatial feature map; the output of the second adaptive selection attention mechanism is the second adaptive rotation attention feature map; The input of the third adaptive selection attention mechanism is the third multi-scale spatial feature map; the output of the third adaptive selection attention mechanism is the third adaptive rotation attention feature map; The dynamic spatial selection mechanism includes a first splicing layer, an average pooling processing layer, a max pooling processing layer, a second splicing layer, a pointwise convolutional layer, a Sigmoid activation function, and a selection fusion processing layer; The input of the dynamic spatial selection mechanism is the first multi-scale spatial feature map, the second multi-scale spatial feature map, and the third multi-scale spatial feature map; The first splicing layer splices the input features to obtain a spliced feature map; The spliced feature map is input into the average pooling processing layer to obtain a first pooled feature map; the spliced feature map is input into the max pooling processing layer to obtain a second pooled feature map; the first pooled feature map and the second pooled feature map are respectively input into the second splicing layer to obtain a fused feature map; The input of the pointwise convolutional layer is the fused feature map; the output of the pointwise convolutional layer is a spatial attention feature map; The input of the Sigmoid activation function is the spatial attention feature map; the output of the Sigmoid activation function is the spatial selection mask; The spatial selection mask is multiplied point - by - point with the first adaptive rotation attention feature map, the second adaptive rotation attention feature map, and the third adaptive rotation attention feature map respectively, and then the results of the point - by - point multiplication are input into the selection fusion processing layer respectively; the selection fusion processing layer adds the input feature maps point - by - point and outputs the first target feature map; the first target feature map and the input of the first batch normalization module are processed through residual connection to obtain the first output feature map; The feed - forward neural network sub - block includes a sixth convolutional layer, a second depth convolutional layer, a second GELU activation function, and a seventh convolutional layer connected in sequence; both the sixth convolutional layer and the seventh convolutional layer have 1×1 convolutional kernels; the input of the feed - forward neural network sub - block is the second normalized image; the output of the feed - forward neural network sub - block is the second target feature map; the second target feature map and the input of the second batch normalization module are processed through residual connection to obtain the second output feature map; The input of the second block module is the second output feature map; the output of the second block module is the first feature map F1; The first convolutional sequence, the second convolutional sequence, and the third convolutional sequence all have 3×3 convolutional kernels; The output of the first convolutional sequence is the first convolutional sequence feature map U1; the output of the second convolutional sequence is the second convolutional sequence feature map U2; the output of the third convolutional sequence is the third convolutional sequence feature map U3; the input of the third convolutional layer is the first convolutional sequence feature map U1; the input of the fourth convolutional layer is the second convolutional sequence feature map U2; the input of the fifth convolutional layer is the third convolutional sequence feature map U3; The output of the third convolutional layer is the first multi - scale spatial feature map; the output of the fourth convolutional layer is the second multi - scale spatial feature map; the output of the fifth convolutional layer is the third multi - scale spatial feature map; The third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer all have 1×1 convolutional kernels; The point - wise convolutional layer includes a number of convolutional kernels; the number of convolutional kernels in the point - wise convolutional layer is the same as the number of convolutional sequences in the depth - dilated convolutional cascade sequence; The first adaptive rotation attention mechanism includes a first depth convolution layer, a Gelu & layer normalization layer, a global average pooling layer, a first linear layer & Softsign activation function, a second linear layer & Sigmoid activation function, and several rotation convolution kernels; the first depth convolution layer, the Gelu & layer normalization layer, and the global average pooling layer are connected in sequence; the first depth convolution layer, the Gelu & layer normalization layer, and the global average pooling layer process the input feature map in sequence to obtain a feature vector; the feature vector is input into the first linear layer & Softsign activation function to obtain the predicted i-th rotation angle; the feature vector is input into the second linear layer & Sigmoid activation function to obtain the scaling factor corresponding to the i-th rotation angle; where the range of i is from 1 to Num; Num is the number of all predicted rotation angles; After rotating the i-th rotation convolution kernel by the i-th rotation angle, it is multiplied by the scaling factor corresponding to the i-th rotation angle to obtain a direction feature vector; all the direction feature vectors are respectively subjected to a convolution operation with the input of the first adaptive rotation attention mechanism to obtain a first adaptive rotation attention feature map; The structure of the second adaptive rotation attention mechanism and the structure of the third adaptive rotation attention mechanism are both the same as the structure of the first adaptive rotation attention mechanism; The output of the second adaptive rotation attention mechanism is a second adaptive rotation attention feature map; The output of the third adaptive rotation attention mechanism is a third adaptive rotation attention feature map; Multi-task loss function of the network model based on spatial perception dynamic feature selection is:
[0023] where, is the target recognition function; is the localization target function; The predictions and targets in and ; is the regression result corresponding to the -th class, is the regression target, is used to adjust the loss weights under multi-task; Using the balanced loss function to replace , that is:
[0024] Balanced loss function is: ; The deviation amount of the predicted box and the ground truth box in geometric parameters for the target; is a factor for controlling the increase of the normal value gradient. A smaller will cause a greater enhancement of the small gradient, while the gradient of the outlier is not affected; is a constant for controlling the gradient enhancement magnification, which adjusts the upper limit of the regression gradient; k is a control parameter used to ensure that when L b =( x =1), ; is a constant; When using the training set in the dataset for training, when the network model based on spatial perception dynamic feature selection reaches the highest mean average precision mAP on the validation set, and remains stable without significant decline for multiple consecutive epochs, and at the same time the multi-task loss function tends to be stable without significant oscillation, that is, the ratio of the standard deviation to the mean of the multi-task loss function value does not exceed 10%, and the average fluctuation amplitude of the multi-task loss function value between adjacent epochs does not exceed 0.005, then the network model at this time is considered the optimal network model based on spatial perception dynamic feature selection; Step three: Use the optimal network model based on spatial perception dynamic feature selection to identify and detect remote sensing images; To enable those skilled in the art to better understand the solution of this application, the following further elaborates on this application in conjunction with the accompanying drawings and specific embodiments.
[0025] Figure 1 is a flowchart of a lightweight remote sensing target detection method based on spatial perception dynamic feature selection provided by an embodiment of this application. As Figure 1 shown, it includes the following steps: S10: Obtain the to-be-detected remote sensing image captured by the photographing device.
[0026] S11: Balance the gradients generated by different samples in the to-be-detected remote sensing image through a balanced loss function to construct a remote sensing target detection network with a stable training and detection process.
[0027] S12: Decompose the large kernel convolution to construct a set of depth dilated convolution cascade sequences with different scale receptive fields in the remote sensing target detection network, and perform feature extraction on the to-be-detected remote sensing image according to the depth dilated convolution cascade sequence to obtain a multi-scale spatial feature map.
[0028] In a specific embodiment, the focus of the present application is to determine the final target detection result based on the remote sensing image to be measured. Simply understood, it is to determine the category and location of the target in the current remote sensing image to be measured according to the specific content displayed in the remote sensing image to be measured. Therefore, it is necessary to obtain the remote sensing image to be measured captured by the imaging device, and then input it into the proposed remote sensing target detection network to obtain the final detection result.
[0029] To enhance the network's ability to capture long-range dependencies of spatial features, it uses a cascade sequence of depth dilated convolutions with different scale receptive fields in the remote sensing target detection network to extract features from the remote sensing image to be measured. The specific steps for constructing a cascade sequence of depth dilated convolutions with different scale receptive fields are as follows: decompose the large kernel convolution to obtain a group of depth convolutions (depth-wise convolution) with increasing dilation rates; among them, different scale receptive fields are determined according to the corresponding dilation rates; cascade connect a group of depth convolutions with increasing dilation rates to obtain a cascade sequence of depth dilated convolutions with different scale receptive fields. By cascading these depth convolutions, it is possible to obtain the same effective receptive field as the standard dynamic large kernel convolution, and then realize the extraction of target features under different scale receptive fields, that is, a cascade sequence of depth convolutions with different scale receptive fields in the remote sensing target detection network extracts features from the remote sensing image to be measured to obtain a multi-scale spatial feature map corresponding to the remote sensing image to be measured.
[0030] In a specific embodiment, the process of using dynamic large kernel convolution for spatial feature extraction is summarized as follows: U0 = X; U i+1 = F i dw (U i ); Among them, F i dw ()represents a depth convolution operation with a dilation rate d i X represents the remote sensing image to be measured, and U0 and U i+1 represent multi-scale spatial feature maps obtained by processing with depth dilated convolution kernels with different degrees of receptive fields.
[0031] It is defined as follows: Let F() be a discrete function and k() be a filter function, then the depth convolution operation can be defined as: ; Let d be the dilation rate parameter, k and p be the running calculation parameters, t be the time parameter, and s be the quantity parameter. Generalizing the above formula, the definition of the depth convolution operation with a dilation rate can be obtained: ; Among them, represents the convolution operation with a dilation rate of d.
[0032] Suppose there are N decomposed convolution kernels, and each convolution kernel is further fused in the channel dimension through a 1x1 convolutional layer F 1x1 () for each extracted spatial feature vector: ; In this application, it is set that The cascaded sequence of depth dilated convolutions after decomposing the large kernel convolution of the i specification is composed of a series of depth convolutions with an increasing dilation rate and a size of 3x3. The dilation rate is set to increase step by step in powers of 2. By cascading these depth convolutions, the same effective receptive field as the large kernel convolution can be obtained, and then the extraction of target features can be realized under receptive fields of different scales. The receptive field f f i =(2 i+1 -1)(2 i+1 -1); By continuously increasing the dilation rate and cascading depth convolutions, the receptive field can be expanded fast enough. At the same time, this exponential growth of the dilation rate also ensures that no gaps are introduced between the remote sensing images to be measured. The number N of depth convolutions after decomposing the large kernel convolution of the specification can be determined by the following formula: N+1 K + 1 = 2 ; Figure 2 As shown in Figure 2 the schematic diagram of the change of its receptive field, Figure 2 it shows that the size of the receptive field grows exponentially and the shape is square. After f it can be seen that after the convolution operation with a dilation rate of Dilation = 1, the receptive field f 1: 3×3, and then through the convolution operation with a dilation rate of Dilation = 2, the receptive field f 2: 7×7, and finally after the convolution operation with a dilation rate of Dilation = 4, the receptive field
[0033] 3: 15×15. It can be seen that the dilation rate can effectively expand the size of the receptive field without losing the coverage range. While the receptive field grows exponentially, the number of parameters only grows linearly. After determining the receptive fields of different scales, the receptive fields of different scales are respectively used to extract features from the remote sensing images to be measured.S13: Capture the fine-grained directional information of the multi-scale spatial feature map through the adaptive rotation attention mechanism in the remote sensing object detection network to enhance the feature representation ability for objects in any direction.
[0034] In a specific embodiment, the present application designs an adaptive rotation attention that can rotate the convolutional kernel according to the object direction, thereby adaptively capturing the fine-grained features of objects in different directions and improving the detection ability of the model for objects in any direction.
[0035] S14: Perform adaptive weighted fusion processing on the multi-scale spatial feature map and the corresponding directional information through the dynamic spatial selection mechanism in the remote sensing object detection network to obtain the target feature map.
[0036] In a specific embodiment, the multi-scale spatial feature map is subjected to splicing fusion processing through the dynamic spatial selection mechanism to obtain the fusion feature map; then, the fusion feature map and the directional information are subjected to weighted convolution fusion processing through the dynamic spatial selection mechanism to obtain the target feature map.
[0037] The specific steps for obtaining the fusion feature map by performing splicing fusion processing on the multi-scale spatial feature map through the dynamic spatial selection mechanism are as follows: perform the first splicing fusion on the multi-scale spatial feature map to obtain the spliced feature map; perform channel-based average pooling processing on the spliced feature map to obtain the first pooled feature map; perform channel-based max pooling processing on the spliced feature map to obtain the second pooled feature map; perform the second splicing fusion on the first pooled feature map and the second pooled feature map to obtain the fusion feature map.
[0038] The specific steps for obtaining the target feature map by performing weighted convolution fusion processing on the fusion feature map and the directional information through the dynamic spatial selection mechanism are as follows: perform a pointwise convolution operation on the fusion feature map to obtain the spatial attention feature map; use the Softsign activation function to perform a non-linear mapping on the spatial attention feature map to obtain the corresponding spatial selection mask; perform adaptive weighted fusion processing on the multi-scale spatial feature map, the spatial selection mask, and the directional information to obtain the target feature map.
[0039] Moreover, the specific steps, algorithms, numerical values, etc. involved in the dynamic spatial selection mechanism provided by the present application are all obtained through a large number of test experiments. The specific steps, etc. can improve the accuracy of the detection results, and the specific algorithms, numerical values, etc. can be adjusted according to the needs of users.
[0040] It should also be noted that the number of channels of the multi-scale spatial feature map in the dynamic spatial selection mechanism of the present application is the same as the number of cascaded depth dilated convolutional kernels, which is used for subsequent dynamic spatial feature selection.
[0041] S15: Perform hierarchical feature extraction on the target feature map according to the pyramid hierarchical structure processing mechanism to obtain the detection result.
[0042] In a specific embodiment, the target feature map obtained through the above steps can only be said to have been preliminarily processed in the remote sensing target detection network. The target feature map obtained through the preliminary processing cannot determine the detection result (i.e., the output picture visualization result) corresponding to the to-be-detected remote sensing image, or it can also be understood that the accuracy rate of obtaining the detection result corresponding to the to-be-detected remote sensing image based on the target feature map is reduced. Therefore, it is necessary to further process the target feature map through the remote sensing target detection network, that is, perform hierarchical feature extraction on the target feature map according to the pyramid hierarchical structure processing mechanism in the remote sensing target detection network.
[0043] A lightweight remote sensing target detection method based on spatial perception dynamic feature selection provided by this application includes: obtaining a to-be-detected remote sensing image captured by a photographing device; balancing the gradients generated by different samples in the to-be-detected remote sensing image through a balanced loss function to construct a remote sensing target detection network with a stable training and detection process; decomposing large kernel convolutions to construct a set of depth dilated convolution cascade sequences with different scale receptive fields in the remote sensing target detection network, and performing feature extraction on the to-be-detected remote sensing image according to the depth dilated convolution cascade sequence to obtain a multi-scale spatial feature map; capturing the fine-grained direction information of the multi-scale spatial feature map through the adaptive rotation attention mechanism in the remote sensing target detection network to enhance the feature representation ability for objects in any direction; performing adaptive weighted fusion processing on the multi-scale spatial feature map and the corresponding direction information through the dynamic spatial selection mechanism in the remote sensing target detection network to obtain a target feature map; performing hierarchical feature extraction on the target feature map according to the pyramid hierarchical structure processing mechanism to obtain the detection result. It can be seen that this application uses a set of depth dilated convolution cascade sequences with different scale receptive fields to perform feature extraction on the to-be-detected remote sensing image, thereby realizing feature extraction of the to-be-detected remote sensing image under large receptive fields of different scales and effectively modeling long-range context information in multiple ranges. And this application further performs weighted and fusion on the generated series of multi-scale spatial feature maps in the spatial dimension through the dynamic spatial selection mechanism to achieve efficient and accurate feature selection, so that the most suitable receptive field can be dynamically selected according to different objects, enhancing the ability of the detection network to focus on its most relevant spatial context area. In addition, by balancing the gradients generated by different samples through the balanced loss function, the training of the network is made more stable. This application effectively improves the accuracy of remote sensing image target detection and achieves the balance between speed and accuracy.
[0044] Based on the above embodiments, as a preferred embodiment, the above S13: capturing the fine-grained direction information of the multi-scale spatial feature map through the adaptive rotation attention mechanism in the remote sensing target detection network, includes: grouping and dynamically rotating the convolutional kernels according to the multi-scale spatial feature map, and multiplying each group of rotated convolutional kernels by the corresponding scaling factor to obtain the convolutional kernels representing the importance of feature extraction at that rotation angle; performing convolution operations on the multi-scale spatial feature map with the convolutional kernels at different rotation angles respectively, and splicing the operation results to obtain the fine-grained direction information of the multi-scale spatial feature map.
[0045] However, grouping and dynamically rotating the convolutional kernels according to the multi-scale spatial feature map, and multiplying each group of rotated convolutional kernels by the corresponding scaling factor, includes: sequentially performing depth convolution processing, normalization processing, and global average pooling processing on the multi-scale spatial feature map to obtain a feature vector; predicting different rotation angles of the objects in the multi-scale spatial feature map according to the first linear layer and the Softsign activation function; predicting the scaling factor corresponding to the rotation angle according to the second linear layer and the Sigmoid activation function; grouping the convolutional kernels according to the number of predicted rotation angles, and dynamically rotating the convolutional kernels in groups; multiplying each group of rotated convolutional kernels by the corresponding scaling factor to obtain the convolutional kernels representing the importance of feature extraction at that rotation angle.
[0046] Based on the above embodiments, as a preferred embodiment, the above S14: performing adaptive weighted fusion processing on the multi-scale spatial feature map and the corresponding direction information through the dynamic space selection mechanism in the remote sensing target detection network to obtain the target feature map, includes: performing splicing fusion processing on the multi-scale spatial feature map through the dynamic space selection mechanism to obtain a fusion feature map; performing weighted convolution fusion processing on the fusion feature map and the direction information through the dynamic space selection mechanism to obtain the target feature map.
[0047] Performing splicing fusion processing on the multi-scale spatial feature map through the dynamic space selection mechanism to obtain a fusion feature map, includes: performing the first splicing fusion on the multi-scale spatial feature map to obtain a spliced feature map; performing channel-based average pooling processing on the spliced feature map to obtain the first pooled feature map; performing channel-based maximum pooling processing on the spliced feature map to obtain the second pooled feature map; performing the second splicing fusion on the first pooled feature map and the second pooled feature map to obtain the fusion feature map.
[0048] It performs weighted convolution fusion processing on the fused feature map and the direction information through a dynamic space selection mechanism to obtain the target feature map, including: performing a pointwise convolution operation on the fused feature map to obtain a spatial attention feature map; using the Softsign activation function to perform a non-linear mapping on the spatial attention feature map to obtain the corresponding spatial selection mask; performing adaptive weighted fusion processing on the multi-scale spatial feature map, the spatial selection mask, and the direction information to obtain the target feature map.
[0049] In a specific embodiment, as Figure 3 shown, first, n rotation angles and corresponding scaling factors are predicted based on the current input. Then, the convolutional kernel is divided into n groups along the output channel dimension. The convolutional kernels of each group are dynamically rotated according to the rotation angle and multiplied by the scaling factor to represent the importance of the features in that direction. Finally, the rotated convolutional kernels of each group are respectively convolved with the current input to capture diverse direction information, and the outputs of each group are concatenated to obtain the current output feature . That is, the present application designs a lightweight angle predictor to predict the rotation angles of different objects in the feature map. First, the current input passes through a depth convolutional layer, and then through the GeLU activation function and layer normalization operation to fully extract the spatial direction features of each multi-scale spatial feature map.
[0050] ; Among them LayerNorm () and DWConv () respectively represent the layer normalization operation and the operation of the depth convolutional layer, where is the i th multi-scale spatial feature map, and N is the number of multi-scale spatial feature maps. Then, the spatial direction features of each multi-scale spatial feature map are compressed into a vector of size C through global average pooling. Subsequently, the vector is passed to two independent prediction branches. The first branch is used to generate n predicted angles , which consists of a first linear layer and Softsign activation function, without setting a bias and multiplying the output of Softsign by a rotation coefficient to expand the rotation range. The second branch is used to predict the scaling factor corresponding to the angle. Similar to the first branch, it is predicted using a second linear layer, but this branch adds a bias and uses Sigmoid activation function.
[0051] ; ; Among them, Lineari () represents the fully connected layer of the $i$-th prediction branch, GPool () represents the global average pooling operation. The generated $n$ predicted angles and the corresponding scaling factors are used for subsequent dynamic rotation.
[0052] For its adaptive rotation processing, first, according to the number of predicted angles $n$, preset convolutional kernel groups are divided into $n$ groups along the channel dimension, that is, each group contains convolutional kernels. By adopting the grouping strategy, different preset convolutional kernels work independently in the subsequent dynamic rotation operation. After completing the rotation angle prediction and grouping the convolutional kernels, each kernel in the $j$-th group of convolutional kernels will be dynamically rotated according to the corresponding predicted angle to generate the rotated convolutional kernel , in order to extract diverse directional information. In addition, each group of rotated convolutional kernels is further multiplied by the corresponding predicted scaling factor to represent the importance of the features extracted by the convolutional kernel group with the rotation angle.
[0053] ; Among them, the weights after the dynamic rotation of the convolutional kernel are determined by the bilinear interpolation method. Specifically, for a $k \times k$ convolutional kernel, the weight value at each position can be regarded as a sampling point in the convolutional kernel space. When rotating the convolutional kernel, first, the original value is extended to a 2D convolutional kernel space by the bilinear interpolation method, and then the original convolutional kernel coordinates are rotated angles around the center point to obtain new sampling coordinates, and the convolutional kernel space is resampled again, and finally the rotated convolutional kernel is obtained. Finally, each group of convolutional kernels is respectively convolved with the input to extract the key features of the target from multiple angles, and finally the output results of each group are preset-connected (Concatenate) to obtain the final output feature .
[0054] On the other side, the multi-scale spatial feature maps are first concatenated and fused to obtain the concatenated feature map, and then the concatenated feature map is respectively subjected to channel-based max pooling MaxPool() and average pooling AvgPool() operations to efficiently extract the spatial position relationship. Channel-based max pooling captures the most significant feature at each spatial position by taking the maximum value among all channels at each spatial position of the input feature map; channel-based average pooling takes the average value among all channels at each spatial position of the input feature map, reflecting the global importance of each spatial position.
[0055] In a specific embodiment, first, the multi-scale spatial feature maps are first concatenated: ; wherein, - represents each multi-scale spatial feature map, represents the concatenated feature map.
[0056] Then, by respectively performing channel-based max pooling MaxPool() processing and average pooling AvgPool() processing on the concatenated feature map image the spatial position relationship is efficiently extracted. Channel-based max pooling captures the most significant feature at each spatial position by taking the maximum value among all channels at each spatial position of the concatenated feature map, that is, the second pooled feature image is obtained; channel-based average pooling takes the average value among all channels at each spatial position of the concatenated feature map to obtain the first pooled feature image, which overall reflects the global importance of each spatial position.
[0057] The expression of its second pooled feature map is: ; The expression of its first pooled feature map is: ; wherein, Z avg represents the first pooled feature map; Z max represents the second pooled feature map.
[0058] Among them, in order to achieve information interaction between different spatial descriptions and enable the network to simultaneously utilize local saliency and global information, the two pooled feature maps after spatial pooling are secondarily concatenated and fused to obtain the fused feature map.
[0059] Next, perform a point-by-point convolution operation on the fused feature image group to obtain a group of spatial attention feature maps, that is, convert the pooled feature maps of 2 channels after splicing into N spatial attention feature maps through 1x1 convolution for further feature selection; then use an activation function to perform a non-linear mapping on the group of spatial attention feature maps to obtain the corresponding spatial selection masks, that is, apply the sigmoid activation function on each spatial attention feature map to obtain the spatial selection masks corresponding to each decomposed deep convolution with different scale receptive fields; finally, perform a weighted processing on the multi-scale spatial feature maps and the corresponding spatial selection masks to obtain a group of weighted feature maps, and perform a selection and fusion processing on each weighted feature map and the direction information in the group of weighted feature maps to obtain the target feature map.
[0060] Among them, the expression of the spatial attention feature map is: ; The expression of the spatial selection mask is: ; Among them, the Sig() function represents the sigmoid activation function, and its expression is: ; represents the spatial attention feature map, represents the spatial selection mask.
[0061] The expression of the target feature map is: ; Among them, Y represents the target feature map at this time.
[0062] It should be noted that the embodiments provided in this application are only one implementable way, but are not limited to only this implementable way, and can be set by users according to their needs.
[0063] In summary, the structural diagram of the above process is as Figure 4 shown.
[0064] This application provides specific steps for adaptively weighted fusion processing of multi-scale spatial feature maps according to the dynamic spatial selection mechanism in the remote sensing target detection network to obtain the target feature map. Under this step, a series of generated multi-scale spatial feature maps are weighted and fused in the spatial dimension to achieve efficient and accurate feature selection, so that the most suitable receptive field can be dynamically selected according to needs, enhancing the ability of different remote sensing images to be measured to focus on their most relevant spatial context areas.
[0065] Based on the above embodiments, as a preferred embodiment, before decomposing the large kernel convolution to construct a set of cascaded sequences of depth dilated convolutions with different scale receptive fields in the remote sensing object detection network, it further includes: performing a first-level batch normalization operation on the to-be-detected remote sensing image based on the remote sensing object detection network to obtain the processed to-be-detected remote sensing image; using the feature selection sub-block in the remote sensing object detection network to perform 1×1 convolution processing corresponding to the first preset number of channels and activation function non-linear mapping processing on the processed to-be-detected remote sensing image in sequence to obtain a multi-scale spatial feature map input to the dynamic space selection mechanism.
[0066] After adaptively weighted fusion processing of the multi-scale spatial feature map and the corresponding direction information through the dynamic space selection mechanism in the remote sensing object detection network to obtain the target feature map, it further includes: using the feature selection sub-block in the remote sensing object detection network to perform 1×1 convolution processing corresponding to the second preset number of channels on the multi-scale spatial feature map input to the dynamic space selection mechanism to obtain a first feature map; establishing a residual connection between the input of the first-level batch normalization operation processing in the remote sensing object detection network and the output of the feature selection sub-block to obtain a first output feature map; performing a second-level batch normalization operation on the first output feature map based on the remote sensing object detection network to obtain the processed first output feature map; using the feed-forward neural network sub-block in the remote sensing object detection network to perform 1×1 convolution processing corresponding to the third preset number of channels, depth convolution processing, activation function non-linear mapping processing, and 1×1 convolution processing corresponding to the fourth preset number of channels on the processed first output feature map in sequence to obtain a second feature map; establishing a residual connection between the input of the second-level batch normalization operation processing in the remote sensing object detection network and the output of the feed-forward neural network sub-block to obtain a second output feature map.
[0067] In a specific embodiment, the remote sensing image object detection network provided by the present application is as Figure 5 shown. The network can establish the dependence relationship of multi-scale long-range spatial information, and dynamically adjust the size of the spatial receptive field during the feature extraction process to more accurately and efficiently model the context information of different objects in the remote sensing scene, which can effectively improve the detection accuracy. Especially for small and dense objects in the remote sensing image, it can significantly improve the detection effect and reduce the occurrence of misdetection and missed detection. And the remote sensing image object detection network provided by the present application has a simple and lightweight structure design, with a small number of parameters, low computational complexity, easy to deploy, and is more suitable for remote sensing software and hardware platforms with limited computing resources. Compared with other mainstream detection network models, the present application can achieve an effective balance between accuracy and inference speed. In addition, the network has good adaptability and can be flexibly extended to other tasks, such as image segmentation and classification.
[0068] Its overall remote sensing image target detection network has a clear pyramid hierarchical structure, which consists of four stages (StageⅠ - StageⅣ). The output spatial resolution of each stage gradually decreases, and at the same time, as the output resolution decreases, the number of output channels gradually increases. The number of output channels C1 - C4 of the four stages are 64, 128, 320, and 512 in sequence. The increased number of channels significantly improves the network's expressive ability, enabling it to better capture and represent complex features. The clear pyramid hierarchical structure not only simplifies the model training and optimization process but also promotes the effective fusion of multi-scale features at different stages, thereby enhancing the performance and robustness of the overall network.
[0069] For each stage, first, downsampling is performed on the input, and the stride is used to control the downsampling rate. After downsampling, the output sizes of all other layers in this stage remain unchanged, that is, the number of channels and the spatial resolution remain unchanged. Next, features are extracted by stacking a certain number of the same dynamic feature encoding blocks (Blocks), as Figure 6 shown. Specifically, first, the remote sensing images to be measured in each batch of input are standardized, adjusting their mean to zero and variance to one to eliminate the offset and scale differences between features. Subsequently, linear transformation is performed on the standardized data using learnable scaling coefficients and offsets to ensure that the model still has sufficient feature expression ability after standardization. Its first-level batch normalization operation processing and second-level batch normalization operation processing normalize the input features within a stable range for each channel, thereby improving the stability of the training process, helping to accelerate the convergence speed, and enhancing the generalization ability of the model.
[0070] The feature selection sub-module can simultaneously extract local detail information and capture long-range dependencies, and dynamically adjust the receptive field of the network according to the input to achieve more accurate and efficient feature extraction. This sub-module consists of 1×1Conv (1×1 convolution processing corresponding to the first preset number of channels), activation function (GELU) (nonlinear mapping processing of the activation function), spatial-aware dynamic feature selection module (a cascade sequence of depthwise dilated convolutions with different scale receptive fields extracts features from the remote sensing images to be measured, the adaptive rotation attention mechanism captures the fine-grained direction information of the multi-scale spatial feature map, and the dynamic spatial selection mechanism performs adaptive weighted fusion processing on the multi-scale spatial feature map), and the second 1×1Conv (1×1 convolution processing corresponding to the second preset number of channels). The input features are added to the output of the sub-module through a residual connection, and this design aims to enhance the gradient flow to alleviate the gradient vanishing problem in deep networks.
[0071] After passing through the feature selection sub-module, the feed-forward neural network sub-module is used for cross-channel information interaction and further processing and refinement of features. First, it also needs to be processed by the second-level batch normalization operation. Channel mixing is performed through a 1×1 Conv (1×1 convolution processing corresponding to the third preset number of channels) to represent rich features. Then, a depth convolution is performed to further capture local spatial features. The activation function (GELU) (non-linear mapping processing of the activation function) is used to introduce non-linearity, enhancing the network's ability to more effectively capture complex patterns and relationships in the data. Finally, a second 1×1 Conv (1×1 convolution processing corresponding to the fourth preset number of channels) is used to remap the features back to the original number of channels, ensuring that the output can be consistent with the input feature size. Finally, the input features are added to the output of the sub-module through a residual connection, and the final second output feature map is obtained at this time.
[0072] In combination with Figure 5 the network shown, the target feature map is subjected to hierarchical feature extraction step by step according to the pyramid hierarchical structure processing mechanism to obtain the detection result, including: performing hierarchical feature extraction on the target feature map based on the network blocks ( Figure 5 Conv Block in) corresponding to each stage in the pyramid hierarchical structure to obtain the detection result; wherein, the output of the upper-level network block is used as the input of the lower-level network block.
[0073] In a specific embodiment, Figure 5 each network block in Figure 6 includes the sub-blocks shown, that is to say, the hierarchical feature extraction is performed as many times as there are network blocks. The target feature map output by the last network block is used as the final target feature map, and then the corresponding detection result is determined.
[0074] Among them, the remote sensing target detection network is constructed within the RetinaNet framework.
[0075] It should be noted that in the entire process of the lightweight remote sensing target detection method based on spatial perception dynamic feature selection, the specific model parameters involved are determined by the model weights with the highest detection accuracy after multiple rounds of training. The training data set uses the open source data set DOTA. The DOTA data set is considered to be the largest public data set for remote sensing target detection. It contains a total of 2806 aerial photos from different sensors and platforms. The pixel size of these images ranges from 800×800 to 20,000×20,000. Before model training, the image needs to be preprocessed. Since the image size in the DOTA data set is large, in order to facilitate network training, the original image is first uniformly adjusted to a size of 1024×1024. Subsequently, the data set is randomly divided into a training set, a validation set, and a test set in a ratio of 6:2:2 for subsequent network model training and testing. In the process of continuous model training, the best deep convolution cascade sequence of receptive fields of different scales, dynamic space selection mechanism, and hierarchical structure processing mechanism are obtained. Finally, the remote sensing image to be tested taken by the shooting device is obtained and input into the network loaded with the optimal weights for detection to obtain the final detection result.
[0076] In existing remote sensing detectors, the multi-task loss function is usually defined as follows: ; in, L cls and L loc They are the objective functions of target recognition and positioning, respectively. L cls The prediction and target in are represented as p and t . s k Corresponding to the regression result of the kth class, v is the return target, Used to adjust the loss weights under multi-task. In order to balance the tasks involved, it is usually necessary to adjust the loss weights. However, for regression tasks, since the regression target is unbounded, directly increasing the weight of the localization loss will make the model more sensitive to outliers. We call samples with a loss greater than or equal to 1 outliers, and other samples normal values. These outliers can be regarded as difficult samples that produce excessively large gradients. During training, samples with large gradients tend to receive more attention, while samples with small gradients are not fully trained, which can lead to an imbalance in training samples and affect the performance of subsequent tasks.
[0077] Based on this, this paper proposes a balanced loss function to balance the gradients generated by different samples, making the training of the network more stable. Specifically, it suppresses the contribution of samples that generate large gradients during training due to the large deviation between the anchor box and the ground-truth, because these samples with large gradients are harmful to the training process; at the same time, it promotes the gradients of high-quality samples that are accurate but contribute less to the overall gradient, achieving an effective balance between samples with large gradients and small gradients during the training process.
[0078] The balanced loss function proposed in this application is derived from the traditional smooth L1 loss function, denoted as L b . Specifically, an inflection point is set to distinguish outliers from other samples and suppress the large gradients generated by outliers, limiting the maximum value of its gradient to 1.0. The localization loss function based on balanced loss L loc is defined as: ; The corresponding gradient satisfies the following formula: ; Based on the above formula, we designed a formula for balancing gradients as follows: ; Among them, as a factor to control the increase of the gradient of normal values, a smaller will result in greater enhancement of small gradients, while the gradients of outliers are not affected. is a constant used to control the magnification of gradient enhancement, which adjusts the upper limit of the regression gradient. k is used to ensure that when L b =( x =1), . Through such operations, the contribution of small-gradient samples to training can be effectively improved, and the harmful effects of large gradients generated by outliers on training can be suppressed, thus achieving a more balanced training between precise classification and accurate positioning.
[0079] By performing integral operations on the above formula, the balanced loss function can be obtained as:
[0080] According to the above formula, balance the gradients generated by different samples in the remote sensing image to be measured, and then, according to the difference between the prediction result and the true label, use the balanced gradients for backpropagation to iteratively optimize the model parameters in order to construct a remote sensing object detection network.
[0081] Therefore, as described above, the overall structural flowchart of its remote sensing target detection method is as Figure 7 shown.
[0082] It can be seen that this application uses a cascaded sequence of depth dilated convolutions with different scale receptive fields to extract features from the remote sensing image to be measured, thereby realizing the feature extraction of the remote sensing image to be measured under large receptive fields of different scales, and effectively modeling long-range context information in multiple ranges. Furthermore, this application further performs weighting and fusion on a series of generated multi-scale spatial feature maps in the spatial dimension through a dynamic spatial selection mechanism to achieve efficient and accurate feature selection, enabling the most suitable receptive field to be dynamically selected according to different objects, and enhancing the ability of the detection network to focus on its most relevant spatial context area. In addition, by balancing the loss function to balance the gradients generated by different samples, the training of the network is made more stable. This application effectively improves the accuracy of remote sensing image target detection and achieves a balance between speed and accuracy.
[0083] The above has introduced in detail a lightweight remote sensing target detection method based on spatial perception dynamic feature selection provided by this application. The various embodiments in the specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0084] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
Claims
1. A lightweight remote sensing target detection method based on spatially aware dynamic feature selection, characterized in that, It includes the following steps: Step 1: Construct a dataset; Preprocess the original remote sensing image; use the preprocessed remote sensing image to construct a dataset; Step 2: Construct a network model based on spatial perception dynamic feature selection; use the data in the dataset to train the network model based on spatial perception dynamic feature selection to obtain the optimal network model based on spatial perception dynamic feature selection; the input of the network model based on spatial perception dynamic feature selection is the remote sensing image in the dataset; The output of the network model based on spatial perception dynamic feature selection is a remote sensing image with detection boxes; the detection boxes include the category and confidence of the image; The network model based on spatial perception dynamic feature selection is based on the RetinaNet framework; the backbone of the network model based on spatial perception dynamic feature selection decomposes the large kernel convolution to construct a set of depth dilated convolution cascade sequences with different scale receptive fields, and performs feature extraction on the to-be-detected remote sensing image according to the depth dilated convolution cascade sequence to obtain multi-scale spatial feature maps; the multi-scale spatial feature maps include a first multi-scale spatial feature map, a second multi-scale spatial feature map, and a third multi-scale spatial feature map; capture the fine-grained direction information of the multi-scale spatial feature maps through an adaptive rotation attention mechanism to enhance the feature representation ability for objects in any direction; perform adaptive weighted fusion processing on the multi-scale spatial feature maps and the corresponding direction information through a dynamic spatial selection mechanism to obtain feature maps of different scales; the feature maps of different scales include a first feature map F1, a second feature map F2, a third feature map F3, and a fourth feature map F4; input the feature maps of different scales into a pyramid feature extraction module to perform hierarchical feature extraction on the feature maps of different scales to obtain the detection result; Step 3: Use the optimal network model based on spatial perception dynamic feature selection to identify and detect remote sensing images.
2. The lightweight remote sensing target detection method based on spatial perception dynamic feature selection according to claim 1, wherein The backbone of the network model based on spatial perception dynamic feature selection includes a Stage I module, a Stage II module, a Stage III module, and a Stage IV module connected in sequence; the input of the backbone is the remote sensing image in the dataset; the structures of the Stage I module, the Stage II module, the Stage III module, and the Stage IV module are the same; The input of the Stage I module is the remote sensing image in the dataset; the output of the Stage I module is the first feature map F1; the output of the Stage II module is the second feature map F2; the output of the Stage III module is the third feature map F3; the output of the Stage IV module is the fourth feature map F4; the pyramid feature extraction module of the network model based on spatial perception dynamic feature selection includes a top layer, a middle layer, and a bottom layer connected in sequence; the input of the top layer is the fourth feature map F4; the input of the middle layer includes the top layer feature map output by the top layer and the third feature map F3; the input of the bottom layer includes the middle layer feature map output by the middle layer and the second feature map F2.
3. The lightweight remote sensing target detection method based on spatially-aware dynamic feature selection according to claim 2, wherein The Stage I module includes a first block module and a second block module connected in sequence; the first block module and the second block module have the same structure; the input of the second block module is the second output feature map; the output of the second block module is the first feature map F1; The first block module includes a first batch normalization module, a feature selection sub-block, a second batch normalization module, and a feedforward neural network sub-block connected in sequence; The input of the first batch normalization module is the remote sensing image in the dataset; the output of the first batch normalization module is the first normalized image; a residual connection is established between the input of the first batch normalization module and the output of the feature selection sub-block to obtain the first output feature map; The input of the second batch normalization module is the first output feature map; A residual connection is established between the input of the second batch normalization module and the output of the feedforward neural network sub-block to obtain the second output feature map; the output of the second batch normalization module is the second normalized image; The first batch normalization module and the second batch normalization module have the same structure; the first batch normalization module and the second batch normalization module normalize the input image features to the same scale on each channel of the image.
4. The lightweight remote sensing target detection method based on spatially-aware dynamic feature selection according to claim 3, wherein, The feature selection sub-block includes a first convolutional layer, a first GELU activation function, a spatial perception dynamic feature selection module, and a second convolutional layer connected in sequence; the input of the feature selection sub-block is the first normalized image; both the first convolutional layer and the second convolutional layer have a 1×1 convolutional kernel.
5. The lightweight remote sensing target detection method based on spatially-aware dynamic feature selection according to claim 4, wherein The spatial perception dynamic feature selection module includes a depthwise dilated convolution cascade sequence, a dynamic spatial selection mechanism, a first adaptive rotation attention mechanism, a second adaptive rotation attention mechanism, and a third adaptive rotation attention mechanism; The structures of the second adaptive rotation attention mechanism and the third adaptive rotation attention mechanism are the same as the structure of the first adaptive rotation attention mechanism; The output of the first adaptive rotation attention feature map is the first adaptive rotation attention feature map; The output of the second adaptive rotation attention mechanism is the second adaptive rotation attention feature map; The output of the third adaptive rotation attention mechanism is the third adaptive rotation attention feature map; The depthwise dilated convolution cascade sequence includes a first convolutional sequence, a second convolutional sequence, a third convolutional sequence, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer connected in sequence; The first convolutional sequence, the second convolutional sequence, and the third convolutional sequence all have a 3×3 convolutional kernel; The output of the first convolutional sequence is the first convolutional sequence feature map U1; the output of the second convolutional sequence is the second convolutional sequence feature map U2; the output of the third convolutional sequence is the third convolutional sequence feature map U3; the input of the third convolutional layer is the first convolutional sequence feature map U1; the input of the fourth convolutional layer is the second convolutional sequence feature map U2; the input of the fifth convolutional layer is the first convolutional sequence feature map U3; The output of the third convolutional layer is the first multi-scale spatial feature map; the output of the fourth convolutional layer is the second multi-scale spatial feature map; the output of the fifth convolutional layer is the third multi-scale spatial feature map; The third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer all have a 1×1 convolutional kernel; The pointwise convolutional layer includes a number of convolutional kernels; the number of convolutional kernels in the pointwise convolutional layer is the same as the number of convolutional sequences in the depth dilated convolutional cascade sequence; The output of the depth dilated convolutional cascade sequence includes the first multi-scale spatial feature map, the second multi-scale spatial feature map, and the third multi-scale spatial feature map; The input of the first adaptive rotation attention mechanism is the first multi-scale spatial feature map; the output of the first adaptive rotation attention mechanism is the first adaptive rotation attention feature map; The input of the second adaptive selection attention mechanism is the second multi-scale spatial feature map; The output of the second adaptive selection attention mechanism is the second adaptive rotation attention feature map; The input of the third adaptive selection attention mechanism is the third multi-scale spatial feature map; the output of the third adaptive selection attention mechanism is the third adaptive rotation attention feature map; The dynamic spatial selection mechanism includes a first concatenation layer, an average pooling processing layer, a max pooling processing layer, a second concatenation layer, a pointwise convolutional layer, a Sigmoid activation function, and a selection fusion processing layer; The input of the dynamic spatial selection mechanism is the first multi-scale spatial feature map, the second multi-scale spatial feature map, and the third multi-scale spatial feature map; The first concatenation layer concatenates the input features to obtain a concatenated feature map; The concatenated feature map is input into the average pooling processing layer to obtain a first pooled feature map; the concatenated feature map is input into the max pooling processing layer to obtain a second pooled feature map; the first pooled feature map and the second pooled feature map are respectively input into the second concatenation layer to obtain a fused feature map; The input of the pointwise convolutional layer is the fused feature map; the output of the pointwise convolutional layer is the spatial attention feature map; The input of the activation function Sigmoid is the spatial attention feature map; The output of the activation function Sigmoid is the spatial selection mask; The spatial selection mask is multiplied pointwise with the first adaptive rotation attention feature map, the second adaptive rotation attention feature map, and the third adaptive rotation attention feature map respectively, and the results of the pointwise multiplications are respectively input into the selection fusion processing layer; the selection fusion processing layer adds the input feature maps pointwise and outputs a first target feature map; the first target feature map and the input of the first batch normalization module are processed through residual connection to obtain a first output feature map.
6. The lightweight remote sensing target detection method based on spatially-aware dynamic feature selection according to claim 3, wherein The feed-forward neural network sub-block includes a sixth convolutional layer, a second depth convolutional layer, a second GELU activation function, and a seventh convolutional layer connected in sequence; both the sixth convolutional layer and the seventh convolutional layer have 1×1 convolutional kernels; the input of the feed-forward neural network sub-block is the second normalized image; the output of the feed-forward neural network sub-block is the second target feature map; the second target feature map and the input of the second batch normalization module are processed through residual connection to obtain the second output feature map.
7. The lightweight remote sensing target detection method based on spatially-aware dynamic feature selection according to claim 5, characterized in that, The first adaptive rotation attention mechanism includes a first depth convolutional layer, a Gelu&layer normalization layer, a global average pooling layer, a first linear layer&Softsign activation function, a second linear layer&Sigmoid activation function, and several rotation convolutional kernels; the first depth convolutional layer, the Gelu&layer normalization layer, and the global average pooling layer are connected in sequence; the first depth convolutional layer, the Gelu&layer normalization layer, and the global average pooling layer process the input feature map in sequence to obtain a feature vector; The The feature vector is input into the first linear layer&Softsign activation function to obtain the predicted i-th rotation angle; the feature vector is input into the second linear layer&Sigmoid activation function to obtain the scaling factor corresponding to the i-th rotation angle; where the range of i is from 1 to Num; Num is the number of all predicted rotation angles. After rotating the i-th rotation convolutional kernel by the i-th rotation angle and then multiplying it by the scaling factor corresponding to the i-th rotation angle, a direction feature vector is obtained; all the direction feature vectors are respectively subjected to convolution operations with the input of the first adaptive rotation attention mechanism to obtain the first adaptive rotation attention feature map.
8. The lightweight remote sensing target detection method based on spatial perception dynamic feature selection according to claim 1, characterized in that Multi-task Loss Function of Network Model Based on Spatial Perception Dynamic Feature Selection is as follows: Among them, is the target recognition function; is the function for locating the target; The prediction and target in and are respectively represented as; is the regression result corresponding to the th class, is the regression target, is used to adjust the loss weight under multi-task; Using a balanced loss function Instead of , that is: Balanced loss function is as follows: ; The deviation of the predicted box from the ground truth box in geometric parameters; is a factor that controls the increase of the normal value gradient. A smaller will result in a greater enhancement of small gradients while the gradients of outliers are not affected; is a constant that controls the gradient enhancement magnification, which regulates the upper limit of the regression gradient; k is a control parameter used to ensure that when L b =( x = 1), ; is a constant; When training using the training set in the dataset, when the network model based on spatial-aware dynamic feature selection reaches the highest mean average precision mAP on the validation set, and remains stable without significant decline for multiple consecutive epochs, and at the same time the multi-task loss function tends to be stable without obvious oscillation, that is, the ratio of the standard deviation to the mean of the multi-task loss function value does not exceed 10%, and the average fluctuation amplitude of the multi-task loss function value between adjacent epochs does not exceed 0.005, then the network model at this time is considered to be the optimal network model based on spatial-aware dynamic feature selection.
9. The lightweight remote sensing target detection method based on spatial perception dynamic feature selection according to claim 1, wherein The preprocessing enables the original remote sensing image to have a unified size; the size is 1024×1024.
10. A computer-readable storage medium storing a computer program therein; characterized in that, When the computer program is executed by a processor, it implements the lightweight remote sensing object detection method based on spatial-aware dynamic feature selection according to any one of claims 1-9.