Remote sensing image rotation high length-width ratio small target detection method and system based on deep learning
By embedding the dynamic refinement rotation convolution module and the anchor point refinement feature alignment module in the remote sensing image object detection model, the problem of the traditional method degradation of accuracy when detecting the rotation high aspect ratio small target is achieved, and higher detection accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510216076.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-17
AI Technical Summary
When traditional remote sensing image object detection methods face rotating or tilted high aspect ratio small targets, the detection accuracy is significantly reduced, mainly due to the problems of feature misalignment and sample imbalance.
By embedding the dynamic refinement rotational convolution module (DRRCM) and anchor point refinement feature alignment module (ARFAM) in the backbone network of the object detection model, the convolution kernel and anchor box are adaptively adjusted to generate variable features such as direction sensitivity feature maps and high-quality rotation.
The detection effect of rotating high aspect ratio small targets in remote sensing images is significantly improved, and the detection accuracy and stability of the model are improved.
Smart Images

Figure CN120163967A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing images, and relates to a method and system for detecting small targets with high aspect ratio and rotation in remote sensing images based on deep learning. Background Technique
[0002] With the continuous development of satellite and aerial photography technologies, the application fields of remote sensing images have been expanding day by day. High-precision remote sensing image analysis is indispensable in many aspects, from environmental monitoring to urban planning and then to military reconnaissance. However, in these applications, accurately identifying and locating targets in remote sensing images, especially those with small sizes, various orientations, and dense arrangements, has become a very challenging task.
[0003] Traditional object detection methods, especially general object detection frameworks based on horizontal anchor boxes, show poor adaptability when facing the above characteristics. These problems are mainly reflected in: 1) It is difficult for horizontal anchor boxes to accurately match rotated or tilted targets; 2) Background information is easily mis-incorporated into candidate regions; 3) A single anchor box may cover multiple targets simultaneously, resulting in a significant decline in detection performance.
[0004] In recent years, the rapid development of deep learning technologies has provided new ideas and tools for solving this problem. By leveraging the powerful feature extraction ability of convolutional neural networks (CNNs), the unique features of targets in complex backgrounds of remote sensing images can be effectively captured. In particular, research on detecting rotated small targets is exploring how to design algorithms that are more adaptable to such specific conditions, including but not limited to developing methods that can adaptively adjust the angle of anchor boxes, optimizing loss functions to reduce the impact of background interference, and adopting multi-scale feature fusion strategies to enhance the recognition ability of small targets, etc.
[0005] Existing methods have good effects when detecting regularly shaped targets, but when facing oriented small targets with large aspect ratio differences, the detection accuracy drops significantly. The main important reasons are as follows:
[0006] (1) Feature misalignment: The convolutional features of traditional backbone networks are usually aligned based on a fixed receptive field direction, which makes it difficult for them to adapt to oriented targets with large aspect ratio differences, thereby resulting in poor feature extraction effects. Even if convolutional alignment or anchor box alignment operations are introduced in subsequent steps, the loss of local information at the target edges caused by the initial fixed convolution method cannot be compensated, thus affecting the overall feature extraction quality. Because all subsequent operations, such as feature fusion, resampling, etc., are based on the feature maps extracted by the initial backbone network. Therefore, in the object detection framework, the backbone network that initially extracts target features is crucial for improving the model accuracy.
[0007] (2) Sample imbalance: During the anchor box regression process, the oriented targets with high aspect ratios are extremely sensitive to the regression of angles. Even a tiny angular deviation may lead to a significant increase in the deviation between the predicted box and the ground truth box. This situation becomes more obvious as the value of the shape information becomes smaller. This will result in the misjudgment of targets that should originally be positive samples as negative samples during the sample selection process. Even if the classification score is high, due to poor regression performance, the overlap area (IoU) between the predicted box and the ground truth box is lower than the preset threshold. This misjudgment will cause imbalance between positive and negative samples, thus having a negative impact on the overall detection effect. Summary of the Invention
[0008] The purpose of the present invention is to provide a method and system for detecting small targets with high aspect ratios and rotations in remote sensing images based on deep learning, so as to improve the detection effect of small targets with high aspect ratios and rotations in remote sensing images.
[0009] To achieve the above purpose, the basic solution of the present invention is: A method for detecting small targets with high aspect ratios and rotations in remote sensing images based on deep learning, comprising the following steps:
[0010] S1, embed the Dynamic Refinement Rotation Convolution Module (DRRCM) into the backbone network ResNet50 of the object detection model, replacing some of its 3×3 convolution modules. DRRCM uses the Data Enhancement Spatial Attention Module (DESAM) to predict the weights and angles of the rotation convolution kernels;
[0011] Combine the predicted weights and angle parameters of the rotation convolution kernels, and use the Dynamic Refinement Rotation Convolution Module (DRRCM) to adaptively adjust the convolution kernels according to the pose information of the oriented targets, generating a direction-sensitive feature map to accurately align it with the target features;
[0012] S2, use the Anchor Refinement Feature Alignment Module (ARFAM) to generate corrected predicted anchor boxes on the direction-sensitive feature map based on the regression branch, serving as a guide to dynamically adjust the positions of the feature sampling points;
[0013] S3, based on the high-quality rotation-equivariant features finally generated by ARFAM and the optimized predicted anchor boxes, use the IoU threshold method to train the object detection model;
[0014] S4, collect the remote sensing images to be detected and input them into the optimized object detection model to obtain the object detection results.
[0015] The working principle and beneficial effects of this basic solution are as follows: In this technical solution, DRRCM is embedded into the backbone network ResNet50, replacing some of the 3×3 convolutional layers, which can effectively extract the preliminary direction perception features of small targets with high aspect ratio and rotation in remote sensing images. DRRCM not only enhances the network's direction perception ability for rotated targets but also improves the quality of feature maps, laying a solid foundation for subsequent processing.
[0016] The high-quality, direction-aware feature maps generated by DRRCM are transmitted to ARFAM in the Head. Here, the regression branch is used to optimize the predicted anchor boxes to ensure that they can more accurately match the position and orientation of the actual targets. ARFAM further adopts the aligned convolution technique, guiding the sampling points through the optimized predicted anchor boxes to achieve more precise feature alignment.
[0017] This dual-aligned convolution mechanism, which combines the advantages of DRRCM and ARFAM, works together to extract high-quality rotation-equivariant features. Finally, the generated high-quality feature representations and the corrected predicted anchor boxes greatly improve the stability and accuracy of model training, significantly enhancing the detection effect for small targets with high aspect ratio and rotation in remote sensing images.
[0018] Furthermore, the method for DRRCM to use the data augmentation spatial attention module DESAM to predict the weights and angles of rotation convolution kernels is as follows:
[0019] Perform channel-based average pooling and max pooling on the feature map F of the original image after deep convolution, denoted as P avg (.), and P max (.), and extract the spatial relationship:
[0020] S avg =P avg (F), S max =P max (F)
[0021] where S avg and S max are the spatial feature descriptors after average pooling and max pooling. To allow information interaction between different spatial descriptors, the spatially pooled features are concatenated and a convolutional layer is used to convert the pooled features into C in spatial attention maps S':
[0022]
[0023] where C indenotes the number of input channels; for each spatial attention map S′, a sigmoid activation function is applied to obtain the individual spatial mask S of each convolutional kernel i ′:
[0024] S i ′ = sigmoid(S′),
[0025] The features after depth convolution are weighted using the spatial mask and compressed into a C in -dimensional feature vector V Cin :
[0026] V Cin = P avg (S i ′ × F),
[0027] The pooled feature vector is fed into two branches respectively. The first is the rotation kernel angle prediction branch. The feature vector is input into this branch and passes through Dropout, a linear layer, Softsign activation, and multiplication by a scale factor to obtain a set of angles θ i :
[0028] θ i = K(Softsign(z θ ))
[0029] where z θ has no bias set in the linear layer to ensure that the angle prediction only depends on the changes in the input features and avoid learning offset angles; is the scale factor to expand the rotation range, and the proportion parameter is used to adjust the range of the angle;
[0030] The second is the rotation kernel weight prediction branch. The feature vector is input into this branch and passes through Dropout, a linear layer, and Sigmoid activation to obtain a set of weights λ i :
[0031] λ i = sigmoid(z a )
[0032] where z α has a bias set in the linear layer to improve the flexibility of the model.
[0033] The weighted and fused features are average-pooled and input into the kernel angle prediction branch and the kernel weight prediction branch. This module can enable the network to more precisely focus on the key feature positions in rotation target detection, thus accurately generating the weights and angles of the predicted rotation kernel.
[0034] Furthermore, the rotation angle θ generated by DESAM predictioni The parameter reparameterizes the weights inside the convolutional kernel, enabling the convolutional kernel W i to be dynamically adjusted according to different input feature maps to achieve adaptive rotation:
[0035] Y i ′ = rotate(Y i , -θ i )
[0036] W i ′ = interpolation(W i , Y i ′)
[0037] where Y i represents the coordinates of the original sampling points, and Y i ′ represents the new sampling point coordinates after rotating the original sampling point Y i counterclockwise by an angle θ i to achieve alignment between the convolutional kernel and the features; rotate represents the rotation operation; W i ′ represents the reparameterized convolutional kernel; interpolation(.) represents bilinear interpolation used to calculate the weights at the new position after the rotation of the W i convolutional kernel;
[0038] Multiply the reparameterized convolutional kernel by the corresponding λ i weights and sum them, then perform a convolution operation with the input feature map to finally generate high-quality orientation-aware features Y:
[0039]
[0040] where n represents the number of rotated convolutional kernels.
[0041] Obtaining high-quality features is beneficial for subsequent use.
[0042] Furthermore, through the Anchor Refinement Feature Alignment Module ARFAM, based on the regression branch, quickly generate corrected predicted anchor boxes on the direction-sensitive feature map as a guide to dynamically adjust the positions of the feature sampling points. The specific steps are as follows:
[0043] The features output by the DRRCM are fused through FPN to obtain a feature map, and only one initial square anchor point is preset at each position of this feature map. It is corrected into a high-quality directional anchor point through the regression branch, predicting the offset of the regression target of the anchor box:
[0044]
[0045] Among them, (x, y, w, h, θ) represent the center coordinates, width, height, and angle parameters of the initial anchor point; (x g , y g , w g , h g , θ g ) represent the center coordinates, width, height, and angle parameters of the true bounding box; (Δx g , Δy g , Δw g , Δh g , Δθ g ) represent the offsets between the true bounding box and the initial anchor point. By regressing these offsets, the model can adjust the initial anchor point to a corrected predicted anchor box closer to the true bounding box; R(θ) represents the rotation transformation matrix used to convert the center point coordinates of the true bounding box into the coordinate system relative to the initial anchor point; k represents the scale coefficient used to adjust the angle value to ensure that the rotation angle is within a reasonable range;
[0046] By using the corrected predicted anchor box as a guide to adjust the feature sampling point positions to achieve dynamic convolution alignment. Based on the original sampling points of the standard convolution, an offset o deduced from the corrected predicted anchor box is added:
[0047]
[0048] Among them, represents the position of the predicted anchor box sampling point; (p0 + p n ) represents the position of the regular sampling point of the standard convolution; p0 and p n represent the two-dimensional coordinates and relative offset of this sampling point respectively, and R represents the regular grid [(p x , p y )] of the standard convolution;
[0049] The dynamic alignment convolution combines the offset o and the input feature X, enabling the sampling point positions to be adjusted according to the shape and direction of the predicted anchor box to match the geometric characteristics of the actual target:
[0050]
[0051] Among them, Y(p) represents the value of the output feature map at position p, w(p n ) is the weight of the convolution kernel at position p n , and x(·) represents the input feature map after feature fusion.
[0052] Based on the high-quality feature maps generated by DRRCM in the backbone network, ARFAM is used to further refine the anchor points through regression operations, and the offset field is calculated according to the refined anchor point parameters to dynamically adjust the positions of the aligned convolutional sampling points, generating a feature representation that is more precisely aligned with the target object.
[0053] Furthermore, based on the high-quality rotation-equivariant features generated by DRRCM and ARFAM, as well as the optimized predicted anchor boxes, a target detection model is trained, specifically as follows:
[0054] The generated high-quality feature representation and the corrected predicted anchor boxes are used for model training. Positive and negative samples are selected for training by a fixed IoU threshold between the corrected predicted anchor boxes and the true ground truth bounding boxes:
[0055]
[0056] Where A∩B represents the intersection of the predicted anchor box and the true ground truth bounding box, the overlapping area; A∪B represents the union of the predicted anchor box and the true ground truth bounding box, the non-overlapping area.
[0057] The generated high-quality feature representation and the corrected predicted anchor boxes greatly improve the stability and accuracy of model training, and significantly improve the detection effect on small targets with high aspect ratios in remote sensing images.
[0058] The present invention also provides a remote sensing image small target detection system with high aspect ratio rotation based on deep learning, including a processing unit, the processing unit includes a dynamic refinement rotation convolution module DRRCM and an anchor point refinement feature alignment module ARFAM, and the processing unit executes the method of the present invention to obtain the target detection result of the remote sensing image.
[0059] This system combines the advantages of DRRCM and ARFAM - working together to extract high-quality rotation-equivariant features with high detection accuracy.
[0060] Furthermore, the dynamic refinement rotation convolution module DRRCM includes a data augmentation spatial attention module DESAM and a convolutional kernel that can adaptively adjust according to the pose information of the oriented target;
[0061] The data augmentation spatial attention module DESAM includes a depth convolutional layer, a max pooling layer, a first average pooling layer, a concatenation layer, a convolutional layer, a second average pooling layer, and two fully connected layers;
[0062] The output end of the depth convolutional layer is respectively connected to the max pooling layer and the first average pooling layer, the output ends of the max pooling layer and the first average pooling layer are connected to the concatenation layer, and the concatenation layer, the convolutional layer, the second average pooling layer, and the two fully connected layers are connected in sequence.
[0063] The Dynamic Refinement Rotation Convolution Module (DRRCM) has a simple structure and is easy to use.
[0064] Furthermore, the Anchor Refinement Feature Alignment Module (ARFAM) includes a regression branch layer and an alignment convolution layer connected in sequence.
[0065] The Anchor Refinement Feature Alignment Module (ARFAM) uses alignment convolution technology to guide sampling points through optimized predicted anchor boxes, achieving more accurate feature alignment. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a schematic flowchart of the method for detecting rotated small targets with high aspect ratios in remote sensing images based on deep learning according to the present invention;
[0067] Figure 2 is a schematic structural diagram of the DRRCM module of the system for detecting rotated small targets with high aspect ratios in remote sensing images based on deep learning according to the present invention;
[0068] Figure 3 is a schematic structural diagram of the Anchor Refinement Feature Alignment Module (ARFAM) of the system for detecting rotated small targets with high aspect ratios in remote sensing images based on deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0070] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0071] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.
[0072] The present invention discloses a remote sensing image rotation high aspect ratio small target detection method based on deep learning (Rotation-Aware Network for Detecting Small Targets with High Aspect Ratios, RAN-HAST), which can generate high-quality rotation-equivariant features and regression prediction anchor boxes, and has good detection effects for detecting rotated and densely distributed small targets. As Figure 1 shown, the remote sensing image rotation high aspect ratio small target detection method based on deep learning includes the following steps:
[0073] S1, alleviating the problem of feature misalignment in the backbone network: embedding the Dynamic Refinement Rotation-aware Convolution Module (DRRCM) into the ResNet50 backbone network of the target detection model, replacing some of its 3×3 convolution modules, and using the Data-Enhanced Spatial Attention Module (DESAM) in DRRCM to predict the weights and angles of the rotation convolution kernels; it can effectively extract the preliminary direction perception features of rotated high aspect ratio small targets in remote sensing images. DRRCM not only enhances the network's direction perception ability for rotated targets but also improves the quality of the feature map, laying a solid foundation for subsequent processing.
[0074] Combining the predicted rotation convolution kernel weights and angle parameters, using the dynamic refinement rotation convolution module DRRCM to adaptively adjust the convolution kernel according to the pose information of the oriented target, generating a direction-sensitive feature map, and accurately aligning it with the target features;
[0075] S2, the high-quality and direction-aware feature map generated by DRRCM is transmitted to the ARFAM in the Head. Using the Anchor Refinement Feature Alignment Module (ARFAM), based on the regression branch, corrected prediction anchor boxes are generated on the direction-sensitive feature map, which are used to guide the dynamic adjustment of the positions of the feature sampling points to further achieve precise feature alignment. Using the regression branch to optimize the prediction anchor boxes to ensure that they can more accurately match the position and orientation of the actual target. ARFAM further adopts the alignment convolution technology, guiding the sampling points through the optimized prediction anchor boxes, and achieving more precise feature alignment.
[0076] S3. Based on the high-quality rotation-equivariant features finally generated by ARFAM (the preliminary direction-sensitive feature map generated after DRRCM undergoes FPN feature fusion, and then the regression branch of ARFAM corrects the anchor points, calculates the offset, and adjusts the alignment convolution according to the offset to further extract feature alignment. It is the final rotation-equivariant feature generated by the combined action of two convolution modules), and the optimized predicted anchor boxes, use the IoU threshold method to train the object detection model;
[0077] The dual alignment convolution mechanism - combining the advantages of DRRCM and ARFAM - acts together to extract high-quality rotation-equivariant features. Finally, the generated high-quality feature representation and the corrected predicted anchor boxes greatly improve the stability and accuracy of model training, and significantly improve the detection effect of small targets with high aspect ratios in remote sensing images. Based on the high-quality rotation-equivariant features generated by DRRCM and ARFAM, and the optimized predicted anchor boxes, use the IoU threshold method to select samples for model training.
[0078] S4. Collect the remote sensing images to be detected and input them into the optimized object detection model to obtain the object detection results.
[0079] In a preferred embodiment of the present invention, a Data-Enhanced Spatial Attention Module (DESAM) is designed in DRRCM. A spatial mask is generated through pooling, splicing, convolutional transformation, and the sigmoid activation function to weighted fuse features and highlight important spatial regions.
[0080] The weighted fused features are average pooled and input into the kernel angle prediction branch and the kernel weight prediction branch. This module enables the network to more precisely focus on the key feature positions in rotation object detection, thereby accurately generating the weights and angles of the predicted rotation kernels.
[0081] The method for DRRCM to use the data enhancement spatial attention module DESAM to predict the weights and angles of rotation convolution kernels is as follows:
[0082] Perform channel-based average pooling and max pooling on the feature map F of the original image after depth convolution (depth convolution is part of DESAM, and DESAM is part of DRRCM), denoted as P avg (.), and P max (.), and extract the spatial relationship:
[0083] S avg =P avg (F), S max =P max (F)
[0084] Among them, S avg and S max are spatial feature descriptors after average pooling and max pooling. To allow information interaction between different spatial descriptors, the features of spatial pooling are concatenated, and a convolutional layer is used to convert the pooled features (with 2 channels) into C in spatial attention maps S′:
[0085]
[0086] Among them, C in represents the number of input channels; for each spatial attention map S′, a sigmoid activation function is applied to obtain an individual spatial mask S i ′:
[0087] S i ′ = sigmoid(S′),
[0088] The features after depth convolution are weighted using the spatial mask and compressed into a C in -dimensional feature vector V Cin :
[0089] V Cin = P avg (S i ′ × F),
[0090] The pooled feature vectors are respectively fed into two branches. The first is the rotation kernel angle prediction branch. The feature vector is input into this branch and passes through Dropout, a linear layer, Softsign activation, and multiplication by a scale factor to obtain a set of angles θ i :
[0091] θ i = K(Softsign(z θ ))
[0092] Among them, z θ has no bias set in the linear layer to ensure that the angle prediction only depends on the changes in the input features and avoid learning offset angles; is the scale factor to expand the rotation range, and the proportion parameter is used to adjust the range of the angle;
[0093] The second is the rotation kernel weight prediction branch. The feature vector is input into this branch and passes through Dropout, a linear layer, and Sigmoid activation to obtain a set of weights λ i :
[0094] λi = sigmoid(z a )
[0095] where z α has a bias set for the linear layer to improve the flexibility of the model.
[0096] DESAM is initialized from a truncated normal distribution with a mean of zero and a standard deviation of 0.2 to help the model converge faster and reduce instability in the initial stage of training.
[0097] In a preferred embodiment of the present invention, the rotation angle θ i predicted by DESAM is used to reparameterize the weights inside the convolutional kernel, so that the convolutional kernel W i can be dynamically adjusted according to different input feature maps to achieve adaptive rotation:
[0098] Y i ′ = rotate(Y i , -θ i )
[0099] W i ′ = interpolation(W i , Y i )
[0100] where Y i represents the coordinates of the original sampling points, and Y i ′ represents the new sampling point coordinates obtained by rotating the original sampling points Y i counterclockwise by the angle θ i to achieve alignment between the convolutional kernel and the features; rotate represents the rotation operation; W i ′ represents the reparameterized convolutional kernel; interpolation(.) represents bilinear interpolation, which is used to calculate the weights at the new positions after the rotation of the W i convolutional kernel;
[0101] Multiply the reparameterized convolutional kernel by the corresponding λ i weights and sum them, and then perform a convolution operation with the input feature map to finally generate high-quality orientation-aware features Y:
[0102]
[0103] where n represents the number of rotated convolutional kernels.
[0104] In a preferred embodiment of the present invention, through the anchor refinement feature alignment module ARFAM, based on the regression branch, corrected prediction anchor boxes are quickly generated on the orientation-sensitive feature map to guide the dynamic adjustment of the positions of the feature sampling points. The specific steps are as follows:
[0105] Based on the high-quality feature maps generated by DRRCM in the backbone network, ARFAM is used to further refine the anchor points through regression operations, and the offset field is calculated according to the refined anchor point parameters to dynamically adjust the positions of the aligned convolutional sampling points and generate a feature representation that is more precisely aligned with the target object.
[0106] The features output by DRRCM are fused through FPN (Feature Pyramid Network) to obtain feature maps, and only one initial square anchor point is preset at each position of the feature map, which is corrected into a high-quality directional anchor point through the regression branch, thereby reducing the large number of anchor points preset on the feature map to reduce the computational load. Predict the offset of the anchor box regression target:
[0107]
[0108] Among them, (x, y, w, h, θ) represent the center coordinates, width, height, and angle parameters of the initial anchor point; (x g , y g , w g , h g , θ g ) represent the center coordinates, width, height, and angle parameters of the true bounding box; (Δx g , Δy g , Δw g , Δh g , Δθ g ) represent the offset between the true bounding box and the initial anchor point. By regressing these offsets, the model can adjust the initial anchor point to a corrected predicted anchor box that is closer to the true bounding box; R(θ) represents the rotation transformation matrix used to convert the center point coordinates of the true bounding box into the coordinate system relative to the initial anchor point; k represents the proportionality coefficient used to adjust the angle value to ensure that the rotation angle is within a reasonable range;
[0109] To achieve feature extraction for directional targets, the position of the feature sampling points is adjusted through the corrected predicted anchor box as a guide to achieve dynamic convolutional alignment. On the basis of the original sampling points of the standard convolution, an offset o deduced from the corrected predicted anchor box is added:
[0110]
[0111] Among them, represents the position of the sampling point of the predicted anchor box; (p0 + p n ) represents the position of the regular sampling point of the standard convolution; p0 and p n represent the two-dimensional coordinates and relative offset of the sampling point respectively, and R represents the regular grid of the standard convolution [(p x , p y)];
[0112] The dynamic alignment convolution combines the offset o and the input feature X, adjusting the sampling point positions according to the shape and orientation of the predicted anchor boxes, so as to better match the geometric characteristics of the actual targets:
[0113]
[0114] Among them, Y(p) represents the value of the output feature map at position p, and w(p n ) is the weight of the convolution kernel at position p n , and x(·) represents the input feature map after FPN feature fusion.
[0115] In a preferred embodiment of the present invention, based on the high-quality rotation equivariant features generated by DRRCM and ARFAM, and the optimized predicted anchor boxes, a target detection model is trained, specifically:
[0116] The generated high-quality feature representations and the corrected predicted anchor boxes are used for model training, and positive and negative samples are selected for training by a fixed IoU threshold between the corrected predicted anchor boxes and the true ground truth bounding boxes:
[0117]
[0118] Among them, A∩B represents the intersection of the predicted anchor box and the true ground truth bounding box, the overlapping area; A∪B represents the union of the predicted anchor box and the true ground truth bounding box, the non-overlapping area.
[0119] The present invention also provides a remote sensing image rotation high aspect ratio small target detection system based on deep learning, including a processing unit, the processing unit includes a dynamic refinement rotation convolution module DRRCM and an anchor point refinement feature alignment module ARFAM, and the processing unit executes the method of the present invention to obtain the target detection result of the remote sensing image.
[0120] In a preferred embodiment of the present invention, as Figure 2 shown, the dynamic refinement rotation convolution module DRRCM includes a data augmentation spatial attention module DESAM and a convolution kernel that can be adaptively adjusted according to the pose information of the oriented target.
[0121] The data augmentation spatial attention module DESAM includes a depth convolution layer, a max pooling layer, a first average pooling layer, a concatenation layer, a convolution layer, a second average pooling layer, and two fully connected layers.
[0122] The output ends of the deep convolutional layers are respectively connected to the max pooling layer and the first average pooling layer. The output ends of the max pooling layer and the first average pooling layer are connected to the concat (concatenation) layer. The concat layer, the convolutional layer, the second average pooling layer, and the two fully connected layers are connected in sequence.
[0123] In a preferred embodiment of the present invention, as Figure 3 shown, the anchor refinement feature alignment module ARFAM includes a regression branch layer and an alignment convolutional layer connected in sequence.
[0124] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0125] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for detecting small targets with high aspect ratio in remote sensing image rotation based on deep learning, characterized in that: The steps include: S1, the dynamic refinement rotation convolution module DRRCM is embedded into the backbone network ResNet50 of the target detection model, replacing part of its 3×3 convolution module. DRRCM uses the data enhancement spatial attention module DESAM to predict the weight and angle of the rotation convolution kernel; The predicted rotation convolution kernel weights and angle parameters are combined, and the dynamic refinement rotation convolution module DRRCM is used to adaptively adjust the convolution kernel according to the posture information of the directional target to generate a direction-sensitive feature map so that it is accurately aligned with the target features; S2, using the anchor refinement feature alignment module ARFAM, generates a modified prediction anchor box on the direction-sensitive feature map based on the regression branch, which serves as a guide to dynamically adjust the position of the feature sampling points; S3, based on the high-quality rotation equivariant features finally generated by ARFAM and the optimized predicted anchor boxes, the object detection model is trained using the IoU threshold method; S4, collect the remote sensing image to be detected, and input it into the optimized target detection model to obtain the target detection result.
2. The method for detecting small targets with high aspect ratio in remote sensing image rotation based on deep learning as claimed in claim 1, characterized in that: The method by which DRRCM uses the data-enhanced spatial attention module DESAM to predict the weights and angles of the rotation convolution kernel is: The feature map F of the original image after deep convolution is used to perform channel-based average pooling and maximum pooling, which are denoted as P avg (.) and P max (.), extract spatial relationships: S avg =P avg (F),S max =P max (F) Among them, S avg and S max It is the spatial feature descriptor after average pooling and maximum pooling. In order to allow information interaction between different spatial descriptors, the spatial pooled features are concatenated and a convolutional layer is used. Convert the pooled features to C in A spatial attention map S′: Among them, C in Represents the number of input channels; for each spatial attention map S′, a sigmoid activation function is applied to obtain the individual spatial mask S of each convolution kernel. i ′: S i ′=sigmoid(S′), The features after deep convolution are weighted using spatial masks and compressed into a C through a global average pooling. in The eigenvector V of Cin : V Cin =P avg (S i ′×F), The pooled feature vector is sent to two branches respectively. The first one is the rotation kernel angle prediction branch. The feature vector is input into this branch and then activated by Dropout, linear layer, Softsign and multiplied by the scale coefficient to obtain a set of angles θ. i : θ i =K(Softsign(z θ )) Among them, z θ No bias is set for the linear layer to ensure that the angle prediction depends only on the changes in the input features to avoid learning biased angles; is the proportional coefficient to expand the rotation range, and the proportion parameter is used to adjust the angle range; The second is the rotation kernel weight prediction branch. The feature vector is input into this branch and then activated by Dropout, linear layer and Sigmoid to obtain a set of weights λ i : λ i =sigmoid(z a ) Among them, z α A bias is set for the linear layer to increase the flexibility of the model.
3. The method for detecting small targets with high aspect ratio in remote sensing image rotation based on deep learning as claimed in claim 2, characterized in that: The rotation angle θ generated by DESAM prediction i The parameters reparameterize the weights inside the convolution kernel so that the convolution kernel W i It can dynamically adjust according to different input feature maps to achieve adaptive rotation: AND i ′=rotate(Y i ,-θ i ) W i ′=interpolation(W i ,Y i ′) Among them, Y i Represents the coordinates of the original sampling point, Y i ′ represents the original sampling point Y i According to the angle θ i The coordinates of the new sampling points after counterclockwise rotation to achieve alignment between the convolution kernel and the feature; rotate indicates the rotation operation; W i ′ represents the re-parameterized convolution kernel; interpolation(.) represents bilinear interpolation, which is used to calculate W i The weight value of the new position after the convolution kernel is rotated; The reparameterized convolution kernel and the corresponding λ i The weights are multiplied and summed, and then a convolution operation is performed with the input feature map to finally generate a high-quality direction-aware feature Y: Where n represents the number of rotated convolution kernels.
4. The method for detecting small targets with high aspect ratio in remote sensing image rotation based on deep learning as claimed in claim 1, characterized in that: Through the anchor refinement feature alignment module ARFAM, based on the regression branch, a corrected prediction anchor box is quickly generated on the direction-sensitive feature map to guide the dynamic adjustment of the position of the feature sampling point. The specific steps are as follows: The features output by DRRCM are fused with FPN features to obtain a feature map, and only an initial square anchor point is preset at each position of the feature map. It is corrected into a high-quality directional anchor point through the regression branch, and the offset of the anchor frame regression target is predicted: Among them, (x, y, w, h, θ) represents the center coordinates, width, height and angle parameters of the initial anchor point; (x g ,y g 、w g 、h g ,θ g ) represents the center coordinates, width, height and angle parameters of the real bounding box; (Δx g , Δy g , Δw g , Δh g , Δθ g ) represents the offset between the true bounding box and the initial anchor point. By regressing these offsets, the model can adjust the initial anchor point to a modified predicted anchor box that is closer to the true bounding box; R(θ) represents the rotation transformation matrix, which is used to transform the coordinates of the center point of the true bounding box to the coordinate system relative to the initial anchor point; k represents the scale factor used to adjust the angle value to ensure that the rotation angle is within a reasonable angle; The corrected prediction anchor box is used as a guide to adjust the position of the feature sampling points to achieve dynamic convolution alignment. On the basis of the original sampling points of the standard convolution, an offset o calculated by the corrected prediction anchor box is added: in, It is expressed as the location of the predicted anchor box sampling point; (p0+p n ) represents the position of the regular sampling point of the standard convolution; p0 and p n Represent the two-dimensional coordinates and relative offset of the sampling point respectively, and R represents the regular grid of standard convolution [(p x ,p y )]; The dynamic alignment convolution combines the offset o and the input feature x so that the sampling point position is adjusted according to the shape and direction of the predicted anchor box to match the geometric characteristics of the actual target: Among them, Y(p) represents the value of the output feature map at position p, w(p n ) is the convolution kernel at position p n , and x(·) represents the feature map of the input after feature fusion.
5. The method for detecting small targets with high aspect ratio in remote sensing image rotation based on deep learning as claimed in claim 1, characterized in that: Based on the high-quality rotation-equivariant features generated by DRRCM and ARFAM, and the optimized predicted anchor boxes, the target detection model is trained as follows: The generated high-quality feature representation and the corrected predicted anchor box are used for model training, and positive and negative samples are selected for training based on a fixed IoU threshold between the corrected predicted anchor box and the true ground bounding box: Among them, A∩B represents the intersection of the predicted anchor box and the true ground bounding box, the overlapping area; A∪B represents the union of the predicted anchor box and the true ground bounding box, the non-overlapping area.
6. A remote sensing image rotation high aspect ratio small target detection system based on deep learning, characterized in that: The method comprises a processing unit, wherein the processing unit comprises a dynamic refinement rotation convolution module DRRCM and an anchor refinement feature alignment module ARFAM. The processing unit executes the method according to any one of claims 1 to 5 to obtain a remote sensing image which is a target detection result.
7. The remote sensing image rotation high aspect ratio small target detection system based on deep learning as claimed in claim 6, characterized in that: The dynamic refinement rotation convolution module DRRCM includes a data enhancement spatial attention module DESAM and a convolution kernel that can be adaptively adjusted according to the posture information of the directional target; The data enhanced spatial attention module DESAM includes a deep convolutional layer, a maximum pooling layer, a first average pooling layer, a splicing layer, a convolutional layer, a second average pooling layer and two fully connected layers; The output ends of the deep convolutional layer are respectively connected to the maximum pooling layer and the first average pooling layer, the output ends of the maximum pooling layer and the first average pooling layer are connected to the splicing layer, and the splicing layer, the convolutional layer, the second average pooling layer and the two fully connected layers are connected in sequence.
8. The remote sensing image rotation high aspect ratio small target detection system based on deep learning as claimed in claim 6, characterized in that: The anchor refinement feature alignment module ARFAM includes a regression branch layer and an alignment convolution layer connected in sequence.
Citation Information
Cited By
Remote sensing image directed target detection method and system based on rotation equivariant structure
CN120823372A
Remote sensing image directed target detection method and system based on rotation equivariant structure
CN120823372B