Target detection method and system for weak and small target in remote sensing image

By exploring the texture and boundary high-frequency details of weak targets in remote sensing images, and using task-decoupled RCNN network for independent training, the problem of low detection accuracy of weak targets in remote sensing images is solved, significantly improving detection performance.

WO2025118826A1PCT designated stage expired Publication Date: 2025-06-12CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2024/125304
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-10-16
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The weak targets in remote sensing images are difficult to detect effectively due to the small proportion of pixels, widespread feature changes, and are susceptible to imaging conditions and context, resulting in low detection accuracy.

Method used

A method of perceived texture and boundary feature object detection for weak targets of remote sensing images is proposed. Through the texture perception enhancement module and boundary perception fusion module, the high-frequency details of the target are explored, and independent training of classification and positioning is carried out through the task-decoupled RCNN network.

Benefits of technology

By improving the information expression and feature extraction of weak targets, the detection accuracy of weak targets is significantly improved, and the problem of insufficient detection performance of weak targets in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024125304_12062025_PF_FP_ABST
    Figure CN2024125304_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of intelligent remote-sensing-image interpretation, and specifically provides a texture-aware and boundary-aware target detection method and system for a weak and small target in a remote sensing image. The method comprises: inputting a remote sensing image into a feature extraction network to extract a basic feature, and then extracting a texture-aware feature; inputting the remote sensing image into a boundary map extraction module to extract a binary boundary map, and then extracting a boundary-aware feature; fusing the boundary-aware feature with the basic feature by means of a boundary-guided feature module, so as to obtain a boundary-guided feature; and decoupling the texture-aware feature and the boundary-guided feature and outputting same, inputting the texture-aware feature and the boundary-guided feature into a task-decoupling RCNN for double-branch decoupling prediction, so as to obtain a classification and positioning result, and completing texture-aware and boundary-aware target detection. By means of the method and system in the present invention, information that is hard for remote-sensing weak and small targets to show can be mined, thereby further improving the expression of the weak and small targets, and thus improving the detection performance regarding the weak and small targets.
Need to check novelty before this filing date? Find Prior Art

Description

Target detection method and system for small targets in remote sensing images

[0001] This application claims priority to the Chinese patent application filed with the Patent Office of China on December 7, 2023, with application number 202311669036.4 and invention name “Target detection method and system for small targets in remote sensing images”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present invention relates to the technical fields of computer vision, deep learning, and intelligent interpretation of remote sensing images, and specifically provides a method and system for perceptual texture and boundary target detection of small targets in remote sensing images. Background Art

[0003] In recent years, with the massive investment in civil and commercial satellites and the rapid development of imaging technology, the quality of remote sensing images has been improved and the cost of image acquisition has been reduced, which has created favorable conditions for remote sensing image target detection. However, small targets in remote sensing images have always been a difficult problem for target detection tasks. These small targets are widely distributed in remote sensing images and have the following properties: (1) These targets only account for a small proportion of pixels in remote sensing images; (2) The high-frequency details of the targets (such as texture details, boundary clues, color, etc.) are not significant, the features vary widely, and are easily disturbed by the imaging conditions and complex context at the time. However, small targets often carry more important information, and effectively detecting these targets is of great significance for real-world scenarios. In addition, in target detection, classification and regression are two highly related and contradictory tasks. Most detectors share the same features to perform classification and regression, but the two tasks have their own information focus. Classification tasks tend to focus on the most semantic features within the target, while regression tasks are more sensitive to the target boundaries. Using coupled features does not produce the best performance, and the model needs to extract suitable features for specific tasks. Summary of the Invention

[0004] To address these issues, the present invention provides a method for detecting small targets in remote sensing images by perceiving texture and boundary features. This method enhances the information representation of small targets by exploring key high-frequency details such as texture and boundaries, thereby enhancing their features and improving their detection accuracy. Furthermore, the present invention proposes a decoupled detection head structure that separates the classification and localization tasks. This decoupled network enables targeted training based on specific tasks.

[0005] The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets provided by the present invention comprises the following steps:

[0006] Step 1: Input the remote sensing image into the feature extraction network to extract basic features, and input the basic features into the texture perception enhancement module to extract texture perception features;

[0007] Step 2: Input the remote sensing image into the boundary map extraction module to extract a binary boundary map, input the boundary map into the boundary feature extraction module to extract boundary perception features; and fuse the boundary perception features with the basic features through the boundary guidance feature module to obtain boundary guidance features.

[0008] Step 3: Decouple the texture perception feature and the boundary guidance feature and output them, and input them into the task-decoupled RCNN network for dual-branch decoupling prediction to obtain classification and positioning results, completing the perception texture and boundary target detection.

[0009] Preferably, in step 1, the basic features are input into a texture perception enhancement module, and extracting texture perception features includes: inputting the basic features into a feature pyramid to obtain fused features, and then inputting the fused features into a texture perception enhancement module to extract texture perception features.

[0010] Preferably, the step 1 comprises:

[0011] Input remote sensing images into the feature extraction network to extract basic features , input the basic features into the feature pyramid to obtain the fusion features ;

[0012] Wherein, c represents the channel dimension of the fused feature; h and w represent the length and width of the fused feature, respectively.

[0013] Preferably, inputting the fused features into a texture perception enhancement module to extract texture perception features comprises:

[0014] The fused feature map undergoes three parallel branches of 1×1 convolution to integrate the features and reduce the dimension to c1 to obtain the feature , and the features Dimensionality reduction to ;

[0015] calculate With the transposed features The covariance matrix between , capturing the correlation between pixels at different positions in the feature dimension space, the calculation formula is:

[0016] ;

[0017] in, represents the intermediate features after convolution processing, (x, y) represents the position of the pixel, Represents element-wise multiplication;

[0018] A spatial attention mechanism is introduced to strengthen the location of the target in the remote sensing image and highlight the intrinsic pixel relationship of the target;

[0019] The fusion features of the input Perform global average pooling in the channel dimension and use 3×3 to remove excess noise;

[0020] Sigmoid activation function is used to normalize features and generate attention weights;

[0021] Convert the weight dimension to Generate attention weight map and with the covariance matrix The channel dimension of is aligned, expressed as:

[0022] ;

[0023] ;

[0024] Among them, GAP represents the global average pooling operation, represents Sigmoid activation, Reconstruct the function for the dimension, To fusion features, is the characteristic of a single channel, is a 3×3 convolution operation, is the result of global average pooling;

[0025] The attention weight map and the covariance matrix Multiply to get the covariance matrix of the target enhancement ;

[0026] The dimension of the calculated target enhancement covariance matrix is ;

[0027] Regularize the target enhanced covariance matrix and convert the dimension of the target enhanced covariance matrix to ;

[0028] Among them, the first dimension wh represents the number of covariance matrices of the target enhancement, which is equal to the feature size; the second dimension w represents the covariance matrix of the target enhancement The third dimension h represents the covariance matrix of the target enhancement width;

[0029] ;

[0030] in, represents the matrix dot product, Reconstruct the function for the dimension, is the target enhanced covariance matrix, is the covariance matrix, is the attention weight map, is the regularization function;

[0031] A 1×1 convolution operation is used to integrate the relationship matrix of pixels between different channels and convert the channels to ;

[0032] By and Multiply and use 1×1 convolution to restore the feature channel to c to obtain texture features;

[0033] By adding the captured texture features to the fusion features In this paper, texture perception features are obtained.

[0034] Preferably, the captured texture features are added to the fusion features In , the texture perception feature is obtained by the following formula:

[0035] ;

[0036] in, represents the addition of matrix elements, is the texture perception feature, is a 1×1 convolution operation, is the mapping feature, is the matrix dot product, To fusion features, is the target augmented covariance matrix.

[0037] Preferably, the step of inputting the remote sensing image into the boundary map extraction module to extract the binary boundary map comprises: performing sliding extraction on the input remote sensing image using an edge extraction operator to generate a gradient map of the remote sensing image.

[0038] Preferably, a gated filtering process is performed on the gradient map, a threshold is set, and pixel values ​​smaller than the threshold are set to zero to remove noise and useless texture information.

[0039] Preferably, the boundary map is input into a boundary feature extraction module, and the extraction of boundary perception features includes:

[0040] Use 3×3 convolution, BN, ReLU, and 3×3 convolution to extract underlying features;

[0041] 1×1 convolution, BN, and ReLU operations are used to reduce the dimension of the underlying features. Three parallel branches including 1×3 convolution, 3×1 convolution, and 3×3 convolution are used to extract edge information from the underlying features. The edge information is aggregated and fused to obtain potential edge features.

[0042] Use 1×1 convolution, BN, and ReLU operations to restore feature channels;

[0043] The underlying features are transferred through jump connections to fuse with the potential edge features, and boundary-aware features are generated through 1×1 convolution and Sigmoid activation.

[0044] Preferably, fusing the boundary perception feature with the basic feature through a boundary guidance feature module to obtain the boundary guidance feature includes:

[0045] The downsampled boundary perception features are added to the basic features to fuse the features, and a skip connection structure and 3×3 convolution are used to amplify the edge structure and integrate the fused information;

[0046] A channel attention mechanism is used to highlight the important feature channels in the fused features, and the channels are added to the fused features output by the feature pyramid to obtain boundary guidance features.

[0047] Preferably, step 3 decouples and outputs the texture perception feature and the boundary guidance feature, and inputs them into the task-decoupled RCNN network for dual-branch decoupling prediction to obtain classification and positioning results. Completing the perception texture and boundary target detection includes:

[0048] Input the texture perception features generated in step 1 into the classification branch in the task decoupling RCNN, and input the boundary guidance features generated in step 2 into the positioning branch in the task decoupling RCNN;

[0049] Setting different thresholds, performing non-maximum suppression (NMS) on the classification branch and the positioning branch to decouple them, and generating proposal regions with different densities;

[0050] Decouple sampling is performed on the proposed region to obtain a decoupled region of interest; for the classification branch, a fixed ratio is used to sample a quantitative positive sample and a quantitative negative sample; for the positioning branch, a quantitative positive sample is sampled;

[0051] The scope of the region of interest of the classification branch is expanded, and the high-quality region of interest is synthesized for the positioning branch using the Weiszfeld algorithm to assist in positioning prediction.

[0052] Preferably, all the proposed regions assigned to each target are synthesized with the upper left corner and lower right corner of the true labeled region, and the geometric center of the discrete sample points is found by iterative reweighted least squares.

[0053] Assume that the reference point to be fitted is The algorithm tries to find a center point y through continuous iteration so that the distance between the point and the reference point is the smallest, which can be expressed as:

[0054] ;

[0055] in, represents the L2 norm, j represents the number of reference points, D represents the sum of the distances between the synthetic center point y and the reference point x, and argmin represents the parameter value y that minimizes the sum;

[0056] Solving in continuous space, according to optimization theory, the above formula will obtain the minimum value at the extreme value. Therefore, find the partial derivative of y:

[0057] ;

[0058] Solve the above equation and find the center point:

[0059] ;

[0060] Where k represents the number of iterations. Through continuous iterations, the center point obtained makes D less than a specified value. This point is the geometric center of the fitting, and this center point does not coincide with any reference point.

[0061] By fitting the upper left and lower right corners of the box, we synthesize a proposed area close to the true annotation box, train the positioning branch network, accelerate the convergence of the model, and improve the stability of the model.

[0062] The present invention also provides a texture perception and boundary target detection system for small and weak targets in remote sensing images, the texture perception and boundary target detection system for small and weak targets in remote sensing images comprising:

[0063] Feature extraction network, used to extract basic features of remote sensing images;

[0064] A texture perception enhancement module, configured to extract texture perception features using the basic features;

[0065] A boundary map extraction module, used to extract a binary boundary map of the remote sensing image;

[0066] A boundary feature extraction module, used to extract boundary perception features of the boundary map;

[0067] A boundary guidance feature module, configured to fuse the boundary perception feature with the basic feature to obtain a boundary guidance feature;

[0068] The decoupling transmission module is used to decouple the texture perception features and the boundary guidance features and output them, and input them into the task-decoupled RCNN network for dual-branch decoupling prediction to obtain classification and positioning results, thereby completing the perception texture and boundary target detection.

[0069] Compared with the prior art, the present invention can achieve the following beneficial effects:

[0070] The present invention provides a method for detecting weak targets in remote sensing images by perceiving texture and boundary features. Specifically, a unique texture and boundary perception network is provided to solve the problem of weak targets in remote sensing images. In order to mine deep information, a texture perception enhancement module and a boundary perception fusion module are proposed. The former can fully explore the texture details inside the target. It constructs the connection between features through the calculated covariance matrix, and at the same time highlights the texture difference between the target and the background. The latter introduces additional boundary clues to highlight the spatial position of the target. By exploring the high-frequency details such as the key texture and boundary of the target, the key features of the weak target are supplemented. The combination of the two modules enhances the network's perception of weak targets. At the same time, in order to resolve the task focus phenomenon caused by the entanglement between classification and regression, a task-decoupled RCNN is proposed to independently train two-branch networks, thereby achieving refined prediction of classification and positioning. The detection method proposed in the present invention can mine information that is difficult to highlight for weak targets in remote sensing, further improve the expression of weak targets, and thus improve the detection performance of weak targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] 1 is a flow chart of a method for detecting targets in remote sensing images according to an embodiment of the present invention;

[0072] FIG2 is an example diagram of a remote sensing image with a small target according to an embodiment of the present invention;

[0073] FIG3 is an overall schematic diagram of a target detection network according to an embodiment of the present invention;

[0074] FIG4 is a schematic diagram of a texture perception enhancement module according to an embodiment of the present invention;

[0075] FIG5 is a schematic diagram of boundary map extraction in a boundary-aware fusion module according to an embodiment of the present invention;

[0076] FIG6 is a schematic diagram of boundary-aware feature extraction in a boundary-aware fusion module according to an embodiment of the present invention;

[0077] FIG7 is a schematic diagram of boundary-aware feature fusion in a boundary-aware fusion module according to an embodiment of the present invention;

[0078] FIG8 is a schematic diagram of a task-decoupled RCNN network according to an embodiment of the present invention;

[0079] FIG9 is a schematic diagram of a proposed synthesis strategy according to an embodiment of the present invention;

[0080] FIG10 is an example diagram of the results of detecting small targets in remote sensing images using the target detection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0081] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the following description, identical modules are denoted by identical reference numerals. In the case of identical reference numerals, their names and functions are also identical. Therefore, their detailed description will not be repeated.

[0082] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation of the present invention.

[0083] In a specific embodiment of the present invention, a method for detecting small targets in remote sensing images based on texture perception and boundary features is provided, comprising the following steps:

[0084] Step 1: Input the remote sensing image into the feature extraction network to extract basic features, and input the basic features into the texture perception enhancement module to further explore the texture pattern of the target in the remote sensing image and extract texture perception features;

[0085] In a specific implementation, inputting the basic features into a texture perception enhancement module, and extracting texture perception features includes: inputting the basic features into a feature pyramid to obtain fused features, and then inputting the fused features into the texture perception enhancement module to extract texture perception features.

[0086] In a specific embodiment, the step 1 specifically includes: inputting the remote sensing image with small targets into the feature extraction network to obtain basic features At the same time, the basic features are input into the feature pyramid to obtain the fused features , where c represents the channel dimension of the fused feature; h and w represent the length and width of the fused feature, respectively.

[0087] The fused features are input into the texture perception enhancement module for texture feature extraction. Specifically, first, the fused feature map undergoes three parallel branches of 1×1 convolution to integrate the features and reduce the dimension to c1, obtaining the feature , and the features Dimensionality reduction to Next, calculate With the transposed features The covariance matrix between To capture the correlation between pixels at different positions in the feature dimension space, the calculation formula is:

[0088] ;(1)

[0089] in, represents the intermediate features after convolution processing, (x, y) represents the position of the pixel, Represents element-wise multiplication. At the same time, a spatial attention mechanism is introduced to strengthen the location of the target in the remote sensing image and highlight the intrinsic pixel relationship of the target. First, the fusion feature of the input Perform global average pooling of the channel dimension and use 3×3 to remove excess noise; then, the Sigmoid activation function is used to normalize the features and generate attention weights; finally, the weight dimension is converted to Generate attention weight map , and with the covariance matrix The channel dimension is aligned, which can be expressed as:

[0090] ;(2)

[0091] ;(3)

[0092] Among them, GAP represents the global average pooling operation, represents Sigmoid activation, Reconstruct the function for the dimension, To fusion features, is the characteristic of a single channel, is a 3×3 convolution operation, is the result of global average pooling; then, the attention map and the covariance matrix Multiply to get the covariance matrix of the target enhancement At this time, the dimension of the calculated target enhanced covariance matrix is , by regularizing these values ​​using SoftMax, that is, regularizing the entirety of the calculated target enhanced covariance matrix and converting the dimension of the target enhanced covariance matrix to , the covariance matrix of the target enhancement before processing is two-dimensional, and the covariance matrix of the target enhancement is converted to three-dimensional after conversion; wherein the first dimension wh represents the number of the covariance matrix of the target enhancement, which is equal to the feature size; the second dimension w represents the covariance matrix of the target enhancement The third dimension h represents the covariance matrix of the target enhancement Therefore, A matrix in the mathematical representation means the relationship between a certain pixel and all pixels at other locations. It can be written as:

[0093] ;(4)

[0094] in, represents the matrix dot product, Reconstruct the function for the dimension, is the target enhanced covariance matrix, is the covariance matrix, is the attention weight map, is the regularization function; at the same time, a 1×1 convolution operation is used to integrate the relationship matrix of pixels between different channels and convert the channels to At this point, a relationship map is generated that can amplify the similar positions of homogeneous pixels and targets, which can be used to emphasize the similar texture representation of the target at the global level and further enhance the valuable texture features; finally, by and Multiply to obtain valuable texture features and use 1×1 convolution to restore the feature channel to c; by adding the captured texture features to the original features In the process of obtaining texture perception features, the specific processing is as follows:

[0095] ;(5)

[0096] in, represents the addition of matrix elements, is the texture perception feature, is a 1×1 convolution operation, is the mapping feature, is the matrix dot product, To fusion features, is the target augmented covariance matrix.

[0097] Step 2: Input the remote sensing image into the boundary map extractor to extract a binary boundary map, and input the boundary map into the boundary perception feature extraction module to further extract the boundary perception features of the target; finally, the extracted features are fused with the basic features extracted in step 1 to obtain boundary guidance features. The step of inputting the remote sensing image into the boundary map extraction module to extract the binary boundary map includes: using an edge extraction operator to perform sliding extraction on the input remote sensing image to generate a gradient map of the remote sensing image; preferably, performing gated filtering on the gradient map, setting a threshold, and setting pixel values ​​less than the threshold to zero to eliminate noise and useless texture information.

[0098] In a specific embodiment of the present invention, step 2 includes:

[0099] The proposed boundary-aware fusion module consists of three main components: boundary map extraction, boundary-aware feature extraction, and boundary-guided feature fusion. The boundary map extraction module extracts the boundary map. Specifically, an edge extraction operator is used to perform sliding extraction on the input image, generating a gradient map of the image. This gradient map may contain a large amount of noise and useless texture details. A hard-threshold-based gated filtering method is used to remove these interferences. Specifically, a threshold is set, and pixels with values ​​below the threshold are set to 0, thereby preserving the image edges as much as possible.

[0100] The boundary feature extraction module extracts boundary-aware features. Specifically, for the generated boundary map, a 3×3 convolutional layer, a batch normalization layer, a ReLU layer, and a 3×3 convolutional layer are used to extract underlying features. A bottleneck structure with asymmetric and symmetric convolutions extracts useful edge features and further suppresses noise. Specifically, 1×1 convolution, a batch normalization layer, and a ReLU layer are first used to reduce the dimensionality of the feature map. Three parallel branches, each containing a 1×3 convolution, a 3×1 convolution, and a 3×3 convolution, are then used to extract potential edge features and fuse them together. Subsequently, a 1×1 convolution, a batch normalization layer, and a ReLU layer are added to restore the feature channels. Finally, the underlying features are fused with the extracted potential edge features via skip connections. A 1×1 convolution and a sigmoid activation are then performed to generate boundary-aware features that contain rich edge information.

[0101] Boundary features are fused using the boundary-guided feature module. Specifically, first, the resolution of the boundary-aware features is reduced using a downsampling operation to make them consistent with the basic features of the feature extraction network in step 1. The boundary-aware features are then fused with the basic features in an element-wise summation. Subsequently, the edge-fused features generated in the previous step are added to the basic features of the feature extraction network using skip connections to amplify the edge structure in the basic features. Important information is then integrated through a 3×3 convolutional layer. Secondly, a channel-wise attention mechanism is used to emphasize important feature channels. This channel-wise attention mechanism consists of a global average pooling layer, a 1×1 convolutional layer, and a sigmoid activation layer to generate channel attention weights. These weights are multiplied with the edge-fused features through skip connections to amplify important information. Finally, the boundary-guided features are added to the fused features output by the feature pyramid to generate the boundary-guided features.

[0102] Step 3: The model decouples the extracted texture perception features and boundary guidance features, outputs them, and inputs them into the task-decoupled RCNN for two-branch decoupled prediction to obtain classification and positioning results;

[0103] In a specific embodiment of the present invention, step 3 includes:

[0104] The texture perception features and boundary guidance features generated in steps 1 and 2 are independently input into the task decoupling RCNN network for decoupled category and position prediction. The texture perception features are input into the classification branch, and the boundary guidance features are input into the localization branch.

[0105] Specifically, the texture perception feature generated in step 1 is input into the classification branch in the task decoupling RCNN, and the boundary guidance feature generated in step 2 is input into the positioning branch in the task decoupling RCNN;

[0106] Setting different thresholds, performing non-maximum suppression (NMS) on the classification branch and the positioning branch to decouple them, and generating proposal regions with different densities;

[0107] Decouple sampling is performed on the proposed region to obtain a decoupled region of interest; for the classification branch, a fixed ratio is used to sample a quantitative positive sample and a quantitative negative sample; for the positioning branch, a quantitative positive sample is sampled;

[0108] The scope of the region of interest of the classification branch is expanded, and the high-quality region of interest of the positioning branch is synthesized using the Weiszfeld algorithm to assist in positioning prediction.

[0109] In a specific implementation, first, the two-branch network is decoupled by NMS by setting different thresholds to generate proposal regions of different densities; then, the proposal regions are decoupled and sampled to obtain decoupled regions of interest. For the classification branch, a fixed ratio is used to sample quantitative positive samples and negative samples. For the positioning branch, only quantitative positive samples are sampled; next, the two-branch network is targeted and enhanced. More contextual information is introduced to assist in classification by expanding the scope of the region of interest. The expanded scope requires a trade-off. A scope that is too small is not enough to introduce sufficient information, and a scope that is too large will introduce too much noise interference. For the positioning branch, a proposal synthesis strategy is used to synthesize high-quality regions of interest to assist in positioning. The proposal synthesis strategy is based on the Weiszfeld algorithm, which synthesizes all the proposal regions assigned to each target with the upper left corner and lower right corner points of the true annotation area, and finds the geometric center of the discrete sample points in the form of iterative reweighted least squares. Specifically, assuming that the reference point to be fitted is , the algorithm tries to find a center point through continuous iteration Make the distance between this point and the reference point as small as possible.

[0110] ;(6)

[0111] in, represents the L2 norm, j represents the number of reference points, D represents the sum of the distances between the synthetic center point y and the reference point x, and argmin represents the parameter value This minimizes the sum. Solving in continuous space, according to optimization theory, the above formula will obtain the minimum value at the extreme value. Therefore, the partial derivative of y is:

[0112] ;(7)

[0113] Solve the above equation to find the center point, which does not coincide with any reference point.

[0114] ;(8)

[0115] Where k represents the number of iterations. Through continuous iterations, the center point obtained makes D less than a specified value, and this point is the fitted geometric center. By fitting the top left and bottom right corners of the bounding box, we synthesize a proposed region that closely matches the ground truth bounding box. This trains the localization branch network, accelerates model convergence, and improves model stability.

[0116] In a specific embodiment of the present invention, a texture perception and boundary target detection system for small and weak targets in remote sensing images is further provided. The texture perception and boundary target detection system for small and weak targets in remote sensing images comprises:

[0117] Feature extraction network, used to extract basic features of remote sensing images;

[0118] A texture perception enhancement module, configured to extract texture perception features using the basic features;

[0119] A boundary map extraction module, used to extract a binary boundary map of the remote sensing image;

[0120] A boundary feature extraction module, used to extract boundary perception features of the boundary map;

[0121] A boundary guidance feature module, configured to fuse the boundary perception feature with the basic feature to obtain a boundary guidance feature;

[0122] The decoupling transmission module is used to decouple the texture perception features and the boundary guidance features and output them, and input them into the task-decoupled RCNN network for dual-branch decoupling prediction to obtain classification and positioning results, thereby completing the perception texture and boundary target detection.

[0123] Compared to existing technologies, the present invention provides a unique texture- and boundary-aware network to address the problem of small and faint targets in remote sensing images. To unlock deeper information, a texture-aware enhancement module and a boundary-aware fusion module are proposed. The former fully explores texture details within the target. It establishes connections between features through a calculated covariance matrix, while highlighting the texture differences between the target and the background. The latter introduces additional boundary cues to highlight the target's spatial position. By exploring key high-frequency details such as texture and boundaries, key features of small and faint targets are enhanced. The combination of these two modules enhances the network's ability to perceive small and faint targets. To resolve the task-distraction phenomenon caused by the entanglement between classification and regression, a task-decoupled RCNN is proposed to independently train two branches of the network, thereby achieving refined predictions for classification and localization. The method proposed in this invention can unlock information that is difficult to highlight in small and faint remote sensing targets, further enhancing the representation of small and faint targets and thereby improving their detection performance.

[0124] The solution of the present invention is further described below with reference to specific embodiments and drawings.

[0125] In a specific embodiment of the present invention, the process of the perceptual texture and boundary target detection method for faint targets in remote sensing images is shown in Figure 1. Specifically, the remote sensing image with faint targets is input into the feature extraction network to extract basic features, and the basic features are input into the texture perception enhancement module to extract texture perception features. At the same time, the remote sensing image with faint targets is input into the boundary map extraction module to extract the boundary map of the image, and the boundary perception fusion module is used to extract boundary guidance features based on the boundary map. The texture perception features and boundary guidance features are independently input into the task decoupled RCNN network, and a decoupled sampling strategy is used to obtain task-specific regions of interest. The decoupled suggestion expansion strategy and suggestion synthesis strategy are used to strengthen the training of the classification and regression networks. The trained network can effectively detect faint targets in remote sensing images.

[0126] Figure 2 is a remote sensing image of a small target according to an embodiment of the present invention. As can be seen from the figure, due to the extremely high imaging altitude and the bird's-eye view, the targets in the image are small in size. At the same time, due to uneven lighting and cloud cover, the information of these targets is suppressed to a great extent, resulting in weak feature expression. These targets are easily submerged by a wide range of backgrounds, resulting in the loss of key information such as texture, boundary, and color. Coupled with further compression processing by the network, these small targets are easily ignored by the network, causing a large-scale missed detection phenomenon. The present invention addresses the key problem of missing features of small targets and proposes key technologies for mining texture and boundary features to effectively improve the detection accuracy of small targets.

[0127] Figure 3 is an overall schematic diagram of a method for detecting small targets in remote sensing images according to a specific embodiment of the present invention. It can be seen from the figure that the method provided by the embodiment of the present invention mainly consists of three parts: first, a texture perception enhancement module: responsible for exploring the texture pattern of the image and extracting texture perception features; second, a boundary perception fusion module: responsible for fusing the extracted boundary perception features with the basic features and generating boundary guidance features; third, a task decoupling RCNN network: responsible for decoupling the texture perception features from the boundary guidance features, and using the decoupled features in conjunction with the decoupling sampling strategy, the decoupling suggestion expansion strategy, and the suggestion synthesis strategy to train the classification and regression networks in a targeted manner to realize the detection of small targets in remote sensing images.

[0128] FIG4 is a schematic diagram of a texture perception enhancement module according to a specific embodiment of the present invention. Specifically, the fusion feature with a resolution of 256×100×100 output by the feature pyramid is shown in FIG4. For example (the fusion feature mentioned above , i is 3, 4, 5, here we take i=3 as an example), first, the feature map will undergo three parallel branches of 1×1 convolution to integrate features and reduce the dimension to 64, and obtain the feature , and the features Dimensionality reduction to Next, calculate With the transposed features The covariance matrix between To capture the correlation between pixels at different positions in the feature dimension space, the calculation formula is:

[0129] ; (1)

[0130] Where (x, y) represents the position of the pixel, Represents element-wise multiplication. However, the above steps can only calculate the correlation of all pixels without difference, but cannot focus on the pixel association within the target. To this end, a spatial attention mechanism is introduced to strengthen the location of the target and highlight the pixel relationship within the target. First, the input feature Perform global average pooling in the channel dimension and use 3×3 convolution to remove excess noise; then, the Sigmoid activation function is used to normalize the features, generate attention weights, and finally convert the weight dimension to Generate attention weight map and compare it with covariance matrix The channel dimension is aligned, which can be expressed as:

[0131] ; (9),

[0132] In this specific embodiment, compared with the original formula (2), this formula mainly clarifies the c value.

[0133] ; (3)

[0134] Among them, GAP represents the global average pooling operation, Represents Sigmoid activation; then, the attention map and the covariance matrix Multiply to get the covariance matrix of the target enhancement At this time, the dimension of the calculated target enhanced covariance matrix is , by regularizing these values ​​using SoftMax, that is, regularizing the entirety of the calculated target enhanced covariance matrix and converting the dimension of the target enhanced covariance matrix to , the covariance matrix of the target enhancement before processing is two-dimensional, and after conversion, the covariance matrix of the target enhancement is converted to three-dimensional; wherein, the first dimension 10000 represents the number of the covariance matrix of the target enhancement, which is equal to the feature size; the second dimension 100 represents the covariance matrix of the target enhancement The third dimension 100 represents the covariance matrix of the target enhancement Therefore, A matrix in the mathematical representation means the relationship between a certain pixel and all pixels at other locations. It can be written as:

[0135] ; (4)

[0136] in, Represents the matrix dot product; at the same time, a 1×1 convolution operation is used to integrate the relationship matrix of pixels between different channels and convert the channels to At this point, a relationship mapping is obtained that can amplify the similar positions of homogeneous pixels and targets, which can be used to emphasize the similar texture representation of the target globally and further enhance the beneficial texture features; finally, by and Multiply to obtain valuable texture features and use 1×1 convolution to restore the feature channel to 256. By adding the captured texture features to the original features In the process of obtaining texture perception features, the following steps are performed:

[0137] ; (5)

[0138] in, Represents the addition of matrix elements.

[0139] Figures 5 to 7 are schematic diagrams of boundary map extraction, boundary-aware feature extraction, and boundary-aware feature fusion in a boundary-aware fusion module according to an embodiment of the present invention. Specifically, assuming that the input image is As an optional solution, the embodiment of the present invention takes the Sobel kernel as an example. Slide in the horizontal and vertical directions to extract the directional gradient map; the gradient is the difference between adjacent pixels. The pixel difference in the area with the sharp gradient change is large, which is generally manifested as the edge inside the image. The horizontal gradient map is superimposed with the vertical gradient map to obtain the boundary map of the image. .However, There are many discontinuous noise points and interfering texture details in the memory. Gating is used to filter these interferences to obtain a fine boundary map. To support the extraction of boundary-aware features, it can be expressed as:

[0140] ; (10)

[0141] in, is a set threshold value, which is set to 0.15 in the embodiment of the present invention. When the pixel value in is less than 0.15 times the maximum grayscale value, it is set to 0. The boundary feature extraction module is further used to extract boundary perception features, and the underlying features are extracted using 3×3 convolution layer, BN layer, ReLU layer, and 3×3 convolution layer. A bottleneck structure with asymmetric convolution and symmetric convolution is used to extract useful edge features and further suppress noise. Specifically, first, the 1×1 convolution layer, BN layer, and ReLU layer are used to reduce the dimension of the feature map, and three parallel branches containing 1×3 convolution, 3×1 convolution, and 3×3 convolution are used to extract and fuse edge features; then, a 1×1 convolution layer, BN layer, and ReLU layer are added to restore the feature channel; then, the basic features are transferred through the jump connection. Fused with the fused edge features, and passed through a 1×1 convolution layer and a Sigmoid activation layer to generate boundary-aware features , this feature contains rich edge information.

[0142] ; (11)

[0143] ; (12)

[0144] Among them, Bottleneck represents the two 1×1 convolutional layers, BN layer, and ReLU layer from front to back. , , Represent the edge features extracted by asymmetric and symmetric convolution respectively, f bou The edge features are further fused using the boundary guidance feature module as a compensation for the extracted edge perception features. With basic features Integration, , to achieve effective edge emphasis and integrate it into the output of FPN To further highlight the edge expression of features. Specifically, first, a downsampling operation is used to reduce the resolution of the boundary-aware features to be consistent with the basic features of the feature extraction network, and the basic features and boundary-aware features are fused in the form of element-wise multiplication. Subsequently, a jump connection is used to add the fused features to the basic features of the feature extraction network to amplify the edge structure in the features, and a 3×3 convolution is used to integrate important information. It can be expressed as:

[0145] ; (13)

[0146] in represents the fused features, and D represents the downsampling operation; secondly, the channel attention mechanism is used to emphasize the important feature channels. The channel attention mechanism includes a global average pooling layer, a 1×1 convolution layer, and a Sigmoid activation layer. The generated channel attention weights are connected to the fused features through jump connections. Multiply and amplify the features of important channels; finally, fuse the features with the feature pyramid Adding features to generate boundary guidance , which can be expressed as:

[0147] ; (14)

[0148] Figure 8 is a schematic diagram of a task-decoupled RCNN network according to a specific embodiment of the present invention, which includes two separate classification and localization branches. After obtaining decoupled features, the task-decoupled RCNN network first performs decoupled sampling to generate specific suggestions for different tasks, thereby preventing information exchange between tasks. Specifically, the task-decoupled RCNN network first performs NMS with different thresholds to obtain proposal regions with different distributions. This embodiment sets the NMS threshold for classification to 0.7 and the NMS threshold for localization to 0.85. After obtaining proposal regions of varying density, these proposal regions are then decoupled sampled to obtain task-specific regions of interest. For the classification branch, this embodiment randomly samples 256 training examples while maintaining a positive-to-negative sample ratio of 1:3 to ensure that the network fully learns foreground and background information. For the localization branch, this embodiment randomly samples only 128 positive examples from the positive sample set to train the regression branch. Next, optimization is performed for the specific task. For the classification branch, a decoupled suggestion expansion strategy is implemented to incorporate more contextual information into the region of interest to assist classification. This embodiment expands the region of interest by 1.3 times its original size. For the localization branch, by proposing a synthesis strategy, more high-quality samples are added to the network to facilitate the training of the localization network. Finally, after a separate ROI Align operation to unify the size, the decoupled region of interest is input into a separate fully connected layer to achieve target classification and localization prediction.

[0149] Figure 9 is a schematic diagram of a suggestion synthesis strategy according to an embodiment of the present invention. The suggestion synthesis strategy is based on the Weiszfeld algorithm, which synthesizes all the suggested regions assigned to each target with the upper left corner and lower right corner of the real annotated region, and finds the geometric center of the discrete sample points through iterative reweighted least squares. Specifically, the embodiment of the present invention assumes that there are 4 reference points to be fitted. , , through the algorithm continuous iteration to try to find a center point Make the distance between this point and the four reference points as small as possible.

[0150] ; (15)

[0151] in, represents the L2 norm, argmin represents the parameter value This minimizes the sum. Solving in continuous space, according to optimization theory, the above formula will obtain the minimum value at the extreme value. Therefore, the partial derivative of y is:

[0152] ; (16)

[0153] Solve the above equation to find the geometric median point, which does not coincide with any reference point.

[0154] ; (17)

[0155] Where k represents the number of iterations. Through continuous iterations, the median point obtained makes Equation 15 less than a specified value. This point is the fitted geometric median point. By fitting the upper left and lower right corners of the box, a high-quality region of interest close to the ground truth box is synthesized to train the localization branch network, accelerating model convergence and improving model stability.

[0156] Figure 10 illustrates the results of detecting small and weak targets in remote sensing images using the target detection method according to an embodiment of the present invention. As shown in the figure, due to light and shadow obstruction, coupled with their small size, these targets exhibit weak details, making them difficult to effectively predict, particularly buildings, ships, cable towers, and self-built houses in the image. The target detection method for small and weak targets in remote sensing images provided by the present invention can still effectively recall these small and weak targets, accurately identifying their category and location, fully demonstrating the method's ability to detect small and weak targets.

[0157] Although the embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0158] The above specific embodiments of the present invention do not constitute a limitation on the scope of protection of the present invention. Any other corresponding changes and modifications made based on the technical concept of the present invention should be included in the scope of protection of the claims of the present invention.

Claims

1. A method for perceptual texture and boundary target detection of small targets in remote sensing images, characterized in that: The steps include: Step 1: Input the remote sensing image into the feature extraction network to extract basic features, and input the basic features into the texture perception enhancement module to extract texture perception features; Step 2: Input the remote sensing image into the boundary map extraction module to extract the binary boundary map, input the boundary map into the boundary feature extraction module to extract the boundary perception feature; fuse the boundary perception feature with the basic feature through the boundary guidance feature module to obtain the boundary guidance feature; Step 3: Decouple the texture perception feature and the boundary guidance feature, and input them into the task-decoupled RCNN network for dual-branch decoupled prediction to obtain classification and positioning results, thereby completing the perception of texture and boundary target detection.

2. The method for detecting small and weak targets in remote sensing images based on perceived texture and boundary targets according to claim 1, characterized in that: In the step 1, the basic features are input into a texture perception enhancement module to extract texture perception features, including: inputting the basic features into a feature pyramid to obtain fused features, and then inputting the fused features into a texture perception enhancement module to extract texture perception features.

3. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 1, characterized in that: In step 1: Input remote sensing images into the feature extraction network to extract basic features , input the basic features into the feature pyramid to obtain the fusion features ; Wherein, c represents the channel dimension of the fused feature; h and w represent the length and width of the fused feature, respectively.

4. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 3, characterized in that: Inputting the fused features into a texture perception enhancement module to extract texture perception features includes: The fusion features The 1×1 convolution of three parallel branches integrates the features and reduces the dimension to c1 to obtain the processed features. , the feature Dimensionality reduction to ; Calculate features With the transposed features The covariance matrix between , capturing the correlation between pixels at different positions in the feature dimension space; The fusion features of the input Perform global average pooling in the channel dimension and use a 3×3 convolutional layer to process the result of global average pooling; The feature map obtained by the 3×3 convolution layer is normalized using the Sigmoid activation function to generate attention weights; Convert the dimension of attention weight to , generate the attention weight map , and the attention weight map is compared with the covariance matrix The channel dimension of is aligned; The attention weight map With the covariance matrix Multiply to get the target enhanced covariance matrix , the dimension of the calculated target enhanced covariance matrix is ; Regularize the target enhanced covariance matrix and transform the dimension of the target enhanced covariance matrix to ; A 1×1 convolution operation is used to integrate the relationship matrix of pixels between different channels and convert the channels to ; By and Multiply and use 1×1 convolution to restore the feature channel to c to obtain texture features; By adding the captured texture features to the fusion features In this paper, texture-aware features are obtained.

5. The method for detecting small and weak targets in remote sensing images based on perceived texture and boundary targets according to claim 4, characterized in that: Covariance matrix The calculation formula is: ; in, represents the intermediate features after convolution processing, (x, y) represents the position of the pixel, Represents element-wise multiplication.

6. The method for detecting small and weak targets in remote sensing images based on perceived texture and boundary targets according to claim 5, characterized in that: The dimensions of the covariance matrix of the target augmentation after the transformation In , the first dimension wh represents the number of covariance matrices of the target enhancement, which is equal to the feature size; The second dimension w represents the covariance matrix of the target enhancement The third dimension h represents the covariance matrix of the target enhancement Width.

7. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 6, characterized in that: The dimension of the covariance matrix of the target enhancement is converted to The calculation formula is: ; in, represents the matrix dot product, Reconstruct the function for the dimension, is the target enhanced covariance matrix, is the covariance matrix, is the attention weight map, is the regularization function.

8. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 7, characterized in that: Convert the dimension of attention weight to , generate the attention weight map The calculation formula is: ; ; Among them, GAP represents the global average pooling operation, represents Sigmoid activation, Reconstruct the function for the dimension, To fusion features, is the characteristic of a single channel, is a 3×3 convolution operation, is the result of global average pooling.

9. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 8, characterized in that: By adding the captured texture features to the fusion features In the above example, texture perception features are obtained by the following formula: ; in, represents the addition of matrix elements, is the texture perception feature, is a 1×1 convolution operation, is the mapping feature, is the matrix dot product, To fusion features, is the target augmented covariance matrix.

10. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 1, characterized in that: In the step S2: inputting the remote sensing image into the boundary map extraction module to extract the binary boundary map includes: using an edge extraction operator to perform sliding extraction on the input remote sensing image to generate a gradient map of the remote sensing image.

11. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 10, characterized in that: The gradient map is subjected to gated filtering processing, a threshold is set, and pixel values ​​less than the threshold are set to zero to remove noise and useless texture information.

12. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 1, characterized in that: In step S2, the step of inputting the boundary map into the boundary feature extraction module to extract boundary perception features includes: Use 3×3 convolution, BN, ReLU, and 3×3 convolution to extract underlying features; 1×1 convolution, BN, and ReLU operations are used to reduce the dimension of the underlying features, and three parallel branches including 1×3 convolution, 3×1 convolution, and 3×3 convolution are used to extract edge information from the underlying features, and the edge information is aggregated and fused to obtain potential edge features; Use 1×1 convolution, BN, and ReLU operations to restore feature channels; The underlying features are transferred through jump connections to be fused with the potential edge features, and boundary-aware features are generated through 1×1 convolution and Sigmoid activation.

13. The method for detecting small and weak targets in remote sensing images based on perceived texture and boundary targets according to claim 1, characterized in that: In the step S2, fusing the boundary perception feature with the basic feature through a boundary guidance feature module to obtain the boundary guidance feature includes: The downsampled boundary perception features are added to the basic features to fuse the features, and a skip connection structure and 3×3 convolution are used to amplify the edge structure and integrate the fused information; A channel attention mechanism is used to highlight important feature channels in the fused features, and the channels are added to the fused features output by the feature pyramid to obtain boundary guidance features.

14. The method for detecting small and weak targets in remote sensing images based on texture perception and boundary targets according to claim 9, characterized in that: The step S3 specifically includes the following steps: Input the texture perception feature generated in step 1 into the classification branch in the task decoupling RCNN, and input the boundary guidance feature generated in step 2 into the positioning branch in the task decoupling RCNN; Setting different thresholds, performing non-maximum suppression (NMS) on the classification branch and the positioning branch to decouple the classification branch and generate proposal regions with different densities; Performing decoupled sampling on the proposed area to obtain a decoupled region of interest; For the classification branch, a fixed ratio is used to sample a quantitative amount of positive samples and negative samples; For the positioning branch, sampling a quantitative positive sample; The scope of the region of interest of the classification branch is expanded, and the Weiszfeld algorithm is used to synthesize a high-quality region of interest for the positioning branch to assist in positioning prediction.

15. A system for detecting small and weak targets in remote sensing images using texture perception and boundary targets, using the method for detecting small and weak targets in remote sensing images using texture perception and boundary targets as claimed in any one of claims 1 to 14, characterized in that: The system for detecting small and weak targets in remote sensing images based on texture perception and boundary targets includes: Feature extraction network, used to extract basic features of remote sensing images; A texture perception enhancement module, used for extracting texture perception features through the basic features; A boundary map extraction module, used for extracting a binary boundary map of the remote sensing image; A boundary feature extraction module, used to extract boundary perception features of the boundary map; A boundary guidance feature module, used to fuse the boundary perception feature with the basic feature to obtain a boundary guidance feature; The decoupled transmission module is used to decouple the texture perception feature and the boundary guidance feature and output them, and input them into the task-decoupled RCNN network to perform dual-branch decoupled prediction, obtain classification and positioning results, and complete the perception texture and boundary target detection.

Citation Information

Patent Citations

  • Deep learning small target detection method and device based on cascade fusion and attention mechanism

    CN112801158A

  • Remote sensing image weak and small target fusion multi-level feature target detection method

    CN113723172A

  • Remote sensing small target detection method, system and device based on fusion cascade attention mechanism and medium

    CN116385896A

  • Satellite remote sensing image small target detection method based on high-resolution characteristic self-attention

    CN117036980A

  • Target detection method and system for weak and small target in remote sensing image

    CN117636172A

Cited By

  • Intelligent landslide identification method and system based on hierarchical frame selection and boundary feature fusion

    CN120339868A

  • Electrocardiosignal analysis-oriented multi-task deep learning method

    CN120345904A

  • Infrared weak and small target detection method

    CN120471928A

  • Real scene three-dimensional processing method and system based on big data

    CN120747400A

  • Remote sensing image segmentation method based on multilayer feature fusion and prior guidance

    CN120852454A