Remote sensing image target detection method based on multi-level feature fusion and joint positioning
By adopting an anchorless object detection method of multi-level feature fusion and joint positioning in remote sensing image object detection, the problem of difficulty in extracting the target features of complex remote sensing images in the prior art is solved, and more efficient and accurate object detection is achieved.
Patent Information
- Application Number
- CN202310586691.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-05-23
AI Technical Summary
The existing remote sensing image object detection methods are difficult to effectively extract the features of remote sensing image objects with deformable shapes, arbitrary directions, tight arrangements and multiple small scales, resulting in missed detection and missed detection, and the target cannot be detected accurately.
An anchor-free object detection method based on multi-level feature fusion and joint positioning is adopted. The feature extraction network and feature fusion network extract and fuse target features, generate multi-scale feature maps, and coarse-fine joint positioning is performed in the detection head network to achieve accurate classification and regression of the target.
It effectively improves the detection performance of remote sensing image objects, improves the detection efficiency and accuracy of targets with tightly arranged and small-scale characteristics, and reduces missed and missed detection.
Smart Images

Figure CN116758263B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and further relates to a target detection method, specifically a remote sensing image target detection method based on multi-level feature fusion and joint positioning, which can be used to accurately detect remote sensing image targets with complex backgrounds. Background Art
[0002] With the rapid development of remote sensing technology, a large amount of remote sensing image data has been acquired. These remote sensing images taken by high-resolution remote sensing satellites contain very rich ground information, which can assist research in many fields, such as agricultural and forestry monitoring, marine surveillance, urban planning, natural resource management, environmental protection and military reconnaissance. Therefore, how to extract the information we need from these massive remote sensing images has become an urgent problem to be solved. Remote sensing image target detection is a very important research direction in the field of remote sensing image information extraction. Remote sensing image target detection has very wide applications in many military and civilian fields, such as military ship reconnaissance, weapon defense, precision guidance, environmental monitoring and other fields. It plays a very important role. Remote sensing image target detection technology is the key direction of development in the field of target detection.
[0003] In recent years, with the rapid development of computer vision technology and deep learning technology, a large number of target detection algorithms based on deep learning technology have shown excellent detection performance in natural image detection tasks. However, unlike natural images, the shooting method and shooting angle of remote sensing images are extremely special, and the imaging quality is easily affected by many factors such as atmospheric absorption, scattering and sensor parameters. In addition, the background of remote sensing image targets is complex, the direction is arbitrary, the scale is different, and there are many small-scale targets and they are closely arranged. Therefore, there have always been many difficulties and challenges in the detection of remote sensing image targets. However, with the widespread application of remote sensing image target detection technology, people are paying more and more attention to the research of remote sensing image target detection methods.
[0004] In 2019, Yang et al. proposed R3Det in the article [R3Det:Refined Single-Stage Detector withFeature Refinement for Rotating Object]. 3 Det, R 3 Det is an improved single-stage target detection network. The network is proposed mainly for the three characteristics of remote sensing image targets, namely large aspect ratio, dense arrangement and arbitrary direction. The network has the following improvements: 1) The rotation box can have better target detection accuracy in dense scenes, while the horizontal box can achieve better recall rate, so R 3Det uses a combination of these two forms of frames. Specifically, the first stage detects horizontal frames to improve detection speed and recall rate. The second stage, which is the refinement stage, detects rotation frames to adapt to the detection of dense targets. 2) In view of the feature misalignment of existing single-stage detectors, R 3 Det designed a feature refinement module that uses feature interpolation to obtain refined position information and reconstruct feature maps to achieve feature alignment. The module also reduces the number of refined bounding boxes and improves the detection speed; 3) In order to solve the problem of non-differentiable Skew IoU calculation, an approximate Skew IoU loss is designed to obtain more accurate rotation estimation. In 2022, researchers from Zhejiang University published an article [Oriented RepPoints for Aerial Object Detection] and proposed Oriented-RepPoints. This network is improved and innovated on the basis of RepPoints. Specifically, three directional transformation functions are improved on the basis of RepPoints to facilitate more accurate directional classification and positioning. At the same time, improvements are made to the weak supervision of key points by RepPoints, and an effective quality assessment and sample allocation scheme for adaptive points is proposed for the supervision of adaptive points to select representative Oriented RepPoints samples during training. The scheme can obtain non-axis-aligned features from adjacent targets or background noise, and introduce spatial constraints to punish abnormal points in the adaptive point learning process.
[0005] Since the development of remote sensing image target detection algorithms, many excellent algorithms have emerged, and good detection effects have been achieved on public remote sensing image datasets. However, due to the many particularities of remote sensing images and remote sensing image targets, how to accurately extract the features of targets in any direction of remote sensing images, how to detect targets with large scale differences in remote sensing images at multiple scales, how to accurately detect densely arranged targets, how to accurately detect small-scale targets in complex background environments, and how to balance the relationship between detection accuracy and detection speed are still the key research directions in the current field of remote sensing image target detection. Summary of the invention
[0006] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a method for anchorless target detection in remote sensing images based on multi-level feature fusion and coarse-fine joint positioning, so as to solve the problem that the existing detection methods have difficulty in extracting features for remote sensing image targets with characteristics such as variable shapes, arbitrary directions, close arrangements and multiple small scales, and are prone to lose a large amount of useful information, resulting in missed detection and false detection, leading to the inability to accurately detect targets. First, the target features are extracted and fused through the feature extraction network and the feature fusion network to generate a feature map with multi-scale feature information, and then the multi-scale feature map is input into the detection head network. In the detection head network, the target is accurately classified and regressed through coarse-fine joint positioning; the present invention can effectively improve the performance of remote sensing image target detection.
[0007] The present invention achieves the above-mentioned purpose by the following specific steps:
[0008] (1) The ResNet50 network is used as the backbone network, and the remote sensing image is input into the backbone network. Through multiple sets of convolution and pooling operations, multi-scale feature maps of different resolutions are output in the last four layers of the ResNet50 network to obtain the remote sensing image feature maps {C2, C3, C4, C5};
[0009] (2) Improve the FPN network by using 3×3 deformable convolution to replace the 1×1 lateral connection in the FPN network, and input the feature map {C3, C4, C5} in the remote sensing image feature map into the improved FPN network for feature fusion to generate the fused three-layer feature map {N3, N4, N5};
[0010] (3) The fused three-layer feature map {N3, N4, N5} is input into a multi-level feature fusion module containing multiple convolutional layers and pooling layers, and convolution and pooling operations are performed on the input feature maps to extract and fuse features. Specifically, the feature maps of the same scale are used as the basic feature maps, and the rest are auxiliary feature maps. They are up-sampled and down-sampled by bilinear interpolation and average pooling to generate feature maps {M3, M4, M5} with multi-scale information;
[0011] (4) The convolution pair containing two convolution kernels is used as a shared convolution network, and the shared convolution network is used to replace the BatchNorm and Non-linear structures in the coordinate attention mechanism. At the same time, the ELU activation function is used to replace the ReLU activation function in the coordinate attention mechanism to obtain an improved balanced coordinate attention mechanism module; the feature map with multi-scale information is input into the improved balanced coordinate attention mechanism module to obtain feature maps of different scales;
[0012] (5) The feature maps of different scales are input into the detection head network Head for detection. The network includes a coarse positioning module and a refined positioning module, which are used together to classify and regress the target. The coarse positioning module is used to obtain the coarse prediction box and coarse classification score of the target. The refined positioning module then performs feature alignment on the coarse prediction box and uses the coarse prediction box and coarse classification score for auxiliary regression to complete the classification and regression of the target and obtain the final detection result.
[0013] Compared with the prior art, the present invention has the following advantages:
[0014] First, since the present invention adopts a multi-level feature fusion module, the problem that a single-layer feature map does not contain different-scale features of other feature maps of the target is solved by fusing the feature maps of each layer of the feature pyramid network (FPN), which is more conducive to the detection of targets with multi-scale features in remote sensing images.
[0015] Second, since the balanced coordinate attention mechanism module is designed in the present invention, by assigning different weights to different feature parts in the fused multi-scale feature map, the feature map can assign greater weights to important feature information when performing detection, thereby improving the detection accuracy of remote sensing image targets.
[0016] Third, the present invention designs a coarse-fine joint positioning method in the detection head network, specifically, by designing a coarse positioning module and a refined positioning module to achieve accurate classification and regression of remote sensing targets; the coarse positioning module predicts the distance to the left, top, right and bottom of the ground-truth box at each positive sample point on the feature map, that is, the predicted distance vector t pt =(l, t, r, b), a coarse prediction box at this positive sample point is generated, and the probability that this coarse prediction box belongs to the target is generated. Here, in order to accurately predict targets with similar appearance characteristics, the distance vector of the coarse prediction box is multiplied by 1.5, and then a coarse prediction box containing some context information is generated. This approach introduces some background information to accurately detect targets with similar appearance. The refined positioning module aligns the axis-aligned convolution feature map with the two coarse prediction boxes respectively by using a feature alignment convolution called AlignConv, and in the final classification and regression, the coarse classification score and coarse prediction box generated by the coarse prediction module are used for auxiliary regression to speed up the speed and accuracy of the final bounding box prediction, thereby improving the target detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flowchart of the implementation process of the method of the present invention;
[0018] Figure 2It is a schematic diagram of the overall network structure of the method of the present invention;
[0019] Figure 3 It is a schematic diagram of the structure of the multi-level feature fusion module of the method of the present invention;
[0020] Figure 4 Schematic diagram of the structure of the balanced coordinate attention mechanism module of the method of the present invention;
[0021] Figure 5 It is a schematic diagram of the structure of the rough positioning module of the method of the present invention;
[0022] Figure 6 It is a structural schematic diagram of a refined positioning module of the method of the present invention;
[0023] Figure 7 This is a partial visualization detection result diagram of the method of the present invention on the DOTAv1.0 data set;
[0024] Figure 8 This is a partial visualization of the detection results of the method of the present invention on the HRSC2016 dataset. DETAILED DESCRIPTION
[0025] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0026] Refer to the attached Figure 1 , Figure 2 In the present invention, MFF represents a multi-level feature fusion module; BCA represents a balanced coordinate attention mechanism module; CPM represents a coarse positioning module; and RPM represents a refined positioning module.
[0027] Some symbols in the drawings of the specification are explained as follows: Conv represents ordinary convolution, 3×3 represents the size of the convolution kernel; Dconv represents deformable convolution; H and W represent the height and width of the feature map respectively; C represents the total number of target categories; c represents the number of channels of the feature map; ReLU and Sigmoid represent two activation functions.
[0028] In view of some problems existing in the prior art, the present invention proposes a remote sensing image target detection method based on multi-level feature fusion and joint positioning. Figure 1 is a flowchart of the implementation process of the method of the present invention, Figure 2 This is a schematic diagram of the overall network structure of the method of the present invention. A multi-scale feature map is generated by fusing multi-level feature fusion modules. The weights of different features in the fused feature map are assigned according to the balanced coordinate attention mechanism to highlight the more important features. Finally, a coarse-fine joint positioning method is designed in the Head, and the feature maps of different scales are input into it for target classification and regression to obtain a more accurate detection result.
[0029] Embodiment 1: Refer to the attached Figure 1 The present invention proposes a remote sensing image target detection method based on multi-level feature fusion and joint positioning, which specifically includes the following steps:
[0030] Step 1. Use the ResNet50 network as the backbone network, input the remote sensing image into the backbone network, and output multi-scale feature maps of different resolutions in the last four layers of the ResNet50 network through multiple sets of convolution and pooling operations to obtain the remote sensing image feature maps {C2, C3, C4, C5};
[0031] Step 2. Improve the FPN network by using 3×3 deformable convolution to replace the 1×1 lateral connection in the FPN network, and input the feature map {C3, C4, C5} in the remote sensing image feature map into the improved FPN network for feature fusion to generate the fused three-layer feature map {N3, N4, N5};
[0032] Step 3. Input the fused three-layer feature map {N3, N4, N5} into a multi-level feature fusion module containing multiple convolutional layers and pooling layers, and perform convolution and pooling operations on the input feature map to extract and fuse features. Specifically, the feature maps of the same scale are used as the basic feature maps, and the rest are auxiliary feature maps. They are up-sampled and down-sampled by bilinear interpolation and average pooling to generate feature maps {M3, M4, M5} with multi-scale information. The implementation is as follows:
[0033] (3.1) Let i = 1, 2 or 3. When generating a feature map Mi with multi-scale information, one of the three fused feature maps is used as the basic feature map Ni, and the other two are auxiliary feature maps Ni. x and Ni y Specifically, Ni undergoes a 3×3 convolution, and the number of channels of the convolution output is set to be consistent with the number of channels of the original feature map to generate the feature map Ni'; Ni in the auxiliary feature map x First, the average pooling method is used to perform 2 times downsampling, and then a 3×3 convolution is performed. The number of output channels of the convolution is set to half of the number of channels of the original feature map to generate the feature map Ni x '; Ni in the auxiliary feature map y First, a 3×3 convolution operation is performed, where the number of output channels of the convolution is set to half of the number of input channels, and then the convolution result is upsampled by 2 times using bilinear interpolation to generate the feature map Ni y ';
[0034] (3.2) The characteristic graphs Ni', Ni x 'And Ni y 'Connect along the channel axis and fuse to generate a multi-scale feature map Mi with 512 channels.
[0035] Step 4. Use the convolution pair containing two convolution kernels as a shared convolution network, use the shared convolution network to replace the BatchNorm and Non-linear structures in the coordinate attention mechanism, and use the ELU activation function to replace the ReLU activation function in the coordinate attention mechanism to obtain the improved balanced coordinate attention mechanism module; input the feature map with multi-scale information into the improved balanced coordinate attention mechanism module to obtain feature maps of different scales. The steps are as follows:
[0036] (4.1) Encode each channel along the horizontal and vertical coordinates respectively, and then aggregate the feature maps with multi-scale information along the horizontal and vertical spatial directions to obtain two direction-aware attention maps;
[0037] (4.2) Set the two convolution kernels contained in the shared convolutional network layer to 5×1, the convolution step size to (1,1), and the padding area to (2,0). Use the shared convolutional network layer to complete the convolution of the attention map and use the ELU activation function to activate it to obtain two activated attention maps.
[0038] (4.3) The two activated attention maps are reactivated using the Sigmoid activation function and multiplied with the input feature map to generate feature maps of different scales.
[0039] Step 5. Input the feature maps of different scales into the detection head network Head for detection, which includes a coarse positioning module and a refined positioning module. The coarse positioning module contains three prediction branches, namely, a category prediction branch, a distance prediction branch, and an angle prediction branch. The category prediction branch is used to generate the probability that the prediction box belongs to the target, the distance prediction branch is used to predict the distance from the positive sample point on the feature map to the left, top, right and bottom of the ground-truth, and the angle prediction branch is used to predict the target angle; the refined positioning module includes a feature alignment convolution AlignConv and two prediction branches, wherein the feature alignment convolution is used to align and fuse the coarse prediction box and the coarse context prediction box features, and the fused feature map is input into two prediction branches, which are a classification branch and a regression branch. The classification branch uses the coarse classification score generated by the coarse positioning module to perform category prediction, and the regression branch completes the regression operation according to the coarse prediction box generated by the coarse positioning module.
[0040] The coarse positioning module and the refined positioning module are used to jointly classify and regress the target. The coarse positioning module is used to obtain the coarse prediction frame and coarse classification score of the target. The refined positioning module then performs feature alignment on the coarse prediction frame and uses the coarse prediction frame and coarse classification score for auxiliary regression to complete the classification and regression of the target and obtain the final detection result. The specific implementation is as follows:
[0041] (5.1) The coarse localization module generates a rough first coarse prediction box and the probability score of the box belonging to the target at each positive sample point in the input feature map, then multiplies the distance vector of the first coarse prediction box by 1.5, and then generates a second coarse prediction box containing some context information;
[0042] (5.2) The refined positioning module uses the feature alignment convolution AlignConv to align the axis-aligned convolution feature map with the two coarse prediction boxes generated in step (5.2) to obtain two aligned feature maps, which are then merged into a final feature map for refined detection. At the same time, the coarse prediction box and coarse classification score are used for auxiliary regression to complete the classification and regression of the target and obtain the final detection result.
[0043] Embodiment 2: The target detection method proposed in this embodiment has the same overall implementation steps as Embodiment 1. Now, how to verify the accuracy of the detection result is further described as follows:
[0044] The loss is calculated based on the loss function for the final test result and the true target. If the loss is close to 1, it means that the test result is very close to the true result, and the test result is judged to be accurate and output. Otherwise, it means that the test result is significantly different from the true result, and the test result may be inaccurate. In this case, the test is judged to be inaccurate, and it is necessary to further optimize the parameters and update the model before retesting.
[0045] The loss that needs to be calculated here is the sum of the losses of the coarse positioning module and the refined positioning module. The losses generated by the coarse positioning module include category prediction loss, distance prediction loss, and angle prediction loss. Here, the category prediction loss uses the Focal loss loss function, and the distance prediction loss and angle prediction loss both use the smooth L1 loss loss function; the losses of the refined positioning module include classification loss and regression loss. The classification loss uses the Focal loss loss function, and the regression loss uses the Gaussian Wasserstein Distance Loss loss function.
[0046] The loss function expression of the coarse positioning module is as follows:
[0047]
[0048] Among them, L cls is the loss function corresponding to the category prediction branch, specifically using the Focal loss function, L reg and L angleThey are the loss functions corresponding to the distance prediction branch and the angle prediction branch, respectively. They both use the smooth L1 loss function, where N is the number of all points on the feature map, λ is the loss balance weight, and N pos Represents the number of positive sample points in the feature map, is a sample point in a feature map. When this sample point is a positive sample point, is 1, otherwise it is 0. and are the ground-truth real distance vector and angle, t and θ are the predicted distance vector and angle respectively;
[0049] The loss function expression of the refined positioning module is as follows:
[0050]
[0051] Among them, N pos is the number of positive sample points in the feature map, λ1 and λ2 are hyperparameters, both of which are 1, and b n is the predicted box, b gt is the ground-truth, c is the predicted category, c gt For the real category;
[0052] The total loss L is calculated according to the following formula:
[0053] L=L C +L R ;
[0054] Among them, L C and L R They represent the losses generated by the coarse positioning module and the refined positioning module respectively.
[0055] Embodiment 3: The overall implementation steps of the target detection method proposed in this embodiment are the same as those in Embodiment 1. Specific parameter settings are now given, and the process of implementing detection using the method of the present invention is further described in detail as follows:
[0056] Step a. Extract the features of the input remote sensing image. The backbone network uses the ResNet50 network. The remote sensing image with a resolution of 1024×1024 and a channel number of 3 is input into the backbone network. Through multiple sets of convolution and pooling operations, the last four layers of the ResNet50 network output multi-scale feature maps {C2, C3, C4, C5} with different resolutions, as shown in Figure 2 As shown in Figure 2, the resolutions of the multi-scale feature map groups C2, C3, C4, and C5 are respectively one-quarter, one-eighth, one-sixteenth, and one-thirty-second of the input remote sensing image, and the number of channels are 256, 512, 1024, and 2048, respectively. Figure 1The c in refers to the number of channels of the feature map.
[0057] Step b. Since the backbone network extracts too many feature channels and contains a lot of redundant information, the FPN network is used for feature fusion. The 3×3 deformable convolution Dconv is used to replace the original 1×1 lateral connection method in FPN. The feature maps C3, C4 and C5 are used as the feature maps of the input FPN. The multi-scale feature maps are fused by using 2x upsampling and 3×3 deformable convolution lateral connection to obtain the fused feature maps N3, N4 and N5.
[0058] Step c. Input the feature maps of the three layers N3, N4 and N5 into the multi-level feature fusion module to perform multi-level feature fusion to generate a feature map with multi-scale information, such as Figure 3 As shown in the figure, it is the internal structure diagram of the multi-level feature fusion module when fusion generates M4, specifically:
[0059] (c-1): When generating the M4 feature map, the MFF module regards N4 as the basic feature map and N3 and N5 as auxiliary feature maps. The basic feature map is used to construct the main features of M4, and the auxiliary feature map provides supplementary features to make up for the shortcomings of N4 in multi-scale target detection in remote sensing images. Specifically, N4 first undergoes a 3×3 convolution to further extract features. The number of output channels of this convolution is set to be consistent with the number of channels of the original feature map. The resolution and number of channels of the feature map N4′ obtained after the convolution are consistent with those of N4. Then N3 is first downsampled by 2 times. The downsampling method here adopts the average pooling method. After 2 times downsampling, the resolution becomes half of the original, and the number of channels remains unchanged. Then it undergoes a 3×3 convolution. This convolution is different from the above convolution. The number of output channels of the convolution needs to be set to half of the original number of channels. Therefore, the resolution and number of channels of the feature map N3′ after the convolution are both half of N3. For feature map N5, a 3×3 convolution operation is first performed. The number of output channels in this convolution also needs to be set to half of the number of input channels. Then the convolution result is upsampled by a factor of 2. The 2-fold upsampling here uses bilinear interpolation. After upsampling, the resolution of feature map N5′ is twice that of the input feature map, and the number of channels is half of the number of channels of the input feature map.
[0060] (c-2): The feature maps N3′, N4′ and N5′ are connected along the channel axis and fused to generate a multi-scale feature map M4 with 512 channels. This feature map contains the features of the target at different scales in the feature maps of each layer and can be well used for the detection of targets with multi-scale characteristics.
[0061] Step d. Since the fused M4 contains features of different scales of the target in different feature maps, the problem that still exists is that the features of different parts of M4 will show different importance for the detection of targets of different scales. The fused feature map still mainly needs the features in its basic feature map to detect the target. Therefore, the fused feature map is input into the balanced coordinate attention mechanism module to adjust the weights of different feature parts in the fused feature map, and give greater weights to important feature parts for more accurate prediction, such as Figure 4 As shown, specifically:
[0062] (d-1): First, the feature map is input into two average pooling layers with kernels (H, 1) and (1, W), encoding each channel along the horizontal and vertical coordinates respectively, aggregating features along two spatial directions, and aggregating the input feature map into two direction-aware attention maps;
[0063] (d-2): The dimension of the attention map of size C×1×W is transposed to facilitate the subsequent convolution operation. The two attention maps are then input into the shared convolutional network layer. The shared convolutional network layer contains two convolutions with a kernel size of 5×1. To ensure that the resolution of the attention map does not change after the convolution, the convolution step size is set to (1,1) and the padding area is set to (2,0). Figure 4 The r in is set to 4 to control the reduction rate of the channel. In this way, the attention map is convolved. After the convolution, the attention map can rely on the information in its corresponding row or column and the information around the corresponding row or column to calculate the importance of each row or column in the input feature map. In the shared convolutional network layer, the ELU activation function is used to solve the possible gradient explosion problem and promote gradient propagation.
[0064] (d-3): Finally, after the attention map with dimension C×W×1 is transformed and the two attention maps are activated by the Sigmoid activation function, both attention maps are multiplied with the input feature map to emphasize the different importance of different features in the input feature map, thereby generating a feature map with more accurate information.
[0065] Step e. Input the feature maps of different scales into the Head for target classification and regression. A coarse-fine joint positioning method is designed in the Head. In essence, the coarse positioning module and the refined positioning module are used to jointly classify and regress the target. The details are as follows:
[0066] (e-1): The feature map is first input into the coarse positioning module, such as Figure 5As shown in the figure, the coarse positioning module generates a coarse prediction box at each positive sample point of the input feature map. For the selection of positive sample points on the feature map, all ground-truths are first assigned to feature maps of different levels in the feature pyramid according to their size, which means that feature maps of different levels only need to be responsible for predicting the ground-truth assigned to them, without having to care about other ground-truths. Figure 2 The multi-scale feature maps of the five levels {P3, P4, P5, P6, P7} have step sizes {s3, s4, s5, s6, s7} of 8, 16, 32, 64 and 128 respectively, and the size of the definition belongs to The ground-truth within the range will be assigned to P i Feature map, where α is the magnification factor of the size range. Here, the minimum ground-truth size assigned to P3 is set to 0, and the maximum ground-truth size assigned to P7 is set to 10000, so that all ground-truths of different sizes can be included. After assigning each ground-truth to the feature map of its corresponding level, the feature map is then mapped back to the input image. If the point on the feature map is located in the center area of the ground-truth it is responsible for predicting, these feature points are recorded as positive sample points, and the other feature points are recorded as negative sample points. At each positive sample point, the distance vector t from the left, top, right, and bottom of the ground-truth is predicted. pt =(l, t, r, b), thereby generating a coarse prediction box coarse boxes at each positive sample point, and generating the probability score coarse cls that this coarse prediction box belongs to the target. In order to accurately detect some targets with similar appearance, the distance vector of the coarse prediction box is multiplied by 1.5 here to generate a coarse prediction box cosrse contextual boxes containing some context information. Cosrse contextual boxes accurately identify targets with similar appearance by introducing appropriate context and background information. At the same time, if the coarse boxes cannot surround the target well, the cosrse contextual boxes will better surround the target and have a better effect because its distance vector is enlarged by 1.5 times on the basis of the coarse boxes.
[0067] (e-2): Since the coarse prediction boxes generated by the coarse localization module are different at different positions, there is a feature misalignment between the feature maps generated by the axis-aligned convolution operation, which leads to a serious inconsistency between the classification score and the localization accuracy finally generated by the model. Therefore, in the refined localization module, the axis-aligned convolution feature map is first aligned with the two coarse prediction boxes using a feature alignment convolution called AlignConv, as shown in the following example: Figure 6 As shown in the figure, the two aligned feature maps are then fused and sent to the subsequent network for refined detection. When the classification and regression are finally performed, the fine classification score predicted by the classification feature map and the coarse cls generated by the coarse positioning module are multiplied, and the result is used as the final classification score. This classification score is jointly generated by the coarse positioning module and the refined positioning module, and has a higher classification proposal value. When performing bounding box regression prediction, the coarseboxes predicted by the coarse positioning module are mapped to the final regression feature map, and then the coarse boxes whose centers are within the ground-truth range are used as the initial positive samples of the ground-truth. The IoU value between the ground-truth and its corresponding initial positive sample is calculated, and then the IoU is sorted, and the coarse boxes whose IoU value is greater than the set threshold are selected as the final positive samples. Next, these final positive samples are used for regression prediction. By using the coarse boxes as the anchor boxes of the final regression feature map, a good initial value is given to the final bounding box regression, thereby speeding up the final bounding box prediction and improving the target detection performance.
[0068] Step f. Design of loss function. The total loss includes the loss generated by the coarse positioning module and the loss generated by the final classification and regression of the refined positioning module. Therefore, the total loss function includes the sum of the loss functions of the coarse positioning module and the refined positioning module, as follows:
[0069] (f-1): The coarse positioning module includes three prediction branches: category prediction, distance prediction, and angle prediction. Therefore, the loss of the coarse positioning module is L C It consists of three parts:
[0070]
[0071] Among them, L cls is the loss function corresponding to the category prediction branch, specifically using the Focal loss function, L reg and L angleThey are the loss functions corresponding to the distance prediction branch and the angle prediction branch, respectively. They both use the smooth L1 loss function, where N is the number of all points on the feature map, λ is the loss balance weight, and N pos Represents the number of positive sample points in the feature map, is a sample point in a feature map. When this sample point is a positive sample point, is 1, otherwise it is 0. and are the ground-truth real distance vector and angle, t and θ are the predicted distance vector and angle respectively;
[0072] (f-2): Loss L of the refined localization module R It consists of two parts, namely classification loss and regression loss. The classification loss also uses the Focal loss loss function, and the regression loss uses the Gaussian Wasserstein DistanceLoss loss function. This loss function can avoid the discontinuity problem of the rotation angle regression interval and the quadrilateral problem, reducing the learning difficulty of the model. The loss function expression is shown in the following formula:
[0073]
[0074] Among them, N pos is the number of positive sample points in the feature map, λ1 and λ2 are hyperparameters, both of which are 1, and b n is the predicted box, b gt is the ground-truth, c is the predicted category, c gt For the real category;
[0075] (f-3): The total loss function L = L C +L R .
[0076] In view of the difficulty in extracting features of remote sensing targets with variable shapes, directions, sizes and multi-scale characteristics, the present invention abandons the method of pre-defining anchor frames and directly predicts the regression frame of the target at each positive sample point in the feature map; designs a multi-level feature fusion module and a balanced coordinate attention mechanism module to efficiently extract and fuse target features, and finally generates a feature map with multi-scale feature information of the target; at the same time, in view of the problem of missed detection in the detection process of targets with dense arrangement and small-scale characteristics, a coarse-fine joint positioning method is designed, including a coarse positioning module and a refined positioning module. The two modules jointly complete the classification and regression of the target and generate accurate detection results. The present invention effectively improves the detection efficiency of targets with dense arrangement and small-scale characteristics and improves the detection performance.
[0077] The effect of the present invention is further described below in conjunction with simulation experiments.
[0078] 1. Experimental conditions:
[0079] The hardware platform of the experiment of the present invention is: Intel(R) Core(TM) processor @4.9GHz, 16G memory, NVIDIAGeForce RTX 3090Ti graphics card, 24G memory;
[0080] The software platform of the experiment of the present invention is: Ubuntu18.04 operating system, the integrated development platform is Pycharm2021, and the deep learning framework used is Pytorch, version 1.12.1.
[0081] 2. Experimental content:
[0082] Using the public remote sensing image datasets DOTAv1.0 and HRSC2016, comparative experiments are conducted with the current mainstream remote sensing image target detection algorithms to compare the performance differences.
[0083] 3. Evaluation indicators:
[0084] The present invention uses mAP as the final indicator for evaluating the performance of the target detection algorithm. mAP is the mean of AP of all target categories. AP refers to the average accuracy of a certain type of target, which is the value of the area enclosed by the PR curve of a certain type of target and the coordinate axis. If the enclosed area is larger, it means that its average accuracy is higher, and if it is smaller, it means that its average accuracy is lower. Among them, P refers to the precision of the predicted target category, and R refers to the recall rate of the predicted target category. The calculation formulas are as follows:
[0085]
[0086]
[0087] Among them, the number of targets detected as positive and detected correctly is recorded as TP, the number of targets detected as positive but detected incorrectly is recorded as FP, and the number of targets detected as negative and detected incorrectly is recorded as FN.
[0088] 4. Experimental results and analysis:
[0089] Table 1 is a table showing the experimental results of the method of the present invention compared with the current mainstream anchor-frame-based remote sensing image target detection algorithm on the DOTAv1.0 dataset. It can be seen from Table 1 that the method of the present invention has the highest detection accuracy compared with the current mainstream anchor-frame-based remote sensing image target detection algorithm, and has excellent detection performance. 2The detection accuracy of the A-Net network is still 0.05 percentage points higher. Now, as a remote sensing image target detection algorithm based on anchor-free frames, the method of the present invention has achieved higher detection accuracy than the detection algorithm based on anchor frames. At the same time, the network structure is simple, and the calculation amount and parameters are small.
[0090] Table 1 Experimental results of this method compared with the Anchor-based method on the DOTAv1.0 dataset
[0091]
[0092] Table 2 is a table showing the experimental results of the method of the present invention compared with the current mainstream remote sensing image target detection algorithm based on no anchor frame on the DOTAv1.0 dataset. It can be seen from Table 2 that compared with the current mainstream remote sensing image target detection algorithm based on no anchor frame, the method of the present invention has the highest detection accuracy, which is higher than all the methods in the figure and 1.97 percentage points higher than the detection accuracy of SCPNet. This shows that compared with most remote sensing image target detection algorithms based on no anchor frame, the method of the present invention has the best detection performance and is highly competitive.
[0093] Table 2 Experimental results of this method compared with the Anchor-free method on the DOTAv1.0 dataset
[0094]
[0095] Figure 7 This is a visualized detection result diagram of the method of the present invention for some test set data in the DOTAv1.0 dataset. It can be seen from the detection result diagram that the method of the present invention has a good detection effect on targets with large aspect ratios, close arrangement, small scales and arbitrary directions, and can accurately detect all targets without missed detections or false detections, proving that the method of the present invention has good practicality.
[0096] Table 3 is a table showing the experimental results of the method of the present invention compared with the current mainstream remote sensing image target detection algorithm on the HRSC2016 dataset. It can be seen from Table 3 that among all the methods in the comparison table, the detection accuracy of the method of the present invention is the highest, which is 0.16 percentage points higher than that of RSDet, indicating that the method of the present invention has excellent detection performance. Figure 8 This is a visualization detection result diagram of part of the test set data in the HRSC2016 dataset using the method of the present invention. From the detection results, it can be seen that the method of the present invention has a good detection effect on ship targets and can be used in the field of military ship reconnaissance.
[0097] Table 3 Comparative experimental results of this method and mainstream methods on the HRSC2016 dataset
[0098]
[0099] The above simulation analysis proves the correctness and effectiveness of the method proposed in the present invention.
[0100] Parts of the present invention that are not described in detail belong to common knowledge among those skilled in the art.
[0101] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Obviously, for professionals in this field, after understanding the content and principles of the present invention, they may make various modifications and changes in form and details without departing from the principles and structures of the present invention. However, these modifications and changes based on the ideas of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A remote sensing image target detection method based on multi-level feature fusion and joint positioning, characterized in that: The multi-level feature fusion module is used to fuse the features of each layer of the feature pyramid network FPN, and then the balanced coordinate attention mechanism module is used to assign weights to the fused feature maps. Finally, the coarse-fine joint positioning method is used to classify and regress the target, thereby improving the target detection performance of remote sensing images. The following steps are included: (1) The ResNet50 network is used as the backbone network, and the remote sensing image is input into the backbone network. Through multiple sets of convolution and pooling operations, multi-scale feature maps of different resolutions are output in the last four layers of the ResNet50 network to obtain the remote sensing image feature maps {C2, C3, C4, C5}; (2) Improve the FPN network by using 3×3 deformable convolution to replace the 1×1 lateral connection in the FPN network, and input the feature map {C3, C4, C5} in the remote sensing image feature map into the improved FPN network for feature fusion to generate the fused three-layer feature map {N3, N4, N5}; (3) The fused three-layer feature map {N3, N4, N5} is input into a multi-level feature fusion module containing multiple convolutional layers and pooling layers, and convolution and pooling operations are performed on the input feature maps to extract and fuse features. Specifically, the feature maps of the same scale are used as the basic feature maps, and the rest are auxiliary feature maps. They are up-sampled and down-sampled by bilinear interpolation and average pooling to generate feature maps {M3, M4, M5} with multi-scale information; (4) The convolution pair containing two convolution kernels is used as a shared convolution network, and the shared convolution network is used to replace the BatchNorm and Non-linear structures in the coordinate attention mechanism. At the same time, the ELU activation function is used to replace the ReLU activation function in the coordinate attention mechanism to obtain an improved balanced coordinate attention mechanism module; The feature map with multi-scale information is input into the improved balanced coordinate attention mechanism module to obtain feature maps of different scales; (5) Inputting feature maps of different scales into the detection head network Head for detection, the network includes a coarse positioning module and a refined positioning module, through which the target is classified and regressed; The coarse localization module is used to obtain the rough prediction box and rough classification score of the target. The refined localization module then performs feature alignment on the coarse prediction box and uses the coarse prediction box and rough classification score for auxiliary regression to complete the classification and regression of the target and obtain the final detection result.
2. The method according to claim 1, characterized in that: The feature map {M3, M4, M5} with multi-scale information in step (3) is obtained according to the following steps: (3.1) Let i = 1, 2 or 3. When generating a feature map Mi with multi-scale information, one of the three fused feature maps is used as the basic feature map Ni, and the other two are auxiliary feature maps Ni. x and Ni y Specifically, Ni undergoes a 3×3 convolution, and the number of channels of the convolution output is set to be consistent with the number of channels of the original feature map to generate the feature map Ni'; Ni in the auxiliary feature map x First, the average pooling method is used to perform 2 times downsampling, and then a 3×3 convolution is performed. The number of output channels of the convolution is set to half of the number of channels of the original feature map to generate the feature map Ni x '; Ni in the auxiliary feature map y First, a 3×3 convolution operation is performed, where the number of output channels of the convolution is set to half of the number of input channels, and then the convolution result is upsampled by 2 times using bilinear interpolation to generate the feature map Ni y '; (3.2) The characteristic graphs Ni', Ni x 'And Ni y 'Connect along the channel axis and fuse to generate a multi-scale feature map Mi with 512 channels.
3. The method according to claim 1, characterized in that: The feature maps of different scales in step (4) are obtained according to the following steps: (4.1) Encode each channel along the horizontal and vertical coordinates respectively, and then aggregate the feature maps with multi-scale information along the horizontal and vertical spatial directions to obtain two direction-aware attention maps; (4.2) Set the two convolution kernels contained in the shared convolutional network layer to 5×1, the convolution step size to (1,1), and the padding area to (2,0). Use the shared convolutional network layer to complete the convolution of the attention map and use the ELU activation function to activate it to obtain two activated attention maps. (4.3) The two activated attention maps are reactivated using the Sigmoid activation function and multiplied with the input feature map to generate feature maps of different scales.
4. The method according to claim 1, characterized in that: The structures of the coarse positioning module and the refined positioning module in step (5) are as follows: The coarse positioning module includes three prediction branches, namely, a category prediction branch, a distance prediction branch, and an angle prediction branch. The category prediction branch is used to generate the probability that the prediction box belongs to the target. The distance prediction branch is used to predict the distance from the positive sample point on the feature map to the left, top, right, and bottom of the ground-truth. The angle prediction branch is used to predict the target angle. The refined positioning module includes a feature alignment convolution AlignConv and two prediction branches, wherein the feature alignment convolution is used to align and fuse the features of the coarse prediction box and the coarse context prediction box. The fused feature map is input into two prediction branches, which are a classification branch and a regression branch. The classification branch uses the coarse classification score generated by the coarse positioning module to perform category prediction, and the regression branch completes the regression operation according to the coarse prediction box generated by the coarse positioning module.
5. The method according to claim 1, characterized in that: The step (5) is specifically implemented as follows: (5.1) The coarse localization module generates a rough first coarse prediction box and the probability score of the box belonging to the target at each positive sample point in the input feature map, then multiplies the distance vector of the first coarse prediction box by 1.5, and then generates a second coarse prediction box containing some context information; (5.2) The refined positioning module uses the feature alignment convolution AlignConv to align the axis-aligned convolution feature map with the two coarse prediction boxes generated in step (5.2) to obtain two aligned feature maps, which are then merged into a final feature map for refined detection. At the same time, the coarse prediction box and coarse classification score are used for auxiliary regression to complete the classification and regression of the target and obtain the final detection result.
6. The method according to claim 1, characterized in that: The final detection result in step (5) is compared with the true target to calculate the loss according to the loss function. If the loss is close to 1, it means that the detection result is accurate and the detection result is output; otherwise, it is judged as inaccurate.
7. The method according to claim 6, characterized in that: The loss is the sum of the losses of the coarse positioning module and the refined positioning module. The loss generated by the coarse positioning module includes category prediction loss, distance prediction loss and angle prediction loss. Here, the category prediction loss uses the Focal loss loss function, and the distance prediction loss and the angle prediction loss both use the smooth L1 loss loss function; the loss of the refined positioning module includes classification loss and regression loss. The classification loss uses the Focal loss loss function, and the regression loss uses the Gaussian Wasserstein Distance Loss loss function.
8. The method according to claim 7, characterized in that: The loss function expression of the coarse positioning module is as follows: Among them, L cls is the loss function corresponding to the category prediction branch, specifically using the Focal loss function, L reg and L angle They are the loss functions corresponding to the distance prediction branch and the angle prediction branch, respectively. They both use the smooth L1 loss function, where N is the number of all points on the feature map, λ is the loss balance weight, and N pos Represents the number of positive sample points in the feature map, is a sample point in a feature map. When this sample point is a positive sample point, is 1, otherwise it is 0. and are the ground-truth real distance vector and angle, t and θ are the predicted distance vector and angle respectively; The loss function expression of the refined positioning module is as follows: Among them, N pos is the number of positive sample points in the feature map, λ1 and λ2 are hyperparameters, both of which are 1, and b n is the predicted box, b gt is the ground-truth, c is the predicted category, c gt For the real category; The total loss L is calculated according to the following formula: L=L C +L R ; Among them, L C and L R They represent the losses generated by the coarse positioning module and the refined positioning module respectively.
Citation Information
Patent Citations
Multi-scale target detection method based on joint recursive feature pyramid
CN115527095A
Method for dim and small object detection based on discriminant feature of video satellite data
US20220067335A1