Target detection method based on partially fused convolutional network

By adopting the object detection method based on partially fusion convolutional network in remote sensing images, an object detection model including partially fusion convolutional network and multi-scale cross-space feature fusion network is constructed, which solves the problem of low detection accuracy of small and medium-sized objects in the prior art, and achieves higher detection accuracy and detailed information capture.

CN120219951APending Publication Date: 2025-06-27XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510254236.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has low accuracy in remote sensing images, making it difficult to effectively process small-sized targets in complex backgrounds, and cross-level residual blocks of the backbone network are easily disturbed by noise, making it difficult to capture the detailed characteristics of the small targets.

Method used

The object detection method based on partially fusion convolutional network is adopted, and the object detection model includes a partially fusion convolutional network backbone network, a multi-scale cross-space feature fusion network neck network and a detection head network are constructed. The model extracts different deep features through multiple partial fusion convolution modules, and dynamically balances the detailed information and semantic features of different levels through multi-scale cross-space feature fusion networks.

Benefits of technology

It improves the perception of the characteristics of the key areas of the channel and space, suppresses noise interference from complex backgrounds, effectively captures the detailed information of small targets, and significantly improves the detection accuracy of remote sensing small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219951A_ABST
    Figure CN120219951A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method based on a partially fused convolutional network. The target detection method comprises the following implementation steps: acquiring a training and testing sample set; constructing a target detection network model based on the partially fused convolutional network and performing iterative training on the target detection network model; and obtaining a remote sensing small target detection result. According to the invention, a plurality of partial fusion convolution modules perform feature extraction of different deep layers on shallow layer features extracted by a convolution feature extraction module, so that the perception capability on channel and space key region features is improved, and noise interference of a complex background in a remote sensing image is suppressed through a two-stage ReLU activation layer. The multi-scale cross space feature fusion network fuses key semantic information and different deep features after channel feature enhancement by selecting a feature fusion module, so that the detail information and semantic information of different level features are dynamically balanced; and the detection accuracy of the remote sensing small target is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and relates to an object detection method based on a partial fusion convolutional network, which can be applied to fields such as geographical object detection and maritime ship detection. Background Art

[0002] Due to the complex imaging environment of remote sensing images, the images usually contain environmental noises such as cloud occlusion, ground object shadows, and seasonal changes. In particular, the similar features between dense building clusters and small objects in urban areas are likely to cause false detections. Since the detection targets in remote sensing images have significant multi-scale characteristics, small-sized targets only occupy a few pixels in remote sensing images, and their characteristic information is extremely easy to be submerged in the complex background. In addition, remote sensing targets often exhibit characteristics such as arbitrary directions, dense arrangements, and partial occlusions, making it difficult to effectively process such special-shaped targets, and the accuracy of object detection is relatively low.

[0003] In order to overcome the missed detections caused by environmental noises in remote sensing images and the impact of the lack of detailed information of small targets on the detection accuracy, for example, Chongqing University of Posts and Telecommunications disclosed a remote sensing object detection method in its patent document "A YOLOV4 Remote Sensing Object Detection Method Combining Feature Transfer and Attention Mechanism" (Patent Application No.: 202211078264.X, Publication No. CN115497005A). This invention preprocesses remote sensing image data through the Mosaic data augmentation method; constructs a YOLOV4 remote sensing object detection model integrating feature transfer and attention mechanism; inputs the remote sensing data into the model for training; obtains the remote sensing image to be detected, preprocesses the remote sensing image to a unified size; inputs the processed remote sensing image into the trained object detection model for detection, and outputs the detection result, that is, the bounding box position and object category of the remote sensing object in the image to be detected. Through YOLOV4 integrating feature transfer and attention mechanism, this invention improves the detection accuracy of objects without increasing the number of model parameters. However, the cross-stage residual blocks of the backbone network it uses are vulnerable to noise interference from the complex background of remote sensing images when extracting features of key regions, and it is difficult to capture the detailed features of small targets; moreover, the feature transfer module it uses is difficult to dynamically balance the detailed information and semantic features at different levels, resulting in relatively low accuracy of the model in remote sensing small object detection. Summary of the Invention

[0004] The purpose of the present invention is to overcome the defects existing in the above-mentioned prior art, and propose an object detection method based on a partial fusion convolutional network to solve the technical problem of relatively low detection accuracy existing in the prior art.

[0005] To achieve the above purpose, the technical solution adopted by the present invention includes the following steps:

[0006] (1) Obtain the training sample set and the test sample set:

[0007] Preprocess N remote sensing small target images including multiple small target categories, and label the bounding boxes of the small targets in more than half of the remote sensing images of each category after preprocessing. Then, use the labeled K preprocessed remote sensing images and their label boxes as the training sample set, and use the remaining N - K preprocessed remote sensing images as the test sample set, where N ≥ 1000;

[0008] (2) Construct the object detection network model O based on the partial fusion convolutional network:

[0009] Construct an object detection model O including a partial fusion convolutional network as the backbone network, a multi-scale cross-space feature fusion network as the neck network, and a detection head network; where the backbone network includes a convolutional feature extraction module for extracting shallow features and S cascaded partial fusion convolutional feature extraction modules for extracting different deep features; the neck network includes three branch networks arranged in parallel for enhancing key semantic information of different deep features; the detection head network includes detection heads for detecting the two-way fused features output by the neck network, where S ≥ 4;

[0010] (3) Iteratively train the object detection model O:

[0011] Iteratively train the object detection network model O through the training sample set to obtain the trained object detection network model O * ;

[0012] (4) Obtain the remote sensing small target detection results:

[0013] Use the test sample set as the input of the trained object detection network model O * to perform forward propagation to obtain the object detection results corresponding to N - K test samples.

[0014] Compared with the prior art, the present invention has the following advantages: The feature transfer module is difficult to dynamically balance the detailed information and semantic features at different levels, resulting in a still low accuracy of the model in remote sensing small target detection.

[0015] (1) In the process of iteratively training the object detection model and obtaining the remote sensing small target detection results of the present invention, multiple partial fusion convolutional modules extract features at different depths from the shallow features extracted by the convolutional feature extraction module, improving the perception ability of the features in the key regions of channels and spaces, and suppressing the noise interference of the complex background in the remote sensing images through two-level ReLU activation layers, which is beneficial to capturing the detailed information of small targets. Compared with the prior art, the detection accuracy of remote sensing small targets is effectively improved.

[0016] (2) The multi-scale cross-space feature fusion network adopted by the neck network of the present invention fuses different deep features after enhancing key semantic information and channel features through a selected feature fusion module, dynamically balancing the detailed information and semantic information of features at different levels, and further improving the accuracy of remote sensing small target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart for the implementation of the present invention.

[0018] Figure 2 It is a schematic structural diagram of the target detection network of the present invention.

[0019] Figure 3 It is a schematic structural diagram of a partial fusion convolution feature extraction module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0021] Refer to Figure 1 , the present invention includes the following steps:

[0022] Step 1) Obtain a training sample set and a test sample set:

[0023] (1a) Obtain N remote sensing small target images including multiple small target categories in the SIMD dataset, perform horizontal or vertical mirror flipping on each image, then enlarge or reduce the flipped image. In this embodiment, it is set to 640×640. Finally, perform Gaussian blur and normalization on the scaled image to obtain N preprocessed remote sensing small target images.

[0024] (1b) Perform border annotation on the small targets in more than half of the remote sensing images of each category after preprocessing, and then use the K preprocessed remote sensing images and their labels as the training sample set, and use the remaining N-K preprocessed remote sensing images as the test sample set. In this embodiment, N = 3000 and K = 2000;

[0025] Step 2) Construct a target detection network model O based on a partial fusion convolution network, and its structure is as Figure 2 shown:

[0026] Construct an object detection model O that includes a partial fusion convolutional network as the backbone network, a multi-scale cross-space feature fusion network as the neck network, and a detection head network; where the backbone network includes a convolutional feature extraction module for extracting shallow features and S cascaded partial fusion convolutional feature extraction modules for extracting different deep features; the neck network includes three branch networks arranged in parallel for enhancing key semantic information of different deep features; the detection head network includes detection heads for respectively detecting two paths of fused features output by the neck network. In this embodiment, S = 4;

[0027] The backbone network, the convolutional feature extraction module it contains includes M stacked 3×3 convolutional layers, batch normalization layers, and ReLU activation layers; the structure of the partial fusion convolutional feature extraction module it contains is as Figure 3 shown, including stacked 3×3 convolutional layers, a feature extraction layer for partial fusion convolution, batch normalization layers, a first ReLU activation layer, two-level 1×1 convolutional layers, and a second ReLU activation layer, and the input and output of the 3×3 convolutional layer are respectively connected to the output of the second ReLU activation layer and the second 1×1 convolutional layer by residual connections. In this embodiment, M = 3;

[0028] The neck network, the three branch networks it contains are respectively connected to the last three partial fusion convolutional feature extraction modules of the backbone network. These three branch networks all include cascaded fusion channel attention modules and 1×1 convolutional layers, and the output of the 1×1 convolutional layer in the first and second branch networks is also cascaded with a selective feature fusion module. The output of the 1×1 convolutional layer in the second and third branch networks is respectively connected to the selective feature fusion module in the first and second branch networks;

[0029] The detection head network, the two detection heads it contains are respectively connected to the output of the first and second branch networks.

[0030] Step 3) Iteratively train the object detection model O:

[0031] (3a) Initialize the iteration number as q, the maximum iteration number as Q, Q ≥ 500, and the weight parameters of the object detection network model O q in the q-th iteration are θ q , and let q = 0. In this embodiment, Q = 500;

[0032] (3b) Take A training samples randomly selected from the training sample set as the input of the object detection network model O. The convolutional feature extraction module in the backbone network extracts shallow features for each training sample; S partial fusion convolutional feature extraction modules perform depth-by-depth feature extraction on the extracted shallow features G to obtain S different deep features;

[0033] Among them, the shallow features G extracted are subjected to feature extraction depth by depth. The specific steps are as follows:

[0034] (3b1) The 3×3 convolutional layer in the S partial fusion convolutional feature extraction modules performs initial feature extraction depth by depth on the extracted shallow features G, obtaining S initially extracted features Z with different depths;

[0035] (3b2) The feature extraction layer of the partial fusion convolution, the batch normalization layer, the first ReLU activation layer, and the two-stage 1×1 convolutional layer perform deep feature extraction, normalization, non-linear activation, and channel dimension increase and decrease on the initially extracted feature Z in sequence, obtaining the deeply extracted feature Z′;

[0036] Among them, the specific steps for the feature extraction layer of the partial fusion convolution to perform deep feature extraction are as follows:

[0037] (3b21) The input feature map with the number of channels H×W×C is evenly divided into four feature maps F1, F2, F3, and F4 according to the number of channels C;

[0038] (3b22) The first feature map F1 obtains the extracted channel-dense feature F through standard convolution SC , and then F SC obtains the spatial-sparse feature F through depthwise separable convolution DSC ;

[0039] (3b23) The channel-dense feature F SC and the spatial-sparse feature F DSC are concatenated, and then the channel order is rearranged through shuffling, improving the channel and spatial perception ability, thereby enhancing the ability to extract key region features, obtaining the fused convolution feature F fus ;

[0040] (3b24) The second, third, and fourth feature maps F2, F3, and F4 after segmentation are concatenated with the fused convolution feature F fus , and then channel information fusion is performed through a 1×1 convolutional layer, obtaining the partial fusion convolution feature F out .

[0041] (3b3) The second ReLU activation layer performs non-linear activation on the residual operation result of the initially extracted feature Z and the deeply extracted feature Z′, suppressing the noise interference of the complex background in the remote sensing image, obtaining the deeply extracted feature Z″ containing more small target detail information;

[0042] (3b4) A residual operation is performed on the shallow feature G extracted by the convolutional feature extraction module and Z″, obtaining S features V with different depths *, the depth features focus on the extraction of high-level semantic information layer by layer, improving the channel and spatial perception ability, enhancing the ability to extract features of key regions, and suppressing the noise interference of complex backgrounds in remote sensing images through two-level ReLU activation layers.

[0043] (3c) The three fusion channel attention modules in the neck network and their cascaded 1×1 convolutional layers respectively enhance the key semantic information of the deep features output by the S-2nd, S-1st, and Sth partial fusion convolutional feature extraction modules, and then perform channel feature enhancement to obtain the enhanced features U S-2 、U S-1 、U S , and the selective feature fusion modules in the first and second branch networks fuse the features of U S-2 and U S-1 、U S-1 and U S to obtain the fusion features J1 and J2;

[0044] The selective feature fusion modules in the first and second branch networks fuse the features of U S-2 and U S-1 、U S-1 and U S The specific steps are as follows:

[0045] (3c1) The transposed convolutional module and bilinear interpolation module of the selective feature fusion module expand and upsample the enhanced features U S-2 、U S-1 to obtain the upsampled features U S ′ -2 、U S ′ -1 ;

[0046] (3c2) The channel fusion attention of the selective feature fusion module enhances the key semantic information of the sampled features U S ′ -2 、U S ′ -1 to obtain the enhanced features U S ″ -2 、U S ″ -1 ;

[0047] (3c3) The selective feature fusion module fuses U S ″ -2 with U S ′ -2 and U S-1 、U S ″ -1 with U S ′ -1 and U SPerform feature fusion to obtain the fused feature J1 = U S-1 ×U S ″ -2 +U S ′ -2 、J2 = U S ×U S ″ -1 +U S ′ -1 The hybrid upsampling method of transposed convolution and bilinear interpolation adopted improves the high-level semantic feature extraction ability, and the fused channel attention module is used to dynamically balance the detailed information and semantic information of different-level features during the fusion process.

[0048] (3d) Two detection heads perform convolution on J1 and J2 to obtain the predicted bounding box B of each training sample.

[0049] (3e) Adopt the L Inner-PIoU loss function, and calculate the loss value L Inner-PIoU of the object detection network model O through the predicted bounding box B and its corresponding labeled bounding box C. Then, adopt the stochastic gradient descent method to update the weight parameter θ Inner-PIoU through L q . The L Inner-PIoU loss function has the calculation formula:

[0050]

[0051]

[0052] where L IOU is the IOU loss function, r is the fusion weight, R λ is the loss function of the adjustment factor λ, P IOU is the penalty factor, ∩ and ∪ are the intersection and union operations respectively, and ln(·) is the logarithmic function.

[0053] Then, adopt the stochastic gradient descent method to update the weight parameter θ Inner-PIoU through L q to obtain the remote sensing small target detection network model O q of this iteration. The update formula is:

[0054]

[0055] where is the result of the update of θ q , η is the learning rate, g is the partial derivative of the loss function L Inner-PIoU with respect to θ q , is the partial derivative operator with respect to θ q .

[0056] (3f) Determine whether q > Q holds. If it holds, obtain the trained remote sensing small target detection network model O * , otherwise, q = q + 1, O q = O, and execute step (3b).

[0057] Step 4) Obtain the remote sensing small target detection result:

[0058] Use the test sample set as the input of the trained target detection network model O * to perform forward propagation and obtain the target detection results corresponding to N - K test samples.

Claims

1. A target detection method based on a partially fused convolutional network, characterized in that: The steps include: (1) Obtain training sample set and test sample set: Preprocess N remote sensing small target images including multiple small target categories, and annotate the small targets in more than half of the remote sensing images of each category after preprocessing, then use the annotated K preprocessed remote sensing images and their label frames as the training sample set, and use the remaining NK preprocessed remote sensing images as the test sample set, where N ≥ 1000; (2) Constructing an object detection network model O based on a partially fused convolutional network: Construct an object detection model O including a partially fused convolutional network as a backbone network, a multi-scale cross-spatial feature fusion network as a neck network and a detection head network; wherein the backbone network includes a convolutional feature extraction module for extracting shallow features and S cascaded partially fused convolutional feature extraction modules for extracting different deep features; the neck network includes three branch networks arranged in parallel for enhancing key semantic information of different deep features; the detection head network includes a detection head for detecting two fused features output by the neck network respectively, wherein S ≥ 4; (3) Iteratively train the target detection model O: The target detection network model O is iteratively trained through the training sample set to obtain the trained target detection network model O * ; (4) Obtain remote sensing small target detection results: The test sample set is used as the trained target detection network model O * The input is forward propagated to obtain the target detection results corresponding to NK test samples.

2. The method according to claim 1, characterized in that The preprocessing of N remote sensing small target images including multiple small target categories described in step (1) is implemented by: Each remote sensing small target image is mirror-flipped horizontally or vertically, and the flipped image is enlarged or reduced, and then the scaled image is Gaussian blurred and normalized to obtain N preprocessed remote sensing small target images.

3. The method according to claim 1, characterized in that The target detection network model O described in step (2), wherein: The backbone network includes a convolutional feature extraction module comprising M stacked 3×3 convolutional layers, a batch normalization layer, and a ReLU activation layer; the partially fused convolutional feature extraction modules include stacked 3×3 convolutional layers, a partially fused convolutional feature extraction layer, a batch normalization layer, a first ReLU activation layer, two levels of 1×1 convolutional layers, and a second ReLU activation layer, and the input and output of the 3×3 convolutional layer are respectively connected to the output residual of the second ReLU activation layer and the second 1×1 convolutional layer, where M≥3; The neck network, wherein the three branch networks contained therein are respectively connected to the last three partial fusion convolution feature extraction modules of the trunk network, and the three branch networks all include a cascaded fusion channel attention module and a 1×1 convolution layer, and the outputs of the 1×1 convolution layers in the first and second branch networks are also cascaded with a selection feature fusion module, and the outputs of the 1×1 convolution layers in the second and third branch networks are respectively connected to the selection feature fusion modules in the first and second branch networks; The detection head network comprises two detection heads connected to the outputs of the first and second branch networks respectively.

4. The method according to claim 3, characterized in that The iterative training of the target detection network model O described in step (3) is implemented as follows: (3a) Initialize the number of iterations to q, the maximum number of iterations to Q, Q ≥ 500, and the target detection network model O of the qth iteration q The weight parameter in is θ q , and set q = 0; (3b) A training samples randomly selected from the training sample set are used as the input of the target detection network model O. The convolutional feature extraction module in the backbone network extracts shallow features for each training sample; S partially fused convolutional feature extraction modules extract features of the extracted shallow features G in depth by depth to obtain S different deep features; (3c) The three fusion channel attention modules in the neck network and their cascaded 1×1 convolutional layers respectively enhance the key semantic information of the deep features output by the S-2, S-1, and S partial fusion convolutional feature extraction modules, and then enhance the channel features to obtain the enhanced features U S-2 , U S-1 , U S , the selection feature fusion module in the first and second branch networks has an effect on U S-2 with U S-1 , U S-1 with U S Perform feature fusion to obtain fused features J1 and J2; (3d) The two detection heads predict the fused features J1 and J2 to obtain the predicted bounding box B of each training sample; (3e) Using L Inner-PIoU Loss function, and calculate the loss value L of the target detection network model O by predicting the bounding box B and its corresponding label bounding box C Inner-PIoU , and then use the stochastic gradient descent method to pass L Inner-PIoU For the weight parameter θ q Update to obtain the remote sensing small target detection network model O of this iteration q ; (3f) Determine whether q>Q is true. If so, obtain the trained remote sensing small target detection network model O * , otherwise, q=q+1, O q =O, and execute step (3b).

5. The method according to claim 4, characterized in that The step (3b) of performing depth-by-depth feature extraction on the extracted shallow features G is implemented as follows: (3b1) The 3×3 convolutional layers in the S partially fused convolutional feature extraction modules perform preliminary feature extraction depth by depth on the extracted shallow features G to obtain S preliminary extracted features Z of different depths; (3b2) The feature extraction layer of the partially fused convolution, the batch normalization layer, the first ReLU activation layer, and the two-stage 1×1 convolution layer sequentially perform deep feature extraction, normalization, nonlinear activation, channel dimension increase and dimension reduction on the initially extracted features Z to obtain the deep extracted features Z′; (3b3) The second ReLU activation layer performs nonlinear activation on the residual operation results of the preliminary extracted features Z and the deep extracted features Z′ to obtain the activated deep extracted features Z″; (3b4) Perform residual operations on the shallow features G and Z″ extracted by the convolutional feature extraction module to obtain S different depth features V * .

6. The method according to claim 4, characterized in that The loss value L of the target detection network model O described in step (3e) Inner-PIoU , the calculation formula is: Among them, L IOU is the IOU loss function, r is the fusion weight, R λ is the loss function of the adjustment factor λ, P IOU is the penalty factor, ∩ and ∪ are the intersection and union operations respectively, and ln(·) is the logarithmic function.

7. The method according to claim 4, characterized in that The weight parameter θ described in step (3e) q Update, the update formula is: in, is θ q Updated result, η is the learning rate, g is the loss function L Inner-PIoU About θ q Find the partial derivative, is about θ q Partial derivative operator.

Citation Information

Patent Citations

  • YOLOV4 remote sensing target detection method fusing feature transfer and attention mechanism

    CN115497005A