A Fine-Grained Oriented Small Target Detection Method

By introducing parallel segmentation branches into the small object detection model, using the refined segmentation task and Focal Loss loss function, the problem of fine-grained features in small object detection in sea surface and coastal scenarios is solved, which improves detection performance and reduces inference costs.

CN115587627BActive Publication Date: 2025-06-27NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211271767.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2025-06-27
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

In the prior art, in small object detection tasks in sea surface and coastal scenarios, it is difficult to effectively retain the fine-grained features of the image, resulting in poor detection performance.

Method used

By introducing a parallel segmentation branch in the training phase of the model, fine-grained features of the network learning input image are extracted using the refined segmentation task-guided features. This segmentation branch generates pseudo-labels, and optimizes the feature extraction network output through the Focal Loss loss function.

Benefits of technology

The model's ability to capture fine-grained features is improved, and the detection performance of small targets in sea surface and coastal scenarios is enhanced, without increasing model inference time and hardware cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115587627B_ABST
    Figure CN115587627B_ABST
Patent Text Reader

Abstract

The present invention provides a fine-grained oriented small target detection method, comprising the following steps: preprocessing a small target detection data set, and cropping training images by using a sliding window; generating segmentation branch pseudo-labels according to the position information of annotation boxes; performing data augmentation on the cropped small target images, and then respectively inputting them into a feature extraction network model; inputting the output feature matrix into a small target detection branch and a segmentation branch; training the small target detection branch and the segmentation branch in parallel and optimizing them independently; until the feature extraction network model converges and the training phase ends; in the testing phase, removing the segmentation branch, and the segmentation branch does not participate in the model inference process. In the network training process of the present invention, a segmentation branch is newly added to guide the feature extraction network to learn the fine-grained features of the input image, and a segmentation branch pseudo-label is designed to eliminate the influence of background noise, and the small target detection ability of the model is specifically improved without increasing the model inference calculation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural networks, and particularly relates to a fine-grained oriented small target detection method. Background Art

[0002] Marine search and rescue includes ship salvage, rescue of distressed ships, rescue of distressed persons, etc. Compared with the general object detection task, the small target detection task in the sea surface and coastal scenarios has higher requirements for fine-grained features. The images processed by the general object detection task have a lower resolution, and the target size is generally larger than 32×32 pixels. However, the small target detection task in the sea surface and coastal scenarios needs to process images with a higher resolution, and the small target size is generally in the range of 2×2 to 32×32 pixels. The optimal parameters of the general object detection model are not applicable to small target detection and cannot be directly applied to small target detection, and the model migration span is large.

[0003] Mainstream target detection models can be divided into anchor-based detection models and non-anchor-based detection models from the perspective of detection basis. An anchor can be understood as a candidate box set artificially. The anchor-based detection model, that is, on the basis of the candidate box set artificially, performs boundary optimization and classification. The most representative detection models include Faster RCNN and SSD, which have occupied the dominant position in the field of target detection research since their publication, and most subsequent research is based on the optimization and modification of these two studies. The non-anchor-based detection model directly predicts the key points of an object, such as the center point, corner points, etc., and generates an object detection box based on the key points. In recent years, it has received more and more attention in the field of target detection. Compared with the anchor-based detection model, it requires fewer hyperparameters, less human intervention, and less redundant calculation.

[0004] To address the problem that small targets are difficult to locate, the following methods are usually adopted in the field of small target detection to retain the fine-grained features of images and optimize the detection performance of the model, including: image super-resolution processing, context feature fusion. The basis for using super-resolution to perform data augmentation on images is that small targets often lack clarity due to the small number of pixels they contain. If only linear interpolation is used to magnify small targets, effective detail information cannot be obtained. Using a super-resolution model for learning can, to a certain extent, restore the fine-grained features of small targets. The context feature fusion method is to make up for the loss of fine-grained features caused by network downsampling, reuse the fine-grained features output by the shallow layer of the network, and fuse the low-resolution feature map containing deep semantic information and the high-resolution feature map containing shallow detail information.

[0005] In the object detection task, the CenterNet detection model in the article "Objects as Points" published by Zhou Xingyi et al. on arxiv in 2019, in order to meet the needs of predicting the center points of objects and maintain a high resolution, through deconvolution operations, magnifies the feature map output by the feature extraction network to 1 / 16 of the size of the input image. However, there are still two major problems for the small object detection task in the sea surface and coastal scenarios. First, all small object regions of Tiny1 size (i.e., the square root of the bounding box area is in the range of [2, 8]) in the sea surface image are mapped to a single pixel on the output feature map, and there is a certain loss of accuracy. Second, there is a significant decrease in the resolution of the feature map output by the feature extraction network at the initial stage, resulting in the lack of fine-grained features. Summary of the Invention

[0006] In view of the technical problems existing in the prior art, from the perspective of feature optimization, the present invention provides a fine-grained oriented small object detection method. Through a more refined segmentation task, during the training stage of the model, it guides the feature extraction network to learn the fine-grained features of the input image, thereby improving the small object detection index.

[0007] The technical solution adopted by the present invention is: a fine-grained oriented small object detection method, including the following steps:

[0008] Step 1: Preprocess the small object detection dataset, and use a sliding window to crop the training images in the dataset;

[0009] Step 2: Generate segmentation branch pseudo-labels according to the position information of the annotation boxes in the dataset;

[0010] Step 3: Perform data augmentation on the cropped small object images, and then input them into the feature extraction network model respectively to obtain a feature matrix with the dimensions of H×W×C, where H and W represent the height and width of the feature matrix, and C represents the number of channels of the feature matrix;

[0011] Step 4: Input the feature matrix output by the feature extraction network model into the small object detection branch and the segmentation branch;

[0012] Step 5: The small object detection branch and the segmentation branch are trained in parallel and optimized independently;

[0013] Among them, for the small object detection branch: in the forward propagation stage, the small object detection branch outputs the detection box of the area where the small object is located, and finally shows the coordinates of the upper left and lower right vertices of the rectangular box; the gradient of the loss layer of the small object detection branch is backpropagated to update the parameters of the detection branch and the feature extraction network, and optimize the position of the detection box.

[0014] Segmentation branch: In the forward propagation stage, the segmentation branch predicts the probability that each pixel on the feature map output by the feature extraction network is inside the region where the small target is located; the gradient of the segmentation branch loss layer is backpropagated to guide the feature extraction network model to learn fine-grained features;

[0015] Step 6: Repeat Step 3 - Step 5 until the feature extraction network model converges and the training stage ends;

[0016] Step 7: In the testing stage, remove the segmentation branch. The segmentation branch does not participate in the model inference process, thus ensuring that while enhancing the model's fine-grained feature extraction ability, the model inference time and hardware cost are not increased. The main role of the segmentation branch is to guide the learning of the feature extraction network and prompt the feature extraction network to focus on the fine-grained features in the input image. After the model training is completed, the segmentation result is no longer needed in the testing stage, and the segmentation branch can be deleted. Therefore, it is possible to improve the model's ability to capture fine-grained features without increasing the cost in the testing stage, and further improve the small target detection performance in the sea surface and coastal scenarios.

[0017] Furthermore, considering that taking all the pixel points inside the entire labeled box area as positive samples of the segmentation branch will inevitably introduce more background noise. In response to this situation, a segmentation branch pseudo-label is designed to screen the pixel point range, which can, on the premise of introducing as little background noise as possible, guide the feature extraction network to extract the fine-grained features in the input image through a refined task. In Step 2, the generation strategy of the segmentation branch pseudo-label is as follows: Considering that the sizes of small targets in the sea surface and coastal areas are small, and the size of the feature map output by the feature extraction sub-network is 1 / 4 of the input image, for the labeled box with a length or width in the pixel interval of [2, 8], all the pixel points inside the labeled box are set as positive examples without radiation range screening. For the labeled boxes of other sizes, with the center of gravity of the labeled box as the center point, set the pixel points within the radiation range from the center point to the edge of the labeled box as positive samples. The radiation range is: an elliptical area with the center of gravity of the labeled box as the center point, the major axis being rl∈[0.5, 1] times the long side of the labeled box, and the minor axis being rs∈[0.5, 1] times the short side of the labeled box. The elliptical radiation area is dynamically adjusted according to the size of the labeled box without loss of generality. The hyperparameters rl and rs are screened through ablation experiments.

[0018] Furthermore, in Step 5, the pixel points outside the radiation range of the segmentation branch pseudo-label generated in Step 2 and inside the labeled box do not participate in the gradient backpropagation of the segmentation branch loss layer. Considering that the position of such pixel points relative to the location of the small target has a high degree of uncertainty, making them not participate in the backpropagation can effectively eliminate the influence brought by background noise.

[0019] Furthermore, the segmentation branch loss uses the Focal Loss loss function, which is suitable for the current task where the positive and negative samples are unbalanced due to the high image resolution and small target size. The formula of the loss function is:

[0020]

[0021] Among them, α is the category balance factor, γ is the difficult sample balance factor, which is used to weaken the impact of category imbalance in model training, enhance the learning of difficult samples, and avoid falling into the local maximum; the probability that the input pixel point is predicted to belong to the small target area The probability that a pixel belongs to the background area Pixel label t ij ∈{0, 1}, 0 means the pixel label is background, 1 means the pixel label is pedestrian, i, j represents each position on the output prediction map.

[0022] Furthermore, in step 5, the small target detection branch loss and the segmentation branch loss are fused with a certain weight, where the segmentation branch loss weight is set to 0.8 and the small target detection branch loss weight is set to 0.2.

[0023] Furthermore, in step 3, data enhancement includes: random scaling, random inversion, random cropping, color space transformation, and mean and variance subtraction.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] 1. Based on the original detection branch of the model, the present invention adds a parallel and independent segmentation branch, and guides the feature extraction network to learn the fine-grained features in the input image through the refined segmentation task, thereby improving the model detection performance and adapting to most existing detection models.

[0026] 2. Since the small target detection dataset in sea and coastal scenes does not provide segmentation labels, the present invention designs a segmentation branch pseudo-label generation strategy based on the data characteristics of this task, while minimizing the influence of background noise and improving the accuracy of the segmentation branch supervision information.

[0027] 3. The segmentation branch of the present invention does not participate in the reasoning process of the model. It specifically improves the model's ability to detect small targets without increasing any model reasoning calculation cost or affecting the actual deployment of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a flow chart of an embodiment of the present invention;

[0029] Figure 2 A schematic diagram of a pseudo-label generation strategy for segmentation branches according to an embodiment of the present invention;

[0030] Figure 3 It is the structure diagram of the segmentation branch model according to the embodiment of the present invention;

[0031] Figure 4 It is the schematic diagram of the small target detection effect according to the embodiment of the present invention. Specific implementation manner

[0032] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0033] The embodiment of the present invention provides a fine-grained oriented small target detection method, as Figure 1 shown, which includes the following steps:

[0034] Step 1: Preprocess the small target detection data set, divide it into a training set and a test set, and use a sliding window to crop the training images in the training set. The overlap between two sliding windows along the current direction is 50 pixels. Finally, the size of all cropped small target images is 640 in width and 512 in height. The small target detection data set is the TinyPerson public data set.

[0035] Step 2: Generate segmentation branch pseudo-labels according to the position information of the annotation boxes in the data set. The generation strategy of the segmentation branch pseudo-labels is as follows: for the annotation boxes in the data set with a length or width in the pixel interval of [2, 8], set all pixel points within the annotation box as positive examples, as Figure 2 (a) shown. For the annotation boxes of the remaining sizes, generate an elliptical region with the center of gravity of the annotation box as the center point, the long axis being rl times the long side of the annotation box, and the short axis being rs times the short side of the annotation box. The pixel points within this elliptical region are positive sample points, where both rl and rs are set to 1.0, as Figure 2 (b) shown. The pixel points within the annotation box range but not within this elliptical region are identified as training-irrelevant points and do not participate in the loss function calculation and backpropagation of the newly added segmentation branch, so as to minimize the interference of background noise as much as possible. The elliptical region is dynamically adjusted according to the annotation box size.

[0036] Step 3: Perform data augmentation on the cropped small target images, and then input them into the feature extraction network model respectively. The data augmentation includes: random scaling, random inversion, random cropping, color space transformation, subtracting the mean and dividing by the variance; among them, random cropping is to crop a region with a width of 640 and a height of 512 on the randomly scaled image. If the size of the scaled image is smaller than this size, it is filled with 0 for expansion. Each small target image after data augmentation is input into the feature extraction network model, and the output feature matrix X ∈ R 256×320×512 .

[0037] Step 4: Input the feature matrix output by the feature extraction network model into the small object detection branch and the segmentation branch; where the number of output channels of the segmentation branch is 1.

[0038] Step 5: The small object detection branch and the segmentation branch are trained in parallel and optimized independently; during the model training process, each iteration includes: one forward propagation and one backward propagation for both the small object detection branch and the segmentation branch.

[0039] Among them, for the small object detection branch: in the forward propagation stage, the small object detection branch outputs the detection box of the area where the small object is located, and finally shows the coordinates of the upper left and lower right vertices of the rectangular box. The gradient of the loss layer of the small object detection branch is backpropagated to update the parameters of the detection branch and the feature extraction network, and optimize the position coordinates of the detection box.

[0040] Segmentation branch: In the forward propagation stage, the segmentation branch predicts the probability that each pixel point is inside the area where the small object is located. The segmentation branch predicts the probability that each pixel point on the feature map output by the feature extraction network is inside the area where the small object is located; the structural diagram of the segmentation branch model is as Figure 3 shown. Taking the feature matrix X output by the feature extraction network as the input, passing through a 3×3 convolution with 64 output channels (accompanied by a ReLU activation layer), and a 1×1 convolution with 1 output channel, the segmentation matrix

[0041] S∈R 1×H′×W′ is output, and then through an element-wise Sigmoid operation, it is mapped to the interval (0, 1), corresponding to the probability of each pixel point within the small object area range.

[0042] The gradient of the loss layer of the segmentation branch is backpropagated to guide the feature extraction network model to learn fine-grained features. The loss of the segmentation branch uses the Focal Loss function:

[0043]

[0044] where α is the class balance factor, γ is the difficult sample balance factor, and α and γ are set to 1.0 and 2.0 respectively; the probability that the input pixel point is predicted to belong to the small object area the probability that the pixel point belongs to the background area the pixel point label t ij ∈{0, 1}, 0 means the pixel point label is the background, 1 means the pixel point label is a pedestrian, and i, j represent each position on the output prediction map.

[0045] The loss of the segmentation branch and the loss of the detection branch are fused with a certain weight to form the training loss of the entire model. Among them, the loss weight of the segmentation branch is set to 0.8, and the loss weight of the small object detection branch is set to 0.2.

[0046] Step 6: Repeat Step 3 - Step 5 until the feature extraction network model converges, and the training phase ends.

[0047] Step 7: In the testing phase, remove the segmentation branch, and the segmentation branch does not participate in the model inference process. Visualize and evaluate the detection effect on the test set.

[0048] The detection effect of small targets in this embodiment is as Figure 4 shown. Figure 4 (a) and Figure 4 (b), the gray rectangular boxes are the detection results of the baseline network model, and the black rectangular boxes are the detection results after fine-grained feature optimization. The dashed boxes are used for auxiliary observation. Figure 4 (c) and Figure 4 (d), the black rectangular boxes are the annotation boxes.

[0049] On the TinyPerson dataset, a set of comparative experiments are conducted based on ResNet18, ResNet34, and ResNet50 respectively. To optimize the detection effect of small targets, according to the existing method, an upsampling module is added to the baseline detection network model, and fine-grained features that reuse the shallow output of the feature extraction network are used, abbreviated as shortConnect. Then, using this method, a segmentation branch is added to shortConnect to assist the feature extraction network in learning fine-grained features, abbreviated as Ellipseseg.

[0050] A comparative experiment is conducted on the TinyPerson dataset using three ResNet series feature extraction networks. The experimental results are shown in Table 1 - Table 3. It can be seen from the table that after adding Ellipseseg, the detection indicators of Tiny-sized targets in the TinyPerson dataset have been improved. Ellipseseg promotes the feature extraction network to learn fine-grained features, thereby increasing the accuracy and recall rate of detecting Tiny-sized targets. Based on ResNet18, the TinyPerson test set and are improved by 2.25 and 2.03 percentage points respectively. Based on ResNet34, the TinyPerson test set and are improved by 1.37 and 1.38 percentage points respectively. Based on ResNet50, the TinyPerson test set and are improved by 0.74 and 0.42 percentage points respectively.

[0051] Table 1 Comparison results of AP of the TinyPerson test set based on the ResNet18 feature extraction network

[0052]

[0053] Table 2 Comparison results of AP of TinyPerson test set based on ResNet34 feature extraction network

[0054]

[0055] Table 3 Comparison results of AP of TinyPerson test set based on ResNet50 feature extraction network

[0056]

[0057] The present invention has been described in detail through the embodiments above. However, the content described above is only an exemplary embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. The protection scope of the present invention is defined by the claims. All those who use the technical solutions described in the present invention, or those skilled in the art who are inspired by the technical solutions of the present invention, within the essence and protection scope of the present invention, design similar technical solutions to achieve the above technical effects, or make equivalent changes and improvements to the application scope, etc., should still fall within the scope of patent coverage protection of the present invention.

Claims

1. A fine-grained oriented small object detection method, characterized in that: It includes the following steps: Step 1: Preprocess the small object detection dataset, and crop the training images in the dataset using a sliding window; Step 2: Generate segmentation branch pseudo-labels according to the position information of the annotation boxes in the dataset. The generation strategy of the segmentation branch pseudo-labels is: taking the center of gravity of the annotation box as the center point, setting radiation from the center point to the edge of the annotation box, and the pixel points within the radiation range are positive samples; Step 3: Perform data augmentation on the cropped small object images, and then input them into the feature extraction network model respectively; Step 4: Input the feature matrix output by the feature extraction network model into the small object detection branch and the segmentation branch; Step 5: The small object detection branch and the segmentation branch are trained in parallel and optimized independently; Among them, for the small object detection branch: in the forward propagation stage, the small object detection branch outputs the detection box of the area where the small object is located; the gradient of the loss layer of the small object detection branch propagates backward to update the parameters and optimize the position of the detection box; For the segmentation branch: in the forward propagation stage, the segmentation branch predicts the probability that each pixel point is inside the area where the small object is located; the gradient of the loss layer of the segmentation branch propagates backward to guide the feature extraction network model to learn fine-grained features; Step 6: Repeat Step 3 - Step 5 until the feature extraction network model converges and the training stage ends; Step 7: In the test stage, remove the segmentation branch, and the segmentation branch does not participate in the model inference process.

2. The fine-grained oriented small target detection method according to claim 1, characterized in that: The radiation range is: an elliptical area with the center of gravity of the annotation box as the center point, the major axis being [0.5, 1] times the long side of the annotation box, and the minor axis being [0.5, 1] times the short side of the annotation box.

3. A fine-grained oriented small target detection method according to claim 1, characterized in that: For the annotation box with the length or width in the pixel interval of [2, 8], set all the pixel points inside the annotation box as positive examples without radiation range screening.

4. A fine-grained oriented small target detection method according to claim 1, characterized in that: In Step 5, the pixel points outside the radiation range of the segmentation branch pseudo-labels generated in Step 2 and inside the annotation box do not participate in the gradient backpropagation of the segmentation branch loss layer.

5. A fine-grained oriented small target detection method according to claim 1, characterized in that: The segmentation branch loss uses the Focal Loss function: ; Among them, is the class balance factor, is the difficult sample balance factor; the probability that the input pixel point is predicted to belong to the small target area is , and the probability that the pixel point belongs to the background area is ; the pixel point label is , 0 indicates that the pixel point label is the background, 1 indicates that the pixel point label is a pedestrian, and i, j represent each position on the output prediction map.

6. A fine-grained oriented small target detection method according to claim 1 or 5, characterized in that: In Step 5, the small object detection branch loss and the segmentation branch loss are fused with certain weights, where the weight of the segmentation branch loss is set to 0.8, and the weight of the small object detection branch loss is set to 0.

2.

7. The fine-grained oriented small target detection method according to claim 1, wherein: In Step 3, the data augmentation includes: random scaling, random inversion, random cropping, color space transformation, subtracting the mean and dividing by the variance.

Citation Information

Patent Citations

  • Parallel method of object detection and semantic segmentation based on end-to-end depth learning

    CN109543754A

  • Multi-task parallel method and system for target detection and semantic segmentation

    CN111680739A