Small target detection method based on scale perception label distribution and context enhancement

By introducing scale-aware label allocation strategy and context interaction feature enhancement module in small object detection technology, the problem of poor detection of extremely small object in the prior art is solved, and more efficient multi-scale object detection performance is achieved.

CN120147802APending Publication Date: 2025-06-13XIDIAN UNIV

Patent Information

Application Number
CN202510228601.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing small-objective detection technology performs poorly when dealing with extremely small targets. The traditional multi-scale detection network lacks feature extraction for extremely small targets, and the label allocation strategy may cause the network to over-optimize small targets, affecting the accuracy of large-objective detection.

Method used

A scale-aware label allocation strategy is adopted to allocate appropriate positive samples to targets of different scales through dynamic scale-aware IoU calculation formulas, and a context interaction feature enhancement module is introduced to enhance the feature response of small targets using a multi-scale self-attention mechanism.

Benefits of technology

It improves the accuracy and robustness of small-object detection, reduces the negative impact on the detection accuracy of other scales of targets, and improves the performance of multi-scale detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147802A_ABST
    Figure CN120147802A_ABST
Patent Text Reader

Abstract

The invention discloses a small target detection method based on scale perception label distribution and context enhancement. The method comprises the following steps of: 1, organizing an image file and a corresponding annotation file to obtain image data, and dividing the image data into a training image and a test image to form a required data set; 2, preprocessing and training the image data through feature extraction and small target detection capability of an optimization model, and storing the weight; and step 3, testing the test image according to the weight stored by training, verifying the small target detection capability of the model in a real scene through evaluation on a test set by using the weight stored by training, and adjusting or optimizing model parameters. According to the invention, sufficient positive samples can be reasonably distributed for the targets of different scales for training according to the target scale change, so that the influence on the detection precision of the targets of other scales is reduced on the basis of improving the small target detection precision; the problem that a traditional multi-scale detection network fails to challenge small targets is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of small target detection, and particularly relates to a small target detection method based on scale-aware label assignment and context enhancement. Background Art

[0002] Small target detection refers to the technology of locating and identifying objects with small areas, usually smaller than 16×16 pixels, in images. Due to their small size, unclear features, and being easily overlooked in complex backgrounds, detecting such targets poses challenges. Small target detection has important value in multiple practical applications, including building and vehicle monitoring in remote sensing images, pedestrian and small obstacle detection in autonomous driving, subtle defect detection in industrial automation, and wildlife monitoring in natural scenes, etc. This technology plays a crucial role in improving safety, automation level, and monitoring efficiency.

[0003] In the field of small target detection, more obvious feature responses and more reasonable label assignment strategies are crucial for improving the robustness of detection and have become a popular research direction. To address the challenges in small target detection, researchers have proposed various innovative methods. FPN and its extended models (such as PANet, BiFPN) improve detection performance through multi-scale feature fusion, especially showing outstanding performance on small targets. EFPN further enhances the detection effect on small targets through cross-resolution detail transfer. At the same time, deep super-resolution methods are used to improve image resolution to provide richer feature details of small targets. Haris et al. improved performance by jointly training super-resolution and the detection framework. Generative adversarial networks (GANs) have also been applied in this field, such as Perceptual GAN and MTGAN. The latter generates high-quality feature representations through multi-task learning, effectively improving the detection accuracy of small targets. In addition, label assignment strategies are crucial for small target detection. To address the sensitivity of traditional IoU-based assignment methods to small targets, novel strategies such as normalized Wasserstein distance and receptive field distance have been proposed to optimize positive sample assignment and improve the performance of detectors on small targets. However, their drawback is that they may lead to over-optimization of small targets in network training, thus affecting the balance of overall detection performance. Since small targets usually occupy a very small proportion in images, these methods may sacrifice the detection accuracy of large targets while ensuring the detection accuracy of small targets.

[0004] A target detection method and system for accurately detecting small objects based on deep learning with the publication number CN118053117A uses deep learning technology to improve the detection accuracy of small objects through multi-layer feature map extraction and innovative convolution strategies. However, since this patent still uses the traditional attention mechanism-based method, it is still necessary to finely tune the entire network parameters. On the one hand, calculating the attention weights for all feature layers results in a relatively large training cost. On the other hand, the failure to utilize the multi-head self-attention mechanism to obtain global-to-local feature information may lead to information redundancy and inaccurate features, so there is still much room for improvement.

[0005] A method for detecting aerial rotating small targets based on feature enhancement with the publication number CN117994493A introduces a method to improve the detection performance of small targets in aerial images, aiming to overcome the detection difficulties caused by small target sizes, complex backgrounds, and randomly arranged directions. However, this patent has the following two defects: (1) This patent still uses the traditional scale multi-scale detection network structure and does not design a specific feature optimization module for extremely small targets. Therefore, the deep network will still lose a lot of information in feature extraction for extremely small targets and cannot highlight the small target areas, resulting in suboptimal results. (2) Although the introduction of NWD loss to replace the traditional CIoU loss function can optimize the problem that small targets do not receive sufficient training to a certain extent, it will also affect the positive sample assignment of other larger-scale targets, thus affecting the overall multi-scale detection accuracy.

[0006] In summary, within the framework of the existing technology, although the development of deep neural networks has brought important progress in target detection, most existing multi-scale detection methods are designed for normally sized targets, and the performance of detecting small targets is still not satisfactory. Specifically, these technologies define small targets as targets with a labeled area less than 32×32. When designing a multi-scale network, they often only consider optimizing and enhancing the features of small targets under this definition. In actual applications, small targets are often targets with a pixel area less than 16×16, and the features of extremely small targets cannot be effectively enhanced under the original setting.

[0007] Since small objects are extremely sensitive to position deviations, even a slight deviation will cause a significant change in IoU (Intersection over Union). Therefore, under the IoU-based label assignment mechanism, it is difficult for the region proposal network to assign sufficient positive samples to small objects. At the same time, experimental results show that the previously existing improved label assignment schemes will have a negative impact on the detection performance of larger-scale targets. Summary of the Invention

[0008] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a small target detection method based on scale-aware label assignment and context enhancement. Through the designed scale-aware label assignment strategy, it can reasonably assign sufficient positive samples to targets of different scales for training according to the change of target scale, thereby reducing the impact on the detection accuracy of other scale targets while improving the detection accuracy of small targets; through the context interaction feature enhancement module, the problem that traditional multi-scale detection networks fail to handle small target challenges is effectively solved.

[0009] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0010] A small target detection method based on scale-aware label assignment and context enhancement, comprising the following steps;

[0011] Step 1: Image data preparation. Organize the image files and the corresponding annotation files to obtain image data, and divide the image data into training images and test images to form the required data set; provide necessary information support for training and testing;

[0012] Step 2:

[0013] Through effective preprocessing, feature extraction, label assignment, and loss calculation of the training images, ensure that the network can accurately identify, locate, and classify targets during the training process;

[0014] Optimize the prediction process from the image to the target box and save the weights;

[0015] Step 3: Test the test images with the weights saved through training. Through the evaluation on the test set, verify the ability of the model to detect small targets in the real scenario, and adjust or optimize the model parameters. In the above Step 1, the image file is the image data containing the target to be detected, stored in common image formats (such as JPEG, PNG, TIFF, etc.), each image contains one or more targets, and perform standardized preprocessing on images with different resolutions and sizes;

[0016] In the above Step 1, the annotation file contains the target position information in the image (such as bounding box, classification label), and is stored in text or binary format;

[0017] The annotation file includes the following contents:

[0018] Bounding box information: The position of the target in the image, represented by the upper left and lower right coordinates, or using the center point coordinates and width and height;

[0019] Category information: The category label of each target object;

[0020] Image Identifier: The annotation file should have a unique correspondence with the image file, and the annotation file name should be the same as the image file name.

[0021] Specifically, step 2 is as follows:

[0022] Step (2.1): Obtain the training image, and adjust the size of the image data to the standard input size in pixels; ensure that the input image maintains a fixed resolution to match the subsequent network structure.

[0023] Step (2.2): Preprocess the image. During the training process, first read the input training image data from the storage medium and ensure that the training image is loaded in RGB format.

[0024] Specifically, step (2.2) is as follows:

[0025] The loaded image will be adjusted to a fixed size set as needed, while maintaining the aspect ratio of the image width and height; to ensure that the input image adapts to the requirements of the detector and avoid geometric distortion. In addition, to enhance the robustness and generalization ability of the model, the image is randomly flipped horizontally with a probability of 50%.

[0026] The size of the processed image is adjusted to a multiple of 32 to adapt to the downsampling operation of the model, ensure the alignment of feature maps during the feature extraction process, and prepare the subsequent target box annotation information.

[0027] Step (2.3): Image feature extraction. After the input image is preprocessed in step (2.2), it is input into the feature extractor of the backbone network ResNet-50 for feature extraction.

[0028] The backbone network ResNet-50 is divided into four stages, and finally outputs 5 feature maps of different scales after feature downsampling in the fourth stage.

[0029] In step (2.3), the backbone network ResNet-50 is divided into four stages, corresponding to different feature layers respectively.

[0030] Each stage will output different feature maps, and the number of output channels of ResNet is 256, 512, 1024, and 2048 in sequence.

[0031] Each convolutional layer of ResNet-50 is regularized using batch normalization, and gradient updates are enabled to ensure the stability of the model during training. To further enhance the multi-scale representation ability of features, the output of the backbone network is fed into the Feature Pyramid Network (FPN). FPN fuses the features of these four stages and captures the spatial information of multiple targets through feature maps of different scales; the output channels of FPN are unified to 256, and finally 5 feature maps of different scales are output after downsampling the features of the fourth stage for subsequent object detection.

[0032] Step (2.4): Context Interaction Feature Enhancement Module: Since multi-scale detection algorithms usually detect small targets at feature levels with relatively large resolutions, the different channels of the feature map obtained in step (2.3) are subjected to shifted convolution to expand the receptive field, and channel splitting multi-scale self-attention calculation is performed. Finally, it is added and fused with the original feature map that has passed through the feature extraction network but not through the feature enhancement module to obtain the context content interaction enhanced feature.

[0033] The context content interaction enhanced feature is realized through the context interaction feature enhancement module, and the context interaction feature enhancement module only uses the highest-level feature pyramid feature layer as the input. Because in the paradigm of multi-scale detection, small targets are often detected on the high-resolution feature map of the highest level, and introducing multi-level feature inputs will also cause a heavy redundant computational burden.

[0034] The specific context content interaction enhancement is as follows:

[0035] First, the feature map is divided into K groups along the channel dimension;

[0036] Then, a window with a size of Mk×Mk is set in each segmented feature map for spatial partitioning, and self-attention calculation is performed within each window. By setting different window sizes in each segmented feature map, context information is captured at different scales. After completing the self-attention calculation at the window level, the segmented feature maps are concatenated and input into a 1×1 convolutional layer to obtain the final output. The complexity of this calculation process is as follows:

[0037]

[0038] Among them, M k represents the window size in the k-th feature partition. By adjusting the number of groups K and the window size M k .

[0039] Step (2.5) Dynamic Scale-Aware Label Assignment Strategy:

[0040] A set of candidate boxes are automatically generated from the one context-enhanced feature and the four original features output from step (2.4), which may contain the target object. Here, the improvement is mainly based on the region candidate network structure, and the ReRPNHead that introduces regression information is used as the core component of RPN (region proposal network), where the input feature channel and the internal feature channel are both set to 256;

[0041] This part includes an anchor generator, which is used to generate a series of candidate boxes on the image. On this basis, the dynamic scale-aware label assignment strategy reshapes the scale of each candidate box through the annotated scale supervision information and calculates the corresponding intersection-over-union ratio, and then jointly supervises the assignment of target positive and negative sample labels through the introduced regression information;

[0042] The step (2.5) is specifically:

[0043] The feature map output by step (2.4) is used as the input of RPN. RPN generates a priori anchor boxes on the input feature map. For each prior anchor box N to be assigned i , first match it to the gt(x,y,w,h) box with the closest spatial position, and introduce network regression parameters to match the gt box with the best spatial position for each prior anchor box to be assigned. First, calculate the regression parameter vector of the detection network regression branch and convert it into a coordinate offset pred to obtain a learnable regression prior anchor box:

[0044] reg_N i =pred(N i ) (2)

[0045] Then use the calculated *N i With gt i Calculate IoU, called reg_IoU, and use reg_IoU for each sample N to be assigned i Find the optimal spatial matching pair (N i ,gt i ), where N i and gt i The following relations are satisfied:

[0046] reg_IoU(N i ,gt i )=max(reg_IoU(N i ,GT)) (3)

[0047] Among them, GT is the set of GT boxes in the same training image. i With gt i , and its IoU calculation formula is:

[0048]

[0049] On this basis, the adaptive scale IoU formula based on spatial addressing is as follows:

[0050]

[0051] where, *N i is obtained by adaptively scaling the center sharing of *N i and its scale is strictly constrained by the gt matched in (3). Specifically, a constraint formula of the following form is designed:

[0052]

[0053] where size is a hyperparameter, and the aspect ratio of the initial anchor box is kept unchanged to maximize the original prior features. On the basis of *N i *N is calculated by scaling the coordinates through scale, i and *IoU is calculated through *N i and remapped back to *N i to assign it the gt with the most matching spatial position and scale;

[0054] Based on the *IoU calculated for each anchor box in the above process, after regression offset, reg_*IoU is calculated, and finally the dynamic scale-aware IoU calculation formula is as follows:

[0055]

[0056] where a is a scaling factor used to regulate the contribution degree of the regression parameter in the assignment strategy to adjust the optimal contribution values of the two.

[0057] Step (2.6): Forward propagation of the basic network. The features output in steps (2.3) and (2.4), including the enhanced features of the first layer and the original features of the remaining four layers, generate candidate regions through step (2.5) and perform training and prediction of positive and negative samples;

[0058] Step (2.7): Calculate the loss value between the predicted value in step 2.6 and the coordinate and class true annotation values of the corresponding target annotation box in the annotation file and perform backpropagation:

[0059] The specific content of the said step (2.7) is:

[0060] First, use the Cross-Entropy Loss function to calculate the loss between the predicted probabilities of the RPN and RoI heads for foreground and background classification and the true labels (foreground or background), and then use the L1 Loss to calculate the predicted value (x,y,w,h) and the true value (gtx,gty,gtw,gth)The loss values are weighted and summed up. The calculated loss value is used for backpropagation to update the network parameters. The training is looped multiple times until the network converges. During the update process, the weights of the base model are frozen and do not participate in the update.

[0061] Specifically, step 3 is as follows:

[0062] Step (3.1): Limit the longest side of the test image to a specified range and input it into the backbone network to extract multi-scale feature maps for use by the region proposal network.

[0063] Step (3.2): Candidate box generation and filtering: Anchor boxes are generated on each feature layer. The top 1000 candidate boxes are calculated based on the object probability, and the candidate boxes that cross the boundary and are too small are removed. Hierarchical non-maximum suppression is used to optimize the candidate boxes, and the final 1000 proposed boxes are reserved for subsequent steps.

[0064] Step (3.3): Final detection and class prediction: The candidate boxes are input into the ROI head for feature pooling and classification, and the class probability and bounding box regression parameters are output.

[0065] Advantages of the present invention:

[0066] (1) Compared with previous advanced technologies, the present invention more carefully designs a multi-scale detection network structure for small targets, calculates the self-similar information between different feature blocks using multi-scale windows, enabling small targets to be well distinguished from the background even in complex scenes. Then, a dynamic scale-aware label assignment strategy is designed, which not only assigns sufficient positive samples for small targets to train, but also maximally avoids the negative impact of small targets' optimized assignment on other scale targets, thereby achieving an improvement in multi-scale detection performance.

[0067] (2) The present invention proposes a context enhancement module to enhance the weak feature responses of small targets using rich context information. First, the receptive field of the input feature is enlarged through shifted convolution without increasing the computational burden. Then, a multi-scale self-attention mechanism is adopted to divide the features into windows of different scales for self-attention calculation. Through these multi-scale windows, small targets can utilize valuable context information from different scales.

[0068] (3) The present invention designs a dynamic scale-aware label assignment strategy to alleviate the problem of insufficient positive samples of small targets in the region proposal network. The dynamic scale-aware label assignment strategy assigns each anchor point to the most matching ground truth object based on the position relationship and adjusts the scale of the anchor point according to the matched ground truth object. This scale adjustment can effectively alleviate the mismatch problem between the preset anchor points and the ground truth boxes, thereby improving the IoU between the preset anchor points and the ground truth boxes, especially for small targets. Description of the Drawings

[0069] Figure 1 It is the implementation flowchart of the present invention.

[0070] Figure 2 It is the model framework diagram of the present invention.

[0071] Figure 3 It is the schematic diagram of the shift convolution module of the present invention.

[0072] Figure 4 It is the schematic diagram of the context interaction feature enhancement module of the present invention.

[0073] Figure 5 It is the example diagram of the detection results of the algorithm of the present invention, the unimproved algorithm, and some excellent algorithms. Specific implementation manners

[0074] The present invention will be further described in detail below with reference to the accompanying drawings.

[0075] As Figure 1 、 Figure 2 shown, the small object detection method based on scale-aware label assignment and context enhancement specifically includes the following steps:

[0076] (1) Image data preparation: Organize the image files and the corresponding annotation files; the purpose and function of this step are to ensure the correct matching of each image with its corresponding annotation data and organize them into a format convenient for model training and evaluation. In this way, it can be ensured that the model can learn the target position and category information from the correct annotations during the training process, thereby improving the detection performance.

[0077] (2) Training stage: The purpose and function of this step are to ensure that the network can learn how to accurately identify, locate, and classify targets during the training process through effective preprocessing, feature extraction, label assignment, and loss calculation of the input images. The training stage lays the foundation for the final detection ability of the model and optimizes the prediction process from the image to the target box.

[0078] (2.1) Obtain the training images, resize the original image size to the standard input size of pixels, and ensure that the input images maintain a fixed resolution to match the subsequent network structure.

[0079] (2.2) Preprocess the image: During the training process, first read the input image data from the storage medium and ensure that the image is loaded in RGB format. The loaded image will be resized to a fixed size of 800x800 pixels while maintaining the aspect ratio of the image width and height unchanged to ensure that the input image adapts to the requirements of the detector and avoid geometric distortion. In addition, to enhance the robustness and generalization ability of the model, the image is randomly flipped horizontally with a probability of 50%. Next, normalize the pixel values of the image and convert the BGR format to RGB format. Finally, resize the image size to a multiple of 32 to adapt to the downsampling operation of the model, ensure the alignment of feature maps during the feature extraction process, and prepare the subsequent object bounding box annotation information.

[0080] (2.3) Image feature extraction:

[0081] After the input image is preprocessed in step 2.2, it is input into a feature extractor using ResNet-50 as the backbone network for feature extraction. The network is divided into 4 stages, corresponding to different feature layers respectively;

[0082] Each stage will output different feature maps, and the number of output channels of ResNet is 256, 512, 1024, and 2048 in sequence;

[0083] Each convolutional layer of ResNet-50 is regularized using batch normalization and gradient update is enabled to ensure the stability of the model during training. To further improve the multi-scale representation ability of features, the output of the backbone network is fed into a Feature Pyramid Network (FPN). FPN fuses the features of these 4 stages and captures the spatial information of multiple objects through feature maps of different scales. The output channels of FPN are unified to 256, and finally 5 different-scale feature maps are output after the feature downsampling in the fourth stage for subsequent object detection.

[0084] (2.4) Context interaction feature enhancement module: As Figure 3 , shown in Figure 4, shift the convolution of different channels of the feature map in the first stage obtained in (2.3) to expand the receptive field, perform channel splitting and multi-scale self-attention calculation, and finally add and fuse with the original features to obtain the context content interaction enhanced features.

[0085] The proposed context interaction feature enhancement module only uses the highest-level feature pyramid feature layer as the input. Because in the multi-scale detection paradigm, small objects are often detected on the high-resolution feature maps of the highest level, and introducing multi-layer feature inputs will also cause a heavy redundant computational burden.

[0086] Specifically, as Figure 4As shown, first, the feature map is divided into K groups along the channel dimension. Then, a window with a size of Mk×Mk is set in each segmented feature map for spatial partitioning, and self-attention calculation is performed within each window. By setting different window sizes in each segmented feature map, context information can be captured at different scales. After completing the self-attention calculation at the window level, the segmented feature maps are concatenated and input into a 1×1 convolutional layer to obtain the final output. The complexity of this calculation process is as follows:

[0087]

[0088] Among them, M k represents the window size in the k-th feature partition. By adjusting the number of groups K and the window size M k , the computational complexity is controlled.

[0089] (2.5) Dynamic scale-aware label assignment strategy: Here, it is mainly improved based on the region candidate network structure, and ReRPNHead introducing regression information is used as the core component of RPN, where both the input feature channels and the internal feature channels are set to 256. This part contains an anchor generator for generating a series of candidate region boxes on the image. On this basis, the dynamic scale-aware label assignment strategy reshapes the scale of each proposed box and calculates the corresponding intersection over union ratio through the labeled scale supervision information, and then jointly supervises the assignment of positive and negative sample labels of the target through the introduced regression information;

[0090] Specifically, for each prior anchor box N i to be assigned, first, it is matched to the gt(x, y, w, h) box with the closest spatial position. And network regression parameters are introduced to match the gt box at the best spatial position for each prior anchor box to be assigned. First, the regression parameter vector of the regression branch of the detection network is calculated and transformed into a coordinate offset pred to obtain a learnable regression prior anchor box:

[0091] reg_N i = pred(N i ) (2)

[0092] Subsequently, the calculated *N i is used to calculate the IoU with gt i , which is called reg_IoU. reg_IoU is used to find the optimal spatial matching pair (N i , gt i ) for each sample Ni to be assigned in the feature space, where N i and gt i satisfy the following relationship:

[0093] reg_IoU(N i , gti ) = max(reg_IoU(N i , GT)) (3)

[0094] Where GT is the set of gt boxes in the same training image. For N i and gt i , its IoU calculation formula is as follows:

[0095]

[0096] On this basis, the adaptive scale IoU formula based on spatial addressing is as follows:

[0097]

[0098] Where, *N i is obtained by adaptively scaling the center sharing of *N i , and its scale is strictly constrained by the gt matched in (3). Specifically, a constraint formula in the following form is designed:

[0099]

[0100] Where size is a hyperparameter, and the aspect ratio of the initial anchor box is kept unchanged to maximize the preservation of the original prior features. Based on *N i , *N i is calculated by scaling the coordinates, and *IoU is calculated through *N i and remapped back to *N i to assign it the gt with the most matching spatial position and scale.

[0101] Based on the *IoU calculated for each anchor box in the above process, after regression offset, reg_*IoU is calculated, and finally the dynamic scale-aware IoU calculation formula is as follows:

[0102]

[0103] Where a is a scaling factor used to regulate the contribution degree of the regression parameter in the assignment strategy to adjust the optimal contribution values of the two.

[0104] (2.6) Forward propagation of the basic network: The enhanced image features obtained in (2.3) and (2.4) are used to generate candidate regions through (2.5) and perform training and prediction of positive and negative samples; finally, the positive sample candidate boxes of different sizes are pooled into a feature map of a fixed size through the RoI head for subsequent classification and regression.

[0105] (2.7) Calculate the loss value and perform backpropagation: First, use the Cross-Entropy Loss function to calculate the loss between the predicted probabilities of the RPN and RoI heads for foreground and background classification and the true labels (foreground or background). Then, use L1Loss to calculate the loss value between the predicted values (x, y, w, h) and the true values (gtx, gty, gtw, gth). Add the above loss values weighted, and use the calculated loss value for backpropagation to update the network parameters. Train in a loop multiple times until the network converges. During the update process, freeze the weights of the base model and do not participate in the update;

[0106] (3) For the testing phase; the purpose and function of this step are to detect the target by processing the input image, generating candidate boxes, and classifying and regressing each candidate box. The goal of the testing phase is to use the trained model for efficient inference, output accurate target categories and positions, and ensure the application effect of the model in the real scenario.

[0107] (3.1) Limit the longest side of the image to the specified range and input it into the backbone network to extract multi-scale feature maps for use by the region candidate network.

[0108] (3.2) Candidate box generation and filtering: Generate anchor boxes on each feature layer, calculate the top 1000 candidate boxes based on the object probability, remove the out-of-bounds and too-small candidate boxes, and use hierarchical non-maximum suppression to optimize the candidate boxes, retaining the final 1000 proposed boxes for subsequent steps.

[0109] (3.3) Final detection and class prediction: The candidate boxes are input into the ROI head for feature pooling and classification, and the class probability and bounding box regression parameters are output.

[0110] Table 1 Test results of the algorithm of the present invention on the AI-TOD dataset

[0111]

[0112]

[0113] Use the commonly used average precision (AP) as the standard metric to evaluate the proposed method. AP is defined as:

[0114]

[0115] where r represents the recall rate, and P(r) represents the corresponding precision rate. The precision rate and recall rate are defined as follows:

[0116]

[0117] Among them, TP, FP, and FN represent true positives, false positives, and false negatives respectively. The recall rate is used as the horizontal axis, and the corresponding precision value is determined as the vertical axis to construct a precision-recall curve. AP is obtained by calculating the area under this precision-recall curve. In particular, AP 50 indicates that the IoU threshold when defining TP is 0.5, while AP 75 indicates that the IoU threshold when defining TP is 0.75. AP represents the average value from AP 50 to AP 95 with the IoU interval set to 0.05. On the AI-TOD dataset, in addition to using AP, AP 50 and AP 75 for evaluation, APvt, APt, APs, and APm are also used to evaluate the performance of very small, tiny, small, and medium-sized targets respectively.

[0118] As can be seen from Table 1, the algorithm proposed in the present invention has improvements in all metrics compared to the Faster R-CNN algorithm, with AP and AP tiny increasing by 10.0% and 14.4% respectively. It proves the feasibility and effectiveness of the small target detection method using scale-aware label assignment and context content enhancement module for the recognition and detection of tiny targets in optical images.

[0119] Att Figure 5 is an example diagram of the detection results of the algorithm of the present invention, the unimproved algorithm, and some excellent algorithms. The red boxes represent the un-detected annotation boxes, and the green boxes represent the correctly detected prediction boxes. It can be seen from the figure that in the scenario with a large number of densely distributed small targets, the unimproved basic algorithm cannot accurately detect the correct positions of the targets. However, through the feature enhancement module and label assignment strategy of the present invention, the present invention has a more accurate feature response to small targets and can accurately locate the targets. That is, the Faster-RCNN algorithm with a scale-aware label assignment strategy has a significantly improved training effect on small targets when dealing with scenarios containing a large number of densely distributed small targets. It proves that the algorithm of the present invention is effective.

[0120] In summary, the method disclosed in the present invention adopts a carefully designed scale-aware label assignment method to replace the traditional IoU label assignment in the region proposal network, thereby providing more positive samples for small targets. In addition, the present invention proposes a context enhancement module to enhance the feature response of small targets through multi-scale context information.

Claims

1. A small target detection method based on scale-aware label assignment and context enhancement, characterized in that: The steps include: Step 1: Organize the image files and the corresponding annotation files, obtain the image data, divide the image data into training images and test images, and form the required data set; Step 2: By effectively preprocessing the training images, extracting features, assigning labels, and calculating losses, we ensure that the network can accurately identify, locate, and classify targets during training. Optimize the prediction process from image to target box and save weights; Step 3: Test the test image with the saved weights from training, and use the saved weights from training to evaluate on the test set to verify the model's ability to detect small objects in real scenes and adjust or optimize the model parameters.

2. The small target detection method based on scale-aware label assignment and context enhancement according to claim 1, characterized in that: In step 1, the image file is image data containing the target to be detected, which is stored in a common image format. Each image contains one or more targets, and standardized preprocessing is performed on images with different resolutions and sizes.

3. The small target detection method based on scale-aware label assignment and context enhancement according to claim 1, characterized in that: In step 1, the annotation file contains the target location information in the image and is stored in text or binary format.

4. The small target detection method based on scale-aware label assignment and context enhancement according to claim 1, characterized in that: The step 2 is specifically as follows: Step (2.1): Get the training image and resize the image data to the standard input size of pixels; Step (2.2): Preprocess the image. During the training process, first read the input training image data from the storage medium and ensure that the training image is loaded in RGB format; Step (2.3): Image feature extraction: input the training image loaded in RGB format into the feature extractor of the backbone network ResNet-50 for feature extraction; Step (2.4): The context interaction feature enhancement module performs shift convolution on different channels of the feature map in the first stage of step (2.3) to expand the receptive field, and performs channel splitting and multi-scale self-attention calculation. Finally, it is added with the original feature map that has passed through the feature extractor but not the feature enhancement module to obtain the context content interaction enhancement feature; Step (2.5): Dynamic scale-aware label assignment strategy, generate a set of candidate boxes from the context content interactively enhanced features and original features output from step (2.4) through RPN, reshape each candidate box with the labeled scale supervision information and calculate the corresponding intersection-over-union ratio, and then jointly supervise the assignment of target positive and negative sample labels through the introduced regression information; Step (2.6): Forward propagation of the basic network: Output features from step (2.3) and step (2.4), generate candidate regions through step (2.5) and perform training and prediction of positive and negative samples; Step (2.7): Compare the predicted value of step (2.6) with the coordinates of the corresponding target annotation box in the annotation file and the actual annotation value of the category, calculate the loss value of the two and back propagate.

5. The small target detection method based on scale-aware label assignment and context enhancement according to claim 4, characterized in that: The step (2.2) is specifically: The loaded training images will be adjusted to a fixed size set on demand, keeping the aspect ratio of the image unchanged; the training images will be randomly flipped horizontally; The processed image is resized to accommodate the downsampling operation of the model, ensuring that the feature maps are aligned during feature extraction and preparing the subsequent target box annotation information.

6. The small target detection method based on scale-aware label assignment and context enhancement according to claim 4, characterized in that: In the step (2.3), each convolutional layer of the backbone network ResNet-50 is regularized using batch normalization, and gradient updates are enabled to ensure that the model remains stable during training; The output of the backbone network ResNet-50 is sent to the Feature Pyramid Network (FPN), which fuses the features and captures the spatial information of multiple targets through feature maps of different scales; the output channels of FPN are unified.

7. The small target detection method based on scale-aware label assignment and context enhancement according to claim 4, characterized in that: The context content interaction enhancement in step (2.4) is specifically as follows: First, feature maps of different scales are divided into K groups along the channel dimension; Then, a window of size Mk×Mk is set in each segmented feature map to perform spatial division and perform self-attention calculation in each window. By setting different window sizes in each segmented feature map, contextual information is captured at different scales. After completing the window-level self-attention calculation, the segmented feature maps are concatenated and input into a 1×1 convolutional layer to obtain the final output. The calculation process is as follows: Among them, M k Represents the window size in the kth feature partition, by adjusting the number of groups K and the window size M k .

8. The small target detection method based on scale-aware label assignment and context enhancement according to claim 4, characterized in that: The step (2.5) is specifically: The feature map output by step (2.4) is used as the input of RPN. RPN generates a priori anchor boxes on the input feature map. For each prior anchor box N to be assigned i , first match it to the gt(x,y,w,h) box with the closest spatial position, and introduce network regression parameters to match the gt box with the best spatial position for each prior anchor box to be assigned. First, calculate the regression parameter vector of the detection network regression branch and convert it into a coordinate offset pred to obtain a learnable regression prior anchor box: reg_N i =pred(N i ) (2) Then use the calculated *N i With gt i Calculate IoU, called reg_IoU (i.e. regression IoU), and use reg_IoU for each sample N to be assigned i Find the optimal spatial matching pair (N i ,gt i ), where N i and gt i The following relations are satisfied: reg_IoU(N i ,gt i )=max(reg_IoU(N i ,GT)) (3) Among them, GT is the set of GT boxes in the same training image. i With gt i , and its IoU calculation formula is: On this basis, the adaptive scale IoU formula based on spatial addressing is as follows: Where, *N i By *N i The adaptive scale center is shared and scaled, and its scale is strictly constrained by the GT matched by (3). Specifically, the constraint formula of the following form is designed: Among them, size is a hyperparameter, and the aspect ratio of the initial anchor box is kept unchanged to maximize the original prior features. i Based on the scaled coordinates, *N is calculated i , through *N i Calculate *IoU and map it back to *N i Assign it the gt with the best matching spatial position and scale; Based on the above process, the *IoU calculated for each anchor box is calculated after regression offset to obtain reg_*IoU. The final dynamic scale-aware IoU calculation formula is as follows: Among them, a is the proportional factor, which is used to regulate the contribution of the regression parameters in the allocation strategy to adjust the optimal contribution value of the two.

9. The small target detection method based on scale-aware label assignment and context enhancement according to claim 4, characterized in that: The step (2.7) is specifically: First, the Cross-Entropy Loss loss function is used to calculate the loss between the predicted probability of foreground and background classification of RPN and the RoI head with RPN output as input and the true label. Then, L1Loss is used to calculate the loss value of the predicted value (x, y, w, h) and the true value (gtx, gty, gtw, gth). The above loss values ​​are weighted added, and the calculated loss value is used to backpropagate to update the network parameters. The training is repeated many times until the network converges. During the update process, the basic model weights are frozen and do not participate in the update.

10. The small target detection method based on scale-aware label assignment and context enhancement according to claim 9, characterized in that: The step 3 is specifically as follows: Step (3.1): Limit the longest side of the test image to a specified range and pass it into the backbone network to extract a multi-scale feature map for use by the region candidate network; Step (3.2): Candidate box generation and filtering, generate anchor boxes on each feature layer, calculate candidate boxes based on target probability, remove out-of-bounds and too small candidate boxes, and use hierarchical non-maximum suppression to optimize candidate boxes, and retain proposal boxes for subsequent steps; Step (3.3): Final detection and category prediction, the candidate box is passed to the ROI head, feature pooling and classification are performed, and the category probability and bounding box regression parameters are output.

Citation Information

Patent Citations

  • Aerial photography rotating small target detection method based on feature enhancement

    CN117994493A

  • Target detection method and system for accurately detecting small object based on deep learning

    CN118053117A

Cited By

  • PCBA element detection method and system based on spatial independence and scale perception

    CN121304689A

  • PCBA component detection method and system based on spatial independence and scale perception

    CN121304689B

  • 3D target detection model, 3D target detection method, 3D target detection device and vehicle

    CN122115980A