Small target detection method and system based on low-frequency samples

By introducing feature enhancement module, dynamic positive sample division strategy and dual-current decoupled detection head in remote sensing object detection, the problem of difficult to capture low-frequency sample features and poor model performance is solved, and more accurate small object detection and stronger model robustness are achieved.

CN119418125BActive Publication Date: 2025-05-13SHANDONG JIANZHU UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411566120.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-05-13
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

Existing remote sensing object detection methods are difficult to effectively capture and learn the features of low-frequency samples, resulting in poor performance of the model for small-object detection, and the feature coupling interference between prediction parameters makes it difficult for the model to obtain accurate bounding boxes.

Method used

A small object detection method based on low-frequency samples is proposed. By introducing feature enhancement modules and dynamic positive sample division strategies, the model's learning ability and feature representation ability of low-frequency samples are enhanced, and interference between features is reduced through dual-current decoupling detection heads.

Benefits of technology

It significantly improves the detection performance of low-frequency samples, achieves more accurate directional bounding box detection and category recognition, and enhances the interpretability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418125B_ABST
    Figure CN119418125B_ABST
Patent Text Reader

Abstract

The present invention discloses a small target detection method and system based on low-frequency samples, which include: obtaining a data set, which is a low-frequency remote sensing image with a known small target label; inputting the data set into a small target detection model, training the model, and obtaining a trained small target detection model; obtaining a remote sensing image to be detected, inputting the remote sensing image to be detected into the trained small target detection model, and obtaining a small target classification label and a position label of the remote sensing image to be detected; wherein, during the training process, the small target detection model is used to: extract multi-scale features from the input image, and perform feature enhancement processing on the extracted multi-scale features; extract an area of ​​interest from the enhanced features, and perform dynamic positive sample division on the area of ​​interest according to a set strategy; use the divided positive and negative sample training model to classify and perform bounding box regression processing on the area of ​​interest, and obtain a small target detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a small target detection method and system based on low-frequency samples. Background Art

[0002] In recent years, remote sensing target detection has attracted widespread attention. This task has important practical value and has broad application prospects in fields such as military identification and urban construction. It can not only replace manpower to complete tasks that are difficult to complete with manual monitoring, but also reduce manpower and material resources and improve work efficiency. Remote sensing target detection is divided into two categories: horizontal detection and directional detection. The recent trend is to develop oriented bounding boxes (OBBs) that can accurately match the direction of the target to achieve more accurate target positioning. At present, there have been a lot of research works at home and abroad, among which deep learning methods have achieved excellent results. In order to obtain accurate oriented bounding boxes, many previous works have been devoted to developing various directional bounding box detectors and improving the angle prediction accuracy of OBBs. However, due to the different semantic interpretations between different prediction parameters (coordinates, angles and shapes for classification and regression), the coupling interference between these parameters makes it difficult for the model to use the same feature map to obtain more accurate oriented bounding boxes and categories.

[0003] Remote sensing targets often present a long-tail distribution, with extremely unbalanced categories and dominated by small targets, which makes it difficult for existing methods to capture enough features from these low-frequency samples to learn their feature representations. The above results are mainly due to the following two reasons: (1) Direct feature fusion will lead to conflicts between different levels, and the shallow feature map lacks semantic information and cannot clearly distinguish the foreground and background, which makes the features of small target low-frequency samples more easily lost and submerged. (2) Since low-frequency samples contain a large number of small target samples, the existing positive sample allocation strategy based on intersection over union (IoU) has been proven to be too strict. This strategy ignores an important fact: high quality does not mean large size, and small size does not mean low quality. This neglect leads to the loss of a large number of high-quality small target samples, exacerbating the challenge of scarce low-frequency samples. Summary of the invention

[0004] In order to effectively mine low-frequency samples in remote sensing targets, alleviate the problem that the lack of low-frequency samples leads to the poor performance of the model in remote sensing target detection and the feature coupling interference between prediction parameters makes it difficult for the model to obtain an accurate oriented bounding box, the present invention provides a small target detection method and system based on low-frequency samples.

[0005] On the one hand, a small target detection method based on low-frequency samples is provided, including: obtaining a data set, which is a low-frequency remote sensing image with a known small target label; inputting the data set into a small target detection model, training the model, and obtaining a trained small target detection model; obtaining a remote sensing image to be detected, inputting the remote sensing image to be detected into the trained small target detection model, and obtaining a small target classification label and a position label of the remote sensing image to be detected; wherein, during the training process, the small target detection model is used to: extract multi-scale features from the input image, and perform feature enhancement processing on the extracted multi-scale features; extract a region of interest from the enhanced features, and dynamically divide the region of interest into positive samples according to a set strategy; use the divided positive and negative sample training model to classify and perform bounding box regression processing on the region of interest to obtain a small target detection result.

[0006] On the other hand, a small target detection system based on low-frequency samples is provided, including: an acquisition module, which is configured to: acquire a data set, wherein the data set is a low-frequency remote sensing image with a known small target label; a training module, which is configured to: input the data set into a small target detection model, train the model, and obtain a trained small target detection model; a detection module, which is configured to: acquire a remote sensing image to be detected, input the remote sensing image to be detected into the trained small target detection model, and obtain a small target classification label and a position label of the remote sensing image to be detected; wherein, during the training process, the small target detection model is used to: extract multi-scale features from the input image, and perform feature enhancement processing on the extracted multi-scale features; extract a region of interest from the enhanced features, and dynamically divide the region of interest into positive samples according to a set strategy; use the divided positive and negative sample training model to classify and perform bounding box regression processing on the region of interest to obtain a small target detection result.

[0007] The above technical solution has the following advantages or beneficial effects: The present invention proposes a low-frequency sample method to improve the learning ability and feature representation ability of the detector for low-frequency samples. First, the present invention introduces a feature enhancement module to enhance the feature representation of low-frequency samples during feature fusion.

[0008] Secondly, the present invention proposes a classification and regression module to customize feature maps with different semantic interpretations for each parameter, thereby obtaining more accurate oriented bounding boxes for low-frequency samples.

[0009] In order to further alleviate the problem of insufficient learning of low-frequency samples, the present invention proposes a dynamic positive sample division strategy ( ). This strategy determines the appropriate positive sample threshold based on the size of each object and the quality of the proposals of the region of interest, thereby mining more small object low-frequency samples. These improvements at the feature and sample levels greatly improve the detection performance of low-frequency samples.

[0010] Experimental results on four representative data sets demonstrate the effectiveness of the low-frequency sample mining method proposed in this invention in remote sensing target detection, achieving significant performance improvement. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0012] Figure 1 It is a schematic diagram of the overall structure of the present invention.

[0013] Figure 2 It is a schematic diagram of the internal structure of the feature extraction module of the present invention.

[0014] Figure 3 It is a schematic diagram of the internal structure of the MCFM of the present invention.

[0015] Figure 4 It is a schematic diagram of the internal structure of the classification module and regression module of the present invention.

[0016] Figure 5 Schematic diagram of the internal structure of the regression branch enhancement module of the present invention.

[0017] Figure 6 It is a schematic diagram of the internal structure of the first SKS module of the present invention.

[0018] Figure 7 Schematic diagram of the internal structure of the fourth SKS module of the present invention. DETAILED DESCRIPTION

[0019] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0020] Embodiment 1. This embodiment provides a small target detection method based on low-frequency samples, including: S101: obtaining a data set, wherein the data set is a low-frequency remote sensing image with a known small target label; S102: inputting the data set into a small target detection model, training the model, and obtaining a trained small target detection model; S103: obtaining a remote sensing image to be detected, inputting the remote sensing image to be detected into the trained small target detection model, and obtaining a small target classification label and a position label of the remote sensing image to be detected; wherein, during the training process, the small target detection model is used to: extract multi-scale features from the input image, and perform feature enhancement processing on the extracted multi-scale features; extract a region of interest from the enhanced features, and dynamically divide the region of interest into positive samples according to a set strategy; use the divided positive and negative sample training model to classify and perform bounding box regression processing on the region of interest to obtain a small target detection result.

[0021] Further, S101: obtaining a data set, wherein the data set is a low-frequency remote sensing image with known small target labels, wherein a small target refers to a target with a size below a set threshold, for example, The target of the pixel, where low frequency means that the frequency of the target appearing in the data set is lower than the set threshold. For example, a target with an appearance frequency lower than one thousandth is a low-frequency target.

[0022] Furthermore, after obtaining the data set, it also includes: performing preprocessing operations on the data set; the preprocessing operations include: cropping the original image into 1024×1024 sub-images, and the overlapping part of each sub-image with the adjacent sub-image is 200×200 pixels; then using methods such as randomly pasting small targets, adding random noise, and changing image brightness to enhance the data set, increase the number of small targets in the data set, simulate special scenes, balance the number of small target positive samples, and improve the model robustness and generalization ability.

[0023] For high-resolution remote sensing images with long-tail distribution, the categories are extremely balanced, and there are a large number of small targets in low-frequency samples. It is difficult for existing general target detectors to extract enough low-frequency samples to learn their feature representation. In addition, since there are currently few small target data sets and they are difficult to label, the existing method of copying small targets and randomly pasting them to other locations is used to increase the number of small targets, increase the number of features of small targets and the number of positive samples, and balance the contribution of small targets during network optimization. By adding noise to the data set to change the brightness to simulate lighting changes and shooting blur problems, the model has better robustness and generalization.

[0024] Further, S102: input the data set into the small target detection model, train the model, and obtain the trained small target detection model. During the training process, when the total loss function of the model or the number of training iterations of the model exceeds the set number, stop the training to obtain the trained small target detection model.

[0025] Furthermore, if Figure 1 As shown, S102: input the data set into the small target detection model, train the model, and obtain the trained small target detection model, wherein the trained small target detection model includes: a feature extraction module, a feature enhancement module, and a region of interest generation module connected in sequence; the output end of the region of interest generation module is respectively connected to the input end of the classification module and the input end of the regression module.

[0026] Furthermore, the feature extraction module uses a first LSKNet (Large Selective Kernel Network), a second LSKNet, a third LSKNet and a fourth LSKNet connected in sequence; the first LSKNet extracts features from the input remote sensing image The second LSKNet extracts features from the output of the first LSKNet The third LSKNet extracts features from the output of the second LSKNet The fourth LSKNet extracts features from the output of the third LSKNet , for the features Perform downsampling to obtain features , the feature As a feature , the feature With features Sum to get the features , the feature With features Sum to get the features , the feature With features Sum to get the features .

[0027] LSKNet is derived from the paper Yuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng, Jian Yang, Xiang Li, Large Selective Kernel Network for Remote Sensing Object Detection, a selective large kernel network for remote sensing target detection.

[0028] It should be understood that existing high-performance remote sensing target detectors usually rely on the RCNN framework, where RPN is responsible for extracting region of interest proposals from the backbone feature map, and the regional CNN detection head is responsible for target classification and bounding box regression. The present invention selects LSKNet as the backbone to obtain a series of feature maps of different resolutions. This method adds appropriate contextual information to targets that require different receptive fields through a spatial selection mechanism, and achieves excellent detection performance.

[0029] Furthermore, if Figure 2 As shown, the feature enhancement module includes four parallel branches: a first branch, a second branch, a third branch and a fourth branch.

[0030] The first branch includes: a first MCFM module, a first adder, a first SKS module and a first EMA module connected in sequence; the input end of the first MCFM module is also connected to the input end of the first adder; the first MCFM input end inputs a characteristic The input end of the first SKS module also inputs the features processed by the first average pooling layer ; The input end of the first SKS module also inputs the upsampled feature ; Output characteristics of the first EMA module .

[0031] The second branch includes: a second SKS module and a second EMA module connected in sequence; the input end of the second SKS module is also connected to the output end of the first SKS module through a second average pooling layer; the input end of the second SKS module also inputs the feature vector processed by upsampling. ; The input of the second SKS module also inputs the characteristic ; Output characteristics of the second EMA module .

[0032] The third branch includes: a third SKS module and a third EMA module connected in sequence; the input end of the third SKS module is also connected to the output end of the second SKS module through a third average pooling layer; the input end of the third SKS module also inputs the feature vector processed by upsampling. The input of the third SKS module also inputs the characteristic ; Output characteristics of the third EMA module .

[0033] The fourth branch includes: a fourth SKS module and a fourth EMA module connected in sequence; the input end of the fourth SKS module is also connected to the output end of the third SKS module through a fourth average pooling layer; the input end of the fourth SKS module also inputs the feature ; Fourth EMA module output characteristics .

[0034] It should be understood that the feature enhancement module is also called a fine-grained context-aware feature pyramid network.

[0035] In order to effectively enhance the representation ability of feature maps for low-frequency samples, the feature enhancement module first enhances the shallow feature maps by learning fine-grained related semantic information. This enhancement helps to distinguish small target features from the background, thereby obtaining enhanced features : , where MCFM is Figure 2 The multi-scale context feature mining module proposed in . This module can obtain the relevant context information of small targets on the shallow feature map, enhance the expression ability of the shallow feature map for small targets, and reduce the risk of small targets being lost and submerged during feature fusion.

[0036] Then, in order to fuse the semantic information of different receptive fields in each layer, the maximum pooling operation is used to obtain the shallower features. , using the bilinear interpolation method to obtain deeper features The SKS module gets the mask of the spatial selection mechanism .

[0037] .

[0038] in, , and Represent the sigmoid function, average pooling operation and maximum pooling operation respectively. The symbol " '' indicates the connection operation. Similarly, the fused feature map is obtained and The mask of . Indicates the use of a convolution kernel size of , the convolution with a stride of 1 aims to increase the number of feature map channels from 2 to 3.

[0039] Finally, by By adding the contextual information of the appropriate receptive fields of targets of different sizes, we obtain an enhanced multi-scale feature map for low-frequency samples, denoted as .

[0040] ;

[0041] in, EMA stands for Efficient Multi-Scale Attention, which can focus the feature map on the foreground target area and reduce the interference of the background area. represents the first feature map; represents an enhanced second feature map; represents the third feature map; Indicates Feature map; Indicates Feature map; Indicates Feature map; represents the fourth characteristic map; represents the fifth feature map.

[0042] Through the above method, the risk of losing and drowning small target features in the feature fusion process is reduced, and more small target low-frequency samples are found. In addition, the context information of targets of different scales is further enhanced, so that the feature map has a stronger representation ability for low-frequency samples.

[0043] Furthermore, if Figure 3 As shown, the first MCFM module includes: a first two-dimensional convolutional layer, a first depth-separable convolutional layer DWConv, a second depth-separable convolutional layer DWConv, a splicing unit, a pooling layer, a second two-dimensional convolutional layer, an activation function layer, a second adder and an output end connected in sequence; the input end of the second adder is connected to the input end of the first multiplier; the input end of the first multiplier is used to input a remote sensing image; the input end of the first multiplier is also connected to the input end of the first two-dimensional convolutional layer.

[0044] Input features of the first 2D convolutional layer ; The output end of the first two-dimensional convolutional layer is also connected to the input end of the splicing unit; the output end of the first depth-separable convolutional layer DWConv is also connected to the input end of the splicing unit; the output end of the first two-dimensional convolutional layer is also connected to the input end of the splicing unit through the third depth-separable convolutional layer DWConv; the output end of the first two-dimensional convolutional layer is also connected to the output end of the activation function layer through the fifth EMA module, the output end of the second depth-separable convolutional layer DWConv is also connected to the output end of the activation function layer, the output end of the third depth-separable convolutional layer DWConv is also connected to the output end of the activation function layer, and the output end of the first depth-separable convolutional layer DWConv is also connected to the output end of the activation function layer.

[0045] Furthermore, the internal structures of the first SKS module, the second SKS module, the third SKS module and the fourth SKS module are consistent, such as Figure 6As shown, the first SKS module includes: a series splicing unit, an average pooling layer, a third two-dimensional convolution layer and an activation function layer; the output end of the series splicing unit is also connected to the input end of the third two-dimensional convolution layer through a maximum pooling layer.

[0046] like Figure 7 As shown, the internal structure of the fourth SKS module is consistent with the internal structure of the first SKS module.

[0047] Furthermore, the internal structures of the first EMA module, the second EMA module, the third EMA module and the fourth EMA module are consistent, and the first EMA module is an attention mechanism layer.

[0048] The first EMA module comes from Daliang Ouyang, Su He, Guozhong Zhang, Mingzhu Luo, Huaiyong Guo, Jian Zhan, Zhijie Huang, Efficient Multi-Scale Attention Module with Cross-Spatial Learning.

[0049] Furthermore, the region of interest generation module is implemented through an RPN (Region Proposal Network).

[0050] Furthermore, the region of interest generation module adopts a dynamic positive sample division strategy; the dynamic positive sample division strategy includes: (1) generating region of interest proposals through the RPN network; (2) using the normalized area metric To indicate the target size The impact of target thresholds: ;in, and Representative The width and height of the ground-truth bounding box of the object, using the maximum value and minimum value The function normalizes the area. Apply the base 2 logarithm function The effect of object size on the larger object threshold can be reduced, making the quality of the region of interest proposals of the current object the main factor.

[0051] (3) Evaluate the quality of the ROI proposals by calculating the mean and median of the ROI proposals Compared with the median value, the higher the average value, the higher the quality of the proposals in the region of interest. To evaluate the The quality of proposals for each target: ;in, The function calculates the median. The function calculates the average value, represents all region of interest proposals for the i-th target whose IoU value is greater than 0.3. 0.3 represents the specified minimum threshold for positive samples in the network.

[0052] (4) In order to avoid losing high-quality samples of large targets in low-frequency samples, the The positive threshold of the target value is set to the maximum positive threshold specified by the network , and derive the conversion ratio To determine the positive sample threshold for other targets. The description is as follows: ;in, To determine The maximum value of the ratio. Generate modules in the region of interest, is 0.7, in the classification module and regression module, is 0.5. The positive sample threshold of a target is expressed as : ;in, Represents the minimum value of the positive sample threshold assigned to each target, adjusted according to the dataset , to determine its optimal value.

[0053] (5) All samples of each target whose IoU value is greater than the positive sample threshold are classified as positive samples.

[0054] By adjusting the positive sample threshold, more samples are allocated to low-frequency sample targets so that the model can better learn their feature expressions during training.

[0055] In order to reduce the loss of high-quality small target samples in low-frequency samples, the present invention believes that high quality does not equal large targets, and small targets do not equal low quality. Therefore, the present invention sets a new evaluation criterion to determine the appropriate positive sample threshold for each target. Subsequently, the present invention proposes that the region of interest proposals generated in the RPN stage can be systematically divided into four groups according to the target size and the quality of the region of interest proposals: "large high quality", "large low quality", "small high quality" and "small low quality".

[0056] When the target area is larger than the first set threshold, but the intersection over union (IoU) score of the region of interest proposal is higher than the second set threshold, the current region of interest is divided into large high-score groups.

[0057] When the target area is larger than the first set threshold, but the IoU score of the ROI proposal is lower than the second set threshold, the current ROI is divided into a large low-score group.

[0058] When the target area is smaller than the first set threshold, but the IoU score of the ROI proposal is higher than the second set threshold, the current ROI is divided into a small high-score group.

[0059] When the target area is smaller than the first set threshold, but the IoU score of the ROI proposal is lower than the second set threshold, the current ROI is divided into a small low-score group.

[0060] Based on the above classification, we can draw the following conclusions: For region of interest proposals classified as “large and high quality”, their targets are usually larger, higher in quality, and easier to detect.

[0061] Therefore, even with a larger positive sample threshold, accurate detection can be achieved. On the contrary, for the "large low quality", "small high quality" and "small low quality" categories, appropriately reducing the positive sample threshold can mine more proposals of the region of interest and enhance the model's effective learning from these samples. Similarly, this strategy can also be applied to the regional CNN stage to increase the training samples of low-frequency samples.

[0062] Large targets are easy to detect, and their region of interest proposals are easy to get high scores. For the large high-score group and the small high-score group, they are undoubtedly high-quality samples and are the key samples used for model training. For the large low-score group, the large targets still get low scores, so they should be classified as low-quality samples and discarded. However, for small targets that are more difficult to detect, it is difficult to obtain high scores under the existing IoU-based strategy, so we should appropriately lower the positive sample threshold of such targets to select some high-quality samples from this group of region of interest proposals to assist training. For low-frequency samples, it is far from enough to rely solely on samples obtained from large targets, highlighting the importance of mining more high-quality region of interest proposals for small targets.

[0063] Obviously, using the target size and the quality of the ROI proposals as the criteria for setting the positive sample threshold can effectively reduce the loss of high-quality small target low-frequency samples.

[0064] This strategy suggests that objects with larger, higher quality ROI proposals have larger value, so a higher positive sample threshold should be assigned. On the contrary, smaller objects and poorer ROI proposal quality require a lower threshold to obtain more positive samples. Choose an appropriate It can minimize the inclusion of low-quality samples, which is conducive to mining more high-quality low-frequency samples and significantly improves its detection performance. If the dataset has more low-frequency samples, the positive sample threshold should be smaller.

[0065] Furthermore, if Figure 4 and Figure 5 As shown, the regression module includes: a position regression branch, a direction regression branch and a shape regression branch.

[0066] Through the three branches of position regression branch, direction regression branch and shape regression branch, we get the feature map customized for different semantic interpretation parameters of the regression branch. .

[0067]

[0068] Among them, the superscript , and Represent the position, direction and shape branches respectively, Represents coordinate convolution. It promotes the model's ability to learn the positional relationship between the target and its surroundings by integrating position information into the convolution operation. refers to the coordinate attention mechanism, represents the regression branch feature map, Indicates the pixel-level supervision obtained by using the prediction result of the previous branch to perform an affine transformation on the proposal. represents the feature map customized for the coordinate branch, Represents the feature map customized for the direction branch.

[0069] Coordinate convolution comes from the paper Liu R, Lehman J, Molino P, et al. An intriguing failing of convolutional neural networks and the coordconv solution[J]. Advances in neural information processing systems, 2018, 31.

[0070] The coordinate attention mechanism comes from Hou Q, Zhou D, Feng J. Coordinate attention for efficient mobile network design[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021: 13713-13722.

[0071] It enables the model to understand complex contextual relationships by focusing on specific areas in the feature map and combining their position information. Represents the enhanced activation mask. The specific approach is to combine the obtained proposal with the current branch prediction through affine transformation to derive the target region, thereby increasing the weight of the target region.

[0072] In order to reduce the instability of regression prediction, the present invention uses a hybrid expert system to gradually improve the regression parameters. In addition, in order to distinguish the pros and cons of each prediction result, the present invention classifies and scores the regression parameter predictions of each branch. These classification scores will then be used as weights to adjust the regression parameters. Taking the fine-tuning of width and height in the shape branch as an example, the present invention first obtains the prediction results from the three branches. , And its corresponding classification score. Then, the present invention applies the softmax function to normalize the classification score into weights , these weights determine the accuracy of the prediction of w and h in each branch. Finally, the present invention obtains the adjusted shape branch result, which is recorded as .

[0073] .

[0074] Among them, the subscript , and Represent the position, direction and shape branches respectively, Represents the prediction results of the position branch for width and height, Indicates the prediction results of the direction branch for width and height, Represents the prediction results of the shape branch for width and height; represents the weight of the position branch, represents the weight of the directional branch, Represents the weight of the shape branch. The higher the weight, the more accurate the prediction result of the current branch.

[0075] The position branch does not need further fine-tuning, while the shape branch needs to be fine-tuned in conjunction with the position branch. Fine-tuning these three branches can obtain a more accurate orientation bounding box.

[0076] like Figure 4 As shown, the position regression branch is used to adopt the coordinate convolution method to meet the model's requirements for accurate spatial position information. This method can obtain additional coordinate information, thereby improving the network's sensitivity to position details.

[0077] The direction regression branch is used to adopt coordinate convolution to enhance the model's ability to perceive the target angle, and utilize the coordinate attention mechanism to focus on the most relevant area to detect changes in the target orientation.

[0078] The shape regression branch is used to capture more complex background information and multi-scale features by using coordinate convolution multiple times. The coordinate attention mechanism enables the model to adaptively focus on the area most relevant to the target size, which is particularly helpful for processing significant changes in target shape or size or complex backgrounds.

[0079] As shown in Figure 2, for the classification branch , a multi-scale contextual feature mining module (MCFM) is proposed. The MCFM module can utilize contextual information to enrich the discriminative features of targets of different scales in low-frequency samples.

[0080] Furthermore, the classification module specifically works as follows: the MCFM module first uses a depth-separable convolution with convolution kernel sizes of 3×3 and 5×5 to learn semantic information of different receptive fields multiple times. . Experimental results show that convolution kernel sizes of 3×3 and 5×5 provide the best accuracy and fastest speed.

[0081] .

[0082] in, Depthwise separable convolution, which has the characteristics of fewer parameters and computational cost than standard convolution, but can still effectively learn feature representation. The 3×3 convolution kernel captures the semantic information of a smaller receptive field, while the 5×5 convolution kernel is used to quickly expand the receptive field to capture the semantic information of a larger receptive field, thereby obtaining different semantic information of objects of different scales in low-frequency samples. They represent the feature maps after the classification branch is convolved with a depthwise separable convolutional layer with convolution kernels of 1×1, 3×3, 5×5, and 5×5, respectively.

[0083] Similarly, inspired by the spatial kernel selection, feature map masks with semantic information of different receptive fields are obtained : ;in, represents the activation function, Represents a convolution that changes the number of channels of the feature map from 2 to 4; represents the average pooling layer; Represents a maximum pooling layer.

[0084] Subsequently, the multi-scale contextual information of the classification features is enhanced by combining the semantic information with the appropriate receptive fields of objects of different scales in low-frequency samples, thus obtaining .

[0085] .

[0086] in, Represents the enhanced classification branch features; express The spatial selection mask of express The spatial selection mask of express The spatial selection mask of express The spatial selection mask of Represents an attention mechanism layer.

[0087] In order to further emphasize the target area and retain more discriminative features of low-frequency samples, the optimal regression parameters are used to perform affine transformation on the region of interest proposals to obtain a pixel-level enhanced activation mask (AM).

[0088] Finally, the multi-scale context features The classification branch features enhanced by pixel-level enhanced activation mask (AM) Integrate to get the final classification feature map : .

[0089] Using feature maps customized according to different semantic interpretations to predict relevant parameters allows each branch to focus on learning beneficial features. This not only solves the feature coupling problem between prediction parameters, provides more accurate oriented bounding boxes for low-frequency samples, but also enhances the interpretability of the model.

[0090] In order to deal with the feature coupling interference between parameters with different semantic interpretations, this paper proposes a dual-stream disentangled detection head (DD-Head). This method tailors the feature maps for parameters with different semantic interpretations, allowing the model to focus more on learning relevant features. This paper verifies the optimal order of prediction regression parameters determined by the STD method and decides to keep the order of position, direction, and shape.

[0091] Specifically, the present invention divides the prediction task into two main branches: classification and regression. The regression branch is further divided into three sub-branches: position, direction, and shape. As shown in Figure 1 (DD-Head), for the regression branch ,The present invention customizes the feature graph by analyzing the semantic interpretation of each branch.

[0092] Furthermore, the multi-scale feature extraction of the input image is achieved through a feature extraction module.

[0093] Existing high-performance remote sensing target detectors usually rely on the R-CNN framework, where the RPN is responsible for extracting region of interest proposals from the backbone feature map, and the region RCNN (Region-based CNN, RCNN) detection head is responsible for target classification and bounding box regression. By extracting features from the image through the backbone network, shallow feature maps with rich detail information and deep feature maps with rich semantic features can be obtained.

[0094] Furthermore, the feature enhancement processing of the extracted multi-scale features is achieved through a feature enhancement module.

[0095] The feature enhancement module is also called the fine-grained context-aware feature pyramid network (FC-FPN): the network first makes it easier to separate the foreground object from the background by learning multi-scale context information on the shallow feature map, reducing the risk of loss and submergence during the feature fusion process. Then, through the interactive fusion of feature maps at different levels, the feature map's ability to represent low-frequency sample features is enhanced.

[0096] Furthermore, the extraction of the region of interest from the enhanced features is achieved through a region of interest generation module.

[0097] Furthermore, the dynamic positive sample division of the region of interest according to the set strategy is achieved through a dynamic positive sample division strategy.

[0098] It should be understood that the dynamic positive sample division strategy ( ): This strategy dynamically determines the appropriate positive sample threshold for each target according to the size of the target and the quality of proposals, mines more small target low-frequency samples to supplement the scarce low-frequency samples, and effectively improves the detection performance of low-frequency samples.

[0099] Furthermore, the classification and bounding box regression processing of the region of interest to obtain the small target detection result is achieved through a classification module and a regression module.

[0100] Two-stream Decoupled Detection Head (DD-Head): This method enhances the interpretability of the model by customizing feature maps with different semantic interpretations to predict relevant parameters, allowing the model to better distinguish and focus on the learning of parameter-related features, reducing interference between features, and thus obtaining more accurate oriented bounding boxes for low-frequency samples.

[0101] Network training: Select appropriate remote sensing and drone aerial photography datasets, such as DOTA-v1.0, DOTA-v1.5, DIOR-R, etc., input the model for training, and iterate until convergence or the specified number of rounds is reached. Save the optimal network model weights for inference.

[0102] Inference phase: Load the optimal model weights, input the test image or video for prediction, and obtain the prediction result, which is an image or video annotated with the oriented bounding box, category, and confidence.

[0103] The present invention can be easily applied to all the most advanced convolution-based frameworks. Take the most advanced method LSKNet as the backbone network as an example. The structure diagram is as follows Figure 1 Select appropriate remote sensing and drone aerial photography datasets and input the datasets preprocessed by data enhancement into Figure 1 In the model shown, after multiple iterations of training, the model parameters are continuously optimized until the loss function no longer fluctuates or the number of training times is reached. The network model weights at the minimum loss value are saved.

[0104] Load the optimal model weights, input the image or video data to be tested for prediction, and obtain the image or video file with the prediction box, category and confidence. Or directly connect the camera for real-time detection.

[0105] The network model has been trained and its parameters are saved. Users only need to input the image data to be tested into the network, and the model will automatically load the optimal model parameters to accurately predict the input image or video, and finally output the prediction results to interact with the user.

[0106] The present invention firstly introduces a fine-grained context-aware feature pyramid network (FC-FPN) to reduce the risk of losing or submerging target features during feature fusion, thereby enhancing the representation capability of low-frequency sample features.

[0107] In order to solve the coupling interference between prediction parameters with different semantic interpretations, the present invention further proposes a dual-stream decoupled detection head (DD-Head). This method predicts relevant parameters by customizing feature maps with different semantic interpretations, so that the model can better distinguish and focus on parameter-related features. Moreover, this method also reduces the interference between features and enhances the interpretability of the model, thereby obtaining a more accurate oriented bounding box for low-frequency samples. However, simply enhancing features is not enough to make up for the fact that low-frequency samples are scarce.

[0108] The existing IoU-based positive sample allocation strategy usually results in low scores for proposals of small objects, which are often classified as low-quality samples. Therefore, this paper proposes a dynamic positive sample partitioning strategy (DPS 3 ). This strategy dynamically determines the appropriate positive sample threshold for each target according to the size of the target and the quality of proposals, mines more small target low-frequency samples to supplement the low-frequency samples, and effectively improves the detection performance of low-frequency samples.

[0109] The low-frequency sample mining method proposed in the present invention achieves outstanding performance that surpasses the existing best methods on four representative remote sensing data sets. It has strong adaptability to a variety of remote sensing scenarios, especially for real scenarios with a large number of low-frequency samples. It lays a solid foundation for promoting the development of remote sensing target detection in real scenarios and its industrialization.

[0110] Embodiment 2. This embodiment provides a small target detection system based on low-frequency samples, including: an acquisition module, which is configured to: acquire a data set, wherein the data set is a low-frequency remote sensing image with a known small target label; a training module, which is configured to: input the data set into a small target detection model, train the model, and obtain a trained small target detection model; a detection module, which is configured to: acquire a remote sensing image to be detected, input the remote sensing image to be detected into the trained small target detection model, and obtain a small target classification label and a position label of the remote sensing image to be detected; wherein, during the training process, the small target detection model is used to: extract multi-scale features from the input image, and perform feature enhancement processing on the extracted multi-scale features; extract a region of interest from the enhanced features, and dynamically divide the region of interest into positive samples according to a set strategy; use the divided positive and negative sample training model to classify and perform bounding box regression processing on the region of interest to obtain a small target detection result.

[0111] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. Small target detection method based on low-frequency samples, characterized by: include: Acquire a data set, wherein the data set is a low-frequency remote sensing image with known small target labels; Input the data set into the small target detection model, train the model, and obtain the trained small target detection model; Obtain a remote sensing image to be detected, input the remote sensing image to be detected into the trained small target detection model, and obtain a small target classification label and a position label of the remote sensing image to be detected; During the training process, the small target detection model is used to: extract multi-scale features from the input image, perform feature enhancement on the extracted multi-scale features; extract the region of interest from the enhanced features, and dynamically divide the region of interest into positive samples according to the set strategy; use the divided positive and negative sample training model to classify the region of interest and perform bounding box regression processing to obtain the small target detection result; The data set is input into the small target detection model, and the model is trained to obtain a trained small target detection model, wherein the trained small target detection model includes: A feature extraction module, a feature enhancement module and an area of ​​interest generation module connected in sequence; the output end of the area of ​​interest generation module is connected to the input end of the classification module and the input end of the regression module respectively; The region of interest generation module adopts a dynamic positive sample division strategy; the dynamic positive sample division strategy includes: (1) Generate region of interest proposals through the RPN network; (2) Using the normalized area measure A i To express the effect of target size on the i-th target threshold: Among them, w i and h i Represents the width and height of the true bounding box of the i-th target, and the area is normalized using the maximum value max and minimum value min functions; (3) Evaluate the quality Q of the proposals of the region of interest by calculating the average and median of the proposals of the region of interest i , Among them, the median function calculates the median, the mean function calculates the mean, and c_p i Represents all region of interest proposals for the i-th target; (4) In order to avoid losing high-quality samples of large targets in low-frequency samples, the positive sample threshold of the target with the highest A / Q value is set to the maximum positive sample threshold τ specified by the network. max , and derive the conversion ratio β to determine the positive sample threshold of other targets; the ratio β is described as follows: Among them, max is used to determine the maximum value of the A / Q ratio; Then, the positive sample threshold of the i-th target is obtained, denoted as T i : Among them, τ min Represents the minimum value of the positive sample threshold assigned to each target, adjusted according to the dataset τ min , to determine its optimal value; (5) All samples of each target whose IoU value is greater than the positive sample threshold are classified as positive samples; The regression module includes: a position regression branch, a direction regression branch and a shape regression branch; Through the three branches of position regression branch, direction regression branch and shape regression branch, we get the feature map customized for different semantic interpretation parameters of the regression branch. It is expressed as: The superscripts l, o, and s represent the position, direction, and shape branches, respectively. represents coordinate convolution, refers to the coordinate attention mechanism, F r represents the feature map of the regression branch, AM represents the pixel-level supervision obtained by using the affine transformation of the proposal using the prediction result of the previous branch, represents a feature map customized for the coordinate branch, Represents the feature map customized for the direction branch.

2. The small target detection method based on low-frequency samples as claimed in claim 1, characterized in that: The feature extraction module uses a first LSKNet, a second LSKNet, a third LSKNet and a fourth LSKNet connected in sequence; the first LSKNet extracts feature P1 from the input remote sensing image, the second LSKNet extracts feature P2 from the result output by the first LSKNet, the third LSKNet extracts feature P3 from the result output by the second LSKNet, the fourth LSKNet extracts feature P4 from the result output by the third LSKNet, downsamples feature P4 to obtain feature F5, uses feature P4 as feature F4, sums feature F4 and feature P3 to obtain feature F3, sums feature F3 and feature P2 to obtain feature F2, and sums feature F2 and feature P1 to obtain feature F1.

3. The small target detection method based on low-frequency samples as claimed in claim 1, characterized in that: The feature enhancement module comprises: Four parallel branches: the first branch, the second branch, the third branch and the fourth branch; The first branch comprises: a first MCFM module, a first adder, a first SKS module and a first EMA module connected in sequence; the input end of the first MCFM module is also connected to the input end of the first adder; the input end of the first MCFM module inputs feature F2; the input end of the first SKS module also inputs feature F1 processed by the first average pooling layer; the input end of the first SKS module also inputs feature F3 processed by upsampling; the first EMA module outputs feature F2'; The second branch includes: a second SKS module and a second EMA module connected in sequence; the input end of the second SKS module is also connected to the output end of the first SKS module through a second average pooling layer; the input end of the second SKS module also inputs the up-sampled feature F4; the input end of the second SKS module also inputs the feature F3; the second EMA module outputs the feature F3'; The third branch includes: a third SKS module and a third EMA module connected in sequence; the input end of the third SKS module is also connected to the output end of the second SKS module through a third average pooling layer; the input end of the third SKS module also inputs the up-sampled feature F5; the input end of the third SKS module also inputs the feature F4; the third EMA module outputs the feature F4'; The fourth branch includes: a fourth SKS module and a fourth EMA module connected in sequence; the input end of the fourth SKS module is also connected to the output end of the third SKS module through a fourth average pooling layer; the input end of the fourth SKS module also inputs feature F5; and the fourth EMA module outputs feature F5'.

4. The small target detection method based on low-frequency samples as claimed in claim 3 is characterized in that: The first MCFM module comprises: A first two-dimensional convolutional layer, a first depth-separable convolutional layer DWConv, a second depth-separable convolutional layer DWConv, a splicing unit, a pooling layer, a second two-dimensional convolutional layer, an activation function layer, a second adder and an output end connected in sequence; the input end of the second adder is connected to the input end of the first multiplier; the input end of the first multiplier is used to input a remote sensing image; the input end of the first multiplier is also connected to the input end of the first two-dimensional convolutional layer; The input end of the first two-dimensional convolutional layer inputs feature F2; the output end of the first two-dimensional convolutional layer is also connected to the input end of the splicing unit; the output end of the first depth-wise separable convolutional layer DWConv is also connected to the input end of the splicing unit; the output end of the first two-dimensional convolutional layer is also connected to the input end of the splicing unit through the third depth-wise separable convolutional layer DWConv; the output end of the first two-dimensional convolutional layer is also connected to the output end of the activation function layer through the fifth EMA module, the output end of the second depth-wise separable convolutional layer DWConv is also connected to the output end of the activation function layer, the output end of the third depth-wise separable convolutional layer DWConv is also connected to the output end of the activation function layer, and the output end of the first depth-wise separable convolutional layer DWConv is also connected to the output end of the activation function layer.

5. The small target detection method based on low-frequency samples as claimed in claim 1, characterized in that: The classification module specifically works as follows: the MCFM module first uses depth-separable convolutions with convolution kernel sizes of 3×3 and 5×5 to learn semantic information of different receptive fields multiple times. Among them, DWConv represents depth-wise separable convolution, They represent the feature maps of the classification branch after convolution using the depthwise separable convolutional layer with convolution kernels of 1×1, 3×3, 5×5, and 5×5 respectively; Get the feature map mask θ with semantic information of different receptive fields i : Among them, σ represents the activation function, F 2→4 represents the convolution that changes the number of channels of the feature map from 2 to 4; P avg represents the average pooling layer; P max represents the maximum pooling layer; Subsequently, the multi-scale contextual information of the classification features is enhanced by combining the semantic information with the appropriate receptive fields of objects of different scales in low-frequency samples, thus obtaining F′ c : Among them, F′ c represents the enhanced classification branch features; θ1 represents The spatial selection mask of , θ2 represents The spatial selection mask of , θ3 represents The spatial selection mask of , θ4 represents The spatial selection mask of ;EMA represents the attention mechanism layer; Use the optimal regression parameters to perform affine transformation on the region of interest proposals to obtain the pixel-level enhanced activation mask AM; finally, the multi-scale context feature F′ c And the classification branch feature F enhanced by the pixel-level enhanced activation mask AM c Integrate to get the final classification feature map 6. The small target detection method based on low-frequency samples as claimed in claim 1, characterized in that: The data set is input into the small target detection model, and the model is trained to obtain the trained small target detection model. During the training process, when the total loss function of the model or the number of training iterations of the model exceeds the set number, the training is stopped to obtain the trained small target detection model.

7. A small target detection system based on low-frequency samples, characterized by: include: An acquisition module is configured to: acquire a data set, wherein the data set is a low-frequency remote sensing image with a known small target label; A training module is configured to: input the data set into the small target detection model, train the model, and obtain a trained small target detection model; The detection module is configured to: obtain a remote sensing image to be detected, input the remote sensing image to be detected into a trained small target detection model, and obtain a small target classification label and a position label of the remote sensing image to be detected; During the training process, the small target detection model is used to: extract multi-scale features from the input image, perform feature enhancement on the extracted multi-scale features; extract the region of interest from the enhanced features, and dynamically divide the region of interest into positive samples according to the set strategy; use the divided positive and negative sample training model to classify the region of interest and perform bounding box regression processing to obtain the small target detection result; The data set is input into the small target detection model, and the model is trained to obtain a trained small target detection model, wherein the trained small target detection model includes: A feature extraction module, a feature enhancement module and an area of ​​interest generation module connected in sequence; the output end of the area of ​​interest generation module is connected to the input end of the classification module and the input end of the regression module respectively; The region of interest generation module adopts a dynamic positive sample division strategy; the dynamic positive sample division strategy includes: (1) Generate region of interest proposals through the RPN network; (2) Using the normalized area measure A i To express the effect of target size on the i-th target threshold: Among them, w i and h i Represents the width and height of the true bounding box of the i-th target, and the area is normalized using the maximum value max and minimum value min functions; (3) Evaluate the quality Q of the proposals of the region of interest by calculating the average and median of the proposals of the region of interest i , Among them, the median function calculates the median, the mean function calculates the mean, and c_p i Represents all region of interest proposals for the i-th target; (4) In order to avoid losing high-quality samples of large targets in low-frequency samples, the positive sample threshold of the target with the highest A / Q value is set to the maximum positive sample threshold τ specified by the network. max , and derive the conversion ratio β to determine the positive sample threshold of other targets; the ratio β is described as follows: Among them, max is used to determine the maximum value of the A / Q ratio; Then, the positive sample threshold of the i-th target is obtained, denoted as T i : Among them, τ min Represents the minimum value of the positive sample threshold assigned to each target, adjusted according to the dataset τ min , to determine its optimal value; (5) All samples of each target whose IoU value is greater than the positive sample threshold are classified as positive samples; The regression module includes: a position regression branch, a direction regression branch and a shape regression branch; Through the three branches of position regression branch, direction regression branch and shape regression branch, we get the feature map customized for different semantic interpretation parameters of the regression branch. It is expressed as: The superscripts l, o, and s represent the position, direction, and shape branches, respectively. represents coordinate convolution, refers to the coordinate attention mechanism, F r represents the feature map of the regression branch, AM represents the pixel-level supervision obtained by using the affine transformation of the proposal using the prediction result of the previous branch, represents a feature map customized for the coordinate branch, Represents the feature map customized for the direction branch.

Citation Information

Patent Citations

  • Remote sensing image small target detection method and system based on multi-scale features

    CN115984712A

  • Anchor-frame-free remote sensing image rotating target detection method under attention mechanism

    CN118379617A