Directed small target detection method based on semantic differentiation
By adopting a directed small object detection method based on semantic differentiation in remote sensing images, the problems of structural information loss, semantic confusion and sample scarcity of small object detection in remote sensing images are solved, and high-precision and high-efficiency small object detection are achieved.
Patent Information
- Application Number
- CN202510220871.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
Small object detection in remote sensing images has problems such as loss of structural information, semantic confusion and scarcity of samples, resulting in inadequate detection accuracy and efficiency.
A directed small object detection method based on semantic differentiation is adopted to build a detection network including backbone network, semantic differentiation branches and directed detection branches. The target response and class feature discrimination of the foreground area are enhanced through semantic differentiation branches, and the model's fitting ability to very small-size targets is improved by using small-objective center metrics and instance-level recalibration strategies.
It significantly improves the detection accuracy of small objects, while maintaining detection efficiency, which can effectively alleviate the problems of semantic confusion and sample scarcity.
Smart Images

Figure CN120071152A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a directed small target detection method based on semantic differentiation. Background Art
[0002] Remote sensing images contain a large number of small-sized target instances, and their accurate detection plays a very important role in remote sensing scene understanding and downstream task applications. However, due to the extremely limited target area and the unique top-down imaging perspective, the structural information of the target is severely lost, resulting in difficulty in extracting discriminative features. In addition, when the target is in a complex background, the semantic confusion problem between foreground categories and between foreground and background is more prominent, further causing frequent false alarms and misdetections, which greatly restricts the development of the small target detection task in remote sensing images.
[0003] To alleviate the above challenges, researchers often assign higher weights to the target area through an attention mechanism, or highlight the feature representation of the foreground area by means of a hierarchical feature fusion strategy. However, due to the lack of direct and effective supervision information, this strategy can only be completed through implicit model optimization, so the generalization performance cannot be guaranteed. In addition, these methods often rely on complex structural designs, which increase the number of parameters while reducing the optimization efficiency of the model. Some scholars also draw on excellent feature extraction paradigms in the field of general object detection to improve the backbone network design for small target detection, such as the Large Selective Kernel Network (LSKNet) based on the long-distance prior knowledge of remote sensing images. Although these algorithms have improved the weak feature representation of small targets to a certain extent, the heuristic design strategy is still difficult to alleviate the detection bottleneck caused by the inherent semantic ambiguity of the target and poses a huge challenge to the current detectors.
[0004] On the other hand, a difficult problem in the task of directed small target detection is the scarcity of training samples. Since very few prior anchor boxes have an overlapping area with the target region greater than the preset positive sample threshold, and the prior feature points falling within the positive sample region are also very scarce due to the limited area of the small target itself, it is difficult to assign enough samples to the small target whether it is a sample assignment strategy based on overlap measurement or a sample assignment strategy based on distance measurement. As a result, due to the lack of training samples, the model is difficult to effectively fit the target region, and the classification branch and the regression branch are also in a sub-optimized state. To address this problem, previous studies have all been committed to selecting more representative samples, that is, selecting higher-quality positive samples from the candidate samples to improve the regression accuracy and classification accuracy of the candidate boxes. Some other methods gradually increase the overlapping area between the candidate box and the target box by adopting iterative regression, which reduces the optimization difficulty of the regression branch to a certain extent. Generally speaking, even if these can select better candidate samples, it cannot be guaranteed that the detection model can be optimized from a limited number of samples, that is, the above methods cannot fundamentally alleviate the challenge of scarce small target samples. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention provides a directed small target detection method based on semantic differentiation. A directed small target detection network based on semantic differentiation is constructed, which mainly includes three parts: a backbone network, a semantic differentiation branch, and a directed detection branch. Among them, the semantic differentiation branch completes differentiation by selecting example features with higher detection quality to guide a custom category-related convolution kernel, and then convolves with the input hierarchical features, which can enhance the target response in the foreground region and the discriminability of target features of different categories; during network training, a small target center metric and an instance-level recalibration strategy are adopted, and by mining potential high-quality samples and dynamically adjusting the regression loss weights of the samples, the model's fitting for extremely small-sized targets is enhanced. The present invention can significantly improve the small target detection accuracy while maintaining the detection efficiency.
[0006] A directed small target detection method based on semantic differentiation, characterized by the following steps:
[0007] Step 1: Construct a training dataset based on a publicly available directed small target detection image dataset;
[0008] Step 2: Construct a directed small target detection network based on semantic differentiation, which mainly includes three parts: a backbone network, a semantic differentiation branch, and a directed detection branch; among them, the backbone network obtains the feature map of the input image, the semantic differentiation branch performs category differentiation by guiding randomly initialized semantic embeddings through loss constraints, and the directed detection branch outputs the target prediction box, including the category and location information of the target;
[0009] Step 3: Use the images in the training dataset obtained in Step 1 as inputs to train the directed small object detection network constructed in Step 2. During training, use the small object center metric and instance-level recalibration strategy.
[0010] Step 4: Input the test dataset images into the trained directed small object detection network. The network outputs target prediction boxes, and use the non-maximum suppression algorithm to delete redundant prediction boxes to obtain the final directed small object detection results of the test set images.
[0011] Specifically, the publicly available directed small object detection image dataset uses the SODA-A dataset or the Tiny-DOTA dataset.
[0012] Specifically, the specific process of Step 1 is as follows: For the images in the publicly available directed small object detection image dataset, perform sliding window cropping on them with a fixed step size to obtain image patches of the same size. Use the annotation boxes cropped in the same way as the annotation boxes of the image patches. All image patches and their annotation boxes constitute the training dataset.
[0013] Specifically, the backbone network mainly includes ResNet-50 and the Feature Pyramid Network. The specific process of the backbone network for obtaining the feature map of the input image is as follows: The input image is extracted through ResNet-50 and the Feature Pyramid Network to obtain hierarchical features, and then the foreground category number is aligned through 1×1 convolution to obtain the aligned feature map. Next, use global average pooling to obtain the channel weight vector. Finally, normalize each category through two fully connected layers to obtain the normalized weights.
[0014] Specifically, the semantic differentiation branch consists of a semantic initialization module, a semantic differentiation loss, and a semantic activation module. Among them, the semantic initialization module converts the input hierarchical features into a feature group distributed according to categories; subsequently, by initializing category-related semantic units and simultaneously selecting sample points with higher classification and regression quality as example features, the semantic differentiation loss completes the semantic unit differentiation by optimizing the similarity between the example features and the semantic units; finally, in the semantic activation module, the differentiated semantic units will be used as dynamic convolution kernels to perform convolution operations on the input features to obtain discriminative category activation features.
[0015] Specifically, the small object center metric and instance-level recalibration strategy are specifically as follows:
[0016] Step (1): Calculate the center metric value centerness of the sample according to the following formula:
[0017]
[0018] where t * 、b * 、l* , r * respectively represent the normalized distance values of the current sample's position from its responsible target in the up, down, left, and right directions, representing the regression target vector of the positive sample
[0019] Step (2): Calculate the improved target center metric value c according to the following formula * :
[0020]
[0021] where γ is a non-linear scale factor, taking positive integer values, and cntr * satisfies cntr * = centerness 2 ; Step (3): Calculate the adaptive regression weight vector w of the i-th instance according to the following formula i :
[0022]
[0023] where represents the center metric vector composed of the center metric values of all positive samples of the i-th instance, and μ i represents the mean of the center metric vector of the i-th instance, and σ i represents the standard deviation of the center metric vector of the i-th instance;
[0024] Step (4): Calculate the regression loss according to the following formula
[0025]
[0026] where P i represents the number of positive samples of the i-th instance, w ij represents the j-th component of the adaptive regression weight vector w of the i-th instance, l i represents a common loss function, reg and represents the true value of the j-th sample regression vector of the i-th instance, and t ij represents the predicted value of the j-th sample regression vector of the i-th instance;
[0027] Step (5): Calculate the total loss of the network model according to the following formula
[0028]
[0029] where α diff , α fcos , α sd respectively represent the weights of the semantic differentiation branch loss, the detection branch loss, and the semantic differentiation loss, Indicates the loss of the semantic differentiation branch, Indicates the loss of the directed detection branch, Indicates the semantic differentiation loss.
[0030] Furthermore, in step 4, first, the images in the test dataset are processed by sliding window cropping with a fixed step size to obtain image patches of the same size. Then, all the image patches are input into the trained directed small object detection network, and the network outputs target prediction boxes. Finally, these prediction boxes are mapped to the original image, and the non-maximum suppression algorithm is used to delete redundant prediction boxes to obtain the final directed small object detection results of the test set images.
[0031] Furthermore, the common loss functions include Smooth-L1 loss and IoU loss.
[0032] Furthermore, the scale factor γ takes a value of 2.
[0033] The beneficial effects of the present invention are as follows: Due to the adoption of the semantic differentiation branch, category-sensitive information is embedded into the semantic units corresponding to each category in a supervised learning manner, which can highlight the feature responses of foreground categories through dynamic convolution, alleviate the problem of semantic confusion of small objects in remote sensing images, and improve the detection accuracy of small objects; due to the adoption of the optimized small object center metric, potential high-quality samples can be further mined, better solving the problem of scarce small object samples and increasing the number of small object samples in the optimization process; due to the construction of an instance-level recalibration strategy based on the optimized small object center metric, dynamically adjusting the loss weights of each sample during the regression process can effectively improve the under-optimization problem for extremely small-sized objects and improve the detection accuracy of small objects; the semantic differentiation branch and the optimized small object center metric of the present invention can also be embedded into the existing single-stage detection framework based on center point modeling, effectively improving the directed small object detection accuracy for remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic diagram of the network structure of the directed small object detection network based on semantic differentiation of the present invention;
[0035] Figure 2 is an example of the result image of directed small object detection on the SODA-A dataset using the method of the present invention;
[0036] Figure 3 is an example of the result image of directed small object detection on the Tiny-DOTA dataset using the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] The present invention will be further described below in conjunction with the drawings and embodiments. The present invention includes but is not limited to the following embodiments.
[0038] Aiming at the semantic confusion between foreground categories and between foreground and background, as well as the scarcity of training samples in the small target detection task of remote sensing scenes, the present invention provides a directed small target detection method based on semantic differentiation. The specific implementation process is as follows:
[0039] 1. Construct a training dataset based on the publicly available directed small target detection image dataset
[0040] Currently, typical publicly available directed small target detection datasets are SODA-A and Tiny-DOTA. Among them, the SODA-A dataset contains 2,513 images, 9 categories, and 872,069 instances, and the Tiny-DOTA dataset contains 2,423 images, 8 categories, and 332,609 instances. The format of the annotation box is a directed rectangle box.
[0041] To construct the training dataset for network training, the original images in SODA-A or Tiny-DOTA are cropped by a fixed step size to obtain image patches of the same size, and the annotation boxes after cropping in the same way are used as the annotation boxes of the image patches. All image patches and their annotation boxes constitute the training dataset. For SODA-A, the original images are cropped to obtain image patches with a size of 800×800, the sliding window size for cropping is 800×800, and the step size is 650; for Tiny-DOTA, the sliding window size used for cropping is 1,024×1,024, and the step size is 824. These image patches are used as network training images.
[0042] 2. Construct a directed small target detection network based on semantic differentiation
[0043] To better achieve directed small target detection, the present invention designs a directed small target detection network model based on semantic differentiation, as Figure 1 shown. It mainly includes three parts: a backbone network, a semantic differentiation branch, and a directed detection branch. Among them, the backbone network obtains the feature map of the input image, the semantic differentiation branch guides the randomly initialized semantic embedding for category differentiation through loss constraints, and the directed detection branch outputs the target prediction box, that is, the small target detection result, including the category and location information of the target. Specifically:
[0044] The backbone network mainly includes ResNet-50 and a feature pyramid network. The specific process of the backbone network obtaining the feature map of the input image is as follows: the input image is extracted by ResNet-50 and the feature pyramid network to obtain hierarchical features, and then the foreground category number is aligned through 1×1 convolution to obtain the aligned feature map. Then, global average pooling is used to obtain the channel weight vector. Finally, each category is normalized through two fully connected layers to obtain the normalized weight. As Figure 1 shown in the box (a) part.
[0045] The semantic differentiation branch consists of a semantic initialization module, a semantic differentiation loss, and a semantic activation module. The semantic initialization module converts the input hierarchical features into a feature group distributed according to categories. Subsequently, by initializing the category-related semantic units, sample points with high classification and regression quality are selected as sample features. The semantic differentiation loss completes the semantic unit differentiation by optimizing the similarity between sample features and semantic units. Finally, in the semantic activation module, the differentiated semantic units will be used as dynamic convolution kernels to perform convolution operations with the input features to obtain discriminative category activation features. This design can not only alleviate the confusion of foreground and background features, but also enhance the feature distinction of targets of different categories. Figure 1 As shown in the middle frame (b).
[0046] The directional detection branch receives the differentiated discriminative features and uses parallel convolutional layers to obtain the final classification score, regression offset and center metric. Figure 1 As shown in the middle frame (c).
[0047] 3. Network training
[0048] Taking the images in the training dataset obtained in step 1 as input, the directed small object detection network constructed in step 2 is trained, and the small object center measurement and instance-level recalibration strategy are adopted during training.
[0049] This paper designs an optimized center metric to improve the response strength of the central area of small targets, thereby mining more potential samples that are relatively deviated from the target center but still have high classification and regression quality, thereby alleviating the challenge of lack of small target samples. At the same time, the instance-level recalibration strategy designed on this basis can dynamically adjust the loss weight of each sample in the regression process and enhance the model's fit to small-sized targets, as follows:
[0050] (1) Calculate the centrality of the sample as follows:
[0051]
[0052] Among them, t * 、b * , l * 、r * Respectively represent the normalized distance value of the current sample location from the target in the up, down, left, and right directions, and represent the regression target vector of the positive sample
[0053] (2) Calculate the optimized target center metric. In order to alleviate the limitation of the above center metric that the activation area is only concentrated in the target center area, an improved target center metric is designed. The improved target center metric value c is calculated as follows:* :
[0054]
[0055] Among them, γ is a non-linear scale factor, which can take any positive integer value, such as 2, and cntr * satisfies cntr * = centerness 2 ;
[0056] (3) Calculate the adaptive regression weight. Considering that when calculating the loss of the regression branch, the central measure of each sample will be used as the loss weight to re-weight the overall regression loss, an adaptive sample weight is designed on the basis of the above-expanded samples. That is, calculate the adaptive regression weight vector w of the i-th instance according to the following formula i :
[0057]
[0058] Among them, represents the central measure vector composed of the central measure values of all positive samples of the i-th instance, and μ i represents the mean of the central measure vector of the i-th instance, and σ i represents the standard deviation of the central measure vector of the i-th instance;
[0059] (4) Calculate the regression loss according to the following formula
[0060]
[0061] Among them, P i represents the number of positive samples of the i-th instance, w ij represents the j-th component of the adaptive regression weight vector w i of the i-th instance, and l reg represents a common loss function (such as Smooth-L1 loss, IoU loss), represents the true value of the j-th sample regression vector of the i-th instance, and t ij represents the predicted value of the j-th sample regression vector of the i-th instance;
[0062] (5) Calculate the total loss of the network model according to the following formula
[0063]
[0064] Among them, α diff , α fcos , α sd represent the weights of the differentiation branch loss, the directed detection branch loss, and the semantic differentiation loss respectively, Represents the loss of the semantic differentiation branch,
[0065] Represents the loss of the baseline detection branch, Represents the semantic segmentation loss.
[0066] Specifically, The calculation formulas of are as follows:
[0067]
[0068] Among them, and respectively represent the classification, regression, and center metric optimization losses in the semantic differentiation branch; and respectively represent the classification, regression, and center metric optimization losses in the directed detection branch; Θ * and Θ respectively represent the true value and the predicted value of the semantic unit to be differentiated.
[0069] 4. Obtain the directed small target detection result
[0070] Input the test dataset image into the trained directed small target detection network. The network outputs target prediction boxes. Use the non-maximum suppression algorithm to delete redundant prediction boxes to obtain the final directed small target detection result of the test set image. It is also possible to first perform sliding window cropping on the images in the test dataset with a fixed step size to obtain image patches of the same size, then input all the image patches into the trained directed small target detection network. Then, map the prediction boxes output by the network to the original image and use the non-maximum suppression algorithm to delete redundant prediction boxes to obtain the final directed small target detection result of the test set image.
[0071] Considering that the input image has a high resolution, it is possible to crop the image both in the training and testing stages. After inputting it into the trained directed small target detection network to obtain the inference result of the sub-image, map it to the original image, and at the same time use the non-maximum suppression algorithm to delete redundant prediction boxes to obtain the detection result of the test set image.
[0072] To verify the effectiveness of the method of the present invention, a simulation experiment is conducted, and the official provided evaluation code is used to calculate the final performance metrics. The operating environment is: a 4-card STH GPU server (CPU is Intel Xeon E5-2698, GPU is RTX 2080Ti with 12G), the operating system of the server is Ubuntu 16.04.5 LTS, and the experimental code is developed based on MMRotate 0.3. Figure 2The result images of directed small target detection on the SODA-A dataset are given. Among them, the first image in the first row is an aircraft target, the second image in the first row is a swimming pool and car targets, the third image in the first row is a helicopter target, and the fourth image in the first row is a ship target; the first image in the second row is a car target, the second image in the second row is a large vehicle and container targets, the third image in the second row is a swimming pool target, and the fourth image in the second row is a windmill target. Figure 3 The result images of directed small target detection on the Tiny-DOTA dataset are given. Among them, the first image in the first row is a ship target, the second image in the first row is an aircraft target, the third image in the first row is a ship target, and the fourth image in the first row is a car target; the first image in the second row is a ship target, the second image in the second row is a ship target, the third image in the second row is an aircraft and car targets, and the fourth image in the second row is an oil storage tank target. Figure 2 and Figure 3 It can be seen that the directed small target detection algorithm designed by the present invention can accurately detect the small-sized targets shown in the figure, and the angle prediction is also very accurate.
[0073] mAP (mean Average Precision) is selected to evaluate the effectiveness of the method of the present invention, and its definition is as follows:
[0074]
[0075] where c represents the target category index, AP c represents the detection accuracy of different categories, C represents the total number of categories included in the dataset to be trained. The SODA-A dataset contains 9 categories, while the Tiny-DOTA dataset has 8 categories. Keeping the backbone network as ResNet-50 and using FPN to obtain hierarchical features, comparing the method of the present invention with the baseline method FCOS, the results are shown in Table 1. It can be seen that the present invention can obtain higher detection accuracy compared with the baseline method.
[0076] Table 1
[0077]
[0078]
Claims
1. A method for detecting directed small objects based on semantic differentiation, characterized in that Here are the steps: Step 1: Build a training dataset based on a publicly available oriented small object detection image dataset; Step 2: Construct a directed small target detection network based on semantic differentiation, which mainly includes three parts: the backbone network, the semantic differentiation branch and the directed detection branch. The backbone network obtains the feature map of the input image, the semantic differentiation branch guides the randomly initialized semantic embedding to perform category differentiation through loss constraints, and the directed detection branch outputs the target prediction box, including the category and location information of the target. Step 3: Take the images in the training dataset obtained in step 1 as input and train the directed small object detection network constructed in step 2. The small object center measurement and instance-level recalibration strategy are used during training. Step 4: Input the test data set images into the trained directed small target detection network. The network outputs the target prediction box. The non-maximum suppression algorithm is used to delete redundant prediction boxes to obtain the final directed small target detection results of the test set images.
2. The method for detecting directed small objects based on semantic differentiation according to claim 1, characterized in that: The public directed small target detection image dataset adopts the SODA-A dataset or the Tiny-DOTA dataset.
3. The method for detecting directed small objects based on semantic differentiation according to claim 1, characterized in that: The specific process of step 1 is as follows: for images in the public directed small target detection image dataset, a sliding window cropping is performed on them using a fixed step size to obtain image blocks of the same size. The annotation boxes cropped in the same way are used as the annotation boxes of the image blocks. All image blocks and their annotation boxes constitute the training dataset.
4. The method for detecting directed small objects based on semantic differentiation according to claim 1, characterized in that: The backbone network mainly includes ResNet-50 and feature pyramid network. The specific process of the backbone network obtaining the feature map of the input image is as follows: the input image is extracted by ResNet-50 and feature pyramid network to obtain hierarchical features, and then the number of foreground categories is aligned through 1×1 convolution to obtain the aligned feature map. Then, global average pooling is used to obtain the channel weight vector. Finally, each category is normalized through two fully connected layers to obtain the normalized weight.
5. The method for detecting directed small objects based on semantic differentiation according to claim 1, characterized in that: The semantic differentiation branch consists of a semantic initialization module, a semantic differentiation loss and a semantic activation module, wherein the semantic initialization module converts the input hierarchical features into a feature group distributed according to categories; subsequently, by initializing the category-related semantic units, sample points with high classification and regression quality are selected as sample features, and the semantic differentiation loss completes the semantic unit differentiation by optimizing the similarity between the sample features and the semantic units; finally, in the semantic activation module, the differentiated semantic units will be used as dynamic convolution kernels to perform convolution operations with the input features to obtain discriminative category activation features.
6. The method for detecting directed small objects based on semantic differentiation according to claim 1, characterized in that: The small object center measurement and instance-level recalibration strategy are specifically as follows: Step (1): Calculate the centerness of the sample as follows: Among them, t * , b * , l * 、r * Respectively represent the normalized distance value of the current sample location from the target in the up, down, left, and right directions, and represent the regression target vector of the positive sample Step (2): Calculate the improved target center measurement value c as follows: * : Among them, γ is a nonlinear scaling factor, which takes a positive integer value, cntr * Meet cntr * =centerness 2 ; Step (3): Calculate the adaptive regression weight vector w of the i-th instance as follows: i : in, Represents the central measurement vector composed of the central measurement values of all positive samples of the i-th instance, μ i represents the mean of the center metric vector of the i-th instance, σ i represents the standard deviation of the center metric vector of the i-th instance; Step (4): Calculate the regression loss as follows Among them, P i represents the number of positive samples of the i-th instance, w ij represents the adaptive regression weight vector w of the i-th instance i The jth component of reg represents the commonly used loss function, represents the true value of the jth sample regression vector of the i-th instance, t ij Represents the predicted value of the jth sample regression vector of the i-th instance; Step (5): Calculate the total loss of the network model as follows: Among them, α diff , α fcos , α sd Represent the weights of semantic differentiation branch loss, detection branch loss and semantic differentiation loss respectively, represents the loss of the semantic differentiation branch, represents the loss of the directed detection branch, represents the semantic differentiation loss.
7. The method for detecting directed small objects based on semantic differentiation according to claim 1, characterized in that: In step 4, a fixed step size is first used to perform sliding window cropping on the images in the test data set to obtain image blocks of the same size, and then all image blocks are input into the trained directed small target detection network. The network outputs the target prediction box. Finally, these prediction boxes are mapped to the original image, and the non-maximum suppression algorithm is used to delete redundant prediction boxes to obtain the final directed small target detection results of the test set images.
8. The method for detecting directed small objects based on semantic differentiation according to claim 6, characterized in that: The commonly used loss functions include Smooth-L1 loss and IoU loss.
9. The method for detecting directed small objects based on semantic differentiation according to claim 6, characterized in that: The scale factor γ is set to 2.