Remote sensing image refined target detection method and system
Through the dual classification head network and error tag filtering module, combined with the Margin cross-entropy loss function, the problem of insufficient data set utilization and error tag interference in remote sensing image object detection is solved, and the accuracy and robustness of refined object detection of remote sensing image is improved.
Patent Information
- Application Number
- CN202510413144.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
Existing remote sensing image object detection methods are difficult to effectively integrate similar but different category definitions, resulting in limited training data scale, error labels interfere with model accuracy, and traditional cross-entropy loss cannot effectively distinguish fine categories, resulting in limited detection accuracy.
The dual classification head network and error label filtering module are used to train the model through mixed samples, filter error labels dynamically, and use the Margin cross-entropy loss function to increase the classification boundary, improving the model's refined object detection capability.
It significantly improves the accuracy and robustness of the refinement object detection of remote sensing images, makes full use of similar data sets, reduces the impact of error labels, enhances classification boundaries, and achieves higher detection accuracy and faster inference speed.
Smart Images

Figure CN120339861A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method and system for fine-grained object detection in remote sensing images. Background Art
[0002] The following problems exist in current remote sensing image object detection: Insufficient utilization of similar data sets. Existing methods are difficult to effectively integrate similar data sets with different category definitions, resulting in limited training data scale. Incorrect labels interfere with the model accuracy. Large-scale manual annotation is prone to introducing label noise, and existing methods lack a dynamic filtering mechanism for incorrect labels. Fine classification is difficult. In scenarios with small inter-class differences and large intra-class differences, the traditional cross-entropy loss cannot effectively distinguish fine categories. Existing technologies such as Oriented R-CNN can detect rotated objects, but do not solve the above problems, resulting in limited detection accuracy. Summary of the Invention
[0003] To solve the technical problems in the background art, the present invention proposes a method and system for fine-grained object detection in remote sensing images.
[0004] A method for fine-grained object detection in remote sensing images proposed by the present invention includes:
[0005] Obtain a remote sensing image data set, where the remote sensing image data set includes a main data set and a reference data set;
[0006] Construct an object detection model, and use the remote sensing image data set to train the object detection model to obtain a trained object detection model;
[0007] Input the to-be-detected remote sensing image into the trained object detection model to obtain an object detection result, where the object detection result includes the corresponding position and category of the object in the to-be-detected remote sensing image.
[0008] Preferably, the object detection model includes a dual-classification head network and an incorrect label filtering module; the dual-classification head network includes a shared feature layer and two fully connected layers connected in parallel, and the two fully connected layers respectively correspond to the output of the categories of the main data set and the reference data set; the incorrect label filtering module is used before calculating the classification loss in the training stage of the dual-classification head network to filter the annotations with category confusion that occur in the actual annotation of remote sensing images.
[0009] Preferably, the training process of the object detection model specifically includes:
[0010] Input the mixed samples of the main data set and the reference data set into the dual-classification head network. After the dual-classification head network extracts general features, perform classification predictions on the main data set and the reference data set through two independent fully connected layers respectively, and calculate the cross-entropy loss of the two classification heads;
[0011] Before the classification loss, the wrong labels are dynamically filtered through a wrong label filtering module;
[0012] The model parameters are jointly optimized through the Margin cross-entropy loss function and the SGD optimizer, and finally the trained object detection model is output.
[0013] Preferably, the mixing ratio of the mixed samples of the main data set and the reference data set is 4:1, the loss weight ratio of the two classification heads in the two-classification head network is 1:1, and the total classification loss of the two-classification head network is the weighted sum of the two.
[0014] Preferably, the calculation formula of the Margin cross-entropy loss function is specifically:
[0015]
[0016] where x i ∈R d represents the depth feature belonging to the y i th class, W and b are the weights and biases of the last fully connected layer; POS indicates that the sample is a positive sample, Δ is a constant.
[0017] Preferably, the dynamic filtering of wrong labels by the wrong label filtering module specifically includes: filtering samples with inconsistent predicted categories and true labels and a confidence level higher than α based on the comparison between the prediction confidence and the threshold α.
[0018] Preferably, the threshold α adjustment strategy is: in the initial training stage, α is set to 1.0 in the first 10 rounds, and then decreased by 0.01 every 1 round until the optimal threshold α = 0.9 is reached.
[0019] Preferably, the class labels of the reference data set and the main data set are isolated by ID mapping.
[0020] A refined object detection system for remote sensing images proposed by the present invention includes:
[0021] A data acquisition module for acquiring a remote sensing image data set, where the remote sensing image data set includes a main data set and a reference data set;
[0022] A data processing module for constructing an object detection model and training the object detection model using the remote sensing image data set to obtain a trained object detection model;
[0023] An object detection module for inputting a to-be-detected remote sensing image into the trained object detection model to obtain an object detection result, where the object detection result includes the corresponding positions and categories of the objects in the to-be-detected remote sensing image.
[0024] In the present invention, a refined object detection method and system for remote sensing images are proposed. A remote sensing image dataset is acquired, and the remote sensing image dataset includes a main dataset and a reference dataset. A target detection model is constructed, and the remote sensing image dataset is used to train the target detection model to obtain a trained target detection model. The to-be-detected remote sensing image is input into the trained target detection model to obtain a target detection result, and the target detection result includes the corresponding position and category of the target in the to-be-detected remote sensing image. A dual classification head is proposed, and different classification heads are trained on different datasets respectively, enabling similar data with different category definitions to participate in training together, thereby effectively utilizing similar data and significantly improving the accuracy of refined object detection for remote sensing images. An error label filtering method is designed to reduce the impact of error labels on the training of the refined object detection model for remote sensing images. By using Margin cross-entropy loss to increase the classification boundary, the robustness of refined object detection for remote sensing images is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic structural diagram of the working process of a refined object detection method for remote sensing images proposed by the present invention;
[0026] Figure 2 It is a schematic comparison diagram between the dual classification head and the ordinary detection head of a refined object detection method for remote sensing images proposed by the present invention;
[0027] Figure 3 It is a schematic diagram of error label filtering of a refined object detection method for remote sensing images proposed by the present invention;
[0028] Figure 4 It is a schematic diagram of the classification boundary between traditional cross-entropy and Margin cross-entropy (right) of a refined object detection method for remote sensing images proposed by the present invention;
[0029] Figure 5 It is a sample diagram of a refined object detection competition dataset of a refined object detection method for remote sensing images proposed by the present invention;
[0030] Figure 6 It is a schematic diagram of the visualized detection result of a refined object detection method for remote sensing images proposed by the present invention on the refined object detection competition dataset;
[0031] Figure 7 It is a schematic diagram of the detection result of a refined object detection method for remote sensing images proposed by the present invention on the competition dataset;
[0032] Figure 8 It is a schematic diagram of the system architecture of a refined object detection system for remote sensing images proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0033] Reference Figure 1-8 , a refined target detection method for remote sensing images proposed by the present invention includes the following steps:
[0034] S1: Obtain a remote sensing image data set, which includes a main data set and a reference data set.
[0035] S2: Construct a target detection model, and use the remote sensing image data set to train the target detection model to obtain a trained target detection model.
[0036] In this embodiment, the target detection model includes a dual-classification head network and an incorrect label filtering module; the dual-classification head network includes a shared feature layer and two fully connected layers connected in parallel, and the two fully connected layers respectively correspond to the outputs of the categories of the main data set and the reference data set; the incorrect label filtering module is used before calculating the classification loss in the training stage of the dual-classification head network, and is used to filter the mislabeled annotations that occur during the actual annotation of remote sensing images.
[0037] In this embodiment, the dual-classification head network uses the LSKNet-t backbone network as the main architecture. Two fully connected layers are connected in parallel at the end of the shared feature layer in the LSKNet-t backbone network. In the training stage of the dual-classification head network, after the end of the shared feature layer is connected to the incorrect label filtering module, two fully connected layers are connected in parallel. The shared feature layer is responsible for extracting the general semantic features of multiple tasks, and the parallel fully connected layers map the features to the exclusive decision spaces of each task respectively, realizing the collaborative optimization of cross-task knowledge transfer and computational efficiency.
[0038] Specifically, during training, the class score obtained by inference is first passed through an additional SoftMax operation to obtain the class and confidence of the target. If the predicted class is different from the ground truth label and the confidence is higher than a certain threshold, then this label is considered an incorrect label, and we use the corresponding predicted value and ground truth after filtering the incorrect label to calculate the classification loss.
[0039] An example of the label filtering process is as follows: As Figure 3 shown in the figure, the ground truth in the figure is the 0th class, the prediction is the 1st class, and the confidence is 0.9 (assuming the given threshold is 0.88). Since 0.9 is higher than the given filtering threshold, this label is considered an incorrect label. By setting different filtering thresholds to control the proportion of filtered targets, this method can filter most of the incorrect labels, thereby reducing the impact of incorrect labels on model training. At the same time, this method will not affect the training of the network in the initial stage, because at this stage, the classification branch of the network starts training from random initialization, and at this time the network is far from converging. After passing through the SoftMax operation, the predicted class confidence is usually lower than the confidence threshold set by the filtering module.
[0040] In this embodiment, the training process of the target detection model specifically includes:
[0041] Input the mixed samples of the main dataset and the reference dataset into the dual-classification head network. After the dual-classification head network extracts the general features, the classification predictions of the main dataset and the reference dataset are respectively performed through two independent fully connected layers, and the cross-entropy losses of the two classification heads are calculated;
[0042] Before the classification loss, dynamically filter the wrong labels through the wrong label filtering module;
[0043] Jointly optimize the model parameters through the Margin cross-entropy loss function and the SGD optimizer, and finally output the trained object detection model.
[0044] In this embodiment, dynamically filtering the wrong labels through the wrong label filtering module specifically includes: based on the comparison between the prediction confidence and the threshold α, filtering the samples whose predicted category is inconsistent with the true label and the confidence is higher than α.
[0045] Specifically, during the training process, perform the SoftMax operation on the prediction results of each sample, and calculate its highest confidence and the corresponding predicted category.
[0046] In this embodiment, the threshold α adjustment strategy is: in the initial training stage, set α = 0.7 in the first 10 rounds, and then increase it by 0.05 every 5 rounds until the best threshold α = 0.9 is reached.
[0047] In this embodiment, the mixing ratio of the mixed samples of the main dataset and the reference dataset is 4:1, the loss weight ratio of the dual-classification heads in the dual-classification head network is 1:1, and the total classification loss of the dual-classification head network is the weighted sum of the two.
[0048] Specifically, in the training stage, filter the class predictions and the true values through the class ID, calculate a classification loss for the class predictions belonging to the original training set and the corresponding true values separately, calculate another classification loss for the class predictions obtained from the reference dataset and the corresponding true values, and the sum of the two classification losses is the final classification loss.
[0049] In this embodiment, the calculation formula of the Margin cross-entropy loss function is specifically:
[0050]
[0051] where x i ∈R d , represents the depth feature belonging to the y i -th class, W and b are the weights and biases of the last fully connected layer; POS indicates that the sample is a positive sample, Δ is a constant.
[0052] Specifically, the schematic diagram of the Margin cross-entropy loss increasing the classification boundary is as follows Figure 4 shown. Green and orange represent two classifications respectively, and the dashed line is the classification decision boundary. After adding Margin (i.e., Δ), the classification boundaries of the classes are artificially widened, which will help improve the discrimination ability for subclasses. This method increases the classification boundaries between different classes, making the model have better fine-grained classification ability. Since this method only improves the classification branch, it can be used for other tasks that require fine-grained classification, including tasks such as horizontal box detection and instance segmentation, and has strong generality. At the same time, this method only improves the loss function in the training stage, so it does not affect the model inference speed at all.
[0053] In this embodiment, the class labels of the reference dataset and the main dataset are isolated through ID mapping.
[0054] S3: Input the remote sensing image to be measured into the trained object detection model to obtain the object detection result, where the object detection result includes the corresponding position and class of the object in the remote sensing image to be measured.
[0055] Specifically, since the performance of the two-stage object detection algorithm is usually better than that of the single-stage object detection algorithm in small object detection, combined with the characteristics that there are a large number of small objects in high-resolution remote sensing images, this paper selects the two-stage rotated object detection algorithm Oriented R-CNN as the basic framework. This method predicts the background of the oriented candidate on the basis of the horizontal bounding box, and the important information of the object shape, direction and other features that the surrounding background environment can provide. Therefore, we adopt LSKNet as the backbone network to dynamically adjust the receptive field of the feature extraction backbone to more effectively detect remote sensing objects.
[0056] In most fine-grained remote sensing object detection tasks, there are some publicly available datasets. These datasets are similar to the training dataset, with similar image scenes and similar objects, but only the definition of object class labels is different, and the refinement degree of class annotation is different. The fine-grained object detection competition dataset contains two similar datasets (training set and reference dataset). The current mainstream object detection methods' strategy for using similar data is to first pre-train with the similar dataset and then use the pre-trained model to initialize the model parameters. This method only uses the model trained on the similar dataset as the initialization parameters for the second training, and the reference dataset does not really participate in the training. Therefore, the utilization efficiency of this method for data is low, and the value of similar data is not fully utilized. For similar datasets, due to the different definitions of object classes, the same object may belong to different classes in the above two datasets. If the two datasets are forced to be merged, then the same object will belong to two different classes, which will mislead the network during training. Therefore, although their scenes are similar and the objects are similar, they still cannot be merged for training, and the dual classification head can well solve this problem.
[0057] Suppose the detection network is trained separately on two datasets with similar class definitions. Then the classification features required to distinguish foreground from background and different classes in the classification branches of these two detection models should also be similar. If we can design a reasonable structure to enable the model to train on both datasets simultaneously, then the data in different datasets will optimize the classification features in the same direction, and the similar datasets can improve the accuracy of the model. Based on the above assumption, we propose a brand-new dual classification head, where different classification heads are trained for different datasets, enabling similar data with different class definitions to participate in the model training throughout the process. Figure 2 The schematic diagrams of the existing ordinary detection head and the dual classification detection head proposed in this paper are given.
[0058] The dual classification head adds an additional fully connected layer (class2) after the features of the last layer in the classification branch to perform the classification task on the reference dataset. The added fully connected layer shares the classification features with the original classification fully connected layer. During the training phase, the class predictions and ground truths are filtered by the class ID. The class predictions belonging to the original training set and the corresponding ground truths are used to calculate a classification loss separately, and the class predictions obtained from the reference dataset and the corresponding ground truths are used to calculate another classification loss. The sum of the two classification losses is the final classification loss. During the inference phase, the additionally added fully connected layer is directly ignored, and only the class predictions of the original classification head are output. This method can make full use of the training model of similar data. For the original detection network, the reference data participates in the training of all layers except the last fully connected layer, and the similar data is fully utilized, effectively alleviating the problem of insufficient data and also improving the accuracy of the model. This method can be used for the co-training of any similar datasets and is also applicable to tasks such as object classification, object detection, instance segmentation, and semantic segmentation, and is expected to become a general method for the utilization of similar data. Since this method directly ignores the additionally added classification head during the inference phase, it has no impact on the inference speed of the model.
[0059] Large-scale high-precision manual annotations often have some annotation errors. In order to fit the wrong labels, the detection network will optimize the network parameters in the opposite direction, which will greatly affect the overall performance of the network. Therefore, it is necessary to reduce the impact of wrong labels on network training. Usually, the trained network can correctly predict the class of the target and obtain a high confidence. Based on this, we designed a prediction-based wrong label filtering module. This module is used before calculating the classification loss during the training phase of the network, mainly to filter a small number of mislabeled annotations that occur during the actual annotation of remote sensing images. During training, the class scores obtained from the inference are first passed through an additional SoftMax operation to obtain the class and confidence of the target. If the predicted class is different from the ground truth label and the confidence is higher than a certain threshold, then this label is considered a wrong label, and we use the corresponding predicted value and ground truth after filtering the wrong label to calculate the classification loss.
[0060] An example of the label filtering process is as follows: As Figure 3 shown, the ground truth in the figure is class 0, the prediction is class 1, and the confidence is 0.9 (assuming the given threshold is 0.88). Since 0.9 is higher than the given filtering threshold, this label is considered a wrong label. By setting different filtering thresholds to control the proportion of filtered targets, this method can filter most of the wrong labels, thereby reducing the impact of wrong labels on model training. At the same time, this method will not affect the training of the network in the initial stage, because in this stage, the classification branch of the network starts training from random initialization, and at this time the network is far from converging. After the SoftMax operation, the confidence of the predicted class is usually lower than the confidence threshold set by the filtering module.
[0061] This method has strong generality and can be used to filter out wrong labels in tasks such as target classification, target detection, and instance segmentation, thereby improving the model accuracy. Since this method only adds an additional SoftMax operation after network prediction, it has almost no impact on the model training speed and does not affect the model inference speed at all.
[0062] In the general object detection task, the algorithm pays more attention to the design of the loss function of the regression branch while ignoring the loss function of the classification branch. Due to the lack of traction from the requirements of fine object detection, almost all detection networks from the Faster RCNN detection network to the recent DINO detection network use the cross-entropy loss as the loss function of the classification branch. However, in fine object detection, there are many target categories, small differences between classes, and large differences within classes, and the traditional cross-entropy loss can no longer meet the requirements of fine detection. Drawing on the SVM idea, the detection accuracy is improved by increasing the classification boundary between different classes. Referring to the ArcFace loss function in face recognition and adding an angular margin, the angle between the deep features and their corresponding weights is punished in an additive manner, thereby enhancing both the intra-class compactness and the inter-class difference. Accordingly, we propose the Margin cross-entropy loss function by artificially subtracting a constant term from the network prediction value of the positive sample and then calculating the cross-entropy loss during training. The formula for the ordinary cross-entropy loss is shown as follows:
[0063]
[0064] The refined object detection competition dataset is the dataset released in the refined object detection track based on sub-meter images in the 2023 National Big Data and Computational Intelligence Challenge. The data source is the fused image of the visible light payload of the Four-dimensional Gaojing satellite (spatial resolution 0.5m). This dataset consists of three parts: a training set, a test set, and reference data. Among them, the training set and the test set contain 98 types of aircraft and ship targets, with 8000 and 3000 images respectively; the reference dataset contains 39 types of coarse-grained targets, with 11000 images. The size of all images is 1024×1024, and the target position consists of eight coordinate values of four corner points. The images and target examples are as Figure 5As shown in the figure. FAIR1M is a large-scale dataset for fine-grained object detection and recognition in remote sensing images. This dataset has 37 fine-grained object categories. For example, airplanes are subdivided into specific models such as Boeing 737, Boeing 747, Boeing 777, Boeing 787, Airbus A220, Airbus A321, Airbus A330, Airbus A350, C919, and ARJ21, and ships are subdivided into liquid cargo ships, dry cargo ships, fishing boats, cruise ships, tugboats, and engineering ships, etc. This dataset contains more than 1 million finely annotated and multi-angled distributed objects, including 15,000 remote sensing images. The size of each image ranges from 1000×1000 to 10,000×10,000 pixels, and the spatial resolution ranges from 0.3m to 0.8m. In the experiment, we used version 2.0 of this dataset, using the training set and the validation set to train the model and verify the accuracy of the algorithm, and using the DOTA dataset as a reference dataset. DOTA (Xia et al., 2018) is a dataset for object detection in large aerial images, including 15 common object categories such as airplanes, ships, storage tanks, baseball fields, tennis courts, basketball courts, ground runways, ports, bridges, large vehicles, small vehicles, helicopters, roundabouts, football fields, and basketball courts. The dataset contains 2,806 satellite or aerial images, and the size of each image ranges from 800×800 to 4000×4000 pixels, containing 188,000 objects of different scales, orientations, and shapes. In this paper, DOTA is used as the reference data for FAIR1M.
[0065] In this experiment, the average precision (mAP) is used to measure the detection performance at an IoU threshold of 0.5. When the intersection over union (IoU) of the target's bounding box and the ground truth is greater than the threshold IoU and the predicted class label is correct, the detection result is considered a true positive (TP).
[0066] The experimental environment includes: Intel i7 13700K CPU, NVIDIA GeForce RTX3090Ti GPU. Ubuntu20.04 system, and the code in this paper is completed based on the MMRotate toolbox.
[0067] In all experiments, we used the Stochastic Gradient Descent (SGD) optimizer with a learning rate of 0.001, a momentum of 0.9, a weight decay of 0.0001, and a batch size of 8. Other settings were kept consistent with the default settings of the Oriented R-CNN (Xie et al., 2021) network in the MMRotate toolbox. In the ablation experiments, to accelerate the experiment progress and verify the effectiveness of the network model as soon as possible, we used the smaller LSKNet-t as the feature extraction backbone network. The model was trained for 12 epochs. The feature fusion network used the Feature Pyramid Network (FPN), and only simple horizontal flipping was used as the data augmentation method. On the FAIR1M (Sun et al., 2022) dataset, LSKNet-s was used as the backbone network. In the performance comparison experiments, to improve the model accuracy and better compare with other methods, we used LSKNet-s as the backbone network. The training settings were the same as those in the performance comparison experiments of other methods, that is, an additional data augmentation method of image rotation was added, and the model was trained for 24 epochs. During the inference process, a preset confidence threshold of 0.02 was used to filter out the background, and Rotated Non-Maximum Suppression (R-NMS) with a threshold of 0.1 was used to filter out redundant detection boxes. The inference speed was obtained with a batch size of 1.
[0068] The dual classification head can effectively solve the problem of utilizing similar datasets, thus improving the model accuracy. As shown in Table 1, by introducing the dual classification detection head based on the baseline model, since this design can fully utilize the reference dataset for training, it is equivalent to increasing the scale of the dataset. At the same time, because there is no artificial class confusion in the reference dataset, it is more conducive to forming stable classification features. Therefore, after adding the dual classification head to the network, the model accuracy on the competition dataset increased from the initial 58.4% to 72.4%, with a significant accuracy improvement of 14%.
[0069] Table 1 Experimental results of the detection network on the refined object detection competition dataset
[0070]
[0071] Table 2 Experimental results of the detection network on the FAIR1M dataset
[0072]
[0073] Due to the large scale of the FAIR1M dataset, there are 37 refined target categories, including more than 1 million refined labeled remote sensing targets; while the DOTA dataset as a reference dataset has a smaller data scale, with only 15 categories and more than 180,000 labeled targets. Previous methods could not jointly train the two datasets to improve the model accuracy. After using the dual classification head we proposed, the DOTA dataset can be added as a reference dataset to the training process. Due to the differences in data scale and annotation fineness, theoretically, even if the DOTA dataset can be added to the training, its improvement in the overall model accuracy will be very limited. As shown in Table 2, on the FAIR1M dataset, after adding the dual classification head, our model still improves the detection accuracy by 0.4% because it can fully utilize the information of the DOTA dataset.
[0074] The wrong label filtering module improves the model accuracy by reducing the impact of wrong labels on training. As shown in Table 1, after adding the wrong label filtering module, the model accuracy on the competition dataset increases from 72.4% to 73.7%, an increase of 1.3%. Since there is no artificial label confusion in the FAIR1M dataset and the problem of wrong labels is not obvious, the wrong label filtering module is not used on this dataset.
[0075] Table 3 Influence of filtering threshold α on the accuracy of the detection network on the refined detection competition dataset
[0076]
[0077]
[0078] According to different filtering thresholds, the proportion of wrong labels filtered by the label filtering module also varies. The accuracy of the model on the competition dataset under different filtering thresholds is shown in Table 3. It can be seen from the table that different filtering thresholds have a great impact on the performance of the model. The best accuracy on the competition dataset is obtained when the filtering threshold is 0.9. As the filtering threshold gradually decreases, the labels that meet the filtering conditions gradually increase, the filtered labels gradually increase, and more wrong labels are filtered. At this time, the model accuracy gradually increases. When the filtering threshold continues to decrease, some correct labels are also filtered, so the model accuracy gradually decreases.
[0079] Table 4 Influence of boundary threshold Δ on the accuracy of the detection network on the refined detection competition dataset
[0080]
[0081] Margin cross entropy loss forces the network to enlarge the classification decision boundary during training, increasing the model detection accuracy. As shown in Table 1, after using Margin cross entropy loss, the model accuracy on the competition dataset continued to increase from 73.7% to 74.7%, an increase of 1%. As shown in Table 2, the model accuracy on the FAIR1M dataset continued to increase from 41.0% to 41.3%, an increase of 0.3%.
[0082] After adding the dual classification head structure and label filtering module to the baseline model, the accuracy of the detection network on the competition dataset under different boundary thresholds is shown in Table 4. It can be seen that the boundary threshold Δ of the loss function has a great influence on the model accuracy. When Δ is relatively small, the distance between the classification boundaries is not large, and the improvement of the model accuracy is limited. As Δ increases, the classification boundary gradually expands, and the accuracy of the model gradually improves. On the refined object detection competition dataset, the best accuracy can be obtained when Δ is 0.12. When Δ continues to increase, the model continues to expand the classification boundary, and the detection accuracy decreases. The accuracy of the detection network on the FAIR1M dataset under different boundary thresholds is shown in Table 5. After adding the Margin cross entropy loss, the detection accuracy is improved by 0.3%, and the best accuracy can be obtained when Δ is 0.10.
[0083] Table 5 The impact of the upper boundary threshold Δ on the detection network accuracy on the FAIR1M dataset
[0084]
[0085] The following is a performance comparison experiment on the refined object detection competition and the FAIR1M dataset after integrating the above three modules.
[0086] Table 6 shows the accuracy of the proposed method on the competition dataset. It can be seen from Table 6 that the detection accuracy of the proposed method is significantly improved, from 58.4% to 74.7%, an increase of 16.3%. After we use a larger backbone network, the inference speed is only slightly reduced, from 29.9 to 25.7 frames per second, while the inference accuracy of 1312 size is increased to 76.5%, and the speed is reduced to 18.1 frames per second. The model still has good real-time performance. The proposed method uses a smaller backbone network and fewer training techniques to win the first place in the "Refined Target Detection Based on Sub-meter Imagery" track of the 2023 National Big Data and Computational Intelligence Challenge. We surpassed all other teams. These teams tried to use many detection methods, use larger backbone networks, and many training techniques. Since the true value annotation of the semi-final test set is not yet public, the accuracy of other methods on this dataset cannot be obtained at present, but it can be seen from the competition that our method has the highest accuracy and faster inference speed.
[0087] As can be seen from Table 6, the method proposed in this paper can stably improve the model accuracy and achieve the best results on the competition dataset, which fully demonstrates the effectiveness of the method in this paper. The visual detection results of the method in this paper on the competition dataset are as Figure 6 shown. Different colors of the detection boxes in the figure represent different categories. It can be seen from this that whether it is for tiny airplanes or densely arranged ships, the method in this paper can achieve accurate detection, further verifying the effectiveness of the method in this paper.
[0088] Table 6 Experimental results of different methods on the refined object detection competition dataset
[0089]
[0090] Table 7 Experimental results of different methods on the FAIR1M dataset
[0091]
[0092] Table 7 shows the comparison results of the method in this paper with other rotated object detector methods on the FAIR1M dataset. The accuracy and speed of other methods are the results obtained by using the default settings based on the DOTA dataset in the MMRotate platform, training on the FAIR1M training set and testing on the validation set. Because the video memory occupancy is too large, STD (Yu et al. 2024) can only be trained using images with a size of 800. It can be seen from the table that the method in this paper simultaneously obtains the second highest accuracy of 41.3% and the highest inference speed of 29.1 frames per second. Since the comparison algorithms can only initialize the training parameters with the pre-trained models on the DOTA dataset, while the algorithm in this paper can directly use the DOTA data for training, the algorithm in this paper effectively improves the model accuracy.
[0093] The visual detection results on the FAIR1M dataset are as Figure 7 shown. Different colors of the detection boxes in the figure represent different target categories. It can be seen from this that whether it is various targets in the dataset or various road facilities and various sites, the method in this paper can achieve accurate detection and obtain good detection results, which fully demonstrates the effectiveness of the method in this paper.
[0094] Aiming at the problems existing in the refined object detection of remote sensing images, this paper proposes a dual-classification head structure, which solves the problem of ineffective utilization of similar data in the refined object detection of remote sensing images; designs an error label filtering method based on prediction, which reduces the impact of error labels on model training; defines the Margin cross-entropy loss, which improves the model accuracy by increasing the classification boundary. Finally, through experiments on the refined object detection competition dataset, FAIR1M dataset and DOTA dataset, it is proved that the method proposed in this paper significantly improves the detection accuracy and robustness of the model.
[0095] Meanwhile, the dual classification head proposed in this paper also has certain limitations. For training sets with relatively small data scales, there is a significant performance improvement after jointly training with similar data sets in the package. For training sets with large data scales themselves, even when adding similar data for joint training, the performance improvement is not obvious (only 0.4% improvement in the FAIR1M data set). In addition, the wrong label filtering module proposed in this paper can only target the situation of label category confusion and cannot filter mislabeled positions. In future research work, we will continue to study more methods for reusing similar data and filtering mislabeled positions to further improve the accuracy of fine-grained object detection in remote sensing images.
[0096] Referring to Figure 1-8 , a fine-grained object detection system for remote sensing images proposed by the present invention includes:
[0097] A data acquisition module for acquiring a remote sensing image data set, the remote sensing image data set including a main data set and a reference data set;
[0098] A data processing module for constructing an object detection model and training the object detection model using the remote sensing image data set to obtain a trained object detection model;
[0099] An object detection module for inputting a to-be-detected remote sensing image into the trained object detection model to obtain an object detection result, the object detection result including the corresponding position and category of the object in the to-be-detected remote sensing image.
[0100] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent replacements or changes, shall be covered by the protection scope of the present invention.
Claims
1. A refined target detection method for remote sensing images, characterized in that, Including: Obtain a remote sensing image dataset, where the remote sensing image dataset includes a main dataset and a reference dataset; Construct an object detection model, and use the remote sensing image dataset to train the object detection model to obtain a trained object detection model; Input the to-be-detected remote sensing image into the trained object detection model to obtain an object detection result, where the object detection result includes the corresponding position and category of the object in the to-be-detected remote sensing image.
2. The refined object detection method for remote sensing images according to claim 1, wherein The object detection model includes a dual classification head network and an incorrect label filtering module; the dual classification head network includes a shared feature layer and two fully connected layers in parallel, and the two fully connected layers respectively correspond to the output of the categories of the main dataset and the reference dataset; the incorrect label filtering module is used before calculating the classification loss in the training stage of the dual classification head network, and is used to filter the annotations with category confusion that occur in the process of actually annotating remote sensing images.
3. The refined target detection method for remote sensing images according to claim 2, characterized in that, The training process of the object detection model specifically includes: Input the mixed samples of the main dataset and the reference dataset into the dual classification head network. After the dual classification head network extracts general features, perform classification predictions on the main dataset and the reference dataset respectively through two independent fully connected layers, and calculate the cross-entropy loss of the two classification heads; Before the classification loss, dynamically filter incorrect labels through the incorrect label filtering module; Jointly optimize the model parameters through the Margin cross-entropy loss function and the SGD optimizer, and finally output the trained object detection model.
4. The refined target detection method for remote sensing images according to claim 2, characterized in that, The mixing ratio of the mixed samples of the main dataset and the reference dataset is 4:1, the loss weight ratio of the dual classification heads in the dual classification head network is 1:1, and the total classification loss of the dual classification head network is the weighted sum of the two.
5. The refined target detection method for remote sensing images according to claim 3, wherein The specific calculation formula of the Margin cross-entropy loss function is: where x i ∈R d represents the depth feature belonging to the y i -th class, W and b are the weights and biases of the last fully-connected layer; POS indicates that the sample is a positive sample, Δ is a constant.
6. The refined target detection method for remote sensing images according to claim 3, wherein The dynamically filtering incorrect labels through the incorrect label filtering module specifically includes: based on the comparison between the prediction confidence and the threshold α, filtering the samples whose predicted category is inconsistent with the true label and the confidence is higher than α.
7. The refined object detection method for remote sensing images according to claim 6, wherein The threshold α adjustment strategy is: in the initial training stage, set α = 1.0 in the first 10 rounds, and then reduce it by 0.01 every 1 round until the optimal threshold α = 0.9 is reached.
8. The refined target detection method for remote sensing images according to claim 1, wherein The category labels of the reference dataset and the main dataset are isolated through ID mapping.
9. A refined target detection system for remote sensing images, characterized in that, Including: A data acquisition module, used to obtain a remote sensing image dataset, where the remote sensing image dataset includes a main dataset and a reference dataset; A data processing module, used to construct an object detection model, and use the remote sensing image dataset to train the object detection model to obtain a trained object detection model; An object detection module, used to input the to-be-detected remote sensing image into the trained object detection model to obtain an object detection result, where the object detection result includes the corresponding position and category of the object in the to-be-detected remote sensing image.