Remote sensing target efficient grading and labeling method based on active learning

Through the efficient hierarchical labeling method of remote sensing targets based on active learning, the problems of inefficiency and inconsistent quality in the construction of remote sensing target data sets are solved, and efficient and low-cost labeling quality improvement is achieved, especially in fine granularity classification, which significantly reduces the labeling error rate.

CN120298917APending Publication Date: 2025-07-11NINGBO INSTITUTE OF TECHNOLOGY BEIHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510437245.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When building large-scale, high-precision remote sensing target data sets, the existing technology has problems such as low labeling efficiency, high cost and inconsistent labeling quality. Especially when fine granularity classification is required, the professional level of the labeling personnel has a direct impact on the labeling quality.

Method used

Efficient hierarchical annotation method based on active learning is adopted, and the labeling method of remote sensing targets is formulated based on hierarchical annotation templates, small-scale annotation data sets are constructed, deep neural network models are used to generate pseudo-labels and screen difficult samples, combine hierarchical annotation by experts and ordinary annotation personnel, optimize the allocation of annotation resources, and improve the labeling quality through iterative training models.

Benefits of technology

It significantly improves the labeling efficiency, reduces the workload of experts in the field, makes resource allocation more reasonable, reduces the labeling error rate, and significantly improves the labeling quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298917A_ABST
    Figure CN120298917A_ABST
Patent Text Reader

Abstract

The invention belongs to the cross technical field of computer vision and remote sensing technology, and particularly discloses a remote sensing target efficient grading labeling method based on active learning, which comprises the steps of labeling template grading formulation, small-scale labeling data set construction, large-scale data set labeling based on active learning and the like. And through a hierarchical labeling strategy and dynamic iterative optimization, the labeling efficiency and the labeling precision are remarkably improved, and the workload of experts is greatly reduced. The method can be widely applied to the fields of military, urban planning and disaster emergency, and low-cost and high-efficiency large-scale data set construction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of computer vision and remote sensing technology, and particularly relates to an efficient hierarchical annotation method for remote sensing targets based on active learning, which is particularly suitable for scenarios requiring large - scale fine annotation. Background Art

[0002] With the rapid development of artificial intelligence technology, especially the wide application of deep learning technology in the field of image recognition, the demand for large - scale high - quality data sets is increasing day by day. In the field of remote sensing, target detection and recognition technology has great application value in military reconnaissance, ocean monitoring, disaster assessment, etc. However, remote sensing image data has characteristics such as high resolution, complex scenes, and a wide variety of target types, making the traditional data set construction methods face huge challenges. Especially for complex targets such as ships, fine - grained classification (such as specific model recognition) not only requires a large amount of labeled data but also a highly professional knowledge background. Therefore, how to construct a large - scale, high - precision fine - grained remote sensing target data set within a limited time has become an urgent problem to be solved.

[0003] The fully manual annotation method is a traditional means of data set construction. It relies on a large number of annotators to carefully annotate the targets in remote sensing images. Although this method can ensure a certain annotation quality, its efficiency is low, the cost is high, and it is difficult to ensure the consistency of annotation. Especially when fine - grained classification is required, the professional level of annotators has a direct impact on the annotation quality.

[0004] It should be noted that the information disclosed in this background art section is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide an efficient hierarchical annotation method for targets based on active learning, so as to solve the problems of low annotation quality and high cost existing in the existing methods.

[0006] To achieve the above - mentioned purpose, the present invention adopts the following technical solutions:

[0007] An efficient hierarchical annotation method for remote sensing targets based on active learning, comprising the following steps:

[0008] S1 Hierarchical formulation of annotation templates: Determine annotation information and divide it into professional knowledge annotation information and common - sense annotation information;

[0009] S2 Construction of a small - scale labeled data set: Randomly extract a small - scale data set from a large - scale unlabeled data set and perform hierarchical manual annotation;

[0010] S3 performs large-scale dataset annotation based on active learning: Train a deep neural network model using a small-scale annotated dataset, then perform inference and prediction on the unannotated dataset to generate pseudo-labels and confidence scores. Screen difficult samples according to the confidence scores for hierarchical manual annotation, update the training set and retrain the model, and iterate until the preset annotation quality threshold is met.

[0011] Further, in step S2, the proportion of the small-scale dataset in the large-scale unannotated dataset is about 10%, which can be dynamically adjusted according to the difficulty of the dataset.

[0012] Further, when performing hierarchical manual annotation in step S2 and step S3, integrate the common-sense annotation information and expert annotation information hierarchically to optimize the allocation efficiency of annotation resources.

[0013] Further, in step S3, one of the following strategies is used to screen difficult samples:

[0014] (1) Least Confidence (LC) strategy: Select the samples with the lowest confidence in the model prediction;

[0015] (2) Margin strategy: Select the samples with the smallest difference in the probabilities of the top two categories predicted by the model;

[0016] (3) Entropy strategy: Select the samples with the largest information entropy of the classification probability distribution.

[0017] Further, in step S1, the hierarchical annotation includes:

[0018] (1) Common-sense annotation: The common-sense annotation information is annotated by ordinary annotators;

[0019] (2) Expert annotation: The professional knowledge annotation information is annotated by domain experts.

[0020] Further, in step S3, a weighted loss function is used during the model training process, and the weight allocation strategy is:

[0021] (1) The weight of expert-annotated samples is higher than that of pseudo-label samples;

[0022] (2) The weight of pseudo-label samples is positively correlated with their confidence scores.

[0023] Further, the method further includes a dynamic threshold adjustment step: Adjust the confidence screening threshold according to the model performance and the distribution complexity of the unannotated samples.

[0024] Further, the method further includes a pseudo-label correction step: Automatically verify the high-confidence pseudo-label samples, and only add the samples that pass the verification to the training set.

[0025] Furthermore, in the iterative optimization process, the block inference and result fusion technology is adopted to improve the processing efficiency of large-size remote sensing images.

[0026] Furthermore, the result also includes the steps of annotation result visualization and interactive correction: providing a visualization interface for experts to quickly correct pseudo-label errors and feedbacking the correction results to model training.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] 1. Compared with full manual annotation, the annotation efficiency is significantly improved;

[0029] 2. Greatly reduce the workload of domain experts and make the resource allocation more reasonable;

[0030] 3. Reduce the annotation error rate and significantly improve the annotation quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 : According to some embodiments of the present invention, a flowchart of an efficient hierarchical annotation method for remote sensing targets based on hierarchical annotation and active learning is shown;

[0032] Figure 2 : According to some embodiments of the present invention, a flowchart of an annotation algorithm for large-scale datasets based on active learning is shown;

[0033] Figure 3 : According to Embodiment 1 of the present invention, a schematic diagram of the annotation result is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The following further describes the technical features and advantages of the present invention in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0035] The basic concept of the embodiments of the present invention is: reducing the burden on experts through hierarchical annotation, and combining active learning to screen difficult samples to achieve a double improvement in annotation efficiency and quality.

[0036] Please refer to Figure 1 , which is a flowchart of an efficient hierarchical annotation method for remote sensing targets based on active learning according to the embodiments of the present invention, and its basic steps include:

[0037] S1 Hierarchical formulation of annotation templates:

[0038] S11 Determine annotation information, such as detection box positioning, occlusion status, and fine-grained categories;

[0039] S12 divides the professional knowledge annotation information and common sense annotation information. For example, the fine-grained categories are defined as professional knowledge annotation information, and the detection box positioning and occlusion status are defined as common sense annotation information.

[0040] S2 Construct a small-scale dataset:

[0041] S21 Randomly extract a small-scale dataset from the large-scale unlabeled dataset;

[0042] S22 Perform hierarchical manual annotation on the small-scale dataset.

[0043] S3 Perform large-scale dataset annotation based on active learning:

[0044] S31 Use the small-scale labeled dataset to train a deep neural network model;

[0045] S32 Perform inference and prediction on the unlabeled dataset to generate pseudo-labels and confidence scores;

[0046] S33 Screen difficult samples according to the confidence scores for hierarchical manual annotation and update the training set;

[0047] S34 Retrain the model and iterate until the preset annotation quality threshold is met.

[0048] Through the collaborative optimization of iterative optimization of the model and annotation data, the quality of pseudo-labels is gradually improved, the dependence on manual annotation is greatly reduced, and the annotation efficiency and annotation accuracy are significantly improved.

[0049] In the embodiments of the present invention, during manual annotation, a hierarchical annotation integration strategy is adopted to reduce the workload of experts. The specific strategy is as follows:

[0050] (1) Ordinary annotators annotate common sense information (such as target location, occlusion status, etc.);

[0051] (2) Domain experts focus on annotating professional knowledge annotation information, such as fine classification (ship model recognition, etc.).

[0052] The application scenarios of the embodiments of the present invention include but are not limited to: large-scale remote sensing dataset annotation.

[0053] Taking the land type patches as an example, a set of large-scale fine-grained remote sensing ground object classification datasets will be constructed to serve the research of deep learning fine-grained scene parsing algorithms. This dataset has the characteristics of wide coverage, complex land type system, and strict boundary accuracy requirements. In order to complete the efficient construction within a limited period, the present invention adopts a collaborative operation mode of "ordinary personnel annotating common sense information + domain experts annotating professional attribute information". Ordinary personnel execute the conventional fully manual annotation process, responsible for drawing the patch boundaries (polygons), recording the surface cover types (vegetation / water / artificial structures), and annotating basic geographical features such as the cloud cover occlusion ratio. "Domain expert annotation" focuses on the high-precision subdivision of land types (such as secondary land types like basic farmland / economic forest land / industrial land, etc.). In addition, in order to further reduce the number of annotations and the annotation cost, the present invention introduces the idea of active learning, and only the most valuable samples need to be annotated manually to improve the recognition ability of the network. Based on the above ideas, the embodiment of the present invention proposes an efficient hierarchical annotation method for remote sensing targets based on active learning.

[0054] Active Learning is a new artificial intelligence technology that emerged to address the problem of the lack of labeled data required for supervised learning in real application scenarios and the difficulty of large-scale annotation in a short period. In the traditional supervised learning framework, the information value of all labeled samples used for training is regarded as equal, so the dataset needs to be of a relatively large scale before the training algorithm can start. Active Learning differentiates the value of samples, queries the most useful unlabeled samples through a certain algorithm, and hands them over to experts for labeling. Then, the samples labeled by experts are used to train the algorithm to improve the model accuracy and also provide the latest algorithm for the next query. Active Learning is a semi-supervised learning framework that iteratively promotes data annotation and model training.

[0055] Learning from a small number of existing labeled samples results in poor model performance. In order to make full use of a large number of existing unlabeled remote sensing image samples and enable the model to learn richer information. Therefore, active learning is introduced. First, a target detector needs to be selected and designed. The embodiment of the present invention first compares the current mainstream single-stage target detectors such as SSD and YOLO series, and selects a target detector with better performance that is more suitable for the known small number of labeled samples and the remote sensing image field.

[0056] The unlabeled sample data is screened using a query strategy. The query stage is mainly divided into two steps: scoring and sampling. Compare the currently more general scoring strategies, compare the aggregation methods from target uncertainty to image uncertainty, and try to improve the existing scoring strategies to design a query strategy more suitable for remote sensing image target detection.

[0057] Sample the scored data and complete the sampling work according to the corresponding sampling strategy. The sampled data is mainly divided into two parts. The data with greater sampling uncertainty is marked manually, and the data with less sampling uncertainty is assigned pseudo-labels. The sampling process needs to consider the diversity of the data. By comparing and analyzing different sampling strategies, relevant sampling strategies are improved to ensure that the data distribution in the sampled samples is as consistent as possible with that in the unlabeled samples.

[0058] Summarize the selected dataset and a small amount of already labeled datasets. Through the adaptive learning mechanism, train the summarized dataset, design the fusion method of the summarized data and the learning parameters for different data and data volumes, and continuously optimize the model until the model performance reaches our expected goal or all the unlabeled data has been completely selected. The flowchart of the large-scale dataset annotation steps based on active learning in the embodiments of the present invention is as Figure 2 shown.

[0059] Among them, the query strategy can adopt the following several methods:

[0060] 1. Least Confidence (LC)

[0061] For the multi-classification problem of images, the least confidence sampling selects the most uncertain samples through the following formula:

[0062]

[0063] where represents the highest-class classification score among the classification scores of each category predicted by the model W for the given x. LC assumes that only the classification score of the single best-predicted category by the model is concerned. If this score is low, it is considered that the model's prediction of this sample is the most uncertain, that is, the confidence is the smallest, and thus this sample is sampled. Introduce the original LC strategy into object detection, and select the most uncertain samples through the following formula:

[0064]

[0065] where represents the category (with the highest classification score) to which the k-th candidate target predicted by the model W for the given input x belongs. We first calculate the uncertainty of all candidate targets, and then select the uncertainty score corresponding to the bounding box with the largest uncertainty among the N b candidate targets as the score corresponding to the entire image. Finally, sample the image sample with the largest uncertainty (the smallest confidence). Different from image classification, by first calculating the uncertainty of candidate targets and then aggregating to represent the uncertainty of image samples, it is more in line with the goal of object detection itself.

[0066] 2. Distance Metric

[0067] Different from the LC strategy that only considers the category information with the highest predicted classification score, the original formula of the sampling method based on distance metric (Margin) is defined as follows:

[0068]

[0069] Among them and respectively represent the top two categories with the highest predicted scores. The Margin-based strategy represents the uncertainty of a sample by considering the absolute value of the predicted score residual. It assumes that the larger this value is, the easier the sample is to predict and the lower the uncertainty; conversely, if this value is very small, it means that the existing model has a similar probability of predicting this sample into two categories, and it is more difficult for the model to accurately classify this sample, so the uncertainty is greater.

[0070] Introduce the original Margin strategy into object detection, and select the most uncertain samples through Equation (4), where and respectively represent the top two categories with the highest predicted scores for the k-th candidate object.

[0071]

[0072] 3. Information Entropy

[0073] Furthermore, as the number of categories in the dataset increases, the Margin method will ignore more output distribution information of the remaining categories. Information Entropy is a common method in information theory to measure the uncertainty of a signal. It measures the uncertainty based on the probability distribution of all output categories. The sampling method is defined formulaically as follows:

[0074]

[0075] Introduce the original Entropy strategy into object detection, and the formulaic definition is shown in Equation (6). The Entropy strategy first calculates the classification information entropy for each candidate object, then selects the entropy value corresponding to the candidate object with the largest entropy as the entropy of the entire image, and finally samples the image sample with the largest entropy.

[0076]

[0077] The following takes a detailed annotation of a land patch fine classification dataset as an example to implement the technical solution of the present invention in detail.

[0078] Example 1

[0079] Step 1: Prepare 100,000 unlabeled remote sensing images of land patches and determine the annotation information: surface cover type (vegetation / water / artificial structures), annotated cloud cover ratio, and land subdivision type (such as secondary land types like cultivated land / forest land / residential land / industrial land, etc.); among them, the surface cover type and the annotated cloud cover ratio are common sense annotation information, and the land subdivision type is professional knowledge annotation information;

[0080] Step 2: Randomly select 2,000 unlabeled images as a small-scale sample;

[0081] Step 3: Complete the annotation of 2,000 samples manually to obtain a small-scale annotated dataset;

[0082] Step 4: Train a target classification deep neural network model EfficientNet based on the small-scale annotated dataset to obtain a pre-trained model;

[0083] Step 5: Based on the pre-trained model, perform inference tests on the remaining unlabeled images, and output the object detection results for each image, including surface cover type, cloud cover ratio, land subdivision type, and confidence information;

[0084] Step 6: Based on the output confidence information, evaluate whether the images are further valuable. If the minimum uncertainty of all images is higher than the set threshold (the threshold is set according to experience), the annotation ends, otherwise go to Step 7;

[0085] Step 7: Sort the value of image annotation. Here, the sorting comprehensively considers the average uncertainty and class balance, and selects the top 200 most valuable images;

[0086] Step 8: Manually annotate the selected images;

[0087] Step 9: Based on the already annotated samples and the samples that are unlabeled but have generated pseudo-labels by model inference, use the supervised learning method with weighted labels to train the deep neural network model EfficientNet to obtain a new round of iterative model. In the Loss function used for training, the loss weight allocation strategy is: for manually labeled samples, the weight is 2; the error weight of pseudo-labeled samples is the confidence corresponding to the pseudo-labels.

[0088] Step 10: Jump to Step 5 to evaluate the new round of iterative model.

[0089] Please refer to Figure 3, which is the annotation result of Example 1. In the figure, the black color represents common sense annotation information, and the red color represents professional knowledge annotation information. In this example, the amount of manual annotation only needs to be 15% of the original data; the efficiency is increased by 3 times compared with full manual annotation.

[0090] In the description of the embodiments of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "center", "top", "bottom", "top part", "bottom part", "inner", "outer", "inner side", "outer side", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation to the present invention. Among them, the "inner side" refers to the internal or enclosed area or space. The "periphery" refers to the area around a specific component or a specific area.

[0091] In the description of the embodiments of the present invention, the terms "first", "second", "third", "fourth" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", "third", "fourth" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0092] In the description of the embodiments of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", "joined", "assembled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0093] In the description of the embodiments of the present invention, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0094] In the description of the embodiments of the present invention, it should be understood that "-" and "~" represent the range between two numerical values, and this range includes the endpoints. For example: "A - B" represents the range greater than or equal to A and less than or equal to B. "A ~ B" represents the range greater than or equal to A and less than or equal to B.

[0095] In the description of the embodiments of the present invention, the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the sole existence of A, the simultaneous existence of A and B, and the sole existence of B. Additionally, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0096] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An efficient hierarchical annotation method for remote sensing targets based on active learning, characterized in that, It includes the following steps: S1 Hierarchical formulation of annotation templates: Determine annotation information and classify it into professional knowledge annotation information and common sense annotation information; S2 Construction of a small-scale annotated dataset: Randomly extract a small-scale dataset from a large-scale unannotated dataset and perform hierarchical manual annotation; S3 Large-scale dataset annotation based on active learning: Use the small-scale annotated dataset to train a deep neural network model, then perform inference and prediction on the unannotated dataset, generate pseudo-labels and confidence scores, screen difficult samples according to the confidence scores for hierarchical manual annotation, update the training set and retrain the model, and iterate until the preset annotation quality threshold is met.

2. The method according to claim 1, characterized in that, In step S3, one of the following strategies is used to screen difficult samples: (1) Least confidence (LC) strategy: Select the samples with the lowest model prediction confidence; (2) Margin strategy: Select the samples with the smallest difference in the probabilities of the top two categories predicted by the model; (3) Entropy strategy: Select the samples with the largest information entropy of the classification probability distribution.

3. The method according to claim 1, wherein In step S1, the hierarchical annotation strategy includes: (1) Common sense annotation: Ordinary annotators complete the annotation of the common sense annotation information; (2) Expert annotation: Domain experts complete the annotation of the professional knowledge annotation information.

4. The method according to claim 1, wherein In step S3, a weighted loss function is used during the model training process, and the weight assignment strategy is: (1) The weight of expert-annotated samples is higher than that of pseudo-label samples; (2) The weight of pseudo-label samples is positively correlated with their confidence scores.

5. The method according to claim 1, wherein The method further includes a dynamic threshold adjustment step: Adjust the confidence screening threshold according to the model performance and the distribution complexity of the unannotated samples.

6. The method according to claim 1, characterized in that, The method further includes a pseudo-label correction step: Automatically verify high-confidence pseudo-label samples, and only add the samples that pass the verification to the training set.

7. The method according to claim 1, wherein During the iterative optimization process, a block inference and result fusion technology is adopted to improve the processing efficiency of large-size remote sensing images.

8. The method according to claim 1, wherein The result also includes an annotation result visualization and interactive correction step: Provide a visualization interface for experts to quickly correct pseudo-label errors and feedback the correction results to model training.