An Efficient Annotation Method for Training Datasets Based on Multi-Class Insect Images

By preprocessing multiple insect image data sets and building a general insect detector, combined with self-supervised learning and hierarchical clustering algorithms, the problems of high difficulty in annotating insect image data sets and high artificial resource consumption are solved, and efficient data set annotation and model performance improvement are achieved.

CN119152318BActive Publication Date: 2025-05-27ZHEJIANG SCI-TECH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411632057.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-05-27
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

In the field of agricultural pest monitoring, it is difficult for the prior art to efficiently label insect image data sets, especially in the images captured by the equipment, there are problems of morphological differences, background complexity and similar appearance pests, resulting in high labeling difficulty and high artificial resource consumption.

Method used

By preprocessing the multi-class insect image dataset, including data filtering, pseudo-label generation and automatic background removal, a general insect detector GInsectD is built, and good feature representation is learned using the self-supervised learning method DEIP, and the target map is classified through the hierarchical clustering algorithm to finally form an efficient target detection dataset.

Benefits of technology

It realizes efficient annotation of multiple insect images, reduces the demand for artificial resources, improves the labeling efficiency of the data set, and improves the accuracy and robustness of the model in feature extraction and classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152318B_ABST
    Figure CN119152318B_ABST
Patent Text Reader

Abstract

The present invention discloses an efficient annotation method based on a multi-class insect image training dataset, which relates to the technical field of image processing. The key points of its technical solution are as follows: It includes the following steps: S0: Preprocess the multi-class insect image dataset; S1: Construct a general insect detector GInsectD; S2: Crop and generate insect target small images with unique indexes; S3: Use the self-supervised learning method DEIP to learn good feature representations; S4: Use the hierarchical clustering algorithm to classify the target small images; S5: Trace back the classified target small images to the original images to form the final object detection dataset. By modifying and fine-tuning the open-vocabulary object detector, the present invention enables it to well detect various multi-class insect images, showing strong generalization performance. It can be used as a general insect detector to locate all insects on multi-class insect images. By means of the general insect detector and self-supervised representation learning, the demand for human resources is greatly reduced, and the labeling efficiency of the dataset is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more specifically, to a method for labeling an efficient training data set based on multiple types of insect images. Background Art

[0002] With the rapid development of deep learning technology, more and more insect counting methods have been transformed from inefficient manual statistics to efficient automatic identification and counting. However, an issue that is easily overlooked is that a robust and excellent deep learning-based algorithm often requires a large amount of labeled data sets to support it. Although deep learning models have shown excellent performance in tasks such as image classification and target detection in public data sets, their success depends largely on large-scale, high-quality, and diverse labeled data.

[0003] In the field of agricultural pest monitoring, with the addition of automatic insect-catching equipment, it has become feasible to collect a large number of original insect image samples, but the problem that comes with it is the difficulty in labeling data sets: compared with the pest images captured in the natural environment, the pests captured by the equipment become darker in color after death, and some pests may be missing limbs, resulting in large morphological differences; among them, the equipment that uses insect phototaxis to attract insects can capture a large number of insects, large and small insect targets are mixed together, high-density insects may block each other, the complexity of the background increases, and the mutual interference between pests with similar appearances increases the difficulty of labeling. High-quality data labeling requires not only professional biological knowledge, but also in-depth understanding of specific types of pests, otherwise it is easy to mislabel or miss labels. At the same time, because the distribution and living habits of agricultural pests vary with location and season, their appearance characteristics in different growth cycles will also vary. In order to make the model adapt to pest identification tasks in various scenarios, it is necessary to ensure that the training data set contains samples over a wide range of time spans and regions, all of which require inestimable and expensive human resources to complete.

[0004] Therefore, new solutions need to be proposed to solve this problem. Summary of the invention

[0005] In view of the shortcomings of the prior art, the object of the present invention is to provide an efficient annotation method based on a multi-class insect image training data set.

[0006] The above technical objectives of the present invention are achieved through the following technical solutions: A method for labeling a multi-class insect image-based efficient training data set, comprising the following steps:

[0007] S0: Preprocessing of multiple insect image datasets, including data filtering, pseudo-label generation and automatic background removal;

[0008] S1: Construct a general insect detector GInsectD to locate all insect targets;

[0009] S2: Crop and generate small insect target images with unique indexes;

[0010] S3: Use the self-supervised learning method DEIP to learn good feature representations;

[0011] S4: Use the hierarchical clustering algorithm to classify the target small images;

[0012] S5: Trace back the classified target small images to the original image to form the final object detection dataset.

[0013] The present invention is further configured that in the step S0, the preprocessing of the multi-class insect image dataset includes the following sub-steps:

[0014] S01: Filter low-quality images from the data.

[0015] S02: Pseudo-label generation and background automatic removal include using a multi-modal open-set object detector to label all insects, and removing image files without target bounding boxes, i.e., background image files, by appropriately reducing the confidence.

[0016] The present invention is further configured that in the step S1, constructing the general insect detector GInsectD includes the following sub-steps:

[0017] S11: Use a multi-modal open-set object detection model architecture with a dual encoder-single decoder;

[0018] S12: Adopt Focal Loss as the classification loss function and L1Loss as the bounding box regression loss function;

[0019] S13: Use Swin-T for the image backbone and BERT for the text backbone;

[0020] S14: The detected class name is 'insect';

[0021] S15: The learning rate scheduling strategy during model training adopts CosineAnnealingLR.

[0022] The present invention is further configured that in the step S2, cropping and generating small insect target images with unique indexes includes using python tools to crop small images according to the detection box coordinate information and saving them as jpg format files.

[0023] The present invention is further configured that in the step S3, using the self-supervised learning method DEIP to learn good feature representations includes the following sub-steps:

[0024] S31: The input image undergoes different data augmentations to obtain a global cropped view and a local cropped view;

[0025] S32: Calculate the cross-entropy loss by comparing the features of different views;

[0026] S33: Only update the parameters of the apprentice network during backpropagation to minimize the loss function.

[0027] The present invention is further configured as follows: In the step S4, classifying the target small images using the hierarchical clustering algorithm includes using a hierarchical clustering method similar to the hierarchical classification structure unique to biology, and setting the number of clusters according to the maximum number of insect species in the image.

[0028] The present invention is further configured as follows: In the step S5, tracing the classified target small images back to the original image to form the final object detection dataset includes the following sub-steps:

[0029] S51: Establish a category database, using the multi-class insect image file name and the target box coordinate information as the small image file name and the unique index of the database;

[0030] S52: Associate and upload the pseudo-classification results obtained by the clustering or classification model to the structured database;

[0031] S53: Import all the cropped images into the dataset annotation software and load the pre-annotation information from the database;

[0032] S54: After the annotating expert corrects and confirms the dataset, update the annotation information to the database, and finally the script automatically traces the categories back to the object detection dataset of the insect-catching device according to the information in the index.

[0033] The present invention is further configured as follows: In the step S1, constructing the general insect detector GInsectD further includes using the method of small-sample iterative training, that is, using the model weights of the previous iteration in subsequent iterations and increasing the number of images to further improve the detection effect.

[0034] The present invention is further configured as follows: In the step S3, learning good feature representations using the self-supervised learning method DEIP further includes using the multi-crop technique, and calculating the cross-entropy loss by comparing the features of different views, where the calculation of the cross-entropy loss involves comparing the similarities of the classification token and the patch tokens.

[0035] The present invention is further configured as follows: In the step S0, data filtering and merging further includes using an image quality assessment algorithm to automatically identify and exclude low-quality images.

[0036] In summary, the present invention has the following beneficial effects: By modifying some training strategies and performing small-sample iterative fine-tuning on the open-vocabulary object detector, the present invention enables it to detect various multi-class insect images well, demonstrating strong generalization performance. It can be used as a general insect detector to locate all insects in multi-class insect images. Through self-supervised representation learning of a large number of small insect target images, the proposed DEIP self-supervised learning method can learn general feature representations of insects, resulting in a significant improvement in accuracy and robustness in feature extraction tasks or downstream classification tasks. At the same time, by leveraging the general insect detector and self-supervised representation learning, an efficient training dataset annotation method based on multi-class insect images is constructed, greatly reducing the demand for human resources and significantly improving the labeling efficiency of the dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a schematic flowchart of the present invention;

[0038] Figure 2 is a schematic diagram of the Focal Loss formula in the present invention;

[0039] Figure 3 is a schematic diagram of the L1 Loss formula in the present invention;

[0040] Figure 4 is a schematic diagram of the cross-entropy loss calculation formula in the present invention;

[0041] Figure 5 is a framework diagram of GInsectD in the present invention;

[0042] Figure 6 is a framework diagram of DEIP in the present invention;

[0043] Figure 7 is a detection result diagram of GInsectD in the present invention;

[0044] Figure 8 is a marked result diagram of the method of the present invention, with different species marked in different colors; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0046] To better describe and illustrate the embodiments of the present application, one or more drawings may be referred to. However, the additional details or examples used to describe the drawings should not be considered as limiting the scope of any of the inventions, the currently described embodiments, or the preferred modes of the present application.

[0047] In the description of the present invention, it should be noted that the orientation or positional relationships indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. are based on the positional relationships shown in the drawings. These are only for the convenience of describing the present invention and do not indicate that the device referred to must have a specific orientation or operate in a specific orientation. Therefore, it should not be construed as a limitation of the present invention.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the description of this application in this specification are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0049] The present invention will be described in detail below with reference to the drawings and embodiments.

[0050] A method for efficiently annotating a training dataset based on multi-class insect images, as Figure 1 shown, includes the following steps:

[0051] S0: Preprocess the multi-class insect image dataset, including data filtering and merging, as well as pseudo-label generation and background automatic removal;

[0052] S1: Build a general insect detector GInsectD to locate all insect targets;

[0053] S2: Crop and generate small insect target images with unique indexes;

[0054] S3: Use the self-supervised learning method DEIP to learn good feature representations;

[0055] S4: Use the hierarchical clustering algorithm to classify the target small images;

[0056] S5: Trace back the classified target small images to the original images to form the final object detection dataset.

[0057] Specifically, in step S0, preprocessing the multi-class insect image dataset includes the following sub-steps:

[0058] S01: Filter low-quality images.

[0059] S02: Pseudo-label generation and background automatic removal include using a multi-modal open-set object detector to label all insects, and removing image files without target bounding boxes, i.e., background image files, by appropriately reducing the confidence level. By the above settings, low-quality images caused by equipment problems are removed, ensuring the quality of the images in the subsequent processed dataset, providing high-quality training data for building an accurate model, and removing unqualified images, which can reduce the processing of these images in subsequent steps, thus saving computing resources and time and improving the overall processing efficiency.

[0060] In step S1, constructing a general insect detector GInsectD includes the following sub-steps:

[0061] S11: Using a multi-modal open-set object detection model architecture with a dual-encoder and single-decoder. By using this architecture, it can simultaneously process image and text information, improve the ability to understand complex scenes, enhance zero-shot detection ability, and improve the adaptability of the model in diverse and dynamically changing environments.

[0062] S12: Adopting Focal Loss as the classification loss function and L1 Loss as the bounding box regression loss function. Specifically, the Focal Loss formula is as Figure 2 shown, where is the predicted probability of the model for the class to which the sample belongs; is the adjustment factor for balancing the positive and negative sample ratios; is the adjustment parameter, called the focal parameter, and the L1 Loss formula is as Figure 3 shown, where is the number of samples. is the true value of the th sample. is the predicted value of the

[0063] th sample;

[0064] S13: Using Swin-T for the image backbone and BERT for the text backbone;

[0065] S14: Using 'insect' as the detection class name;

[0066] Further, in step S1, constructing the general insect detector GInsectD also includes using the method of iterative training with a small sample, that is, using the model weights of the previous iteration in subsequent iterations and increasing the number of images incrementally to further improve the detection effect. Through the above settings, it helps the rapid iteration and performance improvement of new tasks. By freezing the BERT backbone network and fine-tuning it multiple times using the method of iteratively generating self-training labels, the sample size of the dataset increases sequentially. The purpose is to leverage the powerful generalization ability of the pre-trained model and continuously modify the pseudo-labels to improve the model performance.

[0067] In step S2, cropping to generate small insect target images with unique indexes includes using python tools to crop small images according to the detection box coordinate information and saving them as jpg format files. The naming specification of the small insect target image files is that the unique index is obtained by concatenating the file name of the multi-class insect image where it is located and the position of the target box where the target is located. By assigning a unique index based on its position in the original image to each cropped small insect target image, it is easy to track and associate each small image with its source image, thus simplifying the organization and management of data. Moreover, the unique index makes it easier to associate the cropped image with its metadata, which is very useful for data analysis and subsequent machine learning model training.

[0068] In step S3, learning good feature representations using the self-supervised learning method DEIP includes the following sub-steps:

[0069] S31: The input image undergoes different data augmentations to obtain a global cropped view and a local cropped view;

[0070] S32: Calculate the cross-entropy loss by comparing the features of different views;

[0071] S33: Only update the parameters of the apprentice network during backpropagation to minimize the loss function.

[0072] In step S3, learning good feature representations using the self-supervised learning method DEIP also includes using the multicrop technique and calculating the cross-entropy loss by comparing the features of different views. The calculation of the cross-entropy loss involves comparing the similarities of the classification token and patch tokens. The calculation formula of the cross-entropy loss is as Figure 4As shown, the classification token obtained through the apprentice network is projected and converted into a k-dimensional vector, and the softmax is applied to obtain the probability value Pa. Similarly, after the master network features are projected, the softmax is applied, and then the moving average is used for the centering operation to obtain Pm. By comparing the high-level semantic information of the images obtained from different perspectives and enhancement means, the DEIP method can learn more abundant and robust feature representations. These feature representations can capture the essential features of the images without being affected by specific perspectives or enhancement means. At the same time, since the DEIP method learns the consistency of the global features through self-supervised learning, this enables the model to maintain consistent feature representations under different data augmentations and transformations, thereby improving the generalization ability of the model. By forcing the model to learn small-scale dependencies and high-frequency detail information through the DEIP method, this helps the model capture more delicate image features and improve the recognition ability of image details.

[0073] In step S4, using the hierarchical clustering algorithm to classify the target small images includes using a hierarchical clustering method similar to the unique hierarchical classification structure in biology for classification, and the number of clusters is set according to the maximum number of insect species in the image. Hierarchical clustering can display the hierarchy of the data, which helps to understand the internal hierarchical relationship of the data. Through the dendrogram, the internal structure of the data and the clustering situation at each level can be deeply understood, increasing the interpretability of the clustering results. And when the clustering is completed, a cut can be made at any level to obtain a specified number of clusters. This flexibility allows users to select the most appropriate clustering level according to their needs.

[0074] In step S5, backtracking the classified target small images to the original image to form the final object detection dataset includes the following sub-steps:

[0075] S51: Establish a category database, using the multi-class insect image file name and the target box coordinate information as the small image file name and the unique index of the database;

[0076] S52: Associate and upload the pseudo-classification results obtained by the clustering or classification model to the structured database;

[0077] S53: Import all the cropped images into the dataset annotation software and load the pre-annotation information from the database;

[0078] S54: After the annotator corrects and confirms the dataset, update the annotation information to the database, and finally the script automatically backtracks the categories to the object detection dataset of the insect-catching device according to the information in the index.

[0079] By modifying some training strategies and performing small-sample iterative fine-tuning on the open-vocabulary object detector, the present invention enables it to detect various multi-class insect images well, demonstrating strong generalization performance. It can be used as a general insect detector to locate all insects in multi-class insect images. Through self-supervised representation learning of a large number of small insect target images, the proposed DEIP self-supervised learning method can learn general feature representations of insects, resulting in a significant improvement in accuracy and robustness in feature extraction tasks or downstream classification tasks. At the same time, with the help of a general insect detector and self-supervised representation learning, an efficient training dataset annotation method based on multi-class insect images is constructed, greatly reducing the demand for human resources and significantly improving the labeling efficiency of the dataset.

[0080] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.

Claims

1. An efficient annotation method based on a multi-class insect image training dataset, characterized by: The following steps are involved: S0: Preprocessing of multi-category insect image datasets, including data filtering and merging, as well as pseudo-label generation and automatic background removal; S1: Build a general insect detector GInsectD to locate all insect targets; S2: cropping to generate a small insect target image with a unique index; S3: Use the self-supervised learning method DEIP to learn good feature representation; S4: Classify the target image using the hierarchical clustering algorithm; S5: trace the classified target images back to the original images to form the final target detection data set; In the step S3, using the self-supervised learning method DEIP to learn a good feature representation includes the following sub-steps: S31: the input image is subjected to different data enhancements to obtain a global cropped view and a local cropped view; S32: Calculate the cross entropy loss by comparing the features of different views; S33: In back propagation, only the parameters of the apprentice network are updated to minimize the loss function; In step S3, learning a good feature representation using the self-supervised learning method DEIP also includes using the multicrop technique and calculating the cross entropy loss by comparing the features of different views, wherein the calculation of the cross entropy loss involves comparing the similarity of the classification token and the patch tokens. The calculation formula of the cross entropy loss is shown as follows: ; Among them, the classification token obtained through the apprentice network is converted into a k-dimensional vector after projection, and softmax is applied to obtain the probability value Pa. Similarly, the master network features are projected and softmax is applied, and then the moving average is used for centering to obtain Pm.

2. The efficient annotation method based on a multi-class insect image training data set according to claim 1, characterized in that: In the step S0, preprocessing the multi-category insect image dataset includes the following sub-steps: S01: filtering low-quality images; S02: Pseudo-label generation and automatic background removal include using a multimodal open set target detector to mark all insects and removing image files that do not contain target boxes, i.e. background image files, by appropriately lowering the confidence.

3. The efficient annotation method based on a multi-class insect image training data set according to claim 1, characterized in that: In the step S1, constructing a general insect detector GInsectD includes the following sub-steps: S11: using a multimodal open set object detection model architecture with dual encoders and single decoders; S12: Focal Loss is used as the classification loss function and L1Loss is used as the bounding box regression loss function; S13: Swin-T is used as the image backbone and BERT is used as the text backbone; S14: The detection category name uses 'insect'; S15: The learning rate scheduling strategy during model training adopts CosineAnnealingLR.

4. The efficient annotation method based on a multi-class insect image training data set according to claim 1, characterized in that: In the step S2, cropping and generating a small image of the insect target with a unique index includes using a python tool to crop the small image according to the detection frame coordinate information and saving it as an image format file.

5. The efficient annotation method based on a multi-class insect image training data set according to claim 1, characterized in that: In step S4, using a hierarchical clustering algorithm to classify the target thumbnails includes using a hierarchical clustering method similar to a hierarchical classification structure specific to biology for classification, and the number of clusters is set according to the maximum number of insect species in the image.

6. The efficient annotation method based on a multi-class insect image training data set according to claim 1, characterized in that: In the step S5, tracing back the classified target thumbnails to the original images to form the final target detection data set includes the following sub-steps: S51: establishing an insect category database, using the file names of multiple insect images and target frame coordinate information as thumbnail file names and database unique indexes; S52: associating the pseudo classification results obtained by the clustering or classification model and uploading them to a structured database; S53: import all cropped images into the dataset annotation software and load pre-annotated information from the database; S54: After the labeling expert corrects and confirms the data set, the labeling information is updated to the database. Finally, the script automatically traces the category back to the target detection data set based on the information in the index.

7. The efficient annotation method based on a multi-class insect image training data set according to claim 3 is characterized by: In the step S1, constructing the universal insect detector GInsectD also includes using a small sample iterative training method, that is, using the model weights of the previous iteration in subsequent iterations and increasing the number of images to further improve the detection effect.

8. The efficient annotation method based on a multi-class insect image training data set according to claim 2, characterized in that: In step S0, data filtering and merging also includes using an image quality assessment algorithm to automatically identify and exclude low-quality images.

Citation Information

Patent Citations

  • Characterization learning method of digital pathological image

    CN113516181A

  • Pest identification method based on multi-mode self-supervision Transform architecture

    CN116702035A

  • Lightweight pest detection method based on improved YOLOv7 and RKNPU2

    CN117058552A