Small sample pest detection method based on multi-order feature aggregation
By constructing a multi-order feature aggregation module and a central loss function to optimize feature discrimination, the problem of insufficient generalization ability of pest detection in farmland environments is solved, and the accurate detection of small sample pests and the robustness of the model is achieved.
Patent Information
- Application Number
- CN202510357368.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing pest detection methods lack generalization capabilities in complex farmland environments, lack special data sets, and high sample labeling costs, making it difficult to achieve precise prevention and control.
A multi-order feature aggregation module is constructed, multi-layer features are extracted through ResNet101, combined with feature pyramid network and regional suggestion network, and a multi-order feature aggregation module and a central loss function are used to optimize feature discrimination, and small-sample object detection is performed.
It improves the robustness and adaptability of the model in farmland environment, realizes accurate detection of small sample pests, enhances the ability to distinguish features, and is suitable for detection of different scenarios and pest species.
Smart Images

Figure CN120298769A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image detection, relates to the detection of crop pest images in actual scenarios, and specifically relates to a small-sample pest detection method based on multi-order feature aggregation. Background Art
[0002] As a major grain-producing country, food security in China has always been an important cornerstone of national security. However, the threat of pests and diseases has led to a significant reduction in grain production. According to statistics, the annual grain loss caused by pests and diseases in China reaches tens of billions of catties. The current pest control system faces multiple challenges: First, chemical control still dominates, and this single control mode not only pollutes the ecological environment but also accelerates the emergence of pest resistance; Second, the pest monitoring and early warning system is not yet perfect, making it difficult to achieve precise prevention and control.
[0003] In the field of intelligent pest detection, traditional methods face difficulties in collecting samples of some pest species in the complex farmland environment, resulting in insufficient training samples and difficulty in meeting the actual application requirements. In recent years, small-sample detection technology has made breakthrough progress, and among them, the two-stage detection method based on meta-learning has shown outstanding performance. This method adopts a multi-scale positive sample refinement strategy, optimizes predictions at different scales by constructing a target pyramid structure, and combines the integrated architecture of FPN (Feature Pyramid Network) and Faster R-CNN to provide an effective solution for small-sample object detection. However, existing research mainly focuses on the field of general object detection and lacks special research on agricultural pest scenarios, resulting in poor model transfer effects. Therefore, constructing a dedicated small-sample pest dataset and improving the model's adaptability in actual scenarios have become an urgent need for the current development of agricultural intelligence.
[0004] Specifically, the existing technology has the following limitations: First, there is a lack of a dedicated dataset for pest characteristics, making it difficult for the model to accurately identify; Second, the farmland environment is complex and changeable, and the generalization ability of existing models is insufficient; Finally, the cost of sample annotation is high and the cycle is long. To address these problems, there is an urgent need to develop a small-sample detection method suitable for agricultural scenarios, construct a high-quality pest sample database, and improve the robustness and adaptability of the model in the actual farmland environment. Summary of the Invention
[0005] The purpose of the present invention is to provide a small-sample object detection method based on multi-order feature aggregation for the deficiencies and application requirements of the existing technology. By constructing different field pest datasets, using a multi-order feature aggregation module for optimization, and finally using center loss to enhance the discriminability of features by minimizing the distance between the same-class features and their class centers. It provides help for the plant protection prevention and control work to master the incidence of pests and diseases and the small-sample pest monitoring work.
[0006] The present invention provides a small-sample object detection method based on multi-order feature aggregation, which includes the following steps:
[0007] Step 1: Obtain small-sample pest data in the field and construct a comprehensive dataset containing base-class and new-class pests.
[0008] The comprehensive dataset includes image data of 11 common insect species, namely aphids, Adonia variegata, Lygus lucorum, wheat spider mites, tobacco bugs, rice planthoppers, stink bugs, black aphid bugs, Lygus pratensis, Lygus lineolaris, and whiteflies, as the base class. In addition, according to the application scenario, 10 to 30 images of unknown few-sample pests are collected as the new class.
[0009] Step 2: Multi-layer feature extraction and feature pyramid processing;
[0010] Extract multi-layer features from the input image through the ResNet101 convolutional neural network to generate feature maps of different scales, denoted as C2, C3, C4, and C5 respectively. Among them, the size of the feature map is halved layer by layer, and the number of feature channels increases layer by layer (from 256 to 2048).
[0011] Then, send C2 - C5 into the Feature Pyramid Network (FPN) together to integrate features of different levels into a multi-scale feature map with a unified number of channels (256), denoted as P2, P3, P4, and P5 respectively;
[0012] Step 3: Generate candidate boxes and feature alignment;
[0013] Through the Region Proposal Network (RPN), generate candidate box regions on all multi-scale feature maps (P2 - P5) respectively to locate the potential positions of target objects. Subsequently, the features of the candidate box regions pass through the ROI Align layer and are uniformly sampled into a feature map of a fixed size, denoted as F_roi.
[0014] Step 4: Feature optimization of the multi-order feature aggregation module;
[0015] Send F_roi into the MOGA (Multi-Order Gated Aggregation) multi-order feature aggregation module. The MOGA multi-order feature aggregation module dynamically fuses regional features of different scales (P2 - P5) through a hierarchical gating mechanism to generate an optimized aggregated feature, denoted as F_moga. The aggregated feature F_moga is further reduced in dimension through global average pooling to further compress redundant information and generate a feature representation with higher discriminability.
[0016] Step 5: Main detection branch and feature discrimination enhancement;
[0017] The obtained more discriminative feature representations are respectively fed into the classification sub-branch and the regression sub-branch of the main detection branch: The classification sub-branch predicts the target class probabilities through a fully connected layer, and at the same time introduces the Center Loss constraint to gather the feature vectors of the same-class targets towards their class centers, thereby enhancing the discriminative ability of the features. The regression sub-branch finely adjusts the position coordinates of the candidate bounding boxes through a fully connected layer.
[0018] Step Six: Model Training and Optimization;
[0019] In the base class dataset, multiple meta-tasks are generated using the N-way K-shot mechanism. Each meta-task consists of a support set and a query set: The support set is used to construct class prototypes, and the query set is used to evaluate the object detection performance. The network is trained through a hybrid loss function, including classification loss, bounding box regression loss, and Center Loss.
[0020] Step Seven: Model Output;
[0021] After multiple optimizations, the model finally outputs the detected bounding box coordinates, target class labels, and their confidence scores, completing the small sample object detection task.
[0022] The beneficial effects of the present invention are as follows:
[0023] 1. Aiming at the problem that there are many types of field pests and it is difficult to collect them, the present invention adopts a small sample pest detection method. Through the transfer learning mechanism between the base class and the new class, accurate detection of a small amount of pest data is achieved, and the detection performance of the model under limited sample conditions is improved.
[0024] 2. The present invention provides a new theoretical basis and method support for the actual field small sample pest detection model, which helps to expand the traditional pest detection methods. The multi-stage feature aggregation and dynamic optimization methods are introduced, which not only enhance the ability to extract target features under the complex background of field pests, but also further improve the discriminability of the features by introducing the center loss mechanism, providing new ideas and methods for the construction of the small sample pest detection model. By combining with deep learning methods, the performance and efficiency of the small sample pest detection model can be further improved.
[0025] 3. The present invention effectively utilizes multi-scale features under small sample conditions to construct and optimize the detection network, which has strong migration ability and can be widely applied to the detection tasks of different scenarios and pest types, providing reliable technical support for precision agriculture. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is the technical flow chart of the present invention;
[0027] Figure 2 is the schematic diagram of the overall working process of the present invention;
[0028] Figure 3 This is the visualization result of the experimental data of the new class of the present invention. Specific implementation manners
[0029] The present invention will be further explained below with reference to the accompanying drawings;
[0030] As Figure 1 and Figure 2 shown, the small-sample pest detection method based on multi-level feature aggregation includes the following steps:
[0031] Step (1) Dataset construction;
[0032] A comprehensive dataset containing 13 types of pests in different field scenarios is constructed. This dataset includes aphids, green lacewings, multispecies ladybugs, plant bugs, wheat spider mites, tobacco plant bugs, rice planthoppers, stink bugs, black aphid bugs, tarnished plant bugs, alfalfa plant bugs, whiteflies, and cotton bollworms. Among them, aphids, green lacewings, multispecies ladybugs, plant bugs, wheat spider mites, tobacco plant bugs, rice planthoppers, stink bugs, black aphid bugs, tarnished plant bugs, alfalfa plant bugs, and whiteflies are the base classes of the dataset, with a total of 4,577 images. Each image is labeled based on the EasyDL platform, generating 29,625 annotation files in COCO format, including information such as target bounding boxes and class labels. In addition, to solve the problem of few-sample pest detection, in this embodiment, cotton bollworms and green lacewings are taken as examples, and these two types of pests are divided into new classes to simulate the detection scenarios of difficult-to-collect pest samples in the field. Each new class contains 10 and 30 images respectively, and each image has a corresponding annotation file.
[0033] Step (2) Multi-level feature extraction and feature pyramid;
[0034] In the feature extraction stage of the model, the input image is input into the ResNet101 backbone network. After convolutional operations, feature maps at different levels are generated, denoted as C2, C3, C4, and C5 respectively. The resolutions of these feature maps are halved layer by layer, and the number of channels gradually increases from 256 to 2048, forming a multi-level feature representation from low to high levels. In addition, to adapt to the diversity of target sizes, through the Feature Pyramid Network (FPN), the features of C2 to C5 are integrated into multi-scale feature maps P2, P3, P4, and P5 with a unified number of channels (256) (input C2 - C5 into the FPN together, and then generate feature maps with a unified number of channels at multiple scales P2, P3, P4, and P5). Among them, the number of channels of each layer of features is 256 dimensions, ensuring the unity and compatibility of the feature maps in terms of scale. The model can not only extract the deep semantic information of the image but also achieve the fusion of multi-scale features through the feature pyramid, thus effectively coping with the detection tasks of targets of different sizes.
[0035] Step (3) Candidate box generation and feature alignment;
[0036] Based on multi-scale feature extraction, the feature maps P2 to P5 are input into the Region Proposal Network (RPN) to generate candidate box regions (Proposals). The RPN generates multiple anchor boxes with different scales and aspect ratios at each position of the feature map through anchor points, and combines the non-maximum suppression (NMS) algorithm to filter out the candidate boxes with an intersection over union (IoU) exceeding 0.5. Subsequently, the features within the candidate boxes are sampled through ROIAlign: First, according to the coordinates of the candidate boxes, they are accurately mapped to the corresponding feature map (usually a certain layer among P2 to P5 is selected according to the target size); then the candidate region is divided into grids, and bilinear interpolation is used to finely sample the features within each small grid, so as to retain more accurate spatial information; finally, the features of the uniformly sampled small grids are integrated into a feature block of a unified size, denoted as F_roi. This ensures that candidate boxes of different sizes can be represented by features of the same size, providing consistent input features for the subsequent object detection network. By generating candidate regions through the RPN and combining ROI Align for feature alignment, the model can effectively handle the diversity of object sizes.
[0037] Step (4) Feature optimization of the multi-order feature aggregation module;
[0038] The feature map F_roi after ROIAlign is input into the multi-order gated aggregation module MOGA (Multi-Order Gated Aggregation). The MOGA module fuses the regional features of F_roi at different scales through a multi-order feature interaction mechanism and a gated mechanism. Specifically, the MOGA module first extracts the context information at different scales through the multi-order feature interaction method, and then uses the gated unit to adaptively weight and fuse these features to generate an optimized aggregated feature representation. Subsequently, the MOGA module compresses redundant information through global average pooling (Global Average Pooling) to enhance the compactness and discriminability of the features, generating a more discriminative aggregated feature representation, denoted as F_moga. It can more effectively capture the discriminative features of small-sample objects, significantly improve the model's representation ability for few-shot data, and provide better robustness for subsequent classification and localization tasks.
[0039] Step (5) Main detection branch and feature discriminability enhancement;
[0040] The generated F_moga is input into the main detection branch, which is jointly composed of a classification sub-branch and a regression sub-branch. In these two sub-branches, probability prediction of the target category and adjustment of the bounding box are respectively performed. In the classification sub-branch, F_moga predicts the probability distribution of the target category through a fully connected layer. At the same time, through the aggregation mechanism of the feature vector and its category center (that is, constraining the feature vector according to the known category centers of each category to make the feature vectors of the same category approach their centers), the discriminative ability of the features is significantly enhanced, thereby improving the discrimination of categories. In the regression sub-branch, F_moga finely adjusts the coordinates of the candidate boxes through a fully connected layer to generate a more accurate target bounding box. The classification sub-branch and the regression sub-branch operate on F_moga, respectively outputting the category prediction and the bounding box regression result, and realizing end-to-end collaborative optimization through the classification loss and the bounding box regression loss during training. Through the collaborative effect of the classification and regression sub-branches, the model can not only efficiently complete the target detection task, but also further improve the detection performance of the model in the few-shot scenario by enhancing the discrimination ability of the classification features.
[0041] Step (6) Model training and optimization;
[0042] Base class training stage: In the base class training stage, the model takes the base class dataset (including 11 types of pest data such as aphids and rice planthoppers) as input and conducts target detection training through the classification sub-branch and the regression sub-branch of the main detection branch. After obtaining the feature representation F_moga for each input image, it is respectively input into the classification sub-branch and the regression sub-branch of the main detection branch to calculate the classification probability distribution and the regression coordinate prediction; then, according to the true label and the annotated box, the classification loss, the bounding box regression loss, and the CenterLoss are respectively calculated, and the weighted sum of the three is used as the total loss to update all network parameters in one backpropagation. This not only simplifies the training process but also enables all parts of the model to converge together under the same goal, giving full play to the advantages of multi-sub-branch cooperation.
[0043] Specifically, the classification loss is used to improve the recognition accuracy of the target category, and the specific formula is as follows:
[0044]
[0045] where represents the predicted probability of sample i in its true category y i .
[0046] The bounding box regression loss is used to optimize the prediction of the target position, and the specific formula is as follows:
[0047]
[0048] where is the true coordinate, is the predicted coordinate, which is usually defined as:
[0049]
[0050] CenterLoss further enhances the discriminability of features by clustering the feature vectors of the same-class objects towards their class centers. The specific formula is as follows:
[0051]
[0052] where N is the batch size, x i refers to the feature vector of the i-th sample, and refers to the class center vector of the i-th sample.
[0053] Through the training in this stage, the model can learn the general feature representation of the base classes and finally generate a pre-trained model that performs excellently on the base classes.
[0054] New class fine-tuning stage: In the new class fine-tuning stage, the model is further optimized using a small amount of labeled data of the new classes (cotton bollworms and green lacewings). During the fine-tuning process, the model retains the feature extraction part of the base class pre-trained model (i.e., ResNet and FPN) to ensure that the model can reuse the general feature representation ability learned by the base classes. At the same time, the last layer of the classification sub-branch is replaced, and the classification weights of the new classes are newly initialized to adapt to the detection tasks of the new classes. Through the classification and regression tasks of a small number of samples, the model gradually adjusts the network parameters and continues to be optimized using the classification loss, bounding box regression loss, and CenterLoss to improve the detection performance of the new classes. In addition, to further enhance the detection performance of the model in the few-shot and small-object scenarios, the model also includes an auxiliary branch that runs in parallel with the main detection branch. Specifically, the feature extraction network of this auxiliary branch is the same as that of the main detection branch to ensure consistent multi-scale information of the input image. After that, it has two sub-branches of classification and regression that run in parallel with the main detection branch; the structures of these two sub-branches are the same as those of the classification and regression sub-branches of the main detection branch. When the model generates candidate box regions from feature maps at different scales (P2 - P5), it further filters out positive sample candidate boxes with high confidence and inputs these candidate boxes into the auxiliary branch for secondary evaluation and coordinate correction. The auxiliary branch performs more refined feature extraction, classification, and regression on these candidate boxes, re-evaluates and corrects the coordinates, thereby correcting possible biases or missed detections in the main detection branch and further enhancing the detection ability of the model in the few-shot scenario.
[0055] Two-stage training strategy: Through the two-stage training strategy of base class training and new class fine-tuning, the model can not only learn the general object detection ability on the base class data but also quickly adapt to the detection tasks of new classes with a small number of samples.
[0056] Step (7) Model Output
[0057] After the model training is completed, the test set is used for inference to predict the target class and its corresponding bounding box. The detection results output by the model include the coordinates of the target bounding box, the class label, and the confidence score. Use AP, AP 50 , AP 75 , AP S , AP M , AP L and other multiple metrics to jointly evaluate the detection performance of the model on the base classes and new classes. The AP metric refers to judging whether the detection result matches the real target or false detection according to the set IoU threshold. Through the above evaluation, the actual application effect of the model in the few-shot object detection task is verified. The experimental results show that the model can not only maintain a high detection accuracy on the base class data, but also show better generalization ability on the new class data.
[0058] Step (8) The process of training and testing the model based on the constructed dataset
[0059] The experiments of the present invention are carried out under the Ubuntu22.04 operating system and the Pytorch2.1.0 framework, and accelerated by the GeForceGTX4090Ti GPU and CUDA12.2. The AdamW optimizer is used, the initial learning rate is set to 0.0001, and the decay factor is 0.0001. After continuously training for 4 epochs, if the accuracy of the validation set does not improve, the learning rate is multiplied by 0.1. Use the trained model to perform inference on the test set, predict the target class and its bounding box, and calculate the detection performance metrics to verify the effectiveness of the model.
[0060] In this example, by conducting comparative experiments on the effects of different data processing strategies and detection algorithms, the effectiveness of the proposed method in the few-shot object detection task is verified. In the few-shot scenarios of Helicoverpa armigera and Chrysopa pallens (new classes), through few-shot fine-tuning, the proposed method significantly improves the detection accuracy. Compared with the baseline model without MOGA and CenterLoss, the average precision (AP 50 ) is improved by 4.5% in the 10-shot case, and the average precision (AP 50 ) is improved by 3.4% in the 30-shot case. Under the complex field background and diverse lighting conditions, the anti-interference ability of the model is significantly enhanced by the proposed method, solving problems such as sparse target scale distribution and feature imbalance, providing an efficient and reliable solution for field pest monitoring. It provides help for the plant protection control work based on mastering the incidence of plant diseases and insect pests and the few-shot pest monitoring work.
[0061] Table 1 Comparison of detection accuracies
[0062]
[0063] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several alternatives or modifications can be made to these described embodiments, and these alternative or modified forms should all be regarded as falling within the protection scope of the present invention.
[0064] The parts not detailed in the present invention belong to the well-known technologies in the art.
Claims
1. A small-sample object detection method based on multi-order feature aggregation, characterized in that It includes the following steps: Step 1: Obtain the pest data of small field samples and construct a comprehensive data set containing the pest of the base class and the new class; Step 2: Multi-layer feature extraction and feature pyramid processing; Extract multi-layer features of the input image through the ResNet101 convolutional neural network to generate feature maps of different scales, denoted as C2, C3, C4, and C5 respectively, where the size of the feature map is halved layer by layer, and the number of feature channels increases layer by layer; Next, send C2-C5 into the Feature Pyramid Network (FPN) together to integrate the features of different levels into multi-scale feature maps with the same number of channels, denoted as P2, P3, P4, and P5 respectively; Step 3: Generate candidate boxes and feature alignment; Generate candidate box regions on all multi-scale feature maps through the Region Proposal Network (RPN) to locate the potential positions of the target objects; subsequently, the features of the candidate box regions are uniformly sampled into a feature map with a fixed size through the ROI Align layer, denoted as F_roi; Step 4: Feature optimization of the multi-order feature aggregation module; Send F_roi into the MOGA multi-order feature aggregation module to generate a more discriminative feature representation F_moga; Step 5: Main detection branch and feature discrimination enhancement; Send the obtained more discriminative feature representation F_moga into the classification sub-branch and the regression sub-branch of the main detection branch respectively: the classification sub-branch predicts the target category probability through the fully connected layer, and at the same time introduces the Center Loss constraint to gather the feature vectors of the same-class objects towards their class centers, thereby enhancing the discriminative ability of the features; The regression sub-branch finely adjusts the position coordinates of the candidate boxes through the fully connected layer; Step 6: Model training and optimization; In the base class data set, use the N-way K-shot mechanism to generate multiple meta-tasks, and each meta-task consists of a support set and a query set: the support set is used to construct the class prototype, and the query set is used to evaluate the target detection performance; the network is trained through a hybrid loss function, including classification loss, bounding box regression loss, and Center Loss; Step 7: Model output; After multiple optimizations, the model finally outputs the coordinates of the detected bounding boxes, the target class labels, and their confidence scores, completing the small sample target detection task.
2. The small sample object detection method based on multi-order feature aggregation according to claim 1, wherein The comprehensive data set includes the image data of 11 common insect species, such as aphids, Adonia variegata, mirids, wheat spiders, tobacco mirids, rice planthoppers, stink bugs, black aphid bugs, Lygus pratensis, Lygus lucorum, and whiteflies, as the base class, and each image is annotated to generate an annotation file; in addition, according to the application scenario, 10 to 30 unknown few-shot pest images are collected as the new class, and each image has a corresponding annotation file.
3. A few-shot object detection method based on multi-order feature aggregation according to claim 1, characterized in that The specific method of step (3) is as follows: Based on multi-scale feature extraction, the feature maps P2 to P5 are input into the Region Proposal Network (RPN) to generate candidate box regions, Proposal. The RPN generates multiple anchor boxes with different scales and aspect ratios at each position of the feature map through anchor points, and combines the Non-Maximum Suppression (NMS) algorithm to filter out candidate boxes with an Intersection over Union (IoU) exceeding 0.
5. Subsequently, the features within the candidate boxes are sampled through ROI Align: First, according to the coordinates of the candidate box, it is accurately mapped to the corresponding feature map. Then, the candidate region is divided into grids, and bilinear interpolation is used to finely sample the features within each small grid. Finally, the features of the uniformly sampled small grids are integrated into a feature block of a unified size, denoted as F_roi.
4. A few-shot object detection method based on multi-order feature aggregation according to claim 3, characterized in that, The specific method of step (4) is as follows: The feature map F_roi after ROI Align is input into the Multi-Order Gated Aggregation Module (MOGA). The MOGA module fuses the regional features of F_roi at different scales through a multi-order feature interaction mechanism and a gated mechanism. Specifically, the MOGA module first extracts context information at different scales through the multi-order feature interaction method, and then adaptively weights and fuses these features using a gated unit to generate an optimized aggregated feature representation. Subsequently, the MOGA module compresses redundant information through global average pooling to enhance the compactness and discriminability of the features, generating a more discriminative aggregated feature representation, denoted as F_moga.
5. A small-sample object detection method based on multi-order feature aggregation according to claim 4, characterized in that, The specific method of step (5) is as follows: The generated F_moga is input into the main detection branch, which is jointly composed of a classification sub-branch and a regression sub-branch. In these two sub-branches, probability prediction of the target category and adjustment of the bounding box are respectively performed. In the classification sub-branch, F_moga predicts the probability distribution of the target category through a fully connected layer, and at the same time enhances the discriminative ability of the features through the clustering mechanism between the feature vector and its category center, thereby improving the discrimination of the category. In the regression sub-branch, F_moga finely adjusts the coordinates of the candidate box through a fully connected layer to generate a more accurate target bounding box. The classification sub-branch and the regression sub-branch operate on F_moga and output the category prediction and the bounding box regression results respectively.
6. The small sample object detection method based on multi - order feature aggregation according to claim 5, characterized in that, The specific method of step (6) is as follows: Base class training stage: In the base class training stage, the model takes the base class dataset as input and conducts object detection training through the classification sub-branch and regression sub-branch of the main detection branch. After obtaining the feature representation F_moga for each input image, it is respectively input into the classification sub-branch and regression sub-branch of the main detection branch to calculate the classification probability distribution and regression coordinate prediction. Then, according to the ground truth labels and annotation boxes, the classification loss, bounding box regression loss, and Center Loss are calculated respectively, and the weighted sum of the three is used as the total loss to update all network parameters in one backpropagation. Specifically, the classification loss is used to improve the recognition accuracy of target categories, the bounding box regression loss is used to optimize the prediction of target positions, and the Center Loss further enhances the discriminability of features by clustering the feature vectors of the same class towards their class centers. Through the training of this stage, the model can learn the general feature representation of the base class and finally generate a pre-trained model with excellent performance on the base class. New class fine-tuning stage: In the new class fine-tuning stage, the model uses a small amount of annotated data of the new class for further optimization. During fine-tuning, the model retains the feature extraction part of the base class pre-trained model, namely ResNet and FPN. At the same time, the last layer of the classification sub-branch is replaced, and the classification weights of the new class are initialized anew to adapt to the detection task of the new class. Through the classification and regression tasks of a small number of samples, the model gradually adjusts the network parameters and continues to optimize using the classification loss, bounding box regression loss, and Center Loss to improve the detection performance of the new class. In addition, to further enhance the detection performance of the model in few-shot and small object scenarios, the model also includes an auxiliary branch parallel to the main detection branch. Specifically, the feature extraction network of this auxiliary branch is the same as that of the main detection branch to ensure consistent multi-scale information of the input image. After that, it has two sub-branches, classification and regression, parallel to the main detection branch. These two sub-branches have the same structure as the classification and regression sub-branches of the main detection branch. When the model generates candidate box regions from feature maps of different scales, it will further screen out positive sample candidate boxes with high confidence and input these candidate boxes into the auxiliary branch for secondary evaluation and coordinate correction. The auxiliary branch conducts more refined feature extraction, classification, and regression for these candidate boxes, re-evaluates and corrects the coordinates, thereby correcting the biases or missed detections in the main detection branch and further enhancing the detection ability of the model in few-shot scenarios.
Citation Information
Cited By
Marine remote sensing generalized small sample target detection method and system for controlled knowledge migration
CN122090046A
Method and system for generalized small sample target detection of ocean remote sensing with controlled knowledge transfer
CN122090046B