A Continuous Learning Instance Segmentation Method Based on Uncertainty Score Pseudo-Labels

The uncertainty score pseudo-label method in continuous learning instance segmentation addresses catastrophic forgetting and background shift by using a teacher-student framework to integrate old and new labels, enhancing model performance on new tasks while retaining old knowledge.

CN116051918BActive Publication Date: 2025-07-15CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211360327.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2025-07-15
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

Existing instance segmentation methods are difficult to effectively update models in continuous learning, especially facing catastrophic forgetting and background offset problems, and cannot retain old class knowledge without old class datasets.

Method used

A continuous learning instance segmentation method based on uncertainty fraction pseudo-label is adopted, and feature and output distillation is performed to alleviate catastrophic forgetting and background shift through a teacher-student architecture FODS model, combining uncertainty fraction pseudo-label and cross-entropy loss function.

Benefits of technology

Effectively retaining old class knowledge, alleviating catastrophic forgetting and background shifts, improving the robustness and real-time nature of the model, and being able to maintain old class knowledge when new class data is updated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051918B_ABST
    Figure CN116051918B_ABST
Patent Text Reader

Abstract

The present invention discloses a continuous learning instance segmentation detection method based on uncertainty score pseudo-labels, which includes three parts. First is the uncertainty score pseudo-label generation method. By adding uncertainty scores to each pseudo-label, the credibility of the pseudo-labels is dynamically adjusted according to the uncertainty scores and training values during network training, improving the accuracy of the student model's learning. Second is the last-layer feature distillation method, which distills the teacher features (the last-layer features extracted by the teacher model from the input image) and the student features (the last-layer features extracted by the student model from the input image). Third is the output distillation method, which distills the output results (classification, bounding box) of the teacher model into the student model. The present invention can effectively alleviate the catastrophic forgetting and background shift problems encountered in continuous learning for instance segmentation, effectively retain the knowledge of old-class images, and complete the segmentation task for new-class images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and deep learning, and particularly relates to a continuous learning instance segmentation method based on uncertainty score pseudo-labels Background Art

[0002] Instance segmentation is a key problem in computer vision. It aims to identify all instances in an image and classify them at the pixel level, and has wide applications in scenarios such as robotics, autonomous driving, video surveillance, and others. With the availability of a large number of manually annotated datasets and the richness of deep convolutional networks, instance segmentation has developed. Many researchers not only focus on the accuracy of instance segmentation, but also consider the robustness and real-time performance of instance segmentation

[0003] Existing instance segmentation methods all follow the convention that all classes are known in advance and only learned once. This situation does not meet the actual requirements. A real-world system should be able to update the model only with new class data while retaining the knowledge of old classes. However, it is usually difficult to achieve in instance segmentation based on deep learning. This phenomenon is often observed in real life, but updating the model without an old class dataset is a challenge, mainly facing two aspects of problems

[0004] 1. Catastrophic forgetting problem

[0005] Catastrophic forgetting is an inherent problem of forgetting old class knowledge in continuous learning and an inevitable problem in backpropagation. This is because the learning of new classes changes the weights of the network, causing the network to forget what was learned before

[0006] 2. Background shift problem

[0007] There is a special class in instance segmentation, the background class, that is, an instance does not belong to any class. In continuous learning instance segmentation, the background class consists of two parts. The first is the true background class instance. The second is the instance that does not belong to the new class, including the old class instances learned before, also called pseudo-background, and the future class instances that have not been learned yet. That is to say, the instances contained in the pseudo-background are not constants. On the contrary, it changes with the learning task

[0008] In summary, in response to these two challenges, a continuous learning instance segmentation method based on uncertainty score pseudo-labels is proposed Summary of the Invention

[0009] In view of the above problems, the purpose of the present invention is to provide a continuous learning instance segmentation method based on uncertainty score pseudo-labels

[0010] A continuous learning instance segmentation method based on uncertainty score pseudo-labels includes the following steps

[0011] Step 1: Obtain the dataset of task t, including new-class images and new-class instance segmentation labels;

[0012] Step 2: Use the model trained in task t - 1 as the teacher model M for task t t-1 , and input the new-class images into the teacher model M t-1 to obtain the uncertainty score pseudo-labels of old-class instances in the new-class images of task t, and merge the uncertainty score pseudo-labels and the new-class instance segmentation labels to obtain the new dataset of task t;

[0013] Step 3: Construct the student model M for task t t , and based on the teacher model M for task t t-1 , build a Feature-Output Distillation and Score Pseudo Labels (FODS) model based on the teacher-student architecture, where t represents task t;

[0014] Step 4: Input the new dataset at task t into the FODS model, and train the feature extractors of the teacher model and the student model respectively for feature extraction to obtain teacher features and student features Use the L2 loss for the last-layer feature distillation when training the student model;

[0015] Step 5: Obtain the instance results of the new dataset by the teacher model M t-1 and perform output distillation on them using the cross-entropy function;

[0016] Step 6: Combine the loss functions of the last-layer feature distillation in Step 4 and the output distillation in Step 5 to train the FODS model and obtain the student model M t .

[0017] Compared with the prior art, the present invention has the following beneficial effects:

[0018] 1. Reasonably utilize all information. Add the scores corresponding to the pseudo-labels to the RPN network, which endows the network with degrees of freedom and allows the network to independently filter labels according to the training, saving the screening and calculation of the pseudo-label threshold during the network training process.

[0019] 2. Alleviate the problems of catastrophic forgetting and background shift, and be able to retain old-class knowledge in multiple task steps. Description of the Drawings

[0020] Figure 1 is a new-class instance segmentation label map.

[0021] ​Figure 2 It is a schematic diagram of the uncertainty score pseudo-label method.

[0022] Figure 3 It is a diagram of the instance results of the teacher model's prediction of new class images.

[0023] Figure 4 It is a pseudo-label map of the uncertainty scores of new class images.

[0024] Figure 5 It is a task step diagram.

[0025] Figure 6 It is a framework diagram of (Feature-Output Distillation and Score PseudoLabels (FODS)) based on the teacher-student model.

[0026] Figure 7 It is a segmentation effect diagram of the student model for new class images. Detailed implementation manners

[0027] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0028] A continuous learning instance segmentation method based on uncertainty score pseudo-labels specifically includes the following steps:

[0029] Step 1: Obtain the dataset of task t, including new class images and new class instance segmentation labels;

[0030] Step 11: The new class instance segmentation labels do not include the instance segmentation labels of the old classes. As shown in the attached Figure 1 figure, it is the instance segmentation label map of the new class. The figure contains two instances, namely the instance umbrella and the instance person. Among them, the instance umbrella is the new class and the person is the old class;

[0031] Step 2: Judge task t. If t = 1, construct the model M t of task t; if t>1, construct the teacher model M t-1 ; The model construction for task t is carried out according to the task step diagram shown in the attached Figure 2 figure to construct the corresponding model;

[0032] Step 21: Judge task t. If t = 1, then use the Mask RCNN model as the model M t of the current task, and directly input the dataset of task t into M t for training. The obtained M t is the required model; if t>1, then use the model trained in the t-1 task as the teacher model M t-1 ;

[0033] Step 3: Design as shown in the attachedFigure 3 The uncertainty score pseudo-labeling method shown inputs the new-class images of task t into the teacher model M t-1 to obtain instance results Obtain the uncertainty score pseudo-labels of the new-class images in task t, as shown in the appendix Figure 4 The figure shown is the prediction result graph of the teacher model for the input picture, as shown in the appendix Figure 5 The figure shown is the uncertainty score pseudo-label graph of the new-class images, and merge the uncertainty score pseudo-labels of the old-class instances and the segmentation labels of the new-class instances in the new-class images to obtain the new dataset of task t;

[0034] Step 31: Input the new dataset of task t into the teacher model M t-1 to obtain instance results For each instance in the new-class images It is expressed as follows:

[0035]

[0036] where represents the category of instance , represents the bounding box of instance , represents the instance mask of instance . The superscript t - 1 represents task t - 1, and the subscript i represents the iteration variable for all instances in the new-class images;

[0037] Step 32: Input the instance results into the uncertainty score pseudo-labeling method to obtain the uncertainty score pseudo-labels of the new-class images in task t:

[0038]

[0039] where is the label score of instance , is the confidence score that model M t-1 predicts instance to be of class c, represents the training target of the i-th instance in task t, represents that instance belongs to class c0, c0 is the background class, C 1:t-1 ={C 1 , C 2 ,…, C t-1} is the set of old-class categories, C t is the new-class category of task t, C t+1:T ={C t+1 , C t+2 ,…, CT} is the future category, and T represents the number of tasks; the first line of the formula indicates that the instance is a truly new class instance, so the score of its label is 1.0; the second line of the formula indicates that the instance belongs to the pseudo-background, so the label score of it is the confidence score of the prediction result

[0040] Step 33: Combine the uncertainty score pseudo-label and the segmentation label of the new class instance. The combination process is as follows:

[0041]

[0042] Among them, represents the segmentation label of the instance in the new dataset of task t, and y bg represents the label of the background class instance. The first line of the formula indicates that the instance is a truly new class instance, so its label is the segmentation label of the new class instance. The second line indicates that the instance is a pseudo-background, so its label is the uncertainty score pseudo-label;

[0043] Step 34: Since the segmentation label of the instance in the new dataset is different from the segmentation label of the new class instance, the training loss function of the segmentation label of the instance in the new dataset is defined as

[0044] Step Four: Build the student model M t when constructing task t, input the new dataset into the FODS model based on the teacher-student model architecture. The architecture diagram is as attached Figure 6 shown. Then, extract features from the new class images to obtain teacher features and student features, and perform the last layer of feature distillation:

[0045] Step 41: Build the FODS model based on the teacher-student architecture. The architecture diagram is as attached Figure 6 shown;

[0046] Step 42: Train the feature extractor of the teacher model M t-1 to extract features from the new class images, and use the last layer of features extracted as the teacher features Similarly, train the feature extractor of the student model to obtain student features

[0047] Step 43: Input the teacher features and student features into the last layer of feature distillation method. The process of using the L2 loss function as the loss function of the last layer of feature distillation method is:

[0048] ​

[0049] Among them, φ t is a global average pooling operation, and ||·|| 2 is the L2 loss.

[0050] Step Five: Obtain the teacher model M t-1 The instance results predicted by the teacher model M for the new dataset of task t Among them Use the cross-entropy function for output distillation;

[0051] Step 51: Randomly select a proposal box p from the RPN network,

[0052]

[0053] Among them, represents random sampling of x. The specific operation is to sort all proposal boxes in descending order of scores, select the top 128 with the highest scores, and then randomly select 64 from the selected 128 proposal boxes; is the RPN network at task t - 1, represents sequentially performing t on the new-class images X and

[0054] Step 52: For the selected proposal box p, calculate the classification and bounding boxes of the teacher model and the student model respectively;

[0055] Step 53: Calculate the cross-entropy loss for the classification and bounding boxes of the teacher model and the student model. The specific calculation process is as follows:

[0056]

[0057]

[0058] Among them, N p is the number of sampled proposal boxes, set to 64, represents the ROI classifier, where the first parameter is the feature extractor, and the second parameter p is the proposal box for calculating the classification, represents the ROI bounding box regression, and its parameters are the same as represents class distillation, represents bounding box distillation;

[0059] Step 54: Perform output distillation on the teacher model and the student model, including class distillation and bounding box distillation

[0060]

[0061] Step 6: Integrate the loss functions of the last-layer feature distillation in Step 4 and the output distillation in Step 5 to train the FODS model. The loss function is as follows:

[0062]

[0063] where α, β, and γ are hyperparameters; through training, obtain the student model, which is the model required to solve the continuous learning instance segmentation; the effect diagram is as shown in the appendix Figure 7 as follows.

[0064] The complete description of the entire network process is as follows:

[0065] Step 1: Obtain the dataset for task t, including new-class images and new-class instance segmentation labels;

[0066] Step 2: If task t = 1, construct a Mask RCNN model and input the dataset for task t for training to obtain model M 1 , which is the continuous learning instance segmentation model required for task t = 1; if task t > 1, go to Step 3;

[0067] Step 3: Use the model trained in task t - 1 as the teacher model M t-1 for task t, input the new-class images into the teacher model M t-1 to obtain the uncertainty score pseudo-labels for task t, and merge the uncertainty score pseudo-labels and the new-class instance segmentation labels to obtain the new dataset for task t;

[0068] Step 4: Construct the student model M t for task t, and based on the teacher model M t-1 for task t, build a Feature-Output Distillation and Score Pseudo Labels (FODS) model based on the teacher-student architecture;

[0069] Step 5: Input the new dataset for task t into the FODS model, train the feature extractors of the teacher model and the student model respectively for feature extraction to obtain the teacher features and the student features Use the L2 loss for the last-layer feature distillation when training the student model;

[0070] Step 6: Select an appropriate number of proposal boxes, input them into the teacher-student architecture, calculate the corresponding classes and bounding boxes of the teacher model and the student model respectively, and use the cross-entropy loss function for classification distillation and bounding box distillation to complete the output distillation;

[0071] Step 7: Integrate the loss functions of the last-layer feature distillation in Step 5 and the output distillation in Step 6, train the overall network architecture to obtain the student model, and complete the final image segmentation prediction.

Claims

1. A continuous learning instance segmentation method based on uncertainty score pseudo-labels, characterized in that, It includes the following steps: Step 1. Obtain the dataset when the task t, t > 1, including new class images and new class instance segmentation labels; Step 2. Use the model M trained at task t-1 t-1 as the teacher model, construct an uncertainty score pseudo-label method, obtain the uncertainty score pseudo-labels of old-class instances in new-class images at task t, and merge the uncertainty score pseudo-labels of old-class instances in new-class images and the segmentation labels of new-class instances to obtain a new dataset for task t: First, use the model M trained when using task t-1 t-1 as the teacher model; Secondly, input the new-class image at task t into the teacher model to obtain the instance result For each instance in the new-class image Its instance result Is expressed as follows: Among them, represents the category of the instance , represents the bounding box of the instance , represents the instance mask of the instance, and i is an iteration variable; ​ After that, add uncertainty to the old-class instances in the new-class images, that is, the instance results are used as the pseudo-labels of the uncertainty scores of the old-class instances in the new-class images for task t. The calculation method based on the uncertainty score pseudo-label method is as follows: Among them, is the pseudo-label score of the instance , and is the confidence score that the model M t-1 predicts the instance to be of class c. represents the training objective of the i-th instance in task t, indicating that the instance belongs to class c0, where c0 is the background class, and C 1:t-1 ={C 1 , C 2 , …, C t-1} is the set of old-class categories, C t is the new category at task t, and C t+1:T ={C t+1 , C t+2 , …, C T} is the set of future categories. T represents the number of tasks. The first line of the formula indicates that the instance is a truly new-class instance, so the score of its label is 1.

0. The second line of the formula indicates that this is a pseudo-background, that is, an old-class instance in a new-class image, so the label score is the confidence score of the prediction result After that, merge the uncertainty score pseudo-labels of the old class instances and the new class instance segmentation labels in the new class images of task t. The merging process is as follows: Among them, represents the instance segmentation label in the new dataset of task t, y bg represents the label of the background class instance. The first line of the formula indicates that the instance is a truly new class instance, so its label is the new class instance segmentation label; the second line indicates that the instance is a pseudo-background, so its label is the uncertainty score pseudo-label; Finally, since the instance segmentation labels in the new dataset are different from the instance segmentation labels of the new classes, the loss function trained using the instance segmentation labels in the new dataset is defined as Step 3. Student model M when constructing task t t , input the new dataset into the FODS model based on the teacher-student model architecture for feature extraction to obtain teacher features and student features, and perform the last layer feature distillation: First, train the teacher model M t-1 's feature extractor Extract features from the images of new classes in the new dataset to obtain teacher features Similarly, obtain student features where L represents the last layer, t represents task t, and t - 1 represents task t - 1; After that, input the teacher features and student features into the last layer feature distillation method, and use the L2 loss function as the loss function of the last layer feature distillation method. The formula is: where φ t is a global average pooling operation, and ||·|| 2 is the L2 loss; Step 4. Obtain the teacher model M t-1 Instance results predicted by the new dataset for task t Among them Perform output distillation using the cross-entropy function: First, randomly select a proposal box p from the RPN network; Among them, represents the random sampling of x. The specific operation is to sort all the proposal boxes in descending order of scores, select the top 128 with the highest scores, and then randomly select 64 from the selected 128 proposal boxes. is the RPN network at task t - 1. represents successively for the input image X t to perform and Second, for the selected proposal box p, calculate the classification and bounding boxes of the teacher model and the student model respectively; After that, calculate the cross-entropy loss for the classification and bounding boxes of the teacher model and the student model. The specific calculation process is: Among them, N p is the number of sampling suggestion boxes, set to 64, represents the ROI classifier, where the first parameter is the feature extractor, and the second parameter p is the suggestion box for calculating classification. Similarly, represents the ROI bounding box regression, and its parameters are the same as represents class distillation, represents bounding box distillation; Finally, output distillation is performed on the teacher model and the student model including classification distillation and bounding box distillation Step 5. Fuse the loss function of the last layer feature distillation in Step 3 and the output distillation in Step 4 to train the FODS model. Its loss function is: where α, β, γ are hyperparameters; through training, obtain the student model, which is the model required to solve continuous learning instance segmentation.