Incremental small sample instance segmentation method based on transfer learning
By adding lightweight adapters and using knowledge distillation strategies to the instance segmentation model, the problem of the model forgetting old knowledge when learning new knowledge is solved, and better knowledge integration between base class and new class is achieved, and the overall performance and balance of the model are improved.
Patent Information
- Application Number
- CN202510297910.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing example segmentation models are prone to forget old knowledge when learning new knowledge, resulting in sharp deterioration of segmentation performance on old categories, and overfitting on new categories makes it difficult to balance stability and plasticity.
The incremental small sample instance segmentation method based on transfer learning is adopted, and by adding lightweight adapters and knowledge distillation strategies, the basic knowledge and new knowledge are better integrated, and the interference with old knowledge in the learning process of new categories is reduced.
It effectively alleviates the catastrophic forgetting of the model in the old category, reduces the risk of overfitting on the new category, improves the overall performance of the model in the base class and the new class, and balances stability and plasticity.
Smart Images

Figure CN120147646A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of instance segmentation, and particularly to an incremental few-shot instance segmentation method based on transfer learning. Background Art
[0002] Instance segmentation is one of the fundamental tasks in computer vision, integrating the core elements of object detection and semantic segmentation. Object detection aims to predict the class and location of each foreground object in an image, where the location is given in the form of a bounding box, while semantic segmentation focuses on assigning semantic classes to each pixel in the image. In contrast, instance segmentation is more complex and challenging. Instance segmentation aims to segment each foreground object in the image and distinguish different objects belonging to the same class. The output of instance segmentation usually includes the class and segmentation mask of each foreground object.
[0003] With the continuous progress of deep learning, instance segmentation models based on deep learning have achieved significant performance improvements and become the mainstream solution for instance segmentation tasks. However, when applying these models in complex real-world scenarios, many challenges still remain. First, the excellent performance of deep learning-based instance segmentation models depends on training on high-precision, large-scale labeled image datasets. However, in the real world, the difficulty of obtaining large-scale image datasets and the high cost of manual annotation limit the further application of these models. Second, data in the real world often arrives continuously in the form of a stream, requiring the model to have the ability to continuously learn new knowledge without accessing past data. Unfortunately, in the case where previous data cannot be obtained for unified training, current instance segmentation models often forget the old knowledge learned previously while learning new knowledge, resulting in catastrophic forgetting of the model on old knowledge and a sharp deterioration in the segmentation performance on old classes. To address the above challenges, researchers have proposed incremental few-shot instance segmentation.
[0004] Incremental few-shot instance segmentation aims to enable the model to quickly learn how to segment new classes using a small number of new class samples without accessing previous samples, while maintaining the model's performance on old classes and preventing catastrophic forgetting of the model on old classes. There are mainly three challenges in the incremental few-shot instance segmentation task: (1) catastrophic forgetting on old classes; (2) overfitting caused by a small number of new class samples; (3) the model needs to make a trade-off between maintaining the performance of old classes and learning new classes, which is called the stability-plasticity dilemma. Among them, the ability to maintain the performance of old classes is called stability, and the ability to learn new classes is called plasticity.
[0005] In the incremental few-shot instance segmentation task, the entire dataset classes C are usually divided into two disjoint sets: the base class C base and the new class C novel, i.e., C base ∪C novel = C, wherein, each category in the base class has a large number of labeled samples available for training, while each category in the new class has only a small number of K labeled samples available for training (usually K = 1, 5, 10). The dataset composed of the samples of the base class C base is denoted as the base class dataset D base , and the dataset composed of the samples of the new class C novel is denoted as the new class dataset D novel . The task objective of incremental few-shot instance segmentation is to train such a model: the model is first trained on D base containing a large number of labeled samples, and then, without touching the samples in D base , it is trained on D nopvel containing only a small number of samples. Finally, the model performs well on C test = C base ∪C novel .
[0006] The existing incremental few-shot instance segmentation methods can be mainly divided into two categories, namely, the methods based on transfer learning and the methods based on meta-learning. For the methods based on transfer learning, such as iMTFA and iFS-RCNN, these methods usually adopt the method of "pre-training - fine-tuning". The model is trained with a large number of labeled samples on the base class dataset, and then when training on the new class dataset, the model learns new knowledge in the way of freezing the class-agnostic components of the model and fine-tuning the class-specific components of the model, while preventing the model from catastrophic forgetting on the old categories. For the methods based on transfer learning, such as Reft, it uses the method of meta-learning to enable the model to quickly adapt to new categories. These methods usually adopt the way of scenario training during training, and a series of scenarios E i = (I i , S i ) are set. Each scenario contains a training set called the support set and a test set called the query set. When training on the base class dataset, the small sample situation on the new class is simulated in this way, so that the model has the ability of rapid learning, and thus can quickly learn the new category knowledge on the new class dataset.
[0007] However, there are still some limitations in the current methods. In the methods based on transfer learning, when learning new category knowledge through the method of "pre-training - fine-tuning", on the one hand, fine-tuning the class-specific components of the model will damage the original base class knowledge in the model, inevitably resulting in the performance degradation of the model on the base class. On the other hand, it is impossible to perform good feature integration of the base class knowledge and the newly learned new class knowledge in the model, and make full use of the original base class knowledge in the model. Summary of the Invention
[0008] To solve the problems existing in the prior art, the objective of the present invention is to provide an incremental few-shot instance segmentation method based on transfer learning. By adding a lightweight adapter and adopting a knowledge distillation strategy, the present invention aims to overcome the problems that previous methods cannot fully utilize the knowledge of the base classes and the performance of the old classes deteriorates when fine-tuning the model, and to achieve better integration of the base class knowledge and the new class knowledge, and better balance the stability-plasticity dilemma.
[0009] To achieve the above objective, the technical solution adopted by the present invention is as follows: An incremental few-shot instance segmentation method based on transfer learning, comprising the following steps:
[0010] Step 1, construct an incremental few-shot instance segmentation model, where the incremental few-shot instance segmentation model includes a feature extractor, a Transformer Encoder, a Transformer Decoder, and two prediction heads connected in sequence;
[0011] Step 2, train the incremental few-shot instance segmentation model;
[0012] Step 3, use the trained incremental few-shot segmentation model to segment the foreground objects of the base classes and the new classes in the image, and obtain the category and segmentation mask of each foreground object.
[0013] As a further improvement of the present invention, in step 1, the feature extractor takes the image after data augmentation as input and extracts multi-scale features of the image; the Transformer Encoder takes the multi-scale features output by the feature extractor as input, enriches the image features further through the self-attention mechanism, and generates a pixel embedding map using the enriched image features, and at the same time provides the top k tokens (topk tokens) for the Transformer Decoder; the Transformer Decoder initializes its object query with the top k tokens provided by the Transformer Encoder, realizes the interaction between object queries through the self-attention mechanism, and realizes the interaction between object queries and the image features output by the Transformer Encoder through the cross-attention mechanism, and finally outputs the embedding of the object query as the input of two prediction heads; the two prediction heads are a classification head and a segmentation head respectively. The embedding of the object query is input into the classification head to obtain the category of each foreground object. The embedding of the object query is input into the segmentation head, and its output is dot-producted with the pixel embedding map output by the Transformer Encoder to obtain the segmentation mask of each foreground object.
[0014] As a further improvement of the present invention, the multi-scale features include feature maps of four scales, and their sizes are respectively
[0015] As a further improvement of the present invention, the Transformer Encoder first performs dimension and size conversion on the multi-scale feature maps output by the feature extractor through a projection layer, and converts the feature map with a size of to size, converts the feature map with a size of to Converts the feature map with a size of into size and size respectively, and then flattens and concatenates the four-scale feature maps obtained by the conversion as the tokens of the Transformer Encoder and inputs them into the first Transformer Encoder layer; the tokens finally output by the Transformer Encoder are used in three places: (1) the tokens are reshaped into a special The feature map with a size of is upsampled by a factor of two to size and add it to the feature map with size output from the feature extractor to obtain a pixel embedding map; (2) Input these tokens into the classification head, select the top-k tokens according to the class confidence, and use them as the initialization of the object query in the Transformer Decoder; (3) Input the tokens as the enhanced image features into the Transformer Decoder.
[0016] As a further improvement of the present invention, the Transformer Encoder is stacked by multiple Transformer Encoder layers, and each Transformer Encoder includes a multi-head self-attention and a multi-layer perceptron (MLP).
[0017] As a further improvement of the present invention, the projection layer includes four convolutional layers. Among them, three convolutional layers perform dimensional conversion on the three feature maps with smaller scales, and one convolutional layer performs dimensional and size conversion on the feature map with the smallest scale; the four convolutional layers are respectively: (1) The convolutional kernel size is 1×1, the stride is 1, and it is used to convert the feature map with size to size ; (2) The convolutional kernel size is 1×1, the stride is 1, and it is used to convert the feature map with size to size 256; (3) The convolutional kernel size is 1×1, the stride is 1, and it is used to convert the feature map with size to size ; (4) The convolutional kernel size is 3×3, the stride is 2, and it is used to convert the feature map with size to size .
[0018] As a further improvement of the present invention, the pixel embedding map is a tensor with size and is used to perform a dot product with the output of the segmentation head to obtain the segmentation result of the foreground object.
[0019] As a further improvement of the present invention, the classification head is a classifier based on cosine similarity, and its output is the category of the foreground object; the segmentation head is a three-layer multi-layer perceptron, including a hidden layer, and its output is a one-dimensional vector with a dimension of 256. After performing a dot product of this vector with the pixel embedding map generated by the Transformer Encoder, a tensor with size is obtained. This tensor is upsampled to a size of H×W×1 and passed through a sigmoid function to obtain the binary segmentation mask of the foreground object.
[0020] As a further improvement of the present invention, the Transformer Decoder is stacked by multiple Transformer Decoder layers, and each Transformer Decoder layer includes a multi-head self-attention, a multi-head cross-attention, and a multi-layer perceptron (MLP).
[0021] As a further improvement of the present invention, the training phase in step 2 includes a basic training phase and a few-shot incremental fine-tuning phase:
[0022] The basic training phase is to train an incremental few-shot instance segmentation model on the base class dataset, and the model after the training is completed is called the basic model;
[0023] The few-shot incremental fine-tuning phase specifically includes the following steps:
[0024] The first step is to freeze all the parameters of the basic model;
[0025] The second step is to recreate a new incremental few-shot instance segmentation model, which is called the new model, and initialize the new model with the parameters of the basic model;
[0026] The third step is to insert a parallel adapter into the multi-layer perceptron of the first layer of the Transformer Encoder of the new model; the adapter is a bottleneck structure, including a downsampling layer, a non-linear activation function ReLU, and an upsampling layer;
[0027] The fourth step is to freeze all the parameters of the new model except the classifier and the adapter;
[0028] The fifth step is to train the new model on the new class dataset through the knowledge distillation strategy, using the basic model as the teacher model and the new model as the student model, and calculating the knowledge distillation loss on the tokens output by the first layer of the Transformer Encoder;
[0029] The sixth step is to add the classification weights of the basic model to the classifier of the new model. The new model after adding the classification weights of the basic model is the final model in the few-shot incremental fine-tuning phase, which has the ability to detect and segment all foreground objects of the base classes and new classes in the image.
[0030] The beneficial effects of the present invention are:
[0031] 1. Efficient fine-tuning: The present invention enables the model to learn new category knowledge by fine-tuning the adapter. Compared with fine-tuning the class-agnostic components in the model, the present invention requires fewer parameters to be fine-tuned, which can reduce the time and resources required for fine-tuning on the new class dataset, and at the same time reduce the impact on the original model, further alleviating the catastrophic forgetting of the model on the old categories.
[0032] 2. Make full use of the knowledge of the base class: In the present invention, the adapter is inserted in parallel into the multi-layer perceptron of the first layer of the Transformer Encoder. When fine-tuning the adapter, the knowledge of the base class stored in the original multi-layer perceptron is not affected, and the original knowledge of the base class is integrated into the newly learned knowledge of the new class, helping the model to better detect new categories. Compared with the previous methods, the present invention learns new knowledge without damaging the knowledge of the base class and makes more full use of the knowledge of the base class.
[0033] 3. Overcome forgetting: The present invention reduces the interference to the old knowledge during the learning process of new categories by adding an adapter and adopting a knowledge distillation strategy, alleviates the catastrophic forgetting of the model on the old categories, and better balances the stability - plasticity dilemma between the base class and the new class. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of an incremental few-shot instance segmentation model in an embodiment of the present invention;
[0035] Figure 2 It is a schematic diagram of the structure of each layer of the Transformer Encoder in an embodiment of the present invention;
[0036] Figure 3 It is a schematic diagram of the structure of each layer of the Transformer Decoder in an embodiment of the present invention;
[0037] Figure 4 It is a schematic diagram of the two-stage training of the incremental few-shot instance segmentation model in an embodiment of the present invention;
[0038] Figure 5 It is a schematic diagram of the structure of the Transformer Encoder layer after adding an adapter in an embodiment of the present invention;
[0039] Figure 6 It is a schematic diagram of the visualization result in the MS COCO2014 dataset with 10 samples in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0041] As Figure 1 shown, the present invention proposes an incremental few-shot instance segmentation method based on transfer learning. The construction of the incremental few-shot instance segmentation model includes the following steps:
[0042] Step 1: Feature extraction step. Given an initial image, after performing data augmentation operations such as random cropping and flipping and normalization on it, it is used as the input to the feature extractor to extract multi-scale features of the image;
[0043] Step 2: Transformer Encoder. The multi-scale features obtained in Step 1 are input into the Transformer Encoder. The Transformer Encoder first performs dimension and size conversion on the multi-scale features through a projection layer, then flattens and concatenates the converted multi-scale features and inputs them into the Transformer Encoder, and further enriches the features through multiple stacked Transformer Encoder layers. The output of the Transformer Encoder includes three parts: (1) a pixel embedding map for obtaining the final segmentation result; (2) the top k tokens for initializing the object query of the Transformer Decoder object; (3) the enriched image features.
[0044] Step 3: Transformer Decoder. The Transformer Decoder first initializes the object query using the top k tokens obtained in Step 2, and then inputs it into multiple stacked Transformer Decoder layers. In each Transformer Decoder layer, first, the interaction between object queries is realized through the self-attention mechanism, and then the interaction between the object query and the enriched image features in Step 2 is realized through the cross-attention mechanism. The embedding of the object query finally output by the Transformer Decoder will be used as the input to two prediction heads.
[0045] Step 4: Result prediction. The two prediction heads use the embedding of the object query obtained in Step 3 as the input to obtain the segmentation result of the final foreground object. Among them, the classification head will output the category of the foreground object; the segmentation head will output a one-dimensional vector of size 256, and perform a dot product with the pixel embedding map obtained in Step 2, and apply the sigmoid function to the dot product result to obtain the binary segmentation mask of the foreground object.
[0046] As Figure 2 shown, each Transformer Encoder layer includes a multi-head self-attention and a multi-layer perceptron. In this embodiment, the Transformer Encoder is stacked by 6 Transformer Encoder layers.
[0047] As Figure 3 shown, each Transformer Decoder layer includes a multi-head self-attention, a multi-head cross-attention, and a multi-layer perceptron. In this embodiment, the Transformer Decoder is stacked by 9 Transformer Decoder layers.
[0048] As Figure 4 shown, the training of the entire model includes a basic training stage and a few-shot incremental fine-tuning stage. In this embodiment, the training process is as follows:
[0049] (1) Basic training stage:
[0050] In the basic training stage, the model will be trained on the base class training set, and the images in the base class training set only contain the annotations of the base class objects.
[0051] The loss function in the basic training stage is:
[0052]
[0053] where is the classification loss, and sigmoid focal loss is used as the loss function; is the segmentation loss, which is composed of the cross-entropy loss function and the dice loss function, that is is the cross-entropy loss function, is the dice loss function.
[0054] The model obtained in the basic training stage is called the basic model.
[0055] (2) Few-shot incremental fine-tuning stage:
[0056] In the few-shot incremental fine-tuning stage, the model will be trained on the new class dataset, and the images in the new class dataset only contain the annotations of the new class objects.
[0057] As Figure 5 shown, in the few-shot incremental fine-tuning stage, in this embodiment, the model is fine-tuned on the new class dataset by adding an adapter and adopting a knowledge distillation strategy, so that the model can learn new class knowledge without damaging the original old class information in the model, alleviate the catastrophic forgetting of the model on the old classes, and better balance the stability-plasticity dilemma.
[0058] First, in this embodiment, an adapter parallel to the multi-layer perceptron is inserted into the Transformer Encoder layer. The Transformer Encoder layer after adding the adapter is as Figure 3As shown, the adapter is a bottleneck structure, consisting of a lower projection layer, a non-linear activation function ReLU, and an upper projection layer. After adding the adapter, the output of this part is:
[0059] x o = σ(x i W down )W up + LayerNorm(Multi-Layer Perceptron(x i ) + x o )
[0060] Among them, W down represents the lower projection layer, W up represents the upper projection layer, σ represents the ReLU function, x i represents the features input into the multi-layer perceptron, x o represents the final output features, and LayerNorm represents layer normalization.
[0061] In this embodiment, the adapter is only added to the first Transformer Encoder layer, and the intermediate dimension of the adapter is set to 64.
[0062] In the small sample incremental fine-tuning stage, this embodiment uses the base model obtained in the base training stage as the teacher model in the knowledge distillation strategy, and the new model after inserting the adapter as the student model in the knowledge distillation strategy. At the same time, the remaining parameters of the student model except for the adapter and the classifier are initialized with the parameters of the teacher model, and the remaining parameters of the student model except for the adapter and the classifier are frozen.
[0063] This embodiment performs knowledge distillation on the tokens output by the first layer of the Transformer Encoder. The specific steps are as follows:
[0064] Step 1: Reshape the tokens output by the first layer of the Transformer Encoder back into multi-scale feature maps, that is, reshape the tensor with the shape of into four feature maps with sizes of respectively.
[0065] Step 2: Construct an identification mask for identifying the base class object area and the new class object area in the feature map. First, downsample the true segmentation mask of the new category object with size H×W×1 in the image through bilinear interpolation to obtain Binary segmentation masks of four sizes. In this binary segmentation mask, if the value of a pixel in the segmentation mask is 1, it proves that it is in the area of the new class object, and its corresponding identification mask is 0. If the value of a pixel in the segmentation mask is 0, it proves that it is in the area of the base class object, and its corresponding identification mask is 1. In this way, the identification masks corresponding to the four feature maps are obtained.
[0066] Step 3: Apply the four identification masks obtained in Step 2 to the calculation of the knowledge distillation loss, so that the calculation of the knowledge distillation loss is only carried out in the area of the base class object, so as to reduce the constraint on the model to learn new category knowledge while ensuring that the features of the base class object area of the student model do not deviate too much from the features of the base class object area of the teacher model. Calculate the knowledge distillation loss through the L2 loss on the four feature maps, that is:
[0067]
[0068] Among them, flag_mask represents the identification mask, represents the feature map of the new model, represents the feature map of the base model, w i and h i respectively represent the width and height of the feature map f i .
[0069] The total loss in the small sample incremental fine-tuning stage is:
[0070]
[0071] Among them, α is the weight of the knowledge distillation loss and can be used as a trade-off parameter, which is set to 0.05 in this embodiment.
[0072] After the model training is completed, merge the classification weights of the base class in the base model into the classifier of the new model. The model obtained at this time is the final model, and this model can segment the base class objects and new class objects that appear in the image.
[0073] In order to verify the performance of the above method, the following experiments are designed.
[0074] In this embodiment, ResNet-50 is used as the backbone, and experiments are carried out on two benchmark datasets, MS COCO2014 and LVIS, for verification. MS COCO2014 contains a total of 80 categories, among which 20 categories intersecting with the PSCAL VOC dataset are used as new categories, and the remaining 60 categories are used as base categories. LVIS contains a total of 1230 categories. According to the frequency of each category appearing in the images, they are divided into rare categories (less than 10 images), common categories (10 - 100 images), and frequent categories (more than 100 images). During the experiment, rare categories are used as new categories, and common categories and frequent categories are used as base categories.
[0075] On the MS COCO2014 dataset, in this embodiment, the performance of the model on the base categories, new categories, and all categories is considered when the number of samples is 1, 5, and 10 respectively. All experimental results are the average results of 10 random samplings. For the MS COCO2014 dataset, this embodiment uses two metrics, average precision (AP) and average precision with a confidence threshold greater than 0.5 (AP50), as evaluation metrics; for the LVIS dataset, average precision (AP) is used as the evaluation metric.
[0076] The experimental results on the MS COCO 2014 dataset are shown in Table 1. Under the three settings of the number of samples, this embodiment has achieved the best results and far exceeded other methods.
[0077] Table 1 Comparison of the performance of incremental few-shot instance segmentation methods on the MS COCO2014 dataset.
[0078]
[0079] The experimental results on the LVIS dataset are shown in Table 2. It can be seen that on the large dataset with a large number of categories, this embodiment still has great advantages and far exceeds other methods in all indicators.
[0080] Table 2 Comparison of the performance of incremental few-shot instance segmentation methods on the LVIS dataset
[0081]
[0082] To further verify the effectiveness of the proposed method, three additional ablation experiments are designed.
[0083] To verify the effectiveness of the method components, ablation experiments were conducted on the important components of the method, and the results are shown in Table 3. In Table 3, to verify that fine-tuning the lightweight adapter with parallel insertion is a better choice than fine-tuning the other model components, the experimental results of fine-tuning the projection layer and fine-tuning the adapter were compared. It can be seen that fine-tuning the adapter can better alleviate the catastrophic forgetting of the model on the base classes while achieving a competitive effect with fine-tuning the projection layer on the new classes. Therefore, fine-tuning the adapter can achieve a better balance between effectively retaining old knowledge and learning new knowledge and is a better choice than fine-tuning the projection layer. By comparing the experimental results in the second and third rows of the table, it can be seen that knowledge distillation can further alleviate catastrophic forgetting with a slight decrease in the performance on the new classes.
[0084] Table 3 Ablation results of each component of the model with a sample size of 10 on the MS COCO2014 dataset
[0085]
[0086] To verify the impact of the intermediate dimension of the adapter on the model performance, the performance of the model on the base classes and new classes was tested with different intermediate dimensions of the adapter, and the results are shown in Table 4. It can be seen from Table 4 that when the intermediate dimension continuously decreases and the number of tunable parameters gradually decreases, the performance of the model on the base classes continuously increases, while the performance on the new classes continuously decreases. Therefore, the intermediate dimension of the adapter can be used as a trade-off parameter and selected according to the actual situation.
[0087] Table 4 Ablation results of the intermediate dimension of the adapter with a sample size of 10 on the MS COCO2014 dataset
[0088]
[0089] To verify the impact of the weight parameter α of the knowledge distillation loss on the model performance, the model performance with different values of the weight parameter was tested, and the results are shown in Table 5. When the value of the weight parameter is larger, it means that the knowledge distillation restricts the update of the model parameters more strictly. At this time, it is more beneficial to alleviate the performance degradation of the model on the base classes, but at the same time, it also restricts the learning of new knowledge of the model on the new classes. Therefore, the value of the weight parameter can also be used as a trade-off parameter and selected according to the actual situation.
[0090] Table 5 Ablation results of the weight parameter α of the knowledge distillation loss with a sample size of 10 on the MS COCO2014 dataset
[0091]
[0092]
[0093] To further demonstrate the effects of this embodiment, a visualization experiment was conducted under the setting of the MS COCO2014 dataset with a sample size of 10, and the results are as Figure 6 shown. It can be seen from Figure 6 that this embodiment can well segment and classify the base class and new class objects in the image.
[0094] The above-described embodiments only represent specific implementation manners of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. An incremental small sample instance segmentation method based on transfer learning, characterized in that: The following steps are involved: Step 1: construct an incremental small sample instance segmentation model, wherein the incremental small sample instance segmentation model includes a feature extractor, a Transformer encoder, a Transformer decoder, and two prediction heads connected in sequence; Step 2: training the incremental small sample instance segmentation model; Step 3: Use the trained incremental small sample segmentation model to segment the foreground objects of the base class and the new class in the image to obtain the category and segmentation mask of each foreground object.
2. The incremental small sample instance segmentation method based on transfer learning according to claim 1, characterized in that: In step 1, the feature extractor takes the image after data enhancement as input to extract the multi-scale features of the image; the Transformer encoder takes the multi-scale features output by the feature extractor as input, further enriches the image features through the self-attention mechanism, and uses the enriched image features to generate a pixel embedding map, and provides the Transformer decoder with the selected first k tokens; the Transformer decoder uses the first k tokens provided by the Transformer encoder as the initialization of its object query, realizes the interaction between object queries through the self-attention mechanism, realizes the interaction between object queries and the image features output by the Transformer encoder through the cross-attention mechanism, and finally outputs the embedding of the object query as the input of the two prediction heads; the two prediction heads are the classification head and the segmentation head, respectively, the embedding of the object query is input into the classification head to obtain the category of each foreground object, the embedding of the object query is input into the segmentation head, and its output is dot-producted with the pixel embedding map output by the Transformer encoder to obtain the segmentation mask of each foreground object.
3. The incremental small sample instance segmentation method based on transfer learning according to claim 2, characterized in that: The multi-scale features include feature maps of four scales, whose sizes are 4. The incremental small sample instance segmentation method based on transfer learning according to claim 3, characterized in that: The Transformer encoder first converts the dimensions and sizes of the multi-scale feature maps output by the feature extractor through the projection layer, The feature map is converted to size, and set the size to The feature map is converted to The size is The feature maps are converted into Size and The four scale feature maps are then flattened and concatenated as the Transformer encoder tokens and input into the first Transformer encoder layer. The final output token of the Transformer encoder is used in three places: (1) The token is reshaped into a feature map and The feature map of size is upsampled twice to size and compare it with the output size of the feature extractor The feature maps of the first k pixels are added to obtain the pixel embedding map; (2) these tokens are input into the classification head, and the top k tokens are selected according to the category confidence, which are used as the initialization of the object query in the Transformer decoder; (3) the tokens are used as the enhanced image features and input into the Transformer decoder.
5. The incremental small sample instance segmentation method based on transfer learning according to claim 4, characterized in that: The Transformer encoder is composed of multiple stacked Transformer encoder layers, each of which includes a multi-head self-attention and a multi-layer perceptron.
6. The incremental small sample instance segmentation method based on transfer learning according to claim 4, characterized in that: The projection layer includes four convolution layers, three of which convert the dimensions of three feature maps with smaller scales, and one convolution layer converts the dimensions and size of a feature map with the smallest scale. The four convolution layers are: (1) the convolution kernel size is 1×1, the step size is 1, and it is used to convert The feature map of size is converted to size; (2) the convolution kernel size is 1×1 and the step size is 1, which is used to The feature map of size is converted to size; (3) The convolution kernel size is 1×1 and the step size is 1, which is used to The feature map of size is converted to size, (4) the convolution kernel size is 3×3, the stride is 2, which is used to The feature map of size is converted to size.
7. The incremental small sample instance segmentation method based on transfer learning according to claim 4, characterized in that: The pixel embedding map is of size The tensor is used for dot product with the output of the segmentation head to obtain the segmentation result of the foreground object.
8. The incremental small sample instance segmentation method based on transfer learning according to claim 2, characterized in that: The classification head is a classifier based on cosine similarity, and its output is the category of the foreground object; the segmentation head is a three-layer multi-layer perceptron, including a hidden layer, and its output is a one-dimensional vector with a dimension of 256. The vector is dot-producted with the pixel embedding graph generated by the Transformer encoder to obtain The tensor is upsampled to H×W×1 size and passed through the sigmoid function to obtain the binary segmentation mask of the foreground object.
9. The incremental small sample instance segmentation method based on transfer learning according to claim 4, characterized in that: The Transformer decoder is composed of a plurality of stacked Transformer decoder layers, each of which includes a multi-head self-attention, a multi-head cross-attention and a multi-layer perceptron.
10. The incremental small sample instance segmentation method based on transfer learning according to claim 1, characterized in that: The training phase in step 2 includes a basic training phase and a small sample incremental fine-tuning phase: The basic training stage is to train an incremental small sample instance segmentation model on a base class data set, and the model after the training is called a basic model; The small sample increment fine-tuning stage specifically includes the following steps: The first step is to freeze all parameters of the basic model; Step 2: Recreate a new incremental small sample instance segmentation model, call it the new model, and initialize the new model using the parameters of the base model; The third step is to insert a parallel adapter into the multilayer perceptron of the first layer of the Transformer encoder of the new model; the adapter is a bottleneck structure, including a downsampling layer, a nonlinear activation function ReLU and an upsampling layer; Step 4: Freeze all parameters of the new model except the classifier and adapter; Step 5: Train the new model using the knowledge distillation strategy on the new dataset. Use the base model as the teacher model and the new model as the student model. Calculate the knowledge distillation loss on the token output by the first layer of the Transformer encoder. Step 6. Add the classification weights of the base model to the classifier of the new model. The new model after adding the classification weights of the base model is the final model in the small sample incremental fine-tuning stage, which has the ability to detect all base class and new class foreground objects in the segmented image.
Citation Information
Patent Citations
Metalearning-based few-sample instance segmentation method of unified framework
CN116206107A
Cited By
Lightweight star catalogue target detection method and system
CN122090288A