Continuous Semantic Segmentation Method and System Based on Balanced Multi-Granularity Fusion Feature Distillation
Through balanced multi-particle fusion feature distillation technology, combined with parallel self-attention module and moment-based channel attention module, the problems of catastrophic forgetting and prediction bias in continuous semantic segmentation are solved, and more efficient new task learning and feature transfer are achieved.
Patent Information
- Application Number
- CN202510170884.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Existing continuous semantic segmentation methods are prone to catastrophic forgetting when processing new tasks, and randomly initialized classifiers will produce severe prediction bias and cannot effectively migrate important information in multi-grained features.
The continuous semantic segmentation method based on balanced multi-particle size fusion feature distillation is adopted, and feature fusion and balance are performed through parallel self-attention modules and moment-based channel attention modules, and the consistency constraints of features are calculated in combination with knowledge distillation technology, and the total training loss is optimized to train the new model.
It effectively reduces the material and time overhead of retraining the semantic segmentation model in the face of continuous incoming data flow, enhances the learning ability of new tasks, and avoids catastrophic forgetting and prediction bias.
Smart Images

Figure CN119625328B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of continuous semantic segmentation, specifically to a continuous semantic segmentation method and system based on balanced multi-granularity fusion feature distillation. Background Art
[0002] With the rapid development of convolutional neural networks, the accuracy of semantic segmentation tasks has been greatly improved. However, most current semantic segmentation methods are trained based on static large batches of data, while in practical applications, the data that needs to be processed in various industries is mostly dynamic, open, and non-uniformly distributed data streams. The models obtained by this batch learning training can only recognize the categories that have appeared in the training set. For categories that have not appeared in the training set, the model can only be retrained using a new training set containing new categories, which not only causes a large waste of time and material resources, but also catastrophic forgetting will occur, that is, the performance of the model on the previously learned old categories will decrease significantly. To solve this problem, researchers have begun to introduce the continuous learning paradigm into semantic segmentation tasks, trying to enable the model to continuously learn new information using newly arrived data, learn to handle new tasks, and maintain the ability to handle old tasks.
[0003] Recently, many continuous semantic segmentation methods have tried to use knowledge distillation technology to constrain the deep features or output confidence consistency extracted by the old and new models, and randomly initialize the classifiers for new categories. However, these methods have two deficiencies: First, knowledge distillation only constrains the deep semantic features of a single granularity, ignoring the transfer of information crucial for the semantic segmentation task contained in other granularity features; Second, in the face of the non-discriminative features between old and new categories at the beginning of a new task, the randomly initialized classifier will produce serious prediction biases. Therefore, the key to the continuous semantic segmentation task lies in how to retain as much knowledge as possible that plays a key role in preventing the model from catastrophic forgetting, and how to eliminate the prediction biases brought by non-discriminative features at the beginning of a new task to strengthen the learning of new tasks. Summary of the Invention
[0004] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide a continuous semantic segmentation method and system based on balanced multi-granularity fusion feature distillation.
[0005] In the first aspect, the purpose of the present invention can be achieved through the following technical solutions: A continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation, the method includes the following steps:
[0006] Obtain an image, input the image into the feature extraction network of the pre-established old and new models, and output multi-layer features of the image; input the multi-layer features of the image into the pre-established parallel self-attention module, and output preliminary fused multi-granularity features;
[0007] Input the preliminarily fused multi-granularity features into a pre-established moment-based channel attention module to output balanced multi-granularity fused features; calculate the consistency constraint of the balanced multi-granularity fused features based on knowledge distillation, and calculate the total training loss based on the consistency constraint of the balanced multi-granularity fused features;
[0008] Train the new model based on the total training loss to obtain the trained new model, and use the trained new model to provide the optimal initial decision boundary for the new task based on the auxiliary classifier, where the new model is the task-stage model.
[0009] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the output process of the preliminarily fused multi-granularity features:
[0010] The parallel self-attention module consists of several parallel branches, and each branch follows the self-attention network structure in the vision transformer;
[0011] Convert the intermediate features of different layers of the model into the same size using bilinear interpolation;
[0012] Input the features of different layers after bilinear interpolation into different branches of the parallel self-attention module, and each branch corresponds to one layer of features;
[0013] Concatenate the outputs of each branch of the parallel self-attention module in the channel dimension to form the preliminarily fused multi-granularity features.
[0014] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the output process of the balanced multi-granularity fused features:
[0015] Based on the preliminarily fused multi-granularity features, calculate its moment statistics in the spatial dimension, including the first moment: mean, the second moment: variance, and the third moment: skewness, and obtain three feature statistic vectors corresponding to the three statistics;
[0016] Input the three feature statistic vectors into a multi-layer perceptron network to respectively output three channel attention vectors;
[0017] Add the three channel attention vectors and multiply the result by the preliminarily fused multi-granularity features to obtain the balanced multi-granularity fused features.
[0018] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the calculation process of the total training loss:
[0019] Based on the knowledge distillation technology, calculate the distance between the balanced multi-granularity fused features of the new and old models respectively as the training loss, and calculate the consistency constraint of the balanced multi-granularity fused features ;
[0020] Calculate the standard classification loss , and the confidence distillation loss for the final output of the model ;
[0021] Calculate the total loss of model training :
[0022]
[0023] Wherein, and are balance coefficients used to balance each loss.
[0024] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the calculation process of the balanced multi-granularity fusion feature consistency constraint is as follows:
[0025]
[0026] In the formula, and represent the spatial dimension size of the balanced multi-granularity fusion feature, represents the balanced multi-granularity fusion feature at the th row and th column in the spatial dimension, The subscript of indicates that the balanced multi-granularity fusion feature is extracted by the model of the task stage .
[0027] The standard classification loss is:
[0028]
[0029]
[0030] Wherein, represents the class label, and represent the size of the output mask, represents the task stage The confidence of the model for the pixel in the corresponding class output in the picture, represents taking the logarithm, represents the indicator function, which outputs 1 if its parameter is true, otherwise outputs 0, represents the segmentation confidence output by the model, is the pixel in the output mask, Represents a pixel The true label of and Are the dimensions of the output mask in the spatial dimension respectively, Represents the task stage The set of new classes learned by the model, Represents up to the task stage The total number of classes learned by the model;
[0031] The confidence distillation loss Is:
[0032] .
[0033] Wherein, Represents all the old classes learned by the model.
[0034] Combined with the first aspect, in some implementations of the first aspect, the method further includes: The process of the trained new model providing an optimal initialization decision boundary for the new task based on the auxiliary classifier is as follows:
[0035] Based on the classification loss variant Train an auxiliary classifier for each class:
[0036]
[0037] Wherein, Represents the class label, and Represent the dimensions of the output mask, Represents the task stage The model's confidence in the class corresponding to the pixel in the picture Output, Represents taking the logarithm, Represents the indicator function, which outputs 1 if the parameter is true and 0 otherwise, Represents the segmentation confidence output by the model, Is the pixel in the output mask, Represents the pixel The true label of and Are the dimensions of the output mask in the spatial dimension respectively, Represents the task stage The set of new classes learned by the model, Represents up to the task stage The total number of classes learned by the model;
[0038] Calculate and store the feature prototypes of each class based on the model at the end of each task stage training;
[0039] When a new task starts, calculate the feature prototypes of each new class based on a new model, match the old class with the highest similarity to the prototype, and use the old class with the highest similarity to assist in initializing the classifier for the new class.
[0040] Combined with the first aspect, in some implementations of the first aspect, the method further includes: The process of calculating and storing the feature prototypes of each category is as follows:
[0041] For each category, generate a binary mask for each image according to the annotations in the dataset, where the pixels belonging to this category are marked as 1 in the binary mask, and the pixels not belonging to it are marked as 0;
[0042] Based on the model at the end of each task stage training, extract features for the images of each category, multiply the extracted feature map by the binary mask of the corresponding image, and obtain the category-specific feature embedding of each image through global average pooling;
[0043] Calculate the mean of the category-specific feature embeddings of all images of each category to obtain the feature prototype of the category, and store it.
[0044] Combined with the first aspect, in some implementations of the first aspect, the method further includes: The process of using the old class with the highest similarity to assist in initializing the classifier for the new class is as follows:
[0045] At the start of each task stage, initialize the feature extractor of the new model with the feature extractor trained in the previous task;
[0046] For each new category, generate a binary mask for each image according to the annotations in the dataset, where the pixels belonging to this category are marked as 1 in the binary mask, and the pixels not belonging to this category are marked as 0;
[0047] Based on the initialized feature extractor of the new model, extract features for the images of each category, multiply the extracted feature map by the binary mask of the corresponding image, and obtain the category-specific feature embedding of each image through global average pooling;
[0048] Calculate the mean of the category-specific feature embeddings of all images of each new category to obtain the feature prototype of the new category, and calculate the similarity with the feature prototypes of all stored old categories;
[0049] For each new class, select the old class with the highest similarity to its feature prototype, and use the auxiliary classifier of this old class to initialize the classifier for the new class.
[0050] In a second aspect, in order to achieve the above object, a continuous semantic segmentation system based on balanced multi-granularity fusion feature distillation includes:
[0051] A feature processing module, configured to obtain an image, input the image into a feature extraction network of a pre-established new and old model, and output multi-layer features of the image; input the multi-layer features of the image into a pre-established parallel self-attention module, and output preliminary fused multi-granularity features.
[0052] A loss calculation module, configured to input the preliminary fused multi-granularity features into a pre-established moment-based channel attention module, and output balanced multi-granularity fused features; calculate the consistency constraint of the balanced multi-granularity fused features based on knowledge distillation, and calculate the total training loss based on the consistency constraint of the balanced multi-granularity fused features.
[0053] A model training module, configured to train the new model based on the total training loss to obtain a trained new model, and provide an optimal initialization decision boundary for the new task based on an auxiliary classifier through the trained new model, where the new model is a task-phase model.
[0054] In another aspect of the present invention, in order to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The computer program stored in the memory can run on the processor. When the processor loads and executes the computer program, the continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation as described above is adopted.
[0055] Advantages of the present invention:
[0056] The present invention can well solve the problems commonly existing in the existing continuous semantic segmentation methods, that is, only the deep semantic features of a single granularity are constrained, ignoring the migration of information crucial to the semantic segmentation task contained in other granularity features, and the randomly initialized classifier will produce serious prediction biases. And it can reduce the material and time costs required for retraining the semantic segmentation model when facing continuously incoming data streams. Description of the drawings
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0058] Figure 1 is a schematic flowchart of the method of the present invention;
[0059] Figure 2 is a schematic structural diagram of the continuous semantic segmentation method of the present invention;
[0060] Figure 3It is a schematic diagram of the continuous semantic segmentation verification case of the present invention;
[0061] Figure 4 It is a schematic diagram of the system structure of the present invention. Detailed implementation manners
[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0063] Embodiment 1:
[0064] As Figure 1 shown, a continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation, the method includes the following steps:
[0065] S101: Obtain an image, input the image into the feature extraction network of the pre-established new and old models, and output multi-layer features of the image; input the multi-layer features of the image into the pre-established parallel self-attention module, and output the initially fused multi-granularity features;
[0066] In the task stage , given a training image and the old model trained in the previous task stage, the new model is initialized from the old model, and the old model remains frozen during the training process;
[0067] During the training process, select feature maps from specific layers from the intermediate features of each layer of the feature extraction networks of the new and old models (such as: ResNet50 or ResNet101), and represent them as and respectively, where , represents that among all the selected intermediate features, and come from the network layers ranked th from shallow to deep, , and represent the dimensions of the feature Figure 3 respectively.
[0068] The shallow feature maps contain fine-grained local detail information such as color, geometric shape, etc., while the deep feature maps contain high-level semantic information. These information are all important for pixel-level classification tasks such as semantic segmentation. Extract features of various granularities for subsequent balancing and fusion of multi-granularity features, and retain them by proposing knowledge distillation.
[0069] The output process of initially fusing multi-granularity features:
[0070] The parallel self-attention module consists of several parallel branches, and each branch follows the self-attention network structure in the vision transformer;
[0071] Convert the intermediate features of different layers of the model into the same size using bilinear interpolation;
[0072] For all the extracted intermediate features, perform bilinear interpolation uniformly in the spatial dimension to change their spatial size to so that the features of each layer can be concatenated in the channel dimension after outputting from the parallel self-attention module;
[0073] Input the features of different layers after bilinear interpolation into different branches of the parallel self-attention module, and each branch corresponds to one layer of features;
[0074] Concatenate the outputs of each branch of the parallel self-attention module in the channel dimension to form initially fused multi-granularity features.
[0075] Taking the initially fused multi-granularity features of the new model as an example, its size is . Similarly, the old model will generate initially fused multi-granularity features in the same way .
[0076] S102: Input the initially fused multi-granularity features into a pre-established moment-based channel attention module, and output balanced multi-granularity fusion features; calculate the consistency constraint of the balanced multi-granularity fusion features based on knowledge distillation, and calculate the total training loss based on the consistency constraint of the balanced multi-granularity fusion features;
[0077] The output process of balanced multi-granularity fusion features:
[0078] Based on the initially fused multi-granularity features, calculate its moment statistics in the spatial dimension, including the first moment: mean, the second moment: variance, and the third moment: skewness, and obtain three feature statistic vectors corresponding to the three statistics;
[0079] Taking the initially fused multi-granularity features as an example, perform moment operations on each of its channels. For the th channel, the th moment :
[0080]
[0081] Among them, . and represent the spatial dimension size of the preliminarily fused multi-granularity features , and represents the spatial position coordinates of the pixel. Corresponding to order moments, the preliminarily fused multi-granularity features yield a moment vector . Corresponding to the first-order moment, second-order moment, and third-order moment, three moment vectors are obtained in total.
[0082] Input the three feature statistical vectors into a multi-layer perceptron network, and output three channel attention vectors respectively;
[0083] Add the three channel attention vectors and multiply them by the preliminarily fused multi-granularity features to obtain the balanced multi-granularity fusion features and .
[0084] Calculation process of the total training loss:
[0085] Based on the knowledge distillation technology, calculate the distance between the balanced multi-granularity fusion features of the new and old models respectively as the training loss, and calculate the balanced multi-granularity fusion feature consistency constraint ;
[0086] Calculate the standard classification loss , and the confidence distillation loss for the final output of the model ;
[0087] Calculate the total loss of model training :
[0088]
[0089] Among them, and are the balance coefficients used to balance each loss.
[0090] The calculation process of the balanced multi-granularity fusion feature consistency constraint is as follows:
[0091]
[0092] In the formula, and represent the spatial dimension size of the balanced multi-granularity fusion features, represents the balanced multi-granularity fusion features The value at the row and column in the spatial dimension, whose subscript indicates that this balanced multi-granularity fusion feature is extracted by the model at task stage .
[0093] The said standard classification loss is:
[0094]
[0095]
[0096] Wherein, represents the class label, and represent the size of the output mask, represents the task stage the model's confidence in the class corresponding to the pixel in the picture output by denotes taking the logarithm, represents the indicator function, which outputs 1 if its argument is true and 0 otherwise, represents the segmentation confidence output by the model, is the pixel in the output mask, represents the pixel 's true label, and are respectively the sizes of the output mask in the spatial dimension, represents the task stage the set of new classes learned by the model, represents the number of all classes learned by the model up to task stage ;
[0097] The said confidence distillation loss is:
[0098] .
[0099] Wherein, represents all the old classes learned by the model.
[0100] S103: Train the new model based on the total training loss to obtain the trained new model. Use the trained new model to provide the optimal initialization decision boundary for the new task based on the auxiliary classifier. Among them, the new model is the model at task stage .
[0101] The process by which the trained new model provides an optimal initial decision boundary for the new task based on the auxiliary classifier is as follows:
[0102] Based on the classification loss variant Train an auxiliary classifier for each class:
[0103]
[0104] where, represents the class label, and represent the size of the output mask, represents the task stage The confidence of the model for the pixel corresponding to the class output in the picture, represents taking the logarithm, represents the indicator function, which outputs 1 if the parameter is true and 0 otherwise, represents the segmentation confidence output by the model, is the pixel in the output mask, represents the pixel true label. and are the sizes of the output mask in the spatial dimension respectively, represents the task stage The set of new classes learned by the model, represents up to the task stage The total number of classes learned by the model;
[0105] Calculate and store the feature prototype of each class based on the model at the end of each task stage;
[0106] At the start of the new task, calculate the feature prototype of each new class based on the new model, match the prototype with the old class that has the highest similarity to the new class, and initialize the new class classifier with the auxiliary classifier of the old class with the highest similarity.
[0107] The process of calculating and storing the feature prototype of each class is as follows:
[0108] For each class, generate a binary mask for each picture according to the annotation in the dataset, where the pixels belonging to this class are marked as 1 in the binary mask and the pixels not belonging to it are marked as 0;
[0109] Based on the model at the end of each task stage, extract features for the pictures of each class, multiply the extracted feature map by the binary mask of the corresponding picture, and obtain the class-specific feature embedding of each picture through global average pooling;
[0110] Calculate the mean of the class-specific feature embeddings of all images for each class to obtain the feature prototype of the class, and store it.
[0111] The process of initializing the new class classifier with the old class auxiliary classifier with the highest similarity is as follows:
[0112] At the beginning of each task stage, initialize the feature extractor of the new model with the feature extractor trained in the previous task.
[0113] For each new class, generate a binary mask for each image according to the annotations in the dataset, where the pixels belonging to the class are marked as 1 in the binary mask, and the pixels not belonging to the class are marked as 0 in the binary mask.
[0114] Based on the initialized new model feature extractor, extract features for the images of each class, multiply the extracted feature maps by the binary masks of the corresponding images, and obtain the class-specific feature embeddings of each image through global average pooling.
[0115] Calculate the mean of the class-specific feature embeddings of all images of each new class to obtain the feature prototype of the new class, and calculate the similarity with the feature prototypes of all stored old classes.
[0116] For each new class, select the old class with the highest similarity to its feature prototype, and initialize the new class classifier with the auxiliary classifier of this old class.
[0117] Figure 2 Shows the block diagram of the present continuous semantic segmentation method and system. At task stage t, given the new model, the old model at task stage t - 1, and the training images. The training images are respectively input into the feature extractors of the new and old models to obtain a set of multi-granularity features. Then, the multi-granularity features of the new and old models are respectively input into the corresponding parallel self-attention modules according to the layers they are in. Through multiple groups of self-attention layers, task-related regions are mined for the features of different granularities of the new and old models, and then they are concatenated in the channel dimension to obtain the preliminary fusion multi-granularity features of the new and old models respectively. Then, the preliminary fusion multi-granularity features of the new and old models are respectively input into the moment-based channel attention module. Calculate the moment statistics of the preliminary fusion multi-granularity features in the spatial dimension to obtain the moment vector of the preliminary fusion multi-granularity features, and then obtain the corresponding moment-based channel attention vector through an MLP network. Sum the vectors and multiply them by the preliminary fusion multi-granularity features to obtain the balanced multi-granularity fusion features of the new and old models respectively. Finally, input the balanced multi-granularity fusion features of the new and old models into the classifiers of the new and old models, output the confidence corresponding to each class, and calculate the classification loss using binary cross-entropy loss according to the confidence and the true labels of the training data Calculate the feature distillation loss according to the balanced multi-granularity fusion features of the new and old models Calculate the confidence distillation loss based on the output confidence of the old model and the output confidence of the new model. Optimize the new model using the gradient descent algorithm based on the above three losses. Based on a variant of the classification loss Use the gradient descent algorithm to optimize the auxiliary classifier and initialize the new class classifier with the auxiliary classifier at the beginning of the next task phase.
[0118] Figure 3 This is a continuous semantic segmentation verification case for the method and system. The figure shows the segmentation effects of the method and system in various common scenarios and the comparison of the effects with other methods in the same field. Among them, the Image column represents the input image, the GT column represents the ideal segmentation result of the input image, and D 2 UP column represents the segmentation effect of the method and system. The PLOP and EDL columns represent the segmentation effects of two other methods and systems in this field. In the figure, "person", "bottle" and "chair" are the old categories learned previously, and "sofa", "potted plant" and "train" are the new categories. Figure 3 The results show that the effects of the method and system are better than other methods for both old and new categories.
[0119] Example 2: Second aspect, as Figure 4 shown, to achieve the above purpose, a continuous semantic segmentation system based on balanced multi-granularity fusion feature distillation includes:
[0120] A feature processing module 11, configured to obtain a picture, input the picture into a feature extraction network of a pre-established old and new model, and output multi-layer features of the picture; input the multi-layer features of the picture into a pre-established parallel self-attention module, and output preliminary fusion multi-granularity features;
[0121] A loss calculation module 12, configured to input the preliminary fusion multi-granularity features into a pre-established moment-based channel attention module, and output balanced multi-granularity fusion features; calculate the consistency constraint of the balanced multi-granularity fusion features based on knowledge distillation, and calculate the total training loss based on the consistency constraint of the balanced multi-granularity fusion features;
[0122] A model training module 13, configured to train the new model based on the total training loss to obtain a trained new model, and provide an optimal initialization decision boundary for the new task through the trained new model based on the auxiliary classifier, where the new model is a task phase model.
[0123] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.
[0124] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.
[0125] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0126] The foregoing has shown and described the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements fall within the scope of the present disclosure claimed.
Claims
1. A continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation, characterized by: The method comprises the following steps: Get the image, input the image into the pre-established feature extraction network of the new and old models, and output the multi-layer features of the image; The multi-layer features of the image are input into the pre-established parallel self-attention module, and the output is the preliminary fusion of multi-granular features; The output process of the preliminary fusion of multi-granularity features: The parallel self-attention module consists of several parallel branches, each of which follows the self-attention network structure in the visual transformer; Convert the intermediate features of different layers of the model to the same size using bilinear interpolation; The features of different layers after bilinear interpolation are input into different branches of the parallel self-attention module, and each branch corresponds to a layer of features; The outputs of each branch of the parallel self-attention module are spliced in the channel dimension to form a preliminary fusion of multi-granular features; The preliminary fused multi-granularity features are input into the pre-established moment-based channel attention module, and the balanced multi-granularity fused features are output; Based on knowledge distillation, the consistency constraints of balanced multi-granularity fusion features are calculated, and the total training loss is calculated based on the consistency constraints of balanced multi-granularity fusion features; The new model is trained based on the total training loss to obtain a trained new model, and the trained new model provides an optimal initialization decision boundary for the new task based on the auxiliary classifier, wherein the new model is a task stage t model.
2. The continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation according to claim 1 is characterized in that: The output process of the balanced multi-granularity fusion feature: Based on the preliminary fusion of multi-granular features, the moment statistics are calculated in the spatial dimension, including the first-order moment: mean, the second-order moment: variance and the third-order moment: skewness. Three feature statistical vectors are obtained corresponding to the three statistics. Input the three feature statistical vectors into the multi-layer perceptron network and output three channel attention vectors respectively; The three channel attention vectors are added together and multiplied with the preliminary fused multi-granularity features to obtain the balanced multi-granularity fusion features.
3. The continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation according to claim 1 is characterized in that: The calculation process of the total training loss is: Based on the knowledge distillation technology, the distance between the balanced multi-granularity fusion features of the new model and the old model is calculated as the training loss, and the balanced multi-granularity fusion feature consistency constraint L is calculated. mgkd ; Calculate the standard classification loss L mbce , and the confidence distillation loss L for the final output of the model at task stage t kd ; Calculate the total loss L of the model training at task stage t total : L total =L mbce +λ1L mgkd +λ2L kd Among them, λ1 and λ2 are balancing coefficients used to balance various losses.
4. The continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation according to claim 3 is characterized in that: The balanced multi-granularity fusion feature consistency constraint L mgkd The calculation process is as follows: In the formula, and Represents the spatial dimension of balanced multi-granularity fusion features, Represents balanced multi-granularity fusion features The value of the jth row and kth column in the spatial dimension, The subscript t indicates that the balanced multi-granularity fusion feature is extracted by the model of task stage t; The standard classification loss L mbce for: Where c represents the category label, H and W represent the size of the output mask, and p t (i,c) represents the confidence of the model output of the task stage t for the pixel i in the image corresponding to category c, log(·) represents the logarithm, Represents an indicator function that outputs 1 if its argument is true, otherwise it outputs 0. represents the segmentation confidence of the model output, i is the pixel in the output mask, y(i) represents the true label of pixel i, H and W are the sizes of the output mask in the spatial dimension, C t represents the set of new categories learned by the model in task stage t, |C 1:t | represents the number of all categories learned by the model at task stage t; The confidence distillation loss L kd for: Among them, C 1:t-1 Represents all the old categories that the model has learned.
5. The continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation according to claim 1 is characterized in that: The process of providing the optimal initialization decision boundary for the new task based on the auxiliary classifier by the trained new model is as follows: Based on the classification loss variant L neg Train an auxiliary classifier for each class: Where c represents the category label, H and W represent the size of the output mask, and p t (i,c) represents the confidence of the model output of the task stage t for the pixel i in the image corresponding to category c, log(·) represents the logarithm, Represents an indicator function, which outputs 1 if the argument is true, and 0 otherwise. represents the segmentation confidence of the model output, i is the pixel in the output mask, y(i) represents the true label of pixel i; H and W are the sizes of the output mask in the spatial dimension, C t represents the set of new categories learned by the model in task stage t, |C 1:t | represents the number of all categories learned by the model at task stage t; Calculate and store feature prototypes for each category based on the model at the completion of each task stage training; At the beginning of a new task, the feature prototype of each new class is calculated based on the new model, and the old class with the highest similarity between the prototype and the new class is matched, and the new class classifier is initialized with the auxiliary classifier of the old class with the highest similarity.
6. The continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation according to claim 5 is characterized in that: The process of calculating and storing the feature prototype of each category is as follows: For each category, a binary mask is generated for each image based on the annotations in the dataset, where pixels belonging to this category are marked as 1 in the binary mask, and pixels that do not belong to this category are marked as 0 in the binary mask; Based on the model trained at each task stage, features are extracted for each category of images, the extracted feature map is multiplied by the binary mask of the corresponding image, and the category-specific feature embedding of each image is obtained through global average pooling; Calculate the mean of the category-specific feature embeddings of all images in each class, obtain the feature prototype of the category, and store it.
7. The method for continuous semantic segmentation based on balanced multi-granularity fusion feature distillation according to claim 6, characterized in that: The process of initializing the new class classifier with the old class auxiliary classifier with the highest similarity is as follows: At the beginning of each task stage, the feature extractor of the new model is initialized with the feature extractor trained in the previous task; For each new category, a binary mask is generated for each image based on the annotations in the dataset, where pixels belonging to the category are marked as 1 in the binary mask, and pixels not belonging to the category are marked as 0 in the binary mask; Based on the initialized new model feature extractor, features are extracted for each category of images. The extracted feature map is multiplied with the binary mask of the corresponding image, and the category-specific feature embedding of each image is obtained through global average pooling. Calculate the mean of the category-specific feature embeddings of all images of each new category to obtain the feature prototype of the new category, and calculate the similarity with the stored feature prototypes of all old categories; For each new class, select the old class with the highest similarity to its feature prototype, and use the auxiliary classifier of this old class to initialize the new class classifier.
8. A continuous semantic segmentation system based on balanced multi-granularity fusion feature distillation, which adopts the continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation according to any one of claims 1 to 7, characterized in that: include: The feature processing module is used to obtain images, input the images into the pre-established feature extraction network of the new and old models, and output the multi-layer features of the images; The multi-layer features of the image are input into the pre-established parallel self-attention module, and the output is the preliminary fusion of multi-granular features; A loss calculation module is used to input the preliminary fused multi-granularity features into a pre-established moment-based channel attention module, and output a balanced multi-granularity fused feature; Based on knowledge distillation, the consistency constraints of balanced multi-granularity fusion features are calculated, and the total training loss is calculated based on the consistency constraints of balanced multi-granularity fusion features; The model training module is used to train the new model based on the total training loss to obtain the trained new model, and provide the optimal initialization decision boundary for the new task based on the auxiliary classifier through the trained new model, wherein the new model is the task stage t model.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the continuous semantic segmentation method based on balanced multi-granularity fusion feature distillation described in any one of claims 1 to 7 is adopted.
Citation Information
Patent Citations
Continuous semantic segmentation method and system based on spatial vision and statistical relationship distillation
CN118379502A