Incremental Semantic Segmentation Method and Autonomous Driving Image Semantic Segmentation Method
The incremental semantic segmentation model is constructed through a global attention mechanism and an adaptive multi-scale distillation mechanism. Combining embedding loss, distillation loss and background loss, the negative impact of old knowledge distillation on the new model in incremental semantic segmentation is solved, and higher reliability and accuracy are achieved in autonomous driving images.
Patent Information
- Application Number
- CN202510337376.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing incremental semantic segmentation scheme ignores the context information between each task, resulting in a lack of integrity in the learning process of the model, and the direct distillation of the old model's knowledge to the new model may have a negative impact on the new model, affecting the reliability and accuracy of the model, especially in the semantic segmentation of autonomous driving images.
The incremental semantic segmentation model is constructed using the global attention mechanism, the Transformer mechanism and the adaptive multi-scale distillation mechanism, and the training is combined with embedding loss, distillation loss, background loss and cross loss. Through global marking and background compensation strategies, the model is dynamically adjusted to adapt to new tasks, retaining old knowledge and reducing background offsets.
It improves the reliability and accuracy of the model, can maintain the recognition accuracy of old categories in the environment facing dynamic changes, and effectively learn new categories, reduce catastrophic forgetting and background offsets, and improves the accuracy and stability of semantic segmentation of autonomous driving images.
Smart Images

Figure CN119863796B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to an incremental semantic segmentation method and an autonomous driving image semantic segmentation method. Background Art
[0002] The incremental semantic segmentation scheme enables the model to gradually learn the information of new categories without forgetting the categories already learned, so as to adapt to the dynamically changing actual environment. By combining the advantages of incremental learning and semantic segmentation, this scheme solves the limitations of traditional methods in open scenarios and provides a more flexible and efficient solution for practical applications. The incremental semantic segmentation scheme can effectively cope with the challenges of category dynamic changes and open scenarios, and greatly improve the robustness and flexibility of intelligent systems.
[0003] Existing incremental semantic segmentation schemes address key challenges in incremental learning (such as the catastrophic forgetting problem, etc.) through pseudo-label strategies and knowledge distillation, and have achieved certain results. However, existing incremental semantic segmentation schemes ignore the context information between each task, resulting in a lack of integrity in the learning process. Therefore, the reliability and accuracy of the final model are relatively low. In addition, existing incremental semantic segmentation schemes directly distill all the knowledge of the old model into the new model, but some knowledge of the old model may have a non-positive effect on the new model; this situation will also reduce the reliability and accuracy of the final model.
[0004] Similarly, during the autonomous driving process, an incremental semantic segmentation scheme needs to be adopted for the semantic segmentation of autonomous driving images, and the segmentation results will directly provide data support for the autonomous driving process. The low reliability and low accuracy of existing incremental semantic segmentation schemes will also greatly affect the semantic segmentation results of autonomous driving images, and thus affect the reliability of autonomous driving. Summary of the Invention
[0005] One of the purposes of the present invention is to provide an incremental semantic segmentation method with high reliability and good accuracy.
[0006] Another purpose of the present invention is to provide an autonomous driving image semantic segmentation method including the above-mentioned incremental semantic segmentation method.
[0007] The incremental semantic segmentation method provided by the present invention includes the following steps:
[0008] S1. Obtain a target image dataset;
[0009] S2. Preprocess the target image dataset obtained in step S1 to construct a training dataset;
[0010] S3. Based on the global attention mechanism, the Transformer mechanism, and the adaptive multi-scale distillation mechanism, construct a primary incremental semantic segmentation model including an encoder, a decoder, and a distiller;
[0011] Among them, the encoder is used to encode the input image data to extract the feature information of the input image; by replacing the attention module in the Transformer mechanism with a global attention mechanism module, the decoder decodes the feature information extracted by the encoder to obtain global information; the distiller is used to perform feature distillation and fusion operations on the output of the decoder to finally achieve semantic segmentation;
[0012] S4. Use the training dataset constructed in step S2, combined with the embedding loss, the distillation loss, the background loss, and the cross loss, to train the primary incremental semantic segmentation model constructed in step S3 to obtain an incremental semantic segmentation model;
[0013] S5. Use the incremental semantic segmentation model obtained in step S4 to perform semantic segmentation on actual images.
[0014] The said step S3 includes the following steps:
[0015] Construct an encoder based on the Transformer mechanism to extract the feature information of the input image;
[0016] In the decoder, by using the background class label, the learned new label, and the global label as features and inputting them into the Transformer decoder for decoding to obtain global information; among them, replace the attention module in the Transformer decoder with a global attention mechanism module;
[0017] Construct a distiller based on the adaptive multi-scale distillation mechanism to perform feature distillation and fusion operations on the output of the decoder to finally achieve semantic segmentation.
[0018] The said encoder specifically includes the following content:
[0019] The constructed encoder includes an embedding layer, a linear projection layer, and a Transformer encoder layer;
[0020] The embedding layer is used to divide the input image into several visual patches and input the visual patches into the linear projection layer;
[0021] The linear projection layer is used to perform a linear projection on the input visual patches to map the input visual patches into high-dimensional representation vectors and input the obtained data information into the Transformer encoder layer;
[0022] The Transformer encoder layer is used to extract the feature information and dependencies of the input image based on the input data information, and obtain the corresponding encoder result.
[0023] The described decoder specifically includes the following content:
[0024] The constructed decoder includes a feature construction layer and a Transformer decoder layer;
[0025] The feature construction layer is used to construct the feature information required by the Transformer decoder layer; the features constructed by the feature construction layer include the target class token to be learned and the global token to be learned , and the constructed , and the encoder result are input into the Transformer decoder layer together; among them, the target class token is represented as , where is the background class token, is the token of the new category learned during the current t-th round of training; the global token is used to converge and represent the global knowledge of the current task. Before the first training, the global token is initialized to a set value . During the current t-th round of training, the value of the global token is the global token value obtained in the (t - 1)-th round of training ;
[0026] The Transformer decoder layer is used to decode the encoder result according to and to obtain the global information; in the Transformer decoder layer, the attention module in the Transformer decoder is replaced with a global attention mechanism module to achieve the acquisition of global information;
[0027] The output of the decoder includes decoder features and background class token features, where the decoder features are the features obtained by processing the encoder result through the decoder, and the background class token features are the features obtained by processing through the decoder.
[0028] The described global attention mechanism module specifically includes the following content:
[0029] During the current t-th round of training, the encoder result is concatenated with the token of the new category learned during the current t-th round of training to obtain the first concatenated feature ; The Perform normalization to obtain the normalized first splicing feature , and use as the query value; Use and the value of the global label in the current t-th round of training to perform splicing and obtain the first splicing label , and use it as the key value and the value value;
[0030] The processing process of the global attention mechanism module is expressed as: In the formula is the output of the global attention mechanism module; is the weight matrix of the output layer; is the processing function of the multi-layer perceptron; is the processing function of two cascaded masked normalization layers; is the bias term of the output layer; is the intermediate feature, and , is the softmax function, d is the embedding dimension, and h is the number of attention heads.
[0031] The described distiller specifically includes the following content:
[0032] The distiller includes a reshaping layer, a similarity calculation layer, an upsampling layer, and a normalization layer;
[0033] During the current t-th round of training, reshape the decoder feature through the reshaping layer to obtain the reshaped feature ; Use and the background class label feature to calculate the cosine similarity through the similarity calculation layer; After upsampling the obtained cosine similarity through the upsampling layer, perform normalization layer processing through the normalization layer to obtain the final segmentation prediction result;
[0034] Among them, the normalization layer uses the softmax function for normalization processing.
[0035] The training described in step S4 includes the following steps:
[0036] During the current t-th round of training, calculate the embedding loss based on the cosine similarity between the current background class label feature and the background class label feature of the previous round;
[0037] During the current t-th round of training, calculate the distillation loss based on the cosine similarity between the feature segmentation results of the current reshaped feature at different scales and the feature segmentation results of the reshaped feature of the previous round at different scales;
[0038] At the current t-th round of training, based on the KL divergence between the current background output and the background output of the previous round, as well as the difference between the current boundary region and the boundary region of the previous round, the background loss is calculated;
[0039] At the current t-th round of training, based on the current segmentation prediction result, the current ground truth, and the segmentation prediction result of the previous round, the cross loss is calculated;
[0040] Finally, the obtained embedding loss, distillation loss, background loss, and cross loss are weighted and summed as the total loss function for the current t-th round of training;
[0041] The obtained total loss function is used for training.
[0042] The training described in step S4 specifically includes the following steps:
[0043] Embedding loss:
[0044] At the current t-th round of training, the embedding loss is calculated using the following formula : In the formula is the total number of old classes at the current t-th round of training; is the i-th background class label feature at the current time; is the i-th background class label feature of the previous round; is the cosine similarity calculation function;
[0045] Distillation loss:
[0046] The reshaped feature is segmented at the set scale s; then at each segmentation scale, average pooling operations are performed on the segmented feature in the height and width directions, expressed as: In the formula is the pooling embedding result in the height direction at scale s; H is the height of the reshaped feature ; is the value at height h, width w, and channel c; is the pooling embedding result in the width direction at scale s; W is the width of the reshaped feature ;
[0047] At each scale, the obtained pooling embedding result in the height direction and the pooling embedding result in the width direction are concatenated to obtain the pooling embedding result at the corresponding scale;
[0048] The pooling embedding results at each scale are concatenated to obtain a high-dimensional feature representation ;
[0049] At the current t-th round of training, the distillation loss is calculated using the following formula : In the formula is the reliability mapping result, and its value is the cosine similarity value between the current pseudo-label and the segmentation prediction result of the previous round; is the value of the current high-dimensional feature representation at height h and width w; is the value of the high-dimensional feature representation of the previous round at height h and width w; is the Euclidean distance;
[0050] Background loss:
[0051] At the current t-th round of training, the background loss is calculated using the following formula : In the formula is the first weight value; is the second weight value; is the background confidence loss, and , N is the total number of background pixels, bg is the background pixel, is the background confidence weight, is the output probability distribution of the model on pixel i at the current t-th round of training, is the output probability distribution of the model on pixel i at the (t - 1)-th round of training, is the KL divergence calculation amount, and , C is the total number of categories in the current task, is the maximum value of the prediction probability of the new model for category c on pixel i, is the maximum value of the prediction probability of the old model for category c on pixel i; is the boundary region loss, and , M is the total number of boundary region pixels, T is the boundary region pixel, is the reshaped feature of boundary region pixel ii at the current t-th round of training, is the reshaped feature of boundary region pixel ii at the (t - 1)-th round of training;
[0052] Cross loss:
[0053] At the current t-th round of training, the cross loss is calculated using the following formula : In the formula is the pseudo-label, and its value is the fusion value of the segmentation prediction result of the iii-th pixel with the classification category c at the (t - 1)-th round of training and the true value of the iii-th pixel at the current t-th round of training; is the segmentation prediction result that the iii-th pixel belongs to class c during the current t-th round of training;
[0054] The total loss function L is calculated using the following formula: In the formula is the embedding loss weight value; is the distillation loss weight value; is the background loss weight value; is the cross loss weight value;
[0055] Finally, the total loss function L is used for training.
[0056] The present invention also provides an autonomous driving image semantic segmentation method including the above incremental semantic segmentation method, which includes the following steps:
[0057] A. Obtain an autonomous driving image to be segmented;
[0058] B. Use the above incremental semantic segmentation method to perform semantic segmentation on the autonomous driving image obtained in step A;
[0059] C. Complete the semantic segmentation of the autonomous driving image according to the semantic segmentation result obtained in step B.
[0060] The incremental semantic segmentation method and the autonomous driving image semantic segmentation method provided by the present invention, through the construction of an incremental semantic segmentation model including a global attention mechanism and an adaptive multi-scale distillation mechanism, and the training of the model using a loss function including an embedding loss, a distillation loss, and a background loss, not only achieve the semantic segmentation of images, but also have higher reliability and better accuracy. Description of the Drawings
[0061] Figure 1 is a schematic flowchart of the incremental semantic segmentation method of the present invention.
[0062] Figure 2 is a schematic diagram of the comparison effect of an embodiment of the incremental semantic segmentation method of the present invention.
[0063] Figure 3 is a schematic flowchart of the autonomous driving image semantic segmentation method of the present invention. Detailed Embodiments
[0064] As Figure 1 shown is a schematic flowchart of the incremental semantic segmentation method of the present invention: The incremental semantic segmentation method disclosed by the present invention includes the following steps:
[0065] S1. Obtain a target image dataset; for example, if semantic segmentation of remote sensing images is to be performed, obtain an existing image dataset for semantic segmentation of remote sensing images; or if semantic segmentation of fundus images is to be performed, obtain an existing image dataset for semantic segmentation of fundus images.
[0066] S2. Preprocess the target image dataset obtained in step S1 to construct a training dataset.
[0067] S3. Based on the global attention mechanism, the Transformer mechanism, and the adaptive multi-scale distillation mechanism, construct an incremental semantic segmentation primary model including an encoder, a decoder, and a distiller;
[0068] Among them, the encoder is used to encode the input image data to extract the feature information of the input image; by replacing the attention module in the Transformer mechanism with a global attention mechanism module, the decoder decodes the feature information extracted by the encoder to obtain global information; the distiller is used to perform feature distillation and fusion operations on the output of the decoder to finally achieve semantic segmentation.
[0069] Specifically in implementation:
[0070] Construct an encoder based on the Transformer mechanism to extract the feature information of the input image;
[0071] In the decoder, by using the background class label, the learned new label, and the global label as features and inputting them into the Transformer decoder for decoding to obtain global information; among them, the attention module in the Transformer decoder is replaced with a global attention mechanism module;
[0072] Construct a distiller based on the adaptive multi-scale distillation mechanism to perform feature distillation and fusion operations on the output of the decoder to finally achieve semantic segmentation.
[0073] The constructed encoder includes an embedding layer, a linear projection layer, and a Transformer encoder layer;
[0074] The embedding layer is used to divide the input image into several visual patches and input the visual patches into the linear projection layer;
[0075] The linear projection layer is used to perform a linear projection on the input visual patches to map the input visual patches into high-dimensional representation vectors and input the obtained data information into the Transformer encoder layer;
[0076] The Transformer encoder layer is used to extract the feature information and dependencies of the input image according to the input data information to obtain the corresponding encoder result.
[0077] The constructed decoder includes a feature construction layer and a Transformer decoder layer;
[0078] The feature construction layer is used to construct the feature information required by the Transformer decoder layer; the features constructed by the feature construction layer include the target class token to be learned and the global token to be learned and inputs the constructed , and the encoder result into the Transformer decoder layer together; among them, the target class token is represented as , where is the background class token, is the token of the new class learned during the current t-th round of training; the global token is used to converge and represent the global knowledge of the current task. Before the first training, the global token is initialized to a set value . During the current t-th round of training, the value of the global token is the global token value obtained in the (t - 1)-th round of training ; the global token is used to converge and represent the global knowledge of the current task. When entering a new task, by updating the global token, it can incorporate the key information of the new task and be dynamically adjusted to automatically adapt to the requirements of the new task; at the same time, through interaction with local features, the global token can effectively improve the accuracy of semantic segmentation, especially showing significant advantages when dealing with complex scenes and fine-grained features;
[0079] The Transformer decoder layer is used to decode the encoder result according to and to obtain global information; in the Transformer decoder layer, the attention module in the Transformer decoder is replaced with a global attention mechanism module to achieve the acquisition of global information;
[0080] The output of the decoder includes decoder features and background class token features, where the decoder features are the features obtained by processing the encoder result through the decoder, and the background class token features are the features obtained by processing through the decoder.
[0081] The global attention mechanism module in the decoder specifically includes the following content:
[0082] The global attention mechanism module dynamically adjusts the degree of attention of the model to different regions when processing input data by introducing global information and context relationships; the core idea of this module is to continuously update and enhance the utilization of global knowledge over time and with the change of task requirements, ensuring that the model can obtain more comprehensive and accurate information in complex tasks;
[0083] At the current t-th round of training, the encoder result is concatenated with the marker of the new category learned at the current t-th round of training to obtain the first concatenated feature ; is normalized to obtain the normalized first concatenated feature , and is used as the query value; is concatenated with the value of the global marker of the current t-th round of training to obtain the first concatenated marker , and is used as the key value and the value value;
[0084] The processing process of the global attention mechanism module is expressed as: where is the output of the global attention mechanism module; is the weight matrix of the output layer; is the processing function of the multi-layer perceptron; is the processing function of two cascaded masked normalization layers; is the bias term of the output layer; is the intermediate feature, and , is the softmax function, d is the embedding dimension, h is the number of attention heads; the addition of the masked normalization layer and the multi-layer perceptron enhances the expressive ability and stability of the global attention block, and also improves the robustness and learning efficiency of the model.
[0085] The distiller includes a reshaping layer, a similarity calculation layer, an upsampling layer and a normalization layer;
[0086] At the current t-th round of training, the decoder feature is reshaped through the reshaping layer to obtain the reshaped feature ; is concatenated with the background class marker feature The cosine similarity is calculated through the similarity calculation layer; after the obtained cosine similarity is upsampled through the upsampling layer, it is then processed through the normalization layer to obtain the final segmentation prediction result; the distiller ensures that the model can focus on the most beneficial feature representations through the dynamic adjustment of features at different scales, effectively improving the expressiveness and adaptability of the model. At the same time, it can also effectively fuse features from different scales, further alleviating the problem of information loss in traditional incremental semantic segmentation models when dealing with multi-scale features;
[0087] Among them, the normalization layer uses the softmax function for normalization processing.
[0088] S4. Using the training dataset constructed in step S2, combined with the embedding loss, distillation loss, background loss, and cross loss, train the incremental semantic segmentation primary model constructed in step S3 to obtain an incremental semantic segmentation model;
[0089] During the current t-th round of training, based on the cosine similarity between the current background class marker features and the background class marker features of the previous round, the embedding loss is calculated;
[0090] During the current t-th round of training, based on the cosine similarity between the feature segmentation results of the current reshaped features at different scales and the feature segmentation results of the reshaped features of the previous round at different scales, the distillation loss is calculated;
[0091] During the current t-th round of training, based on the KL divergence between the current background output and the background output of the previous round, and the difference between the current boundary region and the boundary region of the previous round, the background loss is calculated;
[0092] During the current t-th round of training, based on the current segmentation prediction result, the current ground truth, and the segmentation prediction result of the previous round, the cross loss is calculated;
[0093] Finally, the obtained embedding loss, distillation loss, background loss, and cross loss are weighted and summed as the total loss function for the current t-th round of training;
[0094] Use the obtained total loss function for training.
[0095] Specifically in implementation, the training process includes the following steps:
[0096] Embedding loss:
[0097] During the current t-th round of training, the embedding loss is calculated using the following formula : In the formula is the total number of old classes during the current t-th round of training; is the current i-th background class marker feature; is the i-th background class marker feature of the previous round; is the cosine similarity calculation function;
[0098] Through the calculation of the embedding loss and the calculation of similarity, it is possible to limit the deviation of the model prediction result towards the old class, thereby reducing the impact of background deviation on the prediction result; at the same time, the embedding loss can not only retain the global knowledge in the previous task, effectively avoid the problem of forgetting the old class knowledge, but also adjust the relationship between class embeddings, reduce the impact of background deviation on the model, and enhance the robustness of the model when dealing with the knowledge transfer between new and old classes.
[0099] Distillation loss:
[0100] The redundant information in the old knowledge often interferes with the current model's learning of new classes, making the prediction results of new classes tend to the old class or the background class. Therefore, the distillation loss is used for correction;
[0101] The reshaped feature is segmented at the set scale s; then at each segmentation scale, average pooling operations are performed on the segmented features in the height direction and the width direction, expressed as: In the formula is the pooling embedding result in the height direction at scale s; H is the height of the reshaped feature ; is the value at height h, width w, and channel c; is the pooling embedding result in the width direction at scale s; W is the width of the reshaped feature ;
[0102] Through the above multi-scale segmentation and pooling operations, the model can adaptively extract reliable features at multiple scales. This method can capture the useful semantic knowledge in the old model more comprehensively and dynamically adjust the distillation process, so that the problems of catastrophic forgetting and background deviation in the incremental learning process are significantly alleviated, further improving the model's ability to retain old class knowledge and its adaptability to new class learning;
[0103] At each scale, the obtained pooling embedding results in the height direction and the width direction are concatenated to obtain the pooling embedding result at the corresponding scale;
[0104] The pooling embedding results at each scale are concatenated to obtain a high-dimensional feature representation ; Through multi-scale feature pooling, while capturing global context information, it can also retain fine-grained local detail information, forming a multi-scale feature expression that takes into account both the global and the local;
[0105] At the current t-th round of training, the distillation loss is calculated using the following formula : In the formula is the reliability mapping result, and its value is the cosine similarity value between the current pseudo-label and the segmentation prediction result of the previous round; is the value of the current high-dimensional feature representation at height h and width w; is the value of the high-dimensional feature representation of the previous round at height h and width w; is the Euclidean distance;
[0106] The calculation of the distillation loss focuses on key knowledge regions by dynamically adjusting the weights between different features, thereby enhancing the model's learning ability for important features; meanwhile, through reliability mapping, a confidence score is assigned to each feature block, enabling the model to adaptively adjust the learning weights according to the reliability of each feature block, thus more precisely focusing on high-confidence regions.
[0107] Background loss:
[0108] Another challenging problem in class-incremental semantic segmentation is the background shift problem. This is because the background pixels predicted by the old model may include new categories in the current step, resulting in confusion and inaccuracy in the segmentation results; this background shift phenomenon will interfere with the learning process of new categories, especially when dealing with complex or dynamic scenes where the boundary between the background and the target becomes more blurred; therefore, the method of the present invention uses the background loss calculation process to solve this problem;
[0109] At the current t-th round of training, the background loss is calculated using the following formula : In the formula is the first weight value; is the second weight value; is the background confidence loss (which enables dynamic adjustment of the consistency of the background output between the old and new models during training), and , N is the total number of background pixels, bg is the background pixel, is the background confidence weight (it increases when the confidence of the old model that pixel i belongs to the background is low and decreases when the confidence of the old model that pixel i belongs to the background is high ), is the output probability distribution of the model on pixel i at the current t-th round of training, is the output probability distribution of the model on pixel i at the (t - 1)-th round of training, is the KL divergence calculation amount, and , where C is the total number of categories in the current task, is the maximum value of the predicted probability of the new model for category c at pixel i, is the maximum value of the predicted probability of the old model for category c at pixel i; is the boundary region loss (aiming to enhance the learning of the boundary region between the background and the new category), and , where M is the total number of pixels in the boundary region, and T is the boundary region pixel, is the reshaped feature of the boundary region pixel ii at the current t-th round of training, is the reshaped feature of the boundary region pixel ii at the (t - 1)-th round of training; Calculate the feature difference between the new and old models in the boundary region, so as to quantify the performance difference between the new and old models in these key regions;
[0110] The background loss can strengthen the model's learning ability at the object boundary, reduce the confusion between the background and the object, and ensure that the model can effectively distinguish the subtle differences between the object and the background; at the same time, the background loss can effectively reduce the interference of the new category learning process on the boundary region, thereby significantly improving the segmentation accuracy of the model; in addition, it can also reduce the offset of the background region, ensure that the model's performance in the boundary region is more stable and accurate, contribute to optimizing the overall segmentation effect, and enhance the model's adaptability to complex scenes; the background loss can help the model identify the background pixels that do not belong to the old categories in incremental learning, so as to appropriately compensate and incorporate them into the new category learning. In this way, the background loss not only effectively alleviates the background offset problem, but also enhances the flexibility and accuracy of the model when dealing with new categories.
[0111] Cross loss:
[0112] At the current t-th round of training, the cross loss is calculated using the following formula : In the formula is the pseudo label, and its value is the fusion value of the segmentation prediction result of the iii-th pixel with the classification category c at the (t - 1)-th round of training and the true value of the iii-th pixel at the current t-th round of training; is the segmentation prediction result of the iii-th pixel with the classification category c at the current t-th round of training.
[0113] Total loss:
[0114] The total loss function L is calculated using the following formula: In the formula is the embedding loss weight value; is the distillation loss weight value; is the background loss weight value; is the cross-loss weight value;
[0115] Finally, the total loss function L is used for training.
[0116] S5. Use the incremental semantic segmentation model obtained in step S4 to perform semantic segmentation on actual images.
[0117] The following uses ablation experiments to illustrate the effectiveness of the incremental semantic segmentation method of the present invention.
[0118] Ablation experiments of each component:
[0119] To verify the effectiveness of the method of the present invention, first, through a series of ablation experiments, the contributions of each module are evaluated, specifically including the global attention mechanism module, the distiller, and the background loss , and the experiments are carried out on two overlapping settings of Pascal VOC 15-1 and ADE20k100-50. The experimental results are shown in Table 1:
[0120] Table 1 Schematic table of component ablation experiment results
[0121]
[0122] The ablation experiment results of each component are shown in Table 1. In row (a), the performance of the baseline model used is presented; in row (b), the global attention mechanism module is introduced into the baseline framework, and the performance of the initial classes has been improved under both settings, with a 0.4%p improvement compared to the baseline. This indicates that the global attention mechanism module of the present invention can effectively capture global information, thereby enhancing the model's ability to retain old-class knowledge and improving the recognition accuracy of old classes when new classes are introduced; in row (c), the distiller is introduced into the baseline framework, and the performance of the model has been improved. Moreover, compared with the initial classes, the performance improvement of the model for new classes is significantly greater. The performance has increased by 1.9%p in the 16 - 20 steps and 2.5%p in the 100 - 150 steps. This proves that the distiller of the present invention can effectively filter out the redundant information in the old model and only retain the knowledge beneficial to the current task, thereby reducing the interference of the model in the distillation process for the learning of new classes, ensuring that the learning of new classes is more focused, and avoiding the negative impact of old classes, thus improving the overall performance of the model; in row (d), the background loss is introduced into the baseline framework, and the overall performance of the model has also been improved to a certain extent. The performance has increased by 0.6%p in the 15 - 1 setting and 0.5%p in the 100 - 50 setting. This indicates that the background loss effectively helps the model better process background information, thereby improving the model's performance under different settings; finally, the results in row (h) show the effect after integrating all modules. Compared with the baseline model, significant improvements are shown under all settings; among them, in the new-class scenario of 15 - 1, the improvement amplitude is the largest, with a 4.3%p performance improvement (16 - 20). The experimental results not only verify the effectiveness of each module but also prove that they are highly compatible with each other and can produce significant synergistic effects when used in combination, thereby improving the overall performance of the model.
[0123] Ablation experiment of background loss:
[0124] An ablation experiment is conducted separately for the background loss. The experiment is carried out on the 15 - 1 overlap setting of Pascal VOC, and the experimental results are shown in Table 2:
[0125] Table 2 Schematic table of ablation experiment results of background loss
[0126]
[0127] Conduct quantitative analysis for different background compensation strategies to prove the effectiveness of the proposed background compensation strategy; specifically, the background compensation strategies are divided into four types. Without means not using background loss, and only adding the global attention mechanism module and the distiller to the baseline framework as a reference for this experiment; "background confidence loss" and "boundary region loss" respectively represent using only background confidence loss and boundary region loss as background compensation strategies, and "background confidence loss + boundary region loss" represents the background loss proposed by the present invention. According to the experimental results shown in Table 2, when only using background confidence loss and boundary region loss as background loss, the performance improvement of the model is relatively small; when using background confidence loss as background loss, the model performance is improved by 0.4%p, and when using boundary region loss as background loss, the model performance is improved by 0.7%p. After combining the two, the overall performance of the model has a greater improvement, with an overall improvement of 1.6%p. This is because simply weighting the background confidence or aligning features cannot effectively separate the complex relationship between the background and the target class, and it is difficult to eliminate the interference between the background information and the target class through a single strategy, resulting in limited performance improvement. Table 2 proves the effectiveness and superiority of the background loss proposed by the present invention. By combining the advantages of background confidence loss and boundary region loss, and through the collaborative work of multiple strategies, while reducing background interference, it can more accurately capture the features of the target class, thus achieving more significant performance improvement in various settings.
[0128] Performance comparison on Pascal VOC dataset:
[0129] The experimental results on Pascal VOC dataset under different incremental learning settings are shown in Tables 3 to 5:
[0130] Table 3 Schematic table of experimental results on Pascal VOC dataset 1
[0131]
[0132] Table 4 Schematic table of experimental results on Pascal VOC dataset 2
[0133]
[0134] Table 5 Schematic table of experimental results on Pascal VOC dataset 3
[0135]
[0136] Among them, the EWC scheme was proposed by Kirkpatrick J in the paper "Overcoming catastrophic forgetting in neural networks" in 2017; the LIT scheme was proposed by Michieli U in the paper "Incremental learning techniques for semantic segmentation" in 2019; the MiB scheme was proposed by Cermelli F in the paper "Modeling the background for incremental learning in semantic segmentation" in 2020; the SDR scheme was proposed by Michieli U in the paper "Continual semantic segmentation via repulsion-attraction of sparse and disentangled latent representations" in 2021; the PLOP scheme was proposed by Douillard A in the paper "Plop: Learning without forgetting for continual semantic segmentation" in 2021; the RECALL scheme was proposed by Maracani A in the paper "Recall: Replay-based continual learning in semantic segmentation" in 2021; the REMIND scheme was proposed by Phan M H in the paper "Class similarity weighted knowledge distillation for continual semantic segmentation" in 2022; the RCLL scheme was proposed by Zhang C B in the paper "Representation compensation networks for continual semantic segmentation" in 2022; the SPPA scheme was proposed by Lin Z in the paper "Continual semantic segmentation via structure preserving and projected feature alignment" in 2022;The RBC scheme was proposed by Zhao H in the paper "Rbc: Rectifying the biased context in continual semantic segmentation" in 2022; the INC scheme was proposed by Shang C in the paper "Incrementer: Transformer for class-incremental semantic segmentation with knowledge distillation focusing on old class" in 2023; the MBS scheme was proposed by Park G in the paper "Mitigating Background Shift in Class-Incremental Semantic Segmentation" in 2025; Joint means training all classes simultaneously;
[0137] In the experiment, the method of the present invention was compared with existing schemes under three different experimental settings of VOC-19-1, VOC-15-5 and VOC-15-1; it can be seen from the table that the method of the present invention is significantly better than the best method in the existing schemes in all incremental scenarios; especially in the multi-step scenario, the performance of the model has been significantly improved; in the multi-step scenario, the introduced global information better helps the model maintain the understanding of global knowledge in each incremental step; compared with the recent MBS method, the method of the present invention exceeds the best model in the previous existing schemes by 2.0%p (points) (disjoint) and 0.9%p (overlap) in the setting of multiple consecutive steps (15-1). In addition, in the Disjoint setting of 15-1, there is a large improvement in the effect compared with the best method in the existing schemes, with an increase of 5.3%p (16-20); in the experimental setting of 15-5, the mIoU of the method of the present invention is improved by 1.5%p (disjoint) and 0.6%p (overlap) compared with the best method in the existing schemes; in the experimental setting of 19-1, the overall performance of the method of the present invention is improved by 0.5%p (overlap) compared with the best method in the existing schemes, and it has approached the experimental results in the case of Joint (indicating that all classes are trained simultaneously).
[0138] Performance comparison on the ADE20k dataset:
[0139] The experimental results on the ADE20k dataset under different incremental learning settings are shown in Tables 6 and 7:
[0140] Table 6 Schematic table of experimental results on the ADE20k dataset 1
[0141]
[0142] Table 7 Schematic Table of Experimental Results of ADE20k Dataset 2
[0143]
[0144] As can be seen from Table 6 and Table 7, the method of the present invention was tested under four different experimental settings: ADE-100-50, ADE-50-50, ADE-100-10, and ADE-100-5; the results show that, compared with the best-performing method in the existing solutions, the average performance of the method of the present invention improved by 0.95%p in all experimental settings. In the 100-50 scenario, it improved by 0.8%p compared with the state-of-the-art method. Especially in the scenario of incremental class sets, the performance of the model was significantly improved, exceeding the best-performing model in the previous existing solutions by 2.0%p (101-150); in the 50-50 scenario, it improved by 1.3%p compared with the state-of-the-art method; from the quantitative results, as the number of task steps increased, the method of the present invention demonstrated the advantage of alleviating catastrophic forgetting and background shift problems when dealing with multi-step incremental learning; generally speaking, the experimental results of these two datasets verified the stability and plasticity of the method of the present invention in incremental semantic segmentation. Through testing on different datasets, the model of the present invention demonstrated strong generalization ability and adaptability, and was able to effectively cope with the challenges brought by class incremental semantic segmentation.
[0145] Performance Comparison in Multi-step Scenarios in Pascal VOC Dataset:
[0146] The experimental results in the multi-step scenarios 5-3 (6 steps) and 10-1 (11 steps) in the Pascal VOC dataset are shown in Table 8:
[0147] Table 8 Schematic Table of Experimental Results of Pascal VOC Dataset
[0148]
[0149] As can be seen from Table 8, compared with the best-performing method in the existing solutions, the overall performance of the method of the present invention improved by 1.62%p in the 10-1 setting, the performance improved by 3.24%p (11-20); in the 5-3 setting, the performance improved by 1.3%p, The performance on [object] has been improved by 1.4%p (6 - 20); From the data in the table, it can be seen that compared with the retention of old - class knowledge, the improvement of the method of the present invention in new - class learning is more significant. This improvement is mainly due to the introduction of global tags; Global tags can effectively transmit cross - step context information during the new - class learning process, significantly reducing the degree to which new tasks are affected by previous tasks and avoiding the problem of knowledge forgetting caused by information interference; At the same time, the adaptive multi - scale distillation technology and the background compensation strategy precisely reduce the interference of old - class knowledge and background noise on new - class learning through dynamic adjustment, effectively improving the accuracy and stability of new - class learning; This shows that the method of the present invention not only promotes the accurate learning of new - class knowledge but also improves the overall robustness of the model to a certain extent.
[0150] Visual comparison of experimental results:
[0151] To better understand the performance of the method of the present invention under different settings, Figure 2 In [reference], the prediction results of the baseline and the method of the present invention in the Pascal VOC 15 - 1 overlap setting are visualized; Through the qualitative visualization results, the performance differences of the model during new - class learning can be intuitively shown, clearly presenting the accuracy and stability of the model during the learning process and revealing its adaptability in dealing with different tasks.
[0152] According to Figure 2 From the qualitative results shown, it can be clearly seen that compared with the results of the baseline, the method of the present invention shows better results in the retention of old classes; Even after multiple learning steps, the present invention still maintains a high accuracy in the segmentation of old classes and is not interfered by new - class learning. Especially in the details, such as the edges and protrusions of objects, the segmentation results are still accurate; In contrast, after multiple steps, the segmentation accuracy of old classes of the baseline method drops significantly, and it is more likely to produce classification errors or blurred boundaries in object boundaries and complex regions, resulting in a significant reduction in the recognition effect of old classes; Moreover, compared with the baseline method, the method of the present invention obtains a more accurate effect in distinguishing background classes. Especially after learning multiple new classes, the details and edges of foreground objects are better retained; Traditional segmentation methods often cause confusion between background classes and foreground classes and even a large number of mis - segmentations when facing background offsets, while the present invention successfully alleviates this problem through the background compensation strategy, ensuring the recognition accuracy of foreground classes.
[0153] The above qualitative results show that the method of the present invention has obvious advantages in the segmentation of background classes compared with traditional methods. In terms of segmentation accuracy, detail presentation, or overall effect, the method of the present invention can provide more accurate and detailed segmentation results. Especially in the segmentation results of the final step, it can well handle the boundary between new classes and old classes, ensuring the balanced and accurate segmentation effects of both. Through comparative analysis, it can be seen that the present invention not only effectively avoids the problem of forgetting old classes but also retains the accuracy of old classes while learning new classes, fully demonstrating its stability and robustness in the long-term learning process. These qualitative results further prove the superiority of the present invention in the task of class-incremental semantic segmentation, and it can provide more accurate segmentation performance when facing multi-step tasks and complex backgrounds.
[0154] As Figure 3 shown in the schematic flowchart of the method of the semantic segmentation method for autonomous driving images of the present invention: The semantic segmentation method for autonomous driving images including the incremental semantic segmentation method disclosed by the present invention includes the following steps:
[0155] A. Obtain the autonomous driving image to be segmented;
[0156] B. Use the incremental semantic segmentation method to perform semantic segmentation on the autonomous driving image obtained in step A;
[0157] C. Complete the semantic segmentation of the autonomous driving image according to the semantic segmentation result obtained in step B. According to this semantic segmentation result, data support can be provided for the control process of the autonomous driving vehicle;
[0158] In addition, the incremental semantic segmentation method provided by the present invention is also applicable to application scenarios such as medical image processing, remote sensing image processing, robot vision, video surveillance, industrial inspection and automation, intelligent transportation systems, and environmental monitoring.
Claims
1. An incremental semantic segmentation method, characterized in that It includes the following steps: S1. Obtain the target image dataset; S2. Preprocess the target image dataset obtained in step S1 to construct a training dataset; S3. Based on the global attention mechanism, Transformer mechanism, and adaptive multi-scale distillation mechanism, construct an incremental semantic segmentation primary model including an encoder, a decoder, and a distiller; Among them, the encoder is used to encode the input image data to extract the feature information of the input image; by replacing the attention module in the Transformer mechanism with a global attention mechanism module, the decoder decodes the feature information extracted by the encoder to obtain global information; the distiller is used to perform feature distillation and fusion operations on the output of the decoder to finally achieve semantic segmentation; The global attention mechanism module specifically includes the following content: At the current t-th round of training, the encoder result is concatenated with the label of the new category learned at the current t-th round of training to obtain the first concatenated feature ; is normalized to obtain the normalized first concatenated feature , and is used as the query value; is concatenated with the value of the global label in the current t-th round of training to obtain the first concatenated label , and is used as the key value and the value value; The processing process of the global attention mechanism module is expressed as: where is the output of the global attention mechanism module; is the weight matrix of the output layer; is the processing function of the multi-layer perceptron; is the processing function of two cascaded masked normalization layers; is the bias term of the output layer; is the intermediate feature, and , is the softmax function, d is the embedding size, and h is the number of attention heads; S4. Use the training dataset constructed in step S2, combined with the embedding loss, distillation loss, background loss, and cross loss, to train the incremental semantic segmentation primary model constructed in step S3 to obtain an incremental semantic segmentation model; S5. Use the incremental semantic segmentation model obtained in step S4 to perform semantic segmentation on actual images.
2. The incremental semantic segmentation method according to claim 1, characterized in that Step S3 includes the following steps: Construct an encoder based on the Transformer mechanism to extract the feature information of the input image; In the decoder, by using the background class label, the learned new label, and the global label as features and inputting them into the Transformer decoder for decoding to obtain global information; among them, the attention module in the Transformer decoder is replaced with a global attention mechanism module; Construct a distiller based on the adaptive multi-scale distillation mechanism to perform feature distillation and fusion operations on the output of the decoder to finally achieve semantic segmentation.
3. The incremental semantic segmentation method according to claim 2, wherein The encoder specifically includes the following content: The constructed encoder includes an embedding layer, a linear projection layer, and a Transformer encoder layer; The embedding layer is used to divide the input image into several visual patches and input the visual patches into the linear projection layer; The linear projection layer is used to perform a linear projection on the input visual patches to map the input visual patches into high-dimensional representation vectors and input the obtained data information into the Transformer encoder layer; The Transformer encoder layer is used to extract the feature information and dependencies of the input image according to the input data information to obtain the corresponding encoder result.
4. The incremental semantic segmentation method according to claim 3, wherein The decoder specifically includes the following content: The constructed decoder includes a feature construction layer and a Transformer decoder layer; The feature construction layer is used to construct the feature information required by the Transformer decoder layer; The features constructed by the feature construction layer include the target class label to be learned and the global label to be learned , and the constructed , and the encoder result are input into the Transformer decoder layer together; among them, the target class label is represented as , where is the background class label, is the label of the new class learned during the current t-th round of training; the global label is used to aggregate and represent the global knowledge of the current task. Before the first training, the global label is initialized to a set value . During the current t-th round of training, the value of the global label is the global label value obtained in the (t - 1)-th round of training ; The Transformer decoder layer is used to decode the encoder results based on and to obtain global information. In the Transformer decoder layer, the attention module in the Transformer decoder is replaced with a global attention mechanism module to achieve the acquisition of global information. The output of the decoder includes decoder features and background class token features, where the decoder features are the features obtained by processing the encoder results through the decoder, and the background class token features are the features obtained by processing through the decoder.
5. The incremental semantic segmentation method according to claim 4, characterized in that The distiller specifically includes the following content: The distiller includes a reshaping layer, a similarity calculation layer, an upsampling layer, and a normalization layer; At the current t-th round of training, the decoder features are reshaped by the reshape layer to obtain reshaped features ; and and the background class token features are calculated to obtain the cosine similarity through the similarity calculation layer; After the obtained cosine similarity is upsampled by the upsampling layer and then processed by the normalization layer through the normalization layer, the final segmentation prediction result is obtained; Among them, the normalization layer uses the softmax function for normalization processing.
6. The incremental semantic segmentation method according to claim 5, characterized in that The training described in step S4 includes the following steps: During the current t-th round of training, based on the cosine similarity between the current background class label features and the background class label features of the previous round, an embedding loss is calculated; During the current t-th round of training, based on the cosine similarity between the feature segmentation results of the current reshaped features at different scales and the feature segmentation results of the reshaped features of the previous round at different scales, a distillation loss is calculated; During the current t-th round of training, based on the KL divergence between the current background output and the background output of the previous round, and the difference between the current boundary region and the boundary region of the previous round, a background loss is calculated; During the current t-th round of training, based on the current segmentation prediction result, the current ground truth, and the segmentation prediction result of the previous round, a cross loss is calculated; Finally, the obtained embedding loss, distillation loss, background loss, and cross loss are weighted and summed as the total loss function for the current t-th round of training; Training is performed using the obtained total loss function.
7. The incremental semantic segmentation method according to claim 6, wherein The training described in step S4 specifically includes the following steps: Embedding loss: At the current t-th round of training, the embedding loss is calculated using the following formula :[[]]END]] In the formula is the total number of old classes at the current t-th round of training; is the i-th background class marker feature at the current time; is the i-th background class marker feature in the previous round; is the cosine similarity calculation function; Distillation loss: Reshape the feature Segment at the set scale s; then, at each segmentation scale, perform average pooling operations on the segmented features in the height and width directions, expressed as: In the formula is the pooling embedding result in the height direction at scale s; H is the reshaped feature height; is the value at height h, width w, and channel c; is the pooling embedding result in the width direction at scale s; W is the reshaping feature width; At each scale, the obtained height-direction pooled embedding results and width-direction pooled embedding results are concatenated to obtain the pooled embedding results at the corresponding scale; Concatenate the pooling embedding results at each scale to obtain a high-dimensional feature representation ; At the current t-th round of training, the distillation loss is calculated using the following formula :[[]] In the formula is the reliability mapping result, and its value is the cosine similarity value between the current pseudo-label and the segmentation prediction result of the previous round; is the value of the current high-dimensional feature representation at height h and width w; is the value of the high-dimensional feature representation of the previous round at height h and width w; is the Euclidean distance; Background loss: At the current t-th round of training, the background loss is calculated using the following formula : In the formula is the first weight value; is the second weight value; is the background confidence loss, and , N is the total number of background pixels, bg is the background pixel, is the background confidence weight, is the output probability distribution of the model at the current t-th round of training on pixel i, is the output probability distribution of the model at the (t - 1)-th round of training on pixel i, is the KL divergence calculation amount, and , C is the total number of categories in the current task, is the maximum value of the predicted probability of the new model for category c on pixel i, is the maximum value of the predicted probability of the old model for category c on pixel i; is the boundary region loss, and , M is the total number of boundary region pixels, T is the boundary region pixel, is the reshaped feature of the boundary region pixel ii at the current t-th round of training, is the reshaped feature of the boundary region pixel ii at the (t - 1)-th round of training; Cross loss: At the current t-th round of training, the cross-loss is calculated using the following formula : In the formula is the pseudo-label, and its value is the fusion value of the segmentation prediction result of the i-th pixel with the classification category c in the (t - 1)-th round of training and the true value of the i-th pixel in the current t-th round of training; is the segmentation prediction result of the i-th pixel with the classification category c in the current t-th round of training; The total loss function L is calculated using the following formula: where is the embedding loss weight value; is the distillation loss weight value; is the background loss weight value; is the cross-entropy loss weight value; Finally, training is performed using the total loss function L.
8. An autonomous driving image semantic segmentation method comprising the incremental semantic segmentation method according to any one of claims 1 to 7, characterized in that It includes the following steps: A. Obtain an autonomous driving image to be segmented; B. Use the described incremental semantic segmentation method to perform semantic segmentation on the autonomous driving image obtained in step A; C. Complete the semantic segmentation of the autonomous driving image according to the semantic segmentation result obtained in step B.