A Continuous Semantic Segmentation Method Based on Internal and External Distillation

Through the combined method of internal and external distillation, feature distillation and multi-scale convolution attention are used to solve the problem of forgetting in the face of new data, and the semantic segmentation effect of continuous learning and enhanced adaptability is achieved.

CN117095172BActive Publication Date: 2025-07-04NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311159231.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-09
Publication Date
2025-07-04
Estimated Expiration
2043-09-09

AI Technical Summary

Technical Problem

The existing semantic segmentation method based on deep learning needs to be retrained from scratch when facing new data, lacks continuous learning ability, cannot effectively utilize the knowledge learned, and lacks flexibility in a dynamic environment.

Method used

Using an internal and external distillation method, the new and old models are characterized by combining internal distillation and external distillation, and the relationship between the features is captured by combining multi-scale convolutional attention, and finally the segmented image is obtained through decoder upsampling.

Benefits of technology

It realizes the ability to continuously learn new knowledge in real scenarios, effectively retain the learned knowledge and perform semantic segmentation, and improves the adaptability and practicality of the model in a dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095172B_ABST
    Figure CN117095172B_ABST
Patent Text Reader

Abstract

The present invention discloses a continuous semantic segmentation method based on internal and external distillation. First, internal distillation of the new and old models is performed using the statistical information of features. For external distillation of features at different scales, multi-scale convolutional attention is used to capture the relationships between features at different scales. Finally, the segmentation image is obtained through upsampling by the decoder. The method of the present invention better meets the actual needs, can retain the learned knowledge and learn new knowledge so as to achieve the purpose of continuously performing semantic segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a continuous semantic segmentation method. Background Art

[0002] Semantic segmentation is a fundamental problem in computer vision, and its goal is to assign labels to each pixel in an image. Existing deep learning-based methods usually require prior knowledge of all classes in the dataset and utilize all available data during training. However, such a setting will be disconnected from reality, and the segmentation model does not have the ability to continuously learn and acquire new knowledge in real-world scenarios. When new data appears, it needs to be retrained from scratch. The literature "Multi-Path Semantic Segmentation Based on Edge Optimization and Global Modeling, Computer Science, 2023, 50(S1), pp 431-437" discloses a multi-path semantic segmentation method based on edge optimization and global modeling. This method proposes a network for multi-path proximity misalignment fusion to perform semantic information blending between the tail of the high-resolution path and the head of the low-resolution path. At the same time, an adaptive edge feature module is used to obtain edge features and integrate them into the intermediate layer and the deep supervision layer of the network to enhance the expression ability of edge features and the segmentation effect of small objects. The method described in the literature is a convolutional neural network method using edge algorithms and attention mechanisms. This method does not have the ability of continuous learning. When facing a new segmentation task, it needs to be trained from scratch, cannot utilize the learned knowledge, and has poor real-time performance. In addition, in a dynamic environment, this method cannot continuously perform segmentation tasks and flexibly respond in a rapidly changing scenario, and its practicability is not high. Summary of the Invention

[0003] To overcome the deficiencies of the prior art, the present invention provides a continuous semantic segmentation method based on internal and external distillation. First, the statistical information of features is used for internal distillation of the old and new models. For external distillation of features at different scales, multi-scale convolutional attention is used to capture the relationship between features at different scales. Finally, the segmentation image is obtained through upsampling by the decoder. The method of the present invention better meets the actual needs, can retain the learned knowledge and learn new knowledge so as to achieve the purpose of continuous semantic segmentation.

[0004] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0005] Step 1: Input the image I of the current task T T into the encoder E of the old model T-1 to obtain intermediate features

[0006] The old model is a model trained through task T-1 and can segment all classes in tasks 0:T-1; the encoder E of the old model T-1It consists of 4 convolutions, and the intermediate features are the outputs of the 4 convolutions, as shown below:

[0007]

[0008] In the formula, is the i-th convolutional layer in the old model encoder E T-1 ; is the image I of the current task T T ;

[0009] Step 2: Input the image I of the current task T T into the new model encoder E T After passing through two convolutional layers, the intermediate features f1 and f2 are obtained respectively;

[0010] The new model is a model that can segment all classes in task 0: T, with the same structure as the old model and initialized with the parameters of the old model;

[0011] The image I of task T T is input into the new model encoder E T After passing through the first convolutional layer in the encoder, the first intermediate feature f1 is obtained, and the feature f1 passes through the second convolutional layer in the encoder to obtain the second intermediate feature f2. The process is as follows:

[0012]

[0013]

[0014] Step 3: Calculate the attention maps of adjacent features of the new and old models and and calculate the L2 distance between them;

[0015] Step 3-1: Introduce an attention mechanism between model layers and perform external feature distillation; for the features f l-1 and f l of the new model, the corresponding attention map is obtained through the following operations: First, adjust their dimensions through 1×1 convolutional operations to make them match and perform a concatenation operation. The process is as follows:

[0016] f l c = Concat(f l , Conv 1×1 (f l-1 )) (4)

[0017] In the formula, l represents the index in the layer, 1×1 represents the convolutional kernel size of 1, and Concat(·) represents the concatenation operation;

[0018] Step 3-2: Implement spatial-aware feature extraction through multi-scale convolutional attention (MSCA) using element-wise multiplication; MSCA is written as:

[0019]

[0020] In the formula, Conv DW (·) represents depth convolution, Scale i (·) performs multi-scale processing on the features, and the features after multi-scale processing pass through a 1×1 convolution to obtain the attention map;

[0021] Step 3-3: Similarly, the features of the old model and go through the same operations to obtain the attention map

[0022] Step 3-4: Calculate the L2 distance between the attention maps and of adjacent features of the new and old models as the external feature distillation loss L ext to prevent forgetting, which is expressed as follows:

[0023]

[0024] Step 4: Multiply the calculated attention map of the new model features with the feature f2 to obtain the feature

[0025] Integrate the calculated attention map of the new model features into the corresponding features. For f1 and f2, the attention maps calculated according to Step 3 are Multiply with the feature f2 to obtain the feature The process is as follows:

[0026]

[0027] where represents matrix multiplication;

[0028] Step 5: The feature enters the third convolution of the new model encoder to obtain f3, and the above process is repeated to obtain f4;

[0029] The feature enters the third convolution of the new model encoder to obtain f3. Calculate the feature attention map of f3 and f2 and multiply it with f3 to obtain After passing through the fourth convolution of the new model encoder, f4 is obtained;

[0030] Step 6: The intermediate feature f of the new modell with the corresponding intermediate features of the old model perform internal distillation and calculate the internal distillation loss L int ;

[0031] Step 6-1: Perform mixed pooling on the features;

[0032] Let X represent a certain feature in f l with a size of H×W×C; use the Φ function to extract the maximum, minimum, and average information of feature X, and concatenate the H×C-width mixed pooling slices and W×C-height mixed pooling slices of X:

[0033]

[0034] where Concat(.) represents concatenation in the channel dimension;

[0035] Ω(X) is an operation that finds the maximum and minimum and combines them according to the corresponding dimensions, and its calculation process is as follows:

[0036] Ω(X[:, w, :]) = Concat(max(X[:, w, :]), min(X[:, w, :])) (9)

[0037] Step 6-2: Calculate the width- and height-mixed pooling slices on multiple regions extracted at different scales to retain local information, where the scales are {1 / 2 s} s=0...S ; given scale s and feature X, the mixed pooling slice Ψ s (X) of feature X at this scale is expressed as follows:

[0038]

[0039] where this is a sub-region of feature X, and for this sub-region has a size of H / s×W / s;

[0040] Step 6-3: Concatenate the mixed feature slices Ψ s (X) of each scale s along the channel dimension to obtain the final mixed feature slice Ψ(X):

[0041] Ψ(X) = Concat(Ψ 1 (X),..., Ψ s (X)) (11)

[0042] Step 6-4: Calculate the mixed feature slices of multiple layers l ∈ {1,..., L} of the old model and the current model, and then minimize the L2 distance between the mixed feature slices calculated at multiple layers during training; the internal feature distillation loss Lint It is defined as:

[0043]

[0044] Step 7: The feature f4 extracted by the new model encoder is upsampled by the decoder to obtain the segmented image

[0045] After the feature f4 extracted by the new model encoder is upsampled, the final segmented image is obtained The process is as follows:

[0046]

[0047] In the formula, Upsample(·) represents the upsampling operation.

[0048] The beneficial effects of the present invention are as follows:

[0049] The present invention combines continuous learning with semantic segmentation, and at the same time uses internal and external distillation methods to solve the forgetting problem in continuous semantic segmentation. First, internal distillation of the new and old models is performed using the statistical information of the features, effectively avoiding interference at the same scale. In addition, for external distillation of features at different scales, multi-scale convolutional attention is used to capture the relationship between features at different scales and ensure their consistency in new and old tasks to retain the learned knowledge. The method of the present invention better meets the actual needs, can retain the learned knowledge and learn new knowledge so as to achieve the purpose of continuous semantic segmentation. Description of the Drawings

[0050] Figure 1 It is a flowchart of continuous semantic segmentation based on internal and external distillation of the present invention.

[0051] Figure 2 It is the image I of the previous task T of the present invention T .

[0052] Figure 3 It is the segmentation map output by the final model of the embodiment of the present invention Detailed Embodiments

[0053] The present invention will be further described below in conjunction with the drawings and embodiments.

[0054] Taking the image I of the given previous task T T as an example, as Figure 2 shown, the specific implementation manner will be described. As Figure 1 shown, the continuous semantic segmentation method based on internal and external distillation in this embodiment includes the following steps:

[0055] Step 1: The image I of the current task T TInput to the encoder E of the old model T-1 to obtain intermediate features

[0056] The old model is a model trained for task T-1 and can segment all classes in task 0: T-1. The encoder E of the old model T-1 consists of 4 convolutions, and the intermediate features are the outputs of the 4 convolutions, as follows:

[0057]

[0058] where is the i-th convolutional layer in the encoder E of the old model T-1 and is the image IT of the current task T.

[0059] Step 2: Input the image I of the current task T T into the encoder E of the new model T After passing through two convolutional layers, intermediate features f1 and f2 are obtained respectively.

[0060] The new model is a model that can segment all classes in task 0: T, has the same structure as the old model, and is initialized with the parameters of the old model. The image I of task T T is input into the encoder E of the new model T After passing through the first convolutional layer in the encoder, the first intermediate feature f1 is obtained, and the feature f1 passes through the second convolutional layer in the encoder to obtain the second intermediate feature f2. The process is as follows:

[0061]

[0062]

[0063] Step 3: Calculate the attention maps of adjacent features of the old and new models and and calculate the L2 distance between them.

[0064] To enhance the distillation operation by considering more context information, the present invention introduces an attention mechanism between model layers and performs external feature distillation. Taking the features f l-1 and f l of the new model as an example, the corresponding attention map is obtained through the following operations: First, adjust their dimensions through 1×1 convolution operations to make them match and perform concatenation operations. The process is as follows:

[0065] f l c = Concat(f l , Conv1×1 (f l-1 )) (4)

[0066] In the formula, l represents the index in the layer, 1×1 represents the convolution kernel size of 1, and Concat(·) represents the concatenation operation.

[0067] The present invention adopts a novel spatial attention method. Different from the traditional self-attention method, this method realizes spatial-aware feature extraction by using multi-scale convolutional attention (MSCA) of element-wise multiplication. Briefly, MSCA can be written as:

[0068]

[0069] In the formula, Conv DW (·) represents depth convolution, Scale i (·) performs multi-scale processing on the features. The features after multi-scale processing pass through a 1×1 convolution to obtain the attention map. Similarly, the features of the old model and can obtain the attention map through the same operation Calculate the L2 distance between the attention maps of adjacent features of the new and old models and as the external feature distillation loss L ext to prevent forgetting, which is expressed as follows:

[0070]

[0071] Step Four: Multiply the calculated attention map of the new model features with the feature f2 to obtain the feature

[0072] Integrate the calculated attention map of the new model features into the corresponding features. Taking f1 and f2 as examples, according to Step Three, the attention map can be calculated as Multiply with the feature f2 to obtain the feature The process is as follows:

[0073]

[0074] Among them, represents matrix multiplication.

[0075] Step Five: The feature enters the third convolution of the new model encoder to obtain f3, and the above process is repeated to obtain f4.

[0076] The feature enters the third convolution of the new model encoder to obtain f3. The attention map of the features of f3 and f2 is calculated and multiply it by f3 to obtain After passing through the fourth convolution of the new model encoder, f4 is obtained.

[0077] Step 6: The intermediate feature f of the new model l and the corresponding intermediate feature of the old model perform internal distillation and calculate the internal distillation loss L int .

[0078] The new model is trained using the data of task T to learn new knowledge. However, training the model only with the data containing the new task categories will cause the model to forget the learned old knowledge. The present invention proposes an internal feature distillation module based on multi-information to perform internal distillation on the features of the same level of the new and old models.

[0079] Sub-step 1: Perform mixed pooling on the features. Let X represent a certain feature in f l , whose size is H×W×C. Use the Φ function to extract the maximum, minimum, and average information of the feature X, and concatenate the H×C width mixed pooling slice and the W×C height mixed pooling slice of X:

[0080]

[0081] In the formula, Concat(·) represents concatenation in the channel dimension, and Ω(X) is an operation of finding the maximum and minimum values and combining them according to the corresponding dimensions. Specifically, taking Ω(X[:, w, :]) as an example, its calculation process is as follows:

[0082] Ω(X[:, w, :]) = Concat(max(X[:, w, :]), min(X[:, w, :])) (9)

[0083] Sub-step 2: Calculate the width and height mixed pooling slices on multiple regions extracted at different scales to retain local information, where the scales are {1 / 2 s} s=0...S . Given the scale s and the feature X, the mixed pooling slice Ψ s (X) of the feature X at this scale is expressed as follows:

[0084]

[0085] In the formula, This is a sub-region of the feature X. For this sub-region, the size is H / s×W / s.

[0086] Sub-step 3: Concatenate the mixed feature slices Ψ s (X) of each scale s along the channel dimension to obtain the final mixed feature slice Ψ(X):

[0087] Ψ(X) = Concat(Ψ 1 (X),..., Ψ s (X)) (11)

[0088] Sub-step 4: Calculate the mixed feature slices of multiple layers l ∈ {1,..., L} of the old model and the current model, and then minimize the L2 distance between the mixed feature slices calculated at multiple layers during training. The internal feature distillation loss L int is defined as:

[0089]

[0090] Step 7: The feature f4 extracted by the new model encoder is upsampled by the decoder to obtain the segmented image As Figure 3 shown.

[0091] After the feature f4 extracted by the new model encoder is upsampled, the final segmented image is obtained The process is as follows:

[0092]

[0093] where Upsample(·) represents the upsampling operation.

Claims

1. A continuous semantic segmentation method based on internal and external distillation, characterized in that It includes the following steps: Step 1: Input the image I of the current task T T into the old model encoder E T-1 to obtain intermediate features i = 1, 2, 3, 4; The old model is a model trained by task T-1 and can segment all classes in task 0: T-1; the encoder E of the old model T-1 consists of 4 convolutions, and the intermediate features are the outputs of the 4 convolutions, as follows: In the formula, is the i-th convolutional layer in the old model encoder E T-1 , is the image I of the current task T T ; Step 2: Input the image I of the current task T T into the new model encoder E T and obtain intermediate features f1 and f2 respectively after passing through two convolutional layers; The new model can segment all categories in task 0: T. It has the same structure as the old model and is initialized with the parameters of the old model; The image I of task T T is input into the new model encoder E T wherein, through the first convolutional layer in the encoder, the first intermediate feature f1 is obtained, and through the second convolutional layer in the encoder, the second intermediate feature f2 is obtained. The process is as follows: Step 3: Calculate the attention maps of adjacent features of the old and new models and calculate the L2 distance between them; Step 3-1: Introduce an attention mechanism between the model layers and perform external feature distillation; for the features f l-1 and f l , the corresponding attention maps are obtained through the following operations: First, adjust their dimensions through 1×1 convolution operations to make them match and perform concatenation operations. The process is as follows: f l c = Concat(f l , Conv 1×1 (f l-1 )) (4) In the formula, l represents the index in the layer, 1×1 represents the convolution kernel size of 1, and Concat(·) represents the concatenation operation; Step 3-2: Implement spatial-aware feature extraction through the multi-scale convolutional attention MSCA using element-wise multiplication; MSCA is written as: where, Conv DW (·) represents depth convolution, Scale i (·) performs multi-scale processing on the features, and the features after multi-scale processing are convolved by 1×1 to obtain the attention map; Step 3-3: Similarly, the features of the old model and are subjected to the same operation to obtain the attention map Step 3-4: Calculate the attention map of adjacent features between the old and new models and The L2 distance between them is used as the external feature distillation loss L ext To prevent forgetting, it is expressed as follows: Step 4: Multiply the calculated new model feature attention map by the feature f2 to obtain the feature Integrate the calculated new model feature attention map into the corresponding feature. For f1 and f2, the attention map calculated according to Step 3 is Multiply by the feature f2 to obtain the feature The process is as follows: Among them, represents matrix multiplication; Step 5: Feature Enter the third convolution of the new model encoder to obtain f3, and repeat the above process to obtain f4; Feature Enter the third convolution of the new model encoder to obtain f3, and calculate the feature attention map between f3 and f2 And multiply it with f3 to obtain Obtain f4 through the fourth convolution of the new model encoder; Step 6: Intermediate feature f of the new model l is internally distilled with the corresponding intermediate feature of the old model to calculate the internal distillation loss L int ; Step 6-1: Perform hybrid pooling on the features; Let X represent a certain feature in f l with a size of H×W×C; use the Φ function to extract the maximum, minimum, and average information of feature X, and concatenate the H×C-width mixed pooling slices and W×C-height mixed pooling slices of X: In the formula, Concat(.) represents the concatenation in the channel dimension; Ω(X) is an operation that finds the maximum and minimum values and combines them according to the corresponding dimensions. Its calculation process is as follows: Ω(X[:, w, :]) = Concat(max(X[:, w, :]), min(X[:, w, :])) (9) Step 6-2: Calculate pooled slices with mixed width and height on multiple regions extracted at different scales to retain local information, where the scales are {1 / 2 s} s=0...S ; given a scale s and a feature X, the mixed pooled slice Ψ s (X) of the feature X at this scale is expressed as follows: wherein, This is a sub-region of Feature X, for the size of this sub-region is H / s × W / s; Step 6-3: Connect the mixed feature slices Ψ s (X) at each scale s along the channel dimension to obtain the final mixed feature slice Ψ(X): Ψ(X) = Concat(Ψ 1 (X),..., Ψ s (X)) (11) Step 6-4: Calculate the mixed feature slices of multiple layers \(l\in\{1,\ldots,L\}\) of the old model and the current model, and then minimize the L2 distance between the mixed feature slices calculated at multiple layers during training; the internal feature distillation loss \(L\) int is defined as: Step 7: The feature f4 extracted by the new model encoder is upsampled by the decoder to obtain the segmentation image After the features f4 extracted by the new model encoder are upsampled, the final segmented image is obtained The process is as follows: In the formula, Upsample(·) represents the upsampling operation.

Citation Information

Patent Citations

  • Continuous image semantic segmentation method based on multilevel knowledge distillation

    CN114120319A

  • Image semantic segmentation method based on attention mechanism and knowledge distillation

    CN116703947A