HRNet Self-Distillation Object Segmentation Method Based on Multi-Scale Pooling Pyramid
By using a self-distillation learning method of multi-scale pooled pyramid and HRNet in the target segmentation task, and integrating multiple loss functions, the problem of improving target segmentation performance in the existing technology is solved, and the effect of significantly improving target segmentation performance without adding parameters is achieved.
Patent Information
- Application Number
- CN202111540428.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-12-16
AI Technical Summary
In the prior art, there are few researches on self-distillation learning methods for target segmentation tasks, and it is difficult to improve target segmentation performance without increasing network parameters.
The HRNet self-distillation target segmentation method based on multi-scale pooled pyramids is adopted. By cascading the multi-scale pooled pyramid module with the branch features of HRNet, a self-distillation learning structure is constructed, integrating KL divergence, cross-soil loss and structured similarity loss, and model training is carried out to improve the target segmentation performance.
Without increasing the HRNet network parameters, the target segmentation performance was significantly improved, and experimental verification was conducted to improve the F value of 1.417 percentage points on the four public data sets on average.
Smart Images

Figure CN114187308B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and machine learning, and particularly relates to an HRNet self-distillation object segmentation method based on a multi-scale pooling pyramid. Background Art
[0002] Knowledge distillation can transfer the representation ability of large-scale networks to small-scale networks and improve the classification and regression performance of lightweight networks, which is mainly used for the lightweight processing of deep neural networks.
[0003] At present, many research institutions are engaged in the study of knowledge distillation methods. For example, Zhang et al. (Zhang L., et al. Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation[C]. IEEE International Conference on Computer Vision, 2019) extracted the features of different layers of the residual network, regarded the deep features as the distillation teacher and the shallow features as the student, constructed a self-distillation learning structure, and combined Softmax and cross-entropy loss for distillation learning to improve the image classification performance of the residual network. Yang et al. (Yang C., et al., Snapshot Distillation: Teacher-Student Optimization in One Generation[C]. IEEE Conference on Computer Vision and Pattern Recognition, 2019) used the network generated in the previous round as the teacher network and the network in the current learning process as the student during the network training process, constructed a distillation structure, and used KL divergence for distillation learning to improve the image classification performance of the network. Duo L. et al. (Duo L., et al., Dynamic Hierarchical Mimicking Towards Consistent Optimization Objectives[C]. IEEE Conference on Computer Vision and Pattern Recognition, 2020) extracted the features of different layers of the residual network, regarded the features of different layers as the distillation teacher and the student for each other, constructed a consistency distillation structure, and used KL divergence for self-distillation learning to improve the image classification performance of the network. The above self-distillation learning methods are mainly for image classification tasks, extracting the features of different layers of the deep neural network, and using loss functions such as KL divergence and cross-entropy to achieve distillation learning and improve the performance of the network. However, there is less research on self-distillation learning methods for target segmentation tasks. Summary of the Invention
[0004] The purpose of the present invention is to provide an HRNet self-distillation target segmentation method based on a multi-scale pooling pyramid to improve the target segmentation performance without increasing the parameters of the HRNet network.
[0005] The technical solution for achieving the object of the present invention is as follows: An HRNet self-distillation object segmentation method based on a multi-scale pooling pyramid, comprising:
[0006] Step 1: Cascade the multi-scale pooling pyramid module with the branch features and output features of HRNet respectively to obtain 4 groups of branch distillation features and 1 group of output-end distillation features;
[0007] Step 2: Construct a self-distillation learning structure, which includes a branch consistency distillation learning structure and a bottom-up distillation learning structure;
[0008] Step 3: Use the original segmentation network of HRNet as a benchmark model, construct a self-distillation learning structure on the benchmark model, fuse the KL divergence, cross-entropy loss, and structural similarity loss to form a self-distillation learning loss function for model training, and use the trained model to obtain the image object segmentation result.
[0009] Compared with the prior art, the present invention has the following remarkable advantages: (1) Cascade the multi-scale pooling pyramid structure with the four groups of sub-branch features of HRNet respectively, which improves the feature representation ability of HRNet; (2) Adopt a self-distillation learning structure, which fuses two modes of consistency distillation and bottom-up distillation, ensuring the consistency and correctness of the optimization direction of the branch structure; (3) Incorporate the structural similarity loss on the basis of the KL divergence and cross-entropy loss to achieve more accurate self-distillation learning; (4) Through experiments on 4 public datasets, it is verified that the present invention can improve the object segmentation performance of HRNet without increasing the parameter scale. Description of the Drawings
[0010] Figure 1 is the overall structure diagram of the self-distillation learning of the present invention.
[0011] Figure 2 is the branch consistency distillation structure diagram of the present invention.
[0012] Figure 3 is the bottom-up distillation structure diagram of the present invention.
[0013] Figure 4 is the multi-scale pooling pyramid feature representation structure diagram of the present invention. Detailed Embodiments
[0014] The present invention realizes an HRNet self-distillation object segmentation method based on a multi-scale pooling pyramid, comprising:
[0015] Step 1: Cascade the multi-scale pooling pyramid module with the branch features and output features of HRNet respectively to obtain 4 groups of branch distillation features and 1 group of output-end distillation features;
[0016] Step 2: Construct a self-distillation learning structure, which includes a branch consistency distillation learning structure and a bottom-up distillation learning structure;
[0017] Step 3: Use the original segmentation network of HRNet as a benchmark model, construct a self-distillation learning structure on the benchmark model, and fuse KL divergence, cross-entropy loss, and structural similarity loss to form a self-distillation learning loss function for model training.
[0018] Step 4: After the training is completed, remove the self-distillation representation structure used for training, and only use the trained HRNet basic network for target segmentation to obtain the image target segmentation result.
[0019] Furthermore, in Step 1, the multi-scale pooling pyramid module is cascaded with the branch features and output features of HRNet respectively to obtain 4 groups of branch distillation features and 1 group of output-end distillation features, specifically as follows:
[0020] (1) For a branch distillation feature, the network structure includes a convolutional layer Subconv, a multi-scale pooling pyramid module PSPModule, a convolutional layer Score, and a Sub softmax layer arranged in sequence. The four parameters of the convolutional layer Subconv are the kernel width, kernel height, number of input channels, and number of output channels, and the two parameters of the multi-scale pooling pyramid module PSPModule1 are the number of input channels and the number of output channels;
[0021] (2) For the output-end distillation feature, the network structure includes a concatenation layer Concat, a multi-scale pooling pyramid module PSPModule, a convolutional layer Score, and a Sub softmax layer arranged in sequence.
[0022] Furthermore, the multi-scale pooling pyramid module PSPModule is specifically as follows:
[0023] Set the size of the input feature InFeat to h×w×n, where h represents the height, w represents the width, and n represents the number of channels. The specific structure of the multi-scale pooling pyramid module PSPModule is: input feature InFeat → parallel four-way pooling feature extraction layer → concatenation layer → convolutional layer → output feature OutFeat; among them, the size of the output feature OutFeat is h×w×n;
[0024] The structure of the first path in the four-way pooling feature extraction layer is: pooling layer 1×1 → convolutional layer 1×1×n×n → normalization layer → bilinear interpolation layer h×w×n → pooling feature; the four-way pooling feature extraction layer is different in the parameters of the pooling layer, and the parameters of the pooling layers of the other three paths are 2×2, 3×3, and 6×6 respectively.
[0025] Furthermore, the self-distillation learning structure in step 2 includes a branch consistency distillation learning structure and a bottom-up distillation learning structure, where:
[0026] Branch consistency distillation learning structure: The four groups of distilled features of HRNet are used as the teacher end and the student end respectively, generating a total of 12 pairs of distilled pairs;
[0027] Bottom-up distillation learning structure: The distilled features at the output end of HRNet are used as the teacher end, and the four groups of branch distilled features are used as the student end respectively, generating 4 pairs of distilled pairs;
[0028] The above 16 pairs of distilled pairs are fused to form the self-distillation learning structure.
[0029] Furthermore, step 3 fuses KL divergence, cross-entropy loss, and structural similarity loss to form a self-distillation learning loss function for model training, specifically as follows:
[0030] Given a training dataset D = {(x i , y i )|i = 1, 2,..., N}, where x i represents the i-th image data in the dataset, N is the number of images in the dataset, y i ∈ (1,..., K) represents the corresponding pixel-level annotation map, and K is the number of prediction categories; W m is the weight matrix of the network body, is the weight matrix of the auxiliary classification network in the self-distillation structure, M is the number of auxiliary classification networks, and the specific positions where the auxiliary classification networks are connected in the overall network are denoted as A = {a i |i = 1, 2,..., M};
[0031] The overall loss function of self-distillation learning is expressed as:
[0032]
[0033] In formula (1), L m is the cross-entropy loss function of the network body:
[0034]
[0035]
[0036] In formula (1), L s is the prediction loss generated by the prediction result of the self-distillation network relative to the annotation map, that is, the fusion of cross-entropy loss and structural similarity loss, specifically as follows:
[0037]
[0038] The first term is the KL divergence loss, specifically as shown in Equation (5), and the second term is the structural similarity loss, specifically as shown in Equation (6):
[0039]
[0040]
[0041] In Equation (6), SSIM(*,*) is the structural similarity metric, which measures the structural differences between two images using the luminance, contrast, and structural differences in the image regions;
[0042] Let be abbreviated as f k (x i ). The structural similarity in Equation (6) is specifically:
[0043]
[0044] where are respectively and the mean of f k (x i ), are respectively and the standard deviation of f k (x i ), is the covariance between and f k (x i ); C 1 = 0.01 2 , C 2 = 0.03 2 are two constants;
[0045] In Equation (1), L k is the distillation loss of the self-distillation structure, specifically as shown in Equation (8). The first term is the KL divergence loss of the outputs of two distillation auxiliary classifiers, where KL(*) is the KL distance; the second term is the structural similarity loss between the two. λ 1 , λ 2 are two weight parameters, which are set to 0.8 and 0.2 respectively,
[0046]
[0047] The above steps of the present invention have the following characteristics:
[0048] (1) Feature extraction. The multi-scale pooling pyramid structure is cascaded with the four groups of sub-branch features of HRNet respectively to enhance the feature representation ability of HRNet.
[0049] (2) Self-distillation learning structure: It integrates two modes of consistency distillation and bottom-up distillation. Consistency distillation generates 12 groups of distillation pairs by taking the 4 sub-branches as the distillation teacher and student ends respectively, ensuring the consistency of the optimization direction of the branch structure. Bottom-up distillation takes the synthetic output end of HRNet as the teacher end and the output ends of the 4 branches as the student ends respectively to form 4 groups of distillation pairs, ensuring the correctness of the optimization direction of the branch structure.
[0050] (3) Self-distillation learning loss. Based on the KL divergence and cross-entropy loss, the structural similarity loss is incorporated to achieve more accurate self-distillation learning. The objective function of self-distillation training includes the distillation loss and segmentation loss of the 4 groups of branches, and the distillation loss and segmentation loss of the main branch. The distillation loss consists of the KL divergence and structural similarity loss between the target class probabilities generated by each group of distillation pairs. The segmentation loss includes the cross-entropy loss and structural similarity loss between the target class probabilities generated by the distillation features and the annotation map.
[0051] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0052] Embodiment
[0053] In combination with Figure 1 The implementation steps of the present invention will be further described.
[0054] Step 1, feature extraction. As Figure 1 shown, Subconv1, Subconv2, Subconv3, and Subconv4 are the four sub-branches of HRNet, and the four sub-branches are merged through a cascade layer to generate the final output features. The present invention constructs a multi-scale pooling pyramid on the basis of the four branches and the output end to obtain the distillation features. The specific structure is as follows:
[0055] (1) For a branch end (such as Subconv1), the network structure for obtaining its distillation features is: convolutional layer Subconv1 (3×3×48×48, the four parameters are the width of the convolutional kernel, the height of the convolutional kernel, the number of input channels, and the number of output channels) → multi-scale pooling pyramid module PSPModule1 (48×48, the two parameters are the number of input channels and the number of output channels) → convolutional layer Score1 (1×1×48×2) → Sub1 softmax layer → distillation features of branch 1. The distillation features of the other three branches are similar to the above structure, specifically as Figure 1 shown.
[0056] (2) For the output - end features, the network structure for obtaining its distilled features is: Concatenation layer Concat → Multi - scale pooling pyramid module PSPModule0(720×48) → Convolution layer Score0(1×1×48×2) → Sub0 softmax layer → Output - end distilled features.
[0057] The structure of the multi - scale pooling pyramid module in the distilled structure is as Figure 2 shown. If the size of the input feature InFeat is h (height)×w (width)×n (number of channels), the specific structure of the multi - scale pooling pyramid module is: Input feature InFeat(h×w×n) → Parallel four - path pooling feature extraction layer → Concatenation layer → Convolution layer(3×3×4n×n) → Output feature OutFeat(h×w×n). The four - path pooling feature extraction layers have similar structures. Taking the first path as an example, its structure is: Pooling layer(1×1) → Convolution layer(1×1×n×n) → Normalization layer → Bilinear interpolation layer(h×w×n, with parameters being the feature size obtained after interpolation: height×width×number of channels) → Pooling feature. The differences in the four - path pooling feature extraction layers lie in the parameters of the pooling layer. The pooling grid parameters of the other three paths are 2×2, 3×3, and 6×6 respectively.
[0058] Step 2: Construct a self - distillation learning structure.
[0059] In Step 1, four groups of branch distilled features and one group of output - end distilled features are obtained. A self - distillation learning structure is constructed based on these five groups of distilled features.
[0060] The self - distillation learning of the present invention includes a branch - consistency distillation learning structure and a bottom - up distillation learning structure.
[0061] The branch - consistency distillation learning structure is as Figure 3As shown in the figure. The four groups of branch distilled features Sub1 softmax, Sub2 softmax, Sub3 softmax, and Sub4 softmax are used as the student side and the teacher side of the distillation pair with each other, generating 12 distillation pairs. In the present invention, the distillation pair is represented as (A→B), where the starting point of the arrow is the teacher side and the end point is the student side. The 12 distillation pairs can be expressed as: (Sub1 softmax→Sub2 softmax), (Sub1 softmax←Sub2 softmax), (Sub1 softmax→Sub3 softmax), (Sub1 softmax←Sub3 softmax), (Sub1 softmax→Sub4 softmax), (Sub1 softmax←Sub4 softmax), (Sub2 softmax→Sub3 softmax), (Sub2 softmax←Sub3 softmax), (Sub2 softmax→Sub4 softmax), (Sub2 softmax←Sub4 softmax), (Sub3 softmax→Sub4 softmax), (Sub3 softmax←Sub4 softmax).
[0062] The bottom-up distillation learning structure is as Figure 4 shown. Taking the output distilled feature Sub0 softmax as the teacher side and the four groups of branch distilled features as the student sides respectively for distillation learning, generating 4 distillation pairs. Specifically: (Sub0 softmax→Sub1 softmax), (Sub0 softmax←Sub2 softmax), (Sub0 softmax→Sub3 softmax), (Sub0 softmax←Sub4 softmax).
[0063] Step 3: Self-distillation training.
[0064] 1. During model training, the sizes of the original images and their segmentation annotation maps are adjusted to 288×288×3. Set the batch size to 8, the number of training iterations to 20 rounds, the initial learning rate to 0.01, and the decay coefficient to 0.0005.
[0065] 2. Define the Loss function. Given the training data set D = {(x i , y i )|i = 1, 2,..., N}, where x i represents the i-th image data in the data set, N is the number of images included in the data set, and y i∈(1,...,K) represents its corresponding pixel-level annotation map, where K is the number of predicted categories, and in the present invention, K = 2. W m is the weight matrix of the network body (excluding the self-distillation network), is the weight matrix of the auxiliary classification network in the self-distillation structure, M is the number of auxiliary classification networks, and in the present invention, M = 5. The specific position where the auxiliary classification network is connected in the overall network is denoted as A = {a i |i = 1, 2,..., M}. The overall loss function of the self-distillation learning in the present invention is expressed as:
[0066]
[0067] In formula (1), L m is the learning loss function of the network body, which is the cross-entropy loss function in the invention:
[0068]
[0069]
[0070] In formula (1), L s is the prediction loss generated by the prediction result of the self-distillation network relative to the annotation map. In the present invention, it is the fusion of the KL divergence loss and the structural similarity loss, specifically as follows:
[0071]
[0072] where the first term is the cross-entropy loss, specifically as shown in formula (5). The second term is the structural similarity loss, specifically as shown in formula (6).
[0073]
[0074]
[0075] In formula (6), SSIM(*,*) is the structural similarity measure, which measures the structural difference between two images using the brightness, contrast, and structural differences of the image regions. Denote as f for short k (x i ). The structural similarity in formula 6 is specifically:
[0076]
[0077] where are respectively and the mean and standard deviation of f k (x i ), is and f k (xi ) The covariance between C 1 = 0.01 2 and C 2 = 0.03 2 are two constants.
[0078] In formula (1), L k The distillation loss of the self-distillation structure of the present invention is specifically shown in formula (8). The first term is the KL divergence loss of the outputs of two distillation auxiliary classifiers, where KL(*) is the KL distance. The second term is the structural similarity loss between the two. λ 1 and λ 2 are two weight parameters, which are set to 0.8 and 0.2 respectively in the present invention.
[0079]
[0080] 3. Model experiment. The original segmentation network of HRNet is used as the benchmark model for training and testing. The self-distillation structure of this article is constructed on the benchmark model for training, and the network is forward-inferred after removing the self-distillation structure to obtain the target segmentation result. By comparing the prediction results of the benchmark model and the model after self-distillation learning, the performance improvement effect of the present invention on the HRNet segmentation task is verified.
[0081] The experimental dataset consists of 4 publicly available object segmentation datasets, namely: (1) COD (Deng-Ping F., et al. Camouflaged Object Detection [C]. CVPR, 2020) is a natural camouflage object dataset, containing 10,000 natural camouflage images. (2) CPD (Fang Z., et al. Camouflage people detection via strong semantic dilation network [C]. The ACM Turing Celebration Conference-China, 2019) is a military camouflage single soldier dataset, containing 2,600 military camouflage single soldier images. (3) DUT-OMRON (Yang C., et al. Saliency Detection via Graph-Based Manifold Ranking [C]. CVPR, 2013) is a saliency object dataset containing 5,168 images. (4) PASCAL-S (Radhakrishna A., et al., Frequency-tuned salient region detection [C]. CVPR, 2009) is a saliency object dataset containing 850 images. The dataset is split into training data and test data according to the ratio of 0.6 and 0.4 for model training and testing.
[0082] The present invention uses the F-value (F-measure) commonly used in object segmentation tasks (Radhakrishna A., et al., Frequency-tuned salient region detection [C]. CVPR, 2009) to evaluate and compare the performance of the baseline model and the self-distillation model. The proportion of the area of the correctly detected target region in the area of the target region in the standard image is the precision, and the precision focuses on measuring the accuracy of the algorithm in detecting the target region. The proportion of the area of the correctly detected target region in all the target regions detected by the algorithm is the recall, and the recall focuses on measuring the completeness of the algorithm in detecting the target region. The F-value is a comprehensive evaluation index that combines detection precision and recall. The calculation formula of the F-value is shown in Equation (9), and β in the formula is set to 0.3.
[0083]
[0084] Table 1 (unit: %)
[0085]
[0086] The experimental results are shown in Table 1. On four typical object segmentation datasets, the present invention improves the performance of the HRNet model by an average of 1.417 percentage points, effectively improving the object segmentation performance of the HRNet baseline model.
Claims
1. An HRNet self-distillation object segmentation method based on a multi-scale pooling pyramid, characterized in that, it includes: Step 1: Cascade the multi-scale pooling pyramid module with the branch features and output features of HRNet respectively to obtain 4 groups of branch distillation features and 1 group of output-end distillation features; Step 2: Construct a self-distillation learning structure, which includes a branch consistency distillation learning structure and a bottom-up distillation learning structure; Step 3: Use the original segmentation network of HRNet as a benchmark model, construct a self-distillation learning structure on the benchmark model, fuse the KL divergence, cross-entropy loss, and structural similarity loss to form a self-distillation learning loss function for model training, and use the trained model to obtain the image object segmentation result; The cascading of the multi-scale pooling pyramid module with the branch features and output features of HRNet in Step 1 to obtain 4 groups of branch distillation features and 1 group of output-end distillation features is specifically as follows: (1) For a branch distillation feature, the network structure includes a convolutional layer Subconv, a multi-scale pooling pyramid module PSPModule, a convolutional layer Score, and a Sub softmax layer arranged in sequence. The four parameters of the convolutional layer Subconv are the convolutional kernel width, convolutional kernel height, number of input channels, and number of output channels respectively. The two parameters of the multi-scale pooling pyramid module PSPModule1 are the number of input channels and the number of output channels respectively; (2) For the output-end distillation feature, the network structure includes a concatenation layer Concat, a multi-scale pooling pyramid module PSPModule, a convolutional layer Score, and a Sub softmax layer arranged in sequence; The multi-scale pooling pyramid module PSPModule is specifically as follows: Set the size of the input feature InFeat to h×w×n, where h represents the height, w represents the width, and n represents the number of channels. The specific structure of the multi-scale pooling pyramid module PSPModule is: input feature InFeat → parallel four-way pooling feature extraction layer → concatenation layer → convolutional layer → output feature OutFeat; among them, the size of the output feature OutFeat is h×w×n; The structure of the first path in the four-way pooling feature extraction layer is: pooling layer 1×1 → convolutional layer 1×1×n×n → normalization layer → bilinear interpolation layer h×w×n → pooling feature; the four-way pooling feature extraction layer is different in the parameters of the pooling layer. The pooling layer parameters of the other three paths are 2×2, 3×3, and 6×6 respectively; The self-distillation learning structure in Step 2 includes a branch consistency distillation learning structure and a bottom-up distillation learning structure, where: Branch consistency distillation learning structure: Use the 4 groups of branch distillation features of HRNet as the teacher end and the student end respectively, and a total of 12 groups of distillation pairs are generated; Bottom-up distillation learning structure: Use the output-end distillation feature of HRNet as the teacher end, and use the 4 groups of branch distillation features as the student end respectively to generate 4 groups of distillation pairs; Fuse the above 16 groups of distillation pairs to form a self-distillation learning structure; In step 3, the KL divergence, cross-entropy loss, and structural similarity loss are combined to form a self-distillation learning loss function for model training, which is specifically as follows: Given a training dataset \(D = \{(x i , y i )|i = 1, 2, \ldots, N\}\), where \(x i \) represents the \(i\)-th image data in the dataset, \(N\) is the number of images contained in the dataset, \(y i \in(1, \ldots, K)\) represents the corresponding pixel-level annotation map, and \(K\) is the number of prediction classes; \(W m \) is the weight matrix of the network body, \) is the weight matrix of the auxiliary classification network in the self-distillation structure, \(M\) is the number of auxiliary classification networks, and the specific positions where the auxiliary classification networks are connected in the overall network are denoted as \(A=\{a i |i = 1, 2, \ldots, M\}\); The overall loss function of self-distillation learning is expressed as: L in formula (1) m is the cross-entropy loss function of the network body: L in Equation (1) s is the prediction loss generated by the prediction result of the self-distillation network relative to the annotation map, that is, the fusion of the cross-entropy loss and the structural similarity loss, which is specifically as follows: where the first term is the KL divergence loss, specifically as shown in Equation (5), and the second term is the structural similarity loss, specifically as shown in Equation (6): In Equation (6), SSIM(*,*) is the structural similarity measure, which measures the structural differences between two images using the luminance, contrast, and structural differences of the image regions; Let be abbreviated as f k (x i ), and the structural similarity in Equation (6) is specifically as follows: in y i k and f k (x i ), y i k and f k (x i ), for and f k (x i ) between the covariance; C 1 =0.01 2 , C 2 =0.03 2 are two constants; L in formula (1) k is the distillation loss of the self-distillation structure, specifically shown in formula (8). The first term is the KL divergence loss of the outputs of two distillation auxiliary classifiers, where KL(*) is the KL distance; the second term is the structural similarity loss between the two. λ 1 , λ 2 are two weight parameters, which are set to 0.8 and 0.2 respectively.