Pseudo-label semi-supervised semantic segmentation method based on diffusion model
By adopting a pseudo-label semi-supervised semantic segmentation method based on diffusion model in remote sensing image processing, combined with student-teacher model and entropy value calculation, the problem of insufficient generalization ability in the existing technology is solved, and higher segmentation accuracy and feature learning complexity are achieved.
Patent Information
- Application Number
- CN202510035028.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing semi-supervised semantic segmentation method lacks generalization ability in remote sensing image processing. The prediction error generated by the teacher model affects the accuracy of unsupervised losses, resulting in insufficient generalization ability of the model.
The pseudo-label semi-supervised semantic segmentation method based on the diffusion model is adopted to extract multi-scale features through the diffusion model, combine with the student-teacher model to generate pseudo-labels, and filter reliable pixels and unreliable pixels through entropy value calculations, and perform unsupervised loss calculations and comparison learning.
It improves the generalization ability of teacher and student models, improves segmentation accuracy, enhances the complexity of feature learning and discriminant ability of feature representation, and reduces the dependence on manual labeled data.
Smart Images

Figure CN119942117A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a pseudo-label semi-supervised semantic segmentation method based on a diffusion model. Background Art
[0002] Remote sensing image semantic segmentation is a key technology that uses different colors to represent different types of objects. It is widely used in road extraction, urban planning, land cover classification and other fields. The acquisition of high-resolution remote sensing images has become more convenient, but it also brings the demand for massive labeled data. Since pixel-level labeling requires a lot of manpower, is costly and extremely time-consuming, the scarcity of labeled data has become a major bottleneck, restricting the implementation of downstream tasks.
[0003] In order to solve the problem of insufficient annotated data for remote sensing images, researchers have proposed unsupervised, weakly supervised and semi-supervised methods. The current semi-supervised semantic segmentation framework can be mainly divided into the following types: adversarial methods, consistency regularization, contrastive learning and hybrid strategies. Among them, the core idea of consistency regularization is to ensure that the output results remain consistent when the input or model changes. Among the consistency regularization methods, the teacher-student model is the most common architecture. The teacher model uses the exponential moving average (EMA) of the student model to generate the target of consistency training; the student model is trained by optimizing the supervision loss of the labeled data and the consistency loss of the unlabeled data. For example, the publication number is: 202310204005.5, a semantic change detection method for high-resolution remote sensing images based on semi-supervised semantic segmentation contrastive learning, using change detection. This training method can help the model learn a more robust feature representation, thereby reducing its dependence on large-scale annotated data. However, relying solely on the teacher-student model for training may face problems: the predictions generated by the teacher model may have errors, which in turn affects the accuracy of the unsupervised loss, resulting in insufficient generalization of the model.
[0004] Another example is the remote sensing image semantic segmentation method and system based on the diffusion model and knowledge distillation, which is published as CN 117152427 A. The sample data set is divided into a support set and a query set; the diffusion model is used to generate an unlabeled image set; the teacher encoder and the student encoder are jointly pre-trained; the teacher encoder extracts features from the support set and obtains the prototype; the teacher encoder extracts features from the unlabeled set to achieve prediction; the distance between the support set and the prototype is calculated as the clustering radius to screen unlabeled samples; the prototype of the unlabeled set is calculated, reverse prototype generation is performed, and the loss is calculated; the samples in the clustering cluster are updated using pseudo labels to achieve prototype update; the student encoder performs multi-scale feature embedding on the query set and calculates the distance from the prototype to achieve category division, comprehensive loss, and iterative training. This method uses a diffusion model to generate an unlabeled image set, and the process is cumbersome. The student and teacher models also need to be pre-trained, and there are efficiency issues in the joint pre-training of the teacher encoder and the student encoder. The teacher model needs sufficient capacity to capture complex features, while the student model needs to reduce the number of parameters while maintaining performance. The balance between the two is difficult to grasp, making the method too complicated when extracting features. When processing high-resolution remote sensing images, this will lead to higher computing resource consumption and longer training time, and poor results. Summary of the invention
[0005] The purpose of the present invention is to provide a pseudo-label semi-supervised semantic segmentation method based on a diffusion model to improve the generalization ability of the teacher-student model, thereby improving the segmentation accuracy, increasing the complexity of feature learning and enhancing the discriminative ability of feature representation.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] A pseudo-label semi-supervised semantic segmentation method based on a diffusion model comprises the following steps:
[0008] S1: inputting remote sensing image data into a pre-trained diffusion model, performing unsupervised training using the diffusion model, and extracting multi-scale features;
[0009] S2: The segmentation head based on the student-teacher model screens and classifies the extracted multi-scale features into a labeled feature set and an unlabeled feature set, the labeled feature set is input into the student model to obtain a labeled semantic segmentation result, the unlabeled feature set is passed through the student model to obtain an unlabeled semantic segmentation result, and the unlabeled feature set is passed through the teacher model to generate a pseudo label;
[0010] S3: Filter out the pixels of the pseudo-label by entropy calculation, divide the pixels in the pseudo-label into reliable pixels and unreliable pixels according to a preset threshold, the reliable pixels participate in the unsupervised loss calculation of the student model, and are compared with the unlabeled semantic segmentation results generated by the student model, and the unreliable pixels are stored in the category memory as negative samples;
[0011] S4: Perform contrastive learning to obtain contrastive loss relative to unreliable labels;
[0012] S5: Determine the loss function optimization target Train the student-teacher model:
[0013]
[0014] in, represents the loss value of supervised training, represents the loss value of unsupervised training, is the contrast loss relative to unreliable labels, λ u is the weight for adjusting the unsupervised loss value, λ c is the weight of the contrastive loss.
[0015] Furthermore, it also includes S6: evaluating the pseudo-label semi-supervised semantic segmentation method based on the diffusion model.
[0016] Furthermore, the S6 evaluation includes: using an average intersection-over-union ratio or an ablation experiment, or combining the average intersection-over-union ratio with the ablation experiment.
[0017] Furthermore, the remote sensing image data is input into a pre-trained diffusion model, and unsupervised training is performed using the diffusion model to extract multi-scale features, specifically including:
[0018] Process the acquired remote sensing image data into a labeled image set and a collection of unlabeled images Then the labeled images in these two image sets and unlabeled images Split into multiple epochs by segmentation head m, batch B of labeled images in the current epoch l and a batch B of unlabeled images u Input into the pre-trained diffusion model Diff(), and extract the labeled multi-scale feature set through the diffusion model training and an unlabeled multi-scale feature set in
[0019]
[0020] Furthermore, the S2 specifically includes:
[0021] The student-teacher model has a consistent structure, including multiple convolutional layers, ReLU activation functions and upsampling layers. Band screening is performed through attention mechanisms and convolutional layers. The convolutional layer and activation function of the first layer are responsible for merging feature sequences, and weights are applied to features in the spatial-channel attention mechanism. The second layer of convolution further screens features, and the last layer of convolution is responsible for pixel classification. The feature map is upsampled to obtain a semantic segmentation result. The segmentation head based on the student-teacher model screens and classifies the extracted multi-scale features. After the labeled feature set is input into the student model, a labeled semantic segmentation result is obtained. After the unlabeled feature set passes through the student model, an unlabeled semantic segmentation result is obtained. At the same time, the unlabeled feature set is input into the teacher model to produce a pseudo label, and the weight of the teacher model is updated using an exponential moving average.
[0022] Furthermore, the S3 specifically includes:
[0023] S301: To prevent the student model from overfitting inaccurate pseudo labels, the probability distribution entropy H(p ij ) as an indicator to filter out low-quality pseudo-labels and thus provide more reliable supervision:
[0024]
[0025] p ij It represents the softmax probability of the pixel at the i-th row and j-th column of the teacher model’s prediction belonging to each category, and C is the number of categories;
[0026] S302: Dynamic partition adjustment: During the training process, the pseudo labels gradually become reliable, and a linear strategy is used in each epoch to adjust the proportion of unreliable pixels.
[0027]
[0028] When the entropy of a pixel is higher than the preset threshold φ t , it is considered that the pixel is highly chaotic, the category distribution is uncertain, and the pixel does not have reliable information. Therefore, the pixel is directly defined as empty in the pseudo label:
[0029]
[0030] for and γ t The relationship is expressed by the following formula:
[0031]
[0032] H is the entropy matrix, which is flattened into a one-dimensional array for percentile calculation. Num(H) is the total number of elements in the entropy matrix H. Num(h∈H|h≤x) represents the number of elements in the entropy matrix H that are less than or equal to x. This means that the entropy matrix H is selected to satisfy the value greater than or equal to Entropy value at percentage γ t ,although It changes in each epoch, but the actual entropy matrix H relative to different pseudo-labels within an epoch is a fixed value;
[0033] S303: Calculate the threshold value according to the value of the current entropy matrix, divide the pixels in the pseudo-label into reliable pixels and unreliable pixels, and perform unsupervised loss calculation on the reliable pixels and the unlabeled semantic segmentation result obtained by the student model to obtain an unsupervised training loss value Unreliable pixels are saved through the category memory bank and used as negative samples for contrastive learning.
[0034] Furthermore, the S4: performing contrastive learning to obtain a contrastive loss relative to unreliable labels specifically includes:
[0035] S401: Sample pixels of each class with marked feature sequence results in each batch as reference pixels. The reference pixel set is expressed as
[0036]
[0037] Among them, y ij is the label of the current feature sequence, δ p is the positive threshold for a specific class, z ij Represents the label value of the i-th row and j-th column in the label; similarly, the pixels of each class of the unlabeled feature sequence results in each batch are sampled as reference pixels, and the reference pixel set can be expressed as
[0038]
[0039] The set G of all qualified pixels of category c c It is defined as the union of all pixels that meet the positive sample conditions in the labeled image and the unlabeled image. This representation combines all qualified pixel features into a central representation as a positive sample of the category c:
[0040]
[0041] In order to construct the positive sample of category c, we calculate the qualified pixel set G cThe mean of all pixels in represents the positive sample of this category:
[0042]
[0043] S402: Negative samples are defined by binary indicator variables n ij (c) is used to indicate whether a pixel is a negative sample of category c. The sampling of negative samples is divided into two cases, labeled images and unlabeled images:
[0044]
[0045] in and are indicators of whether the j-th pixel of the labeled image and the unlabeled image i is qualified to be a negative sample of class c. First, a Kronecker delta function δ(x) and another one are defined. function:
[0046]
[0047] For labeled images, negative samples should satisfy the following two conditions at the same time: (1) the pixel does not belong to category c; (2) the pixel is difficult to distinguish from category c. In order to quantify the degree of “difficult to distinguish”, pixel-level sorting is introduced: where p ij is the predicted confidence that the pixel belongs to category c. The sorted position reflects the difficulty of distinguishing the pixel from category c. The pixel with the maximum confidence is sorted as 0, and the pixel with the minimum confidence is sorted as C-1 (C is the number of categories). The negative sample of category c in the labeled image is defined by the following formula:
[0048]
[0049] is the pixel ranking position, indicating the difficulty of distinguishing the pixel from category c; r l is the low level threshold;
[0050] For unlabeled images, negative samples in unlabeled images should satisfy the following conditions at the same time: (1) the prediction result is unreliable and may not belong to category c; (2) the pixel does not belong to the least likely category. The negative samples of category c in unlabeled images are defined by the following formula:
[0051]
[0052] Among them, γ t is the threshold, which determines whether the pixel is considered a credible negative sample; r h is a high-level threshold used to filter out pixels that are least likely to belong to this category, a set of negative samples of category c Defined as all pixels that meet the negative sample criteria:
[0053]
[0054] S403: Negative samples are stored in the category memory Q c , which is used for subsequent contrastive learning to obtain the contrast loss relative to unreliable labels Further, in said S5:
[0055]
[0056] Among them, F l represents the set of labeled multi-scale features in the diffusion model training step, It's a label. is a labeled multi-scale feature, is the cross loss function, m T is the teacher model, θ T are the parameters of the teacher model;
[0057]
[0058] Among them, F u represents the unlabeled multi-scale feature set in the diffusion model training step, Pseudo labels generated by the teacher model, is the unlabeled multi-scale feature, is the cross loss function, m s is the student model, θ s are the parameters of the student model;
[0059]
[0060] By maximizing the distance between positive and negative sample pairs, the model can learn similar data through self-supervision, where C represents the total number of categories, M represents the number of each category, N represents the number of negative samples, and z ci represents the feature vector representing the i-th sample in the c-th class, and Represents z ci The positive sample pairs and negative sample pairs are represented by the inner product of the positive sample pairs or negative sample pairs with the anchor points to represent the similarity of different pixels. τ represents the temperature coefficient, which controls the scaling of the inner product in the contrast loss, thereby adjusting the sensitivity of the model to the similarity difference.
[0061] Beneficial effects of the present invention:
[0062] In order to reduce the dependence on manual annotation of remote sensing images and effectively utilize the contextual information of remote sensing images, the present invention proposes a semi-supervised remote sensing image semantic segmentation framework, DiifSTMatch, which combines diffusion model, pseudo-label and contrastive learning. The multi-scale features of remote sensing images are extracted using the diffusion model to make up for the global feature extraction of convolutional neural networks; the attention mechanism is then used to perform feature dimensionality reduction and classification. During the feature dimensionality reduction and classification process, the teacher-student model is used to generate labels, and then the pseudo-labels are divided into reliable and unreliable pixels through a preset threshold. Reliable pixels participate in the unsupervised loss calculation, and unreliable pixels will be stored in the category memory as negative samples to participate in the contrastive learning loss calculation.
[0063] At least the following effects have been achieved:
[0064] (1) Compared with other advanced methods, DiffSTMatch achieves the best experimental results.
[0065] (2) This technology realizes semi-supervised semantic segmentation based on the diffusion model.
[0066] (3) DiffSTMatch uses the diffusion model for unsupervised training to model long-range dependencies, strengthen the representation of global image features, and enhance the information interaction between different regions.
[0067] (4) The segmentation head based on the student-teacher model screens and classifies the features. The teacher model uses pseudo labels generated by unlabeled data to alleviate the class imbalance problem caused by sample reduction, improve the generalization ability of the teacher-student model, and thus improve the segmentation accuracy.
[0068] (5) DiffSTMatch distinguishes reliable and unreliable pixels in pseudo-labels by calculating entropy values. Reliable pixels are used for unsupervised loss calculation, and unreliable pixels are used as negative samples for contrastive learning to improve the complexity of feature learning and enhance the discriminative ability of feature representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a flowchart of a specific embodiment 1 of the present invention;
[0070] Figure 2 It is a network structure diagram of the diffusion model in the specific embodiment 1 of the present invention;
[0071] Figure 3 The working principle of the teacher-student model in the specific embodiment 1 of the present invention;
[0072] Figure 4 It is a flowchart of the evaluation stage of the specific embodiment 1 of the present invention;
[0073] Figure 5It is a schematic diagram for comparing the evaluation results in specific embodiment 1 of the present invention. DETAILED DESCRIPTION
[0074] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0075] Embodiment 1:
[0076] See also Figures 1 to 5 As shown in FIG. 1 , a pseudo-label semi-supervised semantic segmentation method based on a diffusion model includes the following steps:
[0077] Step S1: inputting remote sensing image data into a pre-trained diffusion model, performing unsupervised training using the diffusion model, and extracting multi-scale features; specifically comprising:
[0078] Process the acquired remote sensing image data into a labeled image set and a collection of unlabeled images Then the labeled images in these two image sets and unlabeled images Split into multiple epochs by segmentation head m, batch B of labeled images in the current epoch l and a batch B of unlabeled images u Input into Diff() in the pre-trained diffusion model, and use the diffusion model for unsupervised training to extract multi-scale features. The pre-trained diffusion model is used to model long-range dependencies and generate multi-scale features to enhance information interaction between different regions. The principle of the diffusion model and its network structure are as follows;
[0079] S101: training a diffusion model Diff(), wherein the diffusion model adopts an existing diffusion model, and the training process of the diffusion model mainly includes two stages: a forward process and a reverse process, wherein the forward process is to add noise to the original image, gradually destroying the distribution mode of the original image; and the reverse process is to simulate the distribution mode of the original image in the initial disordered noise, so as to reconstruct the original image;
[0080] In the forward process, for a given target image x0, add Gaussian noise of a given step length to x0 Get images with different noise levels. Define these noisy images as a noise sequence {x1,x2,x3…x t}, where t is a positive integer. This noise sequence will be used as the target image in the reverse phase.
[0081]
[0082] The diffusion model Diff() is a latent variable model that uses a fixed Markov chain to map to the latent space. The chain gradually adds noise to the data to obtain an approximate posterior value q(x 1:T |x0), β t is a variance table given in advance.
[0083]
[0084] The reverse process, turning a disordered noise image into a target image, can be seen as the distribution of the noise function simulating the target image.
[0085] Initially, the model learns the joint distribution p θ (x 0:T )for:
[0086]
[0087] This joint probability describes the generation sequence from pure noise to an image that approximates the original image. θ (x t-1 |x t ) indicates that from x t to x t-1 The transition probability of . In the reverse generation stage of the diffusion model, this probability distribution is gradually denoised until an approximation of the original image is generated. The estimation of the model loss during training is to maximize the likelihood function p of the target data x0 θ (x0), which is equivalent to minimizing its negative log-likelihood. Since it is impossible to directly solve the expectation Therefore, a variational lower bound L is introduced for approximate optimization. Rewrite the expectation of the negative log-likelihood and expand the lower bound:
[0088]
[0089] Usually, x T The initial distribution p(x T ) is set to a known Gaussian distribution. Represents the noise x T The KL divergence loss at each step in the process of gradually restoring to data x0. Here p θ (x t-1 |x t ) is the inverse conditional probability distribution of the model, and q(x t |x t-1 ) is the true distribution defined by the addition of Gaussian noise in the forward process. Therefore, each KL divergence term measures the deviation between the inverse distribution generated by the model and the true inverse distribution. Substituting formulas (2) and (3) into formula (4) yields:
[0090]
[0091] Further express the parameters:
[0092]
[0093] For the Gaussian distribution p(x T ), when the variance is known, determine the mean The distribution is completely determined. Therefore, the variance β of the forward process is used t , set it to a constant, q(x T |x0) is the approximate posterior of the forward process. L T There are no parameters to learn. At the same time, the last item L0 of the inverse process is set as an independent discrete decoder
[0094] in,
[0095]
[0096] Where D is the data dimension, i represents the extracted coordinates, and x is the input.
[0097] For L t-1 :
[0098]
[0099] Since the variance β in the forward process is used t , so the model only needs to learn the mean μ θ (x t ,t) is the mean value in the forward process, which is the state x at the previous moment t and the variance of the noise process β t Decide:
[0100]
[0101] Substituting formula (7) and formula (10) into formula (9) yields:
[0102]
[0103] Based on existing work and experience, it is better to use a simplified objective that ignores weighted terms to train the diffusion model:
[0104]
[0105] S102: Figure 2As shown, the diffusion model in this specific embodiment adopts a U-type network structure similar to U-Net, which is an existing U-type network structure. Its encoder is composed of Basic, Basic_I and self-attention mechanism (Self-Attention). The Basic part includes convolution layer, GroupNorm, Identity and Swish activation function, while the Basic_I part replaces Identity with Dropout layer. The decoder stage needs to upsample the features extracted by the encoder, so an Upsample layer is added. The U-type network structure effectively combines low-level and high-level feature information by connecting the feature maps of the upsampling and downsampling paths. This connection method helps to retain richer image details and contextual information, allowing the network to focus on local and global features at the same time, thereby improving the model's ability to understand various parts of the image. In the decoding stage of the diffusion model, we extract features as input for subsequent classification, and we use the diffusion time step t j Set to [50, 100, 400]. According to some previous studies, the features extracted at these time steps are more effective. j The feature group of this stage will be generated
[0106] S103: batch label images l and unlabeled image batch β u Through the diffusion model training, a labeled multi-scale feature set is extracted and an unlabeled multi-scale feature set in
[0107]
[0108] Step S2: The segmentation head based on the student-teacher model screens and classifies the extracted multi-scale features into a labeled feature set and an unlabeled feature set. The labeled feature set is input into the student model to obtain a labeled semantic segmentation result. The unlabeled feature set is passed through the student model to obtain an unlabeled semantic segmentation result. The unlabeled feature set is passed through the teacher model to produce a pseudo label. Specifically, it includes:
[0109] Build as Figure 3The space-channel-based teacher and student models shown have the same structure as the student-teacher model, and both student-teacher models include multiple convolutional layers, ReLU activation functions, and upsampling layers. The segmentation head based on the student-teacher model screens and classifies the extracted multi-scale features; after the labeled feature set is input into the student model, the band screening is mainly performed through the attention mechanism and convolution layer. The convolution layer and activation function of the first layer are responsible for merging feature sequences. In the space-channel attention mechanism, weights are applied to the features. The second layer of convolution further screens the features, and the last layer of convolution is responsible for pixel classification. After upsampling, the feature map obtains the labeled semantic segmentation result; the unlabeled feature set is input into the student model as above to obtain the unlabeled semantic segmentation result; at the same time, the unlabeled feature set is input into the teacher model to produce pseudo labels, and the weight of the teacher model is updated using the exponential moving average (EMA). The pseudo labels generated by the teacher model using unlabeled data alleviate the class imbalance problem caused by sample reduction.
[0110] Step S3: Filter out the pixels of the pseudo-label by entropy calculation, divide the pixels in the pseudo-label into reliable pixels and unreliable pixels according to a preset threshold, the reliable pixels participate in the unsupervised loss calculation of the student model, and are compared with the unlabeled semantic segmentation results generated by the student model, and the unreliable pixels are stored in the category memory as negative samples; specifically including:
[0111] S301: The pseudo-label is used to calculate the unsupervised loss. To prevent the student model from overfitting inaccurate pseudo-labels, the probability distribution entropy H(p ij ) as an indicator to filter out low-quality pseudo-labels and thus provide more reliable supervision:
[0112]
[0113] p ij It represents the softmax probability that the pixel at the i-th row and j-th column of the teacher model belongs to each category, and C is the number of categories.
[0114] S302: Dynamic partition adjustment: During the training process, the pseudo labels gradually become reliable, and a linear strategy is used in each epoch to adjust the proportion of unreliable pixels.
[0115]
[0116] When the entropy of a pixel is higher than a preset threshold , it is considered that the pixel is highly chaotic, the category distribution is uncertain, and the pixel does not have reliable information. Therefore, the pixel is directly defined as empty in the pseudo label:
[0117]
[0118] for and γ t The relationship is expressed by the following formula:
[0119]
[0120] H is the entropy matrix, which is flattened into a one-dimensional array for percentile calculation. Num(H) is the total number of elements in the entropy matrix H. Num(h∈H|h≤x) represents the number of elements in the entropy matrix H that are less than or equal to x. This means selecting the entropy matrix H that satisfies the value greater than or equal to Entropy value at percentage γ t .although It changes in each epoch, but the actual entropy matrix H relative to different pseudo labels within an epoch is a fixed value. This is because the quality of pseudo labels improves along with the improvement of model accuracy.
[0121] S303: Calculate the threshold value according to the value of the current entropy matrix, retain or remove data more flexibly, divide the pixels in the pseudo-label into reliable pixels and unreliable pixels, and directly perform unsupervised loss calculation on the reliable pixels and the unlabeled semantic segmentation result obtained by the student model to obtain an unsupervised training loss value Unreliable pixels are saved in the category memory library, which is a category memory. Unreliable pixels represent the model's uncertainty about these pixels, but these unreliable pixels do not belong to other categories, so they are stored in the category memory library and used as negative samples for contrastive learning.
[0122] Step S4: Perform contrastive learning to obtain contrastive loss relative to unreliable labels; specifically, it includes:
[0123] S401: Sample pixels of each class with marked feature sequence results in each batch as reference pixels. The reference pixel set can be expressed as
[0124]
[0125] y ij is the label of the current feature sequence, δ p is the positive threshold for a particular class. ij Represents the label value of the i-th row and j-th column in the label. Similarly, the pixels of each class of the unlabeled feature sequence results in each batch are sampled as reference pixels, and the reference pixel set can be expressed as
[0126]
[0127] The set G of all qualified pixels of category c c It is defined as the union of all pixels that meet the positive sample conditions in the labeled image and the unlabeled image. This representation combines all qualified pixel features into a central representation as a positive sample of this category:
[0128]
[0129] In order to construct positive samples of category c, we calculate the qualified pixel set G c The mean of all pixels in represents the positive sample of this category:
[0130]
[0131] S402: Negative samples are defined by binary indicator variables n ij (c) indicates whether a pixel is a negative sample of category c. Specifically, the sampling of negative samples is divided into two cases, labeled images and unlabeled images:
[0132]
[0133] in and are indicators of whether the jth pixel of the labeled image and the unlabeled image i is qualified to be a negative sample of class c. First, define a Kronecker delta function δ(x) and another function:
[0134]
[0135] For labeled images, negative samples should satisfy the following two conditions at the same time: (1) the pixel does not belong to category c; (2) the pixel is difficult to distinguish from category c. In order to quantify the degree of “difficult to distinguish”, pixel-level sorting is introduced.
[0136] where p ij is the predicted confidence that the pixel belongs to category c. The sorted position reflects the difficulty of distinguishing the pixel from category c. The pixel with the highest confidence is ranked 0, and the pixel with the lowest confidence is ranked C-1 (C is the number of categories). The negative sample of category c in the labeled image is defined by the following formula:
[0137]
[0138] is the pixel ranking position, indicating the difficulty of distinguishing the pixel from category c; r lThe low level threshold.
[0139] For unlabeled images, the definition of negative samples is more complicated. Due to the lack of label information, negative samples in unlabeled images should meet the following conditions at the same time: (1) The prediction result is unreliable and may not belong to category c; (2) The pixel does not belong to the least likely category. The negative sample of category c in the unlabeled image is defined by the following formula:
[0140]
[0141] Among them, γ t is the threshold, which determines whether the pixel is considered a credible negative sample; r h is a high level threshold, used to filter out pixels that are least likely to belong to this category. Negative sample set of category c Defined as all pixels that meet the negative sample criteria:
[0142]
[0143] S403: Negative samples are stored in the category memory Q c , which is used for subsequent contrastive learning to obtain the contrast loss relative to unreliable labels It improves the complexity of feature learning and enhances the discriminative ability of feature representation.
[0144] Step S5: Determine the loss function optimization target Train the student-teacher model:
[0145]
[0146] in, represents the loss value of supervised training, represents the loss value of unsupervised training, is the contrast loss relative to unreliable labels, λ u is the weight for adjusting the unsupervised loss value, λ c is the weight of the contrast loss,
[0147] Said With the The Cross-Entropy loss function is used:
[0148]
[0149] Among them, F l represents the set of labeled multi-scale features in the diffusion model training step, It's a label. is a labeled multi-scale feature, is the cross loss function, m Tis the teacher model, θ T are the parameters of the teacher model;
[0150]
[0151] Among them, F u represents the unlabeled multi-scale feature set in the diffusion model training step, Pseudo labels generated by the teacher model, is the unlabeled multi-scale feature, is the cross loss function, m s is the student model, θ s are the parameters of the student model; in the representation learning process of contrastive learning, similar samples are trained to be close in the representation space, while different samples are kept away.
[0152]
[0153] The loss function maximizes the distance between positive and negative sample pairs to allow the model to learn similar data in a self-supervised manner. Where C represents the total number of categories, M represents the number of each category, N represents the number of negative samples, and z ci represents the feature vector representing the i-th sample in the c-th class, and Represents z ci The positive sample pairs and negative sample pairs are represented by the inner product of the positive sample pairs or negative sample pairs with the anchor points to represent the similarity of different pixels. τ represents the temperature coefficient, which controls the scaling of the inner product in the contrast loss, thereby adjusting the sensitivity of the student-teacher model to the similarity difference. In this specific embodiment: M = 50, N = 256 and τ = 0.5.
[0154] Step S6: Evaluate the semantic segmentation results obtained by the pseudo-label semi-supervised semantic segmentation method based on the diffusion model.
[0155] The evaluation in S6 includes: using the average intersection-over-union ratio or ablation experiment, or combining the average intersection-over-union ratio with the ablation experiment. In this specific embodiment: the evaluation is performed by combining the average intersection-over-union ratio with the ablation experiment.
[0156] The mean intersection over union (mIoU) is used to evaluate the pseudo-label semi-supervised semantic segmentation method.
[0157]
[0158] At the same time, the following indicators are used as evaluation criteria:
[0159] U2PL: U2PL is an early model trained with unreliable labels.
[0160] Unimatch: Unimatch proposes a two-stream perturbation technique that enables two strong views to jointly guide a weak view. It unifies image- and feature-level perturbations to form a more diverse perturbation space.
[0161] ST++: ST++ constructs a strong self-training (ST) baseline for semi-supervised semantic segmentation by injecting strong data augmentation (SDA) on unlabeled images, eliminating overfitting noisy labels, and disentangling similar predictions between teacher and student.
[0162] CorrMatch: CorrMatch revisits the challenge of accurately assigning pseudo-labels to unlabeled data from the perspective of label propagation.
[0163] LSST: LSST proposes a simple yet effective semi-supervised learning framework based on linear sampling (LS) self-training, which can adaptively assign thresholds to different categories and thus provide a noise-free region for retraining.
[0164] At the same time, ablation experiments are conducted to verify the effectiveness of various components of DiffSTMatch in improving the accuracy of semantic segmentation results of remote sensing images. Taking the segmentation results of the diffusion model as the benchmark, the experimental results of the model with the diffusion model and the teacher-student model (STM) components are compared, and then the experimental results of the model with the contrastive learning component are compared, and then the experimental results of the model with the combination of unreliable pseudo-labels and dynamic partition adjustment are compared to see the effect of the combination of all components on improving the baseline and the MioU value. At the same time, the probability level threshold and the initial threshold α0 of unreliable pseudo-labels are studied based on the ablation experiment results.
[0165] like Figure 4 As shown, in this specific embodiment: the Potsdam dataset is selected for experimentation. The Potsdam dataset is a high-resolution urban object classification dataset provided by the German Aerospace Center (DLR), which is mainly used for semantic segmentation and object detection tasks of remote sensing images. This dataset covers the city of Potsdam, Germany, acquired through aerial photography, with a spatial resolution of 5 cm, and finely depicts the structure and texture of urban objects. It contains RGB, near-infrared (NIR) images and digital surface models (DSM), and annotates six major types of objects such as buildings, roads, trees, low vegetation, and vehicles. The Potsdam dataset is suitable for urban object classification and change detection research in high-resolution remote sensing images, and is an important remote sensing benchmark dataset.
[0166] Step 1: Given a set of labeled images using the Potsdam dataset and a collection of unlabeled images
[0167] Step 2: Use the diffusion model to train the data, the labeled images in the two image sets and unlabeled images Noise destruction and reconstruction are performed through diffusion model.
[0168] Step 3: Train the segmentation head m to label the image and unlabeled images In the current epoch, it is processed into a batch of labeled images B l and a batch B of unlabeled images u These batches of images will enter the diffusion model, and the feature set with labels will be obtained from the diffusion model decoding stage i∈[0,B l ], and an unlabeled feature set i∈[0,B u ]. These low-dimensional latent feature sets are input into the student-teacher model for pixel classification to obtain semantic segmentation results.
[0169] Step 4: Use multiple existing methods for comparison. To verify the feasibility of the present invention, five existing comparison methods are selected:
[0170] U2PL: U2PL is an early model trained with unreliable labels.
[0171] Unimatch: Unimatch proposes a two-stream perturbation technique that enables two strong views to jointly guide a weak view. It unifies image- and feature-level perturbations to form a more diverse perturbation space.
[0172] ST++: ST++ constructs a strong self-training (ST) baseline for semi-supervised semantic segmentation by injecting strong data augmentation (SDA) on unlabeled images, eliminating overfitting noisy labels, and disentangling similar predictions between teacher and student.
[0173] CorrMatch: CorrMatch revisits the challenge of accurately assigning pseudo-labels to unlabeled data from the perspective of label propagation.
[0174] LSST: LSST proposes a simple yet effective semi-supervised learning framework based on linear sampling (LS) self-training, which can adaptively assign thresholds to different categories and thus provide a noise-free region for retraining.
[0175] Step 5: Result Analysis
[0176] (1) Assessment criteria and indicators:
[0177] Use the average intersection-over-union ratio as the accuracy indicator of the experiment:
[0178]
[0179] (2) Accuracy assessment
[0180] Table 1 shows the experimental results of all methods on the Potsdam dataset. From Table 1, we can conclude that DiffSTMatch achieved MIoU scores of 77.61%, 81.15%, and 89.12% when using 12.5% of labels, 25% of labels, and 50% of labels, respectively. CorrMatch achieved scores of 72.13%, 78.27%, and 81.67% when using 12.5% of labels, 25% of labels, and 50% of labels, respectively. UniMatch achieved scores of 75.80%, 80.34%, and 84.85% when using 12.5% of labels, 25% of labels, and 50% of labels, respectively. Compared with UniMatch, DiffSTMatch improved by 1.81% when the label was 12.5%. When the label was 25%, DiffSTMatch improved by 0.81%. When the label was 50%, DiffSTMatch improved by 4.72%.
[0181] Table 1 Summary of quantitative results of Postdam dataset
[0182]
[0183]
[0184] like Figure 5 As shown in the figure, the experimental results of different methods on the Potsdam dataset are visualized. In the figure, e represents the U2PL method, d represents the LSST method, c represents the ST++ method, f represents the CorrMatch method, g represents the UniMatch method, h represents the DiffSTMatch method, and a represents the image. Figure 5 As can be seen in the figure, other methods have more error areas. The experimental results show that DiffSTMatch has the best visual effect and the prediction results are closer to the real images. Some pictures are blocked by shadows, which poses a great challenge to the model's recognition of the contours of objects. For example, in the third row of pictures, the tall building in the lower right corner is illuminated by sunlight, and its roof shadow covers the adjacent shorter buildings and roads, causing the model to confuse the background and the building. The prediction results verify this phenomenon. All methods have different degrees of misjudgment in the shadow area and classify the building as the background. DiffSTMatch significantly reduces such errors. Similar situations occur in the seventh and eighth rows of pictures, and this method also minimizes the number of misclassified pixels. This further proves the superior performance of DiffSTMatch in context information extraction.
[0185] (3) In order to verify the effectiveness of the proposed semi-supervised semantic segmentation framework for remote sensing images, DiffSTMatch, which combines the diffusion model, pseudo-labels and contrastive learning, an ablation study was carried out.
[0186] ① According to Table 2, under the condition of 1 / 2 labeling rate of Potsdam dataset, ablation experiments were conducted on the components of DiffSCMatch. First, segmentation experiments were performed using only the diffusion model, and the result was a score of 79.10%, which was used as the baseline. Subsequently, after adding the teacher-student model (STM), the score increased to 82.11%, an increase of 3.01% over the baseline. Next, contrastive learning was introduced, and the baseline increased by 6.82%. Further adding probability level threshold and dynamic partition adjustment (DPA) increased the baseline by 6.94%. Combining unreliable pseudo labels with dynamic partition adjustment, the baseline increased by 8.15%. When unreliable pseudo labels were combined with probability level thresholds, the baseline increased by 9.01%. Finally, after integrating all components, the baseline increased by 10.02%, and MIoU reached 89.12%.
[0187] Table 2 Component effectiveness ablation experiment
[0188] Diffusion STM CL Unreliable PRT DPA MIoU √ 79.10% √ √ 82.11% √ √ √ 85.92% √ √ √ √ √ 86.04% √ √ √ √ √ 87.25% √ √ √ √ √ 88.11% √ √ √ √ √ √ 89.12%
[0189] ② In the Method's benchmark pixels, a probability level threshold is used to measure the amount of information and the confusion caused by unreliable pixels. l and high level threshold r h As shown in Table 3, the experiment is conducted at 1 / 2 labeling rate of Potsdam dataset. l =3 and r h = 20 achieved the best score, 89.12%. It can effectively balance the amount of information and confusion, and significantly improve the segmentation performance. l =1) will lead to too many false negative examples in the pseudo-label, affecting the distinction of pixels within the class; l =10, the negative examples may be semantically irrelevant to the corresponding anchor pixels, resulting in insufficient information.
[0190] Table 3 Probability level threshold effectiveness ablation experiment
[0191]
[0192]
[0193] ③After the previous ablation experiment, it was found that unreliable pseudo-labels have a significant improvement on the experiment, so the initial threshold α0 of unreliable pseudo-labels was studied. The experiment was conducted at a labeling rate of 1 / 2 of the Potsdam dataset. Experiments were conducted on α0 = {0.1, 0.2, 0.3, 0.4, 0.5} respectively. As shown in Table 4, it was found that α0 = 0.2 achieved the best score. A small α0 cannot correctly filter pseudo-labels. A large α0 cannot make full use of some high-confidence samples.
[0194] Table 4 Hyperparameter ablation experiment
[0195] <![CDATA[α0]]> 0.5 0.4 0.3 0.2 0.1 MIoU 87.80% 88.03% 88.93% 89.12% 87.24%
[0196] The technical solution provided by the present invention is described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A pseudo-label semi-supervised semantic segmentation method based on a diffusion model, characterized in that: The steps include: S1: inputting remote sensing image data into a pre-trained diffusion model, performing unsupervised training using the diffusion model, and extracting multi-scale features; S2: The segmentation head based on the student-teacher model screens and classifies the extracted multi-scale features into a labeled feature set and an unlabeled feature set, the labeled feature set is input into the student model to obtain a labeled semantic segmentation result, the unlabeled feature set is passed through the student model to obtain an unlabeled semantic segmentation result, and the unlabeled feature set is passed through the teacher model to generate a pseudo label; S3: Filter out the pixels of the pseudo-label by entropy calculation, divide the pixels in the pseudo-label into reliable pixels and unreliable pixels according to a preset threshold, the reliable pixels participate in the unsupervised loss calculation of the student model, and are compared with the unlabeled semantic segmentation results generated by the student model, and the unreliable pixels are stored in the category memory as negative samples; S4: Perform contrastive learning to obtain contrastive loss relative to unreliable labels; S5: Determine the loss function optimization target Train the student-teacher model: in, represents the loss value of supervised training, represents the loss value of unsupervised training, is the contrast loss relative to unreliable labels, λ u is the weight for adjusting the unsupervised loss value, λ c is the weight of the contrastive loss.
2. According to claim 1, a pseudo-label semi-supervised semantic segmentation method based on a diffusion model is characterized by: It also includes S6: Evaluation of pseudo-label semi-supervised semantic segmentation methods based on diffusion models.
3. The pseudo-label semi-supervised semantic segmentation method based on diffusion model according to claim 2, characterized in that: The S6 evaluation includes: using the average intersection-over-union ratio or ablation experiment, or combining the average intersection-over-union ratio with the ablation experiment.
4. The pseudo-label semi-supervised semantic segmentation method based on a diffusion model according to claim 1, characterized in that: The remote sensing image data is input into a pre-trained diffusion model, and unsupervised training is performed using the diffusion model to extract multi-scale features, specifically including: Process the acquired remote sensing image data into a labeled image set and a collection of unlabeled images Then the labeled images in these two image sets and unlabeled images Split into multiple epochs by segmentation head m, batch B of labeled images in the current epoch l and a batch B of unlabeled images u Input into the pre-trained diffusion model Diff(), and extract the labeled multi-scale feature set through the diffusion model training and an unlabeled multi-scale feature set in 5. The pseudo-label semi-supervised semantic segmentation method based on diffusion model according to claim 1, characterized in that: The S2 specifically includes: The student-teacher model has a consistent structure, including multiple convolutional layers, ReLU activation functions and upsampling layers. Band screening is performed through attention mechanisms and convolutional layers. The convolutional layer and activation function of the first layer are responsible for merging feature sequences, and weights are applied to features in the spatial-channel attention mechanism. The second layer of convolution further screens features, and the last layer of convolution is responsible for pixel classification. The feature map is upsampled to obtain a semantic segmentation result. The segmentation head based on the student-teacher model screens and classifies the extracted multi-scale features. After the labeled feature set is input into the student model, a labeled semantic segmentation result is obtained. After the unlabeled feature set passes through the student model, an unlabeled semantic segmentation result is obtained. At the same time, the unlabeled feature set is input into the teacher model to produce a pseudo label, and the weight of the teacher model is updated using an exponential moving average.
6. The pseudo-label semi-supervised semantic segmentation method based on diffusion model according to claim 1, characterized in that: The S3 specifically includes: S301: To prevent the student model from overfitting inaccurate pseudo labels, the probability distribution entropy H(p ij ) as an indicator to filter out low-quality pseudo-labels and thus provide more reliable supervision: p ij It represents the softmax probability of the pixel at the i-th row and j-th column of the teacher model’s prediction belonging to each category, and C is the number of categories; S302: Dynamic partition adjustment: During the training process, the pseudo labels gradually become reliable, and a linear strategy is used in each epoch to adjust the proportion of unreliable pixels. When the entropy of a pixel is higher than a preset threshold , it is considered that the pixel is highly chaotic, the category distribution is uncertain, and the pixel does not have reliable information. Therefore, the pixel is directly defined as empty in the pseudo label: for and γ t The relationship is expressed by the following formula: H is the entropy matrix, which is flattened into a one-dimensional array for percentile calculation. Num(H) is the total number of elements in the entropy matrix H. Num(h∈H|h≤x) represents the number of elements in the entropy matrix H that are less than or equal to x. This means that the entropy matrix H is selected to satisfy the value greater than or equal to Entropy value at percentage γ t ,although It changes in each epoch, but the actual entropy matrix H relative to different pseudo-labels within an epoch is a fixed value; S303: Calculate the threshold value according to the value of the current entropy matrix, divide the pixels in the pseudo-label into reliable pixels and unreliable pixels, and perform unsupervised loss calculation on the reliable pixels and the unlabeled semantic segmentation result obtained by the student model to obtain an unsupervised training loss value Unreliable pixels are saved through the category memory bank and used as negative samples for contrastive learning.
7. The pseudo-label semi-supervised semantic segmentation method based on diffusion model according to claim 1, characterized in that: S4: performing contrastive learning to obtain contrastive loss relative to unreliable labels, specifically includes: S401: Sample pixels of each class with marked feature sequence results in each batch as reference pixels. The reference pixel set is expressed as Among them, y ij is the label of the current feature sequence, δ p is the positive threshold for a specific class, z ij Represents the label value of the i-th row and j-th column in the label; similarly, the pixels of each class of the unlabeled feature sequence results in each batch are sampled as reference pixels, and the reference pixel set can be expressed as The set G of all qualified pixels of category c c It is defined as the union of all pixels that meet the positive sample conditions in the labeled image and the unlabeled image. This representation combines all qualified pixel features into a central representation as a positive sample of the category c: In order to construct the positive sample of category c, we calculate the qualified pixel set G c The mean of all pixels in represents the positive sample of this category: S402: Negative samples are defined by binary indicator variables n ij (c) is used to indicate whether a pixel is a negative sample of category c. The sampling of negative samples is divided into two cases, labeled images and unlabeled images: in and are indicators of whether the j-th pixel of the labeled image and the unlabeled image i is qualified to be a negative sample of class c. First, a Kronecker delta function δ(x) and another one are defined. function: For labeled images, negative samples should satisfy the following two conditions at the same time: (1) the pixel does not belong to category c; (2) the pixel is difficult to distinguish from category c. In order to quantify the degree of "difficult to distinguish", pixel-level sorting is introduced: where p ij is the predicted confidence that the pixel belongs to category c. The sorted position reflects the difficulty of distinguishing the pixel from category c. The pixel with the maximum confidence is sorted as 0, and the pixel with the minimum confidence is sorted as C-1 (C is the number of categories). The negative sample of category c in the labeled image is defined by the following formula: is the pixel ranking position, indicating the difficulty of distinguishing the pixel from category c; r l is the low level threshold; For unlabeled images, negative samples in unlabeled images should satisfy the following conditions at the same time: (1) the prediction result is unreliable and may not belong to category c; (2) the pixel does not belong to the least likely category. The negative samples of category c in unlabeled images are defined by the following formula: Among them, γ t is the threshold, which determines whether the pixel is considered a credible negative sample; r h is a high-level threshold used to filter out pixels that are least likely to belong to this category, a set of negative samples of category c Defined as all pixels that meet the negative sample criteria: S403: Negative samples are stored in the category memory Q c , which is used for subsequent contrastive learning to obtain the contrast loss relative to unreliable labels 8. The pseudo-label semi-supervised semantic segmentation method based on a diffusion model according to claim 1 or 4, characterized in that: In S5: Among them, F l represents the set of labeled multi-scale features in the diffusion model training step, It's a label. is a labeled multi-scale feature, l ce is the cross loss function, m T is the teacher model, θ T are the parameters of the teacher model; Among them, F u represents the unlabeled multi-scale feature set in the diffusion model training step, Pseudo labels generated by the teacher model, is the unlabeled multi-scale feature, l ce is the cross loss function, m s is the student model, θ s are the parameters of the student model; By maximizing the distance between positive and negative sample pairs, the model can learn similar data through self-supervision, where C represents the total number of categories, M represents the number of each category, N represents the number of negative samples, and z ci represents the feature vector representing the i-th sample in the c-th class, and Represents z ci The positive sample pairs and negative sample pairs are represented by the inner product of the positive sample pairs or negative sample pairs with the anchor points to represent the similarity of different pixels. τ represents the temperature coefficient, which controls the scaling of the inner product in the contrast loss, thereby adjusting the sensitivity of the model to the similarity difference.
Citation Information
Patent Citations
High-resolution remote sensing image semantic change detection method based on semi-supervised semantic segmentation contrast learning
CN116310812A
Semi-supervised hyperspectral image classification method based on unreliable pseudo-label learning
CN115953621A
Construction method of coal rock microstructure grouping automatic segmentation model
CN117011646A
Remote sensing image semantic segmentation method and system based on diffusion model and knowledge distillation
CN117152427A
Remote sensing image building extraction method and device based on diffusion model
CN117372873A
Cited By
Model training method, image perception method, device and related equipment
CN120526292A
Semi-supervised image segmentation method based on adaptive pixel subdivision
CN120635465A
Passive domain adaptive semantic segmentation method and device for diffusion-guided pseudo-label enhancement
CN121190764A
Semi-supervised spine segmentation method based on global-local semantic constraint visual language model
CN121437534A
Medical image segmentation method and system based on diffusion difference learning
CN121811412A