Self-adaptive image semantic segmentation method based on pixel dual consistency and energy fraction
By adopting pixel dual consistency and energy fraction methods in the semantic segmentation of adaptive street scene images in unsupervised fields, the problems of low pseudo-label reliability and poor generalization capabilities in the prior art are solved, and the semantic segmentation effect with high precision and high reliability is achieved.
Patent Information
- Application Number
- CN202311801255.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-12-26
AI Technical Summary
The existing unsupervised field adaptive street scene image semantic segmentation method is limited when adapting to various scene changes, the pseudo-label reliability is low, and unlabeled data cannot be effectively utilized, resulting in poor generalization capabilities of the model.
Using a method based on pixel dual consistency and energy fraction, the perturbation space is expanded by performing different perturbations at the image level and feature level, and unlabeled samples are screened through energy fractions to generate high-reliability pseudo-labels.
The prediction reliability and generalization ability of the model in the target domain are improved, and the unlabeled data is effectively utilized to achieve high-precision semantic segmentation of street scene images.
Smart Images

Figure CN119992078A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and semantic segmentation, and in particular to an unsupervised domain adaptive street view image semantic segmentation method based on pixel dual consistency and energy score. Background Art
[0002] In the field of image semantic segmentation, the unsupervised domain adaptation (UDA) method plays an increasingly important role. UDA is of particular significance in the field of image semantic segmentation because it can use unlabeled data to improve the generalization ability of the model, thereby reducing the need for labeled data. It has an auxiliary effect on both intelligent transportation systems and autonomous driving systems.
[0003] The existing mainstream solution for unsupervised domain adaptive street view image semantic segmentation is to use a strong or weak consistency framework. The model may be limited in adapting to various scene changes, the pseudo-label reliability is low, the unlabeled data cannot be effectively utilized, and the model generalization ability is poor. Exploring a wider perturbation space can enable the model to better adapt to changes in various scenes; by screening labels through energy scores, unlabeled data can be more fully utilized. It is very necessary to study the above application in UDA semantic segmentation. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides an unsupervised domain adaptive street view image semantic segmentation method based on pixel dual consistency and energy score, which can achieve high-precision semantic segmentation of input images, overcome the problems existing in the current strong and weak consistency framework, and can be widely used in urban smart transportation and autonomous driving systems.
[0005] The technical solution adopted by the present invention to achieve the above-mentioned purpose is:
[0006] An adaptive image semantic segmentation method based on pixel dual consistency and energy score comprises the following steps:
[0007] 1) Obtain labeled and unlabeled street view image data from the source domain and target domain respectively;
[0008] 2) Use labeled data from the source domain to train a semantic segmentation model;
[0009] 3) Perform three different image enhancement operations on the unlabeled data of the target domain, and use the semantic segmentation model to predict the three enhanced images to obtain the corresponding probability distribution and logit value;
[0010] 4) Calculate the energy score of each pixel, filter out pixels that meet the threshold condition according to the energy score threshold, and generate pseudo labels for them;
[0011] 5) Use the pseudo-label of the image after image-level weak enhancement and the prediction results of the other two enhanced images to perform consistency constraints and calculate the consistency loss;
[0012] 6) Use the labeled data of the source domain and the pseudo-labeled data of the target domain to jointly optimize the semantic segmentation model and update the model parameters;
[0013] 7) Repeat steps 3) to 6) until the model converges or reaches a preset number of iterations to obtain the final semantic segmentation model, and use the model for image segmentation.
[0014] The semantic segmentation model adopts the DeepLab-v2 structure, which includes an encoder based on ResNet-101 and a decoder based on void convolution.
[0015] The step 3) comprises the following steps:
[0016] 3.1) Perform translation, cropping, and color transformation operations on the original image to obtain an image-level weakly enhanced image;
[0017] 3.2) Perform noise injection, MixMatch, and CutMatch operations on the original image to obtain an image-level strongly enhanced image;
[0018] 3.3) Randomly discard or add noise to the output features of the encoder to obtain a feature-level enhanced image;
[0019] 3.4) Input the three enhanced images into the semantic segmentation model respectively, and output the logit value representing the relative score of each pixel belonging to each category;
[0020] 3.5) Input the logit value into the softmax function to obtain the corresponding probability distribution.
[0021] The energy fraction E(x, f(x)) is:
[0022]
[0023] Where x is the input data, f i (x) represents the corresponding logit value of the i-th category, K is the number of categories, and T is the adjustable temperature.
[0024] The consistency loss for:
[0025]
[0026] Among them, B u is the batch size of unlabeled data, ω is the image-level weak enhancement operation, Ω is the image-level strong enhancement or feature-level enhancement operation, E is the energy score, τe is the energy score threshold, is the cross entropy loss, p is the probability distribution of strong or feature-level augmentation predictions, is the probability distribution of weak enhancement prediction, is the indicator function, f is the segmentation model, and y is the true value label.
[0027] The step 6) comprises the following steps:
[0028] 6.1) Use cross entropy loss as the supervised loss Ls of the source domain:
[0029]
[0030] Among them, B s is the batch size of the source domain data, x s is the input image of the source domain, y is the true value label of the source domain, p is the predicted probability distribution, o s is the number of samples in the batch, H is the cross entropy loss;
[0031] 6.2) Using consistency loss as the unsupervised loss of the target domain;
[0032] 6.3) Use the joint optimization loss as the total loss function L:
[0033]
[0034] Among them, is the consistency loss of the K-th enhancement method of the target domain, λ k is a hyperparameter;
[0035] 6.4) Update the model parameters using stochastic gradient descent to minimize the joint optimization loss.
[0036] An adaptive image semantic segmentation system based on pixel dual consistency and energy score, comprising:
[0037] An image acquisition module, used to acquire labeled and unlabeled street view image data from a source domain and a target domain respectively;
[0038] A semantic segmentation model training module is used to train a semantic segmentation model using labeled data from the source domain;
[0039] The image enhancement module is used to perform three different image enhancement operations on the unlabeled data of the target domain, and use the semantic segmentation model to predict the three enhanced images to obtain the corresponding probability distribution and logit value;
[0040] The energy score calculation module is used to calculate the energy score of each pixel, filter out pixels that meet the threshold condition according to the energy score threshold, and generate pseudo labels for them;
[0041] A consistency loss calculation module is used to use the pseudo-label of the image after image-level weak enhancement and the prediction results of the other two enhanced images to perform consistency constraints and calculate the consistency loss;
[0042] The semantic segmentation model update module is used to jointly optimize the semantic segmentation model using the labeled data of the source domain and the pseudo-label data of the target domain and update the model parameters.
[0043] An adaptive image semantic segmentation device based on pixel dual consistency and energy score comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement the adaptive image semantic segmentation method based on pixel dual consistency and energy score when executing the computer program.
[0044] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, an adaptive image semantic segmentation method based on pixel dual consistency and energy score is implemented.
[0045] The present invention has the following beneficial effects and advantages:
[0046] 1. This paper proposes an unsupervised domain adaptive semantic segmentation method based on pixel dual consistency. This method expands the perturbation space and improves the prediction reliability of the model in the target domain by performing different perturbations at the image level and feature level.
[0047] 2. The present invention proposes a pseudo-label generation strategy based on energy scores, which calculates the energy scores of unlabeled samples, evaluates their closeness to the current training distribution, selects samples with lower energy scores to generate pseudo-labels, and improves the overall reliability of pseudo-labels. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a schematic diagram of the overall framework of the method of the present invention;
[0049] Figure 2 is a schematic diagram of target domain branches of the method of the present invention;
[0050] Figure 3 is a schematic diagram of a dual consistency framework of the method of the present invention;
[0051] Figure 4 Schematic diagram of energy fractions of the method of the present invention. DETAILED DESCRIPTION
[0052] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0053] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] An unsupervised domain adaptive street scene semantic segmentation method is provided. The PixEL unsupervised domain adaptive network model is established. First, the segmentation network deeplabv2_multi is trained using a source domain dataset to obtain a pre-trained segmentation model. In the target domain, the features of the image obtained by passing through the deeplabv2_multi encoder are perturbed to obtain a feature perturbation enhanced version of the image. The segmentation model is used to perform consistency training on the strong enhancement version, weak enhancement version, and feature perturbation version of the target domain image. In the target domain, the energy scores of unlabeled samples are calculated to evaluate their closeness to the current training distribution, and samples with higher energy scores are selected to generate pseudo labels, and the pseudo labels are used to further optimize the model training. Compared with existing unsupervised domain adaptive methods, the present invention achieves better semantic segmentation results on synthesized real benchmark datasets, with significant improvements both in subjective vision and objective evaluation indicators, effectively solving the unsupervised domain adaptation problem in the field of semantic segmentation of street scene images.
[0055] The method of the present invention mainly includes two parts: source domain branch and target domain branch. Figure 1 As shown in Figure 2. The source domain branch uses the labeled data of the source domain to train an initial semantic segmentation model. The model adopts the DeepLab-v2 structure, including an encoder based on ResNet-101 and a decoder based on dilated convolution. The target domain branch uses the unlabeled data of the target domain for unsupervised domain adaptation. This branch includes two main innovations: one is the consistency framework based on pixel dual consistency, and the other is the pseudo-label generation strategy based on energy score.
[0056] like Figure 2 As shown in the figure, the specific process of the target domain branch is as follows:
[0057] Step 1: For each image in the target domain, we generate three image enhancement versions, namely image-level weak enhancement, image-level strong enhancement, and feature-level enhancement. Image-level weak enhancement is to perform a small degree of perturbation or enhancement operation on the original image, such as translation, cropping, color change, etc. These operations will not change the semantic information of the image. Image-level strong enhancement is to perform a large degree of perturbation or enhancement operation on the original image, such as noise injection, MixMatch, CutMatch, etc. These operations will change the feature expression of the image and increase the diversity of the target domain image. Feature-level enhancement is to perform operations such as random discarding or adding noise on the output features of the encoder. These operations will increase the difference and diversity between features and expand the perturbation space of the features.
[0058] Step 2: Use the semantic segmentation model to predict the three enhanced images and obtain the corresponding probability distribution and logit value. The probability distribution is obtained by applying the softmax function to the logit value, which indicates the probability that each pixel belongs to each category. The logit value is the original output of the model, indicating the relative score of each pixel belonging to each category.
[0059] Step 3: Calculate the energy score of each pixel, filter out pixels close to the current training distribution according to the energy score threshold, and generate pseudo labels for them. The energy score is a metric that measures the distance between an unlabeled sample and the data distribution. It is based on the probability density function and reflects the probability density of a sample in the model. The lower the energy score, the more likely the sample belongs to the data distribution covered by the current model; conversely, it is more likely to be an abnormal sample. The energy score is defined as:
[0060]
[0061] Where x is the input data, f i (x) represents the corresponding logit value of the i-th category, K is the number of categories, and T is the adjustable temperature. Figure 4 As shown in Figure 1, we can set an energy score threshold τe based on the distribution of energy scores, and regard pixels with energy scores lower than the threshold as pixels within the distribution, and generate pseudo labels for them. The pseudo label is obtained by taking the maximum value of the probability distribution, which represents the most likely category of each pixel. In this way, we can filter out samples that are close enough to the current training data and assign pseudo labels, making the model update more robust and avoiding the instability that may be caused by the softmax confidence score.
[0062] Step 4: Use the pseudo-label of the image after image-level weak enhancement and the prediction results of the other two enhanced images to perform consistency constraints and calculate the consistency loss. The purpose of the consistency constraint is to require the model to maintain consistent output for the same target domain image under different perturbations. This step is to allow the model to learn to maintain robustness under different perturbations. Figure 3 As shown in Figure 2, we adopt a dual consistency framework, that is, consistency constraints are performed at both the image level and the feature level. We use cross entropy loss as the consistency loss, defined as:
[0063]
[0064] Where Bu is the batch size of unlabeled data, ω is the image-level weak enhancement operation, Ω is the image-level strong enhancement or feature-level enhancement operation, E is the energy score, and τ e is the energy score threshold, is the cross entropy loss, p is the probability distribution of strong or feature-level augmentation predictions, is the probability distribution of weak enhancement prediction. We only constrain the consistency of pixels whose energy scores are lower than the threshold, which can avoid interference with abnormal samples and improve the quality of consistency.
[0065] Step 5: The segmentation model f can be decomposed into an encoder g and a decoder h. In addition to the p mentioned in the FixMatch framework w and p s In addition, we also added the characteristic perturbation p fp , p fp Obtained by the following formula:
[0066] e w =g(x w ),
[0067]
[0068] Among them, e w For x w The extracted features, It is feature perturbation, which can be dropout or adding noise.
[0069] In summary, the pixel strength consistency constraint framework is as follows Figure 4 As shown, each unlabeled mini-batch maintains three image enhancement methods: (1) Image-level weak enhancement: x w ->f->p w (2) Image-level strong enhancement: x s ->f->p s .
[0070] (3) Feature-level enhancement:
[0071] Finally, the loss of pixel strength and weakness consistency constraints can be simply summarized as follows:
[0072]
[0073] Image-level enhancement and feature-level enhancement each have their own properties and advantages. By combining feature-level perturbations with image-level perturbations, a more effective consistency framework is constructed.
[0074] Step 6: Use the labeled data of the source domain and the pseudo-labeled data of the target domain to jointly optimize the semantic segmentation model and update the model parameters. The purpose of joint optimization is to make the model achieve good performance in both the source domain and the target domain while reducing the difference between the two domains. We use cross entropy loss as the supervised loss of the source domain, which is defined as:
[0075]
[0076] Among them, B s is the batch size of the source domain data, x s is the input image of the source domain, y is the true value label of the source domain, and p is the predicted probability distribution. We use the consistency loss as the unsupervised loss of the target domain, which is defined as:
[0077]
[0078] The meanings of the symbols are the same as above. We use the joint optimization loss as the total loss function, which is defined as:
[0079]
[0080] Among them, L s is the supervised cross entropy loss of the source domain, L uk is the consistency loss of the Kth enhancement method in the target domain, λ k is a hyperparameter that controls the relative weight of the source domain and target domain losses. We use stochastic gradient descent to update the model parameters to minimize the joint optimization loss and improve the performance of the model on both the source and target domains;
[0081] Step 7: Repeat the above steps until the model converges or reaches the preset number of iterations.
[0082] The present invention is oriented to real street scene images. GTA-to-Cityscapes is a widely recognized large-scale UDA segmentation benchmark. It uses the Cityscapes street scene dataset as the target domain. The dataset contains 2975 training images and 500 verification images, and the resolution of each image is 2048*1024. The GTA dataset is used as the source domain. The dataset contains 24966 synthetic images with a resolution of 1914*1052. In addition, the Synthia dataset can also be used, which contains 9400 synthetic images with a resolution of 1280*760.
[0083] This experiment uses the DeepLab-v2 segmentation model with a ResNet-101 backbone. The training batch size is 1, the number of categories is 19, the number of iterations is 150000, and the learning strategy uses the stochastic gradient descent SGD algorithm. The momentum is set to 0.9, the weight decay is 0.0005, and the polynomial learning decay rate is: All experiments are performed on a single NVIDIA RTX 2080TI GPU with 11GB VRAM.
[0084] The experimental results of the present invention are shown in Table 1:
[0085] Table 1
[0086]
Claims
1. An adaptive image semantic segmentation method based on pixel dual consistency and energy score, characterized in that: The following steps are involved: 1) Obtain labeled and unlabeled street view image data from the source domain and target domain respectively; 2) Use labeled data from the source domain to train a semantic segmentation model; 3) Perform three different image enhancement operations on the unlabeled data of the target domain, and use the semantic segmentation model to predict the three enhanced images to obtain the corresponding probability distribution and logit value; 4) Calculate the energy score of each pixel, filter out pixels that meet the threshold condition according to the energy score threshold, and generate pseudo labels for them; 5) Use the pseudo-label of the image after image-level weak enhancement and the prediction results of the other two enhanced images to perform consistency constraints and calculate the consistency loss; 6) Use the labeled data of the source domain and the pseudo-labeled data of the target domain to jointly optimize the semantic segmentation model and update the model parameters; 7) Repeat steps 3) to 6) until the model converges or reaches a preset number of iterations to obtain the final semantic segmentation model, and use the model for image segmentation.
2. The adaptive image semantic segmentation method based on pixel dual consistency and energy score according to claim 1, characterized in that: The semantic segmentation model adopts the DeepLab-v2 structure, which includes an encoder based on ResNet-101 and a decoder based on void convolution.
3. The adaptive image semantic segmentation method based on pixel dual consistency and energy score according to claim 1, characterized in that: The step 3) comprises the following steps: 3.1) Perform translation, cropping, and color transformation operations on the original image to obtain an image-level weakly enhanced image; 3.2) Perform noise injection, MixMatch, and CutMatch operations on the original image to obtain an image-level strongly enhanced image; 3.3) Randomly discard or add noise to the output features of the encoder to obtain a feature-level enhanced image; 3.4) Input the three enhanced images into the semantic segmentation model respectively, and output the logit value representing the relative score of each pixel belonging to each category; 3.5) Input the logit value into the softmax function to obtain the corresponding probability distribution.
4. The adaptive image semantic segmentation method based on pixel dual consistency and energy score according to claim 1, characterized in that: The energy fraction E(x, f(x)) is: Where x is the input data, f i (x) represents the corresponding logit value of the i-th category, K is the number of categories, and T is the adjustable temperature.
5. The adaptive image semantic segmentation method based on pixel dual consistency and energy score according to claim 1, characterized in that: The consistency loss for: Among them, B u is the batch size of unlabeled data, ω is the image-level weak enhancement operation, Ω is the image-level strong enhancement or feature-level enhancement operation, E is the energy score, τ e is the energy score threshold, is the cross entropy loss, p is the probability distribution of strong or feature-level augmentation predictions, is the probability distribution of weak enhancement prediction, is the indicator function, f is the segmentation model, and y is the true value label.
6. The adaptive image semantic segmentation method based on pixel dual consistency and energy score according to claim 1, characterized in that: Step 6) The following steps are involved: 6.1) Use cross entropy loss as the supervised loss L of the source domain s : Among them, b s is the batch size of the source domain data, x s is the input image of the source domain, y is the true value label of the source domain, p is the predicted probability distribution, o s is the number of samples in the batch, H is the cross entropy loss; 6.2) Using consistency loss as the unsupervised loss of the target domain; 6.3) Use the joint optimization loss as the total loss function L: Among them, is the consistency loss of the K-th enhancement method of the target domain, λ k is a hyperparameter; 6.4) Update the model parameters using stochastic gradient descent to minimize the joint optimization loss.
7. An adaptive image semantic segmentation system based on pixel dual consistency and energy score, characterized in that: include: An image acquisition module, used to acquire labeled and unlabeled street view image data from a source domain and a target domain respectively; A semantic segmentation model training module is used to train a semantic segmentation model using labeled data from the source domain; The image enhancement module is used to perform three different image enhancement operations on the unlabeled data of the target domain, and use the semantic segmentation model to predict the three enhanced images to obtain the corresponding probability distribution and logit value; The energy score calculation module is used to calculate the energy score of each pixel, filter out pixels that meet the threshold condition according to the energy score threshold, and generate pseudo labels for them; A consistency loss calculation module is used to use the pseudo-label of the image after image-level weak enhancement and the prediction results of the other two enhanced images to perform consistency constraints and calculate the consistency loss; The semantic segmentation model update module is used to jointly optimize the semantic segmentation model using the labeled data of the source domain and the pseudo-label data of the target domain and update the model parameters.
8. An adaptive image semantic segmentation device based on pixel dual consistency and energy score, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement an adaptive image semantic segmentation method based on pixel dual consistency and energy score as described in any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, an adaptive image semantic segmentation method based on pixel dual consistency and energy score as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Unsupervised pedestrian re-identification method for enhancing sample data
CN111832511A
Semi-supervised learning image classification method based on group representation features
CN113408652A
Unsupervised medical image segmentation method and system based on active contour model
CN113643302A
Vehicle damage detection model training method and vehicle damage identification method
CN113947571A
Domain adaptive identification method for unmanned aerial vehicle aerial image in open scene
CN114220016A
Cited By
Skin disease image classification method based on unsupervised domain adaptation
CN121147625A
A skin disease image classification method based on unsupervised domain adaptation
CN121147625B