A Day-Night Domain Adaptive Semantic Segmentation Method Based on Generative Adversarial Networks
Through the day-night domain adaptive semantic segmentation method based on the generative adversarial network, the daytime data set and the generative adversarial network idea are used to solve the problems of poor image segmentation performance and difficult domain migration at night, and efficient semantic segmentation in different scenarios is achieved.
Patent Information
- Application Number
- CN202211487842.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-11-24
AI Technical Summary
Existing semantic segmentation techniques perform poorly during nighttime image processing, and due to domain differences between data sets, deep learning models perform deteriorate in migration applications, making it difficult to maintain good performance in different scenarios.
Adaptive semantic segmentation method based on the generative adversarial network is adopted. By using the daytime semantic segmentation data set and the generation adversarial network idea, two semantic segmentation branches are designed to predict daytime and night images respectively. The network parameters are optimized through the target loss function and backpropagation to achieve effective semantic segmentation of night images.
It significantly improves the semantic segmentation performance of night images, effectively reduces the dependence on labeled data, and enhances the model's migration ability in different data distribution scenarios.
Smart Images

Figure CN116258849B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image and video processing, and computer vision, and particularly relates to a day-night domain adaptive semantic segmentation method based on a generative adversarial network. Background Art
[0002] Image semantic segmentation is one of the core tasks in the field of computer vision. In semantic segmentation, we need to assign semantic labels to pixel points of people, cars, roads, etc. in the input image. For example, assign a semantic-level label to each car pixel, that is, distinguish all car pixels and mark these pixels with the same color to represent the car category. With the progress of technology, the importance of semantic segmentation for scene understanding has become increasingly prominent. Currently, the application fields of semantic segmentation mainly include autonomous driving technology, geographic information systems, and medical image detection and analysis, etc. Among them, image semantic segmentation is the core algorithm technology for autonomous vehicles. After the on-vehicle camera or lidar detects an image, it is input into the neural network, and the background computer can automatically segment and classify the image to avoid obstacles such as pedestrians and vehicles.
[0003] With the wide application of deep learning and the in-depth research on semantic segmentation by researchers, the performance of semantic segmentation has been greatly improved. However, the success of current semantic segmentation technology highly depends on large-scale densely labeled datasets. Collecting these datasets in reality is very expensive and time-consuming. It takes 1.5 hours to perform semantic annotation on just one image in the commonly used semantic segmentation dataset Cityscapes. This makes it difficult, and sometimes even impossible, to collect a large number of labeled images for semantic segmentation. It is also very difficult for a deep learning model trained on a specific dataset to be well transferred to a new scenario with a different data distribution. This leads to the situation that even if researchers spend a lot of time and effort to establish a labeled visual task dataset, the deep learning model trained on this dataset can only perform well on this dataset. Once this model is applied to other datasets with different data distributions, its performance may drop significantly. This also limits to a certain extent the development of semantic segmentation task analysis and processing.
[0004] Because the differences in style and data distribution between different datasets generally exist between the actual application scenario and a dataset with sufficient label data. People have begun to explore methods that can reduce the impact of domain differences on the model, so that the trained model can also obtain good results on another dataset with no or only a small amount of labels after being trained on one dataset, that is, the domain adaptation method. Domain adaptation aims to study how to transfer the knowledge learned by the model from the source domain to the target domain to reduce the impact of domain differences and improve the generalization performance of the model.
[0005] Domain shift in semantic segmentation tasks mainly comes from two aspects: the difference in overall appearance between different domains and the difference in the distribution of different semantic categories between different domains (or even different semantic categories). The combination of these two factors results in inconsistent marginal distributions in the feature spaces of different domains. The generative adversarial idea is a method applied to domain - adaptive semantic segmentation to reduce domain differences.
[0006] With the progress of deep learning and computing power, the performance of semantic segmentation of natural - scene images taken during the day has been significantly improved recently. However, in various degraded situations, such as segmenting images taken in fog or at night, it still remains a great challenge. To improve the semantic segmentation effect of night - time images, this discovery proposes a day - night domain - adaptive semantic segmentation method based on generative adversarial networks. This method uses a limited daytime semantic segmentation dataset, combines the generative adversarial network idea and domain - adaptive methods, and significantly improves the performance of semantic segmentation of night - time images. Summary of the Invention
[0007] The object of the present invention is to provide a day - night domain - adaptive semantic segmentation method based on generative adversarial networks. This method can achieve better semantic segmentation performance for night - time images by using a daytime semantic segmentation dataset and combining the generative adversarial network idea and domain - adaptive methods.
[0008] To achieve the above object, the technical solution of the present invention is: a day - night domain - adaptive semantic segmentation method based on generative adversarial networks, including the following steps:
[0009] Step A: Pre - process the dataset to be trained, including source - domain images and paired daytime and night - time images in the target domain, including data augmentation and normalization processing;
[0010] Step B: Design two semantic segmentation branches based on generative adversarial networks. One branch A is used to predict source - domain images and target - domain daytime images to obtain corresponding semantic segmentation prediction results \(P_{s}^{d}\) 1 and \(P_{t}^{d}\) 2 , and one branch B is used to predict source - domain images and target - domain night - time images to obtain corresponding semantic segmentation prediction results \(P_{s}^{n}\) 3 and \(P_{t}^{n}\) 4 ;
[0011] Step C: According to the designed objective loss function loss, use the back - propagation method to calculate the gradients of each parameter in the day - night domain - adaptive semantic segmentation network, and use the stochastic gradient descent method to update the parameters to learn the optimal parameters of the day - night domain - adaptive semantic segmentation network;
[0012] Step D: Input the night - time image to be tested into the branch B network with optimized parameters in Step C to obtain the corresponding semantic segmentation prediction result.
[0013] In an embodiment of the present invention, in step A, the dataset to be trained is preprocessed, including data augmentation and normalization, specifically including the following steps:
[0014] Step A1: Divide the dataset to be trained into a source domain training set, a target domain training set, a target domain validation set, and a test set according to data styles and requirements;
[0015] Step A2: Randomly crop the images in the source domain training dataset at a ratio of 0.5 - 1.0, and the size of the cropped area is 512×512; then horizontally flip all the cropped images at a ratio of 0.5 to achieve data augmentation; perform normalization processing on the augmented images, with the mean of normalization being [0.485, 0.456, 0.406] and the standard deviation being [0.229, 0.224, 0.225]; each image in the source domain has a corresponding label, and the image labels also perform the same data augmentation operation;
[0016] Step A3: Randomly crop the images in the target domain training dataset at a ratio of 0.9 - 1.0, and the size of the cropped area is 960×960; then horizontally flip all the cropped images at a ratio of 0.5 to achieve data augmentation; perform normalization processing on the augmented images, with the mean of normalization being [0.485, 0.456, 0.406] and the standard deviation being [0.229, 0.224, 0.225]; paired target domain day images and night images perform the same data augmentation and normalization processing;
[0017] Step A4: For the images in the target domain validation dataset and the test set, first perform a Resize operation of 960×540, and then perform normalization processing on the Resized images, with the mean of normalization being [0.485, 0.456, 0.406] and the standard deviation being [0.229, 0.224, 0.225].
[0018] In an embodiment of the present invention, in step B, design two semantic segmentation branches A and B based on the generative adversarial network, specifically including the following steps:
[0019] Step B1: Design a domain - adaptive semantic segmentation branch A based on the generative adversarial network, using a semantic segmentation backbone network including DeepLab - v2, RefineNet, and PSPNet as the generator of the GAN, and the discriminator D d Consists of 5 convolutional layers, and the number of output channels of each layer is {64, 128, 256, 256, 1} respectively, and the size of the convolutional kernel is 4×4; the stride of the first two convolutional layers is 2, and the stride of the remaining convolutional layers is 1;
[0020] Step B2: Design the domain - adaptive semantic segmentation branch B based on the generative adversarial network, using the low - light enhancement network Zero - DCE and semantic segmentation backbone networks including DeepLab - v2, RefineNet, and PSPNet as the generator of the GAN, and the discriminator structure D n is the same as the discriminator D in branch A d .
[0021] In an embodiment of the present invention, in step B1, designing the domain - adaptive semantic segmentation branch A based on the generative adversarial network specifically includes the following steps:
[0022] Step B11: Assume that the source - domain image is S, the label image corresponding to S is GT, and the target - domain daytime image is T d ; Input S and T d into the backbone network of branch A to obtain the corresponding semantic segmentation prediction results, and then upsample the prediction results to restore them to their respective input sizes. Let the prediction result of the source - domain image after restoration be P 1 , and the prediction result of the target - domain daytime image be P 2 ;
[0023] Step B12: Fix the discriminator to make the gradients of the parameters in the discriminator not updated, and train the generator; Use the weighted cross - entropy loss function to calculate the semantic segmentation loss L 1 between P seg1 and the corresponding label GT, and the formula is as follows:
[0024]
[0025] where N 1 is the total number of pixel values of the prediction result P 1 , represents the total number of categories, that is, the total number of channels; is the k - th channel of the prediction result P 1 from the source - domain image, ω k is the fixed weight corresponding to the k - th channel; GT (k) is the one - hot encoded label of the k - th category; ||·||1 is the L 1 norm;
[0026] The generative loss in the GAN structure of the domain - adaptive semantic segmentation branch A based on the generative adversarial network is calculated using the mean - squared loss function, and the formula is as follows:
[0027] L adv1 = M(D d (softmax(P 2 ))), f fake )
[0028] Among them, M(·) represents the squared loss function, softmax(·) represents the activation function, and f fake is a label with all-zero encoding, and the resolution size of f fake is the same as the discriminator output;
[0029] Step B13: Fix the generator to keep the gradients of the parameters in the generator unchanged and train the discriminator; the objective function of the discriminator D d is defined by the following formula:
[0030]
[0031] Among them, M(·) represents the squared loss function, softmax(·) represents the activation function, and f fake is a label with all-zero encoding, f real is a label with all-one encoding, f fake and f real have the same resolution size as the discriminator output.
[0032] In an embodiment of the present invention, in step B2, a domain adaptation semantic segmentation branch B based on a generative adversarial network is designed, which specifically includes the following steps:
[0033] Step B21: Assume that the target domain night image is T n ; First, input T n into the low-light enhancement network and train the low-light enhancement network; the loss of the low-light enhancement network includes the variation loss L tv , the exposure control loss L exp and the structural similarity loss L ssim ;
[0034] The formula for the variation loss L tv is as follows:
[0035]
[0036] Among them, N is the total number of pixels in the input image, and respectively represent the intensity gradients between adjacent pixel points along the x and y directions, T' n is the output image of the low-light enhancement network, and ||·||1 is the L 1 norm;
[0037] The formula for the exposure control loss L exp is as follows:
[0038]
[0039] Among them, is the average pooling function, and N' is The total number of pixel values, and E is the input image T n The average pixel value of, ||·|| 1 is L 1 Norm;
[0040] The structural similarity loss L ssim The formula is as follows:
[0041]
[0042] where SSIM(·) is the structural similarity index measure function, N is the total number of pixels in the input image, and ||·|| 1 is L 1 Norm;
[0043] Step B22: Input S and T' n into the semantic segmentation backbone network of the domain adaptation semantic segmentation branch B based on the generative adversarial network to obtain the corresponding semantic segmentation prediction results, and then upsample and restore the prediction results to their respective input sizes. Let the prediction result of the source domain image after restoration be P 3 , and the prediction result of the target domain night image be P 4 ;
[0044] Step B23: Fix the discriminator so that the gradients of the parameters in the discriminator do not update, and train the generator; use the weighted cross-entropy loss function to calculate the semantic segmentation loss L 3 between P seg2 and the corresponding label GT. The calculation formula is as follows:
[0045]
[0046] where N 3 is the total number of pixel values of the prediction result P 3 , represents the total number of categories, that is, the total number of channels; is the k-th channel of the prediction result P 3 from the source domain image, ω k is the custom weight corresponding to the k-th channel; GT (k) is the one-hot encoded label of the k-th category; ||·|| 1 is L 1 Norm;
[0047] The generative loss in the GAN structure of branch B is calculated using the mean squared error loss function. The formula is as follows:
[0048] L adv2 = M(D n (softmax(P 4 ))), f fake )
[0049] Among them, M(·) represents the squared loss function, softmax(·) represents the activation function, and f fake is the label with all zeros encoded, and the resolution size of f fake is the same as the discriminator output;
[0050] Step B24: Define eleven categories, namely roads, sidewalks, walls, fences, utility poles, buildings, lights, signboards, grasslands, trees, and the sky, as static categories; given only consider the channels corresponding to the static categories in P 2 and P 4 to calculate the static loss. Let C S be the number of static categories, then the corresponding static prediction results of P 2 and P 4 are respectively To improve the prediction accuracy of small targets, multiply the corresponding fixed weight for the k-th category with the corresponding channels respectively for reweighting to generate the static pseudo-label F td ;
[0051] Step B25: Use the static pseudo-label F td to calculate the static category loss for the corresponding . The calculation formula for the static category loss is as follows:
[0052]
[0053]
[0054] Among them, N 4 represents the total number of pixels in F td whose predictions are static categories, p is the likelihood probability of the static prediction category, o is the one-hot encoding of the static pseudo-label F td , c represents the static category, and j represents each position in the 3×3 region centered on i; within the 3×3 local region, encode the probability of the category c corresponding to each position j, which is the one-hot vector o(c,j); in addition, the probability prediction for the category c corresponding to the central position i of this region is Weight the probabilities of each position in this region to the predicted value of the central position i, and finally return the maximum probability value, and the category probability corresponding to the i position is p(c,i);
[0055] Step B26: Fix the generator so that the gradients of the parameters in the generator are not updated, and train the discriminator; the objective function of the discriminator D n is defined by the following formula:
[0056]
[0057] Among them, M(·) represents the squared loss function, and f fake is a label with all encodings being 0, and f real is a label with all encodings being 1, and the resolution sizes of f fake and f real are the same as the discriminator output.
[0058] In an embodiment of the present invention, in step C, according to the designed target loss function loss, the gradients of the parameters in the day-night domain adaptive semantic segmentation network are calculated by using the backpropagation method, and the parameters are updated by using the stochastic gradient descent method, which specifically includes the following steps:
[0059] Step C1: Repeat steps B11, B12, and B13, use the backpropagation method to calculate the gradients of the parameters in the domain adaptive semantic segmentation branch A based on the generative adversarial network, and use the stochastic gradient descent method to update the neural network parameters; the total loss in the domain adaptive semantic segmentation branch A based on the generative adversarial network is as follows:
[0060]
[0061] Among them, λ seg1 and λ adv1 are the coefficients of L seg1 and L adv1 respectively;
[0062] Step C2: Repeat steps B21, B22, B23, B24, B25, and B26, use the backpropagation method to calculate the gradients of the parameters in the domain adaptive semantic segmentation branch B based on the generative adversarial network, and use the stochastic gradient descent method to update the neural network parameters until the calculated loss value converges and stabilizes; after training, save the network parameters to obtain the trained domain adaptive semantic segmentation branch B based on the generative adversarial network; the joint loss in the domain adaptive semantic segmentation branch B based on the generative adversarial network is as follows:
[0063] L Light = α tv L tv + α exp L exp + α ssim L ssim
[0064]
[0065] Among them, α tv , α exp , α ssim are the coefficients of L tv , L exp and L ssimCoefficient; λ light and λ seg2 and λ adv2 and λ static are the coefficients of L Light and L seg2 and L adv2 and L static respectively.
[0066] In an embodiment of the present invention, in step D, the to-be-tested night image is input into the designed domain adaptation semantic segmentation branch B based on the generative adversarial network to obtain the corresponding semantic segmentation prediction result, which specifically includes the following steps:
[0067] Step D1: Input the night image without label information in the test set into the low-light enhancement network in the trained domain adaptation semantic segmentation branch B based on the generative adversarial network to obtain the image output by the low-light enhancement network;
[0068] Step D2: Input the low-light enhanced image into the semantic segmentation network in the domain adaptation semantic segmentation branch B based on the generative adversarial network for prediction and inference, and the obtained prediction result is the semantic segmentation prediction result of the input image.
[0069] Compared with the prior art, the present invention has the following beneficial effects: The method of the present invention utilizes two generative adversarial networks to achieve domain adaptation from the source domain to the target domain during the day and from the source domain to the target domain at night, improving the accuracy of the semantic segmentation prediction result of the target domain day image. It effectively utilizes the characteristic that there is a high similarity in static categories between the target domain day image and the night image, thereby improving the prediction performance for the unlabeled target domain night image. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is a schematic flow chart of the method of the present invention.
[0071] Figure 2 is a network structure diagram for training the semantic segmentation model of the method of the present invention.
[0072] Figure 3 is an application structure diagram of the semantic segmentation model of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The technical solutions of the present invention will be specifically described below with reference to the drawings.
[0074] The present invention provides a semantic segmentation method for day-night domain adaptation based on a generative adversarial network, as Figures 1-3 shown, including the following steps:
[0075] Step A: Preprocess the dataset to be trained, including source domain images and paired daytime and nighttime images in the target domain, including data augmentation and normalization.
[0076] Step B: Design two semantic segmentation branches based on the generative adversarial network. One branch A is used to predict source domain images and target domain daytime images to obtain corresponding semantic segmentation prediction results P 1 and P 2 , and one branch B is used to predict source domain images and target domain nighttime images to obtain corresponding semantic segmentation prediction results P 3 and P 4 ;
[0077] Step C: According to the designed objective loss function loss, use the backpropagation method to calculate the gradients of the parameters in the day-night domain adaptive semantic segmentation network, and use the stochastic gradient descent method to update the parameters to learn the optimal parameters of the day-night domain adaptive semantic segmentation network.
[0078] Step D: Input the nighttime image to be tested into the branch B network with optimized parameters in Step C to obtain the corresponding semantic segmentation prediction result.
[0079] In this embodiment, in Step A, preprocess the dataset to be trained, including data augmentation and normalization, specifically including the following steps:
[0080] Step A1: Divide the dataset to be trained into a source domain training set, a target domain training set, a target domain validation set, and a test set according to data styles and requirements.
[0081] Step A2: Randomly crop the images in the source domain training dataset at a ratio of 0.5 - 1.0, and the size of the cropped area is 512×512; then horizontally flip all the cropped images with a flipping ratio of 0.5 to achieve data augmentation; perform normalization processing on the augmented images, with a mean of [0.485, 0.456, 0.406] and a standard deviation of [0.229, 0.224, 0.225]; each image in the source domain has a corresponding label, and the image labels also perform the same data augmentation operation.
[0082] Step A3: Randomly crop the images in the target domain training dataset at a ratio of 0.9 - 1.0, and the size of the cropped area is 960×960; then horizontally flip all the cropped images with a flipping ratio of 0.5 to achieve data augmentation; perform normalization processing on the augmented images, with a mean of [0.485, 0.456, 0.406] and a standard deviation of [0.229, 0.224, 0.225]; perform the same data augmentation and normalization processing on the paired target domain daytime and nighttime images.
[0083] Step A4: For the images in the target domain validation dataset and the test set, first perform a Resize operation to 960×540, and then normalize the Resize - after images. The mean for normalization is [0.485, 0.456, 0.406], and the standard deviation is [0.229, 0.224, 0.225].
[0084] In this embodiment, in step B, design two semantic segmentation branches A and B based on the generative adversarial network, which specifically include the following steps:
[0085] Step B1: Design a domain - adaptive semantic segmentation branch A based on the generative adversarial network. Use the semantic segmentation backbone networks including DeepLab - v2, RefineNet, and PSPNet as the generator of the GAN, and the discriminator D d consists of 5 convolutional layers. The number of output channels for each layer is {64, 128, 256, 256, 1} respectively, and the kernel size of the convolutional layer is 4×4; the stride of the first two convolutional layers is 2, and the stride of the remaining convolutional layers is 1;
[0086] Step B2: Design a domain - adaptive semantic segmentation branch B based on the generative adversarial network. Use the low - light enhancement network Zero - DCE and the semantic segmentation backbone networks including DeepLab - v2, RefineNet, and PSPNet as the generator of the GAN, and the discriminator structure D n and the discriminator D in branch A d is the same.
[0087] In this embodiment, in step B1, design a domain - adaptive semantic segmentation branch A based on the generative adversarial network, which specifically includes the following steps:
[0088] Step B11: Assume that the source - domain image is S, the label image corresponding to S is GT, and the target - domain daytime image is T d ; Input S and T d into the backbone network of branch A to obtain the corresponding semantic segmentation prediction results, and then upsample the prediction results to restore them to their respective input sizes. Let the prediction result of the source - domain image after restoration be P 1 , and the prediction result of the target - domain daytime image be P 2 ;
[0089] Step B12: Fix the discriminator to make the gradients of the parameters in the discriminator not updated, and train the generator; use the weighted cross - entropy loss function to calculate the semantic segmentation loss L between p 1 and the corresponding label GT seg1 , and the formula is as follows:
[0090]
[0091] Among them, N 1 is the total number of pixel values of the prediction result P 1 . represents the total number of categories, that is, the total number of channels; is the k-th channel of the prediction result P 1 from the source domain image, ω k is the fixed weight corresponding to the k-th channel; GT (k) is the one-hot encoded label of the k-th category; ||·|| 1 is the L 1 norm;
[0092] The generation loss in the GAN structure of the domain adaptation semantic segmentation branch A based on the generative adversarial network is calculated using the mean squared error loss function, and the formula is as follows:
[0093] L adv1 = M(D d (softmax(P 2 ))), f fake )
[0094] Among them, M(·) represents the mean squared error loss function, softmax(·) represents the activation function, and f fake is a label with all encodings being 0, and the resolution size of f fake is the same as the discriminator output;
[0095] Step B13: Fix the generator so that the gradients of the parameters in the generator are not updated, and train the discriminator; the objective function of the discriminator D d is defined by the following formula:
[0096]
[0097] Among them, M(·) represents the mean squared error loss function, softmax(·) represents the activation function, and f real is a label with all encodings being 1, and the resolution sizes of f fake and f real are the same as the discriminator output.
[0098] In this embodiment, in step B2, a domain adaptation semantic segmentation branch B based on the generative adversarial network is designed, which specifically includes the following steps:
[0099] Step B21: Assume that the target domain night image is T n ; First, input T n into the low-light enhancement network and train the low-light enhancement network; the loss of the low-light enhancement network includes the variation loss L tv , the exposure control loss L expAnd the structural similarity loss L ssim ;
[0100] The variation loss L tv has the following formula:
[0101]
[0102] where N is the total number of pixels in the input image, and respectively represent the intensity gradients between adjacent pixel points along the x and y directions, T' n is the output image of the low-light enhancement network, ||·|| 1 is the L 1 norm;
[0103] The exposure control loss L exp has the following formula:
[0104]
[0105] where, is the average pooling function, N' is the total number of pixel values of , E is the average pixel value of the input image T n , ||·|| 1 is the L 1 norm;
[0106] The structural similarity loss L ssim has the following formula:
[0107]
[0108] where SSIM(·) is the structural similarity index metric function, N is the total number of pixels in the input image, ||·|| 1 is the L 1 norm;
[0109] Step B22: Input S and T' n into the semantic segmentation backbone network of the domain adaptation semantic segmentation branch B based on the generative adversarial network to obtain the corresponding semantic segmentation prediction results, and then upsample and restore the prediction results to their respective input sizes. Let the prediction result of the source domain image after restoration be P 3 , and the prediction result of the target domain night image be P 4 ;
[0110] Step B23: Fix the discriminator to make the gradients of the parameters in the discriminator not updated, and train the generator; use the weighted cross-entropy loss function to calculate the semantic segmentation loss L 3 between P seg2 and the corresponding label GT. The calculation formula is as follows:
[0111]
[0112] Among them, N 3 is the total number of pixel values of the prediction result P 3 . represents the total number of categories, that is, the total number of channels; is the k-th channel of the prediction result P of the source domain image 3 , ω k is the custom weight corresponding to the k-th channel; GT (k) is the one-hot encoded label of the k-th category; ||·|| 1 is the L 1 norm;
[0113] The generation loss in the GAN structure of branch B is calculated using the mean squared error loss function, and the formula is as follows:
[0114] L adv2 = M(D n (softmax(P 4 ))), f fake )
[0115] Among them, M(·) represents the mean squared error loss function, softmax(·) represents the activation function, and f fake is a label with all encoded values being 0, and the resolution size of f fake is the same as the discriminator output;
[0116] Step B24: Define the eleven categories of road, sidewalk, wall, fence, utility pole, building, lamp, sign, grassland, tree, and sky as static categories; Given only consider the channels corresponding to the static categories in P 2 , P 4 to calculate the static loss. Let C S be the number of static categories, then the corresponding P 2 , P 4 The corresponding static prediction results are respectively To improve the prediction accuracy of small targets, multiply the corresponding channel of the fixed weight corresponding to the k-th category by respectively for reweighting to generate the static pseudo-label F td ;
[0117] Step B25: Use the static pseudo-label F td to calculate the static category loss for the corresponding . The calculation formula for the static category loss is as follows:
[0118]
[0119]
[0120] Among them, N 4 represents the total number of pixels with static-class predictions in F td , p is the likelihood probability of the static prediction class, o is the one-hot encoding of the static pseudo-label F td , c represents the static class, and j represents each position in the 3×3 region centered on i; within the 3×3 local region, the probability of the class c corresponding to each position j is encoded, which is the one-hot vector o(c,j); in addition, the probability prediction of the class c corresponding to the central position i of this region is The probability of each position in this region is weighted to the predicted value of the central position i, and finally the maximum probability value is returned, and the class probability corresponding to the i position is p(c,i);
[0121] Step B26: Fix the generator so that the gradients of the parameters in the generator are not updated, and train the discriminator; the discriminator D n 's objective function is defined by the following formula:
[0122]
[0123] Among them, M(·) represents the square loss function, softmax(·) represents the activation function, f fake is a label with all zeros encoded, f real is a label with all ones encoded, f fake and f real have the same resolution size as the discriminator output.
[0124] In this embodiment, in step C, according to the designed objective loss function loss, the gradients of the parameters in the day-night domain adaptive semantic segmentation network are calculated using the backpropagation method, and the parameters are updated using the stochastic gradient descent method, which specifically includes the following steps:
[0125] Step C1: Repeat steps B11, B12, and B13, use the backpropagation method to calculate the gradients of the parameters in the domain adaptive semantic segmentation branch A based on the generative adversarial network, and use the stochastic gradient descent method to update the neural network parameters; the total loss in the domain adaptive semantic segmentation branch A based on the generative adversarial network is as follows:
[0126]
[0127] Among them, λ seg1 and λ adv1 are the coefficients of L seg1 and L adv1 respectively;
[0128] Step C2. Repeat steps B21, B22, B23, B24, B25, and B26. Use the backpropagation method to calculate the gradients of the parameters in the domain adaptation semantic segmentation branch B based on the generative adversarial network, and use the stochastic gradient descent method to update the neural network parameters until the calculated loss value converges and stabilizes. After training, save the network parameters to obtain the trained domain adaptation semantic segmentation branch B based on the generative adversarial network. The combined loss in the domain adaptation semantic segmentation branch B based on the generative adversarial network is as follows:
[0129] L Light = α tv L tv + α exp L exp + α ssim L ssim
[0130]
[0131] where α tv 、α exp 、α ssim are the coefficients of L tv 、L exp and L ssim respectively; λ light 、λ seg2 、λ adv2 、λ static are the coefficients of L Light 、L seg2 、L adv2 、L static respectively.
[0132] In this embodiment, in step D, the to-be-tested night image is input into the designed domain adaptation semantic segmentation branch B based on the generative adversarial network to obtain the corresponding semantic segmentation prediction result, which specifically includes the following steps:
[0133] Step D1. Input the night image without label information in the test set into the low-light enhancement network in the trained domain adaptation semantic segmentation branch B based on the generative adversarial network to obtain the image output by the low-light enhancement network;
[0134] Step D2. Input the low-light enhanced image into the semantic segmentation network in the domain adaptation semantic segmentation branch B based on the generative adversarial network for prediction and inference, and the obtained prediction result is the semantic segmentation prediction result of the input image.
[0135] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects generated do not exceed the scope of the technical solutions of the present invention, fall within the protection scope of the present invention.
Claims
1. A method for day-night domain adaptive semantic segmentation based on generative adversarial networks, characterized in that, it includes the following steps: Step A: Preprocess the dataset to be trained, including source domain images and paired day and night images in the target domain, including data augmentation and normalization; Step B: Design two semantic segmentation branches based on the generative adversarial network. One branch A is used to predict the source domain image and the target domain daytime image to obtain the corresponding semantic segmentation prediction results P 1 and P 2 , and one branch B is used to predict the source domain image and the target domain nighttime image to obtain the corresponding semantic segmentation prediction results P 3 and P 4 ; Step C: According to the designed target loss function loss, use the backpropagation method to calculate the gradients of each parameter in the day-night domain adaptive semantic segmentation network, and use the stochastic gradient descent method to update the parameters to learn the optimal parameters of the day-night domain adaptive semantic segmentation network; Step D: Input the night image to be measured into the branch B network with optimized parameters in Step C to obtain the corresponding semantic segmentation prediction result; In Step B, design two semantic segmentation branches A and B based on generative adversarial networks, specifically including the following steps: Step B1: Design a domain adaptation semantic segmentation branch A based on a generative adversarial network, using a semantic segmentation backbone network including DeepLab-v2, RefineNet, and PSPNet as the generator of the GAN, and a discriminator D d It consists of 5 convolutional layers. The number of output channels for each layer is {64, 128, 256, 256, 1} respectively, and the convolutional kernel size is 4×4; the stride of the first two convolutional layers is 2, and the stride of the remaining convolutional layers is 1; Step B2: Design a domain adaptation semantic segmentation branch B based on a generative adversarial network, using the low-light enhancement network Zero-DCE and semantic segmentation backbone networks including DeepLab-v2, RefineNet, and PSPNet as the generator of the GAN, and a discriminator structure D n and the discriminator D in branch A d are the same; In Step B1, design a domain adaptive semantic segmentation branch A based on generative adversarial networks, specifically including the following steps: Step B11: Assume that the source domain image is S, the label image corresponding to S is GT, and the target domain daytime image is T d ; Input S and T d into the backbone network of branch A to obtain the corresponding semantic segmentation prediction results, and then upsample the prediction results back to their respective input sizes. Let the prediction result of the source domain image after restoration be P 1 , and the prediction result of the target domain daytime image be P 2 ; Step B12: Fix the discriminator so that the gradients of the parameters in the discriminator are not updated, and train the generator; use the weighted cross-entropy loss function to calculate P 1 and the semantic segmentation loss L seg1 of the corresponding label GT. The formula is as follows: Among them, N 1 is the total number of pixel values of the prediction result P 1 . represents the total number of categories, that is, the total number of channels; is the k-th channel of the prediction result P of the source domain image 1 , ω k is the fixed weight corresponding to the k-th channel; GT (k) is the one-hot encoded label of the k-th category; ||·|| 1 is the L 1 norm; The generation loss in the GAN structure of the domain adaptive semantic segmentation branch A based on generative adversarial networks is calculated using the least squares loss function, and the formula is as follows: L adv1 = M(D d (softmax(P 2 )),f fake ) Among them, M(·) represents the squared loss function, softmax(·) represents the activation function, and f fake is a label with all zeros in the encoding, and f fake has the same resolution size as the discriminator output; Step B13: Fix the generator so that the gradients of all parameters in the generator are not updated, and train the discriminator; the objective function of the discriminator D d is defined by the following formula: Among them, M(·) represents the squared loss function, softmax(·) represents the activation function, and f fake is the label with all - zero encoding, and f real is the label with all - one encoding, and the resolution sizes of f fake and f real are the same as the discriminator output; In Step B2, design a domain adaptive semantic segmentation branch B based on generative adversarial networks, specifically including the following steps: Step B21. Assume that the target domain night image is T n ; First, input T n into the low-light enhancement network and train the low-light enhancement network. The loss of the low-light enhancement network includes the variation loss L tv , the exposure control loss L exp and the structural similarity loss L ssim ; Variation loss L tv The formula is as follows: where N is the total number of pixels in the input image, and represent the intensity gradients between adjacent pixel points along the x- and y-directions, respectively, T′ n is the output image of the low-light enhancement network, ||·|| 1 is the L 1 norm; Exposure control loss L exp The formula is as follows: Among them, is the average pooling function, and N' is the total number of pixel values of, and E is the pixel average value of the input image T n , and ||·|| 1 is the L 1 norm; Structural similarity loss L ssim The formula is as follows: where SSIM(·) is the structural similarity index measure function, N is the total number of pixels in the input image, and ||·|| 1 is the 1 L-norm; Step B22: Input S and T' n into the semantic segmentation backbone network of the domain adaptation semantic segmentation branch B based on the generative adversarial network to obtain the corresponding semantic segmentation prediction results, and then upsample and restore the prediction results to their respective input sizes. Let the prediction result of the source domain image after restoration be P 3 , and the prediction result of the target domain night image be P 4 ; Step B23: Fix the discriminator so that the gradients of the parameters in the discriminator are not updated, and train the generator; use the weighted cross-entropy loss function to calculate the semantic segmentation loss L of P 3 and the corresponding label GT seg2 , and the calculation formula is as follows: Among them, N 3 is the total number of pixel values of the prediction result P 3 . represents the total number of categories, that is, the total number of channels; is the k-th channel of the prediction result P of the source domain image 3 , ω k is the custom weight corresponding to the k-th channel; GT (k) is the one-hot encoded label of the k-th category; ||·|| 1 is the L 1 norm; The generation loss in the GAN structure of branch B is calculated using the least squares loss function, and the formula is as follows: L adv2 = M(D n (softmax(P 4 )),f fake ) where M(·) represents the squared loss function, softmax(·) represents the activation function, and f fake is the label with all zeros encoded, and the resolution size of f fake is the same as the discriminator output; Step B24: Define eleven categories of road, sidewalk, wall, fence, utility pole, building, lamp, signboard, grassland, tree, and sky as static categories; given Only consider P 2 , P 4 The channels corresponding to the static categories in are used to calculate the static loss. Let C S be the number of static categories, then the corresponding P 2 , P 4 The corresponding static prediction results are respectively To improve the prediction accuracy of small targets, multiply the corresponding fixed weight of the k-th category by The corresponding channels are multiplied respectively for reweighting to generate the static pseudo-label F td ; Step B25, using the static pseudo-label F td for the corresponding calculate the static class loss; the calculation formula for the static class loss is as follows: Among them, N 4 represents the total number of pixels with static categories in F td , p is the likelihood probability of the static prediction category, o is the one-hot encoding of the static pseudo-label F td , c represents the static category, and j represents each position in the 3×3 region centered on i; within the 3×3 local region, the probability of the category c corresponding to each position j is encoded, which is the one-hot vector o(c, j); in addition, the probability prediction of the category c corresponding to the central position i of this region is The probability of each position in this region is weighted to the predicted value of the central position i, and finally the maximum probability value is returned, and the category probability corresponding to the i position is p(c, i); Step B26: Fix the generator so that the gradients of all parameters in the generator are not updated, and train the discriminator; the discriminator D n has an objective function defined by the following formula: where M(·) represents the squared loss function, softmax(·) represents the activation function, f fake is the label with all - zero encoding, f real is the label with all - one encoding, f fake and f real have the same resolution size as the discriminator output; In Step C, according to the designed target loss function loss, use the backpropagation method to calculate the gradients of each parameter in the day-night domain adaptive semantic segmentation network, and use the stochastic gradient descent method to update the parameters, specifically including the following steps: Step C1: Repeat Steps B11, B12, B13, use the backpropagation method to calculate the gradients of each parameter in the domain adaptive semantic segmentation branch A based on generative adversarial networks, and use the stochastic gradient descent method to update the neural network parameters; The total loss in the domain adaptive semantic segmentation branch A based on generative adversarial networks is as follows: where λ seg1 and λ adv1 are the coefficients of L seg1 and L adv1 respectively; Step C2: Repeat Steps B21, B22, B23, B24, B25, B26, use the backpropagation method to calculate the gradients of each parameter in the domain adaptive semantic segmentation branch B based on generative adversarial networks, and use the stochastic gradient descent method to update the neural network parameters until the calculated loss value converges and stabilizes; After training, save the network parameters to obtain the trained domain adaptive semantic segmentation branch B based on generative adversarial networks; The joint loss in the domain adaptive semantic segmentation branch B based on generative adversarial networks is as follows: L Light = α tv L tv + α exp L exp + α ssim L ssim Among them, α tv 、α exp 、α ssim are the coefficients of L tv 、L exp and L ssim respectively; λ light 、λ seg2 、λ adv2 、λ static are the coefficients of L Light 、L seg2 、L adv2 、L static respectively.
2. The method for day-night domain adaptive semantic segmentation based on generative adversarial networks according to claim 1, characterized in that, in Step A, preprocess the dataset to be trained, including data augmentation and normalization, specifically including the following steps: Step A1: Divide the dataset to be trained into a source domain training set, a target domain training set, a target domain validation set, and a test set according to the data style and requirements; Step A2: Randomly crop the images in the source domain training dataset at a ratio of 0.5 - 1.0, with the cropped area size of 512×512; then horizontally flip all the cropped images with a flipping ratio of 0.5 to achieve data augmentation; perform normalization on the augmented images, with the mean of normalization being [0.485, 0.456, 0.406] and the standard deviation being [0.229, 0.224, 0.225]; each image in the source domain images has a corresponding label, and the image labels also perform the same data augmentation operation. Step A3: Randomly crop the images in the target domain training dataset at a ratio of 0.9 - 1.0, with the cropped area size of 960×960; then horizontally flip all the cropped images with a flipping ratio of 0.5 to achieve data augmentation; perform normalization on the augmented images, with the mean of normalization being [0.485, 0.456, 0.406] and the standard deviation being [0.229, 0.224, 0.225]; the paired target domain daytime images and nighttime images perform the same data augmentation and normalization processing. Step A4: For the images in the target domain validation dataset and test set, first perform a Resize operation of 960×540, and then perform normalization on the Resized images, with the mean of normalization being [0.485, 0.456, 0.406] and the standard deviation being [0.229, 0.224, 0.225].
3. A method for domain - adaptive semantic segmentation between day and night based on a generative adversarial network according to claim 1, characterized in that, in step D, input the to - be - measured nighttime image into the designed domain - adaptive semantic segmentation branch B based on a generative adversarial network to obtain the corresponding semantic segmentation prediction result, which specifically includes the following steps: Step D1: Input the nighttime image without label information in the test set into the low - light enhancement network in the trained domain - adaptive semantic segmentation branch B based on a generative adversarial network to obtain the image output by the low - light enhancement network; Step D2: Input the low - light enhanced image into the semantic segmentation network in the domain - adaptive semantic segmentation branch B based on a generative adversarial network for prediction and inference, and the obtained prediction result is the semantic segmentation prediction result of the input image.
Citation Information
Patent Citations
Recognizing fine-grained objects in surveillance camera images
US20200089966A1