A semantic segmentation method based on adversarial neural network
By using the technical means of the adversarial neural network and Resnet-50 network combined with PSPNet and DeepLab modules in the semantic segmentation method, the accuracy problem of damage area detection in road images is solved, and higher segmentation accuracy and road traffic safety are achieved.
Patent Information
- Application Number
- CN202111474990.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-06
AI Technical Summary
The prior art is difficult to detect damage areas quickly and accurately from road images, especially inadequate segmentation accuracy in small target damage.
Adopting semantic segmentation method based on adversarial neural networks, the Resnet-50 network is built as the benchmark pre-training network, and combining the PPM module in PSPNet network and the ASPP module in DeepLab, adversarial learning is performed to improve segmentation accuracy.
It realizes rapid and accurate detection of road damage, improves the segmentation accuracy of small target damage, reduces the risk of manual detection, and improves road traffic safety.
Smart Images

Figure CN114693922B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image recognition and deep learning, and specifically to a semantic segmentation method based on an adversarial neural network. Background Art
[0002] As one of the important indicators to measure the modernization level of a country, road construction in China is also developing vigorously. However, the road traffic network is extremely large and complex, and the resulting problems are difficult to solve properly. Among them, road maintenance has always been a top priority. The damage to the road is generally caused by factors such as construction methods, climatic conditions, and driving loads, which poses a huge challenge to road traffic safety. Therefore, how to quickly and accurately detect the damaged area from road images has become a difficult problem to be solved urgently. The present invention proposes a semantic segmentation method based on an adversarial neural network, which can not only detect the damage types of roads, but also significantly improve the segmentation accuracy for small target damages. Summary of the Invention
[0003] In order to perform semantic segmentation on the road surface conditions of important sections, and to detect and repair road surface damages in a timely manner if they occur, the present invention provides a semantic segmentation method based on an adversarial neural network.
[0004] The present invention adopts the following technical solutions: A semantic segmentation method based on an adversarial neural network, including the following steps.
[0005] S100: Collect a road damage image dataset, and divide it into a training dataset and a test dataset according to a ratio. The training dataset is used to train the adversarial neural network model, and the test dataset is used to verify the quality of the adversarial neural network model; S200: Construct a Resnet-50 network as a benchmark pre-training network, input the road damage image training set, and perform pre-training; S300: Train the segmentation model of the adversarial neural network, which includes the PPM module in the PSPNet network and the ASPP module in DeepLab; S400: Input the results obtained by the two PPM modules and the ASPP module segmentation networks into the discriminant model, and output a class global probability score map; S500: According to the results obtained by the discriminant network, readjust the two segmentation models for training, and repeat S300 and S400 until the results of the two segmentation networks tend to be consistent, so as to achieve the effect of semantic segmentation.
[0006] In step S100, the road damage image dataset comes from the road scene images taken, including different environments, illuminations, road surfaces, and shapes.
[0007] In step S200, the first convolutional layer of the original Resnet-50 network is changed from a 7×7 convolutional layer to three 3×3 convolutional layers, the last classification layer is removed, then the stride of the last residual module is changed from 2 to 1, and finally, the dilation rate of the network is changed to 2, obtaining a feature map equivalent to 1 / 16 of the input image size.
[0008] In step S300, in the PPM module of PSPNet, pooling is performed on different sub-regions using 1×1, 2×2, , 6×6 respectively, and finally, upsampling and concatenation are performed on this module; in the ASPP module of DeepLab, rates=(1, 6, 12, 18) is used, and the dilated convolution operation consists of 4 parallel convolutional layers; the input image size of the two segmentation networks is H×W×3, and the outputs are two class probability segmentation maps S 1 and , and C 1 and , where C 1 and C 2 are the number of classes output by the two segmentation networks respectively, and adversarial learning is performed on these two segmentation maps.
[0009] In step S300, during the specific training process, the total loss function of the entire adversarial network is divided into two parts, the loss functions of the two segmentation networks themselves and the loss function of the discriminator network. The loss functions of the two segmentation networks are expressed as the following formulas:
[0010] (1)
[0011] (2)
[0012] where L seg-s1 represents the loss function for generating labels by the segmentation network with the PPM module, L seg-s2 represents the loss function for generating labels by the segmentation network with the ASPP module, represents the cross-entropy loss function between the Deeplab-based segmentation network and the true label ; is used to represent the loss function of adversarial learning, which can deceive the discriminator as much as possible and make the discriminator misclassify the segmentation prediction result based on PPM as the prediction result of the Deeplab network; is the loss function of PPM, and this network uses the weak labels generated based on the Deeplab network , is a variable constant used to adjust the training process, and the specific formula is as follows:
[0013] (3)
[0014] (4)
[0015] (5)
[0016] (6)
[0017] Wherein, H, w, and c represent the length, width, and segmentation category of the image, s1 and s2 represent the class probability maps obtained by the two segmentation networks of the PPM module and the ASPP module, Xn represents the predicted label, and Yn represents the true label.
[0018] In step S400, the loss function of the discriminative network is expressed as the following formula:
[0019]
[0020] In the above formula is the learning rate of the discriminative network, =0 indicates that the sample comes from the segmentation network based on PPM, =1 indicates that the sample comes from the segmentation network based on ASPP.
[0021] In step S400, the discriminative network contains 5 convolutional layers, where the size of the convolutional kernel is , and the stride is 2.
[0022] In step S500, after the adversarial process in the generative network, the segmentation result of the Deeplab network is used as the base image, the segmentation result of the PSPNet network is regarded as the generated image, the two obtained images are input into the discriminative network in step S300, and the category of each target object is determined, and a class global probability score map with a size of H×W×1 is output; if the pixel in the image comes from the generated image of the Deeplab network, then P = 1, if it comes from the generated image of the PSPNet network, then P = 0. When the value of P is closer to 0.5, it indicates that the overall adversarial network has achieved the best effect, where P represents the degree of network adversarial.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] 1. The use of the adversarial neural network can automatically extract image features, avoiding the blindness of manual feature extraction.
[0025] 2. It can replace manual detection in some high-risk sections, reducing the possibility of casualties.
[0026] 3. Compared with other semantic segmentation algorithms, the method of the present invention has higher segmentation accuracy and better segmentation effect for small targets. Description of the Drawings
[0027] Figure 1 is the implementation flowchart of the method of the present invention;
[0028] Figure 2 is a partial sample example diagram of the method of the present invention;
[0029] Figure 3 is the segmentation network structure diagram of the method of the present invention;
[0030] Figure 4 is the discrimination network structure diagram of the method of the present invention;
[0031] Figure 5 is the result display diagram of the invention method. Detailed Embodiment
[0032] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0033] As Figure 1 shown, a semantic segmentation method based on an adversarial neural network, the specific steps are as follows:
[0034] S100 - The road damage image dataset comes from road scene images taken by mobile phones, including different environments, lighting, road surfaces, and shapes, and is divided into a training dataset and a test dataset in a ratio of 9:1. Some sample examples are as Figure 2 shown. The training dataset is used to train the adversarial neural network model, and the test dataset is used to verify the quality of the adversarial neural network model. In this embodiment, the dataset provided by the Global Road Damage Detection Challenge competition is used, and this training set is used to train the adversarial neural network model.
[0035] S200 - Build a Resnet-50 network as a benchmark pre-training network, input the road damage image training set, and perform pre-training; change the first convolutional layer of the original Resnet-50 network from 7×7 to three 3×3 convolutional layers, remove the last classification layer, then change the stride of the last residual module from 2 to 1, and finally, change the dilation rate of the network to 2 to obtain a feature map equivalent to 1 / 16 of the input image size.
[0036] S300 - As Figure 3 shown is the structure diagram of the segmentation network. Train the segmentation model of the adversarial neural network, which includes the PPM module in the PSPNet network and the ASPP module in DeepLab. The PPM module in PSPNet uses 1×1, 2×2, , 6×6 represents pooling for different sub-regions respectively, and finally upsampling and concatenation are performed using this module; in the ASPP module of DeepLab, rates = (1, 6, 12, 18) is used, and the dilated convolution operation consists of 4 parallel convolutional layers; the input images of the two segmentation networks are of size H×W×3, and the output is H×W×C 1 and two class probability segmentation maps S 1 and , C 1 and C 2 are the number of classes output by the two segmentation networks respectively, and these two segmentation maps are subjected to adversarial learning.
[0037] During the training process, the loss function of the segmentation network is expressed as the following formula:
[0038] (1)
[0039] (2)
[0040] Among them, L seg-s1 represents the loss function for generating labels by the segmentation network of the PPM module, and L seg-s2 represents the loss function for generating labels by the segmentation network of the ASPP module, represents the cross-entropy loss function between the DeepLab-based segmentation network and the true label ; is used to represent the loss function of adversarial learning, which can deceive the discriminator as much as possible, making the discriminator misclassify the segmentation prediction result based on the PPM as the prediction result of the DeepLab network; is the loss function of the PPM, and this network uses the weak labels generated based on the DeepLab network , is a variable constant used to adjust the training process, and the specific formula is as follows:
[0041] (3)
[0042] (4)
[0043] (5)
[0044] (6)
[0045] In the formula, H, w, c represent the length, width, and segmentation classes of the image, s1, s2 represent the class probability maps obtained by the two segmentation networks of the PPM module and the ASPP module, Xn represents the predicted label, and Yn represents the true label.
[0046] S400~as follows Figure 4 The structural diagram of the discrimination network as shown. The results obtained by the two segmentation networks are input into the discrimination model. The network structure of the discrimination model is similar to the Segnet network, and the output is the class global probability score map. The discrimination network contains 5 convolutional layers, where the size of the convolutional kernel is , and the stride is 2.
[0047] During the training process, the loss function of the discrimination network is expressed by the following formula:
[0048]
[0049] In the above formula is the learning rate of the discrimination network, =0 represents that the sample comes from the PPM-based segmentation network, =1 indicates that the sample comes from the ASPP-based segmentation network.
[0050] S500. According to the results obtained by the discrimination network, readjust the two segmentation models for training, and repeat the above S300 and S400 steps until the results of the two segmentation networks tend to be consistent, making up for the deficiencies of their respective networks, so as to achieve the effect of semantic segmentation.
[0051] The experimental environment of the present invention is an intel COREi7 under windows10 (64-bit), the graphics card is Nvidia GeForceGTX 1660 Ti, the main frequency is 2.6GHz, the memory is 8GB, the IDE is pycharm, and the programming language is Python. Based on the pytorch1.2 framework, the adversarial neural network algorithm is used to train the road damage dataset. mean IoU (the mean intersection-over-union) is used as the evaluation index for semantic segmentation. mean IoU reflects the quality of the model's segmentation effect. The larger the value, the better the effect, and vice versa. The final training results are shown in Table 1.
[0052] Algorithm MIoU (%) PSPNet 62.1 Deeplab 68.4 SegNet 65.8 ours 72.9
[0053] It can be seen from Table 1 that compared with other semantic segmentation methods, the mean IoU value of the algorithm used in this paper is the best, which is 0.729.
[0054] The above examples are used to explain the present invention, rather than limiting the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A semantic segmentation method based on an adversarial neural network, characterized in that: It includes the following steps, S100: Collect a road damage image dataset and divide it into a training dataset and a test dataset according to a ratio. The training dataset is used to train the adversarial neural network model, and the test dataset is used to verify the quality of the adversarial neural network model; S200: Construct a Resnet-50 network as a benchmark pre-training network, input the road damage image training set, and perform pre-training; S300: Train the segmentation model of the adversarial neural network. This model includes the PPM module in the PSPNet network and the ASPP module in DeepLab; The PPM module in PSPNet uses pooling for different sub-regions represented by 1×1, 2×2, 3×3, and 6×6 respectively, and finally upsamples and cascades this module; the ASPP module in DeepLab uses rates=(1, 6, 12, 18), and the dilated convolution operation consists of 4 parallel 3×3 convolutional layers; the input image size of the two segmentation networks is H×W×3, and the output is H×W×C 1 and H×W×C 2 Two class probability segmentation maps S 1 and S 2 , C 1 and C 2 are the number of classes output by the two segmentation networks respectively, and adversarial learning is performed on these two segmentation maps; During the specific training process, the total loss function of the entire adversarial network is divided into two parts, the loss functions of the two segmentation networks respectively and the loss function of the discriminative network. The loss functions of the two segmentation networks are expressed as the following formula: (1) (2) Among them, L seg-s1 represents the loss function for the PPM module to generate labels for the segmentation network, L seg-s2 represents the loss function for the ASPP module to generate labels for the segmentation network, L S represents the cross-entropy loss function between the Deeplab-based segmentation network and the ground truth label ; L adv is the loss function used to represent adversarial learning, which can deceive the discriminator as much as possible, making the discriminator misclassify the PPM-based segmentation prediction result as the prediction result of the Deeplab network; L weakly is the loss function of PPM, and this network uses the weak labels generated based on the Deeplab network , is a variable constant used to adjust the training process, and the specific formula is as follows: (3) (4) (5) (6) In the formula, H, w, and c represent the length, width, and segmentation categories of the image, s1 and s2 represent the class probability maps obtained by the two segmentation networks of the PPM module and the ASPP module, Xn represents the predicted label, and Yn represents the true label; S400: Input the results obtained by the two PPM module and ASPP module segmentation networks into the discriminative model to output a class global probability score map; S500: According to the results obtained by the discriminative network, readjust the two segmentation models for training, and repeat S300 and S400 until the results of the two segmentation networks tend to be consistent, so as to achieve the effect of semantic segmentation.
2. The semantic segmentation method based on an adversarial neural network according to claim 1, characterized in that: In the step S100, the road damage image dataset comes from the road scene images taken, including different environments, lighting, road surfaces, and shapes.
3. The semantic segmentation method based on an adversarial neural network according to claim 2, characterized in that: In the step S200, change the first convolutional layer of the original network Resnet-50 from 7×7 to three 3×3 convolutional layers, remove the last classification layer, then change the stride of the last residual module from 2 to 1, and finally, change the dilation rate of the network to 2 to obtain a feature map equivalent to 1 / 16 of the input image size.
4. The semantic segmentation method based on an adversarial neural network according to claim 3, characterized in that: In the step S400, the loss function of the discriminative network is expressed as the following formula: In the above formula is the learning rate of the discrimination network, and y n = 0 indicates that the sample comes from the PPM-based segmentation network, and y n = 1 indicates that the sample comes from the ASPP-based segmentation network.
5. The semantic segmentation method based on an adversarial neural network according to claim 4, characterized in that: In the step S400, the discriminative network contains 5 convolutional layers, where the size of the convolutional kernel is 4×4 and the stride is 2.
6. The semantic segmentation method based on an adversarial neural network according to claim 5, characterized in that: In the step S500 described above, after the generation network undergoes adversarial training, the segmentation result of the Deeplab network is used as the base image, and the segmentation result of the PSPNet network is regarded as the generated image. The two obtained images are input into the discriminant network in step S300, and the category of each target object is determined, and a class global probability score map with a size of H×W×1 is output; if the pixel in the image comes from the generated image of the Deeplab network, then P = 1, if it comes from the generated image of the PSPNet network, then P = 0. When the value of P is closer to 0.5, it indicates that the overall adversarial network has achieved the best effect, where P represents the degree of network adversarial training.
Citation Information
Patent Citations
Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field
AU2020103901A4
Adversarial-network-based semi-supervised semantic segmentation method
CN108549895A