Road Scene Semantic Segmentation Method Based on GAN and Cross-Modal Feature Fusion

By adopting a GAN and cross-modal fusion method in the semantic segmentation of road scenes, a generator network with cross-modal fusion and mesh context perception is constructed, which solves the problem of insufficient singleness and representativeness of feature maps in the prior art, and achieves a more efficient and more accurate semantic segmentation effect.

CN113378795BActive Publication Date: 2025-06-13ZHEJIANG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110785221.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-12
Publication Date
2025-06-13
Estimated Expiration
2041-07-12

AI Technical Summary

Technical Problem

In the prior art, the feature map obtained by simply using pooling operations and convolution operations is single and not representative, resulting in low segmentation accuracy of road scene semantic segmentation.

Method used

The road scene semantic segmentation method based on generative adversarial network (GAN) and cross-modal fusion is adopted. By constructing a generator convolutional neural network with a cross-modal fusion module and a mesh context-aware module, and training with a discriminator, a more representative feature map is generated.

Benefits of technology

The segmentation accuracy and segmentation efficiency of the semantic segmentation of road scenes are improved, and the generated feature maps are more representative, which significantly improves the segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113378795B_ABST
    Figure CN113378795B_ABST
Patent Text Reader

Abstract

The present invention discloses a road scene semantic segmentation method based on a generative adversarial network and cross-modal feature fusion, which is applied to the field of deep learning technology. The specific steps are as follows: Using the generative adversarial network, the training set is input into the generator convolutional neural network to obtain a predicted map corresponding to each original road scene map in the training set. The original road scene map, the corresponding predicted map, and the corresponding label map are input into the discriminator to obtain an evaluation score. According to the evaluation score, the discriminator parameters are updated; the training set is input into the generator convolutional neural network to obtain a predicted map corresponding to each original road scene map in the training set, the loss function is calculated, and the generator parameters are updated. The present invention solves the problem that the feature maps obtained by simply using pooling operations and convolutional operations in the prior art are single and unrepresentative, which will lead to a reduction in the feature information of the obtained images, and ultimately lead to a relatively rough restored effect information and low segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and more specifically, to a road scene semantic segmentation method based on a generative adversarial network and cross-modal feature fusion. Background Art

[0002] With the rise of the intelligent transportation industry, semantic segmentation has more and more applications in intelligent transportation systems. From traffic scene understanding and multi-object obstacle detection to visual navigation, semantic segmentation technology can be used to achieve them. Currently, the most commonly used semantic segmentation methods include algorithms such as support vector machines and random forests. These algorithms mainly focus on binary classification tasks for detecting and identifying specific objects, such as road surfaces, vehicles, and pedestrians. These traditional machine learning methods often need to be implemented through highly complex features, while using deep learning for semantic segmentation of traffic scenes is simple and convenient. More importantly, the application of deep learning has greatly improved the accuracy of image pixel-level classification tasks.

[0003] The semantic segmentation method using deep learning directly performs pixel-level end-to-end semantic segmentation. It only needs to input the images in the training set into the model framework for training to obtain the weights and the model, and then it can make predictions on the test set. The strength of a convolutional neural network lies in its multi-layer structure that can automatically learn features and can learn features at multiple levels. Currently, the methods for semantic segmentation based on deep learning are divided into two types. The first is the encoder-decoder architecture. In the encoding process, the pooling layer gradually reduces the position information and extracts abstract features; in the decoding process, the position information is gradually restored. Generally, there is a direct connection between the decoding and the encoding. The second architecture is dilated convolutions, which abandons the pooling layer and expands the receptive field through dilated convolutions. Dilated convolutions with smaller values have a smaller receptive field and can learn some specific features of parts; dilated convolution layers with larger values have a larger receptive field and can learn more abstract features, and these abstract features are more robust to the size, position, and orientation of objects, etc.

[0004] Most of the existing road scene semantic segmentation methods adopt deep learning methods, and there are many models that combine convolutional layers and pooling layers. However, the feature maps obtained solely by using pooling operations and convolutional operations are single and unrepresentative, which will lead to a reduction in the feature information of the obtained images, and ultimately result in a relatively rough restored effect information and low segmentation accuracy. Summary of the Invention

[0005] In view of this, the present invention provides a road scene semantic segmentation method based on a generative adversarial network and cross-modal fusion, which has high segmentation efficiency and high segmentation accuracy.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A road scene semantic segmentation method based on a generative adversarial network and cross-modal fusion, comprising the following steps:

[0008] Select a plurality of original road scene images and the corresponding true semantic segmentation images of each original road scene image, and form a training set from the plurality of original road scene images and the label maps corresponding to each original road scene image;

[0009] Construct a generator convolutional neural network with a cross-modal fusion module, a mesh context awareness module, resolution restoration, and enhanced semantic information;

[0010] Construct a discriminator convolutional neural network;

[0011] Input the training set into the generator convolutional neural network to obtain the predicted maps corresponding to each original road scene map in the training set. Input the original road scene map, the corresponding predicted map, and the corresponding label map into the discriminator to obtain an evaluation score. According to the evaluation score, update the discriminator parameters to minimize the score;

[0012] Input the training set into the generator convolutional neural network to obtain the predicted maps corresponding to each original road scene map in the training set. Calculate the loss function between the predicted map and the label map, and update the generator parameters.

[0013] Train the neural network multiple times to obtain a generator convolutional neural network classification training model.

[0014] The label map is a semantic segmentation label map, a foreground label map, and a multi-scale label map.

[0015] The predicted map includes a semantic segmentation predicted map, a foreground-background predicted map, and a multi-scale predicted map.

[0016] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a road scene semantic segmentation method based on a generative adversarial network and cross-modal fusion, which fully utilizes the encoding ability of the GAN and constructs a mapping bridge between data of different modalities, getting rid of the relatively complex network structure in the cross-modal models of existing deep networks, and solving the problem that the feature maps obtained by simply using pooling operations and convolutional operations in the prior art are single and unrepresentative, which will lead to a reduction in the feature information of the obtained images, and ultimately result in a relatively rough restored effect information and low segmentation accuracy. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0018] Figure 1 The attached drawing is a schematic structural diagram of the generator of the present invention;

[0019] Figure 2 The attached drawing is a schematic structural diagram of the first fusion block of the present invention;

[0020] Figure 3 The attached drawing is a schematic structural diagram of the second fusion block, the third fusion block, the fourth fusion block, and the fifth fusion block of the present invention;

[0021] Figure 4 The attached drawing is a schematic structural diagram of the network context awareness module of the present invention;

[0022] Figure 5 The attached drawing is a schematic structural diagram of the decoding block of the present invention;

[0023] Figure 6 The attached drawing is a schematic structural diagram of the discriminator of the present invention. Detailed implementation manners

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0025] The embodiment of the present invention discloses a road scene semantic segmentation method based on a convolutional neural network and a generative adversarial structure. The implementation block diagram of its generator is as Figure 1 shown, and the implementation block diagram of the discriminator is as Figure 2 shown, and it includes two processes: a training stage and a testing stage;

[0026] The specific step 1_1 is as follows:

[0027] Select T initial road scene images and the corresponding heat maps and true semantic segmentation images of each original road scene image, where the t-th initial road scene image is denoted as Denote the heat map corresponding to the t-th road scene image as Denote the true semantic segmentation image corresponding to the t-th original road scene image as The ground truth semantic segmentation image corresponding to the t-th original road scene image is processed into 9 class label images, and the set composed of these 9 class label images is used as the semantic segmentation ground truth label map, denoted as In , the sizes of the pictures are respectively reduced to 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32 of the original size, and 5 groups of semantic segmentation label maps with different sizes are obtained and In , the background category is set to 0, and the non-background category is set to 1. In this way, two foreground label maps for distinguishing the background and the foreground are generated, denoted as Repeat the above three operations T times. A training set is composed of T original road scene images, the corresponding heat maps, semantic segmentation labels, multi-scale semantic segmentation labels, and foreground labels; among them, if T = 784, t is a positive integer, 1 ≤ t ≤ T, 1 ≤ i ≤ W, 1 ≤ j ≤ H, W represents the width of the initial road scene image, H represents the height of the initial road scene image, where W = 480 and H = 640 are taken, and i and j respectively represent the horizontal and vertical coordinates of the pixel point at the coordinate position (i, j), represents the pixel value of the pixel point at the coordinate position (i, j) in the t-th initial road scene image, represents the pixel value of the pixel point at the coordinate position (i, j) in the t-th heat map, represents the pixel value of the pixel point at the coordinate position (i, j) in the ground truth semantic segmentation image, and represent the pixel values of the pixel points at the (i, j) position in the five different-scale ground truth semantic segmentation images, represents the pixel value of the pixel point at the coordinate position (i, j) in the ground truth foreground-background image. In specific implementation, 784 images in the road scene image data InfRecR500 training set of the road scene images are directly selected as the initial road scene images.

[0028] Step 1_2: Construct a generator convolutional neural network and a discriminator convolutional neural network:

[0029] Construct the convolutional neural network of the generator

[0030] The generator convolutional neural network includes two input layers, a hidden layer, a multi-scale output layer, and two regular output layers; the hidden layer consists of a dual-branch structure. The first branch includes a first initial neural network block, a first residual neural network block, a second residual neural network block, a third residual neural network block, and a fourth residual neural network block; the second branch consists of a second initial neural network block, a fifth residual neural network block, a sixth residual neural network block, a seventh residual neural network block, and an eighth residual neural network block. Between the two branches, there are a first fusion block, a second fusion block, a third fusion block, a fourth fusion block, a fifth fusion block, a mesh context awareness block, a first decoding block, a second decoding block, a third decoding block, and a fourth decoding block. The first input layer, the first initial neural network block, the first residual neural network block, the second residual neural network block, the third residual neural network block, and the fourth neural network block are connected in sequence. The second input layer, the second initial neural network block, the fifth residual neural network block, the sixth residual neural network block, the seventh residual neural network block, and the eighth residual neural network block are connected in sequence. The outputs of the first initial neural network block and the second initial neural network block are jointly used as the input of the first fusion block, and the output of the first fusion block is denoted as the first side output. The outputs of the first residual neural network block, the fifth residual neural network block, and the first fusion block are jointly used as the input of the second fusion block, and the output of the second fusion block is denoted as the second side output. The outputs of the second residual neural network block, the sixth residual neural network block, and the second fusion block are jointly used as the input of the third fusion block, and the output of the third fusion block is denoted as the third side output. The outputs of the third residual neural network block, the seventh residual neural network block, and the third fusion block are jointly used as the input of the fourth fusion block, and the output of the fourth fusion block is denoted as the fourth side output. The outputs of the fourth residual neural network block, the eighth residual neural network block, and the fourth fusion block are jointly used as the output of the fifth fusion block, and the output of the fifth fusion block is denoted as the fifth side output. The first side output, the second side output, the third side output, and the fourth side output are respectively used as the input of the multi-scale output layer and the input of the decoding stage in addition to being used as the input of the next fusion block. The output of the fifth fusion block is used as the input of the multi-scale output layer and the input of the mesh context awareness block, and the output of the mesh context awareness block is denoted as the awareness block output;

[0031] The awareness block output that has passed through the mesh context awareness block is used as the input of the first decoding block. The element-wise sum of the fourth side output that has passed through the fourth fusion block and the output of the first decoding block is used as the input of the second decoding block. The element-wise sum of the third side output that has passed through the third fusion block and the output of the second decoding block is used as the input of the third decoding block. The output of the third decoding block is used as the input of the fourth decoding block. The output of the fourth decoding block is denoted as the guiding feature, and the guiding feature is used as the input of the first regular output layer;

[0032] Perform bilinear interpolation on the third side output passing through the third fusion block. Add element-wise the feature map with doubled resolution after bilinear interpolation, the second side output, and the output of the third decoding block. Perform bilinear interpolation on the obtained sum to enlarge the resolution by two times. Add element-wise the enlarged feature map with the first side output. Multiply element-wise the obtained sum with the guiding feature, and use the obtained product as the input of the second output layer;

[0033] The multi-scale output layer consists of five independent convolutional layers. The convolutional kernel size (kernel_size) of the convolutional layer is 1×1, the number of convolutional kernels (filters) is 9, the stride is 1, and the zero-padding (padding) parameter is 0. Corresponding to the inputs of the five convolutional layers, their inputs are the first, second, third, fourth side outputs and the output of the perception block respectively.

[0034] Use the initial road scene image in the training set as the input in the first input layer. The output of the first output layer is the R channel component, G channel component, and B channel component of the original road scene image of this sheet. Use the heat map corresponding to the original road scene image of this sheet as the input of the second input layer. It is required that the width of the initial road scene image and the corresponding heat map received at the input end of the input layer is W, and the height is H.

[0035] For the first branch in the hidden layer:

[0036] The first initial neural network block mainly includes a first convolutional layer (Convolution, Conv) and a first activation layer (Activation, Act), which are connected in sequence. Among them, the convolutional kernel size (kernel_size) of the first convolutional layer is 7×7, the number of convolutional kernels (filters) is 64, the stride is 2, and the zero-padding (padding) parameter is 3. The activation method of the first activation layer is "Relu" for all. Use the R channel, G channel, and B channel components as the input of the first initial neural network block. The output end of the first initial neural network block outputs 64 feature maps, and denote the set of these 64 feature maps as R 1 , R 1 The width of each feature map in The height is

[0037] For the first residual neural network block, it is mainly composed of a first max pooling layer (Maxpooling, Pool) connected to the first residual layer of ResNet50; among them, the pooling size of the first max pooling layer is 2, and the structure of the first residual layer of ResNet50 is the same as the structure of Layer1 in the commonly used neural network architecture ResNet50 that has been publicly disclosed. Taking all the feature maps in R 1 as the input of the first residual neural network block, the output of the first residual neural network block is 256 feature maps, and the set composed of these 256 feature maps is denoted as R l1 ;; R l1 The width of each feature map in and the height is

[0038] For the second residual neural network block, it is mainly composed of the second residual layer of ResNet50, and the structure of the second residual layer of ResNet50 is the same as the structure of Layer2 in the commonly used neural network architecture ResNet50 that has been publicly disclosed. Taking all the feature maps in R l1 as the input of the second residual neural network block, the output end of the second residual neural network block outputs 512 feature maps, and their set is denoted as R l2 . R l2 The width of each feature map in and the height is

[0039] For the third residual neural network block, it is mainly composed of the third residual layer of ResNet50, and the structure of the third residual layer of ResNet50 is the same as the structure of Layer3 in the commonly used neural network architecture ResNet50 that has been publicly disclosed. Taking all the feature maps in R l2 as the input of the third residual neural network block, the output end of the third residual neural network block outputs 1024 feature maps, and their set is denoted as R l3 . R l3 The width of each feature map in and the height is

[0040] For the fourth residual neural network block, it is mainly composed of the fourth residual layer of ResNet50, and the structure of the fourth residual layer of ResNet50 is the same as the structure of Layer4 in the commonly used neural network architecture ResNet50 that has been publicly disclosed. Taking all the feature maps in R l3 as the input of the fourth residual neural network block, the output end of the fourth residual neural network block outputs 2048 feature maps, and their set is denoted as R l4 . R l4 The width of each feature map in The height is

[0041] For the second branch in the hidden layer:

[0042] The second initial neural network block mainly includes a second convolutional layer (Convolution, Conv) and a second activation layer (Activation, Act), which are connected in sequence; among them, the convolutional kernel size (kernel_size) of the second convolutional layer is 7×7, the number of convolutional kernels (filters) is 64, the stride is 2, and the zero-padding (padding) parameter is 3. The activation method of the second activation layer is "Relu"; a single-channel thermal map (Thermal) is used as the input of the second initial neural network block, and the second initial neural network block outputs 64 feature maps. The set of these 64 feature maps is denoted as T 1 , T 1 The width of each feature map in The height is

[0043] For the fifth residual neural network block, it is mainly composed of a second max-pooling layer (Maxpooling, Pool) connected to the first residual layer of ResNet50; among them, the pooling size (pool_size) of the second max-pooling layer is 2, and the structure of the first residual layer of ResNet50 is the same as the structure of layer 1 (Layer1) in the commonly used neural network architecture ResNet50 that has been publicly disclosed. All the feature maps in T 1 are used as the input of the fifth residual neural network block, and the output of the fifth residual neural network block is 256 feature maps. The set composed of these 256 feature maps is denoted as T l1 ;; T l1 The width of each feature map in The height is

[0044] For the sixth residual neural network block, it is mainly composed of the second residual layer of ResNet50, and the structure of the second residual layer of ResNet50 is the same as the structure of layer 2 (Layer2) in the commonly used neural network architecture ResNet50 that has been publicly disclosed. All the feature maps in T l1 are used as the input of the sixth residual neural network block, and 512 feature maps are output at the output end of the sixth residual neural network block. Their set is denoted as T l2 . T l2 The width of each feature map in The height is

[0045] For the seventh residual neural network block, it is mainly composed of the third residual layer of ResNet50, and the structure of the third residual layer of ResNet50 is the same as the structure of Layer3 in the commonly used neural network architecture ResNet50 that has been made public. Taking all the feature maps in T l2 as the input of the seventh residual neural network block, the output end of the seventh residual neural network block outputs 1024 feature maps, and their set is denoted as T l3 . Each feature map in T l3 has a width of and a height of

[0046] For the eighth residual neural network block, it is mainly composed of the fourth residual layer of ResNet50, and the structure of the fourth residual layer of ResNet50 is the same as the structure of Layer4 in the commonly used neural network architecture ResNet50 that has been made public. Taking all the feature maps in T l3 as the input of the eighth residual neural network block, the output end of the eighth residual neural network block outputs 2048 feature maps, and their set is denoted as T l4 . Each feature map in T l4 has a width of and a height of

[0047] For the first fusion block, it is composed of a third convolutional layer and a fourth convolutional layer connected according to the Figure 2 shown structure. The convolutional kernel size (kernel_size) of the third and fourth convolutional layers is 1×1, the number of convolutional kernels (filters) is 64, the stride is 1, and the zero-padding (padding) parameter is 0; inputting the output R 1 of the first initial neural network block and the output T 1 of the second initial neural network block into the first fusion block respectively, the output end of the first fusion block outputs 64 feature maps, denoted as the first side output F 1 . The specific connection method is as follows: first input the output R 1 of the first initial neural network block and the output T 1 of the second initial neural network block into the third convolutional layer and the fourth convolutional layer respectively. The third convolutional layer and the fourth convolutional layer both output 64 feature maps, denoted as A 1 and B 1 respectively; performing element-wise addition on each feature map in A 1 and each feature map in B 1 to obtain 64 preliminary fusion feature maps, denoted as f 1 , and passing each feature map in f 1 through the Sigmoid activation function to obtain 64 feature maps, denoted as S1 ; Respectively, each feature map in S 1 is multiplied element - by - element with each feature map in A 1 , B 1 to obtain two sets of 64 feature maps, denoted as SA 1 and SB 1 respectively; Finally, SA 1 , SB 1 and the preliminary fusion feature map f 1 are added element - by - element to obtain 64 feature maps, denoted as the first - side output F 1 . The width of each feature map in F 1 is and the height is

[0048] For the second fusion block, it is composed of the fifth convolutional layer, the sixth convolutional layer, the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, the tenth convolutional layer, and the eleventh convolutional layer connected in the structure shown in Figure 3 . The fifth and sixth convolutional layers have a kernel size of 1×1, 64 filters, a stride of 1, and a padding parameter of 0; The seventh and eighth convolutional layers have a kernel size of 3×3, 64 filters, a stride of 1, and a padding parameter of 1; The eighth and tenth convolutional layers have a kernel size of 3×3, 64 filters, a stride of 2, and a padding parameter of 1; The eleventh convolutional layer has a kernel size of 1×1, 64 filters, a stride of 1, and a padding parameter of 0; The output R l1 of the first residual neural network block, the output T l1 of the fifth residual neural network block, and the first - side output F 1 are respectively input into the second fusion block, and the output end of the second fusion block outputs 64 feature maps, denoted as the second - side output F 2 . The specific connection method is as follows: First, the output R l1 of the first residual neural network block and the output T l1 of the fifth residual neural network block are respectively input into the fifth convolutional layer and the sixth convolutional layer. Both the fifth convolutional layer and the sixth convolutional layer output 64 feature maps, denoted as A l1 and B l1 respectively; Each feature map in A l1 is multiplied element - by - element with B l1Element-wise summation is performed on each feature map in it to obtain 64 preliminary fusion feature maps, denoted as f l1 , and each feature map in f l1 passes through the Sigmoid activation function

[0049] to obtain 64 feature maps, denoted as S l1 ; Each feature map in S l1 is multiplied element-wise with each feature map in A l1 and B l1 respectively to obtain two sets of 64 feature maps, denoted as SA l1 and SB l1 respectively; Each feature map in the first side output F 1 is input into the seventh convolutional layer, and the 64 feature maps output by the seventh convolutional layer are then input into the eighth convolutional layer, and the eighth convolutional layer outputs 64 feature maps, denoted as U l1 , and the width of each feature map in U l1 is and the height is Each feature map in the first side output F 1 is input into the ninth convolutional layer, and the 64 feature maps output by the ninth convolutional layer are then input into the tenth convolutional layer, and the tenth convolutional layer outputs 64 feature maps, denoted as D l1 , and the width of each feature map in D l1 is and the height is The 64 feature maps in U l1 are superimposed on the 64 feature maps in SA l1 to obtain 128 feature maps, denoted as M l1 ; The 64 feature maps in D l1 are superimposed on the 64 feature maps in SB l1 to obtain 128 feature maps, denoted as N l1 ; Each feature map in M l1 is added element-wise with each feature map in N l1 , and the sum obtained is used as the input of the eleventh convolutional layer. The eleventh convolutional layer outputs 64 feature maps denoted as MN l1 , and finally MN l1 and the preliminary fusion feature map f l1 are added element-wise to obtain 64 feature maps, denoted as the second side output F 2 , and the width of each feature map in F 2 is and the height is

[0050] For the third fusion block, its structure is similar to that of the second fusion block and is composed of the twelfth convolutional layer, the thirteenth convolutional layer, the fourteenth convolutional layer, the fifteenth convolutional layer, the sixteenth convolutional layer, the seventeenth convolutional layer, and the eighteenth convolutional layer connected according to Figure 3 the structure shown. The kernel size of the twelfth and thirteenth convolutional layers (kernel_size) is 1×1, the number of kernels (filters) is 64, the stride is 1, and the padding parameter is 0; the kernel size of the fourteenth and sixteenth convolutional layers (kernel_size) is 3×3, the number of kernels (filters) is 64, the stride is 1, and the padding parameter is 1; the kernel size of the fifteenth and seventeenth convolutional layers (kernel_size) is 3×3, the number of kernels (filters) is 64, the stride is 2, and the padding parameter is 1; the kernel size of the eighteenth convolutional layer (kernel_size) is 1×1, the number of kernels (filters) is 64, the stride is 1, and the padding parameter is 0; the output R l2 of the second residual neural network block, l2 the output T 2 of the sixth residual neural network block, and 3 the second side output F l2 are respectively input into the third fusion block, and 64 feature maps are output at the output end of the third fusion block, denoted as the third side output F l2 . The specific connection method is as follows: First, the output R l2 of the second residual neural network block and l2 the output T l2 of the sixth residual neural network block l2 are respectively input into the twelfth convolutional layer and the thirteenth convolutional layer. Both the twelfth convolutional layer and the thirteenth convolutional layer output 64 feature maps, denoted as A l2 and B l2 respectively; each feature map in A l2 is added element-wise to each feature map in B l2 to obtain 64 preliminary fusion feature maps, denoted as f l2 . Each feature map in f l2 is passed through the Sigmoid activation function to obtain 64 feature maps, denoted as S l2 ; each feature map in S l2 is multiplied element-wise with each feature map in A 2Each feature map in [[]] is input into the fourteenth convolutional layer. The 64 feature maps output by the fourteenth convolutional layer are then input into the fifteenth convolutional layer. The fifteenth convolutional layer outputs 64 feature maps, denoted as U l2 , U l2 The width of each feature map in [[]] is The height is Output the second side F 2 Each feature map in [[]] is input into the sixteenth convolutional layer. The 64 feature maps output by the sixteenth convolutional layer are then input into the seventeenth convolutional layer. The seventeenth convolutional layer outputs 64 feature maps, denoted as D l2 , D l2 The width of each feature map in [[]] is The height is For U l2 The 64 feature maps in [[]] are superimposed with the 64 feature maps in SA l2 to obtain 128 feature maps, denoted as M l2 ; For D l2 The 64 feature maps in [[]] are superimposed with the 64 feature maps in SB l2 to obtain 128 feature maps, denoted as N l2 ; For each feature map in M l2 and each feature map in N l2 perform element-wise addition, and use the obtained sum as the input of the eighteenth convolutional layer. The eighteenth convolutional layer outputs 64 feature maps denoted as MN l2 , and finally add MN l2 and the preliminary fusion feature map f l2 element-wise to obtain 64 feature maps, denoted as the third side output F 3 , F 3 The width of each feature map in [[]] is The height is

[0051] For the fourth fusion block, its structure is similar to that of the second fusion block, consisting of the nineteenth convolutional layer, the twentieth convolutional layer, the twenty-first convolutional layer, the twenty-second convolutional layer, the twenty-third convolutional layer, the twenty-fourth convolutional layer, and the twenty-fifth convolutional layer in accordance with Figure 3The structure shown is connected to form. The 19th convolutional layer and the 20th convolutional layer have a kernel size of 1×1, 64 filters, a stride of 1, and a padding parameter of 0; the 21st convolutional layer and the 23rd convolutional layer have a kernel size of 3×3, 64 filters, a stride of 1, and a padding parameter of 1; the 22nd convolutional layer and the 24th convolutional layer have a kernel size of 3×3, 64 filters, a stride of 2, and a padding parameter of 1; the 25th convolutional layer has a kernel size of 1×1, 64 filters, a stride of 1, and a padding parameter of 0; the output R of the third residual neural network block l3 , the output T of the seventh residual neural network block l3 and the third side output F 3 are respectively input into the fourth fusion block, and 64 feature maps are output at the output end of the fourth fusion block, denoted as the fourth side output F 4 . The specific connection method is as follows: First, the output R of the third residual neural network block l3 , the output T of the seventh residual neural network block l3 are respectively input into the 19th convolutional layer and the 20th convolutional layer. Both the 19th convolutional layer and the 20th convolutional layer output 64 feature maps, denoted as A l3 and B l3 respectively; each feature map in A l3 is added element-wise to each feature map in B l3 to obtain 64 preliminary fusion feature maps, denoted as f l3 . Each feature map in f l3 passes through the Sigmoid activation function to obtain 64 feature maps, denoted as S l3 ; each feature map in S l3 is multiplied element-wise with each feature map in A l3 and B l3 respectively to obtain two groups of 64 feature maps, denoted as SA l3 and SB l3 respectively; each feature map in the third side output F 3 is input into the 21st convolutional layer, and the 64 feature maps output by the 21st convolutional layer are then input into the 22nd convolutional layer. The 22nd convolutional layer outputs 64 feature maps, denoted as U l3 . The width of each feature map in U l3 is The height is Output the third side F 3 Input each feature map in it into the twenty-third convolutional layer, and then input the 64 feature maps output by the twenty-third convolutional layer into the twenty-fourth convolutional layer. The twenty-fourth convolutional layer outputs 64 feature maps, denoted as D l3 , D l3 The width of each feature map in The height is Input U l3 Overlay the 64 feature maps in it with the 64 feature maps in SA l3 to obtain 128 feature maps, denoted as M l3 ; Overlay the 64 feature maps in D l3 with the 64 feature maps in SB l3 to obtain 128 feature maps, denoted as N l3 ; Add each feature map in M l3 to each feature map in N l3 element by element, and use the obtained sum as the input of the twenty-fifth convolutional layer. The twenty-fifth convolutional layer outputs 64 feature maps denoted as MN l3 , and finally add MN l3 and the preliminary fusion feature map f l3 element by element to obtain 64 feature maps, denoted as the fourth side output F 4 , F 4 The width of each feature map in The height is

[0052] For the fifth fusion block, its structure is similar to that of the second fusion block, consisting of the twenty-sixth convolutional layer, the twenty-seventh convolutional layer, the twenty-eighth convolutional layer, the twenty-ninth convolutional layer, the thirtieth convolutional layer, the thirty-first convolutional layer, and the thirty-second convolutional layer in accordance with Figure 3is composed of the structures shown. The twenty-sixth convolutional layer and the twenty-seventh convolutional layer have a kernel size of 1×1, 64 filters, a stride of 1, and a padding parameter of 0; the twenty-eighth convolutional layer and the thirtieth convolutional layer have a kernel size of 3×3, 64 filters, a stride of 1, and a padding parameter of 1; the twenty-ninth convolutional layer and the thirty-first convolutional layer have a kernel size of 3×3, 64 filters, a stride of 2, and a padding parameter of 1; the thirty-second convolutional layer has a kernel size of 1×1, 64 filters, a stride of 1, and a padding parameter of 0; the output R of the fourth residual neural network block l4 , the output T of the eighth residual neural network block l4 and the fourth side output F 4 are respectively input into the fifth fusion block, and 64 feature maps are output at the output end of the fifth fusion block, denoted as the fifth side output F 5 . The specific connection method is as follows: first, the output R l4 of the fourth residual neural network block and the output T l4 of the eighth residual neural network block are respectively input into the twenty-sixth convolutional layer and the twenty-seventh convolutional layer. Both the twenty-sixth convolutional layer and the twenty-seventh convolutional layer output 64 feature maps, denoted as A l4 and B l4 respectively; each feature map in A l4 is added element-wise to each feature map in B l4 to obtain 64 preliminary fusion feature maps, denoted as f l4 . Each feature map in f l4 passes through the Sigmoid activation function to obtain 64 feature maps, denoted as S l4 ; each feature map in S l4 is multiplied element-wise with each feature map in A l4 and B l4 respectively to obtain two groups of 64 feature maps, denoted as SA l4 and SB l4 respectively; each feature map in the fourth side output F 4 is input into the twenty-eighth convolutional layer, and the 64 feature maps output by the twenty-eighth convolutional layer are then input into the twenty-ninth convolutional layer. The twenty-ninth convolutional layer outputs 64 feature maps, denoted as U l4 , U l4The width of each feature map in is Output the fourth side F 4 Input each feature map in l4 to the thirtieth convolutional layer, and then input the 64 feature maps obtained by the output of the thirtieth convolutional layer to the thirty-first convolutional layer. The thirty-first convolutional layer outputs 64 feature maps, denoted as D l4 The width of each feature map in D is Input U l4 The 64 feature maps in l4 are superimposed with the 64 feature maps in SA l4 to obtain 128 feature maps, denoted as M l4 ; The 64 feature maps in D l4 are superimposed with the 64 feature maps in SB l4 to obtain 128 feature maps, denoted as N l4 ; Each feature map in M l4 is element-wise added to each feature map in N l4 , and the obtained sum is used as the input of the thirty-second convolutional layer. The thirty-second convolutional layer outputs 64 feature maps denoted as MN l4 , and finally MN l4 and the preliminary fusion feature map f 5 are element-wise added to obtain 64 feature maps, denoted as the fourth side output F 5 The width of each feature map in F is

[0053] For the mesh context awareness module, it consists of three parallel branches and the thirty-third convolutional layer. The first branch is composed of the first convolutional block, the first dilated convolutional block, the second dilated convolutional block, and the third dilated convolutional block connected in sequence; The second branch is composed of the second convolutional block, the fourth dilated convolutional block, the fifth dilated convolutional block, and the sixth dilated convolutional block connected in sequence; The third branch is composed of the third convolutional block, the seventh dilated convolutional block, the eighth dilated convolutional block, and the ninth dilated convolutional block connected in sequence; The above-mentioned convolutional blocks and dilated convolutional blocks are arranged as Figure 4are connected in the following way. Among them, the first convolutional block is composed of the thirty-fourth convolutional layer and the third activation layer connected in series. The second convolutional block is composed of the thirty-fifth convolutional layer and the fourth activation layer connected in series. The third convolutional block is composed of the thirty-sixth convolutional layer and the fifth activation layer connected in series. The kernel sizes (kernel_size) of the thirty-fourth convolutional layer, the thirty-fifth convolutional layer, and the thirty-sixth convolutional layer are all 1×1, the number of convolutional kernels (filters) is 64, the stride is 1, and the zero-padding (padding) parameter is 0. The activation methods of the third activation layer, the fourth activation layer, and the fifth activation layer are all "Relu". The first dilated convolutional block is composed of the thirty-seventh convolutional layer, the sixth activation layer, the thirty-eighth convolutional layer, the seventh activation layer, the first dilated convolutional layer, and the eighth activation layer connected in sequence. The second dilated convolutional block is composed of the thirty-ninth convolutional layer, the ninth activation layer, the fortieth convolutional layer, the tenth activation layer, the second dilated convolutional layer, and the eleventh activation layer connected in sequence. The third dilated convolutional block is composed of the forty-first convolutional layer, the twelfth activation layer, the forty-second convolutional layer, the thirteenth activation layer, the third dilated convolutional layer, and the fourteenth activation layer connected in sequence. The fourth dilated convolutional block is composed of the forty-third convolutional layer, the fifteenth activation layer, the forty-fourth convolutional layer, the sixteenth activation layer, the fourth dilated convolutional layer, and the seventeenth activation layer connected in sequence. The fifth dilated convolutional block is composed of the forty-fifth convolutional layer, the eighteenth activation layer, the forty-sixth convolutional layer, the nineteenth activation layer, the fifth dilated convolutional layer, and the twentieth activation layer connected in sequence. The sixth dilated convolutional block is composed of the forty-seventh convolutional layer, the twenty-first activation layer, the forty-eighth convolutional layer, the twenty-second activation layer, the sixth dilated convolutional layer, and the twenty-third activation layer connected in sequence. The seventh dilated convolutional block is composed of the forty-ninth convolutional layer, the twenty-fourth activation layer, the fiftieth convolutional layer, the twenty-fifth activation layer, the seventh dilated convolutional layer, and the twenty-sixth activation layer connected in sequence. The eighth dilated convolutional block is composed of the fifty-first convolutional layer, the twenty-seventh activation layer, the fifty-second convolutional layer, the twenty-eighth activation layer, the eighth dilated convolutional layer, and the twenty-ninth activation layer connected in sequence. The ninth dilated convolutional block is composed of the fifty-third convolutional layer, the thirtieth activation layer, the fifty-fourth convolutional layer, the thirty-first activation layer, the ninth dilated convolutional layer, and the thirty-second activation layer connected in sequence. Among them, the kernel sizes (kernel_size) of the thirty-seventh convolutional layer, the forty-third convolutional layer, and the forty-ninth convolutional layer are all 1×3, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameters are all (0,1), the kernel sizes (kernel_size) of the 38th convolutional layer, 44th convolutional layer, and 50th convolutional layer are all 3×1, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameters are all (1, 0). The kernel sizes (kernel_size) of the first dilated convolutional layer, 4th dilated convolutional layer, and 7th dilated convolutional layer are all 3×3, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameter is all 3, and the dilation rates (dilationrate) are all 3. The kernel sizes (kernel_size) of the 39th convolutional layer, 45th convolutional layer, and 51st convolutional layer are all 1×5, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameters are all (0, 2). The kernel sizes (kernel_size) of the 40th convolutional layer, 46th convolutional layer, and 52nd convolutional layer are all 5×1, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameters are all (2, 0). The kernel sizes (kernel_size) of the second dilated convolutional layer, 5th dilated convolutional layer, and 8th dilated convolutional layer are all 3×3, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameter is all 5, and the dilation rates (dilation rate) are all 5. The kernel sizes (kernel_size) of the 41st convolutional layer, 47th convolutional layer, and 53rd convolutional layer are all 1×7, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameters are all (0, 3). The kernel sizes (kernel_size) of the 42nd convolutional layer, 48th convolutional layer, and 54th convolutional layer are all 7×1, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameters are all (3, 0). The kernel sizes (kernel_size) of the third dilated convolutional layer, 6th dilated convolutional layer, and 9th dilated convolutional layer are all 3×3, the number of convolutional kernels (filters) are all 64, the strides (stride) are all 1, and the zero-padding (padding) parameter is all 7, and the dilation rates (dilationrate) are all 7. The activation method of the activation layer in this module is all "Relu". The specific connection description is as follows: Denote the outputs of the four convolutional blocks in the first branch as B, 1-1 , B 1-2 , B 1-3 and B 1-4, denote the outputs of the four convolutional blocks of the second branch as B 2-1 、B 2-2 、B 2-3 and B 2-4 , denote the outputs of the four convolutional blocks of the third branch as B 3-1 、B 3-2 、B 3-3 and B 3-4 , it is worth mentioning that in the second branch, the sum of the input F 5 and B 1-1 is used as the input of the second convolutional block, and the sum of B 2-1 and B 1-2 is used as the input of the fourth dilated convolutional block, the sum of B 2-2 and B 1-3 is used as the input of the fifth dilated convolutional block, and the sum of B 2-3 and B 1-4 is used as the input of the sixth dilated convolutional block. In the third branch, the sum of F 5 , B 2-1 and B 1-2 is used as the input of the third convolutional block, the sum of B 3-1 , B 2-2 and B 1-3 is used as the input of the seventh dilated convolutional block, the sum of B 3-2 , B 2-3 and B 1-4 is used as the input of the eighth dilated convolutional block, and the sum of B 3-3 and B 2-4 is used as the input of the ninth dilated convolutional block. Stack the 64 feature maps in the last convolutional block of the three branches together to generate 192 feature maps, denoted as CB. Use CB as the input of the thirty-third convolutional layer, and the thirty-third convolutional layer outputs 64 feature maps. Then, perform element-wise addition of these 64 feature maps and the input F 5 of the mesh context-aware block to obtain 64 feature maps, denoted as The width of each feature map in is

[0054] For the first decoding block, it is mainly composed of a fifty-fifth convolutional layer, a thirty-third activation layer, a fifty-sixth convolutional layer, a thirty-fourth activation layer, a fifty-seventh convolutional layer, and a thirty-fifth activation layer connected in sequence. The structure is as Figure 5 shown. Use each feature map in the output of the mesh context-aware block as the input of the first decoding block, and the output end of the first decoding block outputs 64 feature maps, denoted as De 1Among them, the convolutional kernel sizes of the fifty-fifth convolutional layer, the fifty-sixth convolutional layer, and the fifty-seventh convolutional layer are all 3×3, the number of convolutional kernels is 64 each, the stride is 1 each, and the zero-padding parameter is "same"; the activation methods of the thirty-third activation layer, the thirty-fourth activation layer, and the thirty-fifth activation layer are all "Relu". The input of the first decoding block is added element-wise to the output of the fifty-seventh convolutional layer in this decoding block, and the sum is upsampled to double the original resolution. The upsampling method is bilinear interpolation. The set of 64 feature maps obtained is the output De of the first decoding block 1 . De 1 The width of each feature map in The height is

[0055] For the second decoding block, it is mainly composed of the fifty-eighth convolutional layer, the thirty-sixth activation layer, the fifty-ninth convolutional layer, the thirty-seventh activation layer, the sixtieth convolutional layer, and the thirty-eighth activation layer connected in sequence, and the structure is as Figure 5 shown. The element-wise sum of each feature map in the output De 1 of the first decoding block and each feature map in the fourth-side output F 4 is used as the input of the second decoding block. Its output end outputs 64 feature maps, denoted as De 2 . Among them, the convolutional kernel sizes of the fifty-eighth convolutional layer, the fifty-ninth convolutional layer, and the sixtieth convolutional layer are all 3×3, the number of convolutional kernels is 64 each, the stride is 1 each, and the zero-padding parameter is "same"; the activation methods of the thirty-sixth activation layer, the thirty-seventh activation layer, and the thirty-eighth activation layer are all "Relu". The element-wise sum of the input of the second decoding block and the output of the sixtieth convolutional layer in this decoding block is upsampled to double the original resolution. The upsampling method is bilinear interpolation, and the set of 64 feature maps obtained is the output De of the second decoding block 2 . De 2 The width of each feature map in The height is

[0056] For the third decoding block, it is mainly composed of the sixty-first convolutional layer, the thirty-ninth activation layer, the sixty-second convolutional layer, the fortieth activation layer, the sixty-third convolutional layer, and the forty-first activation layer connected in sequence, and the structure is as Figure 5 shown. The element-wise sum of each feature map in the output De 2 of the second decoding block and each feature map in the third-side output F 3 is used as the input of the third decoding block. Its output end outputs 64 feature maps, denoted as De 3Among them, the convolutional kernels of the sixty-first convolutional layer, the sixty-second convolutional layer, and the sixty-third convolutional layer are all 3×3 in size, the number of convolutional kernels is 64 for each, the stride is 1 for each, and the zero-padding parameter is "same" for each; the activation methods of the thirty-ninth activation layer, the fortieth activation layer, and the forty-first activation layer are all "Relu". Add the input of the third decoding block element-wise to the output of the sixty-third convolutional layer in this decoding block, and upsample the sum to double the original resolution. The upsampling method is bilinear interpolation. The set of 64 feature maps obtained is the output De of the third decoding block. 3 De 3 The width of each feature map in The height is

[0057] For the fourth decoding block, it is mainly composed of the sixty-fourth convolutional layer, the forty-second activation layer, the sixty-fifth convolutional layer, the forty-third activation layer, the sixty-sixth convolutional layer, and the forty-fourth activation layer connected in sequence. The structure is as Figure 5 shown. Use each feature map in the output De 3 of the third decoding block as the input of the fourth decoding block. Its output end outputs 64 feature maps, denoted as De 4 Among them, the convolutional kernels of the sixty-fourth convolutional layer, the sixty-fifth convolutional layer, and the sixty-sixth convolutional layer are all 3×3 in size, the number of convolutional kernels is 64 for each, the stride is 1 for each, and the zero-padding parameter is "same" for each; the activation methods of the forty-second activation layer, the forty-third activation layer, and the forty-fourth activation layer are all "Relu". Add the input of the fourth decoding block element-wise to the output of the sixty-sixth convolutional layer in this decoding block, and upsample the sum to double the original resolution. The upsampling method is bilinear interpolation. The set of 64 feature maps obtained is the output De of the fourth decoding block. 4 De 4 The width of each feature map in The height is

[0058] For the first regular output layer, it consists of the sixty-seventh convolutional layer and an upsampling operation. The upsampling doubles the resolution through bilinear interpolation. The first regular input end receives each feature map in De 4 After being processed by the output layer, 9 semantic segmentation prediction maps of the same size as the original input image and corresponding to it are output, denoted as P 1 Among them, the convolutional kernel of the sixty-seventh convolutional layer is 1×1 in size, the number of convolutional kernels is 9, the stride is 1, and the zero-padding parameter is "same";

[0059] Add the third side output F 3, after bilinear interpolation for upsampling by a factor of two, the resolution is doubled, and the set of 64 upsampled feature maps is denoted as F 3up , F 3up The width of each feature map in The height is For each feature map in F 3up , the output De of the third decoding block 3 and each feature map in the second side output F 2 are added element-wise to obtain 64 feature maps, and their set is denoted as X. Each feature map in X is bilinearly interpolated to double the original resolution, resulting in 64 feature maps, denoted as Y. The width of each feature map in Y is The height is Each feature map in Y is added element-wise to the first side output F 1 to obtain 64 feature maps denoted as Z. Then, each feature map in Z is multiplied element-wise with each feature map in the output De of the fourth decoding block 4 to obtain 64 feature maps J. The width of each feature map in J is The height is

[0060] For the second regular output layer, it consists of the sixty-eighth convolutional layer and a bilinear interpolation upsampling layer. Among them, the convolutional kernel size of the sixty-eighth convolutional layer is 1×1, the number of convolutional kernels is 2, the stride is 1, and the zero-padding parameter is "same"; the method used in the bilinear interpolation upsampling layer is bilinear interpolation, and the output is feature maps of the same size as the original image. The input end of the second regular output layer receives each feature map in J, and after being processed by the second regular output layer, 2 foreground-background prediction maps corresponding to the original input image are output, denoted as P 2 .

[0061] For the multi-scale output layer, it consists of five independent convolutional layers, namely the sixty-ninth convolutional layer, the seventieth convolutional layer, the seventy-first convolutional layer, the seventy-second convolutional layer, and the seventy-third convolutional layer. The convolutional kernel size of these five convolutional layers is 1×1, the number of convolutional kernels is 9, the stride is 1, and the zero-padding parameter is "same"; the input end of the sixty-ninth convolutional layer receives each feature map in the first side output F 1 and outputs 9 feature maps corresponding to the original image, denoted as M 1 , M 1 The width of each feature map in The height is The input end of the seventieth convolutional layer receives each feature map in the second side output F 2 and outputs 9 feature maps corresponding to the original image, denoted as M 2 , M2 The width of each feature map in is The input end of the seventy - first convolutional layer receives each feature map in the third - side output F 3 and outputs 9 feature maps corresponding to the original image at the output end, denoted as M 3 , M 3 The width of each feature map in is The input end of the seventy - second convolutional layer receives each feature map in the fourth - side output F 4 and outputs 9 feature maps corresponding to the original image at the output end, denoted as M 4 , M 4 The width of each feature map in is The input end of the seventy - third convolutional layer receives the output of the mesh context - aware block and outputs 9 feature maps corresponding to the original image at the output end, denoted as M 5 , M 5 The width of each feature map in is

[0062] Construct a discriminator convolutional neural network:

[0063] The discriminator convolutional neural network is composed of a discriminator input layer, a discriminator first neural network block, a discriminator second neural network block, a discriminator third neural network block, a discriminator first linear transformation block, a discriminator second linear transformation block, and a discriminator third linear transformation block connected in sequence, and the structure is as Figure 6 shown.

[0064] For the discriminator input layer, its input has two cases: one is the predicted map P 1 output by the first conventional output layer in the generator and the original road - scene image. At this time, the output end of the discriminator input layer outputs the superposition of P 1 and the original road - scene image, obtaining 12 feature maps, denoted as P re ; the second case is that the input is the original road - scene image and the corresponding real semantic segmentation image G semantic At this time, the output end of the discriminator input layer outputs the superposition of the two, obtaining 12 feature maps, denoted as G ro .

[0065] For the first neural network block of the discriminator, it is composed of the seventy-fourth convolutional layer, the forty-fifth activation layer, the seventy-fifth convolutional layer, the forty-sixth activation layer, and the third max pooling layer (Max pool) connected in sequence. Among them, the kernel size of the seventy-fourth convolutional layer is 1×1, the number of filters is 3, the stride is 1, and the padding parameter is 0; the kernel size of the seventy-fifth convolutional layer is 3×3, the number of filters is 32, the stride is 1, and the padding parameter is 1; the activation method of the activation layer is "Relu".

[0066] For the second neural network block of the discriminator, it is composed of the seventy-sixth convolutional layer, the forty-seventh activation layer, the seventy-seventh convolutional layer, the forty-eighth activation layer, and the fourth max pooling layer (Max pool) connected in sequence. Among them, the kernel sizes of the seventy-sixth convolutional layer and the seventy-seventh convolutional layer are both 3×3, the numbers of filters are both 64, the strides are both 1, and the padding parameters are both 1; the activation method of the activation layer is "Relu".

[0067] For the third neural network block of the discriminator, it is composed of the seventy-eighth convolutional layer, the forty-ninth activation layer, the seventy-ninth convolutional layer, the fiftieth activation layer, and the fifth max pooling layer (Max pool) connected in sequence. Among them, the kernel sizes of the seventy-eighth convolutional layer and the seventy-ninth convolutional layer are both 3×3, the numbers of filters are both 64, the strides are both 1, and the padding parameters are both 1; the activation method of the activation layer is "Relu".

[0068] For the first linear transformation layer of the discriminator, it is composed of the first linear transformation layer and the forty-ninth activation layer, where the output size of the first linear transformation layer is 100 and the activation method of the forty-ninth activation layer is "Tanh"; for the second linear transformation layer of the discriminator, it is composed of the second linear transformation layer and the fiftieth activation layer, where the output size of the second linear transformation layer is 2 and the activation method of the fiftieth activation layer is "Tanh"; for the third linear transformation layer of the discriminator, it is composed of the third linear transformation layer and the fifty-first activation layer, where the output size of the third linear transformation layer is 1 and the activation method of the fifty-first activation layer is "Sigmoid".

[0069] When the input of the discriminator is P 1 and the original road scene image, the output score of the discriminator is denoted as Score pr ; when the input of the discriminator is G semanticWhen it is the original road scene image, the output score of the discriminator is denoted as Score gt .

[0070] Alternately train the discriminator and the generator.

[0071] Step 1_3:

[0072] First, input the original road scene images and the corresponding heat maps in the training set into the input layer of the generator convolutional neural network. From the first output layer, obtain 9 semantic segmentation prediction maps corresponding to each original road scene image in the training set. Denote the set of semantic segmentation prediction maps composed of the 9 semantic segmentation prediction maps as P 1 t ; Input P 1 t and the original road scene image into the discriminator to obtain the output score of the discriminator Input and the original road scene image into the discriminator to obtain the output score of the discriminator Calculate the final score of the discriminator

[0073] Step 1_4:

[0074] Input the original road scene images and the corresponding heat maps in the training set into the input layer of the generator convolutional neural network. From the first output layer, obtain 9 semantic segmentation prediction maps corresponding to each original road scene image in the training set. Denote the set of semantic segmentation prediction maps composed of the 9 semantic segmentation prediction maps as P 1 t , and obtain 2 segmentation prediction maps corresponding to each original road scene image in the training set from the second output layer. Denote the set composed of these 2 segmentation prediction maps as Output 9 semantic segmentation prediction maps from the first convolutional layer in the multi-scale output layer. Denote the set as Output 9 semantic segmentation prediction maps from the second convolutional layer in the multi-scale output layer. Denote the set as Output 9 semantic segmentation prediction maps from the third convolutional layer in the multi-scale output layer. Denote the set as Output 9 semantic segmentation prediction maps from the fourth convolutional layer in the multi-scale output layer. Denote the set as Output 9 semantic segmentation prediction maps from the fifth convolutional layer in the multi-scale output layer. Denote the set as

[0075] Calculate the set P composed of 9 semantic segmentation prediction maps corresponding to each original road scene image in the training set 1 t and the corresponding semantic segmentation label map set the loss function value between P 1 t and the loss function value between is denoted as obtained by using the Lovász-Softmax loss function; calculate the set consisting of 2 predicted maps corresponding to each original road scene image in the training set and the corresponding foreground-background segmentation label map set the loss function value between, and and the loss function value between is denoted as obtained by using categorical crossentropy; calculate the loss function value between the 9 semantic segmentation predicted maps in and the corresponding multi-scale semantic segmentation label map set the loss function value between, and and the loss function value between is denoted as obtained by using categorical crossentropy; calculate the loss function value between the 9 semantic segmentation predicted maps in and the corresponding multi-scale semantic segmentation label map set the loss function value between, and and the loss function value between is denoted as obtained by using categorical crossentropy; calculate the loss function value between the 9 semantic segmentation predicted maps in and the corresponding multi-scale semantic segmentation label map set the loss function value between, and and the loss function value between is denoted as obtained by using categorical crossentropy; calculate the loss function value between the 9 semantic segmentation predicted maps in and the corresponding multi-scale semantic segmentation label map set the loss function value between, and and the loss function value between is denoted as obtained by using categorical crossentropy; calculate the loss function value between the 9 semantic segmentation predicted maps in and the corresponding multi-scale semantic segmentation label map set the loss function value between, and The loss function value between is denoted as obtained using categorical cross - entropy.

[0076] At this time, the t - th original road scene image, the predicted map P obtained through the generator 1 t and the semantic segmentation label map are input into the discriminator network to obtain the final score of the discriminator The loss function obtained for the t - th image during training is denoted as Loss t ,

[0077] Step 1_5: Alternately execute Step 1_3 and Step 1_4 a total of E times, that is, Step 1_3 and Step 1_4 are each executed E / 2 times. When executing Step 1_3, keep the parameters of the generator unchanged and only update the parameters of the discriminator to make the score minimum. When executing Step 1_4, keep the parameters of the discriminator unchanged and only update the parameters of the generator. Among the E / 2 times of executing Step 1_4, take the time when the training loss Loss t is the minimum, and record the weight vector and bias term in the generator at this time as the optimal weight vector and optimal bias term of the generator training model, denoted as W best and b best respectively; where E > 2, and in this embodiment, E = 200 is taken.

[0078] The specific steps of the above - mentioned test - phase process are as follows:

[0079] Step 2_1: Denote the road scene image to be semantically segmented as T RGB (i', j'), and the corresponding heat map to be semantically segmented as T T (i', j'); where 1 ≤ i' ≤ W', 1 ≤ j' ≤ H', W' represents the width of the road scene image to be semantically segmented, H' represents the height of the road scene image to be semantically segmented, and i', j' respectively represent the horizontal and vertical coordinates of the pixel point at the coordinate position (i', j'). Step 2_2: Input the road scene image to be semantically segmented and its corresponding heat map into the first input layer and the second input layer of the generator training model respectively, and use the optimal weight vector and optimal bias term of the generator obtained in Step 1_5 for prediction, and predict the semantic segmentation image through the first output layer, denoted as Tpret(i', j'), where Tpret(i', j') represents the pixel value of the pixel point at the coordinate position (i', j') in the semantic segmentation prediction image.

[0080] To further verify the feasibility and effectiveness of the method of the present invention, experiments were conducted.

[0081] The architecture of a multi-scale perforated convolutional neural network was built using the deep learning library Pytorch based on python. The public dataset InfRec500 published by Ha Qishen et al. in MFNet was used for training and testing, and 393 road scene images were selected from the test set to test the segmentation effect of the trained model. Here, three commonly used objective parameters for evaluating semantic segmentation methods were used as evaluation indicators, namely Class Acurracy, Mean Pixel Accuracy (MPA), and the ratio of the intersection to the union of the segmented image and the label image (Mean Intersection over Union, MIoU), to evaluate the segmentation performance of the predicted semantic segmentation image. The Class Accuracy (CA), Mean Pixel Accuracy (MPA), and Mean Intersection over Union (MIoU) of the semantic segmentation effect of the method of the present invention are listed in Table 1.

[0082] Table 1 Evaluation results of the method of the present invention on the test set

[0083]

[0084]

[0085] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A road scene semantic segmentation method based on generative adversarial networks and cross-modal feature fusion, characterized in that, the specific steps are as follows: Select multiple original road scene images and the corresponding true semantic segmentation images for each original road scene image, and form a training set from the multiple original road scene images and the label maps corresponding to each original road scene image; Construct a generator convolutional neural network with a cross-modal fusion module, a mesh context awareness module, resolution restoration, and enhanced semantic information; Construct a discriminator convolutional neural network; Input the training set into the generator convolutional neural network to obtain the prediction maps corresponding to each original road scene map in the training set. Input the original road scene map, the corresponding prediction map, and the corresponding label map into the discriminator to obtain an evaluation score. According to the evaluation score, update the discriminator parameters to minimize the score; Input the training set into the generator convolutional neural network to obtain the prediction maps corresponding to each original road scene map in the training set, calculate the loss function between the prediction maps and the label maps, and update the generator parameters; The cross-modal fusion module includes a first fusion module; the first fusion module consists of a third convolutional layer and a fourth convolutional layer; the output R of the first initial neural network block 1 , and the output T of the second initial neural network block 1 are respectively input into the third convolutional layer and the fourth convolutional layer. Both the third convolutional layer and the fourth convolutional layer output 64 feature maps, which are respectively denoted as A 1 and B 1 ; Each feature map in A 1 is element-wise added to each feature map in B 1 to obtain 64 preliminary fusion feature maps, denoted as f 1 . Each feature map in f 1 passes through the Sigmoid activation function to obtain 64 feature maps, denoted as S 1 ; Each feature map in S 1 is element-wise multiplied with each feature map in A 1 and B 1 respectively to obtain two sets of 64 feature maps, which are respectively denoted as SA 1 and SB 1 ; Finally, SA 1 , SB 1 and the preliminary fusion feature map f 1 are element-wise added to obtain 64 feature maps, denoted as the first side output F 1 .

2. The road scene semantic segmentation method based on generative adversarial networks and cross-modal feature fusion according to claim 1, characterized in that, The cross-modal fusion module further includes a second fusion module, a third fusion module, a fourth fusion module, and a fifth fusion module; the second fusion module, the third fusion module, the fourth fusion module, and the fifth fusion module have the same structure; the second fusion module includes a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, and an eleventh convolutional layer; the output R of the first residual neural network block l1 , and the output T of the fifth residual neural network block l1 are respectively input into the fifth convolutional layer and the sixth convolutional layer. Both the fifth convolutional layer and the sixth convolutional layer output 64 feature maps, which are respectively denoted as A l1 and B l1 ; each feature map in A l1 is element-wise added to each feature map in B l1 to obtain 64 preliminary fusion feature maps, denoted as f l1 . Each feature map in f l1 passes through the Sigmoid activation function to obtain 64 feature maps, denoted as S l1 ; each feature map in S l1 is element-wise multiplied with each feature map in A l1 and B l1 respectively to obtain two groups of 64 feature maps, denoted as SA l1 and SB l1 ; each feature map in the first side output F 1 is input into the seventh convolutional layer, and the 64 feature maps output by the seventh convolutional layer are then input into the eighth convolutional layer. The eighth convolutional layer outputs 64 feature maps, denoted as U l1 ; each feature map in the first side output F 1 is input into the ninth convolutional layer, and the 64 feature maps output by the ninth convolutional layer are then input into the tenth convolutional layer. The tenth convolutional layer outputs 64 feature maps, denoted as D l1 . The 64 feature maps in U l1 are superimposed on the 64 feature maps in SA l1 to obtain 128 feature maps, denoted as M l1 ; the 64 feature maps in D l1 are superimposed on the 64 feature maps in SB l1 to obtain 128 feature maps, denoted as N l1 ; each feature map in M l1 is element-wise added to each feature map in N l1 , and the sum obtained is used as the input of the eleventh convolutional layer. The eleventh convolutional layer outputs 64 feature maps denoted as MN l1 . Finally, MN l1 and the preliminary fusion feature map f l1 Perform the summation of the elements to obtain 64 feature maps, denoted as the second side output F 2 。 3. The road scene semantic segmentation method based on generative adversarial networks and cross-modal feature fusion according to claim 1, characterized in that, The said mesh context awareness module includes four branches; the outputs of the four convolutional blocks in the first branch are respectively denoted as B 1-1 、B 1-2 、B 1-3 and B 1-4 , the outputs of the four convolutional blocks in the second branch are respectively denoted as B 2-1 、B 2-2 、B 2-3 and B 2-4 , the outputs of the four convolutional blocks in the third branch are respectively denoted as B 3-1 、B 3-2 、B 3-3 and B 3-4 , in the second branch, the sum of the input F 5 to the mesh context awareness block and B 1-1 is used as the input to the second convolutional block, the sum of B 2-1 and B 1-2 is used as the input to the fourth dilated convolutional block, the sum of B 2-2 and B 1-3 is used as the input to the fifth dilated convolutional block, the sum of B 2-3 and B 1-4 is used as the input to the sixth dilated convolutional block; in the third branch, the sum of F 5 , B 2-1 and B 1-2 is used as the input to the third convolutional block, the sum of B 3-1 , B 2-2 and B 1-3 is used as the input to the seventh dilated convolutional block, the sum of B 3-2 , B 2-3 and B 1-4 is used as the input to the eighth dilated convolutional block, the sum of B 3-3 and B 2-4 is used as the input to the ninth dilated convolutional block; the 64 feature maps in the last convolutional block of the three branches are stacked together to generate 192 feature maps, denoted as CB; CB is used as the input to the thirty-third convolutional layer, the thirty-third convolutional layer outputs 64 feature maps, and then these 64 feature maps are element-wise added to the input F 5 to the mesh context awareness block to obtain 64 feature maps, denoted as F 5 # .

4. The road scene semantic segmentation method based on generative adversarial networks and cross-modal feature fusion according to claim 1, characterized in that, the generator convolutional neural network further includes a decoding block, and the input of the decoding block sequentially passes through three convolutional blocks, is added to the input, and undergoes upsampling to expand the resolution to twice the original.

Citation Information

Patent Citations

  • Unsupervised cross-domain self-adaptive medical image segmentation method based on deep adversarial learning

    AU2020103905A4

  • Multi-modal image fusion method based on generative adversarial network and super-resolution network

    CN109325931A