A road scene semantic segmentation method based on multi-model fusion

By building a multi-classification model and a binary classification model, and combining visual attention and DeepLabV3+ codec structure, multi-model prediction and result fusion is solved, and the problems of insufficient segmentation accuracy and poor connectivity of segmentation results in the existing technology are achieved, achieving higher recognition accuracy and robustness.

CN114693924BActive Publication Date: 2025-06-24NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210246612.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-06-24
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

In the prior art, the road category segmentation accuracy is insufficient, and the segmentation results are poor, especially the segmentation effect is poor for non-linear road surfaces.

Method used

The multi-model fusion method is used to build a multi-classification model and a binary classification model, and the optimal weight value is obtained through end-to-end training, combining visual attention and DeepLabV3+ codec structure to perform multi-model prediction and result fusion.

Benefits of technology

It improves the identification accuracy of road categories and connectivity of segmentation results, enhances the identification accuracy and robustness of road categories, and performs better in non-linear pavement segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693924B_ABST
    Figure CN114693924B_ABST
Patent Text Reader

Abstract

The present invention discloses a road scene semantic segmentation method based on multi-model fusion, and the steps are as follows: building a multi-classification model and a binary classification model; respectively performing end-to-end training on the multi-classification model and the binary classification model to obtain the optimal weight value that minimizes the loss function; using the optimal weight value to perform multi-classification prediction and binary classification prediction on the road scene image to form a preliminary segmentation result map; performing image post-processing on the preliminary segmentation result map formed by the binary classification prediction; and fusing the preliminary segmentation result map formed by the multi-classification prediction and the segmentation result map after image processing. The multi-classification model of the present invention adds visual attention to the feature fusion part on the basis of the original HRNet, so that effective feature maps obtain larger fusion weights, and invalid or poor-effect feature maps obtain smaller fusion weights, improving the pixel representation ability of the multi-classification model and obtaining better segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a road scene semantic segmentation method based on multi-model fusion. Background Art

[0002] Semantic segmentation is an important task in the technical field of computer vision. In the semantic segmentation task, we need to classify the input image into different semantically interpretable categories.

[0003] Traditional semantic segmentation usually adopts methods such as SVM classification and structured random forests. These algorithms usually have disadvantages such as low recognition efficiency, low accuracy, and poor robustness.

[0004] With the increasingly widespread application of deep learning, semantic segmentation methods based on end-to-end training of convolutional neural networks have become more and more common. Using deep learning methods to perform semantic segmentation on images is more convenient and fast, and has gradually become the mainstream method for semantic segmentation. The initial deep learning method applied to image segmentation was an image patch-based classification algorithm. However, in this algorithm, the fully connected layer (FC layer) limits the size of the input image. The fully convolutional network makes it possible to perform semantic segmentation on input images of any size, and is now widely adopted and continuously improved.

[0005] Autopilot is an important application field of semantic segmentation. By performing pixel-level classification on pictures, the computer can understand the semantic information on a picture, such as distinguishing the pixels corresponding to the road surface, vehicles, non-motor vehicles, and pedestrians in the picture, and classifying them into corresponding label categories. These semantic information can be transferred to algorithms for other tasks, such as lane line detection and traffic target detection, for further information extraction.

[0006] Among the numerous recognition categories in the semantic segmentation task in the autopilot scenario, the road surface (road) is an important category. By segmenting the road surface part, the computer can extract the area where the vehicle can travel, so as to make further plans for the vehicle's driving trajectory. Therefore, in the semantic segmentation task, there are higher requirements for the classification accuracy of this category of road surface. Most of the existing road scene semantic segmentation methods have insufficient segmentation effects on this category of road surface, poor connectivity of the road surface segmentation results, and poor performance in segmenting non-straight road surfaces. Summary of the Invention

[0007] Aiming at the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a road scene semantic segmentation method based on multi-model fusion to solve the problems of insufficient segmentation accuracy of road categories and poor connectivity of segmentation results in the existing technologies.

[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A method for semantic segmentation of road scenes based on multi-model fusion according to the present invention comprises the following steps:

[0010] 1) Build a multi-classification model and a binary classification model;

[0011] 2) Perform end-to-end training on the multi-classification model and the binary classification model respectively to obtain the optimal weight values that minimize the loss function;

[0012] 3) Use the optimal weight values to perform multi-classification prediction and binary classification prediction on the road scene image to form a preliminary segmentation result map;

[0013] 4) Perform image post-processing on the preliminary segmentation result map formed by the binary classification prediction in step 3);

[0014] 5) Fuse the preliminary segmentation result map formed by the multi-classification prediction in step 3) and the processed segmentation result map in step 4).

[0015] Further, step 1) specifically includes:

[0016] 11) Build a multi-classification model based on an improved high-resolution network; introduce visual attention, and the multi-classification model outputs a pixel-level label image to predict the category to which the pixel belongs;

[0017] 12) Build a binary classification model based on the encoder-decoder structure of DeepLabV3+; the binary classification model outputs the prediction result of the road category.

[0018] Further, step 11) specifically includes:

[0019] The multi-classification model built based on the improved high-resolution network: remove the last feature fusion unit of the second, third, and fourth sub-networks of the original high-resolution network; introduce visual attention into each feature fusion unit;

[0020] In the original high-resolution network, there are 4 parallel sub-networks. From left to right, the size of the feature map in each sub-network is 1 / 2 of the previous sub-network, and the number of channels of the feature map is 2 times that of the previous sub-network; each sub-network respectively includes repeated multi-resolution units and feature fusion units; before each multi-resolution unit, there is a feature fusion unit; the multi-resolution unit includes 4 repeated convolutional units; the feature fusion unit includes an upsampling / downsampling layer and an addition fusion layer; the input end of the upsampling / downsampling layer is connected to the output end of the multi-resolution unit of each sub-network in the previous layer to perform upsampling or downsampling of the input feature map at the corresponding scale;

[0021] In the last feature fusion unit of each sub-network in the improved high-resolution network, a transposed convolutional unit is added to introduce visual attention, so as to improve the detection accuracy and speed of the multi-classification model; the last feature fusion units of the second, third, and fourth sub-networks are removed, the final output of the first sub-network is connected to the transposed convolutional unit, the number of channels of the feature map is converted into the number of corresponding semantic segmentation categories, and the size of the feature map is restored to the same size as the original input picture; the transposed convolutional unit includes a transposed convolutional layer with a convolution kernel size of 1×1 and a stride of 1 and a bilinear interpolation upsampling layer;

[0022] Visual attention is added between the input end of the feature fusion unit and the upsampling / downsampling layer, which is used to adjust the model weights to strengthen the visual features and weaken other unimportant features, so as to improve the feature extraction ability of the model; specifically, the visual attention is to input the feature map with a size of W×H×C input by the feature fusion unit into the global average pooling layer, the output data with a size of 1x1xC passes through two fully connected layers, and finally passes through the Sigmoid function to limit the value of the data to the interval range of [0, 1], and multiply this value by the data of the C channels of the original input feature map as the input data of the next-level upsampling / downsampling layer.

[0023] Further, step 12) specifically includes:

[0024] The binary classification model built based on the deeplabv3+ encoder-decoder structure includes: an encoder and a decoder, and the encoder includes a feature information extraction unit and an atrous spatial pyramid pooling unit; the atrous spatial pyramid pooling unit is connected to the feature information extraction unit; the decoder includes a skip connection unit, which extracts and fuses multi-scale feature information and shallow feature information as the output of the binary classification model, the multi-scale feature information is extracted by the atrous spatial pyramid unit, and the shallow information is extracted by the shallow part of the feature information extraction unit;

[0025] The feature information extraction unit is based on the lightweight network ShuffleNetV2 and is composed of a sequentially connected convolutional compression unit, 3 Shufflenet units, and a transposed convolutional unit; the convolutional compression unit includes a convolutional layer with a convolution kernel size of 3×3 and a stride of 1 and a pooling layer with a pooling kernel size of 3×3 and a stride of 2, and the pooling layer performs one downsampling on the feature information output by the convolutional layer; each Shufflenet unit performs one downsampling; the transposed convolutional unit is composed of a convolutional layer with a convolution kernel size of 1×1 and a stride of 1;

[0026] The atrous spatial pyramid pooling unit consists of parallel atrous convolutional layers with atrous rates of 1, 6, 12, and 18 in sequence, a global average pooling layer, an upsampling layer, and a splicing and fusion layer. The input end of the upsampling layer is connected to the global average pooling layer for bilinear interpolation upsampling to obtain feature information with the same size as the feature information output by the atrous convolutional layer. The input ends of the splicing and fusion layer are respectively connected to the output segments of the four atrous convolutional layers and the output end of the upsampling layer to splice and fuse the feature information output by the atrous convolutional layer and the upsampling layer.

[0027] Further, the skip connection unit includes: a shallow transposed convolutional layer, a deep transposed convolutional unit, and a fusion unit. The input end of the shallow transposed convolutional layer is connected to the end of the first Shufflenet unit, and the output end is connected to the fusion unit. The deep transposed convolutional unit includes a convolutional layer with a convolution kernel size of 1×1 and a stride of 1 and a bilinear interpolation sampling layer. The input end of the convolutional layer is connected to the end of the atrous spatial pyramid pooling unit, and the output end of the bilinear interpolation is connected to the fusion unit. The fusion unit includes a splicing and fusion layer and a bilinear interpolation upsampling layer.

[0028] Further, step 2) specifically includes:

[0029] 21) Establish datasets for the multi-classification model and the binary classification model, and perform data augmentation on the datasets.

[0030] 22) Use the augmented datasets to perform end-to-end training on the constructed multi-classification model and binary classification model to obtain the optimal weight values when the loss function is minimized.

[0031] Further, step 21) specifically includes:

[0032] Adopt the cityscapes dataset, where the dataset contains 34 categories. Use the one-hot encoding method to convert the real semantic segmentation image into a one-hot encoded real semantic segmentation image. Back up the original image and the corresponding real semantic segmentation image as the initial dataset of the multi-classification model, and perform data augmentation on the initial dataset of the multi-classification model, including horizontal flipping, vertical flipping, and scaling, as the dataset of the multi-classification model.

[0033] Convert the true semantic segmentation images in the initial dataset of the multi-classification model backed up in the above operations into binary-class true semantic segmentation images, set the road category as the foreground, and set other categories as the background; perform threshold screening on the converted image data, retain the images with the pixel area ratio of the road category greater than a certain proportion, and use the screened true semantic segmentation images and their corresponding original images as the initial dataset of the binary-class model; perform data augmentation on the initial dataset of the binary-class model, including horizontal flipping, vertical flipping, and scaling, as the dataset of the binary-class model.

[0034] Further, the step 22) specifically includes:

[0035] Input the original images in the multi-classification model dataset into the multi-classification model for image semantic segmentation prediction to obtain the multi-classification model prediction images; compare them with the true semantic segmentation images in the multi-classification model dataset, calculate the loss value between the predicted value and the true value through the loss function, and according to the calculated loss value, use the gradient descent method of backpropagation and the Adam optimizer to iteratively update the network parameters. Adjust the learning rate using the cosine annealing strategy during each iteration until the network converges or reaches the set number of iterations, and finally obtain the optimal network parameter weight value that minimizes the loss value.

[0036] Among them, the loss function uses the Softmax function combined with the cross-entropy loss function, specifically as follows:

[0037] The Softmax function compresses a K-dimensional real vector into a new K-dimensional real vector in the range [0-1], and the function formula is:

[0038]

[0039] In the formula, K is the number of dataset categories, z c is the predicted value of the multi-classification model in the channel where the c-th semantic segmentation category is located;

[0040] z k is the predicted value of the multi-classification model in the channel where the k-th semantic segmentation category is located, and e is a constant;

[0041] The formula of the cross-entropy loss function is:

[0042]

[0043] In the formula, N is the number of samples in a training batch, M is the number of semantic segmentation categories, y i is the true value of the true semantic segmentation image, is the predicted value, that is, the result obtained by passing the predicted value of the multi-classification model through the above Softmax function.

[0044] Further, step 22) specifically further includes:

[0045] The original images in the binary classification model dataset are passed through the backbone network ShuffleNetV2 and the atrous spatial pyramid pooling unit of the binary classification model to obtain feature maps, and then after encoder upsampling and skip connections, semantic segmentation prediction is performed to obtain the binary classification model prediction images; these are compared with the true semantic segmentation images in the binary classification model dataset, and the loss value between the predicted value and the true value is calculated through a loss function. According to the calculated loss value, the gradient descent method of backpropagation is used and the Adam optimizer is utilized to iteratively update the network parameters. Each time an iteration is performed, the learning rate is adjusted using the cosine annealing strategy until the network converges or reaches the set number of iterations, and finally the optimal network parameter weight value that minimizes the loss value is obtained;

[0046] Among them, the loss function uses the Sigmoid function combined with the binary cross-entropy loss function:

[0047] The Sigmoid function maps the output to the interval [0, 1]. The formula for the Sigmoid function is:

[0048]

[0049] In the formula, x is the predicted value output by the binary classification model;

[0050] The formula for the binary cross-entropy loss function is:

[0051]

[0052] In the formula, N is the number of samples in a training batch, w is a hyperparameter, y n is the true value of the true semantic segmentation image, and x n is the value obtained by passing the predicted value of the binary classification model through the above sigmoid function.

[0053] Further, step 3) specifically includes:

[0054] 31) Load the optimal weight value of the multi-classification model obtained in step 2) into the multi-classification model, input the road scene image to be detected into the multi-classification model, perform semantic segmentation through the neural network to obtain the multi-classification prediction image; use the Argmax function to convert the multi-classification prediction image into a single-channel multi-classification prediction image;

[0055] 32) Load the optimal weight value of the binary classification model obtained in step 2) into the binary classification model, input the road scene image to be detected into the binary classification model, perform semantic segmentation through the neural network to obtain the binary classification prediction image.

[0056] Further, step 4) specifically includes:

[0057] 41) Use the morphologyEx function in the opencv library to perform a closing operation on the binary classification prediction image output in step 3) to connect the broken parts; use the medianBlur function in the opencv library to perform median filtering on the result of the above operation to remove burrs;

[0058] 42) Use the findContours function in the opencv library to extract the contour information output in step 41); screen out isolated pixel clusters by setting the area and length thresholds of the contours, and remove the isolated pixel clusters smaller than the thresholds;

[0059] 43) Extract the point set of the road category in the image output in step 42); use the morphologyEx function in the opencv library to perform a closing operation on the extracted point set; for the output result of the above operation, use the skeletonize function to extract the skeleton of the road category; use the morphologyEx function in the opencv library to perform dilation and erosion operations on the extracted skeleton to ensure connectivity while not exceeding the prediction area of the original binary classification model too much.

[0060] Further, step 5) specifically includes:

[0061] Fuse the binary classification prediction result of the image post - processing obtained in step 4) with the pixels of the corresponding road category in the multi - classification model prediction result obtained in step 3) to obtain a fused prediction result;

[0062] The calculation formula for the fused prediction result is as follows:

[0063]

[0064] In the formula, is the prediction result of the multi - classification model, is the prediction result of the binary classification model of the image post - processing obtained in step 4).

[0065] Advantages of the present invention:

[0066] The present invention uses ensemble learning to fuse the prediction results of different models, improving the recognition accuracy of the road category compared with other semantic segmentation models for road scenes, and at the same time improving the connectivity of the road segmentation result, specifically manifested in several aspects:

[0067] (1) The multi-classification model of the present invention adds visual attention (SEAttention) to the feature fusion part on the basis of the original HRNet, so that effective feature maps obtain larger fusion weights, and ineffective or poor-effect feature maps obtain smaller fusion weights, improving the pixel representation ability of the multi-classification model and obtaining better segmentation results.

[0068] (2) The present invention uses a binary classification model to solve the problems of road category recognition accuracy and recognition result connectivity in road scene semantic segmentation. The binary classification model network is built based on the lightweight ShffleNetV2, improving the operation speed of the model.

[0069] (3) The present invention combines the multi-classification model with the binary classification model, fuses the prediction results, and the binary classification network conducts targeted training for the road category, with higher recognition accuracy. The recognition accuracy and robustness of the network model for road surface categories in road scenes are improved through network integrated prediction.

[0070] (4) The present invention adds a post-processing link after the binary classification neural network prediction, further increasing the connectivity of the recognition results for the road category, and at the same time further improving the recognition accuracy and edge accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is a flowchart of the method of the present invention.

[0072] Figure 2 is the network structure diagram of the original high-resolution network HRNet.

[0073] Figure 3 is the improved HRNet network structure diagram of the present invention.

[0074] Figure 4 is the visual attention structure diagram in the multi-classification model of the present invention.

[0075] Figure 5 is the network structure diagram of the binary classification model of the present invention.

[0076] Figure 6 is the image post-processing flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0077] For the convenience of those skilled in the art, the present invention will be further described below in conjunction with the embodiments and the drawings. The content mentioned in the embodiments does not limit the present invention.

[0078] Referring to Figure 1 as shown, a method for road scene semantic segmentation based on multi-model fusion of the present invention is as follows:

[0079] 1) Build a multi-classification model and a binary-classification model; specifically including:

[0080] 11) Build a multi-classification model based on the improved High-Resolution Network (HRNet); introduce visual attention, and the multi-classification model outputs a pixel-level label image to predict the category to which the pixel belongs;

[0081] 12) Build a binary-classification model based on the encoder-decoder structure of DeepLabV3+; the binary-classification model outputs the prediction result of the road category.

[0082] Among them, the specific steps of step 11) include:

[0083] The multi-classification model built based on the improved High-Resolution Network: remove the last feature fusion unit of the second, third, and fourth sub-networks of the original High-Resolution Network; introduce visual attention (SEAttention) in each feature fusion unit;

[0084] As Figure 2 shown, in the original High-Resolution Network, there are 4 parallel sub-networks. From left to right, the size of the feature map in each sub-network is 1 / 2 of the previous sub-network, and the number of channels of the feature map is 2 times that of the previous sub-network; each sub-network contains repeated multi-resolution units and feature fusion units; before each multi-resolution unit, there is a feature fusion unit; the multi-resolution unit includes 4 repeated convolutional units; the feature fusion unit includes an upsampling / downsampling layer and an addition fusion layer; the input end of the upsampling / downsampling layer is connected to the output end of the multi-resolution unit of each sub-network in the previous layer, and performs upsampling or downsampling of the input feature map at the corresponding scale;

[0085] Add a transposed convolutional unit to the last feature fusion unit of each sub-network in the improved High-Resolution Network, introduce visual attention to improve the detection accuracy and detection speed of the multi-classification model; as Figure 3 shown, remove the last feature fusion unit of the second, third, and fourth sub-networks, connect the last output of the first sub-network to the transposed convolutional unit, convert the number of channels of the feature map to the number of corresponding semantic segmentation categories, and restore the size of the feature map to the same size as the original input picture; the transposed convolutional unit includes a transposed convolutional layer with a convolution kernel size of 1×1 and a stride of 1 and a bilinear interpolation upsampling layer;

[0086] As Figure 4As shown in the figure, visual attention is added between the input end of the feature fusion unit and the upsampling / downsampling layer to adjust the model weights to strengthen the visual features and weaken other unimportant features, so as to improve the feature extraction ability of the model. The visual attention specifically inputs the feature map with a size of W×H×C input by the feature fusion unit into the Global Average Pooling Layer, and the output data with a size of 1x1xC passes through two fully connected layers (FC layer), and finally passes through the Sigmoid function to limit the value of the data to the interval range of [0, 1]. Multiply this value by the data of the C channels of the original input feature map as the input data of the next upsampling / downsampling layer.

[0087] As Figure 5 shown, step 12) specifically includes:

[0088] The binary classification model based on the deeplabv3+ encoder-decoder structure includes an encoder and a decoder. The encoder includes a feature information extraction unit and an Atrous Spatial Pyramid Pooling unit (ASPP); the Atrous Spatial Pyramid Pooling unit is connected to the feature information extraction unit; the decoder includes a skip connection unit that extracts and fuses multi-scale feature information and shallow feature information as the output of the binary classification model. The multi-scale feature information is extracted by the Atrous Spatial Pyramid unit, and the shallow information is extracted by the shallow part of the feature information extraction unit;

[0089] The feature information extraction unit is based on the lightweight network ShuffleNetV2 and is composed of a sequentially connected convolutional compression unit, 3 Shufflenet units, and a transposed convolutional unit. The convolutional compression unit includes a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1 and a pooling layer with a pooling kernel size of 3×3 and a stride of 2. The pooling layer performs a downsampling on the feature information output by the convolutional layer; each Shufflenet unit performs a downsampling; the transposed convolutional unit is composed of a convolutional layer with a convolutional kernel size of 1×1 and a stride of 1;

[0090] The Atrous Spatial Pyramid Pooling unit is composed of parallel atrous convolutional layers with atrous rates of 1, 6, 12, and 18, a global average pooling layer, an upsampling layer, and a splicing and fusion layer. The input end of the upsampling layer is connected to the global average pooling layer, and bilinear interpolation upsampling is performed to obtain feature information with the same size as the feature information output by the atrous convolutional layer; the input end of the splicing and fusion layer is respectively connected to the output segments of the four atrous convolutional layers and the output end of the upsampling layer, and the feature information output by the atrous convolutional layer and the upsampling layer is spliced and fused.

[0091] The jump link unit includes: a shallow transposed convolutional layer, a deep transposed convolutional unit, and a fusion unit; the input end of the shallow transposed convolutional layer is connected to the end of the first Shufflenet unit, and the output end is connected to the fusion unit; the deep transposed convolutional unit includes a convolutional layer with a convolutional kernel size of 1×1 and a stride of 1 and a bilinear interpolation sampling layer, the input end of the convolutional layer is connected to the end of the atrous spatial pyramid pooling unit, and the output end of the bilinear interpolation is connected to the fusion unit; the fusion unit includes a splicing fusion layer and a bilinear interpolation upsampling layer.

[0092] 2) End-to-end train the multi-classification model and the binary classification model respectively to obtain the optimal weight values that minimize the loss function.

[0093] Among them, step 2) specifically includes:

[0094] 21) Establish datasets for the multi-classification model and the binary classification model, and perform data augmentation on the datasets.

[0095] 22) Use the augmented datasets to perform end-to-end training on the constructed multi-classification model and binary classification model, and obtain the optimal weight values when the loss function is minimized.

[0096] Specifically, step 21) specifically includes:

[0097] Adopt the cityscapes dataset, where the dataset contains 34 categories. Use the one-hot encoding method to convert the real semantic segmentation image into a real semantic segmentation image in one-hot encoding form. Back up the original image and the corresponding real semantic segmentation image as the initial dataset of the multi-classification model. Perform data augmentation on the initial dataset of the multi-classification model, including horizontal flipping, vertical flipping, and scaling, as the dataset of the multi-classification model.

[0098] Convert the real semantic segmentation image in the initial dataset of the multi-classification model backed up in the above operation into a binary classification real semantic segmentation image, set the road category as the foreground, and set other categories as the background; perform threshold screening on the converted image data, retain the pictures with the pixel area ratio of the road category greater than a certain proportion, and use the screened real semantic segmentation image and its corresponding original image as the initial dataset of the binary classification model; perform data augmentation on the initial dataset of the binary classification model, including horizontal flipping, vertical flipping, and scaling, as the dataset of the binary classification model.

[0099] Specifically, step 22) specifically includes:

[0100] Input the original images in the multi-classification model dataset into the multi-classification model for image semantic segmentation prediction to obtain the multi-classification model prediction images; compare them with the true semantic segmentation images in the multi-classification model dataset, calculate the loss value between the predicted value and the true value through the loss function, and according to the calculated loss value, use the gradient descent method of backpropagation and the Adam optimizer to iteratively update the network parameters. Adjust the learning rate using the cosine annealing strategy during each iteration until the network converges or reaches the set number of iterations, and finally obtain the optimal network parameter weight value that minimizes the loss value;

[0101] Among them, the loss function adopts the Softmax function combined with the cross-entropy loss function (CrossentropyLoss), specifically as follows:

[0102] The Softmax function compresses a K-dimensional real vector into a new K-dimensional real vector in the range [0-1], and the function formula is:

[0103]

[0104] In the formula, K is the number of dataset categories, and z c is the predicted value of the multi-classification model in the channel where the c-th semantic segmentation category is located;

[0105] z k is the predicted value of the multi-classification model in the channel where the k-th semantic segmentation category is located, and e is a constant;

[0106] The formula of the cross-entropy loss function is:

[0107]

[0108] In the formula, N is the number of samples in a training batch, M is the number of semantic segmentation categories, and y i is the ground truth of the true semantic segmentation image, is the predicted value, that is, the result obtained by passing the predicted value of the multi-classification model through the above Softmax function.

[0109] The specific step 22) further includes:

[0110] The original images in the binary classification model dataset are passed through the backbone network ShuffleNetV2 and the Atrous Spatial Pyramid Pooling (ASPP) unit of the binary classification model to obtain feature maps. After upsampling by the encoder and skip connections, semantic segmentation prediction is performed to obtain the predicted image of the binary classification model. The predicted image is compared with the true semantic segmentation image in the binary classification model dataset, and the loss value between the predicted value and the true value is calculated through the loss function. According to the calculated loss value, the gradient descent method of backpropagation is used and the Adam optimizer is utilized to iteratively update the network parameters. The learning rate is adjusted using the cosine annealing strategy during each iteration until the network converges or reaches the set number of iterations. Finally, the optimal network parameter weight value that minimizes the loss value is obtained.

[0111] Among them, the loss function uses the Sigmoid function combined with the binary cross-entropy loss function:

[0112] The Sigmoid function maps the output to the interval [0, 1]. The formula for the Sigmoid function is:

[0113]

[0114] In the formula, x is the predicted value output by the binary classification model;

[0115] The formula for the binary cross-entropy loss function is:

[0116]

[0117] In the formula, N is the number of samples in a training batch, w is a hyperparameter, y n is the true value of the true semantic segmentation image, and x n is the value obtained by passing the predicted value of the binary classification model through the above sigmoid function.

[0118] 3) Use the optimal weight value to perform multi-classification prediction and binary classification prediction on the road scene image to form a preliminary segmentation result map;

[0119] 31) Load the optimal weight value of the multi-classification model obtained in step 2) into the multi-classification model, input the road scene image to be detected into the multi-classification model, perform semantic segmentation through the neural network to obtain the multi-classification predicted image; use the Argmax function to convert the multi-classification predicted image into a single-channel multi-classification predicted image;

[0120] 32) Load the optimal weight value of the binary classification model obtained in step 2) into the binary classification model, input the road scene image to be detected into the binary classification model, perform semantic segmentation through the neural network to obtain the binary classification predicted image.

[0121] 4) Perform image post - processing on the preliminary segmentation result map formed by the binary classification prediction in step 3).

[0122] 41) Use the morphologyEx function in the opencv library to perform a closing operation on the binary classification prediction image output in step 3) to connect the breaks; use the medianBlur function in the opencv library to perform median filtering on the result of the above operation to remove burrs.

[0123] 42) Use the findContours function in the opencv library to extract the contour information output in step 41); screen out isolated pixel clusters by setting the area and length thresholds of the contours, and remove the isolated pixel clusters smaller than the thresholds.

[0124] 43) Extract the point set of the road category in the image output in step 42); use the morphologyEx function in the opencv library to perform a closing operation on the extracted point set; for the output result of the above operation, use the skeletonize function to extract the skeleton of the road category; use the morphologyEx function in the opencv library to perform dilation and erosion operations on the extracted skeleton to ensure connectivity while not exceeding the prediction area of the original binary classification model excessively.

[0125] 5) Fuse the preliminary segmentation result map formed by the multi - classification prediction in step 3) and the processed segmentation result map in step 4).

[0126] Specifically, fuse the binary classification prediction result of the image post - processing obtained in step 4) with the pixels of the corresponding road category in the multi - classification model prediction result obtained in step 3) to obtain a fused prediction result.

[0127] The calculation formula of the fused prediction result is as follows:

[0128]

[0129] In the formula, is the prediction result of the multi - classification model, is the prediction result of the binary classification model of the image post - processing obtained in step 4).

[0130] There are many specific application ways of the present invention. The above - mentioned is only the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements can still be made, and these improvements should also be regarded as the protection scope of the present invention.

Claims

1. A road scene semantic segmentation method based on multi-model fusion, characterized in that The steps are as follows: 1) Build a multi-classification model and a binary-classification model; 2) Conduct end-to-end training on the multi-classification model and the binary-classification model respectively to obtain the optimal weight values that minimize the loss function; 3) Use the optimal weight values to perform multi-classification prediction and binary-classification prediction on road scene images to form a preliminary segmentation result map; 4) Perform image post-processing on the preliminary segmentation result map formed by the binary-classification prediction in step 3); 5) Fuse the preliminary segmentation result map formed by the multi-classification prediction in step 3) and the processed segmentation result map in step 4); The specific content of step 1) includes: 11) Build a multi-classification model based on an improved high-resolution network; introduce visual attention, and the multi-classification model outputs a pixel-level label image to predict the category to which the pixel belongs; 12) Build a binary-classification model based on the encoder-decoder structure of DeepLabV3+; the binary-classification model outputs the prediction result of the road category; The specific content of step 4) includes: 41) Use the morphologyEx function in the opencv library to perform closing operation on the binary-classification prediction image output in step 3) to connect the broken parts; use the medianBlur function in the opencv library to perform median filtering on the operation result to remove burrs; 42) Use the findContours function in the opencv library to extract the contour information output in step 41); screen out isolated pixel clusters by setting the area and length thresholds of the contours, and remove the isolated pixel clusters smaller than the thresholds; 43) Extract the point set of the road category in the image output in step 42); use the morphologyEx function in the opencv library to perform closing operation on the extracted point set; for the output result of the above operation, use the skeletonize function to extract the skeleton of the road category; use the morphologyEx function in the opencv library to perform dilation and erosion operations on the extracted skeleton to ensure connectivity while not exceeding the prediction area of the original binary-classification model excessively.

2. The method for semantic segmentation of road scenes based on multi-model fusion according to claim 1, characterized in that The specific content of step 11) includes: The multi-classification model built based on the improved high-resolution network: remove the last feature fusion unit of the second, third, and fourth sub-networks of the original high-resolution network; introduce visual attention in each feature fusion unit; In the original high-resolution network, there are 4 parallel sub-networks. From left to right, the size of the feature map in each sub-network is 1 / 2 of the previous sub-network, and the number of channels of the feature map is 2 times that of the previous sub-network; each sub-network contains repeated multi-resolution units and feature fusion units; before each multi-resolution unit, there is a feature fusion unit; the multi-resolution unit includes 4 repeated convolutional units; the feature fusion unit includes an upsampling / downsampling layer and an addition fusion layer; the input end of the upsampling / downsampling layer is connected to the output end of the multi-resolution unit of each sub-network in the previous layer to perform upsampling or downsampling of the input feature map at the corresponding scale; In the last feature fusion unit of each sub-network in the improved high-resolution network, a transposed convolutional unit is added to introduce visual attention to improve the detection accuracy and speed of the multi-classification model. The last feature fusion units of the second, third, and fourth sub-networks are removed, and the last output of the first sub-network is connected to the transposed convolutional unit to convert the number of channels of the feature map into the corresponding number of semantic segmentation categories and restore the size of the feature map to the same size as the original input image. The transposed convolutional unit includes a transposed convolutional layer with a convolutional kernel size of 1×1 and a stride of 1 and a bilinear interpolation upsampling layer. Visual attention is added between the input end of the feature fusion unit and the upsampling / downsampling layer to adjust the model weights to strengthen the visual features and weaken other unimportant features, so as to improve the feature extraction ability of the model. The visual attention specifically means that the feature map with a size of W×H×C input to the feature fusion unit is input to the global average pooling layer, and the output data with a size of 1x1xC is then passed through two fully connected layers. Finally, the data value is restricted to the interval range of [0, 1] through the Sigmoid function, and this value is multiplied by the data of the C channels of the original input feature map as the input data of the next upsampling / downsampling layer.

3. The method for road scene semantic segmentation based on multi-model fusion according to claim 2, wherein The specific content of step 12) includes: The binary classification model based on the deeplabv3+ encoding-decoding structure includes an encoder and a decoder. The encoder includes a feature information extraction unit and an atrous spatial pyramid pooling unit. The atrous spatial pyramid pooling unit is connected to the feature information extraction unit. The decoder includes a skip connection unit that extracts and fuses multi-scale feature information and shallow feature information as the output of the binary classification model. The multi-scale feature information is extracted by the atrous spatial pyramid unit, and the shallow information is extracted by the shallow part of the feature information extraction unit. The feature information extraction unit is based on the lightweight network ShuffleNetV2 and consists of a sequentially connected convolutional compression unit, 3 Shufflenet units, and a transposed convolutional unit. The convolutional compression unit includes a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1 and a pooling layer with a pooling kernel size of 3×3 and a stride of 2. The pooling layer performs one downsampling on the feature information output by the convolutional layer. Each Shufflenet unit performs one downsampling. The transposed convolutional unit consists of a convolutional layer with a convolutional kernel size of 1×1 and a stride of 1. The atrous spatial pyramid pooling unit consists of parallel atrous convolutional layers with atrous rates of 1, 6, 12, and 18, a global average pooling layer, an upsampling layer, and a splicing and fusion layer. The input end of the upsampling layer is connected to the global average pooling layer for bilinear interpolation upsampling to obtain feature information with the same size as the feature information output by the atrous convolutional layer. The input ends of the splicing and fusion layer are respectively connected to the output segments of the four atrous convolutional layers and the output end of the upsampling layer to splice and fuse the feature information output by the atrous convolutional layer and the upsampling layer.

4. The method for semantic segmentation of road scenes based on multi-model fusion according to claim 1, wherein The specific steps of step 2) include: 21) Establish datasets for the multi-classification model and the binary-classification model, and perform data augmentation on the datasets; 22) Use the augmented datasets to conduct end-to-end training on the constructed multi-classification model and binary-classification model, and obtain the optimal weight values when the loss function is minimized.

5. The method for semantic segmentation of road scenes based on multi-model fusion according to claim 4, wherein, The specific steps of step 21) include: Adopt the cityscapes dataset, which contains 34 categories. Use the one-hot encoding method to convert the real semantic segmentation image into a one-hot encoded real semantic segmentation image. Backup the original image and the corresponding real semantic segmentation image as the initial dataset for the multi-classification model. Perform data augmentation on the initial dataset for the multi-classification model, including horizontal flipping, vertical flipping, and scaling, as the dataset for the multi-classification model; Convert the real semantic segmentation image in the initial dataset for the multi-classification model backed up in the above operation into a binary-classification real semantic segmentation image, set the road category as the foreground, and set other categories as the background; perform threshold screening on the converted image data, retain the pictures in which the pixel area ratio of the road category is greater than a certain proportion, and use the screened real semantic segmentation image and its corresponding original image as the initial dataset for the binary-classification model; perform data augmentation on the initial dataset for the binary-classification model, including horizontal flipping, vertical flipping, and scaling, as the dataset for the binary-classification model.

6. The method for semantic segmentation of road scenes based on multi-model fusion according to claim 5, wherein The specific steps of step 22) include: Input the original image in the multi-classification model dataset into the multi-classification model for image semantic segmentation prediction to obtain the multi-classification model prediction image; compare it with the real semantic segmentation image in the multi-classification model dataset, calculate the loss value between the predicted value and the real value through the loss function, and according to the calculated loss value, use the gradient descent method of backpropagation and the Adam optimizer to iteratively update the network parameters. Adjust the learning rate using the cosine annealing strategy during each iteration until the network converges or reaches the set number of iterations, and finally obtain the optimal network parameter weight value that minimizes the loss value; The loss function uses the Softmax function combined with the cross-entropy loss function, specifically as follows: The Softmax function compresses a K-dimensional real vector into a new K-dimensional real vector in the range [0 - 1], and the function formula is: where K is the number of dataset categories, and z c is the predicted value of the multi-classification model in the channel where the c-th semantic segmentation category is located; z k is the predicted value of the multi-classification model in the channel where the k-th semantic segmentation category is located, and e is a constant; The formula for the cross-entropy loss function is: where N is the number of samples in a training batch, M is the number of semantic segmentation categories, and y i is the ground truth of the true semantic segmentation image, is the predicted value, that is, the result obtained by passing the predicted value of the multi-classification model through the above Softmax function.

7. The method for semantic segmentation of road scenes based on multi-model fusion according to claim 6, characterized in that, The specific steps of step 22) also include: Input the original image in the binary-classification model dataset into the backbone network ShuffleNetV2 and the atrous spatial pyramid pooling unit of the binary-classification model to obtain a feature map, and then perform semantic segmentation prediction after upsampling by the encoder and skip connection to obtain the binary-classification model prediction image; compare it with the real semantic segmentation image in the binary-classification model dataset, calculate the loss value between the predicted value and the real value through the loss function, and according to the calculated loss value, use the gradient descent method of backpropagation and the Adam optimizer to iteratively update the network parameters. Adjust the learning rate using the cosine annealing strategy during each iteration until the network converges or reaches the set number of iterations, and finally obtain the optimal network parameter weight value that minimizes the loss value; The loss function adopts a Sigmoid function combined with a binary cross-entropy loss function: The Sigmoid function maps the output to the interval of [0, 1], and the formula of the Sigmoid function is: In the formula, x is the predicted value output by the binary classification model; The formula of the binary cross-entropy loss function is: where N is the number of samples in a training batch, w is a hyperparameter, and y n is the ground truth of the true semantic segmentation image, and x n is the value obtained by passing the prediction value of the binary classification model through the above sigmoid function.

8. The method for semantic segmentation of road scenes based on multi-model fusion according to claim 1, wherein The specific content of step 3) includes: 31) Load the optimal weight value of the multi-classification model obtained in step 2) into the multi-classification model, input the road scene image to be detected into the multi-classification model, perform semantic segmentation through the neural network, and obtain a multi-classification prediction image; use the Argmax function to convert the multi-classification prediction image into a single-channel multi-classification prediction image; 32) Load the optimal weight value of the binary classification model obtained in step 2) into the binary classification model, input the road scene image to be detected into the binary classification model, perform semantic segmentation through the neural network, and obtain a binary classification prediction image.

Citation Information

Patent Citations

  • Unmanned real-time road scene semantic segmentation method based on multi-task supervision

    CN112699889A