Road lightweight feasible region segmentation method based on improved DeepLabV3 + network model

By improving the DeepLabV3+ network model, DLNet is designed for multi-scale feature extraction and lightweight processing, which solves the problem of balance between detection accuracy and resource consumption in traffic scenarios by feasible domain segmentation models, and achieves efficient road segmentation.

CN120279519APending Publication Date: 2025-07-08XI'AN POLYTECHNIC UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510410542.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing feasible domain segmentation model is difficult to achieve high detection accuracy due to occlusion problems and single-scale segmentation models in traffic scenarios. At the same time, the DeepLabV3+ model has problems such as complex structure and difficult to balance prediction efficiency and resource consumption.

Method used

Using the improved DeepLabV3+ network model, the feasible domain segmentation model DLNet is designed, including the Backbone module, MPAM module, feature fusion module and overlay sampling module. Through multi-scale feature extraction and lightweight design, it reduces computational redundancy and information loss.

Benefits of technology

It realizes high-precision feasible domain segmentation, reduces model parameters and calculation amount, has the characteristics of lightweight, strong real-time and wide applicability, and solves the problem of image upsampling edge features loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279519A_ABST
    Figure CN120279519A_ABST
Patent Text Reader

Abstract

The invention discloses a road lightweight feasible region segmentation method based on an improved DeepLabV3 + network model, and the method specifically comprises the steps: preprocessing data, and dividing the data into a training set, a verification set and a test set; designing a feasible region segmentation model DLNet, and processing the training set image by using the feasible region segmentation model DLNet; inputting the training set image, the verification set image and the corresponding picture label into a constructed feasible region segmentation model DLNet for training until convergence; and inputting a test set image into the trained model to generate a prediction image, and outputting the corresponding average pixel accuracy (MPA), average intersection-to-union ratio (MIOU), model parameter quantity and model calculation quantity to a terminal or storing the MPA, the MIOU, the model parameter quantity and the model calculation quantity to a file. According to the method, the prediction precision of the feasible region segmentation task is improved, the model parameters and the calculation amount are reduced, and the problem that the sampling edge features on the image are lost is solved to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for segmenting the lightweight feasible region of a road based on an improved DeepLabV3+ network model. Background Art

[0002] In recent years, with the development of intelligent driving technology, the feasible region segmentation algorithm has received extensive attention. The feasible region segmentation algorithm can usually be divided into two categories: traditional image processing methods and deep learning-based segmentation methods. Traditional image processing methods such as: threshold-based segmentation methods, edge detection-based segmentation methods, and clustering-based segmentation methods. Most of these methods rely on manual operations to determine the threshold and segmentation range, making it difficult to achieve fine-grained operations and unable to adapt to most scenarios. With the development of deep learning, researchers have achieved the feasible region segmentation task by using neural network models to train image feature classifiers. Deep learning-based segmentation methods are mainly divided into three categories, namely fully convolutional neural networks (based on FCN), encoder-decoder structures (based on U-Net), and multi-scale feature extraction and fusion networks. In addition, there are many real-time feasible region segmentation models such as ENet, BiSeNet, etc. These models have achieved excellent performance through training on datasets, greatly promoting the development of the feasible region segmentation task.

[0003] However, in traffic scenarios, due to occlusion problems, there are situations where the same category has different sizes in the image. Therefore, a single-scale segmentation model is difficult to achieve high detection accuracy, and the DeepLabV3+ model of the multi-scale feature extraction and fusion network has problems of complex structure and difficulty in balancing prediction efficiency and resource consumption. Based on this, it is of great significance to study a feasible region segmentation model that can balance prediction efficiency and resource consumption difficulties in the field of intelligent driving. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for segmenting the lightweight feasible region of a road based on an improved DeepLabV3+ network model, realizing lightweight deployment and high-precision segmentation.

[0005] The technical solution adopted by the present invention is a method for segmenting the lightweight feasible region of a road based on an improved DeepLabV3+ network model, which is specifically implemented according to the following steps:

[0006] Step 1, perform data preprocessing, and divide the preprocessed image into a training set, a validation set, and a test set;

[0007] Step 2, design a feasible region segmentation model DLNet, that is, an improved DeepLabV3+ network model, including a Backbone module, an MPAM module, a feature fusion module, and an upsampling module; and use the feasible region segmentation model DLNet to process the training set images;

[0008] Step 3: Input the training set images, validation set images, and their corresponding image labels into the established feasible region segmentation model DLNet for training until convergence;

[0009] Step 4: Input the test set images into the trained feasible region segmentation model DLNet to generate predicted images, and output the corresponding mean pixel accuracy (MPA), mean intersection over union (MIOU), number of model parameters, and model computational cost to the terminal or store them in a file.

[0010] The features of the present invention also lie in that,

[0011] In Step 1, specifically:

[0012] Step 1.1: Use the Segment Anything model to perform segmentation annotation on the images in the dataset, label the feasible region as 1, and label the background region as 0;

[0013] Step 1.2: Save the segmentation results in the form of a two-dimensional array to represent the semantic category to which each pixel belongs, and ensure that the generated labels correspond one-to-one with the images;

[0014] Step 1.3: Adjust the image resolution to 512×1024 pixels, convert the image format to the PyTorchTensor format, and normalize the pixel values to the range of [0,1];

[0015] Step 1.4: Perform standardization processing on the images so that their mean and standard deviation are [0.485, 0.456, 0.406] and [0.229, 0.224, 0.225] respectively to obtain the preprocessed images;

[0016] Step 1.5: Divide the preprocessed images into a dataset according to the ratio of 7:2:1, with 70% as the training set for model training, 20% as the validation set for assisting model training, and 10% as the test set for testing the model results.

[0017] In Step 2, design the feasible region segmentation model DLNet, including the Backbone module, MPAM module, feature fusion module, and upsampling module; specifically:

[0018] Step a: Use the Backbone module to process the images and output the shallow features L;

[0019] Step b: Use the MPAM module to process the deep features extracted by the Backbone module and output the attention features A;

[0020] Step c, use the feature fusion module to fuse the shallow feature L output in step a with the attention feature A in step b;

[0021] Step d, use a transposed convolution and bilinear method superposed upsampling module on the result of step c to restore the original image size.

[0022] In step a, specifically as follows:

[0023] Step a1, pass the training set images through a standard 3×3 convolution with an input channel number of 3, an output channel number of 16, and a stride of 1;

[0024] Step a2, use an atrous convolution module to process the feature map output in step a1;

[0025] The specific process of the atrous convolution module for processing the feature map is as follows: the input feature map size is B×16×512×1024, where B is the batch size, 16 is the input channel number, and 512×1024 is the size of the feature map. Process the feature using an atrous convolution with a convolution kernel size of 3, a dilation rate of 2, a padding of 2, and a stride of 2; the output feature map size is B×48×256×512; then perform a BatchNorm operation on the output result to normalize each batch, and finally use the non-linear ReLU activation function to process the feature;

[0026] Step a3, use two atrous convolution modules to process the feature map in step a2;

[0027] The atrous convolution has a convolution kernel size of 3, a dilation rate of 2, a padding of 2, a stride of 1, an input channel number of 48, and an output channel number of 48; then perform a BatchNorm operation on the output result to normalize each batch so that the features conform to a normal distribution;

[0028] Step a4, use an atrous convolution module to process the feature map in step a3, denoted as the shallow feature L;

[0029] The atrous convolution has a convolution kernel size of 3, a dilation rate of 2, a padding of 2, a stride of 2, an input channel number of 48, and an output channel number of 64; adjust the feature map size again so that the output size is B×64×128×256; then perform a BatchNorm operation on the output result to normalize each batch so that the features conform to a normal distribution;

[0030] Step a5, use two depthwise separable convolution modules to process the feature map in step a4;

[0031] The specific process of the depthwise separable convolution module is as follows: First, perform layer-by-layer convolution on the feature map in step a4. Use a convolution kernel with a size of 3, a padding of 1, and a stride of 1 to adjust the number of channels of the feature map to 96. Subsequently, use BatchNorm to normalize the feature map and use the ReLU function to activate the features. Then, perform pointwise convolution on the feature map, which is intended to adjust the number of channels to 64, and retain key features while ignoring minor features through channel scaling. The pointwise convolution uses a 1×1 convolution kernel, a stride of 1, and a padding of 0. The size of the output feature map after BatchNorm is B×64×128×256;

[0032] Step a6, use a depthwise separable convolution to process the feature map in step a5;

[0033] Specifically, repeat step a5. The parameter settings for layer-by-layer convolution are a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel number of 96; the parameter settings for pointwise convolution are a convolution kernel size of 1, a stride of 1, and an output channel number of 64; the size of the output feature map is B×64×64×128;

[0034] Step a7, repeat steps a5 - a6, and finally the size of the output feature map is B×64×32×64, denoted as the deep feature D.

[0035] In step b, specifically:

[0036] Step b1, use max pooling operation and average pooling operation respectively to downsample the deep feature D in step a7 with a magnification factor of r, and the results are denoted as P max and P avg ;

[0037] Step b2, use 1×1 convolution to optimize the two features obtained in step b1 respectively;

[0038] Step b3, multiply the two feature matrices in step b2, that is, P max ×P avg ;

[0039] Step b4, use the Softmax function to transform the result of step b3, and transform the enhanced features into attention scores S;

[0040] Step b5, repeat steps b1 - b4 with 4 different magnification factors to obtain multi-scale attention scores S r , r ∈ (1, 2, 3, 6);

[0041] Step b6, bilinearly upsample the attention scores S in step b5, so that r

[0042] Step b7, multiply S r by the deep feature D respectively to activate the initial deep feature with the attention score, and the result is denoted as

[0043] Step b8, stack multiple feature maps A r in the channel dimension, and the result is denoted as

[0044] In step c, specifically:

[0045] Step c1, further process the shallow feature L, use 2 times of 3×3 convolution with a stride of 2 and a padding of 1 to reduce the feature size and adjust the number of channels to

[0046] Step c2, stack the result L` in step c1 and the result A in step b in the channel dimension, and the output feature map size is 320×32×64;

[0047] Step c3, use 1×1 convolution to fuse the stacked features and adjust the channel dimension to

[0048] In step d, specifically:

[0049] Step d1, perform 2-fold upsampling using transposed convolution to obtain a feature map

[0050] The transposed convolution is specifically: use a transposed convolution operation with a convolution kernel size of 4, a stride of 2, and a padding of 1, and the output feature map size is 64×64×128;

[0051] Step d2, use 1×1 convolution to smooth the feature edges and adjust the channel dimension, and the output feature map size is 16×64×128;

[0052] Step d3, perform 2-fold upsampling using the bilinear method, and the output feature map size is 16×128×256;

[0053] Step d4, repeat steps d1-d3, and the output feature map size is 2×512×1024.

[0054] In step 3, set the network hyperparameters of the feasible region segmentation model DLNet: the optimizer is Adam, the initial learning rate 1r0 is 5e-3, the momentum is 0.9, the weight decay is 0.0005, and the decay function is cosine; the batch size is 16;

[0055] Input the training set images and their corresponding image labels into the feasible region segmentation model DLNet, calculate the loss function and backpropagate to obtain the gradients, and the optimizer updates the parameters until the loss functions of the training set and the validation set no longer decrease and the output evaluation metrics no longer improve. At this time, the model converges, and the parameters of the model are the finally trained model parameters;

[0056] Use the mean pixel accuracy MPA, mean intersection over union MIoU, number of model parameters, and model computational cost as evaluation metrics.

[0057] The beneficial effects of the present invention are:

[0058] The road lightweight feasible region segmentation method based on the improved DeepLabV3+ network model of the present invention utilizes the feasible region segmentation model DLNet, which combines the Backbone module, MPAM module, feature fusion module, and upsampling module to achieve efficient extraction of multi-scale features, reduce computational redundancy and information loss; through data preprocessing, model design, model training, and finally using the model to perform feasible region segmentation on the images of the test set, the prediction accuracy of the feasible region segmentation task is improved, and the model parameters and computational cost are reduced, achieving an effective balance between prediction accuracy and resource consumption, and solving the problem of edge feature loss in image upsampling to a certain extent. The method of the present invention can quickly and accurately segment the feasible region task, consumes less resources, and has the characteristics of lightweight, strong real-time performance, and wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is the input image in the road lightweight feasible region segmentation method based on the improved DeepLabV3+ network model of the present invention;

[0060] Figure 2 It is the structural diagram of the DLNet model in the road lightweight feasible region segmentation method based on the improved DeepLabV3+ network model of the present invention;

[0061] Figure 3 It is the MobileNetV2 network pruning and optimization diagram in the DLNet model in the road lightweight feasible region segmentation method based on the improved DeepLabV3+ network model of the present invention;

[0062] Figure 4 It is the MPAM module diagram in the DLNet model in the road lightweight feasible region segmentation method based on the improved DeepLabV3+ network model of the present invention;

[0063] Figure 5 It is the upsampling module diagram in the DLNet model in the road lightweight feasible region segmentation method based on the improved DeepLabV3+ network model of the present invention;

[0064] Figure 6 It is the comparison diagram (one) of the output images between the DLNet model of the present invention and the original model;

[0065] Figure 7 It is the comparison diagram (two) of the output images between the DLNet model of the present invention and the original model;

[0066] Figure 8 It is the comparison diagram (three) of the output images between the DLNet model of the present invention and the original model. Detailed implementation manners

[0067] The present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners.

[0068] Embodiment 1

[0069] A method for segmenting the lightweight feasible region of roads based on an improved DeepLabV3+ network model, specifically:

[0070] Step 1, perform data preprocessing, and divide the preprocessed images into a training set, a validation set, and a test set;

[0071] Step 2, design a feasible region segmentation model DLNet, that is, an improved DeepLabV3+ network model, including a Backbone module, an MPAM module, a feature fusion module, and an upsampling module; and use the feasible region segmentation model DLNet to process the training set images;

[0072] Step 3, input the training set images, the validation set images, and the corresponding image labels into the established feasible region segmentation model DLNet for training until convergence;

[0073] Step 4, input the test set images into the trained feasible region segmentation model DLNet to generate prediction images, and output the corresponding mean pixel accuracy MPA, mean intersection over union MIOU, the number of model parameters, the model calculation amount, to the terminal or store them in a file.

[0074] Embodiment 2

[0075] The method for segmenting the lightweight feasible region of roads based on the improved DeepLabV3+ network model of the present invention is specifically implemented according to the following steps:

[0076] Step 1, perform data preprocessing, specifically:

[0077] Step 1.1, use the Segment Anything model to segment and label the images in the dataset, label the feasible region as 1, and label the background region as 0;

[0078] The image is an actual road scene, such as an image dataset generated by a dash cam, or the Cityscapes dataset;

[0079] Step 1.2, save the segmentation result in the form of a two-dimensional array to represent the semantic category to which each pixel belongs, and ensure that the generated label corresponds one-to-one with the image;

[0080] Step 1.3, adjust the image resolution to 512×1024 pixels, convert the image format to the PyTorch Tensor format, and normalize the pixel values to the range of [0,1];

[0081] Step 1.4, perform standardization processing on the image so that its mean and standard deviation are [0.485, 0.456, 0.406] and [0.229, 0.224, 0.225] respectively, to obtain the preprocessed image;

[0082] Step 1.5, divide the preprocessed image into a dataset according to the ratio of 7:2:1. 70% is used as the training set for model training, 20% is used as the validation set to assist model training, and 10% is used as the test set to test the model results;

[0083] Step 2, design a feasible region segmentation model DLNet, that is, improve the DeepLabV3+ network model, as Figure 2 shown, including the Backbone module, the MPAM module, the feature fusion module and the superimposed upsampling module; specifically:

[0084] Step a, process the image using the Backbone module;

[0085] Step a1, pass the training set image in Step 1 through a standard 3×3 convolution, with the input channel number being 3, the output channel number being 16, the stride being 1, while increasing the number of channels and keeping the feature map unchanged;

[0086] Step a2, use an atrous convolution module to process the feature map output in Step a1;

[0087] The specific process of the atrous convolution module processing the feature map is as follows: the input feature map size is B×16×512×1024, where B is the batch size, 16 is the input channel number, and 512×1024 is the size of the feature map. Process the feature using an atrous convolution with a convolution kernel size of 3, a dilation rate of 2, a padding of 2, and a stride of 2; the output feature map size is B×48×256×512; then perform BatchNorm operation on the output result to normalize each Batch so that the features conform to the normal distribution and accelerate the model training speed; finally, use the non-linear ReLU activation function to process the features to prevent feature overfitting;

[0088] Step a3: Process the feature map from step a2 using two dilated convolution modules.

[0089] The dilated convolution has a kernel size of 3, a dilation rate of 2, a padding of 2, a stride of 1, an input channel number of 48, and an output channel number of 48. By setting the parameters of the dilated convolution, the extracted spatial features are preserved in the channels while keeping the feature map size unchanged. Then, perform a BatchNorm operation on the output result to normalize each Batch so that the features conform to a normal distribution.

[0090] Step a4: Process the feature map from step a3 using one dilated convolution module, denoted as the shallow feature L.

[0091] The dilated convolution has a kernel size of 3, a dilation rate of 2, a padding of 2, a stride of 2, an input channel number of 48, and an output channel number of 64. Adjust the feature map size again so that the output size is B×64×128×256. Then, perform a BatchNorm operation on the output result to normalize each Batch so that the features conform to a normal distribution.

[0092] Step a5: Process the feature map from step a4 using two depthwise separable convolution modules.

[0093] The specific process of the depthwise separable convolution module is as follows: First, perform layer-by-layer convolution on the feature map from step a4. Use a convolution with a kernel size of 3, a padding of 1, and a stride of 1 to adjust the number of channels of the feature map to 96. Subsequently, use BatchNorm to normalize the feature map and use the ReLU function to activate the features. Then, perform pointwise convolution on the feature map, which is used to adjust the number of channels to 64. By scaling the channels, key features are retained while minor features are ignored. The pointwise convolution uses a 1×1 convolution with a stride of 1 and a padding of 0. The size of the output feature map after BatchNorm is B×64×128×256.

[0094] Step a6: Process the feature map from step a5 using one depthwise separable convolution.

[0095] Specifically: Repeat step a5. The parameter settings for layer-by-layer convolution are a kernel size of 3, a stride of 2, a padding of 1, and an output channel number of 96. The parameter settings for pointwise convolution are a kernel size of 1, a stride of 1, and an output channel number of 64. The size of the output feature map is B×64×64×128.

[0096] Step a7: Repeat steps a5 - a6. The final output feature map size is B×64×32×64, denoted as the deep feature D.

[0097] Step b: Process the deep features extracted by the Backbone using the MPAM module, as Figure 4 shown. Extract multi-scale features and stack them to meet the requirements of object detection with different sizes. Specifically:

[0098] Step b1: Perform downsampling on the deep feature D in step a7 using max-pooling operation and average-pooling operation respectively, with a magnification factor of r. The results are denoted as P max and P avg respectively. As shown in the following formula:

[0099]

[0100] where maxpool and avgpool are the max-pooling operation and average-pooling operation respectively;

[0101] Step b2: Optimize the two features obtained in step b1 using 1×1 convolution respectively. As shown in the following formula:

[0102] P i = conv 1×1 (p i )

[0103] where conv 1×1 is a two-dimensional convolution with a kernel size of 1, and p i is the result obtained in step b1;

[0104] Step b3: Multiply the two feature matrices in step b2, that is, P max ×P avg . As shown in the following formula:

[0105] M = P max ×P avg

[0106] Step b4: Use the Softmax function to transform the result of step b3 and convert the enhanced features into attention scores S. As shown in the following formula:

[0107]

[0108] where i and j are the pixel coordinates of the feature image respectively;

[0109] Step b5: Repeat steps b1 - b4 with 4 different magnification factors to obtain multi-scale attention scores S r , r ∈ (1, 2, 3, 6);

[0110] Step b6: Bilinearly up-sample the attention scores S r in step b5 so that

[0111] Step b7: Multiply S r with the deep feature D respectively to activate the initial deep feature with the attention score, and the result is denoted as

[0112] Step b8: Stack multiple feature maps A r in the channel dimension, and the result is denoted as

[0113] Step c: Fuse the shallow feature L output in Step a with the attention feature A in Step b; specifically:

[0114] Step c1: Further process the shallow feature L, use 3×3 convolution twice with a stride of 2 and a padding of 1 to reduce the feature size and adjust the number of channels to as shown in the following formula:

[0115] L` = conv 3×3 (L)

[0116] Step c2: Stack the result L` in Step c1 and the result A in Step b in the channel dimension, and the output feature map size is 320×32×64; as shown in the following formula:

[0117] C = concat(L`, A)

[0118] Step c3: Use 1×1 convolution to fuse the stacked features and adjust the channel dimension to as shown in the following formula:

[0119]

[0120] Step d: As Figure 4 shown, use a transposed convolution and bilinear method superposition upsampling module for the result C` in Step c, as Figure 5 shown, to restore the original image size,

[0121] Step d1: Use transposed convolution for 2x upsampling to obtain a feature map

[0122] The transposed convolution is specifically: use a transposed convolution operation with a convolution kernel size of 4, a stride of 2, and a padding of 1, and the output feature map size is 64×64×128;

[0123] Step d2: Use 1×1 convolution to smooth the feature edges and adjust the channel dimension, and the output feature map size is 16×64×128;

[0124] Step d3, perform 2x upsampling using the bilinear method, and the output feature map size is 16×128×256;

[0125] Step d4, repeat steps d1 - d3, and the output feature map size is 2×512×1024.

[0126] Repeat step d1, and the output feature map size is 8×256×512; repeat step d2, and the output feature map size is 2×256×512; repeat step d3, and the output feature map size is 2×512×1024;

[0127] Step 3, input the training set images and the corresponding image labels into the established feasible region segmentation model DLNet for training until convergence;

[0128] Set the network hyperparameters of the feasible region segmentation model DLNet: the optimizer optimizer is Adam, the initial learning rate 1r0 is 5e - 3, the momentum momentum is 0.9, the weight decay weight_decay is 0.0005, and the decay function is cosine. The batch size batch size is 16;

[0129] Input the training set images and the corresponding image labels into the feasible region segmentation model DLNet, calculate the loss function and backpropagate to calculate the gradient, and the optimizer updates the parameters until the training set loss function and the validation set function no longer decrease, and at the same time the output evaluation metrics no longer improve. At this time, the model converges, and the parameters of the model are the finally trained model parameters;

[0130] The loss function specifically uses the BCELoss loss function; as shown in the following formula:

[0131]

[0132] where N is the number of pixels in the image, y i is the prediction result of pixel i, is the true result of pixel i.

[0133] Use the mean pixel accuracy (MPA), mean intersection over union (MIoU), number of model parameters, and model computational complexity as evaluation metrics;

[0134] The pixel accuracy refers to the proportion of correctly classified pixels in the total pixels; the intersection over union is the intersection of the true value and the predicted value of the pixel divided by the union of the true value and the predicted value of the pixel;

[0135] When the test set shows class imbalance (different classes: a large difference in the number of samples), the pixel accuracy cannot objectively reflect the model performance. Therefore, two evaluation metrics, namely the Mean Pixel Accuracy (MPA) and the Mean Intersection over Union (MIoU), are defined as shown in the following formula:

[0136]

[0137] Among them, assuming there are k + 1 classes (including an empty class or background), P ii represents the pixels where the prediction matches the actual value, and P ij represents the pixels that actually belong to class i but are predicted as class j, and P ji represents the pixels that actually belong to class j but are predicted as class i. i represents the true value, j represents the predicted value, and p ij represents the number of pixels that predict i as j;

[0138] The number of model parameters (Parameters) is an evaluation metric that reflects the size of the model. During the training process, the model usually generates weight and bias parameters. The larger the number of model parameters, the larger the weights obtained during training, and more storage space is required;

[0139] GFLOPS is a unit of measurement for floating-point operations, indicating the number of billions of floating-point operations performed per second. During the training process, floating-point operations are usually used to perform the forward and backward propagation of the neural network. The higher the value of GFLOPS, the more powerful the computing performance, the faster the model training and inference speed, which also means that more computing power is required for training;

[0140] Step 4, input the test set images into the trained feasible region segmentation model DLNet to generate prediction images, as Figure 6 shown; and output the corresponding Mean Pixel Accuracy (MPA), Mean Intersection over Union (MIOU), number of model parameters, and model computational amount to the terminal or store them in a file.

[0141] Example 3

[0142] Input the divided test set into the final feasible region segmentation model DLNet, and compare the Mean Pixel Accuracy (MPA), Mean Intersection over Union (MIoU), number of parameters (Parameters), and computational amount (GFLOPS) of the final model for the feasible region segmentation task with the evaluation metrics of the original model DeepLabV3+ to prove the effectiveness of DLNet.

[0143] Example 4

[0144] As shown in Table 1, it is a data comparison table of the feasible region segmentation model DLNet and the original DeepLabV3+ on the dataset. The improved model has obvious improvements in various evaluation metrics. In terms of prediction performance, MPA and MIoU have increased by 0.67% and 0.41% respectively, while the number of parameters is only 39.2% of the original model, and the computational cost is only 12.3% of the original model. It can be seen that the feasible region segmentation model DLNet has achieved better results.

[0145] Table 1 Data comparison table of the feasible region segmentation model DLNet and the original DeepLabV3+ on the dataset

[0146] Model MPA (%) MIoU (%) Parameters (M) GFLOPS (G) DeepLabV3+ 96.50 93.90 5.813 105.717 The DLNet of the present invention 97.17 94.31 2.284 13.069

[0147] Example 5

[0148] As Figures 6-8 shown, it is a visual display of the segmentation results of the input image by the feasible region segmentation model DLNet and the original DeepLabV3+. The improved DeepLabV3+ network model has higher segmentation accuracy in details than the original model. It can be seen that the feasible region segmentation model has achieved better results.

[0149] Example 6

[0150] The road lightweight feasible region segmentation method of the present invention is based on the improved DeepLabV3+ network model, using the feasible region segmentation model DLNet, which combines the Backbone module, the MPAM module, the feature fusion module and the upsampling module to achieve efficient extraction of multi-scale features, reduce computational redundancy and information loss; through data preprocessing, model design, model training, and finally using the model to perform feasible region segmentation on the images of the test set, the prediction accuracy of the feasible region segmentation task is improved and the model parameters and computational cost are reduced, achieving an effective balance between prediction accuracy and resource consumption, and solving the problem of edge feature loss in image upsampling to a certain extent. The method of the present invention can quickly and accurately segment the feasible region task, and consumes less resources, and has the characteristics of lightweight, strong real-time performance, wide applicability, etc.

Claims

1. A method for segmenting the lightweight feasible region of roads based on an improved DeepLabV3+ network model, characterized in that, The implementation is specifically carried out according to the following steps: Step 1, perform data preprocessing, and divide the preprocessed images into a training set, a validation set, and a test set; Step 2, design a feasible region segmentation model DLNet, that is, improve the DeepLabV3+ network model, including a Backbone module, an MPAM module, a feature fusion module, and an upsampling module; And use the feasible region segmentation model DLNet to process the training set images; Step 3, input the training set images, validation set images, and the corresponding picture labels into the established feasible region segmentation model DLNet for training until convergence; Step 4, input the test set images into the trained feasible region segmentation model DLNet to generate predicted images, and output the corresponding mean pixel accuracy MPA, mean intersection over union MIOU, number of model parameters, model computational cost to the terminal or store them in a file.

2. The method for segmenting the lightweight feasible region of a road based on the improved DeepLabV3+ network model according to claim 1, wherein In the said Step 1, specifically: Step 1.1, use the Segment Anything model to segment and label the images in the dataset, label the feasible region as 1, and label the background area as 0; Step 1.2, save the segmentation results in the form of a two-dimensional array to represent the semantic category to which each pixel belongs, and ensure that the generated labels correspond one-to-one with the images; Step 1.3, adjust the image resolution to 512×1024 pixels, convert the image format to the PyTorch Tensor format, and normalize the pixel values to the range of [0,1]; Step 1.4, perform standardization processing on the images so that their mean and standard deviation are [0.485,0.456,0.406] and [0.229,0.224,0.225] respectively to obtain the preprocessed images; Step 1.5, divide the preprocessed images into a dataset according to the ratio of 7:2:1, 70% as the training set for model training, 20% as the validation set for assisting model training, and 10% as the test set for testing the model results.

3. The method for segmenting the lightweight feasible region of a road based on the improved DeepLabV3+ network model according to claim 1, characterized in that In the said Step 2, design a feasible region segmentation model DLNet, including a Backbone module, an MPAM module, a feature fusion module, and an upsampling module; Specifically: Step a, use the Backbone module to process the images and output the shallow features L; Step b, use the MPAM module to process the deep features extracted by the Backbone module and output the attention features A; Step c, use the feature fusion module to fuse the shallow features L output in Step a with the attention features A in Step b; Step d, use the upsampling module of transposed convolution and bilinear method for the result of Step c to restore the original image size.

4. The method for segmenting the lightweight feasible region of a road based on the improved DeepLabV3+ network model according to claim 3, characterized in that, In the said Step a, specifically as follows: Step a1, pass the training set images through a standard 3×3 convolution, with the input channel number being 3, the output channel number being 16, and the stride being 1; Step a2, use an atrous convolution module to process the feature map output in Step a1; The specific process of the dilated convolution module for processing the feature map is as follows: The size of the input feature map is B×16×512×1024, where B is the batch size, 16 is the number of input channels, and 512×1024 is the size of the feature map. The features are processed using a dilated convolution with a convolution kernel size of 3, a dilation rate of 2, a padding of 2, and a stride of 2. The size of the output feature map is B×48×256×512; then, a BatchNorm operation is performed on the output result to normalize each batch, and finally, a non-linear ReLU activation function is used to process the features. Step a3: Use two dilated convolution modules to process the feature map in step a2. The dilated convolution has a convolution kernel size of 3, a dilation rate of 2, a padding of 2, a stride of 1, an input channel number of 48, and an output channel number of 48; then, a BatchNorm operation is performed on the output result to normalize each batch so that the features conform to a normal distribution. Step a4: Use a dilated convolution module to process the feature map in step a3, denoted as the shallow feature L. The dilated convolution has a convolution kernel size of 3, a dilation rate of 2, a padding of 2, a stride of 2, an input channel number of 48, and an output channel number of 64; the size of the feature map is adjusted again so that the output size is B×64×128×256; then, a BatchNorm operation is performed on the output result to normalize each batch so that the features conform to a normal distribution. Step a5: Use two depthwise separable convolution modules to process the feature map in step a4. The specific process of the depthwise separable convolution module is as follows: First, perform layer-by-layer convolution on the feature map in step a4, and use a convolution with a convolution kernel size of 3, a padding of 1, and a stride of 1 to adjust the number of channels of the feature map to 96; then, use BatchNorm to normalize the feature map and use the ReLU function to activate the features; then, perform pointwise convolution on the feature map, the significance of which is to adjust the number of channels to 64, and retain key features and ignore secondary features through channel scaling; the pointwise convolution uses a 1×1 convolution, a stride of 1, and a padding of 0; the size of the output feature map after BatchNorm is B×64×128×256. Step a6: Use a depthwise separable convolution to process the feature map in step a5. Specifically; Repeat step a5, and the parameter settings for layer-by-layer convolution are a convolution kernel size of 3, a stride of 2, a padding of 1, and an output channel number of 96; the parameter settings for pointwise convolution are a convolution kernel size of 1, a stride of 1, and an output channel number of 64; the size of the output feature map is B×64×64×128. Step a7: Repeat steps a5 - a6, and the final output feature map size is B×64×32×64, denoted as the deep feature D.

5. The method for segmenting the lightweight feasible region of a road based on the improved DeepLabV3+ network model according to claim 4, wherein In step b, specifically: Step b1, respectively perform downsampling on the deep feature D in step a7 using max pooling operation and average pooling operation, with a magnification factor of r, and the results are denoted as P max and P avg ; Step b2: Use 1×1 convolutions to optimize the two features obtained in step b1 respectively. Step b3, multiply the two feature matrices in step b2, i.e., P max ×P avg ; Step b4: Use the Softmax function to transform the result of step b3, and transform the enhanced features into attention scores S. Step b5, repeat steps b1 - b4 using 4 different magnifications to obtain the multi-scale attention score S r , where r ∈ (1, 2, 3, 6); Step b6, the attention score S in step b5 r is bilinearly upsampled so that Step b7, multiply S r and the deep feature D respectively to activate the initial deep feature with the attention score, and the result is denoted as Step b8, stack multiple feature maps A r in the channel dimension, and denote the result as 6. The method for segmenting the lightweight feasible region of a road based on the improved DeepLabV3+ network model according to claim 5, wherein In step c, specifically: Step c1, further process the shallow features L, use 2 3×3 convolutions with a stride of 2 and a padding of 1 to reduce the feature size and adjust the number of channels to Step c2: Stack the result L` in step c1 and the result A in step b along the channel dimension, and output a feature map with a size of 320×32×64; Step c3, use 1×1 convolution to fuse and stack features, and adjust the channel dimension to 7. The method for segmenting the lightweight feasible region of a road based on the improved DeepLabV3+ network model according to claim 6, characterized in that, In the said step d, specifically: Step d1, perform 2x upsampling using transposed convolution to obtain a feature map The transposed convolution is specifically: perform a transposed convolution operation with a convolution kernel size of 4, a stride of 2, and a padding of 1, and output a feature map with a size of 64×64×128; Step d2: Use a 1×1 convolution to smooth the feature edges and adjust the channel dimension, and output a feature map with a size of 16×64×128; Step d3: Perform bilinear upsampling by a factor of 2, and output a feature map with a size of 16×128×256; Step d4: Repeat steps d1 - d3, and output a feature map with a size of 2×512×1024.

8. The method for segmenting the lightweight feasible region of a road based on the improved DeepLabV3+ network model according to claim 7, wherein In the said step 3, set the network hyperparameters of the feasible region segmentation model DLNet: the optimizer is Adam, the initial learning rate 1r0 is 5e - 3, the momentum is 0.9, the weight decay is 0.0005, and the decay function is cosine; the batch size is 16; Input the training set images and the corresponding image labels into the feasible region segmentation model DLNet, calculate the loss function and backpropagate to calculate the gradient, and the optimizer updates the parameters until the training set loss function and the validation set function no longer decrease, and at the same time the output evaluation metrics no longer improve. At this time, the model converges, and the parameters of the model are the finally trained model parameters; Use the mean pixel accuracy MPA, the mean intersection over union MIoU, the number of model parameters, and the model computational cost as evaluation metrics.