Improved DeepLabv3 +-based unstructured road travelable area identification method
By improving the DeepLabv3+ model and using the MobileNetv4 network, ASPP module and attention mechanism, the problems of high data labeling cost and high model complexity in unstructured road recognition are solved, and efficient and accurate drivable area recognition is achieved.
Patent Information
- Application Number
- CN202510780350.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-10-03
AI Technical Summary
The existing DeepLabv3+ model has problems in identifying drivable areas on unstructured roads, such as the high cost of supervised learning relying on large-scale labeled data, a large number of model parameters, and a prominent contradiction between recognition efficiency and accuracy. It is difficult to meet the dual requirements of autonomous driving systems for low latency and high precision.
The MobileNetv4 network is used to replace the traditional DeepLabv3+ Xception network, combined with the ASPP module and decoder, and the attention mechanism and hybrid loss function are introduced to improve feature expression ability and recognition efficiency through feature optimization and multi-scale context information extraction.
While reducing the amount of calculation and latency, it significantly improves the recognition effect of unstructured roads, enhances the model's ability to distinguish complex backgrounds and edge areas, improves recognition accuracy and training efficiency, and achieves lightweight and real-time performance.
Smart Images

Figure CN120747894A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of road recognition, and in particular relates to a method for identifying drivable areas on unstructured roads based on an improved DeepLabv3+. Background Art
[0002] One of the core challenges of autonomous driving technology lies in accurately perceiving complex road environments, and identifying drivable areas on unstructured roads is particularly critical. Early research on drivable area identification relied primarily on traditional image processing techniques, such as color space conversion, edge detection, or geometric model matching. In structured road scenarios, these methods achieve high accuracy by extracting lane lines or boundary features. However, in unstructured roads, they are unable to effectively handle dynamic environmental interference and sudden changes in road morphology, often leading to missegmentation or missed detections.
[0003] In recent years, the rise of deep learning has injected new impetus into this field. Semantic segmentation models based on convolutional neural networks (CNNs) achieve automatic extraction of multi-level features through end-to-end learning, significantly improving the adaptability of unstructured road recognition, especially in dealing with fuzzy boundaries and noise interference. Among them, DeepLabv3+ captures multi-scale contextual information through the encoder's ASPP module, and combined with the decoder's progressive upsampling strategy, it significantly improves the boundary accuracy and semantic consistency of the segmentation results. For example, in street scene segmentation for autonomous driving, DeepLabv3+ can accurately distinguish between targets such as road surfaces, traffic signs, and vegetation. Even under the challenges of occlusion or lighting changes, it can still maintain high-quality segmentation of details such as wheel outlines and leaf edges, providing an important guarantee for the reliability of environmental perception systems.
[0004] However, the existing DeepLabv3+ still faces significant limitations: (1) Supervised learning is highly dependent on large-scale labeled data, but the types of unstructured road scenes are complex and the data labeling cost is high. Existing public datasets mostly focus on structured roads and are difficult to cover diverse scenes such as rural dirt roads and mountain trails; (2) The model itself has high parameters, and the contradiction between recognition efficiency and accuracy of existing methods is prominent. Although lightweight network design can improve real-time performance, it often comes at the expense of feature resolution capabilities in complex environments, making it difficult to meet the dual requirements of autonomous driving systems for low latency and high accuracy. Summary of the Invention
[0005] In view of the above-mentioned defects of the prior art, the present invention proposes a method for identifying drivable areas on unstructured roads based on an improved DeepLabv3+. The technical solution designed by the present invention includes the following steps: S1: Build an improved DeepLabv3+ model, including the MobileNetv4 network, ASPP module and decoder; S2: Obtain an image of an unstructured road, perform feature extraction based on the MobileNetv4 network, and obtain a high-dimensional feature map; S3: Performing feature optimization processing based on the attention mechanism on the high-dimensional feature map to obtain a high-dimensional optimized feature map; S4: performing multi-scale context information extraction based on the ASPP module on the high-dimensional optimized feature map to obtain a final high-dimensional optimized feature map; S5: Train the DeepLabv3+ model based on the hybrid loss function; S6: Outputting a recognition image of the unstructured road based on the decoder.
[0006] Preferably, the feature extraction process based on the MobileNetv4 network in S2 includes: The image of the unstructured road is sequentially subjected to initial convolution and downsampling operations, multi-stage UIB module feature mixing operations, and feature fusion and output operations.
[0007] Preferably, the initial convolution and downsampling operations include: The unstructured road image is passed through a standard convolutional layer to perform preliminary feature extraction and spatial compression, reducing the resolution to 1 / 4 of the original image while extracting low-level features of edges and textures. The formula is as follows:
[0008] Where, is the image after convolution, is a two-dimensional convolutional layer, is the step length; The depth convolution and point-by-point convolution are combined into a single-step operation. The depth convolution is responsible for local spatial feature extraction, and the point-by-point convolution is used to expand the channel dimension. The formula is as follows:
[0009]
[0010] Where, After FusedIB processing, After pointwise convolution processing, After depthwise convolution processing.
[0011] Preferably, the multi-stage UIB module feature mixing operation includes: Divide the low-resolution stage and high-resolution stage and process them separately; At the low-resolution stage, two deep convolutions are connected in series to expand the receptive field. Executed before channel expansion, the second Executed after expansion, the formula is as follows:
[0012] At the high-resolution stage, pure point-by-point convolution stacking is used to mine cross-channel semantic associations through nonlinear activation and channel projection. The formula is as follows: .
[0013] Preferably, the feature fusion and output operation includes: The spatial dimension is compressed to 1×1 by global average pooling, and two layers of point-by-point convolution are connected to generate a high-dimensional feature map. The formula is as follows:
[0014] Where, is a high-dimensional feature map, It is the fused feature map of the output results of the multi-stage UIB module feature mixing operation.
[0015] Preferably, the S3 includes: The high-dimensional feature map outputs a high-dimensional optimized feature map through the channel attention submodule and the spatial attention submodule. The channel attention submodule is connected in series with the spatial attention submodule. The channel attention submodule learns the importance weight of each channel through global average pooling and a fully connected layer. The spatial attention submodule generates a spatial weight map through a convolution operation. The formula is as follows:
[0016]
[0017] Where, is the feature after average pooling, is the maximum pooling result, is the Sigmoid function, is the convolution operation, and are two weight matrices, It is batch normalization.
[0018] Preferably, the S4 includes: The high-dimensional optimized feature map is processed by a multi-scale dilated convolution branch and an image-level feature branch, and output to the channel dimension for splicing. A 1×1 convolution layer is used for dimensionality reduction and fusion to output the final high-dimensional optimized feature map.
[0019] Preferably, the multi-scale dilated convolution branch processing includes: For the input high-dimensional optimized feature map, multi-scale contextual information is captured through parallel convolutional layers with different dilation rates. By stacking convolution operations with different dilation rates in parallel, the receptive field of the convolution kernel is expanded, thereby capturing semantic features within different distance ranges and realizing multi-scale contextual information extraction of input features. The formula is as follows:
[0020]
[0021]
[0022]
[0023] Where, is the void ratio, Optimizing feature maps for high dimensions.
[0024] Preferably, the image-level feature branch processing includes: Perform a global average pooling operation on the input high-dimensional optimized feature map to retain the global statistical information of each channel. Use 1×1 convolution to adjust the number of channels of the compressed features to align them with the output channel number of other branches. At the same time, further extract abstract features through linear transformation. Use bilinear interpolation upsampling to restore the 1×1 size feature map to the spatial resolution of the original input feature map to ensure that it is spatially aligned with the output of the multi-scale void convolution branch. Perform batch normalization and ReLU activation operations on the upsampled features in turn. The formula is as follows:
[0025] Where, is the global average pooling, is the convolution kernel, for Corresponding to the bias term, is the upsampling operation, is a non-linear activation function.
[0026] Preferably, the hybrid loss function in S5 includes: Focal Loss and Dice Loss; The Focal Loss introduces an adjustable focusing parameter to dynamically reduce the weight of easy-to-classify samples. The Dice Loss is based on the Dice coefficient to measure the regional overlap between the predicted segmentation result and the true label. The formula is as follows:
[0027]
[0028]
[0029]
[0030] Where, is the Focal Loss function, is the Dice Loss function, is the mixed loss function, is the balance factor, is the predicted probability, is the focus factor, is the set of positive pixels in the predicted segmented image, is the set of positive pixels in the true segmentation image.
[0031] Beneficial effects: 1. This invention uses MobileNetv4 as the backbone network to replace the Xception network structure in the traditional DeepLabv3+, significantly reducing the model's computational complexity and latency while effectively improving feature expression capabilities and cross-platform deployment compatibility. 2. The present invention uses the ASPP module with different dilation rates of the dilated convolution to extract the multi-scale context information of the image in parallel, thereby enhancing the recognition effect of the multi-scale drivable area; 3. This invention introduces CBAM into the attention mechanism enhancement module, combines channel attention with spatial attention mechanism, effectively improves the model's perception of the target area, and enhances the performance of distinguishing complex background and edge areas in unstructured scenes; 4. The present invention evaluates the model through a hybrid loss function, which accelerates the model training efficiency and improves the model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flow chart of a preferred embodiment of the present invention; Figure 2 It is a schematic diagram of the algorithm structure of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0033] The embodiments of the present invention are described in detail below. The following embodiments are implemented based on the technical solutions of the present invention, and provide detailed implementation methods and specific operating procedures. However, the protection scope of the present invention is not limited to the following embodiments.
[0034] The present invention designs a method for identifying drivable areas on unstructured roads based on improved DeepLabv3+. The technical solution includes the following steps: Figure 1-2 As shown, specifically including: S1: Build an improved DeepLabv3+ model, including the MobileNetv4 network, ASPP module and decoder; S2: Obtain an image of an unstructured road, perform feature extraction based on the MobileNetv4 network, and obtain a high-dimensional feature map; S3: Perform feature optimization processing based on the attention mechanism on the high-dimensional feature map to obtain a high-dimensional optimized feature map; S4: Extract multi-scale context information based on the ASPP module from the high-dimensional optimized feature map to obtain the final high-dimensional optimized feature map; S5: Train the DeepLabv3+ model based on the hybrid loss function; S6: Outputting a recognition image of the unstructured road based on the decoder.
[0035] Preferably, the feature extraction process based on the MobileNetv4 network in S2 includes: The unstructured road image is subjected to initial convolution and downsampling operations, multi-stage UIB module feature mixing operations, and feature fusion and output operations in sequence.
[0036] Preferably, the initial convolution and downsampling operations include: The unstructured road image is passed through a standard convolutional layer to perform preliminary feature extraction and spatial compression, reducing the resolution to 1 / 4 of the original image while extracting low-level features of edges and textures. The formula is as follows:
[0037] Where, is the image after convolution, is a two-dimensional convolutional layer, is the step length; The depth convolution and point-by-point convolution are combined into a single-step operation. The depth convolution is responsible for local spatial feature extraction, and the point-by-point convolution is used to expand the channel dimension. The formula is as follows:
[0038]
[0039] Where, After FusedIB processing, After pointwise convolution processing, After depthwise convolution processing.
[0040] Preferably, the multi-stage UIB module feature mixing operation includes: Divide the low-resolution stage and high-resolution stage and process them separately; At the low-resolution stage, two deep convolutions are connected in series to expand the receptive field. Executed before channel expansion, the second Executed after expansion, the formula is as follows:
[0041] At the high-resolution stage, pure point-by-point convolution stacking is used to mine cross-channel semantic associations through nonlinear activation and channel projection. The formula is as follows: .
[0042] Preferably, the feature fusion and output operation includes: The spatial dimension is compressed to 1×1 by global average pooling, and two layers of point-by-point convolution are connected to generate a high-dimensional feature map. The formula is as follows:
[0043] Where, is a high-dimensional feature map, It is the fused feature map of the output results of the multi-stage UIB module feature mixing operation.
[0044] Specifically, for S2, by giving an RGB image of an unstructured road, input , where H and W represent the height and width of the image, respectively, and C = 3 is the number of channels. Image size and content distribution are usually uncertain, especially in unstructured road scenes, where visual elements vary greatly and exhibit obvious spatial irregularities. This makes it difficult for traditional fixed-scale receptive field feature extraction methods to fully capture the target semantic information. The purpose of the entire feature extraction network is to learn the relationship between the input image I and its high-dimensional feature representation. The mapping relationship between them is shown in Figure 1, where H'×W' represents the gradually compressed spatial resolution and D represents the gradually expanded channel dimension. Through phased modular design, the network reduces computational complexity while maximizing the diversity and robustness of feature expression.
[0045] In addition, in the initial convolution and downsampling stage, image I first passes through a standard convolution layer (3×3 kernel, stride 2), depthwise convolution (DWConv) is responsible for local spatial feature extraction, and point-by-point convolution (PWConv) is used to expand the channel dimension to reduce the memory overhead caused by the separation operation; in the multi-stage UIB module feature mixing stage, the multi-level Universal Inverted Bottleneck (UIB) module is used to dynamically balance the strength of spatial and channel mixing. UIB provides four variants (IB, ConvNext-like, ExtraDW, FFN). This invention mainly uses ExtraDW and ConvNext-like, and FFN is used in the high-resolution stage. Preferably, S3 includes: The high-dimensional feature map outputs a high-dimensional optimized feature map through the channel attention submodule and the spatial attention submodule. The channel attention submodule is connected in series with the spatial attention submodule. The channel attention submodule learns the importance weight of each channel through global average pooling and a fully connected layer. The spatial attention submodule generates a spatial weight map through a convolution operation. The formula is as follows:
[0046]
[0047] Where, is the feature after average pooling, is the maximum pooling result, is the Sigmoid function, is the convolution operation, and are two weight matrices, It is batch normalization.
[0048] Specifically, for S3, a Convolutional Block Attention Module (CBAM) attention mechanism is introduced between the model's Backbone network and the ASPP module to dynamically adjust the channel and spatial weights of the feature map, enhancing the model's perception of key areas on unstructured roads. The CBAM module consists of a channel attention submodule and a spatial attention submodule connected in series. The channel attention submodule learns the importance weights of each channel through global average pooling and fully connected layers, suppressing interference from irrelevant feature channels. The spatial attention submodule generates a spatial weight map through convolution operations, focusing on key areas such as road boundaries and obstacle edges. This dual attention mechanism works synergistically to effectively suppress interference from complex backgrounds such as vegetation shadows and gravel in unstructured scenes, while also improving the recognition accuracy of blurred edges.
[0049] Preferably, S4 includes: The high-dimensional optimized feature map is processed by multi-scale dilated convolution branches and image-level feature branches respectively, and then output to the channel dimension for splicing. A 1×1 convolution layer is used for dimensionality reduction and fusion to output the final high-dimensional optimized feature map.
[0050] Preferably, the multi-scale dilated convolution branch processing includes: For the input high-dimensional optimized feature map, multi-scale contextual information is captured through parallel convolutional layers with different dilation rates. By stacking convolution operations with different dilation rates in parallel, the receptive field of the convolution kernel is expanded, thereby capturing semantic features within different distance ranges and realizing multi-scale contextual information extraction of input features. The formula is as follows:
[0051]
[0052]
[0053]
[0054] Where, is the void ratio, Optimizing feature maps for high dimensions.
[0055] Preferably, the image-level feature branch processing includes: Perform a global average pooling operation on the input high-dimensional optimized feature map to retain the global statistical information of each channel. Use 1×1 convolution to adjust the number of channels of the compressed features to align them with the output channel number of other branches. At the same time, further extract abstract features through linear transformation. Use bilinear interpolation upsampling to restore the 1×1 size feature map to the spatial resolution of the original input feature map to ensure that it is spatially aligned with the output of the multi-scale void convolution branch. Perform batch normalization and ReLU activation operations on the upsampled features in turn. The formula is as follows:
[0056] Where, is the global average pooling, is the convolution kernel, for Corresponding to the bias term, is the upsampling operation, is a non-linear activation function.
[0057] Specifically, for S4, the ASPP module, through a parallel multi-branch architecture, achieves the collaborative fusion of multi-scale contextual information while maintaining high-resolution feature representation, significantly improving the prediction accuracy of semantic segmentation. This method combines dilated convolution operations with different dilation rates to effectively address two key issues existing in traditional segmentation networks: insufficient multi-scale information capture and severe loss of spatial detail.
[0058] Preferably, the hybrid loss function in S5 includes: Focal Loss and Dice Loss; Focal Loss introduces an adjustable focus parameter to dynamically reduce the weight of easy-to-classify samples. Dice Loss is based on the Dice coefficient and measures the regional overlap between the predicted segmentation result and the true label. The formula is as follows:
[0059]
[0060]
[0061]
[0062] Where, is the Focal Loss function, is the Dice Loss function, is the mixed loss function, is the balance factor, is the predicted probability, is the focus factor, is the set of positive pixels in the predicted segmented image, is the set of positive pixels in the true segmentation image.
[0063] Specifically, for S5, in order to improve the performance of the model, the present invention adopts a hybrid loss function composed of Focal_LOSS and Dice_Loss loss functions for training the model, wherein Focal Loss dynamically reduces the weight of easy-to-classify samples by introducing an adjustable focusing parameter, so that the model pays more attention to difficult-to-classify samples and minority classes.
[0064] In addition, the experimental settings of the present invention include the use of an Intel(R) Xeon(R) Gold 5218R processor in hardware, equipped with 90GB of running memory, and an NVIDIA GeForce RTX 3090 graphics card with a video memory capacity of 24GB; the software environment is based on the Python 3.10.9 programming language, combined with the CUDA 12.8 parallel computing architecture, and uses the PyTorch 2.6.0 deep learning framework for model training; the dataset is the public dataset IDD and is evaluated, and the detailed information is shown in Table 1 below.
[0065] Table 1 Dataset information
[0066] The Indian Driving Dataset (IDD) is a large-scale autonomous driving dataset focused on Indian road scenarios, designed to provide real-world visual data for complex and diverse traffic environments. Collected from roads across multiple Indian cities, the dataset covers a wide range of driving scenarios, including urban, suburban, and rural areas, and is highly diverse and challenging. Due to the unique characteristics of the Indian traffic environment, such as mixed traffic flows (cars, motorcycles, bicycles, pedestrians, animals, etc.), unstructured roads, complex lighting conditions, and varied weather conditions, the IDD dataset provides a highly realistic testbed for the development and evaluation of autonomous driving algorithms.
[0067] Compared to European and American datasets (such as Cityscapes and KITTI), IDD focuses more on traffic scenarios in developing countries. Its labeled categories and scene complexity more closely reflect the chaotic environments of the real world. For example, IDD includes numerous scenes of dense interactions between non-motorized vehicles and pedestrians, as well as unpaved roads and temporary construction areas, which are less common in more standardized datasets. Therefore, IDD is not only useful for training and evaluating autonomous driving perception models, but also helps study the robustness of algorithms in complex, unstructured environments.
[0068] In addition, this application provides an evaluation index method, including selecting mean intersection-over-union (MIOU), precision (Precision), mean pixel accuracy (MPA), recall (Recall), and frame rate (FPS). As evaluation indicators, higher values indicate better model performance. The formula is as follows:
[0069]
[0070]
[0071]
[0072]
[0073] The experimental results were primarily evaluated on various unstructured road datasets, with comparisons made with various improvements to the original DeepLab v3+. The evaluation metrics used were mean intersection-over-union (MIOU), precision (Precision), mean pixel accuracy (MPA), recall (Recall), and frame rate (FPS). The experimental results are shown in Table 2.
[0074] Table 2 Performance comparison of different improvement methods
[0075] Among them, it can be seen from the data that on the IDD dataset, although the model proposed in this patent is slightly lower in accuracy than the original deeplabv3+, compared with the parameter volume of the original DeepLabv3+ of 54.72M, the model of this patent is only 4.83M, and the parameter volume is reduced by 91.2%. This degree of lightweighting greatly reduces the storage requirements and computing resource consumption of the model while maintaining the core functions; at the same time, the frame rate has increased by nearly 1.5 times, which greatly improves the recognition rate.
[0076] Compared with other improved versions of DeepLabv3+, the method proposed in this patent shows unique advantages in performance balance. Although the improved version using MobileNetV2+CBAM has a parameter volume of 5.81M, its mIoU is only 74.32%, and the frame rate is as low as 32.57FPS; and although the version using MobileNetV3 has a lower parameter volume (2.92M) and a frame rate of 111.12FPS, its mIoU is only 70.45%, which makes a big compromise in accuracy. In contrast, the model of this patent uses 4.83M parameters, while maintaining key accuracy indicators such as 77.28% mIoU, 85.24% precision, and 87.61% mpa, it increases the frame rate to nearly 3 times that of the original version, achieving a better balance between lightweight, recognition rate, and recognition accuracy.
[0077] In addition, this application also provides an analysis of the results of ablation experiments, including an ablation experiment on the IDD dataset, testing the performance comparison of "not modifying BackBone", "removing the attention mechanism" and "doing both". The evaluation indicators include mean intersection-over-union (MIOU), precision (Precision), average pixel accuracy (MPA), recall (Recall) and frame rate (FPS).
[0078] The ablation experiment results for the IDD dataset are shown in Table 3 below. When both BackBone and attention mechanisms are not modified, the model's accuracy is 84.69%, while the accuracy of the model without BackBone modification is 86.21%, the accuracy of the model without attention mechanism modification is 76.19%, and the accuracy of the proposed method is 77.28%. This result indicates that BackBone replacement and the introduction of the attention mechanism do have a certain impact on model accuracy, with BackBone replacement leading to a more significant performance change. Notably, the accuracy of retaining the original BackBone and removing the attention mechanism (90.27%) is lower than that of retaining only the original BackBone (91.13%) and the complete model (85.24%). This suggests that completely removing the attention mechanism may negatively impact the model's feature extraction capabilities.
[0079] In terms of real-time performance, the complete model proposed in this patent achieved 131.58 FPS, significantly higher than the original configuration's 55.37 FPS, validating the effectiveness of the MobileNetV4 backbone network in improving inference speed. The configuration with only the attention mechanism removed achieved the highest frame rate of 133.34 FPS, demonstrating that the attention mechanism does incur a certain computational overhead. In terms of mean pixel accuracy (mPA), the complete model's 87.61% performance outperformed the 86.74% achieved by simply replacing the BackBone, demonstrating that the attention mechanism helps improve pixel-level classification accuracy. Overall, while replacing the BackBone results in some accuracy loss, the resulting lightweighting and speed improvements are significant. While the attention mechanism slightly impacts inference speed, it positively impacts the model's feature selection capabilities and classification accuracy. The synergy between these two modules ensures that the final model achieves a good balance between speed, accuracy, and model complexity.
[0080] Table 3 Ablation experiment results on IDD dataset
[0081] The above describes in detail the preferred embodiments of the present invention. It should be understood that numerous modifications and variations based on the concepts of the present invention are possible by those skilled in the art without inventive effort. Therefore, any technical solution that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A method for identifying drivable areas on unstructured roads based on improved DeepLabv3+, characterized in that: include: S1: Build an improved DeepLabv3+ model, including the MobileNetv4 network, ASPP module and decoder; S2: Obtain an image of an unstructured road, perform feature extraction based on the MobileNetv4 network, and obtain a high-dimensional feature map; S3: Performing feature optimization processing based on the attention mechanism on the high-dimensional feature map to obtain a high-dimensional optimized feature map; S4: performing multi-scale context information extraction based on the ASPP module on the high-dimensional optimized feature map to obtain a final high-dimensional optimized feature map; S5: Train the DeepLabv3+ model based on the hybrid loss function; S6: Outputting a recognition image of the unstructured road based on the decoder.
2. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 1, characterized in that: The feature extraction process based on the MobileNetv4 network in S2 includes: The image of the unstructured road is sequentially subjected to initial convolution and downsampling operations, multi-stage UIB module feature mixing operations, and feature fusion and output operations.
3. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 2, characterized in that: The initial convolution and downsampling operations include: The unstructured road image is passed through a standard convolutional layer to perform preliminary feature extraction and spatial compression, reducing the resolution to 1 / 4 of the original image while extracting low-level features of edges and textures. The formula is as follows: Where, is the image after convolution, is a two-dimensional convolutional layer, is the step length; The depth convolution and point-by-point convolution are combined into a single-step operation. The depth convolution is responsible for local spatial feature extraction, and the point-by-point convolution is used to expand the channel dimension. The formula is as follows: Where, After FusedIB processing, After pointwise convolution processing, After depthwise convolution processing.
4. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 3, characterized in that: The multi-stage UIB module feature mixing operation includes: Divide the low-resolution stage and high-resolution stage and process them separately; At the low-resolution stage, two deep convolutions are connected in series to expand the receptive field. Executed before channel expansion, the second Executed after expansion, the formula is as follows: At the high-resolution stage, pure point-by-point convolution stacking is used to mine cross-channel semantic associations through nonlinear activation and channel projection. The formula is as follows: 。 5. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 2, characterized in that: The feature fusion and output operation includes: The spatial dimension is compressed to 1×1 by global average pooling, and two layers of point-by-point convolution are connected to generate a high-dimensional feature map. The formula is as follows: Where, is a high-dimensional feature map, It is the fused feature map of the output results of the multi-stage UIB module feature mixing operation.
6. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 1, characterized in that: The S3 includes: The high-dimensional feature map outputs a high-dimensional optimized feature map through the channel attention submodule and the spatial attention submodule. The channel attention submodule is connected in series with the spatial attention submodule. The channel attention submodule learns the importance weight of each channel through global average pooling and a fully connected layer. The spatial attention submodule generates a spatial weight map through a convolution operation. The formula is as follows: Where, is the feature after average pooling, is the maximum pooling result, is the Sigmoid function, is the convolution operation, and are two weight matrices, It is batch normalization.
7. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 1, characterized in that: The S4 includes: The high-dimensional optimized feature map is processed by a multi-scale dilated convolution branch and an image-level feature branch, and output to the channel dimension for splicing. A 1×1 convolution layer is used for dimensionality reduction and fusion to output the final high-dimensional optimized feature map.
8. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 7, characterized in that: The multi-scale dilated convolution branch processing includes: For the input high-dimensional optimized feature map, multi-scale contextual information is captured through parallel convolutional layers with different dilation rates. By stacking convolution operations with different dilation rates in parallel, the receptive field of the convolution kernel is expanded, thereby capturing semantic features within different distance ranges and realizing multi-scale contextual information extraction of input features. The formula is as follows: Where, is the void ratio, Optimizing feature maps for high dimensions.
9. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 7, characterized in that: The image-level feature branch processing includes: Perform a global average pooling operation on the input high-dimensional optimized feature map to retain the global statistical information of each channel. Use 1×1 convolution to adjust the number of channels of the compressed features to align them with the output channel number of other branches. At the same time, further extract abstract features through linear transformation. Use bilinear interpolation upsampling to restore the 1×1 size feature map to the spatial resolution of the original input feature map to ensure that it is spatially aligned with the output of the multi-scale void convolution branch. Perform batch normalization and ReLU activation operations on the upsampled features in turn. The formula is as follows: Where, is the global average pooling, is the convolution kernel, for Corresponding to the bias term, is the upsampling operation, is a non-linear activation function.
10. The method for identifying drivable areas on unstructured roads based on improved DeepLabv3+ according to claim 1, characterized in that: The hybrid loss function in S5 includes: Focal Loss and Dice Loss; The Focal Loss introduces an adjustable focusing parameter to dynamically reduce the weight of easy-to-classify samples. The Dice Loss is based on the Dice coefficient to measure the regional overlap between the predicted segmentation result and the true label. The formula is as follows: Where, is the Focal Loss function, is the Dice Loss function, is the mixed loss function, is the balance factor, is the predicted probability, is the focus factor, is the set of positive pixels in the predicted segmented image, is the set of positive pixels in the true segmentation image.
Citation Information
Cited By
Corn phenological period monitoring method based on unmanned aerial vehicle image and deep learning
CN121921688A