A Dual-Branch Lightweight Detectable Driving Area Method Based on Low-Level Features
Through the dual-branch lightweight detection method, combining high-resolution and low-resolution branches to extract features, and using a lightweight guide fusion module, the problems of inference delay and large calculation amount of the feasible area detection model in the prior art are solved, and efficient and real-time detection results are achieved.
Patent Information
- Application Number
- CN202510301647.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Existing deep learning-based feasible area detection methods are difficult to reason in real-time on resource-constrained devices, and the inference time and calculation volume of the model are increasing exponentially, resulting in inference delays and affecting the real-time decision-making and safety of autonomous driving.
The dual-branch lightweight detection method is used to extract the low-level feature and context information of the image through high-resolution branches and low-resolution branches, and the lightweight guide fusion module is used to balance the feature information of the two, reducing the calculation amount and inference time.
It realizes that while ensuring detection accuracy, the real-time and inference speed of the detection model are improved, the calculation amount and inference delay are reduced, and the response speed and decision-making efficiency of the autonomous driving system are improved.
Smart Images

Figure CN119810785B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road drivable area detection, and the present invention is a lightweight dual-branch drivable area detection method based on low-level features. Background Art
[0002] Road drivable area detection is a key link in intelligent driving technology. The main goal is to automatically identify and accurately segment the area suitable for the safe driving of vehicles from a complex road environment. It provides crucial environmental perception data for the path planning and dynamic obstacle avoidance of the automatic driving system, ensuring safe, stable and efficient driving. A clear division of the drivable area helps the vehicle to adjust the driving path in real time and avoid collisions with pedestrians, other vehicles and obstacles, which is particularly important in complex traffic and dynamic environments. Currently, improving the recognition accuracy and inference speed of the network remains the core issue in the research of this task. How to maintain the high real-time performance and fast inference of the detection model while ensuring high accuracy is still a very challenging task.
[0003] In recent years, significant progress has been made in the field of road drivable area detection based on deep learning, especially in image understanding and road feature recognition. Many deep learning networks, especially convolutional neural networks (CNNs) and Transformer architectures, have been widely applied to this task and achieved good detection results. However, most current deep learning-based road drivable area detection methods, especially those aiming for high accuracy, still face some severe challenges. Due to the large volume of existing model files and slow loading speed, it has become a huge obstacle to the deployment of embedded devices and edge devices. Especially on these resource-constrained devices, it is more difficult to perform real-time inference. Even on high-performance GPU platforms, the ability to process sensor data in real time is still limited. Secondly, as the scale of the deep neural network increases, the inference time and computational complexity of the model also increase exponentially. Especially in autonomous driving applications, real-time performance is crucial. The high computational complexity and complex inference process lead to a large inference delay, which poses a serious threat to real-time decision-making and safety. Many tasks in autonomous driving require reacting to sensor data within milliseconds. Any delay may bring potential safety hazards and affect the reaction speed and decision-making efficiency of the vehicle.
[0004] In summary, for the task of road drivable area detection, from the perspective of lightweight, this paper solves the problems of large computational complexity, low accuracy and slow inference speed when the model is deployed in an embedded system, so as to achieve a balance between accuracy and timeliness;
[0005] For the task of drivable area detection, current deep learning methods and training processes vary, but most methods ignore a key feature of roads, that is, roads exist as background elements rather than typical "things". In an image, a road is not as easy to identify as an object with a clear semantic category. It is more like a "thing" (background), that is, part of the environment. Based on the above findings, the primary stage of current mainstream network models is sufficient to represent most of the pixels of the road for detection. Therefore, our network adopts a detection model dominated by low-level features. Specifically, our network adopts a dual-branch structure, namely a high-resolution branch and a low-resolution branch; the high-resolution branch extracts the low-level feature representation information of the drivable area in the image, and the low-resolution branch aims to quickly capture the context information in the image; common ways to combine the two features include element-wise summation and feature concatenation. However, this simple combination ignores the diversity of the two types of feature information and is difficult to balance the differences in context between low-level features and high-level features at the same time. Therefore, a lightweight fusion guidance layer is used to balance the spatial details of low-level features and the semantic context information of high-level features, and suppress irrelevant changes through precise feature guidance; then, the feature information extracted from the two branches is input into the detail module and the segmentation module respectively, and the boundary information is optimized through the detail and segmentation loss functions to improve the detection network's ability to capture small targets and details and the pixel-level classification accuracy of the graph. These two steps only work during training and do not increase the inference cost of the network; finally, the feature information output by the lightweight guidance fusion layer is input into the segmentation head again, and the segmentation head performs stage loss feedback to ensure that the low-resolution branch can learn richer semantic information and the high-resolution branch can improve the model's perception ability of boundary information, and finally obtain the output of the drivable area detection image. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides a dual-branch lightweight drivable area detection method based on low-level features to solve the problems proposed in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solution: A dual-branch lightweight drivable area detection method based on low-level features, comprising the following steps:
[0008] Step S1: Construct a dataset image;
[0009] Step S2: Construct a dual-branch lightweight model, which is composed of a high-resolution branch module, a low-resolution branch module, a detail module, a segmentation module, and a lightweight guidance fusion module; the high-resolution branch module is composed of a high-resolution feature extraction network and an adaptive enhancement module; the low-resolution branch module is composed of a low-resolution feature extraction network and a cross convolution module; the lightweight guidance fusion module includes a guidance restoration module;
[0010] Step S3: Input the dataset image into the high-resolution feature extraction network for processing to obtain feature information, and input the feature information into the adaptive enhancement module for processing to obtain high-resolution branch features;
[0011] Step S4: Input the dataset image into the low-resolution feature extraction network for processing to obtain the extraction information of the second layer, and input the extraction information of the second layer into the cross convolution module for processing to obtain low-resolution branch features;
[0012] Step S5: Input the high-resolution branch features into the detail module for processing to obtain a detail feature map;
[0013] Step S6: Input the low-resolution branch features into the segmentation module for processing to obtain a segmentation feature map;
[0014] Step S7: Optimize the lightweight guidance fusion module through the detail feature map and the segmentation feature map, input the high-resolution branch features and the low-resolution branch features into the optimized lightweight guidance fusion module for processing to obtain the output feature information of the lightweight guidance fusion layer, and input the high-resolution branch features, the low-resolution branch features, and the output feature information of the lightweight guidance fusion layer into the guidance restoration module for processing to obtain the output of the drivable area detection feature information.
[0015] Further, for the dataset image in step S1, the specific process is as follows:
[0016] Obtain a road segmentation dataset, where the road segmentation dataset includes training images and test images; process the training images and test images to obtain the dataset image.
[0017] Further, the high-resolution feature extraction network is a ResNet18 module, and the ResNet18 module includes an L1 layer; the L1 layer is composed of high-resolution feature residual blocks, and each high-resolution feature residual block contains two high-resolution feature residual units, and each residual unit contains two 3×3 convolutional layers;
[0018] Input the dataset image into the high-resolution feature extraction network for processing to obtain feature information F1, expressed as:
[0019] ;
[0020] In the formula, represents the dataset image; represents batch normalization; represents a 3×3 convolutional layer; represents an activation function.
[0021] Further, in step S3, the high-resolution branch features are obtained, and the specific process is as follows:
[0022] Input the feature information F1 into the adaptive enhancement module for processing to obtain the high-resolution branch features;
[0023] The adaptive enhancement module consists of a structure with four dilated convolutional branches; the first dilated convolutional branch consists of a 1×1 standard convolution; the second dilated convolutional branch consists of a 1×1 standard convolution, a 3×1 standard convolution, and a 3×3 dilated convolution; the third dilated convolutional branch consists of a 1×1 standard convolution, a 1×3 standard convolution, and a 3×3 dilated convolution; the fourth dilated convolutional branch consists of a 1×1 standard convolution and a 3×3 standard convolution;
[0024] The feature information F1 only undergoes a 1×1 standard convolution operation in the first dilated convolutional branch to obtain the output of the first dilated convolutional branch; after passing through a 1×1 standard convolution in the second dilated convolutional branch, it performs a 3×1 standard convolution and a 3×3 dilated convolution operation to obtain the output of the second dilated convolutional branch; after passing through a 1×1 standard convolution in the third dilated convolutional branch, it sequentially performs a 1×3 standard convolution and a 3×3 dilated convolution operation to obtain the output of the third dilated convolutional branch; after passing through a 1×1 standard convolution in the fourth dilated convolutional branch, it performs a 3×3 standard convolution operation to obtain the output of the fourth dilated convolutional branch; add the output of the second dilated convolutional branch and the output of the third dilated convolutional branch to obtain the added output, and perform a concatenation operation on the added output with the output of the first dilated convolutional branch and the output of the fourth dilated convolutional branch to obtain the high-resolution branch features .
[0025] Further, the low-resolution feature extraction network is a ResNet18 module, and the ResNet18 module includes layer L1 and layer L2; the ResNet18 module consists of low-resolution feature residual blocks, and each low-resolution feature residual block contains two low-resolution feature residual units, and each residual unit contains two 3×3 convolutional layers;
[0026] Input the dataset image into the low-resolution feature extraction network for processing to obtain the extraction information F2 of the second layer, which is expressed as:
[0027] ;
[0028] ;
[0029] where represents the output of the first low-resolution feature residual unit in the low-resolution feature extraction network.
[0030] Further, the low-resolution branch features in step S4 are specifically obtained as follows:
[0031] Input the extraction information F2 of the second layer into the cross-convolution module for processing to obtain low-resolution branch features;
[0032] The cross-convolution module consists of a 1×5 depthwise row convolution, a 5×1 depthwise column convolution, a 1×1 convolution layer, a batch normalization, and an activation function;
[0033] The extraction information F2 of the second layer first undergoes a 1×5 depthwise row convolution operation in the cross-convolution module to capture the context information in the horizontal direction of the extraction information F2 of the second layer, then undergoes a 5×1 depthwise column convolution operation to extract the context information in the vertical direction of the extraction information F2 of the second layer, uses a 1×1 convolution layer to fuse the context information in the horizontal direction of the extraction information F2 of the second layer and the context information in the vertical direction of the extraction information F2 of the second layer, and completes the non-linear transformation through batch normalization and the activation function to obtain the output of the non-linear transformation. Add the extraction information F2 of the second layer and the output of the non-linear transformation to obtain the low-resolution branch features , which is expressed as:
[0034] ;
[0035] In the formula, represents the 1×1 convolution layer; represents the 1×5 depthwise row convolution; represents the 5×1 depthwise column convolution.
[0036] Further, the detailed feature map in step S5 is specifically obtained as follows:
[0037] Input the high-resolution branch features into the detail module for processing. The detail module includes a detail head; the detail head consists of a 3×3 convolution layer and a 1×1 convolution layer; the high-resolution branch features first pass through the 3×3 convolution layer and the 1×1 convolution layer in the detail head, and then operate through batch normalization and the activation function to generate the detailed feature map , which is expressed as:
[0038] ;
[0039] In the formula, represents the upsampling operation.
[0040] Further, the segmentation feature map in step S6 is specifically obtained as follows:
[0041] Input the low-resolution branch features In the input splitting module, the splitting module includes a splitting head, which is composed of a 3×3 convolutional layer and a 1×1 convolutional layer, and the low-resolution branch features In the splitting head, it first passes through a 3×3 convolutional layer and a 1×1 convolutional layer, and then generates a splitting feature map through batch normalization and activation function operations , expressed as:
[0042] .
[0043] Furthermore, the lightweight guiding fusion module is composed of a high-resolution branch, a low-resolution branch, and a 3×3 convolutional layer;
[0044] The high-resolution branch is composed of a first branch and a second branch; the low-resolution branch is composed of a third branch and a fourth branch; the first branch is composed of a 3×3 depthwise separable convolution and a 1×1 convolutional layer; the second branch is composed of a 3×3 convolutional layer and a downsampling; the third branch is composed of a 3×3 convolutional layer and an upsampling; the fourth branch is composed of a 3×3 depthwise separable convolution and a 1×1 convolutional layer;
[0045] Based on the detailed feature map and the splitting feature map Optimize the lightweight guiding fusion module to obtain the optimized lightweight guiding fusion module;
[0046] Input the high-resolution branch features and the low-resolution branch features into the optimized lightweight guiding fusion module;
[0047] The high-resolution branch features In the optimized lightweight guiding fusion module, input the 3×3 depthwise separable convolution of the first branch and the 3×3 convolutional layer of the second branch respectively, and obtain the outputs of the 3×3 depthwise separable convolution of the first branch and the 3×3 convolutional layer of the second branch respectively;
[0048] The low-resolution branch features In the optimized lightweight guiding fusion module, input the 3×3 convolutional layer of the third branch and the 3×3 depthwise separable convolution of the fourth branch respectively, and obtain the outputs of the 3×3 convolutional layer of the third branch and the 3×3 depthwise separable convolution of the fourth branch respectively;
[0049] Input the output of the 3×3 depthwise separable convolution of the first branch into the 1×1 convolutional layer of the first branch; input the output of the 3×3 depthwise separable convolution of the fourth branch into the 1×1 convolutional layer of the fourth branch, and obtain the outputs of the 1×1 convolutional layer of the first branch and the 1×1 convolutional layer of the fourth branch respectively;
[0050] Downsample the output of the second branch 3×3 convolutional layer to obtain a downsampled feature map, and upsample the output of the third branch 3×3 convolutional layer to obtain an upsampled feature map;
[0051] Multiply the downsampled feature map by the output of the fourth branch 1×1 convolutional layer to obtain low-resolution branch weight features, multiply the upsampled feature map by the output of the first branch 1×1 convolutional layer to obtain high-resolution branch weight features, add the high-resolution branch weight features and the low-resolution branch weight features element-wise to obtain fused features, and then integrate the fused features through a 3×3 convolutional layer to generate the output feature information of the lightweight guiding fusion layer 。
[0052] Further, the output of the drivable area detection feature information in step S7 is as follows:
[0053] The output feature information of the lightweight guiding fusion layer 、high-resolution branch features and low-resolution branch features are input into the guiding restoration module for processing to obtain the output of the drivable area detection feature information;
[0054] The guiding restoration module consists of an upper branch, a middle branch, and a lower branch; the upper branch consists of an upsampling layer, a 1×1 convolutional layer, and a 3×3 convolutional layer;
[0055] The middle branch consists of a global average pooling layer, three 3×3 convolutional layers, and a 1×1 convolutional layer;
[0056] The lower branch consists of an upsampling layer, a 1×1 convolutional layer, and a 3×3 convolutional layer;
[0057] High-resolution branch features In the upper branch, first input the upsampling to obtain the upsampled high-resolution branch features 、and input the upsampled high-resolution branch features sequentially into the 1×1 convolutional layer and the 3×3 convolutional layer to obtain the 3×3 convolutional high-resolution branch features, and add the 3×3 convolutional high-resolution branch features and the upsampled high-resolution branch features element-wise to obtain the upper branch high-resolution branch features ;
[0058] Low-resolution branch features In the lower branch, first input the upsampling to obtain the upsampled low-resolution branch features ,and input the upsampled low-resolution branch features Input the 1×1 convolutional layer and 3×3 convolutional layer in sequence to obtain the low-resolution branch features of the 3×3 convolutional layer, and combine the low-resolution branch features of the 3×3 convolutional layer with the upsampled low-resolution branch features Perform element-wise addition to obtain the low-resolution branch features of the lower branch ;
[0059] The output feature information of the lightweight guiding fusion layer In the middle branch, first input the global average pooling layer to obtain the output feature information of the global average pooling layer for the lightweight guiding fusion layer , and input the output feature information of the global average pooling layer for the lightweight guiding fusion layer into a 3×3 convolutional layer to obtain the output feature information of the 3×3 convolutional layer for the lightweight guiding fusion layer. Multiply the output feature information of the 3×3 convolutional layer for the lightweight guiding fusion layer, the low-resolution branch features of the 3×3 convolutional layer, and the high-resolution branch features of the 3×3 convolutional layer to obtain the output feature information of the multiplied lightweight guiding fusion layer. Then, integrate the multiplied information through another 3×3 convolutional layer to obtain the output feature information of the middle branch for the lightweight guiding fusion layer ;
[0060] Combine the high-resolution branch features of the upper branch , the low-resolution branch features of the lower branch and the output feature information of the middle branch for the lightweight guiding fusion layer through a concatenation operation to obtain the concatenated feature map , and input the concatenated feature map into a 3×3 convolutional layer and a 1×1 convolutional layer in sequence to obtain the output feature map of the 1×1 convolutional layer for the concatenated feature map, which is the output feature map of the guiding recovery module ;
[0061] Input the output feature map of the guiding recovery module into the segmentation head of the segmentation module for processing to obtain the output F of the drivable area detection feature information, expressed as:
[0062] .
[0063] Compared with the existing technologies, the present invention has the following beneficial effects:
[0064] (1)From the perspective of a dual-branch lightweight model, the high-resolution branch module only uses the first layer of the ResNet18 module to extract high-resolution branch features, while the low-resolution branch module uses the first two layers of the ResNet18 module to extract high-level semantic context information. This design significantly reduces the computational load and improves the inference speed of the dual-branch lightweight model. By introducing an adaptive enhancement module, the expressive ability of the high-resolution branch features is enhanced, effectively suppressing noise and background interference, reducing misidentifications, and improving the adaptability and robustness of the dual-branch lightweight model to tasks.
[0065] (2)By replacing the third layer of the ResNet18 module in the low-resolution branch module with a cross-convolution module, the present invention significantly reduces the number of parameters and computational overhead while successfully achieving a receptive field equivalent to that of the third layer of the ResNet18 module, providing the dual-branch lightweight model with stronger feature extraction capabilities.
[0066] (3)The present invention uses a lightweight guiding fusion module to fuse the high-resolution branch features and the low-resolution branch features, balancing the spatial details of the low-resolution branch features and the high-level semantic context information of the high-resolution branch features, guiding the suppression of irrelevant changes, optimizing the fusion effect between features, and improving the detection performance.
[0067] (4)The present invention designs a detail module and a segmentation module to generate a detail feature map and a segmentation feature map, helping the lightweight guiding fusion module learn richer spatial detail features and enhancing the ability to capture high-level semantic context information of the high-resolution branch features and the low-resolution branch features. The detail module and the segmentation module only function during the training of the dual-branch lightweight model and do not work during the inference stage. Therefore, the dual-branch lightweight model does not increase the computational load of the entire model.
[0068] (5)The present invention designs a guiding recovery module to guide and recover the high-resolution branch features, the low-resolution branch features, and the output feature information of the lightweight guiding fusion layer, helping to capture the detailed information of the drivable area detection features output in the dataset images, especially in the identification of edges, small targets, and complex boundaries. At the same time, it enhances the adaptability of the dual-branch lightweight model to different input data, helping the dual-branch lightweight model to cope with environmental changes (such as lighting changes, road conditions, etc.) and improving the adaptability of the dual-branch lightweight model. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a flowchart of the present invention.
[0070] Figure 2 is a schematic structural diagram of the adaptive enhancement module of the present invention.
[0071] Figure 3Schematic diagram of the lightweight guiding fusion module of the present invention.
[0072] Figure 4 Schematic diagram of the guiding restoration module of the present invention. Detailed implementation manners
[0073] As Figure 1 shown, the present invention provides a technical solution: a dual-branch lightweight drivable area detection method based on low-level features, including the following steps:
[0074] Step S1: Construct a dataset image;
[0075] Step S2: Construct a dual-branch lightweight model, which consists of a high-resolution branch module, a low-resolution branch module, a detail module, a segmentation module, and a lightweight guiding fusion module; the high-resolution branch module consists of a high-resolution feature extraction network and an adaptive enhancement module; the low-resolution branch module consists of a low-resolution feature extraction network and a cross convolution module; the lightweight guiding fusion module includes a guiding restoration module;
[0076] Step S3: Input the dataset image into the high-resolution feature extraction network for processing to obtain feature information, and input the feature information into the adaptive enhancement module for processing to obtain high-resolution branch features;
[0077] Step S4: Input the dataset image into the low-resolution feature extraction network for processing to obtain the extraction information of the second layer, and input the extraction information of the second layer into the cross convolution module for processing to obtain low-resolution branch features;
[0078] Step S5: Input the high-resolution branch features into the detail module for processing to obtain a detail feature map;
[0079] Step S6: Input the low-resolution branch features into the segmentation module for processing to obtain a segmentation feature map;
[0080] Step S7: Optimize the lightweight guiding fusion module through the detail feature map and the segmentation feature map, input the high-resolution branch features and the low-resolution branch features into the optimized lightweight guiding fusion module for processing to obtain the output feature information of the lightweight guiding fusion layer, and input the high-resolution branch features, the low-resolution branch features, and the output feature information of the lightweight guiding fusion layer into the guiding restoration module for processing to obtain the output of the drivable area detection feature information.
[0081] Among them, for the dataset image in step S1, the specific process is:
[0082] This paper uses the KITTI-Road dataset. KITTI-Road is a road segmentation dataset, which contains 289 training images and 290 test images;
[0083] The resolution range of the training images and test images is from 370×1224 to 375×1242. The training images and test images with the resolution range from 370×1224 to 375×1242 are processed by unified padding operation. Through the padding operation, the training images and test images are adjusted to 375×1240 to obtain the dataset images after the padding operation.
[0084] Among them, the high-resolution feature extraction network is the ResNet18 module. The ResNet18 module contains L1 layer; The L1 layer contains 64 channels; The L1 layer is composed of high-resolution feature residual blocks, and the high-resolution feature residual block contains two high-resolution feature residual units, and the residual unit contains two 3×3 convolutional layers;
[0085] The dataset images are input into the high-resolution feature extraction network for processing to obtain the feature information F1, which is expressed as:
[0086] ;
[0087] In the formula, represents the feature information; represents the dataset images; represents batch normalization; represents the 3×3 convolutional layer; represents the activation function.
[0088] As Figure 2 shown, among them, the specific process of obtaining the high-resolution branch features in step S3 is as follows:
[0089] The feature information F1 is input into the adaptive enhancement module for processing to obtain the high-resolution branch features;
[0090] The adaptive enhancement module is composed of the structure of four dilated convolution branches; The first dilated convolution branch consists of a 1×1 standard convolution; The second dilated convolution branch consists of a 1×1 standard convolution, a 3×1 standard convolution and a 3×3 dilated convolution; The third dilated convolution branch consists of a 1×1 standard convolution, a 1×3 standard convolution and a 3×3 dilated convolution; The fourth dilated convolution branch consists of a 1×1 standard convolution and a 3×3 standard convolution;
[0091] The feature information F1 obtains the output of the first dilated convolutional branch only through 1×1 standard convolutional operations in the first dilated convolutional branch; after passing through 1×1 standard convolution in the second dilated convolutional branch, the feature information F1 performs 3×1 standard convolution and 3×3 dilated convolution operations to obtain the output of the second dilated convolutional branch; after passing through 1×1 standard convolution in the third dilated convolutional branch, the feature information F1 sequentially performs 1×3 standard convolution and 3×3 dilated convolution operations to obtain the output of the third dilated convolutional branch; after passing through 1×1 standard convolution in the fourth dilated convolutional branch, the feature information F1 performs 3×3 standard convolution operation to obtain the output of the fourth dilated convolutional branch; the output of the second dilated convolutional branch and the output of the third dilated convolutional branch are added to obtain the added output, and the added output is cascaded with the output of the first dilated convolutional branch and the output of the fourth dilated convolutional branch to obtain the high-resolution branch feature , which is expressed as:
[0092] ;
[0093] ;
[0094] ;
[0095] ;
[0096] In the formula, represents the output of the first dilated convolutional branch; represents the added output; represents the output of the fourth dilated convolutional branch; represents 3×3 standard convolution; represents 1×1 standard convolution; represents 3×3 dilated convolution; represents 3×1 standard convolution; represents 1×3 standard convolution; represents the cascading operation; ⊕ represents the addition operation.
[0097] Among them, the low-resolution feature extraction network is a ResNet18 module, and the ResNet18 module includes L1 layer and L2 layer; the ResNet18 module is composed of low-resolution feature residual blocks, and the low-resolution feature residual block contains two low-resolution feature residual units, and the residual unit contains two 3×3 convolutional layers;
[0098] The dataset image is input into the low-resolution feature extraction network for processing to obtain the extraction information F2 of the second layer, which is expressed as:
[0099] ;
[0100] ;
[0101] In the formula, represents the output of the first low-resolution feature residual unit in the low-resolution feature extraction network.
[0102] Among them, the low-resolution branch feature in step S4 is specifically as follows:
[0103] Input the extraction information F2 of the second layer into the cross-convolution module for processing to obtain the low-resolution branch feature;
[0104] The cross-convolution module consists of a 1×5 depthwise row convolution, a 5×1 depthwise column convolution, a 1×1 convolution layer, a batch normalization, and an activation function;
[0105] The extraction information F2 of the second layer first performs a 1×5 depthwise row convolution operation in the cross-convolution module to capture the context information in the horizontal direction of the extraction information F2 of the second layer, then performs a 5×1 depthwise column convolution operation to extract the context information in the vertical direction of the extraction information F2 of the second layer, uses a 1×1 convolution layer to fuse the context information in the horizontal direction of the extraction information F2 of the second layer and the context information in the vertical direction of the extraction information F2 of the second layer, and completes the non-linear transformation through batch normalization and the activation function to obtain the output of the non-linear transformation. Add the extraction information F2 of the second layer and the output of the non-linear transformation to obtain the low-resolution branch feature , which is expressed as:
[0106] ;
[0107] In the formula, represents the 1×1 convolution layer; represents the 1×5 depthwise row convolution; represents the 5×1 depthwise column convolution.
[0108] Among them, the detailed feature map in step S5 is specifically as follows:
[0109] Input the high-resolution branch feature into the detail module for processing. The detail module includes a detail head; the detail head consists of a 3×3 convolution layer and a 1×1 convolution layer; the high-resolution branch feature first passes through the 3×3 convolution layer and the 1×1 convolution layer in the detail head, and then operates through batch normalization and the activation function to generate the detailed feature map , which is expressed as:
[0110] ;
[0111] In the formula, represents the upsampling operation.
[0112] Among them, the specific process of splitting the feature map in step S6 is as follows:
[0113] Input the low-resolution branch feature into the splitting module. The splitting module includes a splitting head. The splitting head consists of a 3×3 convolutional layer and a 1×1 convolutional layer. The low-resolution branch feature first passes through the 3×3 convolutional layer and the 1×1 convolutional layer in the splitting head, and then generates a segmentation feature map through batch normalization and activation function operations , which is expressed as:
[0114] .
[0115] As Figure 3 shown, among them, the lightweight guiding fusion module consists of a high-resolution branch, a low-resolution branch, and a 3×3 convolutional layer;
[0116] The high-resolution branch consists of a first branch and a second branch; the low-resolution branch consists of a third branch and a fourth branch; the first branch consists of a 3×3 depthwise separable convolution and a 1×1 convolutional layer; the second branch consists of a 3×3 convolutional layer and a downsampling; the third branch consists of a 3×3 convolutional layer and an upsampling; the fourth branch consists of a 3×3 depthwise separable convolution and a 1×1 convolutional layer;
[0117] Based on the detailed feature map and the segmentation feature map optimize the lightweight guiding fusion module to obtain the optimized lightweight guiding fusion module;
[0118] Input the high-resolution branch feature and the low-resolution branch feature into the optimized lightweight guiding fusion module;
[0119] The high-resolution branch feature is respectively input into the 3×3 depthwise separable convolution of the first branch and the 3×3 convolutional layer of the second branch in the optimized lightweight guiding fusion module, and the outputs of the 3×3 depthwise separable convolution of the first branch and the 3×3 convolutional layer of the second branch are respectively obtained;
[0120] The low-resolution branch feature is respectively input into the 3×3 convolutional layer of the third branch and the 3×3 depthwise separable convolution of the fourth branch in the optimized lightweight guiding fusion module, and the outputs of the 3×3 convolutional layer of the third branch and the 3×3 depthwise separable convolution of the fourth branch are respectively obtained;
[0121] Input the output of the first-branch 3×3 depthwise separable convolution into the 1×1 convolution layer of the first branch; input the output of the fourth-branch 3×3 depthwise separable convolution into the 1×1 convolution layer of the fourth branch, and obtain the output of the 1×1 convolution layer of the first branch and the output of the 1×1 convolution layer of the fourth branch respectively, which are expressed as:
[0122] ;
[0123] ;
[0124] In the formula, represents the output of the first-branch 3×3 depthwise separable convolution; represents the 3×3 depthwise separable convolution; represents the output of the fourth-branch 3×3 depthwise separable convolution; represents the output of the 1×1 convolution layer of the first branch; represents the output of the 1×1 convolution layer of the fourth branch;
[0125] Downsample the output of the second-branch 3×3 convolution layer to obtain the downsampled feature map, and upsample the output of the third-branch 3×3 convolution layer to obtain the upsampled feature map, which are expressed as:
[0126] ;
[0127] ;
[0128] In the formula, represents the downsampled feature map; represents the upsampled feature map; represents downsampling;
[0129] Multiply the downsampled feature map by the output of the 1×1 convolution layer of the fourth branch to obtain the low-resolution branch weight feature, multiply the upsampled feature map by the output of the 1×1 convolution layer of the first branch to obtain the high-resolution branch weight feature, add the high-resolution branch weight feature and the low-resolution branch weight feature element-wise to obtain the fused feature, and then integrate the fused feature through a 3×3 convolution layer to generate the output feature information of the lightweight guiding fusion layer , which is expressed as:
[0130] ;
[0131] In the formula, represents the activation function; represents the dot product operation.
[0132] As Figure 4 shown, among which, the output of the drivable area detection feature information in step S7 is as follows:
[0133] Input the output feature information of the lightweight guidance fusion layer , the high-resolution branch feature and the low-resolution branch feature into the guidance restoration module for processing to obtain the output of the drivable area detection feature information;
[0134] The guidance restoration module consists of an upper branch, a middle branch, and a lower branch; the upper branch consists of an upsampling layer, a 1×1 convolutional layer, and a 3×3 convolutional layer;
[0135] The middle branch consists of a global average pooling layer, three 3×3 convolutional layers, and a 1×1 convolutional layer;
[0136] The lower branch consists of an upsampling layer, a 1×1 convolutional layer, and a 3×3 convolutional layer;
[0137] The high-resolution branch feature First input the upsampling in the upper branch to obtain the upsampled high-resolution branch feature , and input the upsampled high-resolution branch feature sequentially into the 1×1 convolutional layer and the 3×3 convolutional layer to obtain the 3×3 convolutional high-resolution branch feature. Add the 3×3 convolutional high-resolution branch feature and the upsampled high-resolution branch feature element-wise to obtain the high-resolution branch feature of the upper branch , denoted as;
[0138] ;
[0139] ;
[0140] The low-resolution branch feature First input the upsampling in the lower branch to obtain the upsampled low-resolution branch feature , and input the upsampled low-resolution branch feature sequentially into the 1×1 convolutional layer and the 3×3 convolutional layer to obtain the 3×3 convolutional low-resolution branch feature. Add the 3×3 convolutional low-resolution branch feature and the upsampled low-resolution branch feature element-wise to obtain the low-resolution branch feature of the lower branch , denoted as;
[0141] ;
[0142] ;
[0143] The output feature information of the lightweight guidance fusion layer First, input the global average pooling layer in the middle branch to obtain the output feature information of the global average pooling layer lightweight guiding fusion layer , and input the output feature information of the global average pooling layer lightweight guiding fusion layer into a 3×3 convolutional layer to obtain the output feature information of the 3×3 convolutional layer lightweight guiding fusion layer. Multiply the output feature information of the 3×3 convolutional layer lightweight guiding fusion layer, the low-resolution branch features of the 3×3 convolutional layer, and the high-resolution branch features of the 3×3 convolutional layer to obtain the output feature information of the multiplication lightweight guiding fusion layer. Then, integrate the multiplied information of the output feature information of the multiplication lightweight guiding fusion layer through another 3×3 convolutional layer to obtain the output feature information of the middle branch lightweight guiding fusion layer , denoted as;
[0144] ;
[0145] ;
[0146] In the formula, represents the global average pooling layer;
[0147] Input the high-resolution branch features of the upper branch , the low-resolution branch features of the lower branch , and the output feature information of the middle branch lightweight guiding fusion layer into the concatenation operation Concat to obtain the concatenated feature map . Then, input the concatenated feature map sequentially into a 3×3 convolutional layer and a 1×1 convolutional layer to obtain the output feature map of the 1×1 convolutional layer concatenated feature map, that is, the output feature map of the guiding recovery module , denoted as:
[0148] ;
[0149] ). Input the output feature map of the guiding recovery module into the segmentation head of the segmentation module for processing to obtain the output F of the drivable area detection feature information, denoted as:
[0150] .
[0151] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A dual-branch lightweight drivable area detection method based on low-level features, characterized in that: The following steps are involved: Step S1: constructing dataset images; Step S2: construct a dual-branch lightweight model, which consists of a high-resolution branch module, a low-resolution branch module, a detail module, a segmentation module, and a lightweight guided fusion module; the high-resolution branch module consists of a high-resolution feature extraction network and an adaptive enhancement module; the low-resolution branch module consists of a low-resolution feature extraction network and a cross convolution module; the lightweight guided fusion module includes a guided recovery module; Step S3: input the dataset image into the high-resolution feature extraction network for processing to obtain feature information, and input the feature information into the adaptive enhancement module for processing to obtain high-resolution branch features; Step S4: inputting the dataset image into the low-resolution feature extraction network for processing to obtain the second-layer extracted information, and inputting the second-layer extracted information into the cross convolution module for processing to obtain low-resolution branch features; Step S5: input the high-resolution branch features into the detail module for processing to obtain a detail feature map; Step S6: input the low-resolution branch features into the segmentation module for processing to obtain a segmentation feature map; Step S7: Optimize the lightweight guided fusion module through the detail feature map and the segmentation feature map, input the high-resolution branch features and the low-resolution branch features into the optimized lightweight guided fusion module for processing, obtain the output feature information of the lightweight guided fusion layer, input the high-resolution branch features, the low-resolution branch features and the output feature information of the lightweight guided fusion layer into the guided recovery module for processing, and obtain the output of the drivable area detection feature information.
2. The dual-branch lightweight drivable area detection method based on low-level features according to claim 1 is characterized in that: The specific process of the dataset image in step S1 is as follows: A road segmentation data set is obtained, and the road segmentation data set includes training images and test images; the training images and test images are processed to obtain data set images.
3. The dual-branch lightweight drivable area detection method based on low-level features according to claim 2 is characterized in that: The high-resolution feature extraction network is a ResNet18 module, which includes an L1 layer. The L1 layer consists of a high-resolution feature residual block, which includes two high-resolution feature residual units, and the residual unit includes two 3×3 convolutional layers. The dataset image is input into the high-resolution feature extraction network for processing to obtain feature information F1, which is expressed as: ; In the formula, Represents the dataset image; represents batch normalization; represents a 3×3 convolutional layer; Represents the activation function.
4. The dual-branch lightweight drivable area detection method based on low-level features according to claim 3 is characterized in that: In step S3, high-resolution branch features are obtained. The specific process is as follows: The feature information F1 is input into the adaptive enhancement module for processing to obtain high-resolution branch features; The adaptive enhancement module consists of a structure of four dilated convolution branches; the first dilated convolution branch consists of a 1×1 standard convolution; the second dilated convolution branch consists of a 1×1 standard convolution, a 3×1 standard convolution and a 3×3 dilated convolution; the third dilated convolution branch consists of a 1×1 standard convolution, a 1×3 standard convolution and a 3×3 dilated convolution; the fourth dilated convolution branch consists of a 1×1 standard convolution and a 3×3 standard convolution; The feature information F1 only undergoes a 1×1 standard convolution operation in the first dilated convolution branch to obtain the output of the first dilated convolution branch; after the feature information F1 undergoes a 1×1 standard convolution in the second dilated convolution branch, a 3×1 standard convolution and a 3×3 dilated convolution operation are performed to obtain the output of the second dilated convolution branch; after the feature information F1 undergoes a 1×1 standard convolution in the third dilated convolution branch, a 1×3 standard convolution and a 3×3 dilated convolution operation are performed in sequence to obtain the output of the third dilated convolution branch; the feature information F1 undergoes a 1×1 standard convolution in the fourth dilated convolution branch, and then a 3×3 standard convolution operation is performed to obtain the output of the fourth dilated convolution branch; the output of the second dilated convolution branch and the output of the third dilated convolution branch are added to obtain the added output, and the added output is cascaded with the output of the first dilated convolution branch and the output of the fourth dilated convolution branch to obtain a high-resolution branch feature. .
5. The dual-branch lightweight drivable area detection method based on low-level features according to claim 4 is characterized in that: The low-resolution feature extraction network is a ResNet18 module, which includes an L1 layer and an L2 layer. The ResNet18 module consists of a low-resolution feature residual block, which includes two low-resolution feature residual units, and the residual unit includes two 3×3 convolutional layers. The dataset image is input into the low-resolution feature extraction network for processing to obtain the second-layer extracted information F2, which is expressed as: ; ; In the formula, Represents the output of the first low-resolution feature residual unit in the low-resolution feature extraction network.
6. The dual-branch lightweight drivable area detection method based on low-level features according to claim 5 is characterized in that: The low-resolution branch features in step S4 are specifically: The extracted information F2 of the second layer is input into the cross convolution module for processing to obtain low-resolution branch features; The cross convolution module consists of a 1×5 depthwise row convolution, a 5×1 depthwise column convolution, and a 1×1 convolution layer, a batch normalization, and an activation function; The extracted information F2 of the second layer is first subjected to a 1×5 depth row convolution operation in the cross convolution module to capture the context information of the extracted information F2 of the second layer in the horizontal direction, and then a 5×1 depth column convolution operation is performed to extract the context information of the extracted information F2 of the second layer in the vertical direction. A 1×1 convolution layer is used to fuse the context information of the extracted information F2 of the second layer in the horizontal direction and the context information of the extracted information F2 of the second layer in the vertical direction, and a nonlinear transformation is completed through batch normalization and activation function to obtain the output of the nonlinear transformation. The extracted information F2 of the second layer and the output of the nonlinear transformation are added to obtain the low-resolution branch features. , expressed as: ; In the formula, represents a 1×1 convolutional layer; Represents 1×5 depthwise row convolution; Represents a 5×1 depthwise column convolution.
7. The dual-branch lightweight drivable area detection method based on low-level features according to claim 6 is characterized in that: The detailed feature map in step S5, the specific process is: High-resolution branch features The input is processed in the detail module, which includes a detail head; the detail head consists of a 3×3 convolution layer and a 1×1 convolution layer; high-resolution branch features In the detail head, it first passes through a 3×3 convolution layer and a 1×1 convolution layer, and then operates through batch normalization and activation functions to generate a detail feature map. , expressed as: ; In the formula, Represents an upsampling operation.
8. The dual-branch lightweight drivable area detection method based on low-level features according to claim 7 is characterized in that: In step S6, the feature map is segmented, and the specific process is as follows: The low-resolution branch features In the input segmentation module, the segmentation module includes a segmentation head, which consists of a 3×3 convolutional layer and a 1×1 convolutional layer. The low-resolution branch feature In the segmentation head, the 3×3 convolution layer and the 1×1 convolution layer are first passed, and then the segmentation feature map is generated through batch normalization and activation function operations. , expressed as: 。 9. The dual-branch lightweight drivable area detection method based on low-level features according to claim 8, characterized in that: The lightweight guided fusion module consists of a high-resolution branch, a low-resolution branch, and a 3×3 convolutional layer; The high-resolution branch consists of the first branch and the second branch; the low-resolution branch consists of the third branch and the fourth branch; the first branch consists of a 3×3 depth-separable convolution and a 1×1 convolution layer; the second branch consists of a 3×3 convolution layer and a downsampling; the third branch consists of a 3×3 convolution layer and an upsampling; the fourth branch consists of a 3×3 depth-separable convolution and a 1×1 convolution layer; Based on detail feature map and segmentation feature map Optimizing the lightweight guidance fusion module to obtain an optimized lightweight guidance fusion module; High-resolution branch features and low-resolution branch features Input into the optimized lightweight guidance fusion module; High-resolution branch features The 3×3 depth-wise separable convolution of the first branch and the 3×3 convolution layer of the second branch are respectively input into the optimized lightweight guided fusion module to obtain the output of the 3×3 depth-wise separable convolution of the first branch and the output of the 3×3 convolution layer of the second branch respectively; Low-resolution branch features The 3×3 convolution layer of the third branch and the 3×3 depth-separable convolution of the fourth branch are respectively input into the optimized lightweight guided fusion module to obtain the output of the 3×3 convolution layer of the third branch and the output of the 3×3 depth-separable convolution of the fourth branch respectively; Input the output of the 3×3 depth-separable convolution of the first branch to the 1×1 convolution layer of the first branch; input the output of the 3×3 depth-separable convolution of the fourth branch to the 1×1 convolution layer of the fourth branch, and obtain the output of the 1×1 convolution layer of the first branch and the output of the 1×1 convolution layer of the fourth branch respectively; Downsample the output of the second branch 3×3 convolutional layer to obtain a downsampled feature map, and upsample the output of the third branch 3×3 convolutional layer to obtain an upsampled feature map; The downsampled feature map is multiplied by the output of the fourth branch 1×1 convolution layer to obtain the low-resolution branch weight feature. The upsampled feature map is multiplied by the output of the first branch 1×1 convolution layer to obtain the high-resolution branch weight feature. The high-resolution branch weight feature and the low-resolution branch weight feature are added element by element to obtain the fusion feature. The fusion feature is then integrated through a 3×3 convolution layer to generate the output feature information of the lightweight guided fusion layer. .
10. The dual-branch lightweight drivable area detection method based on low-level features according to claim 9, characterized in that: In step S7, the drivable area detection feature information is output, and the specific process is as follows: The output feature information of the lightweight guide fusion layer , high-resolution branch features and low-resolution branch features The input is processed in the guidance recovery module to obtain the output of the drivable area detection feature information; The guided recovery module consists of an upper branch, a middle branch, and a lower branch; the upper branch consists of an upsampling layer, a 1×1 convolution layer, and a 3×3 convolution layer; The middle branch consists of a global average pooling layer, three 3×3 convolutional layers, and a 1×1 convolutional layer; The lower branch consists of an upsampling layer, a 1×1 convolution layer, and a 3×3 convolution layer; High-resolution branch features In the upper branch, the upsampled data is first input to obtain the upsampled high-resolution branch features. , upsample the high-resolution branch features Input the 1×1 convolution layer and the 3×3 convolution layer in sequence to obtain the high-resolution branch features of the 3×3 convolution layer, and combine the high-resolution branch features of the 3×3 convolution layer with the upsampled high-resolution branch features. Element-by-element addition to obtain high-resolution branch features of the upper branch ; Low-resolution branch features In the lower branch, the upsampled data is first input to obtain the upsampled low-resolution branch features. , upsample the low-resolution branch features Input the 1×1 convolution layer and the 3×3 convolution layer in sequence to obtain the low-resolution branch features of the 3×3 convolution layer, and combine the low-resolution branch features of the 3×3 convolution layer with the upsampled low-resolution branch features. Element-by-element addition to obtain the low-resolution branch features of the lower branch ; Lightweight guide fusion layer output feature information In the middle branch, the global average pooling layer is first input to obtain the output feature information of the lightweight guided fusion layer of the global average pooling layer. , lightweight the global average pooling layer to guide the output feature information of the fusion layer Input the 3×3 convolutional layer to obtain the output feature information of the 3×3 convolutional layer lightweight guided fusion layer, multiply the output feature information of the 3×3 convolutional layer lightweight guided fusion layer, the 3×3 convolutional layer low-resolution branch features and the 3×3 convolutional layer high-resolution branch features to obtain the output feature information of the multiplied lightweight guided fusion layer, and then pass the multiplied information through a 3×3 convolutional layer to obtain the output feature information of the middle branch lightweight guided fusion layer. ; The high-resolution branch features of the upper branch , low-resolution branch features of the lower branch The output feature information of the light-weight guide fusion layer in the middle branch Perform cascade operation to obtain the spliced feature map , the concatenated feature map Input the 3×3 convolution layer and the 1×1 convolution layer in sequence to obtain the 1×1 convolution layer splicing feature map, which is the output feature map of the guided recovery module. ; The output feature map of the boot recovery module The input is processed into the segmentation head of the segmentation module to obtain the drivable area detection feature information output F, which is expressed as: 。
Citation Information
Patent Citations
High-resolution saliency target detection method based on graffiti supervision
CN114332490A
High-resolution image saliency target detection method based on deep learning
CN115294359A