Travelable area detection method based on progressive gating decoder

Through the progressive gated decoder combined with the visual field heterogeneity extractor and the cross-scale feature fusion module, the problems of low accuracy and low computational efficiency of driving area detection in weak feature scenarios are solved, and efficient and accurate complex scene detection is achieved.

CN120388342APending Publication Date: 2025-07-29CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510623202.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing feasible area detection technology has low detection accuracy and low computing efficiency in weak characteristic scenarios, making it difficult to generalize in complex scenarios, and the existing model has high computational cost, making it difficult to meet real-time requirements.

Method used

The progressive gated decoder is adopted, combining the visual field heterogeneity extractor and the cross-scale feature fusion module, through heterogeneity feature extraction and adaptive integration of multi-scale information, the efficient multi-scale convolution module is used to reduce the calculation amount, enhance the key target area and suppress background noise.

Benefits of technology

High-precision driving area detection is realized in complex scenarios, reducing calculation complexity, improving detection efficiency, and meeting real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005403107280000026
    Figure BDA0005403107280000026
  • Figure BDA0005403107280000101
    Figure BDA0005403107280000101
  • Figure FDA0005403107270000011
    Figure FDA0005403107270000011
Patent Text Reader

Abstract

The invention discloses a drivable area detection method based on a progressive gating decoder, and mainly solves the problems of low drivable area detection precision, complex network structure and low calculation efficiency in a weak feature scene in the existing drivable area detection technology. The implementation scheme is as follows: 1) acquiring a data set and a detection label; 2) constructing a drivable area detection model; 3) constructing a loss function; 4) training a drivable area detection model; and 5) obtaining a drivable area detection result. According to the driving area detection model constructed by the invention, multi-dimensional modeling of a target is realized by combining a visual field heterogeneity extractor and extracting heterogeneity features, self-adaptive integration of feature information of different scales is realized through a cross-scale feature fusion module, and self-adaptive integration of feature information of different scales is realized through a progressive gating decoder. And by utilizing the efficient multi-scale convolution module design, the calculation amount is reduced, and meanwhile, an excellent segmentation effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a drivable area detection method based on a progressive gated decoder. Background Art

[0002] The drivable area detection technology identifies the drivable road areas in images through vision algorithms, providing visual information support for path planning in autonomous driving, thereby ensuring the safe driving of vehicles. However, there are still challenges in achieving accurate drivable area detection in weak feature scenarios such as at night, occlusion, and mine roads. For example, at night, due to insufficient light, the image quality deteriorates, and the visual features of the drivable area are often not obvious enough, resulting in the network's difficulty in extracting effective features; in the case of road occlusion, the lack of context information makes the model unable to infer the correct drivable area; in special environments such as mines and mine roads, the complex terrain, the lack of obvious texture features, and the ambiguity of boundary information further exacerbate the detection difficulty. The root cause of these problems is that existing general semantic segmentation models are mostly designed for standard road scenes and are not optimized for the characteristics of complex scenes, restricting the generalization ability of the model in diverse scenes. At the same time, currently excellent drivable area detection algorithms usually rely on deep network structures and complex multi-scale feature fusion strategies, significantly increasing the computational cost and resource consumption and making it difficult to meet the real-time requirements. Therefore, while improving the model performance, how to reduce the network complexity and optimize the computational efficiency has become another challenge to be solved. Summary of the Invention

[0003] The present invention fully considers the disadvantages of existing methods, and its purpose is to provide a drivable area detection method based on a progressive gated decoder, which realizes multi-dimensional feature extraction of targets in complex environments through a visual field heterogeneity extractor; through a cross-scale feature fusion module, it realizes the dynamic fusion of features at different scales and improves the network's perception ability for targets at different scales; through a progressive gated decoder, it uses an upsampling convolution module to retain richer fine-grained information, and uses a saliency enhancement module to enhance key target areas and suppress background noise.

[0004] I. Technical Principle

[0005] Most of the current drivable area detection algorithms use four parts: a backbone network, a neck network, a decoder, and a detection head to complete the detection of the drivable area in the image. The drivable area detection algorithm extracts features from the input image through the backbone network, and the neck network realizes the extraction and enhancement of multi-scale features. The decoder mainly completes the tasks of upsampling and feature fusion. Finally, the detection head classifies and makes pixel-level predictions on the feature map output by the decoder to generate a segmentation map of the drivable area. In the application of autonomous driving technology, detecting the drivable area in weak feature scenarios and enhancing the computational efficiency of the algorithm are currently two major difficulties. Existing general semantic segmentation models are mostly designed for standard road scenarios and are not optimized for the characteristics of complex scenarios, which limits the generalization ability of the models in diverse scenarios. To achieve accurate detection of the drivable area in weak feature scenarios, the present invention proposes a visual field heterogeneity extractor and a cross-scale feature fusion module. The former realizes multi-dimensional modeling of the target through heterogeneous feature extraction, and the latter adaptively integrates feature information of different scales through an internal feature fuser, so that multi-scale information can be efficiently transmitted throughout the network. Aiming at the problem of low computational efficiency of the current drivable area detection algorithm, the present invention proposes a progressive gated decoder. This decoder is designed based on an efficient multi-scale convolution module and can still achieve excellent segmentation results without relying on traditional complex computational modules, thus effectively reducing the overall computational amount of the network.

[0006] II. According to the above principle, the present invention is implemented through the following solutions:

[0007] A drivable area detection method based on a progressive gated decoder, comprising the following steps:

[0008] (1) Obtain a dataset and detection labels:

[0009] Obtain a drivable area dataset and corresponding detection labels;

[0010] (2) Construct a drivable area detection model: This model consists of a backbone network, a neck network, a decoder, and a detection head. The specific construction process includes the following steps:

[0011] (2-a) Construct a backbone network: Use ResNet-101 as the backbone network. The input image is processed by the backbone network to obtain four feature maps f1, f2, f3, and f4, where the scales of f3 and f4 are the same, and the scales of the other feature maps decrease in turn;

[0012] (2-b) Construct the neck network: This network consists of four Visual Field Heterogeneity Extractors (VFHE) and a Cross-Scale Feature Fusion Module (CSFM); the feature maps f1, f2, f3, and f4 obtained in step (2-a) are respectively used as the inputs of the four Visual Field Heterogeneity Extractors VFHE1, VFHE2, VFHE3, and VFHE4 to obtain the processing results and Take and as the inputs of the Cross-Scale Feature Fusion Module (CSFM) to obtain the output results and

[0013] The Visual Field Heterogeneity Extractor (VFHE) and the Cross-Scale Feature Fusion Module (CSFM) in this step are constructed as follows:

[0014] The Visual Field Heterogeneity Extractor (VFHE) consists of a feature grouping module, a feature concatenation module, a convolutional module Conv, four branch modules, and a residual path; the Visual Field Heterogeneity Extractor first inputs the input feature I in parallel to the residual path and the feature grouping module for processing, respectively obtaining the residual path output O and four groups of feature representations G1, G2, G3, and G4; the feature grouping module evenly divides the input feature in the channel dimension to obtain four groups of features. When the number of channels of the input feature cannot be divided evenly by four, then N channels of features are discarded so that the number of channels of the remaining features is an integer multiple of four, where N is a positive integer and N ∈ [1, 3]; the first group of features G1 is processed by branch module 1 to obtain the processing result R1, the second group of features G2 is processed by branch module 2 to obtain the processing result R2, the third group of features G3 is processed by branch module 3 to obtain the processing result R3, and the fourth group of features G4 is processed by branch module 4 to obtain the processing result R4; among them, branch modules 1 and 2 consist of four convolutional layers, branch module 3 consists of three convolutional layers, and branch module 4 consists of a max pooling layer and a convolutional layer; the four groups of processing results R1, R2, R3, and R4 are first processed by the feature concatenation module, concatenated in the channel dimension to obtain the output feature, then processed by the convolutional module Conv, and finally feature-added to the residual path output O to obtain the output feature; the convolutional module Conv consists of a convolutional layer and a Relu activation function;

[0015] The cross-scale feature fusion module CSFM consists of three cross-scale fusers CSF1, CSF2, and CSF3; to ensure that the three inputs of the cross-scale feature fusion module have the same size scale before being fed into each cross-scale fuser, the cross-scale feature fusion module will perform horizontal connection, upsampling, or downsampling processing on the three inputs in advance; each cross-scale fuser CSF has three inputs I1, I2, and I3, and the three inputs are respectively processed through a convolution module Conv, feature concatenation, a convolution module Conv, and SoftMax in the cross-scale fuser to obtain three feature weights weight1, weight2, and weight3. The feature weights are multiplied element-wise with the inputs to obtain the processing results O1, O2, and O3, and the three processing results are subjected to feature addition and a convolution module Conv processing to obtain the output result of the cross-scale fuser; the convolution module Conv in this step has the same structure, and each convolution module consists of a convolution layer and a Relu activation function;

[0016] (2-c) Build a progressive gating decoder, which consists of three upsampling convolution modules UCB and three significance enhancement modules SEM; input the processing result obtained in step (2-b) and into the decoder for processing; input into UCB3 to obtain the processing result Input and into SEM3 to obtain the processing result Input and perform element-wise addition to obtain the fusion result Input into UCB2 to obtain the processing result Input and into SEM2 to obtain the processing result Input and perform element-wise addition to obtain the fusion result Input into UCB1 to obtain the processing result Input and into SEM1 to obtain the processing result Input and perform element-wise addition to obtain the output result of the decoder

[0017] The upsampling convolution module UCB and the significance enhancement module SEM in this step are constructed as follows:

[0018] The upsampling convolution module UCB has one input. After the input is processed by the upsampling module, the convolution module Conv, batch normalization, the Relu activation function, and the convolution module Conv, the output result of the upsampling convolution module is obtained. The convolution modules Conv in this step have the same structure, and each convolution module consists of a convolutional layer and a Relu activation function.

[0019] The saliency enhancement module SEM has two inputs I1 and I2. The two inputs are respectively processed by the convolution module Conv and batch normalization and then multiplied element-wise to obtain the result F1. After F1 is processed by Relu, the convolution module Conv, batch normalization, and the Sigmoid activation function, the output result of the saliency enhancement module is obtained. The convolution modules Conv in this step have the same structure, and each convolution module consists of a convolutional layer and a Relu activation function.

[0020] (2-d) Construct the detection head: The detection head consists of a detection module Head1 and an upsampling module. The obtained in step (2-c) is respectively processed by the detection module Head1 and the upsampling module to obtain the final detection result.

[0021] The detection module Head is constructed as follows in this step:

[0022] The detection module Head has one input. After the input is processed by the convolution module Conv, batch normalization, Dropout, and the convolution module Conv, the output result of the detection module Head is obtained. The convolution modules Conv in this step have the same structure, and each convolution module consists of a convolutional layer and a Relu activation function.

[0023] (3) Construct the loss function:

[0024] Construct the following loss function L:

[0025] L = L BCE (DAD GT , DAD Pre )

[0026] where DAD GT represents the label of the drivable area detection, DAD Pre represents the drivable detection result, and L BCE represents the binary cross-entropy loss function.

[0027] (4) Train the drivable area detection model:

[0028] Train the drivable area detection model constructed in step (2) using the dataset obtained in step (1); calculate the error between the predicted result output by the model and the label using the loss function L constructed in step (3); during the training process, use the Adam algorithm to update the model parameters and use L-2 regularization as a constraint until the loss no longer decreases to obtain the trained drivable area detection model;

[0029] (5) Drivable area detection:

[0030] After normalizing the test image, input it into the trained drivable area detection model, and the output result of the model is the final drivable area detection result.

[0031] In steps (2-c) and (2-d), the upsampling module preferably uses linear interpolation upsampling;

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] (1) The present invention constructs a visual field heterogeneity extractor, which mimics the attention and processing methods for different spatial regions in the human visual system through a multi-branch structure, and realizes the extraction of heterogeneous spatial features of the input image.

[0034] (2) The present invention constructs a cross-scale feature fusion module, which uses its feature fusion device with adaptive integration ability to integrate information of different scales, realizes the efficient transmission of multi-scale features in the entire network information flow, and enhances the interaction between the network fine-grained and global information.

[0035] (3) The present invention constructs a progressive gating decoder, which consists of an upsampling convolution module and a saliency enhancement module, enhances the feature detail information after upsampling through convolution, and uses a simple gating mechanism to achieve efficient decoding with less computational effort. Description of the Drawings

[0036] Figure 1 Flowchart of the drivable area detection method based on the progressive gating decoder according to the embodiment of the present invention;

[0037] Figure 2 Structure diagram of the drivable area detection model according to the embodiment of the present invention;

[0038] Figure 3 Structure diagram of the visual field heterogeneity extractor VFHE according to the embodiment of the present invention;

[0039] Figure 4 Structure diagram of the cross-scale feature fusion module CSFM according to the embodiment of the present invention;

[0040] Figure 5Structure diagram of the upsampling convolutional module UCB according to an embodiment of the present invention;

[0041] Figure 6 Structure diagram of the saliency enhancement module SEM according to an embodiment of the present invention;

[0042] Figure 7 Structure diagram of the detection module Head according to an embodiment of the present invention;

[0043] Figure 8 Comparison chart of the detection results of the drivable area detection method according to an embodiment of the present invention and the detection results of other methods. Detailed implementation manners

[0044] The following describes the detailed implementation manners of the present invention:

[0045] Example 1

[0046] Figure 1 The following is a flowchart of the drivable area detection method based on a progressive gated decoder according to an embodiment of the present invention. The specific steps are as follows:

[0047] Step 1, obtain a data set and detection labels:

[0048] Obtain a drivable area data set and corresponding detection labels;

[0049] Step 2, construct a drivable area detection model: This model consists of a backbone network, a neck network, a decoder, and a detection head. The specific construction process includes the following steps.

[0050] Figure 2 The following is a structure diagram of the drivable area detection model constructed according to an embodiment of the present invention. The specific steps are as follows:

[0051] (2-a) Construct a backbone network: Use ResNet-101 as the backbone network. An input image with a size of 3×720×1280 is processed by the backbone network to obtain feature maps f1, f2, f3, and f4 with scales of 64×180×320, 128×90×160, 256×45×80, and 512×45×80 respectively;

[0052] (2-b) Construct a neck network: This network consists of four visual field heterogeneity extractors VFHE and a cross-scale feature fusion module CSFM; The feature maps f1, f2, f3, and f4 obtained in step (2-a) are used as the inputs of four visual field heterogeneity extractors VFHE1, VFHE2, VFHE3, and VFHE4 respectively to obtain processing results with scales of 64×180×320, 64×90×160, 64×45×80, and 64×45×80 and will and As the input of the cross-scale feature fusion module CSFM, the output results with scales of 64×180×320, 64×90×160, and 64×45×80 are obtained and

[0053] Figure 3 The structure diagram of the visual field heterogeneity extractor VFHE according to the embodiment of the present invention is shown as follows. The specific construction is as follows:

[0054] The visual field heterogeneity extractor VFHE is composed of a feature grouping module, a feature splicing module, a convolutional module Conv, four branch modules, and a residual path. The visual field heterogeneity extractor first inputs the input feature I with a scale of C×H×W into the residual path and the feature grouping module in parallel for processing, and respectively obtains the output O of the residual path with a scale of 64×H×W and the four groups of feature representations G1, G2, G3, and G4 with scales of and The feature grouping module evenly divides the input feature in the channel dimension to obtain four groups of features. When the number of channels of the input feature cannot be divided by four, the features of N channels are discarded, so that the number of channels of the remaining features is an integer multiple of four, where N is a positive integer and N∈[1,3]. The first group of features G1 is processed by branch module 1 to obtain the processing result R1 with a scale of The second group of features G2 is processed by branch module 2 to obtain the processing result R2 with a scale of The third group of features G3 is processed by branch module 3 to obtain the processing result R3 with a scale of The fourth group of features G4 is processed by branch module 4 to obtain the processing result with a scale of The processing result R4; among the four branch modules, branch module 1 consists of four convolutional layers with kernel sizes of 1×1, 5×1, 1×5, and 5×5, and dilation rates of 0×0, 0×0, 0×0, and 3×3 respectively. Branch module 2 consists of four convolutional layers with kernel sizes of 1×1, 3×1, 1×3, and 3×3, and dilation rates of 0×0, 0×0, 0×0, and 3×3 respectively. Branch module 3 consists of three convolutional layers with kernel sizes of 1×1, 3×3, 3×3, and dilation rates of 0×0, 0×0, 0×0 respectively. Branch module 3 consists of three convolutional layers with kernel sizes of 1×1, 3×3, 3×3, and dilation rates of 0×0, 0×0, 0×0 respectively. Branch module 4 consists of a max pooling layer and a convolutional layer with a pooling kernel size of 3×3 and a convolutional kernel size of 1×1, and a dilation rate of 0×0. The four sets of processing results R1, R2, R3, and R4 are first processed by the feature concatenation module, concatenated in the channel dimension to obtain the output feature with a scale of C×H×W, then processed by the convolutional module Conv, and finally added to the output O of the residual path in the feature dimension to obtain the output feature with a scale of 64×H×W. The convolutional module Conv consists of a convolutional layer with a kernel size of 3×3 and a Relu activation function.

[0055] Figure 4 The following is the structural diagram of the cross-scale feature fusion module CSFM according to the embodiment of the present invention, which is specifically constructed as follows:

[0056] The cross-scale feature fusion module CSFM is composed of three cross-scale fusers CSF1, CSF2, and CSF3. To ensure that the three inputs of the cross-scale feature fusion module have the same scale before being fed into each cross-scale fuser, the cross-scale feature fusion module will perform horizontal connection, upsampling, or downsampling on the three inputs in advance. Each cross-scale fuser CSF has three inputs I1, I2, and I3 with a scale of 64×H×W. The three inputs are respectively processed by the convolutional module Conv, feature concatenation, convolutional module Conv, and SoftMax in the cross-scale fuser to obtain three feature weights weight1, weight2, and weight3 with a scale of H×W. The convolutional module Conv consists of a convolutional layer with a kernel size of 3×3 and a Relu activation function. The feature weights are multiplied element by element with the corresponding inputs to obtain three processing results O1, O2, and O3 with a scale of 64×H×W. The three processing results are added in the feature dimension and processed by the convolutional module Conv to obtain the output result of the cross-scale fuser with a scale of 64×H×W. The convolutional module Conv consists of a convolutional layer with a kernel size of 1×1 and a Relu activation function.

[0057] (2-c) Build a progressive gated decoder, which consists of three upsampling convolutional modules UCB and three saliency enhancement modules SEM; input the processing result obtained in step (2-b) and into the decoder for processing; input into UCB3 to obtain a processing result with a scale of 64×45×80 Input and into SEM3 to obtain a processing result with a scale of 64×45×80 Input and perform element-wise addition to obtain a fusion result with a scale of 64×45×80 Input into UCB2 to obtain a processing result with a scale of 64×90×160 Input and into SEM2 to obtain a processing result with a scale of 64×90×160 Input and perform element-wise addition to obtain a fusion result with a scale of 64×90×160 Input into UCB1 to obtain a processing result with a scale of 64×180×320 Input and into SEM1 to obtain a processing result with a scale of 64×180×320 Input and perform element-wise addition to obtain the output result of the decoder with a scale of 64×180×320

[0058] Figure 5 The following shows the structural diagram of the upsampling convolutional module UCB according to the embodiment of the present invention, and the specific construction is as follows:

[0059] The upsampling convolutional module UCB has an input with a scale of 64×H×W. After being processed by linear interpolation upsampling, convolutional module Conv, batch normalization, Relu activation function, and convolutional module Conv, the output result of the upsampling convolutional module with a scale of is obtained. In this step, each convolutional module Conv consists of a convolutional layer with a kernel size of 3×3 and a Relu activation function. is the upsampling ratio;

[0060] Figure 6The following is the structural diagram of the saliency enhancement module SEM according to the embodiments of the present invention, and the specific construction is as follows:

[0061] The saliency enhancement module SEM has two inputs I1 and I2 with a scale of 64×H×W. After the two inputs are respectively processed by the convolutional module Conv and batch normalization, element-wise multiplication is performed to obtain a processing result F1 with a scale of 64×H×W. After F1 is processed by Relu, the convolutional module Conv, batch normalization, and the Sigmoid activation function, the output result of the saliency enhancement module with a scale of 64×H×W is obtained; wherein the convolutional module Conv in this step is composed of a convolutional layer with a kernel size of 3×3 and a Relu activation function;

[0062] (2-d) Construct the detection head: The detection head is composed of a detection module Head1 and an upsampling module; the obtained in step (2-c) is respectively processed by the detection module Head1 and the upsampling module to obtain a final detection result with a scale of 2×720×1280; wherein the upsampling module uses linear interpolation upsampling with an upsampling ratio of 4;

[0063] Figure 7 The following is the structural diagram of the detection module Head according to the embodiments of the present invention, and the specific construction is as follows:

[0064] The detection module Head has an input with a scale of 64×H×W. After the input is processed by the convolutional module Conv, batch normalization, Dropout, and the convolutional module Conv, the output result of the detection module Head with a scale of 2×H×W is obtained;

[0065] Step 3, construct the loss function:

[0066] Construct the following loss function L:

[0067] L = L BCE (DAD GT , DAD Pre )

[0068] where DAD GT represents the label of the drivable area detection, DAD Pre represents the drivable detection result, and L BCE represents the binary cross-entropy loss function;

[0069] Step 4, train the drivable area detection model:

[0070] Train the drivable area detection model constructed in step (2) using the dataset obtained in step (1); calculate the error between the predicted result output by the model and the label using the loss function L constructed in step (3); update the model parameters using the Adam algorithm during the training process and use L-2 regularization as a constraint until the loss no longer decreases to obtain the trained drivable area detection model;

[0071] Step 5, drivable area detection:

[0072] After normalizing the test image, input it into the trained drivable area detection model, and the output result of the model is the final drivable area detection result.

[0073] Example 2

[0074] Use the method in Example 1 to conduct an experiment on drivable area detection for the underground mine tunnel dataset MineDrive. The operating system of this experiment is Ubuntu 18.04, based on the PyTorch 1.2.0 framework of CUDA 10.0 and cuDNN 7.6.0, and a personal computer equipped with an Intel Core i5-12600KF CPU (3.70GHz) and an NVIDIA GeForce RTX 3090 (24GB) hardware is used for training and testing.

[0075] In this example, two metrics, mIoU and mPA, are used to conduct an experimental comparison of three testing methods, PSPNet, Deeplabv3+, and HRNet, with the method of the present invention on the MineDrive dataset.

[0076] The comparison results are shown in Table 1. It can be found that compared with other methods, the present invention can obtain accurate detection results on the MineDrive dataset and reaches the optimal level in the evaluation metrics of mIoU and mPA.

[0077] Figure 8 The figure shows a comparison diagram of the drivable area detection results of the embodiment of the present invention and the detection results of other methods. The results show that the model designed by the present invention can well handle various challenging scenarios, including scenarios where the road boundary and the background are blurred ( Figure 8 lines 1 and 2), and scenarios where the road is dimly lit ( Figure 8 lines 3 and 4). Compared with other methods, the drivable area detection results obtained by this method are more accurate.

[0078] The above-described embodiments are only the preferred embodiments of the present invention and do not limit the scope of implementation of the present invention. Therefore, any changes made according to the structure and principle of the present invention should be covered within the protection scope of the present invention.

[0079] Table 1

[0080]

Claims

1. A drivable area detection method based on a progressive gating decoder, characterized in that It includes the following steps: (1) Obtain the dataset and detection labels: Obtain the drivable area dataset and the corresponding detection labels; (2) Construct a drivable area detection model: This model consists of a backbone network, a neck network, a decoder, and a detection head. The specific construction process includes the following steps: (2-a) Construct the backbone network: Use ResNet-101 as the backbone network. The input image is processed by the backbone network to obtain four feature maps f1, f2, f3, and f4, where the scales of f3 and f4 are the same, and the scales of the other feature maps decrease sequentially; (2-b) Construct the neck network: This network consists of four Visual Field Heterogeneity Extractors (VFHE) and one Cross-Scale Feature Fusion Module (CSFM); the feature maps f1, f2, f3, and f4 obtained in step (2-a) are respectively used as the inputs of the four Visual Field Heterogeneity Extractors VFHE1, VFHE2, VFHE3, and VFHE4 to obtain the processing results and The and are used as the inputs of the Cross-Scale Feature Fusion Module (CSFM) to obtain the output results and The visual field heterogeneity extractor VFHE and the cross-scale feature fusion module CSFM in this step are constructed as follows: The visual field heterogeneity extractor VFHE consists of a feature grouping module, a feature splicing module, a convolutional module Conv, four branch modules, and a residual path; The visual field heterogeneity extractor first inputs the input feature I into the residual path and the feature grouping module in parallel for processing, and respectively obtains the residual path output O and four groups of feature representations G1, G2, G3, and G4; The first group of features G1 is processed by branch module 1 to obtain the processing result R1, the second group of features G2 is processed by branch module 2 to obtain the processing result R2, the third group of features G3 is processed by branch module 3 to obtain the processing result R3, and the fourth group of features G4 is processed by branch module 4 to obtain the processing result R4; The four groups of processing results R1, R2, R3, and R4 are first processed by the feature splicing module, then processed by the convolutional module Conv, and finally added to the residual path output O for feature addition to obtain the output feature; The cross-scale feature fusion module CSFM consists of three cross-scale fusers CSF1, CSF2, and CSF3; To ensure that the three inputs of the cross-scale feature fusion module have the same size scale before being fed into each cross-scale fuser, the cross-scale feature fusion module will perform horizontal connection, upsampling, or downsampling processing on the three inputs in advance; Each cross-scale fuser CSF has three inputs I1, I2, and I3. The three inputs are respectively processed by a convolutional module Conv, feature splicing, a convolutional module Conv, and SoftMax in the cross-scale fuser to obtain three feature weights weight1, weight2, and weight3. The feature weights are multiplied element by element with the inputs to obtain the processing results O1, O2, and O3, and the three processing results are added for feature addition and processed by a convolutional module Conv to obtain the output result of the cross-scale fuser; (2-c) Construct a progressive gated decoder, which consists of three upsampling convolution modules UCB and three saliency enhancement modules SEM; the processing results obtained in step (2-b) are converted into and Input to the decoder for processing; Enter UCB3 to get the processing result Will and Send to SEM3 to get the processing results Will and Perform element-by-element addition to obtain the fusion result Will Enter UCB2 to get the processing result Will and Send to SEM2 to get the processing results Will and Perform element-by-element addition to obtain the fusion result Will Enter UCB1 to get the processing result Will and Send to SEM1 to get the processing results Will and Perform element-by-element addition to get the output of the decoder The upsampling convolutional module UCB and the saliency enhancement module SEM in this step are constructed as follows: The upsampling convolutional module UCB has one input. After the input is processed by an upsampling module, a convolutional module Conv, batch normalization, a Relu activation function, and a convolutional module Conv, the output result of the upsampling convolutional module is obtained; The saliency enhancement module (SEM) has two inputs, I1 and I2. After the two inputs are processed by the convolution module Conv and batch normalization, they are element-wise multiplied to obtain the result F1. After F1 is processed by Relu, convolution module Conv, batch normalization and Sigmoid activation function, the output result of the saliency enhancement module is obtained. (2-d) Build the detection head: The detection head consists of a detection module Head1 and an upsampling module; After the obtained in step (2-c) is processed by the detection module Head1 and the upsampling module respectively, the final detection result is obtained; After being processed by the detection module Head1 and the upsampling module respectively, the final detection result is obtained; The detection module Head in this step is constructed as follows: The detection module Head has an input, which is processed by the convolution module Conv, batch normalization, Dropout, and convolution module Conv to obtain the output result of the detection module Head; (3) Construct loss function: Construct the following loss function L: L = L BCE (DAD GT , DAD Pre ) Among them DAD GT Indicates the label of drivable area detection, DAD Pre Indicates the drivable test result, L BCE represents the binary cross entropy loss function; (4) Training the drivable area detection model: The drivable area detection model constructed in step (2) is trained using the data set obtained in step (1); the error between the prediction result output by the model and the label is calculated using the loss function L constructed in step (3); during the training process, the model parameters are updated using the Adam algorithm and L-2 regularization is used as a constraint until the loss no longer decreases, thereby obtaining a trained drivable area detection model; (5) Driving area detection: After normalization, the test image is input into the trained drivable area detection model, and the model output is the final drivable area detection result.

2. The method for detecting a drivable area based on a progressive gating decoder according to claim 1, wherein The feature grouping module in step (2-b) divides the input features evenly in the channel dimension to obtain four groups of features. When the number of channels of the input features cannot be divided by four, the features of N channels are discarded so that the number of channels of the remaining features is an integer multiple of four, where N is a positive integer and N∈[1,3].

3. The method for detecting a drivable area based on a progressive gating decoder according to claim 1, wherein, The convolution modules Conv in step (2-b) all have the same structure, and each convolution module consists of a convolution layer and a Relu activation function.

4. The method for detecting a drivable area based on a progressive gated decoder according to claim 1, wherein: The feature concatenation module in step (2-b) concatenates the four input feature sets in the channel dimension to obtain output features.

5. The method for detecting a drivable area based on a progressive gating decoder according to claim 1, wherein The four branches in step (2-b) are composed of four convolutional layers, branch module 1 and 2, three convolutional layers, and branch module 4, a maximum pooling layer and one convolutional layer.

6. The method for detecting a drivable area based on a progressive gating decoder according to claim 1, wherein, The convolution modules Conv in step (2-c) all have the same structure, and each convolution module consists of a convolution layer and a Relu activation function.

7. The method for detecting a drivable area based on a progressive gating decoder according to claim 1, wherein The convolution modules Conv in step (2-d) all have the same structure, and each convolution module consists of a convolution layer and a Relu activation function.

8. The method for detecting a drivable area based on a progressive gating decoder according to claim 1, wherein The upsampling modules in steps (2-c) and (2-d) both use linear interpolation upsampling.