An autonomous driving method based on convolutional neural network and attention mechanism
By introducing a method based on convolutional neural network and attention mechanism in autonomous driving technology, the problem of degradation of image data quality in complex road environments is solved, and the accuracy and robustness of detection and recognition are improved.
Patent Information
- Application Number
- CN202311192579.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-09-15
AI Technical Summary
The existing autonomous driving technology is difficult to effectively identify and process low-quality image data caused by changes in light intensity and bad weather in complex and changing road environments, resulting in the challenge of the robustness and accuracy of the algorithm model.
The autonomous driving method based on convolutional neural network and attention mechanism is adopted to build a semantic segmentation network and object detection network based on attention mechanism, so as to enhance the network's attention to important image areas, suppress noise and interference, and improve the perception of details and key goals.
It improves the adaptability of network models under complex factors, enhances the accuracy and robustness of detection and recognition, and can better deal with the decline in image data quality caused by changes in light intensity and bad weather.
Smart Images

Figure CN117115770B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to an autonomous driving method based on a convolutional neural network and an attention mechanism. Background Art
[0002] In recent years, with the rapid development of the field of artificial intelligence, how to use artificial intelligence to accelerate the empowerment of all walks of life has become a new wave of development. In the automotive industry, autonomous driving technology has led an important development direction in the future transportation field, attracted extensive attention at home and abroad, and has broad development prospects. How to transform autonomous driving technology from science fiction into reality has become a hot technology that countries around the world are competing to study. During driving, how to understand complex road scenarios has become one of the most difficult tasks in autonomous driving and assisted driving. At present, vehicles that want to achieve autonomous driving must accurately perceive and identify road information, such as elements like lane lines, traffic signs, pedestrians, vehicles, etc., and extract effective feature information from this complex environment. By using artificial intelligence algorithms such as deep convolutional neural network (DCNN), it is possible to automatically extract and learn effective features by learning a large amount of driving image data, thereby achieving the perception and recognition capabilities required in autonomous driving. In addition, artificial intelligence algorithms have excellent adaptability and iterability. The development and application of autonomous driving technology will face challenges such as constantly changing traffic environments, road conditions, and user needs. However, artificial intelligence algorithms can be flexibly optimized and adjusted according to different road scenario requirements, and have strong adaptability and iterability.
[0003] In image processing, although semantic segmentation algorithms and object detection algorithms based on traditional deep learning have achieved remarkable results, there are still some dilemmas in the field of autonomous driving: Firstly, it is how to solve the impact brought by the complex and changeable road environment. Secondly, there are still some defects in the existing algorithms themselves that need to be improved. Firstly, it is how to cope with the complex and changeable road scenes. (1) The different illumination intensities at different time periods, such as the impact brought by the illumination conditions during the day and at night. (2) In bad weather conditions, such as foggy days and rainy days. The above complex road scenes will reduce the quality of the image data collected by in-vehicle cameras during driving. When inputting these low-quality and noisy image data into the algorithm model, it will pose great challenges to the robustness and accuracy of the algorithm model. Secondly, for the existing algorithm models in autonomous driving tasks, such as neural network models for semantic segmentation of road backgrounds and lane lines, and neural network models for object detection of spatial objects such as cars, pedestrians, and traffic signs, these algorithms still have problems such as accuracy and real-time performance. For example, the deeplabv3plus neural network model commonly used for semantic segmentation cannot extract effective feature information through the feature extraction network under some complex conditions, resulting in a decrease in segmentation accuracy. Another example is the yolov7 neural network model used for object detection. Because the entire yolo series pays more attention to how to improve the real-time performance and processing speed of object detection, the accuracy of object detection is inferior to other neural network models. Summary of the Invention
[0004] To solve the above problems existing in the prior art, the present invention proposes an autonomous driving method based on a convolutional neural network and an attention mechanism, including: constructing an autonomous driving model; inputting road surface information into the trained autonomous driving model to obtain a road surface information recognition result; performing autonomous driving of the vehicle according to the road surface information recognition result; wherein the autonomous driving model includes a semantic segmentation network based on an attention mechanism and an object detection network based on an attention mechanism;
[0005] The process of training the autonomous driving model includes:
[0006] S1. Collect road image data, label the road image data; divide the labeled data into a training set, a validation set, and a test set;
[0007] S2. Input the data in the training set into the semantic segmentation network based on an attention mechanism to obtain a lane line recognition prediction map;
[0008] S3. Input the data in the training set into the object detection network based on an attention mechanism to obtain an object detection map; fuse the lane line recognition prediction map and the object detection map to obtain a recognition result;
[0009] S4: Calculate the loss function of the model according to the recognition result;
[0010] S5: Input the validation set into the autonomous driving model for verification, and use the test set to test the verified autonomous driving model, continuously adjust the parameters, and complete the model training when the loss function converges.
[0011] Advantages of the present invention:
[0012] Based on the traditional semantic segmentation network model (deeplabv3plus) and object detection network model (yolov7), the present invention proposes a deep convolutional neural network model based on the attention mechanism. By improving the attention mechanism, the network model pays more attention to important image regions, suppresses noise and interference, and improves the perception ability of details and key targets. Therefore, it complements the disadvantages of the traditional convolutional neural network model, can better enhance the adaptability of the network model to complex factors such as light intensity and bad weather, and can make the model more focused on the target feature region, improving the accuracy and robustness of detection and recognition. Description of the drawings
[0013] Figure 1 It is a flowchart of the autonomous driving method based on convolutional neural network and attention mechanism of the present invention;
[0014] Figure 2 It is a structural diagram of the optimized channel attention mechanism module of the present invention;
[0015] Figure 3 It is a structural diagram of the hybrid attention mechanism module of the present invention;
[0016] Figure 4 It is a structural diagram of the semantic segmentation network model of the present invention;
[0017] Figure 5 It is a structural diagram of the feature extraction network based on the dual attention mechanism of the present invention;
[0018] Figure 6 It is a structural diagram of the enhanced feature extraction network based on the channel attention mechanism and ASPP of the present invention;
[0019] Figure 7 It is a structural diagram of the object detection network model of the present invention;
[0020] Figure 8 It is a structural diagram of the feature extraction network based on deformable convolution and channel attention mechanism of the present invention. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0022] An autonomous driving method based on a convolutional neural network and an attention mechanism, as Figure 1 shown, the method includes: constructing an autonomous driving model; inputting road surface information into the trained autonomous driving model to obtain a road surface information recognition result; performing autonomous driving of the vehicle according to the road surface information recognition result; wherein the autonomous driving model includes a semantic segmentation network based on an attention mechanism and an object detection network based on an attention mechanism;
[0023] The process of training the autonomous driving model includes:
[0024] S1. Collect road image data and label the road image data; divide the labeled data into a training set, a validation set, and a test set;
[0025] S2. Input the data in the training set into the semantic segmentation network based on an attention mechanism to obtain a lane line recognition prediction map;
[0026] S3. Input the data in the training set into the object detection network based on an attention mechanism to obtain an object detection map; fuse the lane line recognition prediction map and the object detection map to obtain a recognition result;
[0027] S4: Calculate the loss function of the model according to the recognition result;
[0028] S5. Input the validation set into the autonomous driving model for validation, and use the test to test the validated autonomous driving model, and continuously adjust the parameters. When the loss function converges, the model training is completed.
[0029] The autonomous driving method proposed by the present invention can be roughly divided into two stages: the first stage is the training stage, in which the image data is input into the deep convolutional neural network based on the attention mechanism for training; the second stage is the test stage, in which the optimal network model saved in the training stage is used to predict the image data that has not been trained, so as to detect the performance and accuracy of the entire network model. Specifically, it includes:
[0030] Step 1: Divide the data set
[0031] Specifically, it includes: labeling the collected road image data, and then splitting the standard data set into a training set, a validation set, and a test set according to the ratio of 8:1:1. The training set participates in the training of the entire network model. The validation set does not participate in the training of the network model. Its role is to detect the state of the entire network model during the training process, such as whether it converges, etc. It is generally used to adjust hyperparameters and check whether the network model has overfitting phenomena. The test set does not participate in the training of the network model, and the entire training process has nothing to do with the test set. It is used to evaluate the parameters of the finally saved network model.
[0032] Step 2: Train and adjust the network model parameters
[0033] Input the training set and the validation set into the improved semantic segmentation network and object detection network. The main task of the optimized semantic segmentation network is to separate the image data into lane line information and background information to complete the lane line recognition task. The optimized object detection network is mainly used to detect spatial objects such as cars, pedestrians, traffic signs, and bicycles in the image data to complete the spatial object detection task.
[0034] The core of the entire autonomous driving algorithm is jointly composed of a semantic segmentation network and an object detection network. The training set and the validation set used in the two are the same during the training process, but the problems to be solved are different, and the execution order of the two is parallel, that is, the training tasks are executed simultaneously, so as to reduce the training time of the entire network model and improve the efficiency.
[0035] Step 3: Test the performance of the network model
[0036] When the validation set performs stably in the semantic segmentation network and the object detection network respectively, the training can be stopped. At this time, an optimal network model parameter can be obtained. Input the divided test set into this model to evaluate the saved optimal network model. Since the entire model has never been in contact with the test set from beginning to end, this test set can well test the generalization ability of the network model.
[0037] In this embodiment, an optimized attention mechanism is disclosed, specifically including: the purpose of this attention mechanism is to assign corresponding weights to each channel on the feature map, so that the neural network can focus on certain feature channels. The optimized channel attention mechanism is as Figure 2 shown, and the operation of this attention mechanism will be described below.
[0038] Squeeze operation: Assume that the input feature map is X, and its size is C*H*W, where C represents the number of channels of the input feature map, and H and W represent the height and length of the input feature map. Using a global maximum pooling (GlobalMaxPooling), the input feature map is compressed into a 1*1*C feature vector, which can represent the importance of each channel.
[0039] Excitation operation: The excitation operation mainly includes two full connections and two activation functions. The main purpose is to convert the importance obtained by the squeeze operation into a normalized weight value. Specifically, the feature vector obtained in the previous step is first fully connected and then activated with the Relu activation function, and then activated with a full connection and HardSigmoid activation function, and finally a weight vector representing each feature channel is obtained.
[0040] Feature weighting operation: Multiply each weight value of the learned weight vector by the channel feature on the corresponding original feature map to obtain the weighted feature map X ′ .
[0041] The optimization strategy is to replace the global average pooling in the original squeezing operation with global maximum pooling to generate the initial channel weight value. The purpose of doing this is to deal with the low-quality image data collected in different light intensities and bad weather environments mentioned above. The reason for replacing it with global maximum pooling is that it is more sensitive to edge and detail features. Global maximum pooling can highlight the edges, textures, and local detail information in the feature map, thereby effectively suppressing the interference caused by noise in low-quality images. Secondly, the global average pooling used in the original channel attention mechanism will average the features in each channel, blur the differences between features, and cause some information loss. It is difficult to deal with the negative impact of low-quality images. The global maximum pooling operation selects the maximum value, which saves the significant information in the feature map and reduces the averaging effect.
[0042] The Sigmoid activation function in the original channel attention mechanism is replaced by the HardSigmoid activation function. Compared with the two, the Sigmoid function contains exponential calculations and is slow, while the HardSigmoid function only has multiplication operations, which improves the calculation efficiency of the entire function and can effectively reduce the network training time when facing a large number of training sets. The expression of the entire channel attention mechanism is:
[0043] X ′ =Scale(X)=X*HardSigmoid(W2*Relu(W1*MaxPool(X)))
[0044] Among them, X ′It represents the weighted feature map. Scale(X) represents the multiplication operation between the feature vector and the feature channel. X represents the feature map. HardSigmoid represents the activation function. W2 represents the weight parameter generated by the second fully connected layer. Relu represents the activation function. W1 represents the weight parameter generated by the first fully connected layer. MaxPool represents the max pooling operation.
[0045] As Figure 3 shown, the hybrid domain attention mechanism is used in the present invention. As a simple and effective attention module, this attention mechanism is also often used in the training process of convolutional neural networks. For this attention mechanism, the optimized channel attention module is replaced with the original channel attention module. While using the hybrid domain attention mechanism to extract more discriminative and distinguishable features, since the hybrid domain attention mechanism will add more parameters, it will affect the calculation efficiency of the network model parameters. Therefore, weighing the performance improvement brought by using this attention mechanism and the consumption of computing resources, the present invention will only use this attention mechanism at two places.
[0046] In this embodiment, the semantic segmentation network based on the attention mechanism is as Figure 4 shown. The entire semantic segmentation network model still adopts the Encoder-Decoder structure as the main framework. The Encoder part is the focus of the innovation of this network model, mainly including the feature extraction network based on the dual attention mechanism and the enhanced feature extraction network based on the channel attention mechanism and ASPP. The Decoder part is inherited from the original network model.
[0047] Encoder part: When the image is input into the Encoder part, it will first pass through a feature extraction network based on the dual attention mechanism. This network part contains many deep convolution modules. Two feature maps will be generated from this feature extraction network. The first feature map is the low-level feature map that has not experienced all convolution modules, and the second feature map is the feature map that has experienced all convolution modules. The obtained first feature map will be directly sent to the Decoder module, and the second feature map will be sent to the enhanced feature extraction network based on the channel attention mechanism and ASPP to generate a high-level feature map. In the enhanced feature extraction network, the feature map will successively pass through the ASPP module and the channel attention module, mainly by increasing the depth and receptive field of the network to expand the context range of the feature, and then combining the weights given by the attention mechanism to learn deeper feature information, thereby improving the performance and generalization ability of the model.
[0048] Decoder part: This part uses the Decoder module in the original model. First, the low-level feature map is subjected to channel dimensionality reduction using a 1*1 convolution, and at the same time, the high-level feature map is bilinearly interpolated and upsampled. At this time, the two feature maps have become the same size in terms of dimensions. Then, the two feature maps are concatenated together and sent to a 3*3 convolution for processing, and then upsampled once to obtain the prediction map for lane line recognition.
[0049] In this embodiment, the feature extraction network based on the dual attention mechanism is introduced with the improved channel attention module and the hybrid domain attention module mentioned above on the Resnet50 network structure. The feature extraction network based on the dual attention mechanism is as Figure 5 shown. The Resnet50 structure first undergoes a 7*7 convolution and max pooling, and this part is called the initialization block. Then, it goes through four large residual blocks (ResBlock). The number of repetitions of each large residual block is different, but the operations therein are generally the same, including multiple convolutions and identity mappings. The hybrid domain attention module is added in the initialization block, and a hybrid domain attention module is also added after the last residual block (ResBlock4). This approach is equivalent to wrapping the entire feature extraction network with two large hybrid domain attention modules, making the entire feature extraction network form a whole, and improving the network's ability to focus on feature details and context from a macroscopic perspective, forming an up-down mapping relationship. For each residual block (ResBlock), the improved channel attention module is introduced on each residual block. Introducing this module can assign weights to the channel dimension of the feature map generated by each Residual, enhancing the feature representation ability of each residual block. The improved ResBlock module can be said to enhance the network's perception ability of key features from a microscopic perspective, reducing the interference of noise information and unimportant features to the network model.
[0050] The two feature maps generated from the entire feature extraction network mentioned above, the first one is the low-level feature map output after ResBlock1, and the second one is the feature map output from the last hybrid domain attention mechanism.
[0051] In this embodiment, the enhanced feature extraction network based on the channel attention mechanism and ASPP is as Figure 6As shown in the figure, the improvement of this network lies in adding an improved channel attention module after the original ASPP structure, which specifically includes: First, the feature map obtained in the previous step is input into the ASPP module. ASPP will perform multiple parallel dilated convolutions with different dilation rates and average pooling on the feature map, and then splice the generated five feature maps into a large feature map. After that, the feature map is passed to the improved channel attention module for squeezing and excitation operations, so that each small feature map is given weights on the channels, enhancing the representation ability of the features. Then, a 1*1 convolution is performed on the feature map to compress it, and finally a high-level feature map is obtained. This high-level feature map will be sent to the Decoder module.
[0052] The original network structure directly performs 1*1 convolution on the feature map generated by ASPP, that is, compresses the feature map. However, this ignores the inherent importance of each feature channel after splicing five different feature maps. After introducing the attention mechanism, it can dynamically learn the correlation between channels and adaptively adjust the weights of feature channels, making the context information represented by the entire feature map more compact.
[0053] In this embodiment, as Figure 7 shown, the object detection network based on the attention mechanism is an improvement on the yolov7 network model. The training of the object detection network based on the attention mechanism includes:
[0054] Step 1: Feature extraction; The image data will first perform feature extraction in the optimized feature extraction network. As the feature extraction network deepens, three effective feature maps are obtained, which can be called low-level, middle-level, and high-level features.
[0055] Step 2: Feature enhancement; The highest-level effective feature map will be input into the SPPCSPC structure for processing. Using this structure can make the network adapt to images of different resolutions and reduce the computational amount by half. The three effective feature maps are sent into the enhanced feature extraction double towers of FPN+PAN. First, upsampling is performed on the three feature maps to achieve feature fusion, and then downsampling is performed to achieve feature fusion.
[0056] Step 3: Output prediction results; Then, three enhanced effective feature maps will be output, and after passing through RepConv once respectively, multi-scale (large, medium, and small sizes) prediction of the same class of objects can be achieved.
[0057] The feature extraction network is optimized to a feature extraction network based on deformable convolution and channel attention mechanism; This feature extraction network introduces deformable convolution and the improved channel attention module mentioned above on the original network structure. The overall feature extraction network is as Figure 8 shown.
[0058] Specifically, the entire feature extraction network consists of multiple convolutional layers, pooling layers, and ELAN feature extraction units. The main operation is to continuously stack these modules to deepen the feature extraction of the input image. The improvement mainly focuses on the ELAN feature extraction unit, and a channel attention module is added before outputting features at different levels. The detailed improved structure is shown in Figure 8 as follows. The original ELAN feature extraction unit is stacked by three ordinary 1*1 convolutional layers and four ordinary 3*3 convolutional layers. Its main function is to perform feature extraction and control the number of feature channels. All ordinary 3*3 convolutional layers are replaced with deformable 3*3 convolutional layers.
[0059] In this embodiment, the object detection network based on the attention mechanism includes an optimized feature extraction network, an SPPCSPC structure, a strengthened feature extraction double-tower module of FPN+PAN, and three RepConv layers; the object detection network based on the attention mechanism processes the feature map as follows: input the image into the optimized feature extraction network for feature extraction to obtain a low-level feature map, a middle-level feature map, and a high-level feature map; input the high-level feature map into the SPPCSPC structure; input the output result of the SPPCSPC structure, the low-level feature map, and the middle-level feature map into the strengthened feature extraction double-tower module of FPN+PAN for sampling and fusion to obtain an effective feature map; input the effective feature map into the three RepConv layers respectively to obtain large-object recognition results, medium-object recognition results, and small-object recognition results.
[0060] The optimized feature extraction network extracts features from the image as follows: convolutional layers, pooling layers, improved ELAN feature extraction units, and improved channel attention modules; the optimized feature extraction network processes the image as follows: perform feature extraction on the input image through a 3*3 convolutional layer and an improved ELAN feature extraction unit, and output a feature map each time the input image passes through a round of three convolutional layers, one pooling operation, and an improved ELAN feature extraction unit; pass the output feature map through an improved channel attention module respectively to obtain a low-level feature map, a middle-level feature map, and a high-level feature map. The improved ELAN feature extraction unit includes: 3 ordinary 1*1 convolutional operations, 4 deformable 3*3 convolutional operations, and its processing process includes: concatenate the results of two 1*1 convolutional operations, the results of two 3*3 deformable convolutional operations, and 4 3*3 deformable convolutional operations, and finally adjust the number of channels through one 1*1 convolutional operation.
[0061] The enhanced feature extraction two - tower module of FPN + PAN samples and fuses the output results of the SPPCSPC structure, the low - level feature map, and the middle - level feature map, including: upsample the high - level feature map, and gradually stack the upsampled features with the middle - level and low - level features to generate a feature pyramid that goes down layer by layer, where each layer of the feature pyramid that goes down layer by layer is a fused feature map of a different scale; downsample the low - level fused feature map, and gradually stack the downsampled features with the middle - level and high - level fused features to generate a feature pyramid that goes up layer by layer, where each layer of the feature pyramid that goes up layer by layer is a fused feature map of a different scale.
[0062] During driving, according to the collected image data, it can be seen that the shapes of cars are inconsistent, and pedestrians also have characteristics such as tall, short, fat, and thin. However, in traditional ordinary convolution operations, the sampling positions of the convolution kernels are fixed, so it is impossible to well fit the features of irregular targets. However, deformable convolution introduces learnable offset parameters, enabling the convolution kernel to be fine - tuned at each sampling position to adapt to different deformations of the target. Such feature extraction will contain more local details and structural information.
[0063] The feature map output by the ELAN feature extraction unit passes through an improved channel attention module. At this time, each feature channel of the feature map is assigned a corresponding weight, making it more focused on important features and reducing the influence of redundant features. Through such a series of improvements, the finally output low - level, middle - level, and high - level features well reduce the negative impacts brought by complex scenes and different scales, improving the accuracy and robustness of the entire object detection.
[0064] The loss function of the model consists of a semantic segmentation network loss function based on the attention mechanism and an object detection network loss function based on the attention mechanism.
[0065] The semantic segmentation network loss function based on the attention mechanism includes:
[0066] L = L cross +L dice
[0067] where L cross represents the cross - entropy loss function, which is used when the semantic segmentation platform classifies pixel points using Softmax; L dice represents the Dice coefficient loss function. The Dice coefficient is a function for measuring the similarity of sets and is generally used to calculate the similarity between two samples.
[0068] L cross The loss includes:
[0069]
[0070] Where N represents the number of samples, C represents the number of classes, and y ij is the true label of sample i, which is 1 if sample i belongs to class j, and 0 otherwise; is the probability that the network model predicts sample i belongs to class j. This loss function can minimize the difference between the model prediction value and the true label, enable the model to better fit the data, and improve the generalization ability of the model.
[0071] L dice The loss includes:
[0072]
[0073] Where X represents the prediction result and Y represents the true result. The entire value range of L dice is between [0, 1]. The closer it is to 0, the higher the similarity between the prediction result and the true result, and the smaller the loss.
[0074] The loss function of the object detection network based on the attention mechanism includes:
[0075] L = L loc + L conf + L class
[0076] Where L loc represents the localization loss, L conf represents the confidence loss, and L class represents the classification loss. Both the confidence loss and the classification loss adopt the cross-entropy loss function, and the localization loss adopts the CIoU loss function.
[0077] L loc The localization loss includes:
[0078]
[0079] Where IoU represents the intersection over union, b represents the predicted bounding box, and b gt represents the ground truth bounding box, ρ represents the distance between the predicted bounding box and the ground truth bounding box, c represents the diagonal distance of the smallest bounding rectangle that can contain the predicted bounding box and the ground truth bounding box, α is a balancing parameter, and v is used to measure whether the aspect ratios are consistent. L loc The localization loss takes into account the distance, overlapping area, and aspect ratio between the ground truth bounding box and the predicted bounding box, which can make the network model better fit the training data and further improve the object detection effect.
[0080] The above-described embodiments further illustrate in detail the objectives, technical solutions, and advantages of the present invention. It should be understood that the above-described embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An autonomous driving method based on a convolutional neural network and an attention mechanism, characterized in that, Including: Construct an autonomous driving model; input road surface information into the trained autonomous driving model to obtain a road surface information recognition result; Perform vehicle autonomous driving according to the road surface information recognition result; wherein the autonomous driving model includes a semantic segmentation network based on an attention mechanism and an object detection network based on an attention mechanism; The process of training the autonomous driving model includes: S1. Collect road image data and annotate the road image data; divide the annotated data into a training set, a validation set, and a test set; S2. Input the data in the training set into the semantic segmentation network based on the attention mechanism to obtain a lane line recognition prediction map; the semantic segmentation network based on the attention mechanism adopts an Encoder-Decoder structure, including an Encoder module and a Decoder module; wherein the Encoder module consists of a feature extraction network based on a dual attention mechanism and an enhanced feature extraction network based on a channel attention mechanism and ASPP; The Decoder module consists of a first convolutional layer, a bilinear interpolation upsampling layer, a splicing layer, a second convolutional layer, and an upsampling layer; The feature extraction network based on the dual attention mechanism includes an initialization module, four residual module groups, and a first hybrid domain attention mechanism module; wherein the initialization module consists of a convolutional layer, a second hybrid domain attention mechanism module, and a max pooling layer, each residual module group consists of different residual modules, and each residual module introduces an improved channel attention mechanism module; the improved channel attention mechanism includes: X′ = Scale(X) = X * HardSigmoid(W2 * Relu(W1 * MaxPool(X))) wherein, X′ represents the weighted feature map, Scale(X) represents the multiplication operation between the feature vector and the feature channel, X represents the feature map, HardSigmoid represents the activation function, W2 represents the weight parameter generated by the second fully connected layer, Relu represents the activation function, W1 represents the weight parameter generated by the first fully connected layer, and MaxPool represents the max pooling operation; The processing of the feature map by the hybrid domain attention mechanism includes: inputting the feature map into the improved channel attention module to obtain a channel feature map; fusing the channel feature map with the input feature map to obtain a fused feature map; using a spatial attention module to perform spatial feature extraction on the fused feature map; fusing the spatial feature map and the fused feature map to obtain an output feature map; S3. Input the data in the training set into the object detection network based on the attention mechanism to obtain an object detection map; fuse the lane line recognition prediction map and the object detection map to obtain a recognition result; S4: Calculate the loss function of the model according to the recognition result; S5: Input the validation set into the autonomous driving model for validation, use the test to test the validated autonomous driving model, continuously adjust the parameters, and complete the model training when the loss function converges.
2. The autonomous driving method based on a convolutional neural network and an attention mechanism according to claim 1, characterized in that, Processing the image using a semantic segmentation network based on the attention mechanism includes: The feature extraction network based on the dual attention mechanism consists of multiple deep convolutional modules; Input the road image into the feature extraction network based on the dual attention mechanism for feature extraction to obtain the first feature map and the second feature map; Input the second feature map into the enhanced feature extraction network based on the channel attention mechanism and ASPP to obtain the high-level feature map; Input the first feature map and the high-level feature map into the Decoder module; Use the first convolutional layer to reduce the channels of the first feature map, and use the bilinear interpolation upsampling layer to perform bilinear interpolation upsampling on the high-level feature map; Concatenate the downsampled feature map and the sampled feature map, and input the concatenated feature map into the second convolutional layer and the upsampling layer to obtain the prediction map for lane line recognition.
3. The autonomous driving method based on a convolutional neural network and an attention mechanism according to claim 1, characterized in that, The object detection network based on the attention mechanism includes an optimized feature extraction network, an SPPCSPC structure, an enhanced feature extraction double tower module of FPN+PAN, and three RepConv layers; Processing the feature map by the object detection network based on the attention mechanism includes: Input the picture into the optimized feature extraction network for feature extraction to obtain the low-level feature map, the middle-level feature map, and the high-level feature map; Input the high-level feature map into the SPPCSPC structure; Input the output result of the SPPCSPC structure, the low-level feature map, and the middle-level feature map into the enhanced feature extraction double tower module of FPN+PAN for sampling and fusion to obtain the effective feature map; Input the effective feature map into the three RepConv layers respectively to obtain the large object recognition result, the medium object recognition result, and the small object recognition result.
4. The autonomous driving method based on a convolutional neural network and an attention mechanism according to claim 3, characterized in that, The optimized feature extraction network performs feature extraction on the picture including: convolutional layer, pooling layer, improved ELAN feature extraction unit, and improved channel attention module; The optimized feature extraction network processes the picture including: Input the input picture through a 3*3 convolution and an improved ELAN feature extraction unit for feature extraction, and output a feature map once after the input image goes through three convolutions, one pooling operation, and one improved ELAN feature extraction unit in each round; Respectively pass the output feature map through an improved channel attention module to obtain the low-level feature map, the middle-level feature map, and the high-level feature map; The improved ELAN feature extraction unit includes: 3 times of 1*1 ordinary convolution, 4 times of 3*3 deformable convolution, and its processing process includes: Concatenate the results of two 1*1 convolutions, and the results of two 3*3 deformable convolutions and four 3*3 deformable convolutions, and finally adjust the number of channels through one 1*1 convolution.
5. The autonomous driving method based on a convolutional neural network and an attention mechanism according to claim 3, characterized in that, The enhanced feature extraction dual-tower module of FPN+PAN samples and fuses the output results of the SPPCSPC structure, low-level feature maps, and intermediate-level feature maps, including: performing upsampling on the high-level feature map, gradually stacking the upsampled features with intermediate-level features and low-level features to generate a feature pyramid that descends layer by layer, where each level of the feature pyramid that descends layer by layer is a fused feature map of a different scale; performing downsampling on the low-level fused feature map, gradually stacking the downsampled features with intermediate-level fused features and high-level fused features to generate a feature pyramid that ascends layer by layer, where each level of the feature pyramid that ascends layer by layer is a fused feature map of a different scale.
6. The autonomous driving method based on a convolutional neural network and an attention mechanism according to claim 1, characterized in that, The loss function of the model consists of a semantic segmentation network loss function based on the attention mechanism and an object detection network loss function based on the attention mechanism.
Citation Information
Patent Citations
Ground background small target identification method based on cascade combination network
CN115861756A
Drivable area division method for unmanned logistics vehicle
CN116279592A