A method and system for detecting drivable areas and lane lines on traffic roads

Through the improved TwinLiteNet model, the combination of downsampling layers and other levels can be separated by reparameterization depth, the accuracy and efficiency problems of road driving areas and lane lines detection in complex traffic scenarios are solved, and efficient and robust detection effects are achieved.

CN120088752BActive Publication Date: 2025-07-08NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510570642.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-08
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In the complex traffic scenarios, the accuracy and efficiency of road travelable areas and lane lines are low, especially in low light conditions, the detection accuracy is insufficient and the robustness is insufficient, making it difficult to adapt to a variable environment and meet real-time reaction needs.

Method used

Using the improved TwinLiteNet model, the downsampling layer, position attention layer, depth separation convolution layer, partially decomposed self-attention layer and image decomposition fusion layer are used to reparameterize depth, enhance feature extraction and detection accuracy, and reduce calculation amount and reasoning delay.

Benefits of technology

It improves the accuracy and efficiency of traffic roads and lane line detection, reduces the calculation amount and reasoning delay, and enhances the model's adaptability and real-time response capabilities in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088752B_ABST
    Figure CN120088752B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting a drivable area and lane line of a traffic road, and relates to the technical field of automatic driving and assisted driving, including: obtaining an initial driving video image, and dividing it into a training image and an image to be tested; improving the existing TwinLiteNet model based on a preset re-parameterized deep separable downsampling layer, a position attention layer, a deep separable convolution layer, a partial decomposition self-attention layer, and an image decomposition and fusion layer, and building an improved TwinLiteNet model; using the training image to train the improved TwinLiteNet model, and outputting the detection results of the drivable area and lane line of the traffic road. The present invention can improve the accuracy and efficiency of the drivable area and lane line detection by building an improved TwinLiteNet model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of autonomous driving and assisted driving, and particularly relates to a method and system for detecting drivable areas and lane lines on traffic roads. Background Art

[0002] In the automotive industry, autonomous driving and assisted driving technologies are regarded as the key to alleviating and even solving traditional traffic problems. Among them, the accurate detection of various drivable areas and lane lines in the traffic environment is the basis and key step for completing autonomous driving tasks. Only by accurately identifying and understanding the lane environment can autonomous vehicles make correct decisions and ensure driving safety.

[0003] Although current vision-based algorithms for detecting drivable areas and lane lines on roads have developed rapidly and achieved remarkable results in detecting drivable areas and lane lines in general traffic scenarios, there are still many challenges in more complex traffic scenarios. For example, in dense traffic, vehicles or pedestrians on the road may block the road and lane lines, thus affecting the detection accuracy. Especially in harsh environments, such as at night or under low light conditions, the quality of the captured photos is often low, and some areas may become blurred due to insufficient light, which further exacerbates the difficulty of detecting the road surface and lane lines. When the environment is too dark and the lighting conditions are insufficient, the road and the surrounding environment may be confused during the detection process, and it is difficult to detect the lane lines. In addition, many existing algorithms lack robustness in actual applications and are difficult to adapt to changing environments and meet the requirements of real-time response. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method and system for detecting drivable areas and lane lines on traffic roads, which can solve the problem of low accuracy and efficiency in detecting drivable areas and lane lines on traffic roads in the prior art.

[0005] To solve the above technical problems, this application is implemented as follows:

[0006] In a first aspect, the embodiments of this application provide a method for detecting drivable areas and lane lines on traffic roads. The method includes: obtaining an initial driving video image and dividing it into training images and images to be tested; improving the existing TwinLiteNet model based on a preset reparameterized depthwise separable downsampling layer, position attention layer, depthwise separable convolutional layer, partially decomposed self-attention layer, and image decomposition fusion layer, and building an improved TwinLiteNet model; training the improved TwinLiteNet model using the training images to obtain a trained TwinLiteNet model; inputting the images to be tested into the trained TwinLiteNet model for detection, and outputting the detection results of drivable areas and lane lines on traffic roads.

[0007] As an alternative implementation of the first aspect of the present application, the image decomposition and fusion layer is used to obtain enhanced features. The specific process includes: replacing the downsampling layer in the existing TwinLiteNet model with a reparameterized depthwise separable downsampling layer to obtain the number of channels of each branch in the TwinLiteNet model. The number of channels consists of a first number of channels and a second number of channels. Each first number of channels is the ratio of the number of output channels to the number of branches, and the second number of channels is the difference between the number of output channels and the total number of the first number of channels; using a depthwise separable reparameterized convolutional layer to replace the dilated convolution of each branch in the downsampling layer, and inputting the input feature map into multiple depthwise separable reparameterized convolutional layers according to the number of channels for convolutional processing to obtain multiple receptive field features. Among them, the number of depthwise separable reparameterized convolutional layers on each branch increases sequentially; adding a position attention layer behind each branch in the reparameterized depthwise separable downsampling layer to perform feature enhancement processing on the multiple receptive field features to obtain the enhanced features output by multiple branches.

[0008] As an alternative implementation of the first aspect of the present application, the mathematical expression of the first number of channels is:

[0009] ;

[0010] The mathematical expression of the second number of channels is:

[0011] ;

[0012] Among them, represents the first number of channels, represents the number of output channels, represents the second number of channels.

[0013] As an alternative implementation of the first aspect of the present application, part of the decomposed self-attention layer is placed at the end of the attention mechanism layer in TwinLiteNet to obtain a simple global attention feature map. Among them, the process of obtaining the simple global attention feature map includes: performing resolution scaling processing on multiple feature maps with different resolution sizes in the feature pyramid to obtain multiple feature maps with the same resolution; sequentially performing channel splicing and convolutional aggregation feature processing on the multiple feature maps with the same resolution to obtain a first feature map; performing channel segmentation on the first feature map to obtain a second feature map and a third feature map, and the number of channels of the second feature map and the third feature map is the same; respectively performing horizontal self-attention and vertical self-attention calculations on the second feature map to obtain a horizontal feature map and a vertical feature map; performing channel splicing processing on the horizontal feature map, the vertical feature map, the second feature map and the third feature map to obtain a simple global attention feature map.

[0014] As an alternative implementation of the first aspect of the present application, the enhancement features include lane line upsampling features and drivable area upsampling features. Among them, the specific process of obtaining the lane line upsampling features and the drivable area upsampling features includes: performing a 3×1 convolution on the input feature map to obtain a first output feature; performing a deconvolution on the image features in the backbone that have the same resolution as the image features in the current decoder upsampling stage to obtain convolution features; performing channel max pooling, channel average pooling, and activation processing on the convolution features to obtain a first spatial weight; multiplying the first spatial weight with the upsampled lane line feature map and adding the result residually to obtain the lane line upsampling features; performing a 1×3 convolution on the first output feature to obtain a second output feature; performing channel max pooling, channel average pooling, and activation processing on the second output feature to obtain a second spatial weight; multiplying the second spatial weight with the drivable area features and adding the result residually to obtain the drivable area upsampling features.

[0015] In a second aspect, an embodiment of the present application provides a traffic road drivable area and lane line detection system, the system includes:

[0016] An image acquisition module, configured to acquire an initial driving video image and divide it into training images and images to be tested;

[0017] A reparameterized depthwise separable downsampling module, configured to replace the downsampling layer in the existing TwinLiteNet model to obtain feature information with different receptive fields;

[0018] A lightweight processing module, configured to replace the ordinary convolution layers in the backbone of the existing TwinLiteNet model with depthwise separable convolution layers for lightweight processing to obtain a lightweight backbone;

[0019] A position attention module, configured to perform feature enhancement processing on the input features to obtain enhanced features;

[0020] A partial decomposition self-attention module, configured to perform resolution scaling processing on feature maps with multiple different resolution sizes in the feature pyramid to obtain feature maps with the same resolution;

[0021] An image decomposition fusion module, configured to fuse multiple feature maps with the same resolution and the input of the decoder features to enhance the decoding effect;

[0022] A model training module, configured to perform iterative optimization training on the improved TwinLiteNet model using the training images to obtain the improved TwinLiteNet model;

[0023] The detection module is used to input the image to be detected into the improved TwinLiteNet model for detection, so as to output the detection results of the drivable area and the lane lines.

[0024] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method in the first aspect are implemented.

[0025] In a fourth aspect, an embodiment of the present application provides a readable storage medium, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method in the first aspect are implemented.

[0026] Compared with the prior art, the beneficial effects of a method for detecting the drivable area and lane lines of a traffic road proposed by the present invention are as follows: First, by stacking depthwise separable reparameterized convolutional layers, the advantages of few parameters, small computational complexity, and large receptive field are effectively combined. Different numbers of stacked depthwise separable reparameterized convolutional layers in different branches can effectively obtain features with different receptive field sizes, and the accuracy during training can be effectively guaranteed. Second, the newly added position attention layer in each branch first obtains the average pixel of each point in the feature map through average pooling of the feature map channels, calculates the weights through reshaping and softmax, can obtain the attention weights of each pixel point in the feature map, and multiplies the original image element by element, which can strengthen the original feature image; replacing the original standard convolutional layer with a depthwise separable convolutional layer can effectively reduce the overall parameters and computational complexity of the model. Third, by adding a partially decomposed self-attention layer at the end of the attention mechanism layer in the TwinLiteNet model, the channel data can be effectively utilized, perform attention calculation on half of the data to obtain simple global attention, and finally perform channel-level splicing on the output of each stage and then pass through a 1x1 convolution, which not only obtains global attention information but also reduces the computational complexity of obtaining attention. Fourth, adding an image decomposition and fusion layer in the decoder upsampling stage can enhance the decoding effect.

[0027] In summary, a method for detecting the drivable area and lane lines of a traffic road proposed by the present application can detect large-scale traffic road drivable areas and small-scale lane lines, reduce the computational complexity and inference latency in the segmentation process, and improve the accuracy and efficiency of detecting the drivable area of a traffic road. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flowchart of a method for detecting the drivable area and lane lines of a traffic road proposed in the first embodiment of the present application;

[0029] Figure 2 The structural diagram of the reparameterized depthwise separable downsampling layer proposed in the first embodiment of this application;

[0030] Figure 3 The structural diagram of the depthwise separable reparameterized convolutional layer proposed in the first embodiment of this application;

[0031] Figure 4 The structural diagram of the position attention layer proposed in the first embodiment of this application;

[0032] Figure 5 The structural diagram of the lightweight backbone proposed in the first embodiment of this application;

[0033] Figure 6 The structural diagram of the partial decomposition self-attention layer proposed in the first embodiment of this application;

[0034] Figure 7 The structural diagram of the image decomposition and fusion layer proposed in the first embodiment of this application;

[0035] Figure 8 The structural schematic diagram of a traffic road drivable area and lane line detection system proposed in the second embodiment of this application. Detailed implementation manners

[0036] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0037] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects. The character " / " generally indicates an "or" relationship between the associated objects before and after.

[0038] Next, in conjunction with the accompanying drawings, a traffic road drivable area and lane line detection method and system provided by the embodiments of this application will be described in detail through specific embodiments and their application scenarios.

[0039] Embodiment 1

[0040] Please refer to Figure 1, which is a flowchart of a method for detecting drivable areas and lane lines on traffic roads proposed in the first embodiment of this application. The proposed method includes steps S01 to S03.

[0041] S01: Obtain the initial driving video image and divide it into training images and images to be tested.

[0042] Specifically, the present invention screens images from a publicly available driving video dataset to ensure coverage of diverse scenarios, weather conditions, and time periods. Based on a ratio of 7:1, the image dataset is divided into a training image set and an image set to be tested, ensuring a reasonable data distribution between the two. In terms of weather, the dataset covers six conditions: sunny, cloudy, overcast, rainy, snowy, and foggy, which helps the model adapt to different weather conditions. In terms of scenarios, the dataset covers six common scenarios: residential areas, highways, urban streets, parking lots, gas stations, and tunnels, enhancing the generalization ability of the model in practical applications. In terms of time, the dataset covers three time periods: dawn / dusk, daytime, and night, with a focus on daytime and night, further increasing the diversity of the dataset.

[0043] To obtain a more accurate and efficient training image set, the present invention performs data cleaning on the training image set to remove duplicate annotation boxes. The image set to be tested does not require additional processing and can be directly used as a validation set. During the process of optimizing the training image set, the present invention adopts a data augmentation strategy to improve the generalization ability and robustness of the model. Specifically, the present invention implements two main data augmentation techniques, which are directly based on the provided code logic: color space transformation and random geometric transformation (including perspective transformation).

[0044] In terms of color processing, the present invention introduces a method of random color transformation. Through the augment_hsv function, the present invention randomly adjusts the three channels of hue, saturation, and value in the HSV color space. This process is achieved by generating a random gain value for each channel, applying it to the lookup table, and then using these lookup tables to transform the color attributes of the original image. This not only increases the diversity of the images but also simulates the scene changes under different lighting conditions, which helps improve the adaptability of the model to various environmental conditions. Secondly, to increase geometric diversity, the present invention applies random geometric transformations, and the present invention realizes various transformations including rotation, scaling, shearing, translation, and perspective distortion. The augment_hsv function receives a combination of color images, grayscale images, and line drawings and performs a series of affine transformations on them. The transformation matrix is constructed by combining operations such as centering, perspective transformation, rotation and scaling, shear transformation, and translation. Finally, if the image has changed, the cv2.warpPerspective or cv2.warpAffine function is selected to transform the image and its corresponding grayscale image and line drawing according to whether the perspective transformation is enabled. This method can effectively simulate various poses and perspective changes of objects in three-dimensional space, enhancing the complexity and authenticity of the dataset. In addition, to further expand the diversity of the dataset, the present invention also performs a horizontal flipping operation with a probability of 50% during the training phase. This means that for each training image, there is a 50% chance of being mirrored along the horizontal axis, thereby introducing an additional dimension of variation and enabling the model to learn a more comprehensive feature representation.

[0045] By combining the above data augmentation means of color space transformation and geometric transformation, the present invention not only increases the quantity of training data, but more importantly, significantly improves the quality and diversity of the data. Such a dataset can more effectively support the learning process of computer vision models and enable them to perform better when facing various challenges in the real world in the future.

[0046] S02: Based on the preset reparameterized depthwise separable downsampling layer, position attention layer, depthwise separable convolutional layer, partially decomposed self-attention layer, and image decomposition and fusion layer, improve the existing TwinLiteNet model and build an improved TwinLiteNet model.

[0047] Among them, the process of building the improved TwinLiteNet model includes steps S021 to S025.

[0048] S021: Replace the downsampling layer in the existing TwinLiteNet model with a reparameterized depthwise separable downsampling layer to obtain feature information of different size scales;

[0049] Specifically, the specific process of obtaining feature information of different scales includes: replacing the downsampling layer in the existing TwinLiteNet model with a reparameterized depthwise separable downsampling layer to obtain the number of channels in each branch within the TwinLiteNet model. The number of channels consists of a first number of channels and a second number of channels. Each first number of channels is the ratio of the number of output channels to the number of branches, and the second number of channels is the difference between the number of output channels and the total number of the first number of channels; replacing the dilated convolution in each branch within the downsampling layer with a depthwise separable reparameterized convolution layer, and inputting the input feature map into multiple depthwise separable reparameterized convolution layers for convolution processing according to the number of channels to obtain multiple receptive field features. Among them, the number of depthwise separable reparameterized convolution layers on each branch increases sequentially; adding a position attention layer behind each branch within the reparameterized depthwise separable downsampling layer to perform feature enhancement processing on the multiple receptive field features to obtain enhanced features output by multiple branches; after sequentially performing channel concatenation, normalization, and activation processing on the multiple enhanced features, obtaining feature information of different scales.

[0050] Among them, the mathematical expression of the first number of channels is:

[0051] ;

[0052] Among them, the mathematical expression of the second number of channels is:

[0053] ;

[0054] Among them, represents the first number of channels, represents the number of output channels, represents the second number of channels.

[0055] Please refer to Figure 2 for the structure diagram of the reparameterized depthwise separable downsampling layer. The input channels are input into the depthwise separable convolution layer, batch normalization processing, and Relu activation function processing, and the obtained number of channels is n. Among them, the size of the convolution kernel within the depthwise separable convolution layer is 3, and the stride is 2. In the internal structure of the reparameterized depthwise separable downsampling layer, there are five channel branches, and the number of channels in each branch is , from left to right, it is divided into the first branch, the second branch, the third branch, the fourth branch and the fifth branch. Each branch contains two convolutional layers, two batch normalization layers, a position attention layer and two convolutional kernels with a size of 1 and a stride of 1. In addition, each branch also contains a reparameterized depthwise separable convolutional layer. The number of reparameterized depthwise separable convolutional layers in each branch from left to right increases sequentially, which are 1, 2, 3, 4, 5 respectively. When each branch performs activation function (to obtain non-linearity) processing, the number of channels changes from n to 2n. When the receptive field features are input into the position attention module for feature enhancement processing to obtain the enhanced features output by multiple branches, the number of channels of the first branch changes from 2n to n1 (n1 < 2n), and the number of channels of the remaining branches is n (n1 < n). Finally, after adding and concatenating the enhanced features output by multiple branches, it passes through a batch normalization layer with the number of channels nOut and a relu activation function.

[0056] Please refer to Figure 3 , which is the structural diagram of the depthwise separable reparameterized convolutional layer. When the input features are input into the depthwise separable reparameterized convolutional layer, it will be divided into 3 branches. The first branch is the pointwise convolution with a kernel size of 3 and the number of groups equal to the number of channels of the input features. The second branch is the pointwise convolution with a kernel size of 1 and the number of groups equal to the number of channels of the input features. The third branch is the residual connection. The input features are respectively input into the pointwise convolution with a kernel size of 3 and the number of groups equal to the number of channels of the input features and the pointwise convolution with a kernel size of 1 and the number of groups equal to the number of channels of the input features and then added together. The result obtained by adding is subjected to batch normalization processing and activation function processing.

[0057] The pointwise convolution used in the present invention can reduce the number of parameters and the amount of calculation, and the combination of multiple branches can improve the accuracy of the model. Using multiple branches during training can improve the performance of the model.

[0058] S022: After inserting the position attention layer into each branch of the reparameterized depthwise separable downsampling layer in the TwinLiteNet model, it is used to capture the weight scores of each pixel in different receptive field feature maps.

[0059] Specifically, please refer to Figure 4 , which is the structural diagram of the position attention layer. The position attention layer has two branches. One branch is responsible for calculating the attention weights. This branch first performs channel average pooling on the input features with a size of to obtain the average feature between channels. Then, it reshapes the average feature between channels, reshaping the feature map with dimensions of 1 H W into HW 1 1, and after passing through the softmax function, according to the pixel attention weights, for HW 1 The feature map of 1 is reshaped again to the original dimension 1 H W. Another branch is responsible for element-wise multiplying the input features with the pixel attention weights and performing residual addition to obtain the final output result.

[0060] S023: Lightweight processing is performed on the standard convolution of the backbone in the existing TwinLiteNet model to obtain a lightweight backbone.

[0061] It should be noted here that the number of parameters of the original 3×3 convolution kernel is small, but excessive stacking will still affect the number of parameters of the model. To effectively reduce the number of parameters and computational complexity of the model, the present invention performs lightweight processing on the standard convolution of the backbone in the existing TwinLiteNet model, that is, uses a depthwise separable convolution layer to replace the standard convolution layer.

[0062] Please refer to Figure 5 , which is the structure diagram of the lightweight backbone. There are two convolutions in it, namely, a depthwise convolution with the number of groups equal to the number of input channels and a pointwise convolution with the number of output channels equal to the number of output channels of the original standard convolution. The depthwise separable convolution layer proposed by the present invention decomposes the traditional convolution into two simpler steps: depthwise convolution (applying a 3x3 convolution kernel independently to each input channel) and pointwise convolution (using a 1x1 convolution kernel to combine the results and adjust the number of output channels). This decomposition method greatly reduces the number of parameters and computational complexity. For example, when the number of input and output channels is 64, the number of parameters of the standard 3×3 convolution is: 9×64×64 = 36864, while the number of parameters of the depthwise separable convolution layer is: 9×64 + 64×64 = 4608, that is, the number of parameters is reduced by approximately 87%. At the same time, the computational complexity is also reduced accordingly, improving the running efficiency of the model. By using a depthwise separable convolution layer to replace the standard convolution layer, not only the lightweight of the model is achieved, but also the efficient deployment and performance optimization of the model in resource-constrained environments are ensured.

[0063] S024: Partially decompose the self-attention layer and insert it at the end of the attention mechanism layer in the TwinLiteNet model.

[0064] Specifically, the partially decomposed self-attention layer is placed at the end of the attention mechanism layer in TwinLiteNet to obtain a simple global attention feature map. Among them, the process of obtaining the simple global attention feature map includes: performing resolution scaling processing on multiple feature maps with different resolution sizes in the feature pyramid to obtain multiple feature maps with the same resolution;

[0065] Successively perform channel concatenation and convolutional aggregation feature processing on multiple feature maps with the same resolution to obtain a first feature map; perform channel splitting on the first feature map to obtain a second feature map and a third feature map, where the second feature map and the third feature map have the same number of channels; perform horizontal self-attention and vertical self-attention calculations on the second feature map respectively to obtain a horizontal feature map and a vertical feature map; perform channel concatenation processing on the horizontal feature map, the vertical feature map, the second feature map, and the third feature map to obtain a simple global attention feature map.

[0066] Please refer to Figure 6 , which is the structural diagram of a partially decomposed self-attention layer. First, perform dimensional reshaping on feature maps with different resolution sizes in the feature pyramid, then perform channel concatenation and 3 convolution processing, and then perform average channel splitting processing on the processed feature map (i.e., evenly splitting the input channels into two input channels ) to obtain two feature maps X1 and Y1. Pass X1 through two parallel route calculations, namely horizontal multi-head self-attention calculation and vertical multi-head self-attention calculation, and input the calculation results into two feed-forward neural networks respectively to obtain two output feature maps X2 and X3. Finally, perform channel concatenation on X1, Y1, X2, and X3. In this way, the computational complexity can be effectively reduced.

[0067] S025: Insert the image decomposition and fusion layer into the upsampling stage of the decoder in the TwinLiteNet model to obtain the lane upsampling feature and the drivable area upsampling feature.

[0068] Specifically, please refer to Figure 7 , which is the structural diagram of the image decomposition and fusion layer. The process of obtaining the lane upsampling feature and the drivable area upsampling feature includes: performing 3×1 convolution on the input feature map to obtain a first output feature; performing decomposition convolution on the image feature in the backbone with the same resolution as the image feature in the current decoder upsampling stage to obtain a convolutional feature; performing channel max pooling, channel average pooling, and activation processing on the convolutional feature to obtain a first spatial weight; multiplying the first spatial weight with the upsampled lane line feature map and adding them residually to obtain the lane upsampling feature; performing 1×3 convolution on the first output feature to obtain a second output feature; performing channel max pooling, channel average pooling, and activation processing on the second output feature to obtain a second spatial weight; multiplying and adding residually based on the second spatial weight and the drivable area feature to obtain the drivable area upsampling feature.

[0069] The present invention performs deconvolution on the image features in the backbone that have the same resolution as the image features in the upsampling stage of the current decoder to obtain convolutional features, which can better obtain the feature information in the vertical direction, thereby facilitating the segmentation of lane lines.

[0070] S03: Use the training images to train the improved TwinLiteNet model, obtain the trained TwinLiteNet model, and input the image to be tested into the trained TwinLiteNet model for detection, and output the drivable area of the traffic road and the lane line detection results.

[0071] Specifically, the training images used in the present invention are the training image set obtained after data cleaning in step S01. When the parameters of the improved TwinLiteNet model remain unchanged, the trained TwinLiteNet model is obtained.

[0072] In summary, a method for detecting the drivable area of a traffic road and lane lines provided by the present application is implemented through a trained TwinLiteNet model. Among them, the reparameterized depthwise separable downsampling layer is used to replace the original downsampling layer, effectively reducing the number of calculation parameters and the amount of calculation. The per-channel convolution can effectively reduce the number of parameters of the convolution, and the multi-branch can ensure the accuracy of the model. In addition, stacking depthwise separable reparameterized convolutional layers in each branch and using depthwise separable reparameterized convolutional layers to replace the dilated convolutions in each branch of the downsampling layer can reduce the grid problems existing in the dilated convolutions.

[0073] The position attention layer proposed by the present invention can effectively capture the weight scores of each pixel in the feature maps of different receptive fields, so that the model can more effectively notice the important features in the feature image, exclude redundant features, thereby achieving the purpose of enhancing the image features. The lightweight backbone can minimize the total number of parameters of the model, so that the model can better run on edge devices.

[0074] The present invention uses a partial decomposition self-attention layer to effectively fuse the backbone features of different stages and perform channel splicing on them, and uses half of the number of channels for horizontal self-attention calculation and vertical self-attention calculation. This can better combine features of different scales, making the segmentation more accurate, and applying the original half of the channel data to horizontal self-attention and vertical self-attention can obtain global attention weights with less computational effort.

[0075] The image decomposition and fusion layer proposed by the present invention fuses the inputs of multiple feature maps with the same resolution and decoder features, which can enhance the decoding effect.

[0076] Embodiment 2

[0077] Please refer to Figure 8, which shows a schematic structural diagram of a traffic road drivable area and lane line detection system proposed in the second embodiment of the present application, including:

[0078] An image acquisition module, configured to acquire an initial driving video image and divide it into training images and images to be tested;

[0079] A reparameterized depthwise separable downsampling module, configured to replace the downsampling layer in the existing TwinLiteNet model to obtain feature information with different receptive field sizes;

[0080] A lightweight processing module, configured to replace the ordinary convolution layers in the backbone of the existing TwinLiteNet model with depthwise separable convolution layers for lightweight processing to obtain a lightweight backbone;

[0081] A position attention module, configured to perform feature enhancement processing on the input features to obtain enhanced features;

[0082] A partially decomposed self-attention module, configured to perform resolution scaling processing on feature maps with multiple different resolution sizes in the feature pyramid to obtain feature maps with the same resolution;

[0083] An image decomposition and fusion module, configured to fuse the input of feature maps with the same resolution and decoder features to enhance the decoding effect;

[0084] A model training module, configured to perform iterative optimization training on the improved TwinLiteNet model using the training images to obtain the improved TwinLiteNet model;

[0085] A detection module, configured to input the images to be tested into the improved TwinLiteNet model for detection to output the drivable area detection result and the lane line detection result.

[0086] The beneficial effects of a traffic road drivable area and lane line detection system provided by this application are as follows. First, through the stacking of reparameterized depthwise separable downsampling modules, the advantages of few parameters, small computational complexity, and large receptive field are effectively combined, enabling the effective acquisition of features with different receptive field sizes, and the accuracy during training can be effectively guaranteed. Second, for the newly added position attention module in each branch, first, the average pixel of each point in the feature map can be obtained by performing average pooling on the channels of the feature map, and the attention weights of each pixel point in the feature map can be obtained through reshaping and softmax calculation of weights, and then the original image is multiplied element-wise to strengthen the original feature image. Replacing the original standard convolution layer with a depthwise separable convolution layer can effectively reduce the overall parameters and computational complexity of the model. Third, by introducing a partially decomposed self-attention module, channel data can be effectively utilized, and attention calculation is performed on half of the data to obtain simple global attention. Finally, the outputs of each stage are concatenated at the channel level and then passed through a 1x1 convolution, which not only obtains global attention information but also reduces the computational complexity of obtaining attention. Fourth, by adding an image decomposition and fusion module in the decoder upsampling stage of the TwinLiteNet model, the decoding effect can be enhanced.

[0087] A traffic road drivable area and lane line detection system in an embodiment of this application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of this application do not make specific limitations.

[0088] A traffic road drivable area and lane line detection system in an embodiment of this application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of this application do not make specific limitations.

[0089] A traffic road drivable area and lane line detection system provided by an embodiment of this application can achieve Figures 1 to 7For the sake of avoiding repetition, the processes implemented in the method embodiments of a traffic road drivable area and lane line detection method are not elaborated here.

[0090] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a traffic road drivable area and lane line detection method and can achieve the same technical effect. For the sake of avoiding repetition, it is not elaborated here.

[0091] An embodiment of the present application further provides a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a traffic road drivable area and lane line detection method and can achieve the same technical effect. For the sake of avoiding repetition, it is not elaborated here.

[0092] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0093] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be executed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0095] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A method for detecting drivable areas and lane lines on traffic roads, characterized in that, Including: Obtain the initial driving video image and divide it into training images and images to be tested; Improve the existing TwinLiteNet model based on the preset reparameterized depthwise separable downsampling layer, position attention layer, depthwise separable convolutional layer, partially decomposed self-attention layer, and image decomposition and fusion layer, and build an improved TwinLiteNet model; the partially decomposed self-attention layer is placed at the end of the attention mechanism layer in TwinLiteNet to obtain a simple global attention feature map. Among them, the process of obtaining the simple global attention feature map includes: performing resolution scaling processing on multiple feature maps with different resolution sizes in the feature pyramid to obtain multiple feature maps with the same resolution; sequentially performing channel splicing and convolutional aggregation feature processing on multiple feature maps with the same resolution to obtain a first feature map; performing channel segmentation on the first feature map to obtain a second feature map and a third feature map, and the second feature map and the third feature map have the same number of channels; respectively performing horizontal self-attention and vertical self-attention calculations on the second feature map to obtain a horizontal feature map and a vertical feature map; performing channel splicing processing on the horizontal feature map, vertical feature map, second feature map, and third feature map to obtain a simple global attention feature map; the image decomposition and fusion layer is used to obtain enhanced features. The specific process includes: replacing the downsampling layer in the existing TwinLiteNet model with a reparameterized depthwise separable downsampling layer to obtain the number of channels of each branch in the TwinLiteNet model; using a depthwise separable reparameterized convolutional layer to replace the dilated convolution in each branch of the downsampling layer, and inputting the input feature map into multiple depthwise separable reparameterized convolutional layers according to the number of channels for convolutional processing to obtain multiple receptive field features; adding a position attention layer behind each branch in the reparameterized depthwise separable downsampling layer to perform feature enhancement processing on multiple receptive field features to obtain enhanced features output by multiple branches, where the number of depthwise separable reparameterized convolutional layers on each branch increases sequentially; Use the training images to train the improved TwinLiteNet model to obtain a trained TwinLiteNet model; Input the images to be tested into the trained TwinLiteNet model for detection, and output the traffic road drivable area and lane line detection results.

2. The method for detecting a drivable area and lane lines of a traffic road according to claim 1, characterized in that The number of channels consists of a first number of channels and a second number of channels. Each first number of channels is the ratio of the output number of channels to the number of branches, and the second number of channels is the difference between the output number of channels and the total number of the first number of channels.

3. A method for detecting a traffic road drivable area and lane lines according to claim 2, wherein The mathematical expression of the first number of channels is: ; The mathematical expression of the second number of channels is: ; Among them, represents the number of first channels, represents the number of output channels, represents the number of second channels.

4. A method for detecting drivable areas and lane lines on a traffic road according to claim 1, characterized in that, The enhanced features include lane line upsampling features and drivable area upsampling features. Among them, the specific process of obtaining the lane line upsampling features and drivable area upsampling features includes: Perform a 3×1 convolution on the input feature map to obtain the first output feature; Perform a decomposition convolution on the image features in the backbone that have the same image feature resolution as the current decoder upsampling stage to obtain convolution features; Perform channel max pooling, channel average pooling, and activation processing on the convolution features to obtain the first spatial weight; Multiply the first spatial weight with the upsampled lane line feature map and add the results residually to obtain the lane line upsampled feature; Perform a 1×3 convolution on the first output feature to obtain the second output feature; Perform channel max pooling, channel average pooling, and activation processing on the second output feature to obtain the second spatial weight; Based on the second spatial weight, multiply it with the drivable area feature and add the results residually to obtain the drivable area upsampled feature.

5. A traffic road drivable area and lane line detection system, characterized in that, The system includes: An image acquisition module for acquiring an initial driving video image and dividing it into training images and images to be tested; A reparameterized depthwise separable downsampling module for replacing the downsampling layer in the existing TwinLiteNet model to obtain feature information with different receptive field sizes; A lightweight processing module for replacing the ordinary convolution layers in the backbone of the existing TwinLiteNet model with depthwise separable convolution layers for lightweight processing to obtain a lightweight backbone; A position attention module for performing feature enhancement processing on the input features to obtain enhanced features; A partial decomposition self-attention module placed at the end of the attention mechanism layer in TwinLiteNet for obtaining a simple global attention feature map. The process of obtaining the simple global attention feature map includes: performing resolution scaling processing on multiple feature maps with different resolution sizes in the feature pyramid to obtain multiple feature maps with the same resolution; sequentially performing channel concatenation and convolution aggregation feature processing on the multiple feature maps with the same resolution to obtain the first feature map; performing channel segmentation on the first feature map to obtain a second feature map and a third feature map, where the second feature map and the third feature map have the same number of channels; respectively performing horizontal self-attention and vertical self-attention calculations on the second feature map to obtain a horizontal feature map and a vertical feature map; performing channel concatenation processing on the horizontal feature map, the vertical feature map, the second feature map, and the third feature map to obtain the simple global attention feature map; An image decomposition and fusion module, which is used to obtain enhanced features. The specific process includes: replacing the downsampling layer in the existing TwinLiteNet model with a reparameterized depthwise separable downsampling layer to obtain the number of channels in each branch of the TwinLiteNet model; using depthwise separable reparameterized convolutional layers to replace the dilated convolutions in each branch of the downsampling layer, and inputting the input feature map into multiple depthwise separable reparameterized convolutional layers for convolutional processing according to the number of channels to obtain multiple receptive field features; adding a position attention layer behind each branch in the reparameterized depthwise separable downsampling layer to perform feature enhancement processing on the multiple receptive field features to obtain enhanced features output by multiple branches, where the number of depthwise separable reparameterized convolutional layers on each branch increases sequentially; A model training module, which is used to iteratively optimize and train the improved TwinLiteNet model using training images to obtain the improved TwinLiteNet model; A detection module, which is used to input the image to be detected into the improved TwinLiteNet model for detection to output the detection results of the drivable area and lane lines.

6. A computer device, characterized in that, The computer device includes a memory, a processor, and a processing program stored on the memory and executable on the processor. When the processing program is executed by the processor, it implements a method for detecting the drivable area and lane lines of a traffic road as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, A processing program is stored on the computer-readable storage medium. When the processing program is run by the processor, it implements a method for detecting the drivable area and lane lines of a traffic road as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Complex scene lane line detection method based on deep learning

    CN117935195A

  • Traffic target detection method and system based on improved YOLOv8n

    CN118552929A