Lane line detection system based on multi-task convolutional neural network

By adopting a multi-task convolutional neural network architecture in the lane line detection system, combined with technical means such as deep separable convolution and attention mechanism, the problem of insufficient performance of the existing lane line detection model in complex environments is solved, and efficient and accurate lane line detection and multi-task collaborative training are achieved.

CN120148004AInactive Publication Date: 2025-06-13SUZHOU FEIQUELAI TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510220595.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing deep learning lane line detection models are difficult to achieve efficient and accurate lane line detection in complex road environments, especially in the case of changes in light, severe weather and insufficient multitasking synergy.

Method used

The lane line detection system based on multi-task convolutional neural network is adopted, including a shared feature extraction layer, lane line positioning subnet, lane line type identification subnet and feasible area segmentation subnet. Through technical means such as deep separation convolution, hollow space pyramid pooling, attention mechanism and timing feature aggregation module, multi-task collaborative training and information sharing are realized.

Benefits of technology

It significantly improves the accuracy and adaptability of lane line detection, can maintain stable performance in complex environments, reduces computing resource requirements, and improves the system's real-time deployment capabilities.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent auxiliary driving, and discloses a lane line detection system based on a multi-task convolutional neural network, and the system comprises a multi-task convolutional neural network architecture which is composed of a shared feature extraction layer, a lane line positioning sub-network, a lane line type recognition sub-network and a drivable region segmentation sub-network. Through a cascade structure of depth separable convolution and void space pyramid pooling, the model calculation complexity is significantly reduced, and the expression ability of multi-scale road scene features is enhanced at the same time. Through the combination design of shallow grouping convolution and deep channel rearrangement dense connection, the feature fusion efficiency is improved on the basis of reducing the parameter quantity, and the problem that an existing model is difficult to deploy in real time due to large computing resource requirements is solved; a feature fusion module based on an attention mechanism is combined with an improved preprocessing technology, shadow, uneven illumination and night low-light interference are effectively suppressed, and the lane line positioning precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent assisted driving, and specifically to a lane line detection system based on a multi-task convolutional neural network. Background Art

[0002] At present, with the booming development of autonomous driving and intelligent assisted driving systems, lane line detection, as a key basic link for the safe driving of vehicles, the level of its technology directly affects the safety and reliability of driving.

[0003] Traditional lane line detection methods mainly rely on handcrafted features and machine learning algorithms. These methods rely on manually designed features, such as edge detection, color features, etc. When facing complex real-world road environments, there are obvious limitations. When there are drastic changes in lighting, such as suddenly entering a tunnel from a bright sunny day, or when driving at night, due to insufficient light, the brightness and contrast of the image will change significantly, which makes it difficult for traditional detection methods based on color and edges to accurately identify lane lines. Under adverse weather conditions, such as the reflection of the road surface and water accumulation on rainy days causing lane lines to be blurred, and lane lines being covered by snow on snowy days, the detection accuracy of traditional methods will drop significantly. Moreover, problems such as wear and stains that occur after long-term use of the road will also interfere with the accurate detection of lane lines by traditional methods.

[0004] In recent years, lane line detection technologies based on deep learning have gradually emerged. Although there has been a certain performance improvement compared to traditional methods, there are still many challenges. On the one hand, the existing deep learning model structures are often relatively complex, containing a large number of parameters and computational layers. This not only results in the model training and inference processes consuming huge computational resources, having extremely high requirements for hardware devices, increasing the cost and difficulty of practical applications; but also complex models are prone to overfitting, performing well on the training set, but having weak generalization ability and unstable detection accuracy in unseen real-world scenario data. On the other hand, the lack of multi-task collaboration is a major problem in current deep learning lane line detection technologies. Lane line detection usually involves multiple tasks, such as lane line localization, type recognition, and drivable area segmentation. However, many existing models are difficult to efficiently coordinate these tasks, lacking an effective information sharing and interaction mechanism between tasks, and unable to make full use of the correlation between different tasks, resulting in the overall detection performance being difficult to reach the optimal. For example, some models perform well in the lane line localization task, but perform poorly in the type recognition and drivable area segmentation tasks, and cannot meet the multi-faceted requirements of autonomous driving and intelligent assisted driving systems for lane line detection at the same time.

[0005] In summary, it is of great practical significance to develop an efficient, accurate and adaptable lane detection system, which not only helps to promote the further development of autonomous driving and intelligent assisted driving technologies, but also significantly improves the level of road traffic safety. The present invention is based on such a background to carry out research and innovation. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides a lane detection system based on a multi-task convolutional neural network, which solves the problems that many existing models are difficult to efficiently coordinate these tasks, there is a lack of effective information sharing and interaction mechanisms between tasks, and the relevance between different tasks cannot be fully utilized, resulting in the overall detection performance being difficult to reach the optimal level.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A lane detection system based on a multi-task convolutional neural network, comprising:

[0008] The multi-task convolutional neural network architecture: It is composed of a shared feature extraction layer, a lane line localization sub-network, a lane line type recognition sub-network and a drivable area segmentation sub-network;

[0009] The shared feature extraction layer: Adopts a cascaded structure of depthwise separable convolution and atrous spatial pyramid pooling module. A grouped convolution structure is used in the shallow network, and a densely connected structure with channel rearrangement is used in the deep network, effectively extracting multi-scale road scene features, reducing the amount of calculation and enhancing feature expression;

[0010] The lane line localization sub-network: Includes a feature fusion module based on an attention mechanism. Through a spatial attention unit (generating an attention weight map using deformable convolution), a channel attention unit (adopting a squeeze-and-excitation network structure to dynamically adjust the feature channel weights), and a cross-modal fusion unit (fusing visible light image and infrared image features to improve night detection ability), integrating semantic features and outputting a pixel-level lane line position heat map;

[0011] The lane line type recognition sub-network: Processes the continuous frame image features through a temporal feature aggregation module composed of an inter-frame motion compensation unit based on optical flow estimation, a temporal modeling unit constructed by a long short-term memory network (LSTM), and a dynamic key frame selection mechanism (adapting the processing frame rate according to the scene complexity), and combines the histogram of oriented gradients (HOG) and color space transformation for solid and dashed line classification;

[0012] The drivable area segmentation sub-network: Adopts an adaptive edge constraint loss function, fuses lidar point cloud projection data, and generates a high-precision road area mask.

[0013] Post - processing module: It includes a Bezier curve fitting algorithm based on curvature constraint to smooth the lane line position; a dynamic filtering algorithm based on vehicle motion state, including a lane line curvature tracking unit based on extended Kalman filter, a motion state compensation unit based on vehicle CAN - bus data, and an anomaly detection unit that identifies misdetected lane lines by constructing a Mahalanobis distance feature space to optimize the spatio - temporal consistency of the detection results.

[0014] Preferably, the multi - task convolutional neural network architecture adopts an asymmetric parameter sharing mechanism:

[0015] The shared feature extraction layer adopts a grouped convolution structure in the shallow network and a densely connected structure with channel rearrangement in the deep network;

[0016] Cross - task feature interaction is carried out between each sub - network through a learnable feature gating unit;

[0017] In the training stage, a multi - task loss function with dynamic weight adjustment is adopted to automatically adjust the optimization weight according to the gradient magnitude of each sub - task.

[0018] Preferably, the feature fusion module of the attention mechanism includes:

[0019] Spatial attention unit, which generates an attention weight map through deformable convolution;

[0020] Channel attention unit, which adopts a squeeze - excitation network structure to dynamically adjust the feature channel weights;

[0021] Cross - modal fusion unit, which fuses the features of visible light images and infrared images to enhance the robustness of night - scene detection.

[0022] Preferably, the temporal feature aggregation module includes:

[0023] An inter - frame motion compensation unit based on optical flow estimation;

[0024] A temporal modeling unit constructed by a long short - term memory network (LSTM);

[0025] A dynamic key - frame selection mechanism, which adaptively adjusts the processing frame rate according to the scene complexity.

[0026] Preferably, the lane line detection system further includes:

[0027] Pre - processing module, which performs illumination compensation using an improved Retinex theory and combines adaptive gamma correction to eliminate shadow interference;

[0028] Inverse perspective transformation module, which generates a three - dimensional space mapping relationship through a depth estimation network;

[0029] The embedded deployment optimization module uses neural architecture search (NAS) technology to generate a lightweight model adapted to the device.

[0030] Preferably, the dynamic filtering algorithm includes:

[0031] A lane line curvature tracking unit based on the extended Kalman filter;

[0032] A motion state compensation unit based on vehicle CAN bus data;

[0033] An anomaly detection unit that identifies misdetected lane lines by constructing a Mahalanobis distance feature space.

[0034] Preferably, the training process of the lane line detection system uses:

[0035] A progressive curriculum learning strategy to train the network in stages according to the scene complexity;

[0036] An adversarial sample generation mechanism to enhance data diversity through a generative adversarial network (GAN);

[0037] A knowledge distillation framework to compress a multi-task model into a single-task inference model.

[0038] The present invention provides a lane line detection system based on a multi-task convolutional neural network. It has the following beneficial effects:

[0039] 1. Through the cascaded structure of depthwise separable convolution and atrous spatial pyramid pooling, the present invention significantly reduces the model calculation complexity and at the same time enhances the expression ability of multi-scale road scene features. The combined design of shallow grouped convolution and deep channel rearrangement dense connection improves the feature fusion efficiency while reducing the number of parameters, solving the problem that existing models are difficult to be deployed in real time due to large computational resource requirements.

[0040] 2. The feature fusion module based on the attention mechanism of the present invention, combined with the improved preprocessing technology, effectively suppresses the interference of shadows, uneven illumination and low light at night, and improves the lane line positioning accuracy. The cross-modal fusion unit enables the system to maintain stable detection performance even at night or in bad weather through the feature fusion of visible light and infrared images.

[0041] 3. The temporal feature aggregation module of the present invention accurately captures the dynamic changes between consecutive frames of lane lines through optical flow estimation, LSTM temporal modeling and dynamic key frame selection mechanism, enhancing the system's adaptability to vehicle motion blur and scene mutations. The asymmetric parameter sharing mechanism and the multi-task loss function with dynamic weight adjustment optimize the collaborative training efficiency of lane line positioning, type recognition and drivable area segmentation, solving the problem of performance degradation caused by task conflicts in traditional multi-task models.

[0042] 4. The present invention fuses lidar point cloud data through an adaptive edge constraint loss function, significantly improving the geometric accuracy of the road area mask, especially in scenarios such as curves or blurred lane lines. The post-processing module is based on a dynamic tracking algorithm of Bezier curve fitting with curvature constraint and extended Kalman filter, effectively eliminating detection noise, ensuring the smoothness and spatio-temporal consistency of the lane line output, and reducing the false detection rate and missed detection rate.

[0043] 5. The present invention generates a lightweight model adapted to the device through a neural network architecture search, combines knowledge distillation and progressive curriculum learning strategies, significantly reducing the computational resource requirements while ensuring the detection accuracy, and supporting efficient deployment on multiple platforms. The adversarial sample generation mechanism enhances data diversity, significantly improving the generalization ability of the system in complex scenarios (such as rainy days and road wear). Detailed implementation manners

[0044] The technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] An embodiment of the present invention provides a lane line detection system based on a multi-task convolutional neural network, including:

[0046] Data acquisition and preparation:

[0047] Collect a large amount of road image data in different scenarios (including sunny days, rainy days, nights, urban roads, highways, etc.), and at the same time obtain the corresponding lidar point cloud data and vehicle CAN bus data. Label the image data, marking the positions, types (solid lines, dashed lines, etc.) of the lane lines and the drivable areas.

[0048] Preprocessing stage:

[0049] For the collected images, use the improved Retinex theory for illumination compensation, adjust the brightness and contrast according to the illumination distribution of the images, and enhance the image details.

[0050] Then adopt adaptive gamma correction to further eliminate shadow interference, make the image illumination uniform, and provide high-quality images for subsequent neural network processing.

[0051] Network training stage:

[0052] Training of the shared feature extraction layer: Input the preprocessed images into the shared feature extraction layer. First, use a shallow grouped convolution structure to initially extract features and reduce the computational amount; then, through a deep channel rearrangement dense connection structure, strengthen feature fusion and transmission, and extract multi-scale road scene features.

[0053] Lane line localization sub - network training: The features output by the shared feature extraction layer are input into this sub - network. The spatial attention unit generates an attention weight map through deformable convolution, focusing on the lane line area; the channel attention unit dynamically adjusts the feature channel weights using a squeeze - excitation network structure; if there is infrared image data, the cross - modal fusion unit fuses it with the visible light image features. Taking the labeled lane line positions as a reference, it is trained through an appropriate loss function (such as the mean square error loss function) to make the network output an accurate pixel - level lane line position heat map.

[0054] Lane line type recognition sub - network training: The features of consecutive frame images are input into this sub - network. The inter - frame motion compensation unit based on optical flow estimation compensates for inter - frame motion, reducing the influence of motion blur; the temporal modeling unit constructed by LSTM learns temporal relationships to capture the dynamic changes of lane lines; the dynamic key - frame selection mechanism adjusts the processing frame rate according to the scene complexity. Combining HOG and color space transformation to extract lane line features, using the labeled lane line types as labels, and training with a cross - entropy loss function, etc., to achieve the classification of solid and dashed lines.

[0055] Drivable area segmentation sub - network training: The features output by the shared feature extraction layer are fused with the lidar point cloud projection data and input into this sub - network. Using an adaptive edge - constraint loss function, with the labeled drivable area as the target, training the network to generate a high - precision road area mask.

[0056] Multi - task joint training: Using an asymmetric parameter sharing mechanism, cross - task feature interaction between sub - networks is achieved through a learnable feature gating unit. Adopting a multi - task loss function with dynamic weight adjustment, automatically adjusting the optimization weights according to the gradient magnitudes of each sub - task. At the same time, following a progressive curriculum learning strategy, training the network in stages from simple scenarios to complex scenarios; using GAN to generate adversarial samples to enhance data diversity; applying a knowledge distillation framework to compress the multi - task model into a single - task inference model to improve the training effect and model performance.

[0057] Post - processing stage:

[0058] Bézier curve fitting based on curvature constraint: Processing the heat map output by the lane line localization sub - network, through a Bézier curve fitting algorithm based on curvature constraint, fitting the discrete lane line position points into a smooth and continuous curve to improve the accuracy and aesthetics of lane line detection.

[0059] Dynamic filtering based on vehicle motion state: The lane line curvature tracking unit based on the extended Kalman filter tracks the changes in lane line curvature; the motion state compensation unit based on vehicle CAN - bus data compensates according to the vehicle motion state; the anomaly detection unit identifies and eliminates misdetected lane lines by constructing a Mahalanobis distance feature space to optimize the spatio - temporal consistency of the detection results.

[0060] Inverse perspective transformation and deployment phase:

[0061] Inverse perspective transformation: Generate a three-dimensional space mapping relationship through a depth estimation network, perform inverse perspective transformation on the detected lane lines and drivable areas, convert them from a two-dimensional plane to a three-dimensional space, and provide more accurate information for subsequent path planning and vehicle control.

[0062] Embedded deployment optimization: Adopt neural network architecture search (NAS) technology to generate an adapted lightweight model according to the resource limitations (such as computing power, storage capacity, etc.) of the target embedded device. Perform optimization operations such as quantization and pruning on the lightweight model to reduce the computational amount and storage requirements and achieve efficient embedded deployment.

[0063] Through the above implementation manners, the lane line detection system of the present invention can accurately and real-time detect lane lines and drivable areas in a complex and changeable road environment, and provide reliable support for autonomous driving and intelligent assisted driving. The present invention meets the requirements of the Patent Law, has novelty, creativity and practicability, and completely covers aspects such as the technical solution, implementation steps and application effects of the invention.

[0064] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A lane detection system based on a multi-task convolutional neural network, characterized in that: include: A multi-task convolutional neural network architecture consisting of a shared feature extraction layer, a lane localization subnetwork, a lane type recognition subnetwork, and a drivable area segmentation subnetwork; The shared feature extraction layer adopts a cascade structure of depth-separable convolution and dilated spatial pyramid pooling modules to extract multi-scale road scene features; The lane positioning subnetwork includes a feature fusion module based on an attention mechanism, which is used to integrate semantic features at different levels and output a pixel-level lane position heat map; The lane line type recognition subnetwork processes the continuous frame image features through the temporal feature aggregation module, and classifies the virtual and real lines by combining the directional gradient histogram and color space transformation; The drivable area segmentation subnetwork adopts an adaptive edge constraint loss function to generate a high-precision road area mask by fusing the laser radar point cloud projection data; The post-processing module, including a Bezier curve fitting algorithm based on curvature constraints and a dynamic filtering algorithm based on the vehicle motion state, is used to optimize the spatiotemporal consistency of the detection results.

2. The lane detection system based on a multi-task convolutional neural network according to claim 1, characterized in that: The multi-task convolutional neural network architecture adopts an asymmetric parameter sharing mechanism: The shared feature extraction layer adopts a grouped convolution structure in the shallow network and a dense connection structure with channel rearrangement in the deep network; Each sub-network interacts with other tasks through learnable feature gating units. During the training phase, a multi-task loss function with dynamic weight adjustment is used to automatically adjust the optimization weights according to the gradient amplitude of each subtask.

3. The lane line detection system based on a multi-task convolutional neural network according to claim 2, characterized in that: The feature fusion module of the attention mechanism includes: Spatial attention unit, which generates attention weight map through deformable convolution; Channel attention unit, which uses compression-excitation network structure to dynamically adjust feature channel weights; The cross-modal fusion unit fuses the features of visible light images and infrared images to enhance the robustness of night scene detection.

4. The lane detection system based on a multi-task convolutional neural network according to claim 1, characterized in that: The time series feature aggregation module includes: Inter-frame motion compensation unit based on optical flow estimation; Temporal modeling units constructed from long short-term memory networks; Dynamic keyframe selection mechanism, adaptively adjusting the processing frame rate according to the complexity of the scene.

5. The lane detection system based on a multi-task convolutional neural network according to claim 1, characterized in that: The lane line detection system also includes: The pre-processing module uses the improved Retinex theory to perform illumination compensation and combines it with adaptive gamma correction to eliminate shadow interference; The inverse perspective transformation module generates a three-dimensional spatial mapping relationship through a depth estimation network; The embedded deployment optimization module uses neural network architecture search technology to generate lightweight models adapted to devices.

6. The lane detection system based on a multi-task convolutional neural network according to claim 1, characterized in that: The dynamic filtering algorithm comprises: Lane curvature tracking unit based on extended Kalman filter; Motion state compensation unit based on vehicle CAN bus data; The anomaly detection unit identifies misdetected lane lines by constructing the Mahalanobis distance feature space.

7. The lane detection system based on a multi-task convolutional neural network according to claim 1, characterized in that: The training process of the lane detection system adopts: Progressive curriculum learning strategy, training the network in stages according to the complexity of the scene; Adversarial sample generation mechanism, enhancing data diversity by generating adversarial networks; Knowledge distillation framework, compresses multi-task models into single-task reasoning models.

Citation Information

Cited By

  • Simulation test system and method for visual perception module of automatic driving system

    CN121386456A