Lane line detection method based on multi-scale mixed feature-Transform

Through the end-to-end model MHFS-FORMER and ShuffleLaneNet modules based on Transformer, the shortcomings of traditional CNN lane line detection models in long-distance dependency capture and complex morphological processing are solved, high-precision and low-power lane line detection are realized, and the environmental perception capability of the autonomous driving system is improved.

CN120356170APending Publication Date: 2025-07-22KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510411470.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The traditional CNN lane line detection model lacks long-distance dependency capture capabilities, is difficult to deal with complex and changeable lane line patterns, and has poor end-to-end performance, which limits the development of lane line detection technology.

Method used

The end-to-end model MHFS-FORMER based on Transformer is adopted, combining multi-scale features with Transformer Encoder, and a multi-reference deformable attention module is introduced, lane line parameters are decoded through a multi-layer decoder, and a ShuffleLaneNet module is designed for feature enhancement and feature expression optimization.

Benefits of technology

It significantly improves the global perception capability and detection accuracy in complex scenarios, improves the reliability and real-time nature of lane line detection, and provides high-precision and low-power environmental perception solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356170A_ABST
    Figure CN120356170A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of lane line detection, in particular to a lane line detection method based on multi-scale mixed feature-Transform, which comprises the following steps of: S1, preprocessing an input image, mapping a ground plane point into an aerial view coordinate, and compensating a tilt influence of a camera; s2, multi-scale feature extraction: extracting features by adopting a ResNet backbone network; s3, performing multi-scale feature fusion, performing cascade fusion on the features of adjacent levels through a multi-scale feature fusion network, and generating enhanced mixed features by combining and using adaptive space fusion; s4, enhancing mixed lane features, and carrying out dual-channel grouping attention calculation and channel shuffling operation; s5, decoding lane line parameters through a multi-reference deformable attention mechanism of a multi-layer decoder; and S6, loss calculation. The problems that a traditional CNN lane line detection model is insufficient in long-distance dependence capture, complex forms are difficult to process, and end-to-end performance is poor are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lane line detection, and in particular to a lane line detection method based on multi-scale hybrid features-Transformer. Background Art

[0002] As intelligent transportation is developing rapidly, lane detection is a key part of autonomous driving and assisted driving systems. Its accuracy and real-time performance play a decisive role in the safety and efficiency of road travel. For a long time, models based on convolutional neural networks (CNNs) have dominated the field of lane detection. CNNs have shown certain performance advantages in many scenarios due to their powerful local feature extraction capabilities.

[0003] However, as the complexity of application scenarios continues to increase, traditional CNN models have gradually exposed many limitations. They are unable to capture long-distance dependencies, are unable to handle complex and changeable lane line shapes, and the end-to-end performance of the model is also poor. These shortcomings have become a serious obstacle to the further development of lane line detection technology.

[0004] Inspired by the DETR architecture, we innovatively proposed a Transformer-based end-to-end model, MHFS-FORMER, to overcome the above difficulties. In order to effectively deal with the interference of lane lines in complex scenarios, we carefully designed MHFNet, which fuses multi-scale features with Transformer Encoderder to obtain enhanced multi-scale features. Subsequently, these enhanced multi-scale features are input into Transformer Decoder, and a new multi-reference deformable attention module is introduced. This module can distract attention to the surrounding of the object, greatly enhancing the representation ability of the model during training, so as to better capture the slender structure of the lane and the full environment information.

[0005] In addition, we designed ShuffleLaneNet, which deeply and meticulously mines the channel and spatial information of multi-scale lane features, significantly improving the accuracy of object recognition. Summary of the invention

[0006] In order to solve the above problems, the present invention provides a lane line detection method based on multi-scale hybrid features-Transformer, which solves the problems of traditional CNN lane line detection model, such as insufficient capture of long-distance dependence, difficulty in handling complex forms and poor end-to-end performance.

[0007] The technical implementation scheme of the present invention is: a lane detection method based on multi-scale hybrid features-Transformer, comprising the following steps:

[0008] S1. Input image preprocessing: Map the ground plane points to bird's-eye view coordinates through the camera parameter conversion formula to compensate for the influence of camera tilt.

[0009] S2. Multi-scale feature extraction: Use the ResNet backbone network to extract {C3, C4, C5} features, and the number of channels is unified to 256.

[0010] S3. Multi-scale feature fusion: Cascade and fuse the features of adjacent levels through the multi-scale feature fusion network (MHFNet), and generate enhanced hybrid features by combining adaptive spatial fusion (ASF).

[0011] S4. Hybrid lane feature enhancement: Use ShuffleLaneNet for dual-channel grouped attention calculation and channel shuffle operations.

[0012] S5. Transformer processing: Decode the lane line parameters through the multi-reference deformable attention mechanism of the multi-layer Transformer decoder.

[0013] S6. Loss calculation: Optimize the model parameters based on the bipartite matching loss of the Hungarian algorithm.

[0014] Optionally, the lane line modeling in the input image preprocessing of step S1 includes:

[0015] (a) Define the lane shape prior model as a polynomial structure:

[0016]

[0017] Among them, (X, Y) represents the points of the lane line on the ground plane, N is the polynomial order, a n and b are coefficients. For a single-lane line on a flat ground, it is represented by a cubic curve as:

[0018] X = kZ 3 + mZ 2 + nZ + b

[0019] Among them, k, m, n, and b are real parameters, k ≠ 0, and (X, Z) indicates the points on the ground plane;

[0020] (b) To avoid interference from uneven road surfaces, convert the cubic polynomial to the bird's-eye view form, and the conversion formula is:

[0021]

[0022] Among them, (X, Y) are the pixel points on the converted phase plane, f x is the pixel width on the focal plane divided by the focal length, f yis the focal plane pixel height divided by the focal length, and H is the height at which the camera is mounted;

[0023] (c) Considering the camera tilt angle θ, the non-tilted and tilted transformation formulas are:

[0024]

[0025] where f is the camera focal length, f′ is the camera focal length after pitch transformation, and (x′, y′) are the pixel points corresponding to those after pitch transformation;

[0026] (d) Combining the above formulas, the converted expression is obtained:

[0027] where

[0028] m′ = m cos 2 θ × H / f x × f y , n′ = n / f x , b′ = b × f y / fxcosθ × H,

[0029] b″ = b × f y / ftanθ / f x × H;

[0030] (e) The lane line output parameters are parameterized as: p i = (c, k′, m′, n′, f″, b′, b″), where i ∈ {0, 1,..., M}, and M is the total number of lane lines.

[0031] Optionally, the multi-scale feature fusion network (MHFNet) in the step S3 includes:

[0032] (a) Perform 1×1 convolutions and dimension alignment on the {C3, C4, C5} features respectively to generate {S3, S4, S5}; (b) Iteratively fuse S4 and F5 through a cascaded fusion module to generate {P4, P5}, and the cascaded fusion module contains two 1×1 convolutional layers, N RepBlocks, and an element-wise addition operation; (c) Use adaptive spatial fusion (ASF) to perform dynamic weight allocation on {S3, P4, P5}, and its calculation formula is:

[0033] where

[0034] Optionally, the ShuffleLaneNet in the step S4 includes:

[0035] (a) Divide the input features into two groups along the channels, and perform spatial and channel attention fusion on the first group;

[0036] (b) The second group is further divided into two subgroups to calculate spatial attention and channel attention respectively;

[0037] (c) Channel attention is achieved through global average pooling and Sigmoid gating, and the calculation formula is:

[0038]

[0039] The calculation formula of Sigmoid gating is: X k ′ = s(F c (c)) · X k = s(W1s + b1) · X k ;

[0040] (d) Spatial attention is calculated through group normalization (GN) and activation function: X m ′ = σ(W2 · GN(X m )) + b2) · Xm;

[0041] (e) Cross-group information interaction is achieved through channel shuffle operation.

[0042] Optionally, the calculation formula of the multi-reference deformable attention module in step S5 is:

[0043]

[0044] Among them, is the normalized reference point coordinate, Δp mlqk is the sampling offset, A mlpk is the normalized weight.

[0045] Optionally, the loss calculation in step S6 specifically includes:

[0046] (a) Define the parameter set, and the calculation formula is:

[0047] Among them, M is the number of predicted lane lines,

[0048] The calculation formula of the marker set representation is:

[0049] (b) Through the true lane line search function l: S → P, the matching between the optimal injective parameter set of the predicted lane line and the true lane line is transformed into a minimum cost problem, and is calculated by the following formula:

[0050] Among them, L bmc is the matching cost between the i-th true lane line and the predicted parameter set;

[0051] (c) The regression loss function is defined as: where g(c i ) is the probability of class c i , is the fitted lane line sequence, μ1 and μ2 are the loss function coefficients, L mae is the mean absolute error, and Z(·) represents the indicator function.

[0052] Optionally, the input image preprocessing in the S1 step includes: randomly scaling the image and performing color jitter processing. During the random scaling of the image, the scaling ratio is 0.5 - 2.0; during the color jitter processing, first adjust the brightness, then randomly crop according to the size of 640×360 or 820×295, and finally perform a horizontal flip operation with a probability of 0.5 to enhance the generalization ability through multi-dimensional data augmentation.

[0053] Optionally, the Transformer decoder in the S5 step includes a processing structure composed of 6 layers of multi-head self-attention modules, receives multi-scale feature inputs {S3, S4, F5}, and simultaneously realizes global context modeling by learning lane embedding vectors to strengthen the global correlation analysis and perception ability of lane line features.

[0054] Optionally, the ResNet backbone network outputs features {C3, C4, C5} in the S2 step specifically as follows: the C3 feature has a size of H / 8×W / 8 and 256 channels, the C4 feature has a size of H / 16×W / 16 and 256 channels, and the C5 feature has a size of H / 32×W / 32 and 256 channels, providing multi-dimensional information for subsequent processing by outputting different-scale features in layers.

[0055] The present invention has the following advantages:

[0056] Through innovative multi-module collaborative design, this technical solution effectively solves the core problems existing in traditional lane line detection technologies: constructing a full-length long-distance dependence model using the Transformer architecture to break through the limitations of local modeling of convolutional neural networks and significantly improve the global perception ability in complex scenarios; designing a multi-scale feature fusion module MHFNet to achieve efficient integration of lane line detail features and global semantics and enhance adaptability to complex conditions such as lighting changes and occlusions; introducing a multi-reference deformable attention mechanism to dynamically focus on key areas of lane lines and significantly improve the detection accuracy of slender structures and complex shapes; developing the ShuffleLaneNet module to optimize feature representation and enhance feature discriminability through joint channel and spatial recombination. On the premise of ensuring real-time performance, this solution greatly improves the reliability of lane line detection in multiple scenarios, provides a high-precision and low-power environment perception solution for autonomous driving systems, and effectively solves technical problems such as weak long-distance detection ability, poor adaptability to complex scenarios, and insufficient real-time performance of traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a framework diagram of the lane detection network of the present invention;

[0058] Figure 2 It is a structural diagram of the multi-scale feature fusion network MHFNet of the present invention;

[0059] Figure 3 It is a structural diagram of the Fusion network of the present invention;

[0060] Figure 4 It is a structural diagram of the hybrid lane network of the present invention;

[0061] Figure 5 It is a visualization result diagram of the MHFS-Former of the present invention on the TuSimple dataset;

[0062] Figure 6 It is a visualization result diagram of the MHFS-Former of the present invention on the TuSimple dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. It is hereby declared that the azimuth terms such as up, down, left, right, front, back, inside, and outside that appear or will appear in the text of the present invention are only based on the accompanying drawings of the present invention, and they do not specifically limit the present invention.

[0064] Example 1:

[0065] A method for lane line detection based on multi-scale hybrid features-Transformer includes the following steps:

[0066] S1. Input image preprocessing: Map ground plane points to bird's-eye view coordinates through the camera parameter conversion formula to compensate for the influence of camera tilt.

[0067] S2. Multi-scale feature extraction: Use the ResNet backbone network to extract {C3, C4, C5} features, and the number of channels is unified to 256.

[0068] S3. Multi-scale feature fusion: Cascade and fuse adjacent hierarchical features through the multi-scale feature fusion network (MHFNet), and generate enhanced hybrid features by combining adaptive spatial fusion (ASF).

[0069] S4. Hybrid lane feature enhancement: Use ShuffleLaneNet for two-channel grouped attention calculation and channel shuffle operations.

[0070] S5. Transformer processing: Decode lane line parameters through the multi-reference deformable attention mechanism of multiple layers of Transformer decoders.

[0071] S6. Loss calculation: Optimize the model parameters based on the bipartite matching loss of the Hungarian algorithm.

[0072] It should be noted that after the input image is preprocessed in this technical solution, different hierarchical features are generated through the multi-scale feature extraction module; the MHFNet module realizes cross-scale feature fusion and enhancement; the multi-reference deformable attention module dynamically focuses on the key areas of the lane lines; ShuffleLaneNet performs feature recombination optimization; the Transformer decoder generates lane line coordinates; the post-processing module completes curve fitting and result smoothing. Each module works together to achieve high-precision lane line detection in complex scenarios.

[0073] Example 2:

[0074] In the calculation of the lane line parameter model, inspired by the DETR architecture, we propose an end-to-end model MHFS-FORMER based on Transformer to solve these problems. Among them, the prior model of the lane shape in MHFS-FORMER is defined as a polynomial structure on the road:

[0075]

[0076] where (X, Y) represents the points of the lane line on the ground, N is the order of the polynomial X, a n and b are the coefficients of the lane polynomial function. In the actual driving scenario, the shape of the lane line is generally not too complex. In view of this, it is a more appropriate choice to use a cubic curve to approximate the single lane line on the flat ground. Therefore, it can be expressed as:

[0077] X=kZ 3 +mZ 2 +nZ+b Formula 2

[0078] Wherein, k, m, n and b are real number parameters, k≠0, and (X, Z) indicates a point on the ground plane.

[0079] It should be noted that in real driving environments, road conditions are complex and the road surface is often bumpy. This bumpy condition will interfere with the polynomial model and affect its accurate expression and analysis of road elements such as lane lines. In order to avoid the adverse effects of uneven roads, we consider converting the cubic polynomial applicable to flat roads into a bird's-eye view form for representation. In this way, the interference caused by the ups and downs of the road surface can be effectively reduced, and the accuracy and stability of the model can be improved.

[0080] The conversion formula is as follows:

[0081]

[0082] Among them, (X, Y) is the pixel point on the phase plane after transformation, and f x The focal plane is often calculated by dividing the width of the pixels on the surface by the focal length, f y is the height of the pixel on the focal plane divided by the focal length, H is the height of the camera, considering that there are bumps during driving, which makes the installed camera tilted, assuming the tilt angle is θ, the transformation between non-tilted and tilted is as follows:

[0083]

[0084] Among them, f is the relative focal length of the camera. f' is the relative focal length of the camera after pitch transformation, (x', y') is the corresponding pixel point after pitch transformation, and combining formula 3 and formula 4, we can get:

[0085]

[0086] Where, f″=fsinθ, m′=m cos 2 θ×H / f x ×f y , n′=n / f x , b′=b×f y / f x cosθ×H,b″=b×f y ×f tanθ / f x×H. Additionally, the model also considers the confidence c of the lane line, where c ∈ [0, 1], with 0 being the background and 1 being the lane line marking.

[0087] In summary, the output of the i-th lane line is parameterized by the following calculation formula:

[0088] p i = (c, k′, m′, n′, f″, b′, b″) Formula 6

[0089] where i ∈ {0, 1,..., M}, and M is the total number of lane lines.

[0090] Embodiment 3:

[0091] In the overall network architecture process, multi-scale feature fusion includes the following steps:

[0092] Step 1, construction of MHFS-FORMER:

[0093] Based on the geometric prior features of the lane and the global information of the image, MHFS-FORMER is designed. By enhancing the multi-scale features, the recognition effect of the Transformer model on slender structures is strengthened, and at the same time, the computational complexity is reduced. As shown in the appendix, it includes a CNN backbone for extracting feature maps from the input image, a hybrid lane network, a multi-scale feature fusion network, a lane detection head for detecting lane lines through Transformer deformable convolution, and a bipartite matching loss for model training. Figure 1 As shown, it includes a CNN backbone for extracting feature maps from the input image, a hybrid lane network, a multi-scale feature fusion network, a lane detection head for detecting lane lines through Transformer deformable convolution, and a bipartite matching loss for model training.

[0094] Step 2, multi-scale feature fusion network MHFNet:

[0095] We introduce a multi-scale feature extraction and fusion network into the Transformer model, performing cascading fusion on the multi-level features extracted by the backbone network, and combining the low-level features in the early stage to guide the learning of high-level features in subsequent stages, thereby enhancing the effectiveness of multi-scale feature fusion. However, current popular vision Transformer methods usually use single-scale feature detection to reduce the excessive computational cost, or directly introduce multi-scale features into the Transformer model by reducing the computational complexity of the self-attention mechanism. Since these methods allocate most of the computational resources to the Transformer module, the feature fusion in these networks occurs relatively late, which may lead to insufficient feature fusion. Inspired by the existing literature reports, we propose a multi-scale fusion feature network rich in semantic information, called MHFNet. It starts by fusing two adjacent high-level features, combines multiple cascading stages to iteratively extract and fuse multi-scale features, and gradually adds low-level features to the fusion process. This makes the semantic information of different-level features closer, enhances the entire feature hierarchy with the rich semantic information of the low level, and achieves a lightweight module design. Given that information conflicts are likely to occur during the feature fusion process at each spatial position, an adaptive spatial fusion operation is further introduced to alleviate such inconsistency problems.

[0096] As Figure 2 shown, it is the overall architecture of MHFNet. The features output by the last residual block in each stage are used as the multi-scale inputs of the Transformer module. For the outputs of conv 3, conv 4, and conv 5, we denote the outputs of these last residual blocks as {C3, C4, C5}, and note that they have strides of {8, 16, 32} pixels relative to the input image, and unify their number of channels to 256 through a 1×1 convolutional layer (stride of 1). Since the memory occupancy of conv 1 and conv2 is very large, we do not include them in the pyramid. This set of final feature maps is called {S3, S4, S5}, corresponding to spatial sizes of H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32 respectively, and the F5 obtained through single-scale encoder calculation, along with S3 and S4, are used as the input representations of MHFNet, denoted as {S3, S4, F5}.

[0097] Since the encoder-decoder of the Transformer module uses features with a unified number of channels, we fix the feature dimension in all feature maps to d. In this technical solution, we set d = 256. Therefore, all additional convolutional layers have an output of 256 channels. First, we fuse two adjacent high-level features S4 and F5 through the Fusion module to obtain the fused features P4 and P5 at their respective scales. The fusion block contains two 1x1 convolutions to adjust the number of channels and uses N RepBlocks composed of RepConv for feature fusion. RepBlocks are residual blocks, that is, modules that implement residual learning through skip connections;

[0098] Secondly, the two-way outputs are fused by element-wise addition. One part of the features extracts semantic information through multiple convolutional layers, and the other part of the features directly retains the original information;

[0099] Finally, these two parts of the features are fused by addition, so that the fused features contain both the abstracted high-level features and the original low-level detailed features. Its structure is as Figure 4 shown. In addition, to achieve dimension alignment and lay the foundation for subsequent feature fusion work, we use 1×1 convolution combined with bilinear interpolation to perform upsampling on the features. In another dimension, for different downsampling requirements, we will flexibly select appropriate convolution kernels and strides to perform downsampling. Specifically, when 2-fold downsampling is required, a 2×2 convolution with a stride of 2 is used; if 4-fold downsampling is to be achieved, a 4×4 convolution with a stride of 4 is used.

[0100] After that, we will fuse the features at the three levels of S3, P4, and P5. We use the ASF-related literature to assign different spatial weights to features at different levels during the multi-level feature fusion process, strengthen the expression of key information, reduce the interference between features of different objects, and improve the fusion effect. We fuse the features at three levels. Let represent the feature vector at position (i, j) from level n to level l. The obtained feature vector is denoted as obtained through the adaptive spatial fusion of multi-level features and is defined by the linear combination of the feature vectors and as follows:

[0101]

[0102] It should be noted that and represent the spatial weights of the three-level features at level l and need to satisfy the constraint condition Finally, through the concat operation, we obtain the mixed multi-scale features {P3`, P4`, P5`}, whose feature maps are feature maps with scales of H / 16×W / 16, H / 32×W / 32, and H / 64×W / 64 respectively, and the number of channels is 256.

[0103] Step 3, the mixed lane network ShuffleLaneNet:

[0104] Due to the narrow lane structure, there is relatively little useful information in the image. To enhance the accuracy of lane localization, we propose a mixed lane network shuffleLaneNet used between the backbone and the Transformer decoder. As Figure 3 shown, this mechanism fully considers the information flow between the feature channels and spatial positions, effectively fuses the interaction between two types of features, not only maintains the lightweight of the model, but also significantly improves the performance of the model. Finally, the ShuffleNet unit is used to achieve cross-group information flow along the channel dimension, and the channel and spatial information in the lane features are carefully mined to enhance important information and provide more valuable information for subsequent modules.

[0105] This algorithm takes the feature map as the input, where C, H, and W represent the number of channels, the height of the space, and the width respectively. shuffleLaneNet first divides the input feature X into two groups along the channel dimension. The first group performs spatial attention calculation and channel attention fusion calculation on the feature map, and its feature The second group is further divided into two groups along the channel dimension. One group focuses on spatial attention calculation, and its feature The other group focuses on channel attention calculation, and its feature Finally, they are merged into the original dimension according to the number of channels, and the channel shuffle operation is used to suppress noise. In channel attention, we compress the global spatial information into the channel descriptor by using global average pooling to generate channel statistics For X1, n is 2, and for X3, n is 4. The form can be calculated by shrinking X in the spatial dimension H×W:

[0106]

[0107] In addition, a compact function is created to guide precise and adaptive selection. This is achieved through a simple gating mechanism with Sigmoid activation. Then, the final output of channel attention can be obtained as follows:

[0108] X′ k = s(Fc(c))·X k = s(W1s + b1)·X kFormula 9

[0109] Among them, and are parameters for scaling and moving s.

[0110] Furthermore, by utilizing the spatial relationship of features to generate spatial attention, what steps are adopted to calculate the spatial attention:

[0111] First: Perform channel pooling on the feature map. We use group normalization (GN, i.e., GroupNormalization in the literature) on X1 and X2 to obtain spatial statistics. The n of X1 is 2, and the n of X2 is 4. Then activate the obtained vector to obtain the attention weights in the spatial dimension, and use F c (·) to enhance the representation of X m . This method is used to associate long-term dependencies. The final output of the spatial attention is obtained as follows:

[0112] X m ′ = σ(W2·GN(X m )) + b2)·X m Formula 10

[0113] Among them, and W2 is a matrix with a dimension of , and b2 is a matrix with a dimension of .

[0114] It should be noted that is a learnable weight matrix and belongs to the parameters of the convolution operation. Among them, C represents the total number of channels of the feature map, n represents the number of groups (such as the grouping of group normalization GN), and 1×1 represents the convolution kernel size. Its role is to add the bias b2 on the basis of passing through group normalization GN(X m ), so that the model can learn more flexible feature expressions, is the bias term and also participates in the linear transformation process.

[0115] Secondly, aggregate all sub-features. First, connect X2′ and X3′, and finally connect them with X1′. We use the ShuffleNet unit to achieve cross-group information flow along the channel dimension. The final output of ShuffleLaneNet is the same size as X.

[0116] Example 4:

[0117] In the design of the Transformer detection module, it includes the following steps:

[0118] Step 1, as shown in the appendix Figure 1The shown Transformer architecture consists of a simplified Transformer encoder, MHFNet, a Transformer decoder using a deformable self-attention mechanism, several feed-forward networks (FFNs) for parameter prediction, and a Hungarian loss.

[0119] The encoder is composed of a single-layer Transformer encoder structure, which consists of a multi-head self-attention module and a feed-forward network (FFN). It takes the output of the last layer of ResNet18, which is reshaped to 256 channels after channel rearrangement, and the feature map S5 with a size of H / 32×W / 32. It encodes position information through absolute-position-based sine embedding and calculates it with the self-attention mechanism. The self-attention mechanism is as follows:

[0120]

[0121] where Q, K, and V represent sequences of queries, keys, and values linearly transformed through each input row, A represents the attention map that measures non-local interactions to capture elongated structures and global context, and O represents the output of self-attention.

[0122] Step 2: Obtain the output F5 by following the FFN and connecting it to the residual layer with layer normalization.

[0123] Step 3: Calculate the loss through the loss function.

[0124] In Step 1, the decoder is composed of a multi-scale n-layer Transformer decoder structure, using a multi-head self-attention module, a multi-scale deformable attention module, and a feed-forward layer. It uses the mixed multi-scale features of MHFNet {S3, S4, F5} as another attention module inserted into each layer, and their feature maps are feature maps with scales of H / 16×W / 16, H / 32×W / 32, and H / 64×W / 64 respectively, with 256 channels. The input of the decoder is set as an empty N×C matrix of the encoder, and all curves are directly decoded for parameters at once. In addition, we also introduce a learned lane embedding vector of size N×C as an implicit learning of global lane information for position global lane information embedding. The attention mechanism uses the same formula 11 for global lane information as in the encoder, and sequentially obtains a decoded sequence of shape N×C for global lane information, which is calculated with the multi-scale deformable attention module for the mixed multi-scale features of MHFNet output for global lane information. The formula for the multi-scale deformable attention module is:

[0125]

[0126] It should be noted that is the input feature multi-scale map, where Let is the normalized reference point coordinate of each query element q, where m represents the head, L represents the input feature attention level, k represents the point, and Δp mlqk and A mlqk denote the lth feature sampling level and the mth attention head.

[0127] The sampling offset and weight of k sampling points, scalar attention weight A mlqk Depend on Normalize and use normalized coordinates The coordinates (0, 0) and (1, 1) represent the upper left corner and lower right corner of the normalized image, respectively.

[0128] In step 2, the function in formula 12 Normalize The input feature maps rescaled to L coordinates are then connected through the following FFN, residual hierarchy with layer normalization, and then they are independently decoded by (FFNs) into feed-forward network lane line parameters to obtain N final predictions and class labels.

[0129] It should be noted that the FFNs prediction module uses a collection of three parts to generate the prediction curve. The single decoder output is directly projected into N×2 linear operations, and then the softmax layer operates on it to obtain the last dimension c of the predicted label. i , i∈{1,...,N}, meanwhile, for background or lane, one with ReLU activation and hidden dimension C projects the output of the decoder into N×4, where dimension 4 represents four sets of parameters specific to the lane of the 3-layer perceptron, and the other first projects the features into N x 4 and then averages them in the first dimension of the 3-layer perceptron, resulting in four shared parameters.

[0130] In step 3, the model is trained to calculate the matching degree between the predicted main difficulty parameters and the real lane line. The proposed Hungarian method is used for fitting, and a two-part matching is performed between the predicted loss parameters and the real lane line. The matching result lane line is used to optimize the regression loss of the specific lane. First, the parameter set is defined as:

[0131]

[0132] Where M is defined as the number of lanes larger than the lane lines predicted in common cases, and the label set is expressed as:

[0133]

[0134] Then, through the real lane line search function: l:S→P, the predicted optimal single-shot parameter set of the l lane line and the real one are defined as a problem with minimum cost:

[0135]

[0136] where l bmc is the matching cost between the prediction parameter set with index li of the bipartite matching problem between the i-th true lane line and the lane line marker set, which matches the class prediction and also considers the similarity between the prediction parameter set and the pre-cost.

[0137] Finally, the regression loss function for the true lane line is defined as:

[0138]

[0139] where g(c i ) is the probability of class c i , is the fitted lane line sequence μ1, μ2 is the coefficient of the loss function, L mae is the mean absolute error, and Z(·) represents the indicator function.

[0140] Example 5:

[0141] Experimental verification includes the following steps:

[0142] First, to widely evaluate the proposed method, we conducted experiments on two representative lane detection benchmarks: the TuSimple dataset and the CULane dataset for trial verification. Among them, the TuSimple dataset is an autonomous driving dataset consisting of 6408 annotated images, which are photos recorded by a high-resolution (720x1280) front-view camera under different road conditions and weather conditions day and night on US highways, specifically focusing on real highway scenarios, and consists of 3268 images for training, 358 images for validation, and 2782 images for testing, and the size of all images is 1280x720 pixels; CULane is a large open-source dataset collected by in-vehicle cameras in Beijing, China, from urban and highway scenarios. It consists of 88,880 training images, 9675 validation images, and 34,680 test images, and all images are 1640x5920 pixels. This dataset covers various traffic scenarios, including Normal, Crowd, Dazzle, Shadow, Noline, Arrow, Curve, CrossRoad, and Night.

[0143] Second, for the TuSimple dataset, there are three official metrics: false positive rate (FPR), false negative rate (FNR), and accuracy. We follow the literature and use the TuSimple metrics to calculate accuracy. The prediction accuracy is calculated as:

[0144]

[0145] Where S clip is the number of true prediction points in the video clip lane points, and C alip is the number of correctly predicted lane points. A prediction point within 20 pixels of the ground truth point is considered correct, and if the accuracy is greater than 85%, it is considered a true positive. Otherwise, it will be regarded as a false positive (FP) or a false negative (FN).

[0146] For the CULane dataset, the lane is considered 30 pixels wide. If the intersection over union (IoU) between the prediction and the ground truth is greater than 0.5, the predicted lane is regarded as a true positive. We also use the false negative rate (FN) and the false positive rate (FP) to evaluate our method. Another evaluation metric we use is the F1 score, and its formula is as follows:

[0147]

[0148] Where

[0149] It should be noted that TP represents the number of correctly detected lane lines, and FP represents the number of misdetected lane lines. This calculation model reflects the proportion of truly correct results among the results where the model predicts "there is a lane line", and measures the "precision" of the prediction.

[0150] It should be further noted that the F1 score is the harmonic mean of the precision and the recall, balancing the performance of both. The closer the value is to 1, the more precise and comprehensive the model is in detecting lane lines, and the better the comprehensive performance.

[0151] Finally, analysis of the experimental results:

[0152] The experimental environment established in this paper uses Ubuntu 22.04 as the operating system, PyTorch 1.13 as the deep learning framework, and Python 3.8 as the main development language. In terms of computing devices, GeForce RTX 4090 is selected. In the training parameter settings of the experimental model, ADAM is selected as the optimizer, the learning rate is set to 0.001, and it decays 10 times every 400k iterations. The batch size is set to 16. The used ResNet-18 and ResNet-34 are used as the backbone networks to create different versions of the proposed MHFS-Former. The input resolution of TuSimple is set to 640×360, and that of CULane is set to 820×295. The original data is enhanced by random scaling, color jittering, cropping, rotation, and horizontal flipping. The number of encoding layers and decoding layers is both set to 1, the number of training iterations is set to 400k, and the fixed number N of the prediction curve is set to 7. It should be noted that ResNet-18 and ResNet-34 are reported in existing literature.

[0153] Table 1 shows the Pareto solution optimization objective values of the NSGA-II algorithm and the prediction performance indicators of the DESE model.

[0154]

[0155] Table 2 shows the test results of MHFS-Former on the CULane dataset and the comparison with other methods

[0156]

[0157] Table 3 shows the ablation experiment results of the model based on the ResNet18 version on the CULane dataset

[0158]

[0159] It should be noted that the experimental results on the TuSimple dataset are shown in Table 1 and are attached Figure 5The visualization of the experimental results is given in Table 1. The proposed method achieved the highest F1 score. We note that CondLaneNet based on ResNet18 had the highest false negative rate (FNR) of 1.80%, but its false positive rate (FPR) was relatively poor at 6.17%. Similarly, SCNN also had the highest false negative rate of 1.80% and a relatively poor false positive rate of 6.17%. In contrast, our ResNet18-based version achieved a commendable false negative score of 2.28%, and the false positive score remained at a good level of 2.70%. In addition, our ResNet34-based version showed a balanced performance, with a false positive score of 2.34% and a false negative score of 2.43%, indicating a strong balance between false positives and false negatives.

[0160] Since this dataset mainly presents real highway scenes recorded at different times of day and night on US highways, these scenes are relatively simple and other methods have achieved good results, but the performance gap is small. However, our method still achieved a remarkable accuracy of up to 96.88%, significantly outperforming other methods. The visualization results also demonstrate the robustness of the method. It can be clearly seen from the figure that the lanes predicted by our method almost perfectly match the actual lane markings, proving the effectiveness of our proposed MHFNet in helping the Transformer network accurately locate the lanes.

[0161] It should be further noted that in the appendix Figure 5 the image in the middle of the first column shows the situation where a vehicle blocks the lane at a distant turn; the image at the bottom of the second column shows that a vehicle blocks the lane on the left side of the field of view; the images in the third column indicate that the lane cannot be fully seen due to vehicle occlusion. Since our proposed ShuffleLaneNet fully considers the information flow between feature channels and spatial positions, and our MHFS-Former can utilize the context information from the global environment, despite these visual obstacles, our method can still successfully and accurately identify the presence of the lane, demonstrating its strong robustness and accuracy.

[0162] It should be noted that the test results of our MHFS-Former on the CULane dataset and the comparison with other methods are shown in Table 2. As shown in the table, the method we proposed is significantly better than other methods. The R34 version of our MHFS-Former achieved 77.38% in terms of F1 score. In addition, our method demonstrated commendable performance under normal, strong light, shadow, and no lane line conditions. Under normal conditions, our method reached an amazing 92.89%, 68.86% under strong light conditions, 78.80% under shadow conditions, and 53.78% under no lane line conditions, all exceeding other methods. This excellent performance indicates that MHFS-Former can effectively utilize the rich high-dimensional feature information embedded near the lanes under complex scene conditions. Therefore, our method significantly improves the accuracy and reliability of detecting and locating these lanes in various challenging and complex environments.

[0163] The results of testing on the CULane dataset are as attached Figure 6 As shown, our method can effectively detect and locate lanes in all scenarios. Obviously, in strong light and shadow scenarios, the method we proposed can also accurately capture the positions of the lanes, achieving a high level of detection and location, which demonstrates its excellent adaptability and reliability.

[0164] Since ShuffleLaneNet fully considers the information flow between feature channels and spatial positions, and our MHFS-Former can proficiently utilize global information to accurately infer occluded and difficult-to-detect lanes, even in complex scenarios such as severe vehicle congestion (heavy traffic flow) and shadow occlusion, it can achieve high-precision lane detection and present good detection results.

[0165] Experimental results:

[0166] To verify the effectiveness of the proposed module, we conducted ablation experiments on the CULane dataset using a model based on the ResNet18 version. The empirical results of these experiments are systematically shown in Table 3 to comprehensively evaluate the role and effectiveness of the proposed module.

[0167] As shown in Table 3, in the basic model, the F1 score was initially 77.11%. After adopting ShuffleLaneNet, the score of the basic model increased to 77.16%, achieving a significant increase of 0.05%. In addition, after adding MHFNet to the basic model, the score rose to 77.30%, a substantial increase of 0.19%. When both MHFNet and ShuffleLaneNet are added to the basic model, the MHFS-Former model is obtained, and its performance has been significantly improved. The score soars to 77.38%, with a net increase of 0.27%, demonstrating its excellent performance. Since our MHFNet improves the accuracy of lane localization, the detection performance and robustness of this model have also been significantly enhanced.

[0168] The present invention provides an end-to-end lane detection network MHFS-FORMER based on Transformer, and its working principle is as follows: taking multi-scale features as input, through a multi-reference deformable attention module, for the slender structure and global environment of the lane, attention is allocated around the object, eliminating the complex post-processing process; enhancing the hybrid multi-scale features with MHFNet, fusing the multi-scale features with the output of the Transformer encoder to improve the effectiveness of multi-scale feature fusion; designing ShuffleLaneNet to mine the channel and spatial information in the lane features, strengthening feature extraction and enhancing the overall detection ability of the network. This method achieves an accuracy of 96.88% on the TuSimple dataset and an F1 score of 76.32% on the CULane dataset, verifying the effectiveness of the algorithm in lane detection.

[0169] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the gist of the present invention.

Claims

1. A lane line detection method based on multi-scale hybrid features-Transformer, characterized in that, It includes the following steps: S1. Input image preprocessing: Map the ground plane points to the bird's-eye view coordinates through the camera parameter conversion formula to compensate for the influence of camera tilt; S2. Multi-scale feature extraction: Use the ResNet backbone network to extract {C3, C4, C5} features, and the number of channels is unified to 256; S3. Multi-scale feature fusion: Cascade and fuse the features of adjacent levels through the multi-scale feature fusion network (MHFNet), and generate enhanced hybrid features by combining adaptive spatial fusion (ASF); S4. Hybrid lane feature enhancement: Use ShuffleLaneNet to perform dual-channel grouped attention calculation and channel shuffle operation; S5. Transformer processing: Decode the lane line parameters through the multi-reference deformable attention mechanism of the multi-layer Transformer decoder; S6. Loss calculation: Optimize the model parameters based on the bipartite matching loss of the Hungarian algorithm.

2. A method for lane line detection based on multi-scale hybrid features-Transformer according to claim 1, characterized in that, In the input image preprocessing of step S1, the lane line modeling includes the following steps: (a) Define the lane shape prior model as a polynomial structure: Among them, (X, Y) represents the points of the lane line on the ground plane, N is the polynomial order, and a n and b are coefficients. For a single-lane line on a flat ground, it is represented by a cubic curve as follows: X = kZ 3 + mZ 2 + nZ + b Formula 2 where k, m, n, and b are real parameters, k≠0, and (X, Z) indicates the points on the ground plane; (b) To avoid the interference of uneven road surfaces, convert the cubic polynomial to the bird's-eye view form, and the conversion formula is: where (X, Y) are the pixel points of the transformed phase plane, f x is the pixel width on the focal plane divided by the focal length, f y is the focal plane pixel height divided by the focal length, H is the height at which the camera is installed; (c) Considering the camera tilt angle θ, the non-tilt and tilt transformation formulas are: where, f is the camera focal length, f is the camera focal length after pitch transformation, (x′, y′) are the corresponding pixel points after pitch transformation; (d) Combining the above formulas, the transformed expression is obtained: where f″ = f sinθ, m′ = m cos 2 θ × H / f x × f y , n′ = n / f x , b′ = b × f y / f x cosθ × H, b″ = b × f y × f tanθ / f x × H; (e) The lane line output is parameterized as: p i = (c, k′, m′, n′, f″, b′, b″), Formula 6 where i ∈ {0, 1,..., M}, and M is the total number of lane lines.

3. A method for lane line detection based on multi-scale hybrid features-Transformer according to claim 1, characterized in that The multi-scale feature fusion network (MHFNet) in step S3 includes: (a) Perform 1×1 convolution and dimension alignment on the {C3, C4, C5} features respectively to generate {S3, S4, S5}; (b) Iteratively fuse S4 and F5 through the cascade fusion module to generate {P4, P5}, and the cascade fusion module includes two 1×1 convolutional layers, N RepBlocks, and element-wise addition operations; (c) Use adaptive spatial fusion (ASF) to perform dynamic weight allocation on {S3, P4, P5}, and its calculation formula is: Among them, 4. A method for lane line detection based on multi-scale hybrid features-Transformer according to claim 1, characterized in that The ShuffleLaneNet in step S4 includes: (a) Divide the input features into two groups along the channels, and perform spatial and channel attention fusion on the first group; (b) The second group is further divided into two subgroups, and spatial attention and channel attention are calculated respectively; (c) The channel attention is realized through global average pooling and Sigmoid gating, and the calculation formula is: The calculation formula of the Sigmoid gating is: X k ′ = s(F c (c))·X k = s(W1s + b1)·X k Formula 9 (d) The spatial attention is calculated through group normalization (GN) and activation function; X m ′ = σ(W2·GN(X m ) + b2)·X m Equation 10 (e) Realize cross-group information interaction through channel shuffle operation.

5. A method for lane line detection based on multi-scale hybrid features-Transformer according to claim 1, characterized in that The calculation formula of the multi-reference deformable attention module in step S5 is: Among them, is the normalized reference point coordinate, and Δp mlqk is the sampling offset, and A mlqk is the normalized weight.

6. A method for lane line detection based on multi-scale hybrid features-Transformer according to claim 1, characterized in that, The calculation of the loss tool in step S6 specifically includes: (a) Define a parameter set, and the calculation formula is: where M is the number of predicted lane lines, The calculation formula of the marker set representation is: (b) Through the true lane line search function l: S→P, transform the matching between the optimal injective parameter set of the predicted lane line and the true lane line into a problem of minimizing the cost, and calculate it through the following formula: where L bmc is the matching cost between the i-th true lane line and the prediction parameter set; (c) The regression loss function is defined as: Among them, g(c i ) is the probability of category c i . is the fitted lane line sequence, μ1 and μ2 are the coefficients of the loss function, and L mae is the mean absolute error, and Z(·) represents the indicator function.

7. A method for lane line detection based on multi-scale hybrid features-Transformer according to claim 1, characterized in that The preprocessing of the input image in step S1 includes: randomly scaling and color jittering the image. During the random scaling of the image, the scaling ratio is 0.5 - 2.0; during the color jittering process, first adjust the brightness, then randomly crop according to the size of 640×360 or 820×295, and finally perform a horizontal flip operation with a probability of 0.5 to improve the generalization ability through multi-dimensional data augmentation.

8. A method for lane line detection based on multi-scale hybrid features-Transformer according to claim 1, characterized in that The Transformer decoder in step S5 includes a processing structure composed of 6 layers of multi-head self-attention modules, receives multi-scale feature inputs {S3, S4, F5}, and at the same time realizes global context modeling by learning lane embedding vectors to strengthen the global correlation analysis and perception ability of lane line features.

9. A method for lane line detection based on multi-scale hybrid features-Transformer according to claim 1, characterized in that, The ResNet backbone network output features {C3, C4, C5} in step S2 are specifically: The C3 feature size is H / 8×W / 8 and the number of channels is 256, the C4 feature size is H / 16×W / 16 and the number of channels is 256, the C5 feature size is H / 32×W / 32 and the number of channels is 256. Outputting features of different scales through layering provides multi-dimensional information for subsequent processing.

Citation Information

Cited By

  • Small target detection method and system based on deformable recurrent neural network

    CN122574597B