Lane line detection method based on multi-scale fusion
By adopting a multi-scale fusion method in lane line detection technology, DCNResNet and AGFPN feature extraction networks are built, and combined with cross-layer refinement and multi-loss functions, the problem of low lane line detection accuracy in complex traffic environments is solved, and higher detection accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510010099.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-16
AI Technical Summary
The existing lane line detection technology has low detection accuracy in complex traffic environments, high computing resource costs, and insufficient generalization capabilities.
Using a lane line detection method based on multi-scale fusion, a new feature extraction backbone network DCNResNet and an improved progressive feature pyramid network AGFPN are used to improve feature extraction and multi-scale feature fusion capabilities by constructing a new feature extraction backbone network and an improved progressive feature pyramid network, and combining a cross-layer refinement network and four loss functions.
It significantly improves the accuracy and generalization ability of lane line detection, enhances the stability and robustness of the model, and adapts to lane line detection in complex traffic environments.
Smart Images

Figure CN120014573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving technology, and in particular to a lane line detection method based on multi-scale fusion. Background Art
[0002] Lane detection technology is an important component of advanced driver assistance systems (ADAS) and self-driving cars. It is used to identify lane lines on the road and help vehicles identify driving paths during real-time driving through technical means such as image processing and machine learning. The core goal of lane detection is to ensure that vehicles drive safely within the lane and prevent them from deviating from the lane, thereby improving driving safety and efficiency.
[0003] Traditional lane line detection methods are based on image processing to solve the lane line detection problem. This type of method mainly uses edge detection, Hough transform, morphological processing and other technologies to identify lane lines. Traditional methods based on image processing play a basic role in lane line detection, but they also have some inherent defects, such as poor adaptability to complex scenes. If the weather changes, lighting conditions are poor (such as at night, backlight) or the lane lines are blurred, the performance will be significantly reduced.
[0004] Compared with traditional methods, deep learning-based methods have more powerful feature learning and extraction capabilities. Deep learning models can automatically learn features from data, avoiding the process of manually designing features in traditional methods. Convolutional neural networks (CNNs) can identify different shapes and details of lane lines through multi-level feature extraction, improving the accuracy and robustness of detection. However, there are still some key problems in the actual application of deep learning-based methods, such as unsatisfactory detection accuracy in complex environments such as strong light, damaged lane lines, and vehicle congestion, high computing resource costs, and insufficient generalization capabilities in different environments. Summary of the invention
[0005] Purpose of the invention: In view of the defects of the prior art, the present invention discloses a lane line detection method based on multi-scale fusion, which aims to improve the lane line detection accuracy and generalization ability in complex road scenes. By constructing a new feature extraction backbone network and an improved progressive feature pyramid network multi-scale feature extraction model, the global information of the image can be better utilized for lane line detection, thereby achieving the purpose of lane line detection in complex road scenes.
[0006] Technical solution: The present invention discloses a lane line detection method based on multi-scale fusion, comprising the following steps:
[0007] Step 1: Obtain a lane line image and preprocess the lane line image;
[0008] Step 2: Lane feature extraction: Input the image obtained in step 1 into the backbone network for feature extraction to obtain a multi-scale feature map; the backbone network uses the ResNet residual structure, and introduces the deformable convolution DCNV3 to construct a new backbone network DCNResNet;
[0009] Step 3: Design a new progressive feature pyramid network AGFPN for multi-scale feature fusion. In the bottom-up feature extraction process of the backbone network, the progressive feature pyramid network AFPN progressively integrates low-level features, high-level features and top-level features. By integrating the skip layer connection method in the generalized feature pyramid network GFPN on the basis of the progressive feature pyramid network AFPN, feature layers of different resolutions are directly connected.
[0010] Step 4: Build a cross-layer refinement network, realize feature fusion and collect global context with rich semantic information through ROI aggregation network and lane prior network, and output the predicted foreground and background probability, length, axis and angle information of lane lines through convolutional layer and fully connected layer;
[0011] Step 5: Construct the loss function. Four loss functions are introduced, namely regression loss, classification loss, segmentation loss, and line segment intersection loss, to determine the total training loss function.
[0012] Step 6: Model training: Calculate and update network parameters through the error back propagation algorithm to minimize the loss function. Continue to improve model performance by optimizing classification loss, regression loss, segmentation loss, and segment intersection loss until the model reaches convergence. Use the trained model for lane line detection.
[0013] Furthermore, in step 2, the original 3x3 convolutional layer in the ResNet residual structure is replaced with a DCNV3 convolutional layer to construct a new backbone network DCNResNet. The expression of DCNV3 is:
[0014]
[0015] Where G is the total number of aggregation groups, K is the total number of sampling points, k is the enumeration number of sampling points, and p is 0 is the current pixel, p k is the kth grid sampling position, for the gth group, w g ∈R C×C′ represents the position-independent projection weight of the group, where C′=C / G represents the group dimension, m gk ∈R represents the modulation scalar of the k-th sampling point in the g-th group, normalized by the softmax function along the k-th dimension, x g ∈R C′×H×W represents the input feature map after slicing, Δpgk is the sampling position p of the g-th group of grids k The corresponding offset.
[0016] Furthermore, the progressive feature pyramid network AFPN also introduces an adaptive spatial fusion operation in the multi-level feature fusion process, as follows:
[0017] In the progressive feature pyramid network AFPN, there are multiple scale feature maps, namely F 1 ,F 2 ,...,F n , where F i Representing feature maps at different levels, the AFPN cross-scale feature fusion process is:
[0018] Downsampling from a higher resolution feature layer to a lower resolution feature layer, the downsampling path expression is:
[0019] F i+1 =DownSample(F i )
[0020] Upsampling from a lower resolution feature layer to a higher resolution feature layer, the upsampling path expression is:
[0021] F n-1 =UpSample(F n )
[0022] Feature fusion expression:
[0023]
[0024] In the formula, Resize(F j ,size(F i )) indicates that the feature map F is upsampled or downsampled j Adjust to F i The same spatial size, α ij is the weight coefficient of adaptive fusion, which indicates the contribution of the j-th layer feature to the i-th layer feature. i final It is the fused feature map used for prediction tasks.
[0025] Furthermore, in step 2, the skip layer connection mode of GFPN is integrated on the basis of the progressive feature pyramid network AFPN to construct a new progressive feature pyramid network AGFPN, as follows:
[0026] The expression of the feature map of the first layer receiving all previous layers is:
[0027]
[0028] In the formula, Concat refers to the concatenation of the feature maps generated in the previous layer, Conv refers to 3x3 convolution, Represents each scale feature in level k;
[0029] The new progressive feature pyramid network AGFPN fusion process expression is:
[0030]
[0031] In the formula, Resize(F k ,size(F i )) indicates that the feature map F is converted by upsampling or downsampling. k Adjust to F i The same space size, F k represents the feature map of the kth layer, α ij represents the adaptive fusion weight of adjacent level features in AFPN, β ik Represents the adaptive fusion weights of the GFPN skip layer connections, which are used to directly connect non-adjacent features across layers. Represents the final feature map after feature fusion of the i-th layer.
[0032] Furthermore, the specific process of the cross-layer refinement network in step 4 is as follows: set the new progressive feature pyramid network AGFPN, combined with the backbone network DCNResNet to generate feature levels {L0, L1, L2, L3}, cross-layer refinement starts from the highest level L0 and gradually approaches L3, and uses {R0, R1, R2, R3} to represent the corresponding refinement. The cross-layer refinement network expression is:
[0033]
[0034] Where t = 1, ···T, T is the total number of optimizations, detection is performed from the highest level high semantic layer, P t are the lane prior parameters, starting point coordinates x, y and angle θ;
[0035] For the first layer L0, P0 is evenly distributed on the image plane, and the refined R t P t As input, we obtain the channel features of the region of interest (ROI), and then perform two fully connected layers to obtain the refinement parameters P t , perform convolution on the extracted ROI features to collect nearby features of each channel pixel, and use the fully connected layer to further extract channel prior features:
[0036]
[0037] The feature map is resized to:
[0038]
[0039] and flattened to:
[0040]
[0041] In order to collect the global context of channel prior features, we first calculate the ROI channel prior feature χ p and the global feature map χ f The attention matrix between Its expression is:
[0042]
[0043] In the formula, f is the normalization function softmax, C is the number of input channels, and then the features are aggregated according to the attention matrix, which is expressed as:
[0044]
[0045] Output Reflects from Select from all locations arrive Finally, the output is added to the original input middle.
[0046] Furthermore, the loss function in step 5 is specifically as follows:
[0047] The regression loss uses Smooth L1 loss to predict the position of the lane line, and its expression is:
[0048]
[0049] Among them, x represents the error between the predicted value and the target value, and |x| represents the absolute value of the error;
[0050] Focal classification loss is used to deal with the problem of category imbalance, and its expression is:
[0051] FL(p t )=-α t (1-p t ) γ log(p t )
[0052] In the formula, p t is the model’s predicted probability for the target class, α t is a balancing factor used to adjust the influence between positive and negative samples, and γ is a focus factor used to adjust the weight of difficult and easy samples;
[0053] Segmentation loss, an auxiliary cross entropy loss for each pixel segmentation mask, is expressed as:
[0054]
[0055] Where N is the number of samples, y i is the classification label of the target mask, p i is the probability or score of the predicted mask;
[0056] The segment intersection loss regresses the lane prior as a whole unit and its expression is:
[0057]
[0058] In the formula, H represents the total number of slices in the entire height direction, I i and U i is defined as:
[0059]
[0060] In the formula, represents the x-coordinate of the center position of lanes p and q on slice i, represents the virtual lane width of lanes p and q on slice i, I i represents the intersection width of lanes p and q on slice i, U i Represents the union width of lanes p and q on slice i.
[0061] Furthermore, during the training process, each ground truth lane is dynamically assigned one or more predicted lanes as a positive sample, and the predicted lanes are sorted according to the assigned cost, which is defined as:
[0062]
[0063] in, is the focal cost between prediction and label, is the similarity cost between the predicted lane and the real lane, which consists of three parts, represents the average pixel distance of all valid lane points, Indicates the distance of the starting point coordinates, represents the difference in angle θ between the starting point of the lane line and the x-axis of the prior lane, normalized to [0,1]; w cls and w sim is the weight coefficient for each defined component, and each ground truth lane is calculated according to Assigned with a dynamic number of prediction lanes;
[0064] The total training loss is:
[0065] L=λ0 L reg +λ 1 L cls +λ 2 L seg +λ 3 L LaneIoU
[0066] Where, L reg , L cls , L seg , L LaneIoU They are regression loss, classification loss, segmentation loss, line segment intersection loss, λ 0 , 1 , 2 , 3 They are the weight parameters of regression loss, classification loss, segmentation loss, and line segment intersection loss respectively;
[0067] Finally, effective inference is performed by setting a threshold with classification scores to filter out background lanes, i.e., low-class lane priors, and using non-maximum suppression to remove subsequent high-overlap lanes.
[0068] Furthermore, the pre-processing in step 1 includes:
[0069] The lane line image is grayed out. In the RGB color scheme, each pixel in the image is composed of the values of three channels: red, green, and blue. The grayscale image is obtained by the weighted average method. Edge detection is performed on the grayscale image to accurately detect the edges in the image.
[0070] Beneficial effects:
[0071] The present invention adopts a lane detection method based on multi-scale fusion, which can significantly improve the accuracy and generalization ability of lane detection in complex traffic scenes. The present invention introduces deformable convolution DCNV3 into the backbone network ResNet, and dynamically adjusts the position of the convolution kernel to enable it to better adapt to image features of different scales and improve the feature extraction ability of the backbone network. And on the basis of the progressive feature pyramid network (AFPN), the jump layer connection mode of GFPN is integrated, and the cross-layer information interaction ability is enhanced by directly connecting feature layers of different resolutions, which effectively solves the problem of feature information loss during up and down sampling, and can better integrate global information to ensure cross-scale connection while multi-scale feature fusion. In summary, the present invention significantly improves the feature extraction and multi-scale feature fusion capabilities by improving the model, thereby improving the accuracy and generalization ability of lane detection in challenging environments. At the same time, the stability and robustness of the model are improved through cross-layer refinement network and four loss functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 This is a flow chart of the lane line detection method in the present invention;
[0073] Figure 2 This is a structural diagram of the lane detection algorithm in the present invention;
[0074] Figure 3 It is a structural diagram of the novel backbone network residual unit in the present invention;
[0075] Figure 4 It is a diagram of the structure of the progressive feature pyramid network in the present invention;
[0076] Figure 5 This is a skip layer structure diagram of the generalized feature pyramid network in the present invention;
[0077] Figure 6 This is a structural diagram of lane prior information output in the present invention. DETAILED DESCRIPTION
[0078] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0079] The purpose of the present invention is to solve the problem that the lane line detection accuracy of the prior art in complex traffic environments is not high, and to provide a lane line detection method and system based on multi-scale fusion, which improves the lane line detection accuracy through deep learning methods, and further improves the safety and reliability of autonomous driving.
[0080] To achieve the above purpose, refer to Figure 1 and Figure 2 , the present invention specifically comprises the following steps:
[0081] Step 1: Preprocess the lane line image: grayscale the input image, perform Canny edge detection and other image preprocessing.
[0082] The lane line image is grayed out. Each image consists of different color channels. In the RGB color scheme, each pixel in the image consists of the values of the three channels: red, green, and blue. A more reasonable grayscale image can be obtained by the weighted average method. The conversion algorithm is:
[0083] I gray =0.299×I R +0.587×I G +0.114×I B
[0084] In the formula, I R , I G , I Bare the pixel values of the red, green, and blue channels respectively, I gray is the pixel value of the grayscale image.
[0085] Edge detection is performed on grayscale images to accurately detect the edges in the image and have good resistance to noise. The detailed calculation steps are as follows:
[0086] 1. Use Gaussian filter to smooth the grayscale image to reduce the influence of noise. The expression is:
[0087]
[0088] Where G(x,y) represents the Gaussian function and σ represents the standard deviation of the filter.
[0089] 2. Use the Sobel operator to calculate the gradient values of the grayscale image in the horizontal and vertical directions, so as to obtain the gradient strength and direction of each pixel.
[0090] The Sobel operator expression is:
[0091]
[0092] In the formula, G x and G y Represent the gradient values in the horizontal and vertical directions respectively, and * represents the convolution operator.
[0093] Gradient magnitude and direction formula:
[0094]
[0095] Where M(x,y) represents the gradient strength of the pixel, and θ(x,y) represents the gradient direction of the pixel.
[0096] 3. Through non-maximum gradient pixel suppression, compare the gradient values of the two adjacent pixels in the gradient direction, and only retain the pixels with the largest gradient values in the gradient direction. Then use double threshold processing to determine the edge strength. Finally, use the edge connection algorithm to connect the strong edge pixels and the adjacent weak edge pixels to form a complete edge.
[0097] Step 2: Lane feature extraction: Input the image obtained in step 1 into the new backbone network for feature extraction, introduce deformable convolution DCNV3 into the backbone network, and construct a new backbone network DCNResNet, so that it can better adapt to image features of different scales and improve the feature extraction ability of the backbone network.
[0098] The residual network ResNet-101 is used as the backbone network to extract features from the input image. ResNet is implemented through the residual module, which retains the better one of the original branch and the learned branch. The residual module avoids network degradation through identity mapping.
[0099] The identity mapping expression is:
[0100]
[0101] In the formula, Represents the residual mapping function learned by the network layer, and x and y are the input and output layers of the residual network respectively.
[0102] In order to better adapt to image features of different scales and improve the feature extraction capability of the backbone network, the deformable convolution DCNV3 is introduced into the residual structure of the backbone network. The original 3x3 convolution layer in the ResNet residual structure is replaced with the DCNV3 convolution layer to construct a new backbone network DCNResNet.
[0103] Specifically, the input features enter the convolution module and pass through the lightweight convolution subnetwork f offset Generate dynamic offset Δp gk Used to adjust the sampling point position through another sub-network f modulation Generate dynamic blur weight m gk Used to adjust the importance of sampling points, and then according to p k +Δp gk The offset sampling points are calculated, and bilinear interpolation is used to extract values from the input feature map. Finally, weighted convolution is performed on each sampling point to generate context-enhanced features and output dynamically adjusted features, thereby enhancing the network's receptive field and adaptability to irregular targets.
[0104] The expression of DCNV3 is:
[0105]
[0106] Where G is the total number of aggregation groups, K is the total number of sampling points, k is the enumeration number of sampling points, and p is 0 is the current pixel, p k is the kth grid sampling position, for the gth group, w g ∈R C×C′ represents the position-independent projection weight of the group, where C′=C / G represents the group dimension, m gk ∈R represents the modulation scalar of the k-th sampling point in the g-th group, normalized by the softmax function along the k-th dimension, x g ∈R C′×H×W represents the input feature map after slicing, Δp gkis the sampling position p of the g-th group of grids k The corresponding offset.
[0107] Step 3, design of progressive feature pyramid network: design a new progressive feature pyramid network (AGFPN). In the bottom-up feature extraction process of the backbone network, the progressive feature pyramid network (AFPN) gradually integrates low-level features, high-level features and top-level features to achieve feature fusion. By integrating the skip layer connection method in the generalized feature pyramid network (GFPN) on the basis of the progressive feature pyramid network (AFPN), the design directly connects feature layers of different resolutions, enhances the cross-layer information interaction capability, and can better integrate global information.
[0108] The progressive feature pyramid network (AFPN) is adopted. In the bottom-up feature extraction process of the backbone network, AFPN gradually fuses low-level and high-level features to achieve direct interaction between non-adjacent layers. In order to suppress the information contradiction between features at different levels, AFPN introduces adaptive spatial fusion operation in the multi-level feature fusion process.
[0109] In the AFPN network, we have feature maps of multiple scales, namely F 1 ,F 2 ,...,F n , where F i Representing feature maps at different levels, the AFPN cross-scale feature fusion process is:
[0110] 1. Downsample from higher resolution feature layers (such as F1, F2) to lower resolution feature layers (such as F3 or F4) to provide more detailed information. Downsampling path expression:
[0111] F i+1 =DownSample(F i )
[0112] 2. Upsample from lower resolution feature layers (such as F3, F4) to higher resolution feature layers (such as F2 or F1) to provide more context information. Upsampling path expression:
[0113] F n-1 =UpSample(F n )
[0114] 3. Feature fusion expression:
[0115]
[0116] In the formula, Resize(F j ,size(F i )) represents the feature map Fj Adjust to F i Same spatial size (by upsampling or downsampling). ij is the weight coefficient of adaptive fusion, which indicates the contribution of the j-th layer features to the i-th layer features. These weights are usually learned through the network. It is the fused feature map used for prediction tasks.
[0117] In order to enhance AFPN, the present invention integrates the skip layer connection mode of GFPN on the basis of AFPN to construct a new progressive feature pyramid network (AGFPN), which can better integrate global information through multi-scale information exchange. The feature map expression of the first layer receiving all the previous layers is:
[0118]
[0119] In the formula, Concat refers to the concatenation of the feature maps generated in the previous layer, Conv refers to 3x3 convolution, Represents each scale feature in level k.
[0120] The AGFPN fusion process expression is:
[0121]
[0122] In the formula, α ij represents the adaptive fusion weight of adjacent level features in AFPN, β ik Represents the adaptive fusion weights of GFPN skip layer connections, which are used to directly connect non-adjacent features across layers.
[0123] Step 4: Implementation of cross-layer refinement structure: Through the ROI aggregation network and lane prior network, feature fusion can be effectively realized and global context with rich semantic information can be collected.
[0124] Lane line modeling is in the form of a point set. Specifically, the lane is represented as a series of points, and its expression is:
[0125] P={(x 1 ,y 1 ),…,(x N ,y N )}
[0126] In the formula, the y coordinate of the point is sampled evenly in the vertical image, and its expression is:
[0127]
[0128] Where H is the height of the image, and N=72, which means that the input image is divided into 72 equal parts in height.
[0129] Using the improved progressive feature pyramid network (AGFPN), combined with the backbone network DCNResNet to generate feature levels {L0, L1, L2, L3}, cross-layer refinement starts from the highest level L0 and gradually approaches L3. We use {R0, R1, R2, R3} to represent the corresponding refinement. The cross-layer refinement network expression is:
[0130]
[0131] Where t = 1, ... T, T is the total number of optimizations. Our method performs detection from the highest level high semantic layer, P t are the lane prior parameters (starting point coordinates x, y and angle θ).
[0132] For the first layer L0, P0 is evenly distributed on the image plane. t P t As input, we obtain the region of interest (ROI) channel features, and then perform two fully connected layers to obtain the refinement parameters P t , convolution is performed on the extracted ROI features to collect nearby features of each channel pixel. To save memory, we use a fully connected layer to further extract channel prior features:
[0133]
[0134] The feature map is resized to:
[0135]
[0136] and flattened to:
[0137]
[0138] In order to collect the global context of channel prior features, we first calculate the ROI channel prior feature χ p and the global feature map χ f The attention matrix between , whose expression is:
[0139]
[0140] In the formula, f is the normalization function softmax, C is the number of input channels, and then the features are aggregated according to the attention matrix, which is expressed as:
[0141]
[0142] Output Reflects from Select from all locations arrive Finally, the output is added to the original input middle.
[0143] Step 5. Construct loss function: Improve model performance by using four specific loss functions (classification loss, regression loss, segmentation loss, and segment intersection loss).
[0144] The present invention introduces four loss functions, namely regression loss, classification loss, segmentation loss and line segment intersection-to-union ratio loss.
[0145] The regression loss uses Smooth L1 loss to predict the position of the lane line, and its expression is:
[0146]
[0147] Focal classification loss is used to deal with the problem of category imbalance, and its expression is:
[0148] FL(p t )=-α t (1-p t ) γ log(p t )
[0149] In the formula, p t is the model’s predicted probability for the target class, α t is a balancing factor used to adjust the influence between positive and negative samples, and γ is a focus factor used to adjust the weight of difficult and easy samples.
[0150] The segmentation loss is an auxiliary cross entropy loss for the per-pixel segmentation mask, which is expressed as:
[0151]
[0152] Where N is the number of samples, y i is the classification label of the target mask, p i is the probability or score of the predicted mask.
[0153] The LaneIoU loss can regress the lane prior as a whole unit to improve the positioning accuracy. Its expression is:
[0154]
[0155] In the formula, H represents the total number of slices in the entire height direction, I i and U i is defined as:
[0156]
[0157] In the formula, represents the x-coordinate of the center position of lanes p and q on slice i, represents the virtual lane width of lanes p and q on slice i, I i represents the intersection width of lanes p and q on slice i, U i Represents the union width of lanes p and q on slice i.
[0158] Step 6: Model training: Calculate and update network parameters through the error back propagation algorithm to minimize the loss function. Continue to improve model performance by continuously optimizing classification loss, regression loss, segmentation loss, and segment intersection loss until the model reaches a convergence state.
[0159] Reference Figure 6 , each lane prior will be predicted by the network and consists of four parts
[0160] (1) Foreground and background probabilities.
[0161] (2) Lane length.
[0162] (3) The angle between the starting point of the lane line and the x-axis of the prior lane (called x, y, and θ).
[0163] (4) N offsets, i.e., the horizontal distance between the prediction and its true value.
[0164] During training, each ground truth lane is dynamically assigned one or more predicted lanes as a positive sample. In particular, the predicted lanes are ranked according to the assigned cost, which is defined as:
[0165]
[0166] in, is the focal cost between prediction and label, is the similarity cost between the predicted lane and the real lane, which consists of three parts, represents the average pixel distance of all valid lane points, Indicates the distance of the starting point coordinates, represents the difference in theta angle, normalized to [0,1]; w cls and w sim is the weight coefficient for each defined component, and each ground truth lane is calculated according to A dynamic number of prediction lanes are assigned.
[0167] Secondly, there is the training loss, and the total loss is:
[0168] L=λ 0 L reg +λ 1 L cls+λ 2 L seg +λ 3 L LaneIoU
[0169] Where, L reg , L cls , L seg , L LaneIoU They are regression loss, classification loss, segmentation loss, line segment intersection loss, λ 0 , 1 , 2 , 3 They are the weight parameters of regression loss, classification loss, segmentation loss, and line segment intersection loss. Finally, effective inference is performed by setting a threshold with classification scores to filter background lanes (low-class lane priors) and using non-maximum suppression to delete subsequent high-overlap lanes.
Claims
1. A lane line detection method based on multi-scale fusion, characterized in that: The steps include: Step 1: Obtain a lane line image and preprocess the lane line image; Step 2: Lane feature extraction: Input the image obtained in step 1 into the backbone network for feature extraction to obtain a multi-scale feature map; the backbone network uses the ResNet residual structure, and introduces the deformable convolution DCNV3 to construct a new backbone network DCNResNet; Step 3: Design a new progressive feature pyramid network AGFPN for multi-scale feature fusion. In the bottom-up feature extraction process of the backbone network, the progressive feature pyramid network AFPN progressively integrates low-level features, high-level features and top-level features. By integrating the skip layer connection method in the generalized feature pyramid network GFPN on the basis of the progressive feature pyramid network AFPN, feature layers of different resolutions are directly connected. Step 4: Build a cross-layer refinement network, realize feature fusion and collect global context with rich semantic information through ROI aggregation network and lane prior network, and output the predicted foreground and background probability, length, axis and angle information of lane lines through convolutional layer and fully connected layer; Step 5: Construct the loss function. Four loss functions are introduced, namely regression loss, classification loss, segmentation loss, and line segment intersection loss, to determine the total training loss function. Step 6: Model training: Calculate and update network parameters through the error back propagation algorithm to minimize the loss function. Continue to improve model performance by optimizing classification loss, regression loss, segmentation loss, and segment intersection loss until the model reaches convergence. Use the trained model for lane line detection.
2. The lane line detection method based on multi-scale fusion according to claim 1, characterized in that: In step 2, the original 3x3 convolution layer in the ResNet residual structure is replaced with the DCNV3 convolution layer to construct a new backbone network DCNResNet. The expression of DCNV3 is: Where G is the total number of aggregation groups, K is the total number of sampling points, k is the enumeration number of sampling points, p0 is the current pixel, and p k is the kth grid sampling position, for the gth group, w g ∈R C×C′ represents the position-independent projection weight of the group, where C′=C / G represents the group dimension, m gk ∈R represents the modulation scalar of the k-th sampling point in the g-th group, normalized by the softmax function along the k-th dimension, x g ∈R C′×H×W represents the input feature map after slicing, Δp gk is the sampling position p of the g-th group of grids k The corresponding offset.
3. The lane line detection method based on multi-scale fusion according to claim 1, characterized in that: The progressive feature pyramid network AFPN also introduces an adaptive spatial fusion operation in the multi-level feature fusion process, as follows: In the progressive feature pyramid network AFPN, there are multiple scale feature maps, namely F1, F2, ..., F n , where F i Representing feature maps at different levels, the AFPN cross-scale feature fusion process is: Downsampling from a higher resolution feature layer to a lower resolution feature layer, the downsampling path expression is: F i+1 =DownSample(F i ) Upsampling from a lower resolution feature layer to a higher resolution feature layer, the upsampling path expression is: F n-1 =UpSample(F n ) Feature fusion expression: In the formula, Resize(F j ,size(F i )) indicates that the feature map F is converted by upsampling or downsampling. j Adjust to F i The same spatial size, α ij is the weight coefficient of adaptive fusion, which indicates the contribution of the j-th layer feature to the i-th layer feature. i final It is the fused feature map used for prediction tasks.
4. The lane line detection method based on multi-scale fusion according to claim 3 is characterized in that: In step 2, the skip layer connection mode of GFPN is integrated on the basis of the progressive feature pyramid network AFPN to construct a new progressive feature pyramid network AGFPN, as follows: The expression of the feature map of the first layer receiving all previous layers is: In the formula, Concat refers to the concatenation of the feature maps generated in the previous layer, Conv refers to 3x3 convolution, Represents each scale feature in level k; The new progressive feature pyramid network AGFPN fusion process expression is: In the formula, Resize(F k ,size(F i )) indicates that the feature map F is converted by upsampling or downsampling. k Adjust to F i The same space size, F k represents the feature map of the kth layer, α ij represents the adaptive fusion weight of adjacent level features in AFPN, β ik Represents the adaptive fusion weights of the GFPN skip layer connections, which are used to directly connect non-adjacent features across layers. Represents the final feature map after feature fusion of the i-th layer.
5. The lane line detection method based on multi-scale fusion according to claim 1, characterized in that: The specific process of the cross-layer refinement network in step 4 is as follows: set up a new progressive feature pyramid network AGFPN, combined with the backbone network DCNResNet to generate feature levels {L0, L1, L2, L3}, cross-layer refinement starts from the highest level L0 and gradually approaches L3, and uses {R0, R1, R2, R3} to represent the corresponding refinement. The cross-layer refinement network expression is: Where t = 1, ···T, T is the total number of optimizations, detection is performed from the highest level high semantic layer, P t are the lane prior parameters, starting point coordinates x, y and angle θ; For the first layer L0, P0 is evenly distributed on the image plane, and the refined R t P t As input, we obtain the channel features of the region of interest (ROI), and then perform two fully connected layers to obtain the refinement parameters P t , perform convolution on the extracted ROI features to collect nearby features of each channel pixel, and use the fully connected layer to further extract channel prior features: The feature map is resized to: and flattened to: In order to collect the global context of channel prior features, we first calculate the ROI channel prior feature χ p and the global feature map χ f The attention matrix between Its expression is: In the formula, f is the normalization function softmax, C is the number of input channels, and then the features are aggregated according to the attention matrix, which is expressed as: Output Reflects from Select from all locations arrive Finally, the output is added to the original input middle.
6. The lane line detection method based on multi-scale fusion according to claim 1, characterized in that: The loss function in step 5 is as follows: The regression loss uses Smooth L1 loss to predict the position of the lane line, and its expression is: Among them, x represents the error between the predicted value and the target value, and |x| represents the absolute value of the error; Focal classification loss is used to deal with the problem of category imbalance, and its expression is: FL(p t )=-a t (1-p t ) γ log(p t ) In the formula, p t is the model’s predicted probability for the target class, α t is a balancing factor used to adjust the influence between positive and negative samples, and γ is a focus factor used to adjust the weight of difficult and easy samples; Segmentation loss, an auxiliary cross entropy loss for each pixel segmentation mask, is expressed as: Where N is the number of samples, y i is the classification label of the target mask, p i is the probability or score of the predicted mask; The segment intersection loss regresses the lane prior as a whole unit and its expression is: In the formula, H represents the total number of slices in the entire height direction, I i and U i is defined as: In the formula, represents the x-coordinate of the center position of lanes p and q on slice i, represents the virtual lane width of lanes p and q on slice i, I i represents the intersection width of lanes p and q on slice i, U i Represents the union width of lanes p and q on slice i.
7. The lane line detection method based on multi-scale fusion according to claim 6, characterized in that: During the training process, each ground truth lane is dynamically assigned one or more predicted lanes as a positive sample, and the predicted lanes are sorted according to the assigned cost, which is defined as: in, is the focal cost between prediction and label, is the similarity cost between the predicted lane and the real lane, which consists of three parts, represents the average pixel distance of all valid lane points, Indicates the distance of the starting point coordinates, represents the difference in angle θ between the starting point of the lane line and the x-axis of the prior lane, normalized to [0,1]; w cls and w sim is the weight coefficient for each defined component, and each ground truth lane is calculated according to Assigned with a dynamic number of prediction lanes; The total training loss is: L=λ0L reg +λ1L cls +λ2L seg +λ3L LaneIoU Where, L reg , L cls , L seg , L LaneIoU are regression loss, classification loss, segmentation loss, and line segment intersection loss, respectively. λ0, λ1, λ2, and λ3 are weight parameters for regression loss, classification loss, segmentation loss, and line segment intersection loss, respectively. Finally, effective inference is performed by setting a threshold with classification scores to filter out background lanes, i.e., low-class lane priors, and using non-maximum suppression to remove subsequent high-overlap lanes.
8. The lane line detection method based on multi-scale fusion according to claim 1, characterized in that: The pre-processing in step 1 includes: The lane line image is grayed out. In the RGB color scheme, each pixel in the image is composed of the values of three channels: red, green, and blue. The grayscale image is obtained by the weighted average method. Edge detection is performed on the grayscale image to accurately detect the edges in the image.