Feature fusion and row anchor point classification driven fast lane line detection method

By grid-based preprocessing and feature fusion of lane line detection methods, it is transformed into row anchor point classification tasks, and a multi-layer network model is built, which solves the problem of insufficient real-time and accuracy of existing lane line detection methods in complex environments, and achieves fast and efficient lane line detection.

CN119992497APending Publication Date: 2025-05-13SICHUAN POLICE COLLEGE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071631.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing lane line detection methods are difficult to meet the real-time and accuracy requirements in complex traffic environments, especially in the case of light changes, occlusion and blurring. The traditional methods have poor results. Although the deep learning methods have high accuracy, their calculation complexity is high, and it is difficult to meet the real-time requirements.

Method used

The fast lane line detection method driven by feature fusion and row anchor point classification is adopted. By grid-based preprocessing of the input images, the lane line detection task is converted into row anchor point classification task, and a lane line detection model including backbone network, neck network, classification network and auxiliary segmentation network is constructed. Multi-scale feature fusion and global feature extraction are used to improve detection accuracy and speed.

Benefits of technology

It significantly reduces the amount of calculation, improves the real-time and accuracy of lane line detection, and can effectively detect lane lines in complex environments to meet the requirements of real-time and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992497A_ABST
    Figure CN119992497A_ABST
Patent Text Reader

Abstract

The invention discloses a feature fusion and row anchor classification driven fast lane line detection method, and relates to the technical field of computer vision and intelligent driving, and the method comprises the following steps: S1, carrying out the grid preprocessing of an input image; s2, constructing a lane line detection model; s3, using the gridded image to train a lane line detection model; and S4, inputting an image to be detected into the trained lane line detection model to obtain a row anchor point classification result, namely a lane line detection result. According to the invention, a lane line detection task is converted into a row anchor point classification task, so that the calculation amount is obviously reduced and the detection speed is improved. A lane line detection model is constructed and trained to complete a row anchor point classification task, and the lane line detection model has the capability of sensing global information, so that the lane line detection effect in a complex environment is improved; on the basis of ensuring the lane line detection accuracy, the detection time is remarkably shortened, and the dual requirements of real-time performance and accuracy are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and intelligent driving technology, and in particular to a fast lane line detection method driven by feature fusion and line anchor point classification. Background Art

[0002] With the acceleration of urbanization and the development of intelligent driving technology, vehicle assisted driving systems play a key role in ensuring traffic safety. Among them, lane line detection, as one of the core technologies of intelligent driving, directly affects the safety and driving stability of the vehicle. However, the complex urban traffic environment puts forward higher real-time and accuracy requirements for lane line detection.

[0003] The traditional method mainly obtains the manual features of lane lines and models and fits lane curves based on parameters. The mainstream method is lane line detection based on Hough transform, which is only applicable to simple environments. In actual scenes, such as complex situations such as light changes, occlusion, and blur, the effect is poor. In recent years, deep learning methods have made significant progress in the field of lane line detection. The detection methods based on deep learning mainly include segmentation and classification. The segmentation method regards lane line detection as a pixel-level semantic segmentation task. Although it has high accuracy, the computational complexity is large and it is difficult to meet real-time requirements. The classification method effectively reduces the amount of calculation by dividing the image into multiple grids and classifying each grid, but it lacks the ability to extract multi-scale information and global features in complex scenes. Summary of the invention

[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a fast lane line detection method driven by feature fusion and row anchor point classification, so as to improve the real-time and accuracy of lane line detection and adapt to the needs of complex traffic scenes.

[0005] The objective of the present invention is achieved through the following technical solutions:

[0006] A fast lane line detection method driven by feature fusion and line anchor point classification includes the following steps:

[0007] S1. Perform a gridding preprocessing operation on the input image, that is, divide the image into a specific number of rows, and divide each row into a specific number of cells, where one cell is a row anchor point; an additional identification anchor point is preset in each row to indicate whether the anchor point of this row contains lane line information;

[0008] S2. Construct a lane detection model; the lane detection model includes a backbone network, a neck network, a classification network and an auxiliary segmentation network, the backbone network includes two aggregation modules and four ResNet residual blocks, the neck network includes three FPN modules of different levels, the classification network includes two fully connected layers, and the auxiliary segmentation network includes an ASPP module; the backbone network is connected to the neck network, the neck network is connected to the classification network, the auxiliary segmentation network is connected to the neck network, and is in parallel with the classification network;

[0009] S3. Use the gridded images to train the lane detection model. The specific process is as follows: the input image is subjected to feature extraction and feature fusion by the ResNet residual block and aggregation module in the backbone network to obtain feature maps of three different scales; the feature maps of three different scales are respectively input into the three FPN modules of the neck network for multi-scale feature fusion to obtain fused feature maps of three different levels; the fused feature maps of three different levels are input into the auxiliary segmentation network, and are subjected to multi-scale feature fusion processing by the ASPP module to obtain three global features containing more global information, and then the three global features are concatenated in the channel dimension, and then 1×1 convolution is used for dimensionality reduction to generate segmentation output; the highest-level fused feature map obtained by the neck network is sent to the two fully connected layers in the classification network for dimensionality reduction, and the fused feature map after dimensionality reduction is scaled to obtain the classification output;

[0010] S4. Input the image to be tested into the lane line detection model obtained after training to obtain the line anchor point classification result, that is, the lane line detection result.

[0011] Furthermore, in the training process of the lane detection model, the optimizer used is the SDG optimizer, the learning rate is 0.0001, and the total loss function of the model consists of two parts: classification loss and segmentation loss, specifically:

[0012] Loss = L cls +L seg

[0013] Among them, the classification loss L cls Used to represent the error between the classification output and the classification label in the classification network, the segmentation loss L seg It is used to represent the error between the segmentation output and the segmentation label in the auxiliary segmentation network.

[0014] Furthermore, the aggregation module includes a CANCAT layer, a 1×1 convolution layer, and a BN+ReLU layer arranged in sequence, and the input of the next module is the output of the previous module.

[0015] Furthermore, the ASPP module includes a parallel 1×1 convolution layer, three 3×3 convolution layers and a pooling layer, followed by a CANCAT layer and a 1×1 convolution layer, and the expansion rates of the three parallel 3×3 convolution layers are 6, 12, and 18, respectively.

[0016] Furthermore, the input image is subjected to feature extraction and feature fusion respectively through the ResNet residual block and the aggregation module in the backbone network to obtain feature maps of three different scales. The specific process is as follows:

[0017] The input image is sent to the first residual block to obtain output X1, X1 is processed by the second residual block to obtain output X2, X1 and X2 are input to the aggregation module for feature fusion to obtain output M1;

[0018] Perform a maximum pooling operation on M1 and adjust the size to be the same as X2. Then, add M1 and X2 vectors and input them into the third residual block to obtain output X3. X3 and M1 are input into the aggregation module for feature fusion to obtain output M2.

[0019] Similarly, a maximum pooling operation is performed on M2, and the size is adjusted to be the same as X3. After vector addition of M2 and X3, they are input into the fourth residual block to obtain the output X4, and finally three feature maps of different scales X2, X3, and X4 are obtained.

[0020] Furthermore, the three feature maps of different scales are respectively input into the three FPN modules of the neck network for multi-scale feature fusion to obtain three fused feature maps of different levels. The specific process is as follows:

[0021] The three feature maps X2, X3, and X4 of different scales extracted from the backbone network are input into the neck network; the neck network is composed of the first FPN block, the second FPN block, and the third FPN block connected in sequence;

[0022] After upsampling X4, the output fused feature map P1 of the same size as the previous level feature X3 is obtained. The fused feature map P1 is spliced ​​with X3 in the channel dimension and subjected to a convolution to obtain a new fused feature map P2. The fused feature map P1 is then upsampled to obtain a feature map of the same size as X2, which is then spliced ​​with X2 in the channel dimension and subjected to a convolution to obtain a new fused feature map P3.

[0023] Furthermore, when the three global features are concatenated in the channel dimension, the sizes of the three global feature maps are adjusted to the size of the lowest-level feature map image through bilinear interpolation.

[0024] The beneficial effects of the present invention are:

[0025] The present invention performs gridding preprocessing operations on the input image, divides multiple line anchor points in the image, and converts the lane line detection task into a line anchor point classification task, thereby significantly reducing the amount of calculation and improving the detection speed. A lane line detection model is constructed and trained to complete the line anchor point classification task, wherein the backbone network integrates the aggregation module and the ResNet residual block, extracts different scale feature maps from the image, and can deeply obtain the shallow texture features and deep abstract features of the image; the auxiliary segmentation network further performs multi-scale feature fusion processing on the different scale feature maps through the ASPP module to obtain a global feature map containing more global information, so that the lane line detection model has a better ability to perceive global information and improves the effect of lane line detection in complex environments; the segmentation loss is introduced into the total loss function during the training process to promote the accuracy of the classification network. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A schematic diagram of row anchor point classification provided by an embodiment of the present invention;

[0027] Figure 2 A schematic diagram of a network structure of a lane detection model provided by an embodiment of the present invention;

[0028] Figure 3 A network structure diagram of an aggregation module provided for an embodiment of the present invention;

[0029] Figure 4 An ASPP module network structure diagram provided for an embodiment of the present invention;

[0030] Figure 5 This is an example of the lane line detection process of the present invention, including detection result diagrams in multiple scenarios;

[0031] Figure 6 This is a comparison chart of the reasoning time of the present invention and the mainstream method. DETAILED DESCRIPTION

[0032] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0033] See also Figure 1-Figure 6 , the present invention provides a technical solution:

[0034] A fast lane line detection method driven by feature fusion and line anchor point classification includes the following steps:

[0035] S1. Perform grid preprocessing on the input image, that is, divide the image into a specific number of rows, and divide each row into a specific number of cells, where one cell is a row anchor point; an additional identification anchor point is preset in each row to indicate whether the anchor point of this row contains lane line information.

[0036] In the gridding process, assuming that the image size is H×W, the number of rows divided is much smaller than the image size H; the number of row anchor points divided in each row is much smaller than the image size W. Figure 1 , after being gridded, the image is composed of many row anchor points, and the lane line part is also included in the row anchor points. When classifying anchor points, the anchor points are classified row by row. The classification identifies the anchor points containing lane line information in each row, and merges the row anchor points belonging to the same lane line. If the anchor points in the same row contain multiple lane line information, the anchor points are classified and the anchor points belonging to the same lane line are classified into one category. If no anchor points in a row contain lane line information, the identification anchor point is set to a specific value to indicate that no anchor points in the row contain lane line information.

[0037] In a specific implementation case of the present invention, the data used comes from the CULane dataset. The image is gridded by a preset number of row anchor points to reduce the amount of calculation and improve the lane line detection speed. The purpose of the image gridding operation is to use row anchor points to represent lane lines, thereby converting the lane line detection task into a classification problem of row anchor points in the image. Compared with pixel-by-pixel segmentation, this classification method greatly reduces the amount of calculation and effectively improves the lane line detection speed.

[0038] S2. Construct a lane line detection model. The lane line detection model is as follows: Figure 2 As shown, it includes a backbone network, a neck network, a classification network and an auxiliary segmentation network. The backbone network includes two aggregation modules and four ResNet residual blocks, the neck network includes three FPN modules at different levels, the classification network includes two fully connected layers, and the auxiliary segmentation network includes an ASPP module; the backbone network is connected to the neck network, the neck network is connected to the classification network, the auxiliary segmentation network is connected to the neck network, and is in parallel with the classification network.

[0039] The purpose of the backbone network is to extract global features from the image and obtain feature maps of different scales. This example uses the ResNet-18 model pre-trained for the ImageNet image classification task as the feature extraction network. The aggregation module is integrated into ResNet-18 to enhance feature extraction. The four outputs obtained by processing the image through the aggregation network and four residual blocks represent the shallow texture features and deep abstract features of the image from shallow to deep.

[0040] Furthermore, the network structure diagram of the aggregation module is as follows: Figure 3 As shown in the figure, it includes the CANCAT layer, 1×1 convolution layer, BN+ReLU, and the input of the next module is the output of the previous module. The aggregation module receives two inputs, concatenates the two inputs in channel dimension, and then performs 1×1 convolution to reduce the dimension to achieve the purpose of feature fusion.

[0041] The network structure diagram of the ASPP module in the auxiliary segmentation network is as follows Figure 4 As shown, it includes a parallel 1×1 convolution layer, three 3×3 convolution layers and a pooling layer, followed by a CANCAT layer and a 1×1 convolution layer. The expansion rates of the three parallel 3×3 convolution layers are 6, 12, and 18, respectively.

[0042] S3. Use the gridded images to train the lane detection model. The specific process is as follows: the input image is subjected to feature extraction and feature fusion by the ResNet residual block and aggregation module in the backbone network to obtain feature maps of three different scales; the feature maps of three different scales are respectively input into the three FPN modules of the neck network for multi-scale feature fusion to obtain fused feature maps of three different levels; the fused feature maps of three different levels are input into the auxiliary segmentation network, and are subjected to multi-scale feature fusion processing by the ASPP module to obtain three global features containing more global information, and then the three global features are concatenated in the channel dimension, and then 1×1 convolution is used for dimensionality reduction to generate segmentation output; the highest-level fused feature map obtained by the neck network is sent to the two fully connected layers in the classification network for dimensionality reduction, and the fused feature map after dimensionality reduction is scaled to obtain the classification output;

[0043] In this embodiment, during the training process of the lane detection model, the optimizer used is the SDG optimizer, the learning rate is 0.0001, and the total loss function of the model consists of two parts: classification loss and segmentation loss, specifically:

[0044] Loss = L cls +L seg

[0045] Among them, the classification loss L cls Used to represent the error between the classification output and the classification label in the classification network, the segmentation loss L seg It is used to represent the error between the segmentation output and the segmentation label in the auxiliary segmentation network. Segmentation loss L seg The Loss function is used to improve the accuracy of the classification network. The optimization goal is to obtain more accurate lane line texture details. The segmentation output in the auxiliary segmentation network is only used during training and is disabled during inference.

[0046] The input image in step S3 is subjected to feature extraction and feature fusion respectively by the ResNet residual block and the aggregation module in the backbone network to obtain feature maps of three different scales. The specific process is as follows:

[0047] The input image is sent to the first residual block to obtain output X1, X1 is processed by the second residual block to obtain output X2, X1 and X2 are input to the aggregation module for feature fusion to obtain output M1;

[0048] Perform a maximum pooling operation on M1 and adjust the size to be the same as X2. Then, add M1 and X2 vectors and input them into the third residual block to obtain output X3. X3 and M1 are input into the aggregation module for feature fusion to obtain output M2.

[0049] Similarly, a maximum pooling operation is performed on M2, and the size is adjusted to be the same as X3. After vector addition of M2 and X3, they are input into the fourth residual block to obtain the output X4, and finally three feature maps of different scales X2, X3, and X4 are obtained.

[0050] The three feature maps of different scales in step S3 are respectively input into the three FPN modules of the neck network for multi-scale feature fusion to obtain three fused feature maps of different levels. The specific process is as follows:

[0051] The three feature maps X2, X3, and X4 of different scales extracted from the backbone network are input into the neck network; the neck network is composed of the first FPN block, the second FPN block, and the third FPN block connected in sequence;

[0052] After upsampling X4, the output fused feature map P1 of the same size as the previous level feature X3 is obtained. The fused feature map P1 is spliced ​​with X3 in the channel dimension and subjected to a convolution to obtain a new fused feature map P2. The fused feature map P1 is then upsampled to obtain a feature map of the same size as X2, which is then spliced ​​with X2 in the channel dimension and subjected to a convolution to obtain a new fused feature map P3.

[0053] Furthermore, when the three global features are concatenated in the channel dimension, the sizes of the three global feature maps are adjusted to the size of the lowest-level feature map image through bilinear interpolation.

[0054] In a specific embodiment, when the input image size in the backbone network is 288×800, the sizes of the four feature maps X1, X2, X3, and X4 are 288×800, 144×400, 72×200, and 36×100, respectively. The sizes of the feature maps output by the four ResNet residual blocks are reduced by 2 times, 4 times, 8 times, and 16 times respectively relative to the input image size, and the corresponding number of channels is 64, 128, 256, and 512, respectively. When the three feature maps X2, X3, and X4 of different scales are input into the neck network for upsampling, the scaling factor used is 2, and the sampling method is bilinear interpolation. Finally, the number of channels of the fused feature maps P1, P2, and P3 are 128, 256, and 512, respectively.

[0055] S4. Input the image to be tested into the lane line detection model obtained after training to obtain the line anchor point classification result, that is, the lane line detection result. The detection result is as follows Figure 5 As shown in Figure 2, the time consumption of reasoning of the present invention and the mainstream method is compared. Figure 6 The results show that the method proposed in the present invention significantly reduces the detection time while ensuring the accuracy of lane line detection, and meets the requirements of real-time performance and accuracy.

[0056] The present invention first performs a gridding operation on the image. After the gridding operation, the image becomes composed of multiple grids, each of which is a row anchor point. The lane line is represented by the row anchor point, thereby converting the lane line detection from a pixel-by-pixel segmentation task to a row anchor point classification task, significantly reducing the amount of calculation and improving the detection speed. Subsequently, a neural network is used to extract image features, which mainly includes four parts: a backbone network, a neck network, a classification network, and an auxiliary segmentation network. The backbone network receives image input and extracts features of multiple different levels; the features of different levels are input into the neck network, and the neck network performs multi-scale feature fusion on the features of different levels, and outputs multiple different level features from shallow to deep. Subsequently, all the different level features output by the neck network are input into the auxiliary segmentation network for segmentation to improve the detection accuracy of the overall network. The deepest layer features output by the neck network are directly input into the classification network, and after dimension adjustment, row anchor point classification is performed to obtain a classification output. The auxiliary segmentation network is only used during model training to help improve model accuracy; it is disabled during model inference in order to improve the actual inference speed of the model.

[0057] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.

Claims

1. A fast lane line detection method driven by feature fusion and line anchor point classification, characterized in that: The following steps are involved: S1. Perform a gridding preprocessing operation on the input image, that is, divide the image into a specific number of rows, and divide each row into a specific number of cells, where one cell is a row anchor point; an additional identification anchor point is preset in each row to indicate whether the anchor point of this row contains lane line information; S2. Construct a lane detection model; the lane detection model includes a backbone network, a neck network, a classification network and an auxiliary segmentation network, the backbone network includes two aggregation modules and four ResNet residual blocks, the neck network includes three FPN modules of different levels, the classification network includes two fully connected layers, and the auxiliary segmentation network includes an ASPP module; the backbone network is connected to the neck network, the neck network is connected to the classification network, the auxiliary segmentation network is connected to the neck network, and is in parallel with the classification network; S3. Use the gridded images to train the lane detection model. The specific process is as follows: the input image is subjected to feature extraction and feature fusion by the ResNet residual block and aggregation module in the backbone network to obtain feature maps of three different scales; the feature maps of three different scales are respectively input into the three FPN modules of the neck network for multi-scale feature fusion to obtain fused feature maps of three different levels; the fused feature maps of three different levels are input into the auxiliary segmentation network, and are subjected to multi-scale feature fusion processing by the ASPP module to obtain three global features containing more global information, and then the three global features are concatenated in the channel dimension, and then 1×1 convolution is used for dimensionality reduction to generate segmentation output; the highest-level fused feature map obtained by the neck network is sent to the two fully connected layers in the classification network for dimensionality reduction, and the fused feature map after dimensionality reduction is scaled to obtain the classification output; S4. Input the image to be tested into the lane line detection model obtained after training to obtain the line anchor point classification result, that is, the lane line detection result.

2. The fast lane line detection method driven by feature fusion and line anchor point classification according to claim 1 is characterized by: In the training process of the lane detection model, the optimizer used is the SDG optimizer, the learning rate is 0.0001, and the total loss function of the model consists of two parts: classification loss and segmentation loss, specifically: Loss=L cls +L seg Among them, the classification loss L cls Used to represent the error between the classification output and the classification label in the classification network, the segmentation loss L seg It is used to represent the error between the segmentation output and the segmentation label in the auxiliary segmentation network.

3. The fast lane line detection method driven by feature fusion and line anchor point classification according to claim 1, characterized in that: The aggregation module includes a CANCAT layer, a 1×1 convolution layer, and a BN+ReLU layer which are arranged in sequence, and the input of the next module is the output of the previous module.

4. The fast lane line detection method driven by feature fusion and line anchor point classification according to claim 1 is characterized in that: The ASPP module includes a parallel 1×1 convolution layer, three 3×3 convolution layers and a pooling layer, followed by a CANCAT layer and a 1×1 convolution layer, and the expansion rates of the three parallel 3×3 convolution layers are 6, 12, and 18, respectively.

5. The fast lane line detection method driven by feature fusion and line anchor point classification according to claim 1, characterized in that: The input image is subjected to feature extraction and feature fusion respectively through the ResNet residual block and the aggregation module in the backbone network to obtain feature maps of three different scales. The specific process is as follows: The input image is sent to the first residual block to obtain output X1, X1 is processed by the second residual block to obtain output X2, X1 and X2 are input to the aggregation module for feature fusion to obtain output M1; Perform a maximum pooling operation on M1 and adjust the size to be the same as X2. Then, add M1 and X2 vectors and input them into the third residual block to obtain output X3. X3 and M1 are input into the aggregation module for feature fusion to obtain output M2. Similarly, a maximum pooling operation is performed on M2, and the size is adjusted to be the same as X3. After vector addition of M2 and X3, they are input into the fourth residual block to obtain the output X4, and finally three feature maps of different scales X2, X3, and X4 are obtained.

6. The fast lane line detection method driven by feature fusion and line anchor point classification according to claim 1, characterized in that: The three feature maps of different scales are respectively input into the three FPN modules of the neck network for multi-scale feature fusion to obtain three fused feature maps of different levels. The specific process is as follows: The three feature maps X2, X3, and X4 of different scales extracted from the backbone network are input into the neck network; the neck network is composed of the first FPN block, the second FPN block, and the third FPN block connected in sequence; After upsampling X4, the output fused feature map P1 of the same size as the previous level feature X3 is obtained. The fused feature map P1 is spliced ​​with X3 in the channel dimension and subjected to a convolution to obtain a new fused feature map P2. The fused feature map P1 is then upsampled to obtain a feature map of the same size as X2, which is then spliced ​​with X2 in the channel dimension and subjected to a convolution to obtain a new fused feature map P3.

7. The fast lane line detection method driven by feature fusion and line anchor point classification according to claim 1, characterized in that: When the three global features are spliced ​​in the channel dimension, the sizes of the three global feature maps are adjusted to the size of the lowest level feature map image through bilinear interpolation.