A Precise Detection Method for Lane Lines in Complex Environments Based on Multi-Scale Collaborative Enhancement and Semantic Compensation

Through the multi-scale collaborative enhancement and semantic compensation methods, combined with multi-level feature maps, the robustness and computational cost problems of lane line detection in complex environments are solved, and efficient lane line detection is achieved.

CN120032334BActive Publication Date: 2025-07-22SHIJIAZHUANG TIEDAO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510190179.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-22
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing lane line detection methods are poorly robust in complex environments, insufficient integration of multi-scale features, high calculation costs, and insufficient generalization ability in unknown environments.

Method used

The multi-scale collaborative enhancement and semantic compensation methods are adopted, and the multi-scale collaborative enhancement module, a hybrid scene semantic compensation module and a boundary correction guidance module are combined with multi-level feature maps to enhance feature robustness and detailed information, and improve lane line boundary detection.

Benefits of technology

It improves the accuracy and robustness of lane line detection, adapts to detection in different scenarios, and reduces calculation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032334B_ABST
    Figure CN120032334B_ABST
Patent Text Reader

Abstract

The present invention discloses a precise lane line detection method based on multi-scale collaborative enhancement and semantic compensation. The method includes the following steps: obtaining a lane line data set and inputting it into a trained precise lane line detection network; using a backbone network to obtain multi-level features; using a multi-scale collaborative enhancement module to obtain features with enhanced interaction between spatial and semantic information; using a hybrid scene semantic compensation module to obtain supplementary semantic features to enhance the feature robustness in different scenarios; using a boundary correction guidance module to refine the features from the spatial and channel dimensions to improve the lane boundary detection effect; and optimizing and adjusting the lane line anchors layer by layer through the enhanced features to obtain the final lane line coordinate prediction sequence. The method combines multi-scale collaborative enhancement and semantic compensation, improves the accuracy of lane line detection, and enhances the detection robustness in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a precise lane line detection method in complex environments based on multi-scale collaborative enhancement and semantic compensation, belonging to the field of computer vision technology. Background Art

[0002] The rapid development of autonomous driving technology and intelligent transportation systems is profoundly changing modern transportation modes. In the successful application of these systems, the ability to accurately detect and interpret lane markings on the road is an important factor. The lane line detection system enables the vehicle to maintain the correct driving trajectory by accurately identifying lane boundaries and the road center line, and provides support for functions such as lane keeping assistance, lane departure warning, and automatic lane cruising. It is an important research direction in the field of computer vision.

[0003] Lane line detection is an important part of the autonomous driving system. Its core task is to accurately extract lane marking information from road images or sensor data to ensure that the vehicle travels along the correct trajectory. Traditional methods mainly rely on techniques such as edge detection, color segmentation, and geometric fitting, but are easily interfered in complex environments such as changes in lighting, road wear, and occlusion, resulting in poor robustness. In recent years, the successful application of convolutional neural networks in the field of computer vision has promoted the development of lane line detection technology. Compared with traditional methods based on handcrafted features, neural networks can automatically learn and extract deep features of lane markings, improving the accuracy and robustness of detection. By constructing a deep network structure, neural networks can extract information at different levels from the input image, such as edge, texture, and shape features, so as to more effectively identify lane lines and still maintain good detection performance even in complex environments such as changes in lighting, occlusion, or blurred lane markings.

[0004] Although the lane line detection method based on deep learning has made certain progress in terms of accuracy and adaptability, there are still the following problems. First, the detection difficulties brought about by the diversity of road scenes, including changes in lane morphology, marking styles, and occlusion or interference of environmental factors on lane markings. Since model training is usually based on a limited road dataset, their generalization ability in unknown environments is poor, resulting in an inability to stably handle various actual road conditions. Second, existing methods ignore the integration of global context information. Shallow features lack high-level semantic information, and deep features lack detailed information. The effective integration of multi-scale features still needs to be explored. Finally, the computational cost of existing lane line detection models is relatively high. The autonomous driving system needs to process a large amount of sensor data, but the computing resources are limited. Summary of the Invention

[0005] The purpose of the present invention is to solve the above problems in existing methods and propose a precise lane line detection method in complex environments based on multi-scale collaborative enhancement and semantic compensation.

[0006] To achieve the above object, the technical solution of the present invention is as follows:

[0007] A precise lane line detection method based on multi-scale collaborative enhancement and semantic compensation in complex environments, characterized by including the following steps:

[0008] S1: Obtain a lane line detection data set and input it into the backbone network to obtain the feature maps of the last three levels of the network, denoted as Fi, where i represents the level of the feature, 3 ≤ i ≤ 5;

[0009] S2: Use the multi-scale collaborative enhancement module to perform spatial information enhancement and semantic information interaction aggregation on the feature maps of different levels through a top-down spatial compensation path and a bottom-up semantic filling path to improve the feature information of the feature maps extracted by the feature network. This module includes a soft upsampling combination sub-module SUpCom constructed by GSConv convolutional layers and a GSConv convolutional combination sub-module GSConvCom for feature extraction;

[0010] S3: Use the hybrid scene semantic compensation module to adaptively route and select the propagation path and extract compensation semantic information through the progressive compensation extraction layer, and fuse it with the improved feature map to enhance the feature robustness in different scenarios and improve the scene adaptation ability of the feature map. This module includes an adaptive routing sub-module and a progressive compensation extraction layer sub-module;

[0011] S4: Input the feature map with improved scene robustness into the boundary correction guidance module to further focus on the important information in the feature in terms of spatial and channel dimensions, refine the lane line feature information, and improve the lane line boundary detection effect. This module includes an improved spatial attention sub-module and a channel attention sub-module;

[0012] S5: Pass the refined feature maps of different levels through the detection head Ha, where a represents the level of the feature, 1 ≤ a ≤ 3, and transfer the prefabricated prior lane anchor Pc, where c represents the number of optimization times, 0 ≤ c ≤ 3, to the first detection head H1;

[0013] S6: Gradually adjust the parameters of the prior lane anchor Pc downward. Finally, the prior lane anchor is adjusted by the last detection head and outputs the optimized anchor P3 as the predicted output result of the lane line.

[0014] A further technical solution is that the multi-scale collaborative enhancement module performs spatial information enhancement and semantic information interaction aggregation on the feature maps of different levels extracted; the multi-scale collaborative enhancement module consists of a soft upsampling combination sub-module SUpCom constructed based on the GSConv convolutional layer, a GSConv convolutional combination sub-module GSConvCom, and a top-down spatial compensation path and a bottom-up semantic filling path built, so as to obtain a feature map containing richer semantic and detailed feature information.

[0015] Furthermore, the GSConv convolutional layer captures image features through a standard convolutional layer Conv, divides the feature map into channels, obtains feature maps X1 and X2, then performs convolution on feature map X1 and channel relationship modeling on X2 respectively, and then performs channel splicing and channel rearrangement Shuffle on the feature maps, so as to effectively improve the expression ability of the feature map, enhance the feature interaction between channels and reduce the computational overhead; the soft upsampling combination module SUpCom includes C3k2, C2fGS2 and SNI sub-modules, where C3k2 is a convolutional layer that replaces the Bottleneck module in C2f with a C3k convolutional layer, C2fGS2 is a convolutional layer that replaces the Bottleneck module in C2f with a GSConv convolutional layer, and SNI is the weighting of the upsampling layer and the influencing factor, which can effectively retain gradient information and interact with feature information of different levels to avoid feature loss during the upsampling process; the GSConv convolutional combination sub-module GSConvCom includes C3k2 convolutional layer and C2fGS2 convolutional layer sub-modules, which can enrich gradient information and extract higher-level semantic information; Split represents channel division, Conv represents the standard convolutional layer, DCN represents deformable convolution, DSConv represents depthwise separable convolution, SENet represents squeeze-and-excitation network, [·] represents channel splicing, Shuffle represents channel rearrangement operation, X and R represent input features, and their specific calculation formulas are as follows:

[0016] X1,X2 = Split(Conv(X)),

[0017] GSConv = Shuffle([DSConv(DCN(X1)), SENet(X2)]),

[0018] GSConvCom = C2fGS2(C3k2(C2fGS2(R))),

[0019] SUpCom = SNI(C2fGS2(C3k2(R))).

[0020] Further, the top-down spatial compensation path adjusts the size of the feature map F5 to the size of the feature map F4 through the soft upsampling combination module SUpCom and performs channel concatenation with the feature map F4 to model multi-scale information and obtain a feature map. Adjust it to the size of the feature map F3 to obtain a feature map. To supplement the semantic information lost during propagation, the feature map Extracts the aggregated multi-scale feature information through the GSConv convolution combination module GSConvCom and adjusts it to the size of the feature map F3 through the soft upsampling combination module SUpCom and concatenates it with the feature map F3 and the feature map Performs channel concatenation to generate a feature map. Thus, while retaining high-level semantic information, the detailed information of low-level features is enhanced; Cat represents the channel concatenation operation, and its specific calculation formula is as follows:

[0021]

[0022] Further, the bottom-up semantic filling path processes the feature map Captures more context and semantic information through the GSConv convolution combination module GSConvCom to obtain a feature map As the feature output of the third level and passes it upward; the feature map The feature map generated by GSConv convolution is concatenated with the feature map generated by the GSConv convolution combination module GSConvCom Performs channel concatenation and generates a feature map For the feature map The feature map obtained through the GSConv convolution combination module GSConvCom As the feature output of the fourth level and passes it upward, for the feature map Feature map Feature map Captures rich spatial information through GSConv convolution and performs channel concatenation with the feature map F5 to generate a feature map For the feature map The feature map generated by the GSConv convolution combination module GSConvCom As the feature output of the last level, realizing the gradual transmission of semantic information to the bottom layer; Cat represents the channel concatenation operation, and its specific calculation formula is as follows:

[0023]

[0024] Furthermore, the technical solution lies in that the hybrid scene semantic compensation module includes an adaptive routing sub-module and a progressive compensation extraction layer sub-module, which outputs appropriate path weights for samples through adaptive routing and weights them onto the feature maps extracted by the progressive compensation extraction layer sub-module, so as to obtain rich scene feature information to improve the scene robustness effect.

[0025] Furthermore, the hybrid scene semantic compensation module is used to improve the context and semantic information in different scenes and make up for the feature information lost in the feature extraction process.

[0026] Furthermore, for the adaptive routing, the input feature map X is subjected to feature extraction through a convolutional layer and combined with channel attention CA to perform channel modeling on the features to generate multi-scale features X′, so as to enhance the global correlation between features; subsequently, through global average pooling GAP, linear mapping Linear, and uniform noise Noise processing, a probabilistic output Prob for feature selection is generated n where n represents the number of blocks, 0 ≤ n ≤ K, to determine the optimal semantic compensation path; its specific calculation formula is as follows:

[0027] X′ = Conv3(X) + CA(Conv3(X)),

[0028] Prob n = Linear(GAP(X′)) + Noise, n ∈ 0, 1, 2…, K.

[0029] Furthermore, the semantic compensation path is a progressive compensation extraction layer, including two shared blocks S m where m represents the number of layers, 1 ≤ m ≤ 2 and multiple independent blocks where K represents the number of blocks, m represents the number of layers, 1 ≤ K ≤ 3, 1 ≤ m ≤ 2, to build a two-layer progressive enhancement mechanism; the first layer contains one shared block S 1 and several independent blocks where 1 represents the first layer, K represents the number of blocks, 1 ≤ K ≤ 3, including K groups of residual blocks composed of CBR and C3k2, and the second layer contains one shared block S 2 and several independent blocks where K represents the number of blocks, 1 ≤ K ≤ 3, including a convolutional block composed of a group of CBR and C3k2; the input feature map passes through the first layer to obtain feature maps with different depths where the subscript bi represents the i-th block, and the superscript 1 represents the first layer, and the feature maps with different depths are concatenated in channels, and then added to the feature information extracted by the shared block to obtain the feature information Out2, so as to enhance the basic features of different paths and compensate for the information loss of other paths with different depths; the feature information extracted by the first layer is passed to the second layer to capture multi-scale feature information, by It is represented and weighted and fused with the path probability generated by adaptive routing to obtain the supplementary semantic information Z with the highest sample correlation; Sum represents the summation operation, [·] represents channel splicing, and Convk represents a convolutional layer with a convolution kernel of k×k. The specific calculation formula is as follows:

[0030]

[0031] A further technical solution lies in that the boundary correction guidance module effectively optimizes the feature detail information by improving the spatial attention mechanism and introducing the channel attention mechanism.

[0032] Furthermore, the boundary correction guidance module is used to refine the lane line feature information and improve the lane line boundary detection effect.

[0033] Furthermore, the spatial attention mechanism is improved. The input feature map X is divided into r groups through channels, and each group of feature maps is denoted as X r , where r represents the number of groups. Through an X-axis average pooling branch and a Y-axis average pooling branch, and then for the outputs of the two branches, denoted as X x and X y , where x represents the X-axis direction or the horizontal direction, and y represents the Y-axis direction or the vertical direction, are channel-spliced to generate a feature map with horizontal and vertical dimensions, and important spatial features are extracted through a multi-core one-dimensional convolution module. The extracted features are split and restored to the shapes of the horizontal and vertical dimension feature maps, and after channel-splicing each group of feature maps, the spatial weights in the horizontal and vertical dimensions are obtained through the Sigmoid layer, and then multiplied with the input feature map to obtain the spatially enhanced feature maps X2 in the horizontal and vertical directions; the input feature map X is used to generate a feature map S through a 3×3 convolutional layer and multiplied with the feature map X2 after group normalization GroupNorm and global average pooling GAP to obtain the spatial weight W1; the feature map X2 after group normalization GroupNorm is multiplied with the feature map S after global average pooling GAP to obtain the spatial weight W2; the spatial weight W1 and the spatial weight W2 are added and the superimposed spatial weight is obtained through Sigmoid, and then weighted to the feature map X2 to generate the spatially enhanced feature map T, so as to enhance the correlation between different spatial positions through the cross-spatial interaction enhancement network; σ is the sigmoid function operation, [·] represents the channel-splicing operation, ⊙ represents the Hadamard product operation, CS r represents the operation of dividing the channels into r groups, GAP x represents the average pooling operation in the horizontal direction, GAP Y represents the average pooling operation in the vertical direction, GAP represents global average pooling, ConvGroup [3,5,7,9]Represents one-dimensional convolutional layer operations with multiple convolutional kernel sizes of 3, 5, 7, and 9, Cat group Represents the operation of concatenating channels by group, CBR 3×3 Represents the combined operation of 3×3 convolution, batch normalization, and rectified linear unit. Reshape represents the shape reshaping operation, and its specific calculation formula is as follows:

[0034] X r = CS r (X),

[0035] X x = GAP x (X r ),

[0036] X y = GAP Y (X r ),

[0037] X2 = σ(Cat group (Reshape(ConvGroup [3,5,7,9] ([X x , X y )))) ⊙ X,

[0038] S = CBR 3×3 (X),

[0039] W1 = S × GAP(GroupNorm(X2)),

[0040] W2 = GAP(S) × GroupNorm(X2),

[0041] T = X2 × σ(W1 + W2).

[0042] Furthermore, a channel attention mechanism is introduced. The input feature X is passed through the GhostConv layer to obtain the response in the channel dimension, and the channel relationship is modeled through the partial self-attention mechanism PSA to generate the channel dimension attention weights, which are weighted channel by channel into the feature map to generate the feature map C after channel attention. Then, the feature maps after spatial and channel dimension attention are added to generate the final feature map Z; its specific calculation formula is as follows:

[0043] C = σ(GAP(PSA(GhostConv(GAP(X))))) ⊙ T,

[0044] Z = C + T.

[0045] A further technical solution lies in that the training steps of the trained lane line detection network include:

[0046] Construct a lane line detection network;

[0047] Construct a training set, where the training set is a sequence of video frames and a sequence of true coordinates of lane lines;

[0048] Input the training set into the lane line detection network for training;

[0049] The lane line detection network outputs a sequence of predicted lane coordinates;

[0050] Calculate the difference between the predicted sequence and the true coordinate sequence and backpropagate;

[0051] When the loss value reaches the minimum, the model converges, stops training, and obtains a trained lane line detection network.

[0052] The beneficial effects of adopting the above technical solutions are as follows: The present invention provides a multi-scale collaborative enhancement module, which fully combines the semantic information and detailed information of objects, and helps with positioning and detection; the present invention designs a hybrid scene semantic compensation module for compensating the spatial and semantic information lost during the feature network extraction process and improving the detection robustness problem in different scenes; the present invention develops a boundary correction guidance module, which improves the effect of captured feature detailed information and the detection effect of lane line boundaries. The three modules adopted are integrated in the network, greatly improving the accuracy of lane line detection and reflecting the advantages of the proposed technical solution. Brief Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be further described in detail below with reference to the drawings.

[0054] Figure 1 It is the overall flowchart of the network of the embodiment of the present invention;

[0055] Figure 2 It is the overall architecture diagram of the network of the embodiment of the present invention;

[0056] Figure 3 It is the sub-module structure diagram in the multi-scale collaborative enhancement module of the embodiment of the present invention;

[0057] Figure 4 It is the structure diagram of the multi-scale collaborative enhancement module of the embodiment of the present invention;

[0058] Figure 5 It is the structure diagram of the hybrid scene semantic compensation module of the embodiment of the present invention;

[0059] Figure 6 It is the boundary correction guidance module of the embodiment of the present invention;

[0060] Figure 7 It is the result diagram of the embodiment of the present invention. Detailed Embodiments

[0061] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0062] The present invention provides a precise lane line detection method in a complex environment based on multi-scale collaborative enhancement and semantic compensation, as Figure 1 shown, including the following steps:

[0063] S1: Construct a backbone network to obtain multi-level features; use DLA34 as the backbone network to obtain multi-level features from the input image, denoted as F i , where 3 ≤ i ≤ 5, and i represents the level of the feature.

[0064] S2: Construct a multi-scale collaborative enhancement module, the structure of which is shown in Figure 3 and Figure 4 ;

[0065] S2-1: The multi-scale collaborative enhancement module enhances the features of the feature maps at different levels extracted and captures the features of different receptive fields, including a soft upsampling combination sub-module SUpCom constructed based on the GSConv convolutional layer, a GSConv convolutional combination sub-module GSConvCom, and a top-down spatial compensation path and a bottom-up semantic filling path built.

[0066] S2-2: The GSConv convolutional layer, the structure of which is shown in Figure 3 a, captures the image features through a standard convolutional layer Conv, divides the feature map into channels to obtain feature maps X1 and X2, then performs convolution on the feature map X1 and models the channel relationship of X2 respectively, and then splices and rearranges the channels of the feature maps Shuffle, so as to effectively improve the expression ability of the feature map, enhance the feature interaction between channels and reduce the computational overhead; Split represents channel division, Conv represents the standard convolutional layer, DCN represents the deformable convolutional layer, DSConv represents the depthwise separable convolutional layer, SENet represents the squeeze-and-excitation network, [·] represents channel splicing, Shuffle represents the channel rearrangement operation, X represents the input feature, and its specific calculation formula is as follows:

[0067] X1,X2 = Split(Conv(X)),

[0068] GSConv = Shuffle([DSConv(DCN(X1)),SENet(X2)]).

[0069] S2-3: The soft upsampling combination module SUpCom includes the C3k2, C2fGS2, and SNI sub-modules. Its structure is shown in Figure 3 Figure b, where C3k2 is a convolutional layer that replaces the convolutional layer in the Bottleneck module in C2f, C2fGS2 is a GSConv convolutional layer that replaces the convolutional layer in the Bottleneck module in C2f, and SNI is the weighting of the upsampling layer and the influencing factor, which can effectively retain gradient information and interact with feature information at different levels to avoid feature loss during the upsampling process; the GSConv convolutional combination sub-module GSConvCom, its structure is shown in Figure 3 Figure c, which includes the C3k2 convolutional layer and the C2fGS2 convolutional layer sub-module, and can enrich gradient information and extract higher-level semantic information; R represents the input feature, and its specific calculation formula is as follows:

[0070] GSConvCom = C2fGS2(C3k2(C2fGS2(R))),

[0071] SUpCom = SNI(C2fGS2(C3k2(R))).

[0072] S2-4: The top-down spatial compensation path, its structure is shown in Figure 4 Figure a. The feature map F5 is adjusted to the size of the feature map F4 through the soft upsampling combination module SUpCom and is concatenated with the feature map F4 in the channel dimension to model multi-scale information to obtain the feature map Adjusted to the size of the feature map F3 to obtain the feature map To supplement the semantic information lost during propagation, the feature map Extracts the aggregated multi-scale feature information through the GSConv convolutional combination module GSConvCom and is adjusted to the size of the feature map F3 through the soft upsampling combination module SUpCom and is concatenated with the feature map F3 and the feature map In the channel dimension to generate the feature map Thus, while retaining high-level semantic information, the detailed information of low-level features is enhanced; its specific calculation formula is as follows:

[0073]

[0074] S2-5: The bottom-up semantic filling path, its structure is shown in Figure 4 Figure b. The feature map Captures more context and semantic information through the GSConv convolutional combination module GSConvCom to obtain the feature map As the feature output of the third level and is passed upward; the feature map The feature map generated by the GSConv convolution and the feature map generated by the GSConv convolution combination module GSConvCom are concatenated in channels and a feature map is generated The feature map The feature map obtained through the GSConv convolution combination module GSConvCom is used as the feature output of the fourth level and passed upward. The feature map Feature map Feature map captures rich spatial information through the GSConv convolution and is concatenated in channels with the feature map F5 to generate a feature map The feature map The feature map generated by the GSConv convolution combination module GSConvCom is used as the feature output of the last level to realize the gradual transmission of semantic information to the bottom layer; Cat represents the channel concatenation operation, and its specific calculation formula is as follows:

[0075]

[0076] S3: Construct a hybrid scene semantic compensation module, and its structure is shown in Figure 5 ;

[0077] S3-1: Adaptive routing. The input feature map X is subjected to feature extraction through a convolutional layer and the feature is modeled in channels by combining channel attention CA to generate multi-scale features X′ to enhance the global correlation between features; subsequently, through global average pooling GAP, linear mapping Linear and uniform noise Noise processing, a probability output Prob n is generated, where n represents the number of blocks, 0 ≤ n ≤ K, to determine the optimal semantic compensation path; Convk represents a convolutional layer with a convolution kernel of k×k, and its specific calculation formula is as follows:

[0078] X′ = Conv3(X) + CA(Conv3(X)),

[0079] Prob n = Linear(GAP(X′)) + Noise, n ∈ 0, 1, 2…, K.

[0080] S3-2: Progressive compensation extraction layer, including two shared blocks S m , where m represents the number of layers, 1 ≤ m ≤ 2 and multiple independent blocks where K represents the number of blocks, m represents the number of layers, 1 ≤ K ≤ 3, 1 ≤ m ≤ 2, to construct a two-layer progressive enhancement mechanism; the first layer contains a shared block S 1 , and several independent blocks Among them, 1 represents the first layer, K represents the number of blocks, 1 ≤ K ≤ 3, including K groups of residual blocks composed of CBR and C3k2. The second layer includes a shared block S 2 and several independent blocks Among them, K represents the number of blocks, 1 ≤ K ≤ 3, including a convolutional block composed of a group of CBR and C3k2; the input feature map is used to obtain feature maps of different depths through the first layer Among them, the subscript bi represents the i-th block, and the superscript 1 represents the first layer. The feature maps of different depths are concatenated along the channel dimension, and then the feature information extracted by the shared block is added to the feature maps to obtain the feature information Out2, so as to enhance the basic features of different paths and compensate for the information loss of other paths with different depths; the feature information extracted by the first layer is passed to the second layer to capture multi-scale feature information, which is represented by and weighted fusion with the path probability generated by the adaptive routing to obtain the supplementary semantic information Z with the highest sample correlation; Sum represents the summation operation, [·] represents the channel concatenation, and Convk represents the convolutional layer with a convolution kernel of k×k. The specific calculation formula is as follows:

[0081]

[0082] S4: Construct a boundary correction guidance module, the structure of which is shown in Figure 6 ;

[0083] S4-1: Improve the spatial attention mechanism. The input feature map X is divided into r groups along the channel dimension. The feature maps of each group are denoted as X r , where r represents the number of groups. Through an X-axis average pooling branch and a Y-axis average pooling branch, and then the outputs of the two branches are denoted as X x and X y, where x represents the X-axis direction or the horizontal direction, and y represents the Y-axis direction or the vertical direction. Channel concatenation is performed to generate a feature map with horizontal and vertical dimensions, and important spatial features are extracted through a multi-core one-dimensional convolution module. The extracted features are split and restored to the shapes of the horizontal and vertical dimension feature maps, and after channel concatenation of each group of feature maps, spatial weights in the horizontal and vertical dimensions are obtained through a Sigmoid layer. Then, the spatial weights are multiplied with the input feature map to obtain the spatially enhanced feature maps X2 in the horizontal and vertical directions; the input feature map X is passed through a 3×3 convolutional layer to generate a feature map S, and the feature map S is multiplied with the feature map X2 after group normalization (GroupNorm) and global average pooling (GAP) to obtain the spatial weight W1; the feature map X2 after group normalization (GroupNorm) is multiplied with the feature map S after global average pooling (GAP) to obtain the spatial weight W2; the spatial weight W1 and the spatial weight W2 are added and passed through a Sigmoid to obtain the superimposed spatial weight, and then the superimposed spatial weight is weighted to the feature map X2 to generate the spatially enhanced feature map T, thereby enhancing the correlation between different spatial positions through the cross-spatial interaction enhancement network; σ represents the sigmoid function operation, [·] represents the channel concatenation operation, and ⊙ represents the Hadamard product operation, CS r represents the operation of dividing the channels into r groups, GAP x represents the average pooling operation in the horizontal direction, GAP Y represents the average pooling operation in the vertical direction, GAP represents global average pooling, GonvGroup [3,5,7,9] represents the operation of one-dimensional convolutional layers with multiple convolution kernel sizes of 3, 5, 7, and 9, Cat group represents the operation of channel concatenation by group, CBR 3×3 represents the combined operation of 3×3 convolution, batch normalization, and rectified linear unit, Reshape represents the shape reshaping operation, and its specific calculation formula is as follows:

[0084] X r =CS r (X),

[0085] X x =GAP x (X r ),

[0086] X y =GAP Y (X r ),

[0087] X2=σ(Cat group (Reshape(ConvGroup [3,5,7,9] ([X x ,X y ))))⊙X,

[0088] S = CBR 3×3 (X),

[0089] W1 = S × GAP(GroupNorm(X2)),

[0090] W2 = GAP(S) × GroupNorm(X2),

[0091] T = X2 × σ(W1 + W2).

[0092] S4-2: Introduce the channel attention mechanism. Obtain the channel-dimensional response of the input feature X through the GhostConv layer of phanton convolution, and model the channel relationship through the partial self-attention mechanism PSA to generate the channel-dimensional attention weights, which are weighted channel by channel into the feature map to generate the feature map C with channel attention. Then, add the feature maps with spatial and channel attention to generate the final feature map Z; its specific calculation formula is as follows:

[0093] C = σ(GAP(PSA(GhostConv(GAP(X))))) ⊙ T,

[0094] Z = C + T.

[0095] S5: Construct a detection head, which includes a self-attention mechanism and a fully connected layer. Focus on the lane line features of the lane line anchors through the self-attention mechanism and map them to the coordinate sequence parameters through the fully connected layer.

[0096] S6: Construct a lane line detection network, and its structure is shown in Figure 2 , and conduct training;

[0097] S6-1: Construct a training set, which is a sequence of video frames and its corresponding ground truth coordinate sequence of lane lines. Two widely used datasets are used for training: CULane and CurveLanes. Among them, CULane is a commonly used public dataset for lane line detection, which contains 88,880 training set images and a test set of 8 challenging scenarios. CurveLanes is a more challenging public dataset, which contains 150,000 images in real urban and highway environments.

[0098] S6-2: Input the training set into the lane line detection network and train the network. The resolution of the input image is adjusted to 320×800, and data augmentation is performed by horizontal flipping, randomly adjusting the image brightness, and blurring. The AdamW algorithm is used to train the network with a batch size of 32 and an initial learning rate of 6e-4.

[0099] S6-3: The lane line detection network outputs the detection results of the current image.

[0100] S6-4: Calculate the loss between the detection results and the true lane line coordinate sequences. Use KorniaFocalLoss, SmoothL1Loss, LineIOU, and NLLLoss as loss functions, A p and G cls_t represent the predicted value and the true value of the foreground-background classification result respectively, B p and G param_t represent the predicted value and the true value of the anchor parameter value respectively, C seg and G seg_t represent the predicted mask map and the true mask map of per-pixel segmentation respectively, D p and G coord_t represent the predicted value and the true value of the coordinate sequence. Then the expression of the final loss function is as follows:

[0101] Loss = L KornialFocalLoss (A p , G cls_t ) + L SmoothF1Loss (B p , G param_t ) + L NLLLoss (C seg , G seg_t )

[0102] + L LineIou (D p , G coord_t ).

[0103] S6-5: When the loss value reaches the minimum, the model converges, stops training, saves the parameters, and obtains the trained lane line detection network.

[0104] S7: Input the image to be detected into the trained lane line detection model, so as to output the predicted coordinate sequence of the image to be detected and generate the final lane line prediction map.

[0105] To verify the effectiveness of the above examples, the method of the present invention is compared with other advanced methods in terms of performance on two datasets, CULane and Curvelanes. Eleven metrics are selected on the CULane dataset: F1 50 , F1 75 , Normal, Crowd, Dazzle, Shadow, Noline, Arrow, Curve, Cross, Night. Among these eleven metrics, except for Cross, F1 50 , F1 75, Normal, Crowd, Dazzle, Shadow, Noline, Arrow, Curve, Night. The larger the value, the better the performance. Three metrics were selected on the Curvelanes dataset: F1, Precision, Recall. The experimental results are shown in Table 1.

[0106] Table 1 Comparison results of detection accuracy on the CULane dataset

[0107]

[0108] As can be seen from Table 1, this embodiment leads the existing methods in multiple metrics on the CULane dataset, proving the effectiveness of the method of this embodiment.

[0109] Table 2 Comparison results of detection accuracy on the Curvelanes dataset

[0110]

[0111] As can be seen from Table 2, the F1 score of this embodiment on the Curvelanes dataset has improved compared with the existing methods, proving the effectiveness of the method of this embodiment.

[0112] Figure 7 This is the result comparison diagram of the method of the present invention. The first column is the original image, the second column is the Mask image, which is a binary mask image of the lane ground truth. The third column is the ground truth image, which is the visualization display of the true label on the original image. The fourth column is the method result output image, which is the prediction map displayed after the prediction sequence output by the network is post-processed. It can be seen from the comparison that the solution provided by this example can accurately locate the lane line object, precisely locate the lane line target, and dynamically adapt to the changes in different scenarios.

[0113] The above has described the embodiments of the present invention in detail with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations made to these embodiments still fall within the protection scope of the present invention.

Claims

1. A precise lane line detection method in complex environments based on multi-scale collaborative enhancement and semantic compensation, characterized in that It includes the following steps: S1: Obtain a lane line detection dataset and input it into the backbone network to obtain the feature maps of the last three levels of the network, denoted as Fi, where i represents the level of the feature. ; S2: Using the multi-scale collaborative enhancement module, perform spatial information enhancement and semantic information interaction aggregation on feature maps at different levels through a top-down spatial compensation path and a bottom-up semantic filling path, so as to improve the feature information of the feature maps extracted by the feature network. This module includes a soft upsampling combination sub-module SUpCom constructed by GSConv convolutional layers and a GSConv convolutional combination sub-module GSConvCom for feature extraction; The top-down spatial compensation path takes the feature map and adjusts its size to that of the feature map through the soft upsampling combination sub-module SUpCom, and then performs channel concatenation with the feature map to model multi-scale information and obtain the feature map , which is then adjusted to the size of the feature map to obtain the feature map to supplement the semantic information lost during propagation. The feature map extracts the aggregated multi-scale feature information through the GSConv convolution combination sub-module GSConvCom and adjusts it to the size of the feature map through the soft upsampling combination sub-module SUpCom, and then performs channel concatenation with the feature map and the feature map to generate the feature map , thereby enhancing the detailed information of the low-level features while retaining the high-level semantic information. The bottom-up semantic filling path captures more context and semantic information from the feature map through the GSConv convolution combination sub-module GSConvCom to obtain the feature map as the feature output of the third level and passes it upward; the feature map is obtained by convolving the feature map generated by GSConv with the feature map generated by the GSConv convolution combination sub-module GSConvCom and channel concatenation is performed to generate the feature map . The feature map is obtained by passing the feature map through the GSConv convolution combination sub-module GSConvCom as the feature output of the fourth level and passed upward. The feature map , the feature map , and the feature map capture rich spatial information through GSConv convolution and perform channel concatenation with the feature map to generate the feature map . The feature map is obtained by passing the feature map through the GSConv convolution combination sub-module GSConvCom as the feature output of the last level, realizing the gradual transmission of semantic information to the bottom layer; S3: Using the hybrid scene semantic compensation module, adaptively route and select the propagation path, extract compensation semantic information through the progressive compensation extraction layer, and fuse it with the improved feature maps to enhance the feature robustness in different scenes and improve the scene adaptation ability of the feature maps. This module includes an adaptive routing sub-module and a progressive compensation extraction layer sub-module; S4: Input the features with improved scene robustness effect into the boundary correction guidance module, further focus on the important information in the features in terms of spatial and channel dimensions, refine the lane line feature information, and improve the lane line boundary detection effect. This module includes an improved spatial attention sub-module and a channel attention sub-module; S5: Pass the refined feature maps of different levels through the detection head Ha, where a represents the level of the feature. , and pass the prefabricated prior lane anchor Pc, where c represents the number of optimization times. , to the first detection head H1. S6: Gradually adjust the prior lane anchor Pc parameters downward level by level. Finally, after being adjusted by the last detection head, the prior lane anchor outputs the optimized anchor P3 as the predicted output result of the lane line.

2. The accurate lane line detection method in complex environment based on multi-scale collaborative enhancement and semantic compensation according to claim 1, characterized in that, Use the multi-scale collaborative enhancement module to perform spatial information enhancement and semantic information interaction aggregation on the extracted feature maps at different levels; the multi-scale collaborative enhancement module includes a soft upsampling combination sub-module SUpCom constructed based on GSConv convolutional layers, a GSConv convolutional combination sub-module GSConvCom, and a top-down spatial compensation path and a bottom-up semantic filling path built, so as to obtain feature maps containing richer semantic and detailed feature information; The GSConv convolutional layer captures image features through a standard convolutional layer Conv, divides the feature maps in terms of channels to obtain feature maps X1 and X2, then performs convolution on feature map X1 and models the channel relationship of X2 respectively, and then splices and rearranges the channels of the feature maps; the soft upsampling combination sub-module SUpCom includes C3k2, C2fGS2, and SNI sub-modules, where C3k2 is a convolutional layer that replaces the Bottleneck module in C2f with a C3k convolutional layer, C2fGS2 is a convolutional layer that replaces the Bottleneck module in C2f with a GSConv convolutional layer, and SNI is the weighting of the upsampling layer and the influence factor, which can effectively retain gradient information and interact with feature information at different levels to avoid feature loss during the upsampling process; The GSConv convolutional combination sub-module GSConvCom includes a C3k2 convolutional layer and a C2fGS2 convolutional layer sub-module.

3. The precise lane line detection method in a complex environment based on multi-scale collaborative enhancement and semantic compensation according to claim 1, characterized in that, The described hybrid scenario semantic compensation module includes an adaptive routing sub-module and a progressive compensation extraction layer sub-module, which outputs appropriate path weights for samples through adaptive routing and weights them to the feature maps extracted by the progressive compensation extraction layer sub-module, so as to obtain rich scene feature information to improve the scene robustness effect.

4. The precise lane line detection method in a complex environment based on multi-scale collaborative enhancement and semantic compensation according to claim 1, characterized in that, The described boundary correction guidance module effectively optimizes the feature detail information by improving the spatial attention mechanism and introducing the channel attention mechanism.

5. The accurate lane line detection method in complex environment based on multi-scale collaborative enhancement and semantic compensation according to claim 3, characterized in that, The described hybrid scenario semantic compensation module is used to improve the context and semantic information in different scenarios and make up for the feature information lost in the feature extraction process; Adaptive routing, the input feature map Extract features through the convolutional layer and perform channel modeling on the features by combining channel attention CA to generate multi-scale features , to enhance the global correlation between features; Subsequently, through global average pooling (GAP), linear mapping (Linear), and uniform noise (Noise) processing, a probabilistic output for feature selection is generated. , where n represents the number of blocks. , to determine the optimal semantic compensation path; the semantic compensation path is a progressive compensation extraction layer, including two shared blocks , where m represents the number of layers. and multiple independent blocks , where K represents the number of blocks and m represents the number of layers. , , to construct a two-layer progressive enhancement mechanism; the first layer contains one shared block , and several independent blocks , where 1 represents the first layer and K represents the number of blocks. , containing K groups of residual blocks composed of CBR and C3k2. The second layer contains one shared block and several independent blocks , where K represents the number of blocks. , containing a convolutional block composed of a group of CBR and C3k2; passing the input feature map through the first layer to obtain feature maps of different depths , where the subscript bi represents the i-th block and the superscript 1 represents the first layer, and splicing the feature maps of different depths in channels, and then adding the feature information extracted by the shared block to obtain the feature information , to enhance the basic features of different paths and compensate for the information loss of other paths with different depths. The feature information extracted from the first layer is passed to the second layer to capture multi-scale feature information, which is represented by and weighted and fused with the path probability generated by the adaptive routing, so as to obtain the supplementary semantic information Z with the highest sample correlation.

6. The accurate lane line detection method in complex environment based on multi-scale collaborative enhancement and semantic compensation according to claim 4, characterized in that The boundary correction guidance module is used to refine the lane line feature information, improve the lane line boundary detection effect; improve the spatial attention mechanism, and input the feature map is divided into r groups through channels, and the feature maps of each group are denoted as , where r represents the number of groups. Through a axis average pooling branch and a axis average pooling branch, and then the outputs of the two branches are denoted as and , where x represents the X-axis direction or the horizontal direction, and y represents the Y-axis direction or the vertical direction. Channel splicing is performed to generate a feature map with horizontal and vertical dimensions, and important spatial features are extracted through a multi-core one-dimensional convolution module. The extracted features are split and restored to the shapes of the horizontal and vertical dimension feature maps, and after channel splicing the feature maps of each group, spatial weights in the horizontal and vertical dimensions are obtained through the Sigmoid layer. Then, point multiplication with the input feature map is performed to obtain the spatially enhanced feature maps X2 in the horizontal and vertical directions; the input feature map X is passed through a 3×3 convolution layer to generate the feature map S, and the product is taken with the feature map X2 after group normalization GroupNorm and global average pooling GAP to obtain the spatial weight W1; the product is taken with the feature map X2 after group normalization GroupNorm and the feature map S after global average pooling GAP to obtain the spatial weight W2; the spatial weight W1 and the spatial weight W2 are added and passed through Sigmoid to obtain the superimposed spatial weight, and then weighted to the feature map X2 to generate the spatially enhanced feature map T, thereby enhancing the correlation between different spatial positions through the cross-spatial interaction enhancement network; the channel attention mechanism is introduced, and the input feature X passes through the phantom convolution layer GhostConv to obtain the response in the channel dimension, and the channel relationship is modeled through the partial self-attention mechanism PSA to generate the attention weight in the channel dimension and weighted to the feature map channel by channel to generate the feature map C after channel attention. Then, the feature maps after spatial and channel dimension attention are added to generate the final feature map Q.

7. The accurate lane line detection method in a complex environment based on multi-scale collaborative enhancement and semantic compensation according to claim 1, characterized in that, The training steps of the trained lane line detection network include: Construct a lane line detection network; Construct a training set, which is a video frame sequence and its lane line ground truth coordinate sequence; Input the training set into the lane line detection network for training; The lane line detection network outputs a lane coordinate prediction sequence; Calculate the difference between the prediction sequence and the ground truth coordinate sequence and perform backpropagation; When the loss value reaches the minimum, the model converges, stops training, and obtains the trained lane line detection network.

Citation Information

Patent Citations

  • Lane line detection system based on geometric attention perception

    CN111582201A

  • Image semantic segmentation method and device, computer readable storage medium and chip

    CN112529904A