Complex environment lane line accurate detection method based on multi-scale cooperative enhancement and semantic compensation

By introducing multi-scale collaborative enhancement and semantic compensation modules into the lane line detection method, the problem of poor robustness of existing methods in complex environments is solved, and higher detection accuracy and adaptability are achieved, reducing calculation costs.

CN120032334AActive Publication Date: 2025-05-23SHIJIAZHUANG TIEDAO UNIV

Patent Information

Application Number
CN202510190179.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing lane line detection methods are poorly robust in complex environments, insufficient generalization capabilities, high computing costs, and neglect the integration of global context information.

Method used

A lane line detection method based on multi-scale collaborative enhancement and semantic compensation is proposed. Through multi-scale collaborative enhancement module, hybrid scene semantic compensation module and boundary correction guidance module, multi-scale integration and semantic compensation of feature information are improved, and the robustness and accuracy of detection are improved.

Benefits of technology

It significantly improves the accuracy and robustness of lane line detection, enhances the adaptability of the detection model, reduces the calculation cost, and can work stably in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032334A_ABST
    Figure CN120032334A_ABST
Patent Text Reader

Abstract

The invention discloses a complex environment lane line accurate detection method based on multi-scale cooperative enhancement and semantic compensation. The method comprises the following steps: obtaining a lane line data set, and inputting the lane line data set into a trained lane line accurate detection network; adopting a backbone network to obtain multi-level features; utilizing a multi-scale collaborative enhancement module to obtain features of spatial and semantic information interaction enhancement; a mixed scene semantic compensation module is utilized to obtain supplementary semantic features so as to enhance feature robustness in different scenes; a boundary correction guiding module is adopted, features are refined from space and channel dimensions, and the lane boundary detection effect is improved; and optimizing and adjusting the lane line anchor layer by layer through the enhanced features to obtain a final lane line coordinate prediction sequence. According to the method, multi-scale cooperative enhancement and semantic compensation are combined, the lane line detection precision is improved, and the detection robustness in a complex scene is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for accurately detecting lane lines in a complex environment based on multi-scale collaborative enhancement and semantic compensation, and belongs to the technical field of computer vision. Background Art

[0002] The rapid development of autonomous driving technology and intelligent transportation systems is profoundly changing modern transportation. In the successful application of these systems, the ability to accurately detect and interpret lane markings on the road is an important factor. The lane detection system enables the vehicle to stay on the correct driving trajectory by accurately identifying lane boundaries and road centerlines, and provides support for functions such as lane keeping assist, lane departure warning, and automatic lane cruise control. It is an important research direction in the field of computer vision.

[0003] Lane detection is an important part of the autonomous driving system. Its core task is to accurately extract lane marking information from road images or sensor data to ensure that the vehicle drives along the correct trajectory. Traditional methods mainly rely on edge detection, color segmentation, and geometric fitting techniques, but they are easily disturbed in complex environments such as illumination changes, road wear and occlusion, resulting in poor robustness. In recent years, the successful application of convolutional neural networks in the field of computer vision has promoted the development of lane detection technology. Compared with traditional methods based on manual features, neural networks can automatically learn and extract deep features of lane markings, improving the accuracy and robustness of detection. By constructing a deep network structure, the neural network can extract different levels of information from the input image, such as edge, texture, and shape features, so as to more effectively identify lane lines, and maintain good detection performance even in complex environments such as illumination changes, occlusion, or blurred lane markings.

[0004] Lane detection methods based on deep learning have made some progress in accuracy and adaptability, but the following problems still exist. First, the diversity of road scenes brings difficulties to detection, including changes in lane shape and marking style, and occlusion or interference of lane markings by environmental factors. Since model training is usually based on limited road data sets, their generalization ability in unknown environments is poor, resulting in an inability to stably cope with various actual road conditions. Secondly, existing methods ignore the integration of global contextual information, shallow features lack high-level semantic information, deep features lack detailed information, and the effective integration of multi-scale features still needs to be explored. Finally, the computational cost of existing lane detection models is high, and autonomous driving systems need to process a large amount of sensor data, while computing resources are limited. Summary of the invention

[0005] The purpose of the present invention is to solve the above problems in the existing methods and propose a method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation.

[0006] To achieve the above object, the technical solution of the present invention is:

[0007] A method for accurately detecting lane lines in complex environments based on multi-scale collaborative enhancement and semantic compensation, characterized by comprising the following steps:

[0008] S1: Obtain the lane detection dataset and input it into the backbone network to obtain the feature maps of the last three layers of the network, denoted as Fi, where i represents the level of the feature, 3≤i≤5;

[0009] S2: Using a multi-scale collaborative enhancement module, the feature maps of different levels are enhanced in spatial information and interactively aggregated in semantic information through a top-down spatial compensation path and a bottom-up semantic filling path to improve the feature information of the feature maps extracted by the feature network. This module includes a soft upsampling combination submodule SUPCom constructed by a GSConv convolutional layer and a GSConv convolution combination submodule GSConvCom for feature extraction.

[0010] S3: Using the hybrid scene semantic compensation module, the propagation path is selected through adaptive routing and the compensation semantic information is extracted through the progressive compensation extraction layer, and then fused with the improved feature map to enhance the feature robustness in different scenarios and improve the scene adaptability of the feature map. This module includes an adaptive routing submodule and a progressive compensation extraction layer submodule;

[0011] S4: The features with improved scene robustness are passed to the boundary correction guidance module, which further focuses on the important information in the features through the spatial and channel dimensions, refines the lane feature information, and improves the lane boundary detection effect. This module includes an improved spatial attention submodule and a channel attention submodule;

[0012] S5: The refined feature maps of different levels are passed through the detection head Ha, where a represents the level of the feature, 1≤a≤3, and the prefabricated prior lane anchor Pc, where c represents the number of optimizations, 0≤c≤3, is passed to the first detection head H1;

[0013] S6: By adjusting the parameters of the prior lane anchor Pc step by step, the prior lane anchor is finally adjusted by the last detection head and outputs the optimized anchor P3 as the predicted output result of the lane line.

[0014] A further technical solution is that the multi-scale collaborative enhancement module performs spatial information enhancement and interactive aggregation of semantic information on feature maps extracted at different levels; the multi-scale collaborative enhancement module includes a soft upsampling combination submodule SUPCom constructed based on the GSConv convolution layer, a GSConv convolution combination submodule GSConvCom and a top-down spatial compensation path and a bottom-up semantic filling path, thereby obtaining a feature map containing richer semantics and detailed feature information.

[0015] Furthermore, the GSConv convolution layer captures image features through the standard convolution layer Conv, and divides the feature map into channels to obtain feature maps X1 and X2, then convolves the feature map X1 and X2 for channel relationship modeling, and then performs channel splicing and channel rearrangement Shuffle on the feature map, thereby effectively improving the expressiveness of the feature map, enhancing the feature interaction between channels and reducing the computational overhead; the soft upsampling combination module SUPCom includes C3k2, C2fGS2 and SNI sub-modules, wherein C3k2 is the convolution layer of the Bottleneck module in C2f replaced by the C3k convolution layer, and C2fGS2 is the GSConv convolution layer replacing the Bottl in C2f The convolution layer of the eneck module, SNI is the weighted sum of the upsampling layer and the influencing factor, which can effectively retain the gradient information and interact with the feature information of different levels to avoid feature loss in the upsampling process; the GSConv convolution combination submodule GSConvCom includes the C3k2 convolution layer and the C2fGS2 convolution layer submodule, which can enrich the gradient information and extract more advanced semantic information; Split represents channel division, Conv represents the standard convolution layer, DCN represents the deformable convolution, DSConv represents the depth separable convolution, SENet represents the compression and excitation network, [·] represents the channel splicing, Shuffle represents the channel rearrangement operation, X and R represent the input features, and the specific calculation formula is as follows:

[0016] X1,X2=Split(Conv(X)),

[0017] GSConv=Shuffle([DSConv(DCN(X1)),SENet(X2)]),

[0018] GSConvCom=C2fGS2(C3k2(C2fGS2(R))),

[0019] SUpCom=SNI(C2fGS2(C3k2(R))).

[0020] Furthermore, the top-down spatial compensation path converts the feature map F 5Adjust the size to the feature map F through the Soft Upsampling Combination Module SUpCom 4 and concatenate channels with the feature map F 4 to model multi-scale information and obtain a feature map Adjust to the feature map F 3 size to obtain a feature map to supplement the semantic information lost during propagation. For the feature map Extract the aggregated multi-scale feature information through the GSConv Convolution Combination Module GSConvCom and adjust to the feature map F through the Soft Upsampling Combination Module SUpCom 3 size and concatenate channels with the feature map F 3 and the feature map to generate a feature map Thus, while retaining high-level semantic information, the detailed information of low-level features is enhanced; Cat represents the channel concatenation operation, and its specific calculation formula is as follows:

[0021]

[0022] Furthermore, the bottom-up semantic filling path processes the feature map to capture more context and semantic information through the GSConv Convolution Combination Module GSConvCom, thereby obtaining a feature map as the feature output of the third level and passing it upward; Process the feature map the feature map generated by GSConv convolution and the feature map generated by the GSConv Convolution Combination Module GSConvCom to concatenate channels and generate a feature map Process the feature map the feature map obtained through the GSConv Convolution Combination Module GSConvCom as the feature output of the fourth level and pass it upward. Process the feature map feature map feature map to capture rich spatial information through GSConv convolution and concatenate channels with the feature map F 5 to generate a feature map Process the feature map the feature map generated by the GSConv Convolution Combination Module GSConvCom as the feature output of the last level, realizing the gradual transmission of semantic information to the bottom layer; Cat represents the channel concatenation operation, and its specific calculation formula is as follows:

[0023]

[0024] A further technical solution is that the hybrid scene semantic compensation module includes an adaptive routing submodule and a progressive compensation extraction layer submodule, and the appropriate path weights of the samples output by the adaptive routing are weighted to the feature map extracted by the progressive compensation extraction layer submodule, thereby obtaining rich scene feature information to improve the scene robustness effect.

[0025] Furthermore, the mixed scene semantic compensation module is used to improve the context and semantic information in different scenes to compensate for the feature information lost during the feature extraction process.

[0026] Furthermore, adaptive routing extracts features from the input feature map X through a convolutional layer and combines channel attention CA to model the feature into a channel to generate a multi-scale feature X′ to enhance the global correlation between features; then, through global average pooling GAP, linear mapping Linear and uniform noise processing, the probability output Prob of feature selection is generated. n , where n represents the number of blocks, 0≤n≤K, to determine the optimal semantic compensation path; the specific calculation formula is as follows:

[0027] X′=Conv3(X)+CA(Conv3(X)),

[0028] Prob n =Linear(GAP(X′))+Noise,n∈0,1,2…,K.

[0029] Furthermore, the semantic compensation path is a progressive compensation extraction layer, including two shared blocks S m , where m represents the number of layers, 1≤m≤2 and multiple independent blocks Where K represents the number of blocks, m represents the number of layers, 1≤K≤3, 1≤m≤2, to construct a two-layer progressive enhancement mechanism; the first layer contains a shared block S 1 , and several independent blocks Where 1 represents the first layer, K represents the number of blocks, 1≤K≤3, contains K groups of residual blocks composed of CBR and C3k2, and the second layer contains a shared block S 2 and several independent blocks Where K represents the number of blocks, 1≤K≤3, including a set of convolution blocks composed of CBR and C3k2; the input feature map is passed through the first layer to obtain feature maps of different depths. The subscript bi represents the i-th block, the superscript 1 represents the first layer, and the feature maps of different depths are channel-joined and then combined with the feature information extracted by the shared block. Feature information Out2 is obtained by adding features to enhance the basic features of different paths and compensate for the information loss of other different depth paths; the feature information extracted by the first layer is passed to the second layer to capture multi-scale feature information. It is represented by and weightedly fused with the path probability generated by adaptive routing to obtain the supplementary semantic information Z with the highest sample relevance; Sum represents the summation operation, [·] represents channel concatenation, and Convk represents a convolutional layer with a convolution kernel of k×k. The specific calculation formula is as follows:

[0030]

[0031] A further technical solution is that the boundary correction guidance module effectively optimizes feature detail information by improving the spatial attention mechanism and introducing the channel attention mechanism.

[0032] Furthermore, the boundary correction guidance module is used to refine lane line feature information and improve lane line boundary detection effect.

[0033] Furthermore, the spatial attention mechanism is improved, and the input feature map X is divided into r groups by channels, and the feature map of each group is recorded as X r , where r represents the number of groups, through an X-axis average pooling branch and a Y-axis average pooling branch, and then the output of the two branches is recorded as X x With X y , where x represents the X-axis direction or horizontal direction, and y represents the Y-axis direction or vertical direction. Channel splicing is performed to generate feature maps with horizontal and vertical dimensions, and important spatial features are extracted through a multi-core one-dimensional convolution module. The proposed features are split and restored to the shape of horizontal and vertical feature maps, and the feature maps of each group are spliced ​​through the channel. After the sigmoid layer is passed through the sigmoid layer, the spatial weights of the horizontal and vertical dimensions are obtained, and then the feature map X2 is obtained after spatial enhancement in the horizontal and vertical directions by point multiplication with the input feature map; the input feature map X is passed through a 3×3 convolution layer to generate a feature map S and is normalized with the group N orm and the feature map X2 after global average pooling GAP are multiplied to obtain the spatial weight W1; the feature map X2 after group normalization GroupNorm is multiplied with the feature map S after global average pooling GAP to obtain the spatial weight W2; the spatial weight W1 is added to the spatial weight W2 and the superimposed spatial weight is obtained through Sigmoid, and then weighted to the feature map X2 to generate the spatially enhanced feature map T, thereby enhancing the correlation between different spatial positions of the network through cross-space interaction; σ is the sigmoid function operation, [·] represents the channel splicing operation, ⊙ represents the Hadamard product operation, CS r Indicates that the channel is divided into r groups of operations, GAP x Represents the average pooling operation in the horizontal direction, GAP Y Represents the average pooling operation in the vertical direction, GAP represents the global average pooling, ConvGroup [3,5,7,9]Represents multiple one-dimensional convolution layer operations with kernel sizes of 3, 5, 7, and 9, Cat group Indicates channel splicing operation by group, CBR 3×3 Reshape represents the combined operation of 3×3 convolution, batch normalization, and rectified linear unit. Reshape represents the shape reshaping operation. The specific calculation formula is as follows:

[0034] X r =CS r (X),

[0035] X x =GAP x (X r ),

[0036] X y =GAP Y (X r ),

[0037] X2=σ(Cat group (Reshape(ConvGroup [3,5,7,9] ([X x ,X y ]))))⊙X,

[0038] S=CBR 3×3 (X),

[0039] W1=S×GAP(GroupNorm(X2)),

[0040] W2=GAP(S)×GroupNorm(X2),

[0041] T=X2×σ(W1+W2).

[0042] Furthermore, the channel attention mechanism is introduced, the input feature X is passed through the ghost convolution layer GhostConv to obtain the response of the channel dimension, and the channel relationship is modeled through the partial self-attention mechanism PSA to generate the attention weight of the channel dimension and weight it channel by channel to the feature map to generate the feature map C after the channel attention, and then the feature maps paid attention to by the space and channel dimensions are added to generate the final feature map Z; the specific calculation formula is as follows:

[0043] C=σ(GAP(PSA(GhostConv(GAP(X)))))⊙T,

[0044] Z=C+T.

[0045] A further technical solution is that the training steps of the trained lane detection network include:

[0046] Build a lane detection network;

[0047] Constructing a training set, wherein the training set is a video frame sequence and a lane line true value coordinate sequence;

[0048] Input the training set into the lane detection network for training;

[0049] The lane detection network outputs a sequence of lane coordinate predictions;

[0050] Calculate the difference between the predicted sequence and the true value coordinate sequence and back propagate;

[0051] When the loss value reaches the minimum, the model converges, the training stops, and the trained lane line detection network is obtained.

[0052] The beneficial effects of adopting the above technical solution are: the present invention provides a multi-scale collaborative enhancement module, which fully combines the semantic information and detail information of the object, which is helpful for positioning and detection; the present invention designs a hybrid scene semantic compensation module to compensate for the spatial and semantic information lost in the feature network extraction process, and improves the detection robustness problem in different scenes; the present invention develops a boundary correction guidance module, which improves the effect of capturing feature detail information and improves the detection effect of lane line boundaries. The three modules adopted are integrated in the network, which greatly improves the accuracy of lane line detection and reflects the advantages of the proposed technical solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0054] Figure 1 The network overall flow chart of the embodiment of the present invention;

[0055] Figure 2 This is a diagram of the overall network architecture of an embodiment of the present invention;

[0056] Figure 3 This is a submodule structure diagram of a multi-scale collaborative enhancement module in an embodiment of the present invention;

[0057] Figure 4 This is a structural diagram of a multi-scale collaborative enhancement module in an embodiment of the present invention;

[0058] Figure 5 This is a structural diagram of a hybrid scene semantic compensation module in an embodiment of the present invention;

[0059] Figure 6 A boundary correction guiding module in an embodiment of the present invention;

[0060] Figure 7 This is a result diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] The present invention provides a method for accurately detecting lane lines in complex environments based on multi-scale collaborative enhancement and semantic compensation. Figure 1 As shown, the following steps are included:

[0063] S1: Construct a backbone network to obtain multi-level features; use DLA34 as the backbone network to obtain multi-level features from the input image, represented by F i , where 3≤i≤5, i represents the level of the feature.

[0064] S2: Construct a multi-scale collaborative enhancement module, whose structure is shown in Figure 3 and Figure 4 ;

[0065] S2-1: The multi-scale collaborative enhancement module enhances the extracted feature maps at different levels and captures the features of different receptive domains. It consists of a soft upsampling combination submodule SUPCom built based on the GSConv convolution layer, a GSConv convolution combination submodule GSConvCom, and a top-down spatial compensation path and a bottom-up semantic filling path.

[0066] S2-2: GSConv convolution layer, its structure can be found in Figure 3 a, capture image features through the standard convolution layer Conv, and divide the feature map into channels to obtain feature maps X1 and X2, then perform convolution on feature map X1 and channel relationship modeling on X2 respectively, and then perform channel splicing and channel rearrangement Shuffle on the feature map, so as to effectively improve the expression ability of the feature map, enhance the feature interaction between channels and reduce the computational overhead; Split represents channel division, Conv represents standard convolution layer, DCN represents deformable convolution, DSConv represents depth separable convolution, SENet represents compression and excitation network, [·] represents channel splicing, Shuffle represents channel rearrangement operation, X represents input feature, and its specific calculation formula is as follows:

[0067] X1,X2=Split(Conv(X)),

[0068] GSConv=Shuffle([DSConv(DCN(X1)),SENet(X2)]).

[0069] S2-3: The soft upsampling combination module SUPCom includes C3k2, C2fGS2 and SNI submodules. Its structure is shown in Figure 3 b, where C3k2 is the convolutional layer of the C3k convolutional layer replacing the convolutional layer of the Bottleneck module in C2f, C2fGS2 is the convolutional layer of the GSConv convolutional layer replacing the convolutional layer of the Bottleneck module in C2f, SNI is the weighted upsampling layer and the influencing factor, which can effectively retain the gradient information and interact with the feature information of different levels to avoid feature loss during the upsampling process; the GSConv convolution combination submodule GSConvCom, its structure see Figure 3 c, including C3k2 convolutional layer and C2fGS2 convolutional layer submodules, can enrich gradient information and extract more advanced semantic information; R represents the input feature, and its specific calculation formula is as follows:

[0070] GSConvCom=C2fGS2(C3k2(C2fGS2(R))),

[0071] SUpCom=SNI(C2fGS2(C3k2(R))).

[0072] S2-4: Top-down spatial compensation path, its structure is shown in Figure 4 a, the feature map F 5 The size is adjusted to the feature map F through the soft upsampling combination module SupCom 4 size and feature map F 4 Perform channel splicing to model multi-scale information and obtain feature maps Adjust to feature map F 3 Size to obtain feature map To supplement the semantic information lost in the propagation, the feature map The GSConv convolution combination module GSConvCom extracts the aggregated multi-scale feature information and adjusts it to the feature map F through the soft upsampling combination module SupCom 3 size and feature map F 3 and feature map Perform channel concatenation to generate feature maps Thus, while retaining high-level semantic information, the detail information of low-level features is enhanced; the specific calculation formula is as follows:

[0073]

[0074] S2-5: Bottom-up semantic filling path, its structure is shown in Figure 4 b, the feature map The GSConv convolution combination module GSConvCom captures more context and semantic information to obtain the feature map As the feature output of the third level and passed upward; the feature map The feature map generated by GSConv convolution and the feature map generated by GSConv convolution combination module GSConvCom Perform channel splicing and generate feature maps The feature map Feature map obtained by GSConv convolution combination module GSConvCom As the fourth-level feature output and passed upward, the feature map Feature Map Feature Map Rich spatial information is captured by GSConv convolution and combined with the feature map F 5 Perform channel concatenation to generate feature maps The feature map Feature map generated by GSConv convolution combination module GSConvCom As the feature output of the last layer, it realizes the step-by-step transmission of semantic information to the bottom layer; Cat represents the channel splicing operation, and its specific calculation formula is as follows:

[0075]

[0076] S3: Construct a hybrid scene semantic compensation module, the structure of which can be found in Figure 5 ;

[0077] S3-1: Adaptive routing, extracting features from the input feature map X through the convolutional layer and combining the channel attention CA to model the feature into the channel to generate multi-scale features X′, so as to enhance the global correlation between features; then, through global average pooling GAP, linear mapping Linear and uniform noise processing, the probability output Prob of feature selection is generated n , where n represents the number of blocks, 0≤n≤K, to determine the optimal semantic compensation path; Convk represents the convolution layer with a convolution kernel of k×k, and its specific calculation formula is as follows:

[0078] X′=Conv3(X)+CA(Conv3(X)),

[0079] Prob n =Linear(GAP(X′))+Noise,n∈0,1,2…,K.

[0080] S3-2: progressive compensation extraction layer, including two shared blocks S m , where m represents the number of layers, 1≤m≤2 and multiple independent blocks Where K represents the number of blocks, m represents the number of layers, 1≤K≤3, 1≤m≤2, to construct a two-layer progressive enhancement mechanism; the first layer contains a shared block S 1 , and several independent blocks Where 1 represents the first layer, K represents the number of blocks, 1≤K≤3, contains K groups of residual blocks composed of CBR and C3k2, and the second layer contains a shared block S 2 and several independent blocks Where K represents the number of blocks, 1≤K≤3, including a set of convolutional blocks composed of CBR and C3k2; the input feature map is passed through the first layer to obtain feature maps of different depths. The subscript bi represents the i-th block, the superscript 1 represents the first layer, and the feature maps of different depths are channel-joined and then combined with the feature information extracted by the shared block. Feature information Out2 is obtained by adding features to enhance the basic features of different paths and compensate for the information loss of other different depth paths; the feature information extracted by the first layer is passed to the second layer to capture multi-scale feature information. It is represented by and weightedly fused with the path probability generated by adaptive routing to obtain the supplementary semantic information Z with the highest sample relevance; Sum represents the summation operation, [·] represents channel concatenation, and Convk represents a convolutional layer with a convolution kernel of k×k. The specific calculation formula is as follows:

[0081]

[0082] S4: Construct the boundary correction guidance module, whose structure is shown in Figure 6 ;

[0083] S4-1: Improve the spatial attention mechanism, divide the input feature map X into r groups through channels, and record the feature map of each group as X r , where r represents the number of groups, through an X-axis average pooling branch and a Y-axis average pooling branch, and then the output of the two branches is recorded as X x With X y, where x represents the X-axis direction or horizontal direction, and y represents the Y-axis direction or vertical direction. Channel splicing is performed to generate feature maps with horizontal and vertical dimensions, and important spatial features are extracted through a multi-core one-dimensional convolution module. The proposed features are split and restored to the shape of horizontal and vertical feature maps, and the feature maps of each group are spliced ​​through the channel. After the sigmoid layer is passed through the sigmoid layer, the spatial weights of the horizontal and vertical dimensions are obtained, and then the feature map X2 is obtained after spatial enhancement in the horizontal and vertical directions by point multiplication with the input feature map; the input feature map X is passed through a 3×3 convolution layer to generate a feature map S and is normalized with the group N orm and the feature map X2 after global average pooling GAP are multiplied to obtain the spatial weight W1; the feature map X2 after group normalization GroupNorm is multiplied with the feature map S after global average pooling GAP to obtain the spatial weight W2; the spatial weight W1 is added to the spatial weight W2 and the superimposed spatial weight is obtained through Sigmoid, and then weighted to the feature map X2 to generate the spatially enhanced feature map T, thereby enhancing the correlation between different spatial positions of the network through cross-space interaction; σ is the sigmoid function operation, [·] represents the channel splicing operation, ⊙ represents the Hadamard product operation, CS r Indicates that the channel is divided into r groups of operations, GAP x Represents the average pooling operation in the horizontal direction, GAP Y Represents the average pooling operation in the vertical direction, GAP represents the global average pooling, and GonvGroup [3,5,7,9] Represents multiple one-dimensional convolution layer operations with kernel sizes of 3, 5, 7, and 9, Cat group Indicates channel splicing operation by group, CBR 3×3 Reshape represents the combined operation of 3×3 convolution, batch normalization, and rectified linear unit. Reshape represents the shape reshaping operation. The specific calculation formula is as follows:

[0084] X r =CS r (X),

[0085] X x =GAP x (X r ),

[0086] X y =GAP Y (X r ),

[0087] X2=σ(Cat group (Reshape(ConvGroup [3,5,7,9] ([X x ,X y ]))))⊙X,

[0088] S=CBR 3×3 (X),

[0089] W1=S×GAP(GroupNorm(X2)),

[0090] W2=GAP(S)×GroupNorm(X2),

[0091] T=X2×σ(W1+W2).

[0092] S4-2: Introduce the channel attention mechanism, pass the input feature X through the ghost convolution layer GhostConv to obtain the response of the channel dimension, and use the partial self-attention mechanism PSA to model the channel relationship to generate the attention weight of the channel dimension and weight it channel by channel to the feature map to generate the feature map C after the channel attention, and then add the feature maps paid attention to by the space and channel dimensions to generate the final feature map Z; the specific calculation formula is as follows:

[0093] C=σ(GAP(PSA(GhostConv(GAP(X))))))⊙T,

[0094] Z=C+T.

[0095] S5: Build a detection head, including a self-attention mechanism and a fully connected layer. The lane anchor focuses on the lane features through the self-attention mechanism and maps them to the coordinate sequence parameters through the fully connected layer.

[0096] S6: Build a lane detection network. Its structure is shown in Figure 2 , conduct training;

[0097] S6-1: Construct a training set, which is a sequence of video frames and their true lane coordinate sequences. Two widely used datasets are used for training: CULane and CurveLanes. CULane is a commonly used public dataset for lane detection, containing 88,880 training set images and test sets of 8 challenging scenes. CurveLanes is a more challenging public dataset, containing 150,000 images in real urban and highway environments.

[0098] S6-2: Input the training set into the lane detection network to train the network. The resolution of the input image is adjusted to 320×800 and data enhancement is performed by horizontal flipping, random adjustment of image brightness, and blurring. The AdamW algorithm is used to train the network with a batch size of 32 and an initial learning rate of 6e-4.

[0099] S6-3: The lane detection network outputs the detection result of the current image.

[0100] S6-4: Calculate the loss of the detection results and the true lane line coordinate sequence. Use KorniaFocalLoss, SmoothL1Loss, LineIOU and NLLLoss as loss functions. p and G cls_t Represent the predicted value and true value of the foreground and background classification results, respectively, p and G param_t Represent the predicted value and true value of the anchor parameter value, C seg and G seg_t Denote the predicted mask map and the real mask map of pixel-by-pixel segmentation, respectively, and D p and G coord_t Represents the predicted value and true value of the coordinate sequence, then the expression of the final loss function is as follows:

[0101] Loss = L KornialFocalLoss (A p ,G cls_t )+L SmoothF1Loss (B p ,G param_t )+L NLLLoss (C seg ,G seg_t )

[0102] +L LineIou (D p ,G coord_t )).

[0103] S6-5: When the loss value reaches the minimum, the model converges, the training stops, the parameters are saved, and the trained lane line detection network is obtained.

[0104] S7: Input the image to be detected into the trained lane line detection model, thereby outputting the predicted coordinate sequence of the image to be detected and generating the final lane line prediction map.

[0105] In order to verify the effectiveness of the above examples, the performance of the proposed method is compared with other advanced methods on two datasets, CULane and Curvelanes. Eleven indicators are selected on the CULane dataset: F1 50 、F1 75 , Normal, Crowd, Dazzle, Shadow, Noline, Arrow, Curve, Cross, Night. Among these eleven indicators, except Cross, F1 50 、F1 75,Normal,Crowd,Dazzle,Shadow,Noline,Arrow,Curve,Night,The larger the value, the better the performance.,Three indicators were selected on the Curvelanes dataset: F1, Precision, and Recall, and the experimental results are shown in Table 1.

[0106] Table 1 Comparison of detection accuracy on the CULane dataset

[0107]

[0108] As shown in Table 1, this embodiment is ahead of the existing methods in multiple indicators on the CULane dataset, which proves the effectiveness of the method of this embodiment.

[0109] Table 2 Comparison of detection accuracy on the Curvelanes dataset

[0110]

[0111] As shown in Table 2, the F1 score of this embodiment on the Curvelanes dataset is improved compared with the existing method, which proves the effectiveness of the method of this embodiment.

[0112] Figure 7 This is a comparison chart of the results of the method of the present invention. The first column is the original image, the second column is the Mask image, which is the binary mask image of the lane truth value, the third column is the truth image, which is the visualization of the real label on the original image, and the fourth column is the output image of the method result, which is the prediction image displayed after the prediction sequence output by the network is post-processed. Through comparison, it can be seen that the solution provided in this example can accurately locate the lane line object, finely locate the lane line target, and dynamically adapt to changes in different scenarios.

[0113] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.

Claims

1. A method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation, characterized in that The steps include: S1: Obtain the lane detection dataset and input it into the backbone network to obtain the feature maps of the last three layers of the network, denoted as Fi, where i represents the level of the feature, 3≤i≤5; S2: Using a multi-scale collaborative enhancement module, the feature maps of different levels are enhanced in spatial information and interactively aggregated in semantic information through a top-down spatial compensation path and a bottom-up semantic filling path to improve the feature information of the feature maps extracted by the feature network. This module includes a soft upsampling combination submodule SUPCom constructed by a GSConv convolutional layer and a GSConv convolution combination submodule GSConvCom for feature extraction. S3: Using the hybrid scene semantic compensation module, the propagation path is selected through adaptive routing and the compensation semantic information is extracted through the progressive compensation extraction layer, and then fused with the improved feature map to enhance the feature robustness in different scenarios and improve the scene adaptability of the feature map. This module includes an adaptive routing submodule and a progressive compensation extraction layer submodule; S4: The features with improved scene robustness are passed to the boundary correction guidance module, which further focuses on the important information in the features through the spatial and channel dimensions, refines the lane feature information, and improves the lane boundary detection effect. This module includes an improved spatial attention submodule and a channel attention submodule; S5: The refined feature maps of different levels are passed through the detection head Ha, where a represents the level of the feature, 1≤a≤3, and the prefabricated prior lane anchor Pc, where c represents the number of optimizations, 0≤c≤3, is passed to the first detection head H1; S6: By adjusting the parameters of the prior lane anchor Pc step by step, the prior lane anchor is finally adjusted by the last detection head and outputs the optimized anchor P3 as the predicted output result of the lane line.

2. The method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation as claimed in claim 1, characterized in that: The multi-scale collaborative enhancement module is used to perform spatial information enhancement and semantic information interactive aggregation on the feature maps of different levels extracted; the multi-scale collaborative enhancement module includes a soft upsampling combination submodule SUPCom constructed based on the GSConv convolution layer, a GSConv convolution combination submodule GSConvCom, and a top-down spatial compensation path and a bottom-up semantic filling path, so as to obtain a feature map containing richer semantics and detailed feature information; The GSConv convolution layer captures image features through the standard convolution layer Conv, and divides the feature map into channels to obtain feature maps X1 and X2, then convolves the feature map X1 and X2 for channel relationship modeling, and then performs channel splicing and channel rearrangement Shuffle on the feature map, thereby effectively improving the expressiveness of the feature map, enhancing the feature interaction between channels and reducing the computational overhead; the soft upsampling combination module SUPCom includes C3k2, C2fGS2 and SNI submodules, wherein C3k2 is the convolution layer of the Bottleneck module in C2f replaced by the C3k convolution layer, and C2fGS2 is the GSConv convolution layer replaced by the Bottlene in C2f The convolution layer of the ck module, SNI is the weighted sum of the upsampling layer and the influencing factor, which can effectively retain the gradient information and interact with the feature information of different levels to avoid feature loss in the upsampling process; the GSConv convolution combination submodule GSConvCom includes the C3k2 convolution layer and the C2fGS2 convolution layer submodules, which can enrich the gradient information and extract more advanced semantic information; Split represents channel division, Conv represents the standard convolution layer, DCN represents the deformable convolution, DSConv represents the depth separable convolution, SENet represents the compression and excitation network, [·] represents the channel splicing, Shuffle represents the channel rearrangement operation, X and R represent the input features, and the specific calculation formula is as follows: X1,X2=Split(Conv(X)), GSConv=Shuffle([DSConv(DCN(X1)),SENet(X2)]), GSConvCom=C2fGS2(C3k2(C2fGS2(R))), SUpCom=SNI(C2fGS2(C3k2(R))).

3. The method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation as claimed in claim 1, characterized in that: The hybrid scene semantic compensation module includes an adaptive routing submodule and a progressive compensation extraction layer submodule. The appropriate path weight of the sample output by the adaptive routing is weighted to the feature map extracted by the progressive compensation extraction layer submodule, thereby obtaining rich scene feature information to improve the scene robustness effect.

4. The method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation as claimed in claim 1, characterized in that: The boundary correction guidance module effectively optimizes feature detail information by improving the spatial attention mechanism and introducing the channel attention mechanism.

5. The method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation as claimed in claim 2, characterized in that: The top-down spatial compensation path adjusts the size of the feature map F5 to the size of the feature map F4 through the soft upsampling combination module SupCom and performs channel splicing with the feature map F4 to model the multi-scale information to obtain the feature map Adjust to the feature map F3 size to obtain the feature map To supplement the semantic information lost in the propagation, the feature map The GSConv convolution combination module GSConvCom extracts the aggregated multi-scale feature information and adjusts it to the size of feature map F3 through the soft upsampling combination module SupCom and combines it with feature map F3 and feature map Perform channel concatenation to generate feature maps Thus, while retaining high-level semantic information, the detailed information of low-level features is enhanced; the bottom-up semantic filling path transforms the feature map The GSConv convolution combination module GSConvCom captures more context and semantic information to obtain the feature map As the feature output of the third level and passed upward; the feature map The feature map generated by GSConv convolution and the feature map generated by GSConv convolution combination module GSConvCom Perform channel splicing and generate feature maps The feature map Feature map obtained by GSConv convolution combination module GSConvCom As the fourth-level feature output and passed upward, the feature map Feature Map Feature Map Rich spatial information is captured through GSConv convolution and channel concatenated with feature map F5 to generate feature map The feature map Feature map generated by GSConv convolution combination module GSConvCom As the feature output of the last layer, it realizes the step-by-step transmission of semantic information to the bottom layer; Cat represents the channel splicing operation, and its specific calculation formula is as follows:

6. The method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation as claimed in claim 3, characterized in that: The mixed scene semantic compensation module is used to improve the context and semantic information in different scenes to compensate for the feature information lost during the feature extraction process; Adaptive routing: extract features from the input feature map X through the convolutional layer and combine channel attention CA to model the feature channel to generate multi-scale features X′ to enhance the global correlation between features; Then, the probability output Prob of feature selection is generated through global average pooling GAP, linear mapping Linear and uniform noise Noise processing n , where n represents the number of blocks, 0≤n≤K, to determine the optimal semantic compensation path; the semantic compensation path is a progressive compensation extraction layer, including two shared blocks S m , where m represents the number of layers, 1≤m≤2 and multiple independent blocks Where K represents the number of blocks, m represents the number of layers, 1≤K≤3, 1≤m≤2, to construct a two-layer progressive enhancement mechanism; the first layer contains a shared block S 1 , and several independent blocks Where 1 represents the first layer, K represents the number of blocks, 1≤K≤3, contains K groups of residual blocks composed of CBR and C3k2, and the second layer contains a shared block S 2 and several independent blocks Where K represents the number of blocks, 1≤K≤3, including a set of convolutional blocks composed of CBR and C3k2; the input feature map is passed through the first layer to obtain feature maps of different depths. The subscript bi represents the i-th block, the superscript 1 represents the first layer, and the feature maps of different depths are channel-joined and then combined with the feature information extracted by the shared block. Feature information Out2 is obtained by adding features to enhance the basic features of different paths and compensate for the information loss of other different depth paths; the feature information extracted by the first layer is passed to the second layer to capture multi-scale feature information. It is represented by and weightedly fused with the path probability generated by adaptive routing to obtain the supplementary semantic information Z with the highest sample relevance; Sum represents the summation operation, [·] represents channel concatenation, and Convk represents a convolutional layer with a convolution kernel of k×k. The specific calculation formula is as follows: X ′ =Conv3(X)+CA(Conv3(X)), Prob n =Linear(GAP(X ′ ))+Noise,n∈0,1,2…,K, 7. The method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation as claimed in claim 4, characterized in that: The boundary correction guidance module is used to refine the lane feature information and improve the lane boundary detection effect; improve the spatial attention mechanism, divide the input feature map X into r groups through channels, and record the feature map of each group as X r , where r represents the number of groups, through an X-axis average pooling branch and a Y-axis average pooling branch, and then the output of the two branches is recorded as X x With X y , where x represents the X-axis direction or horizontal direction, y represents the Y-axis direction or vertical direction, and channels are spliced ​​to generate feature maps with horizontal and vertical dimensions, and important spatial features are extracted through a multi-core one-dimensional convolution module. The proposed features are split and restored to the shape of horizontal and vertical feature maps, and the feature maps of each group are spliced ​​through the channel. After the sigmoid layer is passed, the spatial weights of the horizontal and vertical dimensions are obtained, and then the feature map X2 is obtained after the point multiplication with the input feature map in the horizontal and vertical directions. The input feature map X is passed through a 3×3 convolution layer to generate a feature map S and the feature map X2 after group normalization GroupNorm and global average pooling GAP is multiplied to obtain the spatial weight W1; the feature map X2 after group normalization GroupNorm and the feature map X2 after global average pooling GAP are multiplied to obtain the spatial weight W1. The spatial weight W2 is obtained by multiplying the feature map S of GAP; the spatial weight W1 is added to the spatial weight W2 and the superimposed spatial weight is obtained through Sigmoid, which is then weighted to the feature map X2 to generate the spatially enhanced feature map T, thereby enhancing the correlation between different spatial positions of the network through cross-spatial interaction; the channel attention mechanism is introduced, the input feature X is passed through the ghost convolution layer GhostConv to obtain the response of the channel dimension and the channel relationship is modeled through the partial self-attention mechanism PSA to generate the attention weight of the channel dimension and weighted to the feature map channel by channel to generate the feature map C after channel attention, and then the feature maps paid attention to by the spatial and channel dimensions are added to generate the final feature map Z; σ is the sigmoid function operation, [·] represents the channel concatenation operation, ⊙ represents the Hadamard product operation, CS r Indicates that the channel is divided into r groups of operations, GAP x Represents the average pooling operation in the horizontal direction, GAP Y Represents the average pooling operation in the vertical direction, GAP represents the global average pooling, and GonvGroup [3,5,7,9] Represents multiple one-dimensional convolution layer operations with kernel sizes of 3, 5, 7, and 9, Cat group Indicates channel splicing operation by group, CBR 3×3 Reshape represents the combined operation of 3×3 convolution, batch normalization, and rectified linear unit. Reshape represents the shape reshaping operation. The specific calculation formula is as follows: X r =CS r (X), X x =GAP x (X r ), X y =GAP Y (X r ), X2=σ(Cat group (Reshape(ConvGroup [3,5,7,9] ([X x ,X y ]))))⊙X, S=CBR 3×3 (X), W1=S×GAP(GroupNorm(X2)), W2=GAP(S)×GroupNorm(X2), T=X2×σ(W1+W2), C=σ(GAP(PSA(GhostConv(GAP(X))))))⊙T, Z=C+T.

8. The method for accurate lane line detection in complex environments based on multi-scale collaborative enhancement and semantic compensation as claimed in claim 1, characterized in that: The training steps of the trained lane detection network include: Build a lane detection network; Constructing a training set, wherein the training set is a video frame sequence and a lane line true value coordinate sequence; Input the training set into the lane detection network for training; The lane detection network outputs a sequence of lane coordinate predictions; Calculate the difference between the predicted sequence and the true value coordinate sequence and back propagate; When the loss value reaches the minimum, the model converges, the training stops, and the trained lane line detection network is obtained.

Citation Information

Patent Citations

  • Lane line detection system based on geometric attention perception

    CN111582201A

  • Automobile drivable area planning method based on multi-task neural network

    CN112418236A

  • Image semantic segmentation method and device, computer readable storage medium and chip

    CN112529904A

  • RGBD semantic segmentation method based on semantic flow network

    CN114596322A

  • Complex scene-oriented asymmetric double-branch real-time semantic segmentation network method

    CN115082928A

Cited By

  • Lightweight lane line detection method in automatic driving scene

    CN121121682A

  • Lightweight lane line detection method in autonomous driving scenario

    CN121121682B

  • Unmanned aerial vehicle aerial image target detection method based on high-frequency detail feature compensation

    CN121999395A