3D lane line generation method and system based on frequency perception feature fusion

By adopting the frequency-aware feature fusion method in the 3D lane line generation technology and combining multi-scale and frequency pyramid networks, the shortcomings of 2D lane line generation technology in complex environments are solved, and more accurate and robust 3D lane line generation is achieved, which is suitable for autonomous driving systems.

CN119992502BActive Publication Date: 2025-06-06YANTAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510465508.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-06
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In a complex and changeable road environment, traditional 2D lane line generation technology is difficult to effectively deal with factors such as lighting changes, road occlusion and bad weather, and cannot capture the three-dimensional information of the lane, affecting the stable driving of the vehicle.

Method used

The 3D lane line generation method based on frequency-aware feature fusion is adopted. Through the combination of multi-scale feature extraction network, frequency pyramid network, spatial transformation fusion network and lane line extraction network, lane line extraction network is extracted and fused, and lane features of different scales and frequencies are generated to generate more accurate and robust 3D lane lines.

Benefits of technology

It improves the accuracy and robustness of 3D lane line generation, can better cope with changes in complex environments, maintain good real-time, and is suitable for autonomous driving and assisted driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992502B_ABST
    Figure CN119992502B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of 3D image generation, and specifically to a 3D lane line generation method and system based on frequency perception feature fusion. The method improves the accuracy and robustness of 3D lane line generation by designing a multi-scale feature extraction network for extracting features of different depth scales of an image, a frequency pyramid network for extracting high-frequency information from high-resolution features and fusing them with low-resolution features to enhance the expression capability of the low-resolution features, a spatial transformation fusion network for extracting more robust and advanced BEV features, and a 3D lane line generation model of a lane line extraction network for generating 3D lane lines. The method can better cope with changes in complex environments while maintaining good real-time performance, is suitable for autonomous driving and assisted driving systems, and shows broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 3D image generation, and in particular to a 3D lane line generation method and system based on frequency perception feature fusion. Background Art

[0002] In practical applications, lane line generation faces many complex environmental factors, such as lighting changes, road occlusion and bad weather. These factors often pose severe challenges to lane generation technology. With the advancement of autonomous driving technology, only 2D lane line generation can no longer meet the needs of practical applications. In order to enable vehicles to drive stably in complex and changing road environments, lane generation has gradually expanded from traditional 2D generation to 3D tasks. This transformation makes lane generation not only limited to lane line markings on the plane, but also able to take into account the three-dimensional information of the lane, including position, width, curvature and height.

[0003] The BEV perspective can provide road information from above, so that road lane lines can be more intuitively represented as geometric figures. However, the geometry of lane lines may change significantly with the road type, the curvature of the road section, and the motion state of the vehicle. Smaller-scale lane lines require more refined feature descriptions, while larger-scale lane lines require the network to capture a more macroscopic road structure. In order to effectively cope with the scale changes of lane lines on different road sections, the lane generation model needs to pay attention to both the details (such as the endpoints and bifurcations of the lane lines) and the global parts (such as the overall shape of the lane, curvature changes, etc.). The raw data captured by sensors such as cameras usually have differences in perspective and coordinate system. How to convert it into a unified bird's-eye view and effectively extract spatial features is a difficult point. Especially in complex urban environments, the resolution, perspective differences and noise problems of sensors affect the accuracy of feature descriptions, thereby reducing the accuracy of 3D lane line generation. Summary of the invention

[0004] The purpose of the present invention is to provide a 3D lane line generation method and system based on frequency perception feature fusion.

[0005] The technical solution of the present invention is as follows:

[0006] A 3D lane line generation method based on frequency perception feature fusion includes the following operations:

[0007] S1. Acquire a number of road images and corresponding lane line marking position data to form a lane image training set; use the lane image training set to train a 3D lane line generation model including a multi-scale feature extraction network, a frequency pyramid network, a spatial transformation fusion network, and a lane line extraction network to obtain a trained 3D lane line generation model;

[0008] In the frequency pyramid network, several lane feature maps of different scales are respectively processed by feature dimension alignment to obtain several image dimension perception feature maps of different scales; the low-scale image dimension perception feature map is processed by frequency perception feature fusion with the medium-scale image dimension perception feature map, and then the high-scale image dimension perception feature map is processed by frequency perception feature fusion, and the fusion results are extracted to obtain several frequency perception cross-fusion feature maps, which are used to perform operations in the spatial transformation fusion network;

[0009] The operation of frequency-aware feature fusion processing is as follows: the high-frequency features of the first-scale image dimension perception feature map and the second-scale image dimension perception feature map are cross-processed with low-frequency features to obtain a low-frequency cross-feature map; the first-scale image dimension perception feature map and the cross-low-frequency feature map are fused with low-frequency features to obtain a low-frequency fusion feature map; the convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map are cross-processed with high-frequency features to obtain a first frequency-domain perception fusion feature map; the first-scale image dimension perception feature map and the low-frequency cross-feature map are processed with a low-pass filter, and then added to the first frequency-domain perception fusion feature map by element-by-element addition to obtain a shallow-medium-scale frequency perception cross-fusion feature map;

[0010] S2. The lane image to be processed is processed by the trained 3D lane line generation model to obtain 3D lane lines.

[0011] The processing operations in the multi-scale feature extraction network are as follows: the lane image is processed by the residual network to obtain a shallow-scale lane feature map; the shallow-scale lane feature map is processed by the residual network to obtain a middle-scale lane feature map; the middle-scale lane feature map is processed by convolution to obtain a deep-scale lane feature map.

[0012] The operation of the spatial transformation fusion network is as follows: several frequency-aware cross-fusion feature maps are respectively processed by view conversion and attention enhancement, and then spliced ​​to obtain a 3D fusion feature map; the view conversion processing can be achieved through several full connections, ReLU activation functions, and several convolutions, normalization processing, and ReLU activation functions; the operation of attention enhancement processing is as follows: the view conversion map is split into two feature maps to obtain a first split map and a second split map; the first split map is subjected to group normalization and self-attention processing to obtain an attention feature map; the attention feature map and the second split map are convolved and normalized after element-by-element addition to obtain a feature map to be spliced.

[0013] The method for acquiring the high-frequency features of the second-scale image dimension perceptual feature map is specifically as follows: the second-scale image dimension perceptual feature map is subjected to 1×1 convolution processing to obtain a second-scale initial convolution map; the second-scale initial convolution map is subjected to 3×3 convolution processing to obtain a second-scale reconvolution map; the second-scale reconvolution map is subjected to high-pass filtering processing, and then subjected to high-pass filtering and convolution processing with the second-scale initial convolution map to obtain the high-frequency features of the second-scale image dimension perceptual feature map as the second-scale high-frequency feature map.

[0014] The operation of the high-pass filter is specifically as follows: the second-scale high-pass filter image and the second-scale initial convolution image are processed by a low-pass filter to obtain a low-pass filter cross feature map; the difference map between the second-scale initial convolution image and the low-pass filter cross feature map is added element by element to the second-scale initial convolution image, and then a convolution operation is performed; the second-scale high-pass filter image is obtained by high-pass filtering the second-scale reconvolution image.

[0015] The operation of cross-processing of high-frequency features is specifically as follows: the convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map are element-by-element added and high-pass filtered, and then processed with the second-scale image dimension perception feature map by a high-pass filter to obtain the first frequency domain perception fusion feature map.

[0016] The specific operation of low-frequency feature cross processing is as follows: the high-frequency features of the second-scale image dimension perception feature map are processed by low-pass filtering, and then the convolution features of the first-scale image dimension perception feature map are processed by low-pass filtering, and then the high-frequency features of the second-scale image dimension perception feature map are added element by element to obtain a low-frequency cross feature map.

[0017] A 3D lane line generation system based on frequency perception feature fusion, used to implement the above-mentioned 3D lane line generation method based on frequency perception feature fusion, comprising:

[0018] The training 3D lane line generation model generation module is used to obtain several road images and corresponding lane line mark position data to form a lane image training set; the lane image training set is used to train a 3D lane line generation model including a multi-scale feature extraction network, a frequency pyramid network, a spatial transformation fusion network and a lane line extraction network to obtain a trained 3D lane line generation model; in the frequency pyramid network, several lane feature maps of different scales are respectively processed by feature dimension alignment to obtain several image dimension perception feature maps of different scales; the low-scale image dimension perception feature map is subjected to frequency perception feature fusion processing with the medium-scale image dimension perception feature map, and then the frequency perception feature fusion processing is performed with the high-scale image dimension perception feature map, and the fusion result is extracted to obtain several The frequency-aware cross-fusion feature map is used to perform operations in the spatial transformation fusion network; the frequency-aware feature fusion processing operation is specifically as follows: the high-frequency features of the first-scale image dimension perception feature map and the second-scale image dimension perception feature map are cross-processed with low-frequency features to obtain a low-frequency cross-feature map; the first-scale image dimension perception feature map and the cross-low-frequency feature map are fused with low-frequency features to obtain a low-frequency fusion feature map; the convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map are cross-processed with high-frequency features to obtain a first frequency domain perception fusion feature map; the first-scale image dimension perception feature map and the low-frequency cross-feature map are processed by a low-pass filter, and then the first frequency domain perception fusion feature map is processed by element-by-element addition to obtain a shallow medium-scale frequency perception cross-fusion feature map;

[0019] The 3D lane line generation module is used to process the lane image to be processed through the trained 3D lane line generation model to obtain the 3D lane line.

[0020] A 3D lane line generation device based on frequency perception feature fusion includes a processor and a memory, wherein the processor implements the above-mentioned 3D lane line generation method based on frequency perception feature fusion when executing a computer program stored in the memory.

[0021] A computer-readable storage medium is used to store a computer program, wherein when the computer program is executed by a processor, the above-mentioned 3D lane line generation method based on frequency-aware feature fusion is implemented.

[0022] The beneficial effects of the present invention are:

[0023] The present invention provides a 3D lane line generation method based on frequency-aware feature fusion. The method improves the accuracy and robustness of 3D lane line generation by designing a multi-scale feature extraction network for extracting features of different depth scales of an image, a frequency pyramid network for extracting high-frequency information from high-resolution features and fusing them with low-resolution features to enhance the expressiveness of low-resolution features, a spatial transformation fusion network for extracting more robust and advanced BEV features, and a 3D lane line generation model of a lane line extraction network for generating 3D lane lines. The method can better cope with changes in complex environments while maintaining good real-time performance. The method is suitable for autonomous driving and assisted driving systems, and shows broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By reading the detailed description of the preferred embodiment below, the scheme and advantages of the present application will become clear to those skilled in the art. The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention.

[0025] In the attached picture:

[0026] Figure 1 Schematic diagram of the effect of the method of this embodiment in road condition 1 in the embodiment; Figure 1 In the figure, (a) is the front view result, (b) is the 3D fusion feature map result, and (c) is the 3D lane line three-dimensional space result; Figure 1 The green line in the middle is the 3D lane line generated by the method of this embodiment, and the blue line is the actual lane line;

[0027] Figure 2 Schematic diagram of the effect of the method of this embodiment in road condition 2 in the embodiment; Figure 2 In the figure, (a) is the front view result, (b) is the 3D fusion feature map result, and (c) is the 3D lane line three-dimensional space result; Figure 2 The green line in the middle is the 3D lane line generated by the method of this embodiment, and the blue line is the actual lane line;

[0028] Figure 3 Schematic diagram of the effect of the method of this embodiment in road condition three in the embodiment; Figure 3 In the figure, (a) is the front view result, (b) is the 3D fusion feature map result, and (c) is the 3D lane line three-dimensional space result; Figure 3 The green line in the middle is the 3D lane line generated by the method of this embodiment, and the blue line is the actual lane line;

[0029] Figure 4 Schematic diagram of the effect of the method of this embodiment in road condition 4 in the embodiment; Figure 4 In the figure, (a) is the front view result, (b) is the 3D fusion feature map result, and (c) is the 3D lane line three-dimensional space result; Figure 4The green line in the middle is the 3D lane line generated by the method of this embodiment, and the blue line is the actual lane line;

[0030] Figure 5 Schematic diagram of the effect of the method of this embodiment in road condition 5 in the embodiment; Figure 5 In the figure, (a) is the front view result, (b) is the 3D fusion feature map result, and (c) is the 3D lane line three-dimensional space result; Figure 5 The green line in the middle is the 3D lane line generated by the method of this embodiment, and the blue line is the actual lane line. DETAILED DESCRIPTION

[0031] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.

[0032] This embodiment provides a 3D lane line generation method based on frequency perception feature fusion, including the following operations:

[0033] S1. Acquire a number of road images and corresponding lane line marking position data to form a lane image training set; use the lane image training set to train a 3D lane line generation model including a multi-scale feature extraction network, a frequency pyramid network, a spatial transformation fusion network, and a lane line extraction network to obtain a trained 3D lane line generation model;

[0034] In the frequency pyramid network, several lane feature maps of different scales are respectively processed by feature dimension alignment to obtain several image dimension perception feature maps of different scales; the low-scale image dimension perception feature map is processed by frequency perception feature fusion with the medium-scale image dimension perception feature map, and then the high-scale image dimension perception feature map is processed by frequency perception feature fusion, and the fusion results are extracted to obtain several frequency perception cross-fusion feature maps, which are used to perform operations in the spatial transformation fusion network;

[0035] The operation of frequency-aware feature fusion processing is as follows: the high-frequency features of the first-scale image dimension perception feature map and the second-scale image dimension perception feature map are cross-processed with low-frequency features to obtain a low-frequency cross-feature map; the first-scale image dimension perception feature map and the cross-low-frequency feature map are fused with low-frequency features to obtain a low-frequency fusion feature map; the convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map are cross-processed with high-frequency features to obtain a first frequency-domain perception fusion feature map; the first-scale image dimension perception feature map and the low-frequency cross-feature map are processed with a low-pass filter, and then added to the first frequency-domain perception fusion feature map by element-by-element addition to obtain a shallow-medium-scale frequency perception cross-fusion feature map;

[0036] S2. The lane image to be processed is processed by the trained 3D lane line generation model to obtain 3D lane lines.

[0037] S1. Acquire several road images and corresponding lane line marking position data to form a lane image training set; use the lane image training set to train a 3D lane line generation model including a multi-scale feature extraction network, a frequency pyramid network, a spatial transformation fusion network and a lane line extraction network to obtain a trained 3D lane line generation model.

[0038] First, obtain several road images and corresponding lane marking position data, mark the lane marking position in the road image, and perform data enhancement on each road image, including but not limited to scaling and color transformation, to ensure that the lane marking is consistent with the image, and unify all road images to the same size and color space to form a lane image training set.

[0039] Then a 3D lane line generation model is constructed, which includes an input network for image input, a multi-scale feature extraction network for extracting features of different depth scales of the image, a frequency pyramid network for extracting high-frequency information from high-resolution features and fusing them with low-resolution features to enhance the expressiveness of low-resolution features, a spatial transformation fusion network for extracting more robust and advanced BEV features, a lane line extraction network for generating 3D lane lines, and an output network for output.

[0040] The operation steps in the representative network of the 3D lane line generation model are as follows.

[0041] The processing operations in the multi-scale feature extraction network are as follows: the lane image is processed by the residual network to obtain a shallow-scale lane feature map; the shallow-scale lane feature map is then processed by the residual network to obtain a medium-scale lane feature map; the medium-scale lane feature map is processed by convolution to obtain a deep-scale lane feature map. The residual network is used in the multi-scale feature extraction network to extract multi-scale features and generate lane feature maps with 1 / 16 and 1 / 32 resolutions, thereby capturing rich spatial information at different scales and enhancing the perception ability of the model. In order to further extract deeper semantic information, a convolution module is added after the residual network to enhance the expressiveness of the multi-scale feature extraction network, so that the multi-scale feature extraction network can effectively extract feature maps with a resolution of 1 / 64, thereby further improving the generation accuracy and efficiency of the model.

[0042] In the frequency pyramid network, several lane feature maps of different scales, namely, shallow-scale lane feature maps, middle-scale lane feature maps, and deep-scale lane feature maps, are respectively processed by feature dimension alignment to enhance the expression ability of lane feature maps of different scales in the horizontal and vertical directions, and realize the precise dimension alignment of features of different spatial scales, so as to obtain several image dimension perception feature maps of different scales, namely, shallow-scale image dimension perception feature maps (first-scale image dimension perception feature maps), middle-level image dimension perception feature maps (second-scale image dimension perception feature maps), and deep-level image dimension perception feature maps (third-scale image dimension perception feature maps); the low-scale image dimension perception feature map is subjected to frequency perception feature fusion processing with the middle-scale image dimension perception feature map, and then the high-scale image dimension perception feature map is subjected to frequency perception feature fusion processing to extract the fusion result, that is, the image dimension perception feature maps of adjacent scales are subjected to frequency perception feature fusion processing in order from low to high scales, and then the image dimension perception feature maps of the next scale are subjected to frequency perception feature fusion processing. The feature map is subjected to frequency-aware feature fusion processing; specifically, the first-scale image dimension-aware feature map and the second-scale image dimension-aware feature map are first subjected to frequency-aware feature fusion processing, and after obtaining the fusion result, the fusion result is subjected to frequency-aware feature fusion processing with the third-scale image dimension-aware feature map, high-frequency information is extracted from the high-resolution features, and is fused with the low-resolution features to enhance the expression ability of the low-resolution features. At the same time, under the guidance of the low resolution, the high-resolution features can highlight the key features at the high resolution, thereby further strengthening their important information. Finally, the enhanced two are fused, effectively combining the semantic information of the low-resolution features and the detail information of the high-resolution features, to obtain several frequency-aware cross-fusion feature maps (including shallow-medium-scale frequency-aware cross-fusion feature maps and medium-deep-scale frequency-aware cross-fusion feature maps), which are used to perform operations in the spatial transformation fusion network with the shallow-scale image dimension-aware feature map (the first-scale image dimension-aware feature map), so as to facilitate more accurate extraction of lane features.

[0043] The above-mentioned feature dimension alignment processing operation is specifically as follows: taking the shallow-scale lane feature map as an example, the shallow-scale lane feature map is subjected to average pooling and maximum pooling, respectively, to obtain a shallow-scale average pooling feature map and a shallow-scale maximum pooling feature map; the shallow-scale average pooling feature map and the shallow-scale maximum pooling feature map are respectively processed by a multi-layer perceptron MLP (including but not limited to being implemented by 1×1 convolution, ReLU activation function, and 1×1 convolution), firstly compressing the number of channels of the pooling feature map to 1 / r (Reduction, reduction rate) times of the original number of channels, and then expanding it to the original number of channels, and then performing element-by-element addition and Sigmoid activation function processing to obtain a shallow-scale nonlinear feature map; the shallow-scale nonlinear feature map and the shallow-scale lane feature map are processed by element-by-element addition, 1×1 convolution, and ReLU activation function to obtain a shallow-scale image dimension perception feature map, that is, a first-scale image dimension perception feature map is obtained.

[0044] The above operations of obtaining the second-scale image dimension perceptual feature map and the third-scale image dimension perceptual feature map are the same as the method for obtaining the first-scale image dimension perceptual feature map, and will not be repeated here to save space.

[0045] The operation of frequency-perceptual feature fusion processing is specifically as follows: taking the first-scale image dimension perceptual feature map and the second-scale image dimension perceptual feature map as examples, the high-frequency features of the first-scale image dimension perceptual feature map and the second-scale image dimension perceptual feature map are cross-processed with low-frequency features to obtain a low-frequency cross-feature map; the first-scale image dimension perceptual feature map and the cross-low-frequency feature map are fused with low-frequency features to obtain a low-frequency fusion feature map; the convolution features of the low-frequency fusion feature map and the second-scale image dimension perceptual feature map are cross-processed with high-frequency features to obtain a first frequency domain perceptual fusion feature map; the first-scale image dimension perceptual feature map and the low-frequency cross-feature map are processed by a low-pass filter, and then element-by-element addition is performed with the first frequency domain perceptual fusion feature map to obtain a shallow medium-scale frequency perceptual cross-fusion feature map.

[0046] The shallow-medium scale frequency-aware cross-fusion feature map and the third-scale image dimension-aware feature map are subjected to frequency-aware feature fusion processing to obtain the medium-deep scale frequency-aware cross-fusion feature map. The operation is the same as above and will not be repeated here to save space.

[0047] Among them, the high-frequency features of the second-scale image dimension perceptual feature map, that is, the high-frequency features of the middle-level image dimension perceptual feature map, are obtained as follows: the second-scale image dimension perceptual feature map is processed by 1×1 convolution to obtain a second-scale initial convolution map; the second-scale initial convolution map is processed by 3×3 convolution to obtain a second-scale reconvolution map; the second-scale reconvolution map is processed by high-pass filtering, and then the second-scale initial convolution map is processed by high-pass filtering and 3×3 convolution to obtain the high-frequency features of the second-scale image dimension perceptual feature map as the second-scale high-frequency feature map.

[0048] The specific operation of high-pass filtering is as follows: after the second-scale reconvolution image is processed by mapping adjustment and Softmax function processing, the convolution kernel weight of the first neighborhood area of ​​each pixel (the 3×3 convolution kernel weight in the four neighborhood areas) is normalized to enhance the high-frequency details in the image. After resizing, the second-scale high-pass filter image is obtained.

[0049] The operation of the high-pass filter is as follows: the second-scale high-pass filter image and the second-scale initial convolution image are processed by a low-pass filter to obtain a low-pass filter cross feature map; the difference map of the second-scale initial convolution image and the low-pass filter cross feature map is added element by element to the second-scale initial convolution image to perform a convolution operation. The core of the high-pass filter is to apply a low-pass filter to extract low-frequency components, and then subtract these components from the original signal to separate the high-frequency components; then the original signal is combined with the high-frequency enhanced signal to obtain the high-pass filter processing result. The second-scale high-pass filter image is obtained by high-pass filtering the second-scale reconvolution image.

[0050] The specific operation of low-frequency feature cross processing is as follows: the high-frequency features of the second-scale image dimension perceptual feature map are processed by low-pass filtering, and then the convolution features of the first-scale image dimension perceptual feature map (the first-scale image dimension perceptual feature map is obtained after 1×1 convolution and 3×3 convolution) are processed by low-pass filtering, and then the high-frequency features of the second-scale image dimension perceptual feature map are added element by element to obtain a low-frequency cross feature map.

[0051] The specific operation of low-pass filtering is as follows: after the high-frequency features of the second-scale image dimension perception feature map are adjusted by mapping and resized by the Softmax function, the convolution kernel weights of the second neighborhood area of ​​each pixel (5×5 convolution kernel weights in 12 neighborhood areas) are normalized to remove noise and retain the low-frequency information of the image. After resizing, the second-scale low-pass filter map is obtained.

[0052] The operation of low-pass filter processing is as follows: the convolution features of the first-scale image dimension perception feature map are processed by pixel padding, feature expansion, size adjustment, nearest neighbor interpolation processing (to achieve upsampling) and size adjustment, and then the convolution kernel aggregation processing and size adjustment are performed with the second-scale low-pass filter map to obtain the initial low-frequency cross feature map. The second-scale low-pass filter map is obtained by low-pass filtering the high-frequency features of the second-scale image dimension perception feature map. The low-pass filter allows the low-frequency components of the feature map to pass through while attenuating the high-frequency components. The mask helps the filter generate and reduce the high-frequency components that may be inconsistent during the feature fusion process. As needed, including but not limited to using interpolation to upsample the expanded feature map to expand its size, then adjusting the feature map and mask to an appropriate shape, performing element-wise multiplication, generating a weighted feature map, and then adding the results of all convolution kernels to generate a processed feature map.

[0053] The specific operation of low-frequency feature fusion processing is as follows: the cross low-frequency feature map is processed by low-pass filtering to obtain a cross low-frequency low-pass filtered feature map; the first-scale image dimension perception feature map is processed by 3×3 convolution, and then processed with the cross low-frequency low-pass filtered feature map by a low-pass filter to obtain a low-frequency fusion feature map.

[0054] The specific operation of high-frequency feature cross-processing is as follows: the convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map (the second-scale image dimension perception feature map is obtained after 1×1 convolution and 3×3 convolution) are added element by element and high-pass filtered, and then processed with the second-scale image dimension perception feature map through a high-pass filter to obtain the first frequency domain perception fusion feature map.

[0055] The operation of the spatial transformation fusion network is as follows: several frequency-aware cross-fusion feature maps, including the first-scale image dimension-aware feature map, the shallow-medium-scale frequency-aware cross-fusion feature map, and the deep-scale frequency-aware cross-fusion feature map, are respectively processed by view conversion (generating BEV feature maps) and attention enhancement, and then spliced ​​to obtain a 3D fusion feature map. Through the spatial transformation mechanism, the multi-level features of the network are no longer simply processed in parallel, but can dynamically adjust the spatial alignment of features of different scales, thereby enhancing the fusion of multi-level features and the efficiency of information transmission. Through precise spatial registration, the problem of information loss caused by perspective differences or complex scenes can be effectively solved, and the attention module is combined to extract more robust and high-level features, further improving the overall performance of the network.

[0056] The view conversion process can be implemented through several fully connected, ReLU activation functions, and several convolution, normalization, and ReLU activation functions. Specifically, the view conversion process can be implemented through full connection, ReLU activation function, full connection, ReLU activation function, 1×1 convolution, normalization, ReLU activation function, 3×3 convolution, normalization, and ReLU activation function. The view conversion process is processed by two fully connected layers, whose main function is to map features of different sizes into a fixed-size feature space. This operation enables the network to handle inputs of different resolutions while providing feature maps of uniform size for subsequent convolution operations. Subsequently, the 1×1 and 3×3 convolution layers further extract and process the features.

[0057] The operation of attention enhancement processing is as follows: split the view transformation map into two feature maps to obtain the first split map and the second split map; the first split map is subjected to group normalization and self-attention processing to obtain the attention feature map; the attention feature map and the second split map are convolved and normalized after element-by-element addition to obtain the feature map to be spliced. The input in attention enhancement processing is divided into two parts and processed by different branches. One part is first group normalized and then divided into three parts: query (Q), key (K) and value (V). Next, the attention score is calculated using Q and K, and the attention weight is applied to V. The core is to calculate the inner product of the Q, K and V vectors, and use the calculation results as the weights of different feature vectors to realize the "focusing" function of the attention mechanism. The weighted features are fused with the other part of the input. Finally, the fused features are further refined through a 1×1 convolutional layer to ensure the best feature representation in the output.

[0058] The operation of the lane extraction network can be implemented through the detection head.

[0059] S2: The lane image to be processed is processed by the trained 3D lane line generation model to obtain the 3D lane line. Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 .in, Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 The horizontal coordinate (x-axis), vertical coordinate (y-axis) and vertical coordinate (z-axis) in (b) all mean distance, and the unit is m; Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 The horizontal axis (x-axis) and vertical axis (y-axis) in (c) both represent distance, with the unit being m.

[0060] This embodiment further provides a 3D lane line generation system based on frequency perception feature fusion, which is used to implement the above-mentioned 3D lane line generation method based on frequency perception feature fusion, including:

[0061] The training 3D lane line generation model generation module is used to obtain several road images and corresponding lane line mark position data to form a lane image training set; the lane image training set is used to train a 3D lane line generation model including a multi-scale feature extraction network, a frequency pyramid network, a spatial transformation fusion network and a lane line extraction network to obtain a trained 3D lane line generation model; in the frequency pyramid network, several lane feature maps of different scales are respectively processed by feature dimension alignment to obtain several image dimension perception feature maps of different scales; the low-scale image dimension perception feature map is subjected to frequency perception feature fusion processing with the medium-scale image dimension perception feature map, and then the frequency perception feature fusion processing is performed with the high-scale image dimension perception feature map, and the fusion result is extracted to obtain several The frequency-aware cross-fusion feature map is used to perform operations in the spatial transformation fusion network; the frequency-aware feature fusion processing operation is specifically as follows: the high-frequency features of the first-scale image dimension perception feature map and the second-scale image dimension perception feature map are cross-processed with low-frequency features to obtain a low-frequency cross-feature map; the first-scale image dimension perception feature map and the cross-low-frequency feature map are fused with low-frequency features to obtain a low-frequency fusion feature map; the convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map are cross-processed with high-frequency features to obtain a first frequency domain perception fusion feature map; the first-scale image dimension perception feature map and the low-frequency cross-feature map are processed by a low-pass filter, and then the first frequency domain perception fusion feature map is processed by element-by-element addition to obtain a shallow medium-scale frequency perception cross-fusion feature map;

[0062] The 3D lane line generation module is used to process the lane image to be processed through the trained 3D lane line generation model to obtain the 3D lane line.

[0063] This embodiment also provides a 3D lane line generation device based on frequency perception feature fusion, including a processor and a memory, wherein the processor implements the above-mentioned 3D lane line generation method based on frequency perception feature fusion when executing a computer program stored in the memory.

[0064] This embodiment also provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-mentioned 3D lane line generation method based on frequency-aware feature fusion.

[0065] The present embodiment provides a 3D lane line generation method based on frequency-aware feature fusion. The method improves the accuracy and robustness of 3D lane line generation by designing a multi-scale feature extraction network for extracting features of different depth scales of an image, a frequency pyramid network for extracting high-frequency information from high-resolution features and fusing them with low-resolution features to enhance the expressiveness of low-resolution features, a spatial transformation fusion network for extracting more robust and advanced BEV features, and a 3D lane line generation model of a lane line extraction network for generating 3D lane lines. The method can better cope with changes in complex environments while maintaining good real-time performance. The method is suitable for autonomous driving and assisted driving systems, and shows broad application prospects.

Claims

1. A 3D lane line generation method based on frequency perception feature fusion, characterized in that: The following operations are included: S1. Acquire a number of road images and corresponding lane line marking position data to form a lane image training set; use the lane image training set to train a 3D lane line generation model including a multi-scale feature extraction network, a frequency pyramid network, a spatial transformation fusion network, and a lane line extraction network to obtain a trained 3D lane line generation model; In the frequency pyramid network, several lane feature maps of different scales are respectively processed by feature dimension alignment to obtain several image dimension perception feature maps of different scales; the low-scale image dimension perception feature map is processed by frequency perception feature fusion with the medium-scale image dimension perception feature map, and then the high-scale image dimension perception feature map is processed by frequency perception feature fusion, and the fusion results are extracted to obtain several frequency perception cross-fusion feature maps, which are used to perform operations in the spatial transformation fusion network; The operation of the frequency perception feature fusion processing is specifically as follows: the high-frequency features of the first-scale image dimension perception feature map and the second-scale image dimension perception feature map are cross-processed with low-frequency features to obtain a low-frequency cross-feature map; The first scale image dimension perception feature map and the cross low-frequency feature map are processed by low-frequency feature fusion to obtain a low-frequency fusion feature map; The convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map are cross-processed with high-frequency features to obtain the first frequency domain perception fusion feature map; the first-scale image dimension perception feature map and the low-frequency cross feature map are processed by a low-pass filter, and then added element by element with the first frequency domain perception fusion feature map to obtain a shallow-medium scale frequency perception cross fusion feature map; S2. The lane image to be processed is processed by the trained 3D lane line generation model to obtain 3D lane lines.

2. The 3D lane line generation method based on frequency perception feature fusion according to claim 1, characterized in that: The processing operations in the multi-scale feature extraction network are as follows: The lane image is processed by the residual network to obtain a shallow-scale lane feature map; the shallow-scale lane feature map is processed by the residual network to obtain a mid-scale lane feature map; The middle-scale lane feature map is processed by convolution to obtain the deep-scale lane feature map.

3. The 3D lane line generation method based on frequency perception feature fusion according to claim 1, characterized in that: The operation of the spatial transformation fusion network is specifically as follows: Several frequency-aware cross-fusion feature maps are processed by view conversion and attention enhancement respectively, and then spliced ​​to obtain a 3D fusion feature map; The view conversion process can be implemented through several fully connected, ReLU activation functions, as well as several convolution, normalization, and ReLU activation functions; The operation of the attention enhancement processing is specifically as follows: splitting the view transformation map into two feature maps to obtain a first split map and a second split map; the first split map is subjected to group normalization processing and self-attention processing to obtain an attention feature map; The attention feature map and the second split map are added element by element, then convolved and normalized to obtain the feature map to be spliced.

4. The 3D lane line generation method based on frequency perception feature fusion according to claim 1, characterized in that: The method for obtaining the high-frequency features of the second-scale image dimension perceptual feature map is specifically as follows: The second-scale image dimension perception feature map is processed by 1×1 convolution to obtain a second-scale initial convolution map; the second-scale initial convolution map is processed by 3×3 convolution to obtain a second-scale reconvolution map; the second-scale reconvolution map is processed by high-pass filtering, and then processed by high-pass filtering and convolution with the second-scale initial convolution map to obtain the high-frequency features of the second-scale image dimension perception feature map as the second-scale high-frequency feature map.

5. The 3D lane line generation method based on frequency perception feature fusion according to claim 4 is characterized in that: The operation of the high-pass filter is as follows: The second scale high-pass filter image and the second scale initial convolution image are processed by a low-pass filter to obtain a low-pass filter cross feature map; the difference image between the second scale initial convolution image and the low-pass filter cross feature map is added element by element to the second scale initial convolution image to perform a convolution operation; The second-scale high-pass filtered image is obtained by performing high-pass filtering on the second-scale reconvolution image.

6. The 3D lane line generation method based on frequency perception feature fusion according to claim 1, characterized in that: The specific operation of high-frequency feature cross processing is: The convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map are added element by element and high-pass filtered, and then processed with the second-scale image dimension perception feature map by a high-pass filter to obtain the first frequency domain perception fusion feature map.

7. The 3D lane line generation method based on frequency perception feature fusion according to claim 1, characterized in that: The specific operation of low-frequency feature cross processing is: The high-frequency features of the second-scale image dimension perception feature map are processed by low-pass filtering, and then the convolution features of the first-scale image dimension perception feature map are processed by low-pass filtering, and then added element by element with the high-frequency features of the second-scale image dimension perception feature map to obtain a low-frequency cross feature map.

8. A 3D lane line generation system based on frequency perception feature fusion, used to implement the 3D lane line generation method based on frequency perception feature fusion according to claim 1, characterized in that: include: A 3D lane line generation model generation module is trained to obtain a number of road images and corresponding lane line marking position data to form a lane image training set; A lane image training set is used to train a 3D lane line generation model including a multi-scale feature extraction network, a frequency pyramid network, a spatial transformation fusion network and a lane line extraction network to obtain a trained 3D lane line generation model; in the frequency pyramid network, several lane feature maps of different scales are respectively processed by feature dimension alignment to obtain several image dimension perception feature maps of different scales; a low-scale image dimension perception feature map is subjected to frequency perception feature fusion processing with a medium-scale image dimension perception feature map, and then subjected to frequency perception feature fusion processing with a high-scale image dimension perception feature map, and the fusion result is extracted to obtain several frequency perception cross-fusion feature maps for executing operations in the spatial transformation fusion network; the frequency perception feature fusion processing operation is specifically as follows: the high-frequency features of the first-scale image dimension perception feature map and the second-scale image dimension perception feature map are subjected to low-frequency feature cross-processing to obtain a low-frequency cross-feature map; The first scale image dimension perception feature map and the cross low-frequency feature map are processed by low-frequency feature fusion to obtain a low-frequency fusion feature map; The convolution features of the low-frequency fusion feature map and the second-scale image dimension perception feature map are cross-processed with high-frequency features to obtain the first frequency domain perception fusion feature map; the first-scale image dimension perception feature map and the low-frequency cross feature map are processed by a low-pass filter, and then added element by element with the first frequency domain perception fusion feature map to obtain a shallow-medium scale frequency perception cross fusion feature map; The 3D lane line generation module is used to process the lane image to be processed through the trained 3D lane line generation model to obtain the 3D lane line.

9. A 3D lane line generation device based on frequency perception feature fusion, characterized in that: The method comprises a processor and a memory, wherein the processor implements the 3D lane line generation method based on frequency-aware feature fusion as described in any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the 3D lane line generation method based on frequency-aware feature fusion as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Face deep false detection method based on multi-modal feature fusion

    CN115880749A

  • Detection method using fusion network based on attention mechanism, and terminal device

    US11222217B1