Lane line segmentation method and device, electronic equipment and storage medium

By introducing multi-scale entanglement learning module and global adaptive attention mechanism module into the lane line segmentation network, the problem of low segmentation accuracy in traditional methods in complex scenarios is solved, and more efficient lane line segmentation and stronger robustness are achieved.

CN119942128AActive Publication Date: 2025-05-06STREAMAP TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510429114.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

When traditional lane line segmentation methods deal with complex shape characteristics, dynamic lighting changes and disturbances in scenes, the segmentation accuracy is not high and it is difficult to effectively deal with diverse scenarios.

Method used

A lane line segmentation method is adopted. By obtaining road image information and inputting it into a pre-established lane line segmentation network, multi-scale local features and frequency features are extracted using the backbone network, the neck network performs feature fusion, and the detection head network outputs lane lines. The network includes five convolution modules, multi-scale entanglement learning modules and global adaptive attention mechanism modules to enhance feature extraction and fusion capabilities.

Benefits of technology

It improves the segmentation accuracy of lane lines, can handle lane lines in complex scenarios more effectively, and enhances the robustness and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942128A_ABST
    Figure CN119942128A_ABST
Patent Text Reader

Abstract

The invention provides a lane line segmentation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining the image information of a road, and enabling the image information to comprise a lane line; the image information is input into a pre-established lane line segmentation network, lane lines in the image information are obtained, the lane line segmentation network comprises a backbone network, a neck network and a detection head network, the input of the backbone network is the image information, the backbone network is used for extracting multi-scale local features and frequency features in the image information, and the detection head network is used for detecting the lane lines in the image information; the local feature and the frequency feature are fused to obtain an initial fusion feature, the initial fusion feature is output to the neck network, the neck network is used for carrying out feature fusion on the initial fusion feature to obtain a fusion feature and outputting the fusion feature to the detection head network, and the detection head network is used for outputting a lane line in the image information based on the fusion feature. The segmentation precision of the lane line can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of lane line segmentation, and in particular, relates to a lane line segmentation method, device, electronic device and storage medium. Background Art

[0002] At present, with the rapid development of intelligent driving and advanced driver assistance systems, lane segmentation, as an important part of road scene understanding, has received extensive attention. The accuracy and robustness of lane segmentation directly affect the path planning and decision-making control of the vehicle. However, due to the complex shape characteristics of lanes, dynamic lighting changes, and the influence of interference objects in the scene (such as shadows, worn markings, occlusions, etc.), traditional methods have limitations in processing diverse scenes and the segmentation accuracy is not high. Summary of the invention

[0003] In response to the above problems, embodiments of the present application provide a lane line segmentation method, device, electronic device and storage medium, which can improve the lane line segmentation accuracy.

[0004] In a first aspect, an embodiment of the present application provides a lane line segmentation method, the method comprising: Acquire image information of a road, wherein the image information includes: lane lines; The image information is input into a pre-established lane line segmentation network to obtain the lane lines in the image information. The lane line segmentation network includes: a backbone network, a neck network and a detection head network. The input of the backbone network is the image information. The backbone network is used to extract multi-scale local features and frequency features in the image information, and fuse the local features with the frequency features to obtain initial fused features, and output the initial fused features to the neck network. The neck network is used to fuse the initial fused features to obtain fused features, and output the fused features to the detection head network. The detection head network is used to output the lane lines in the image information based on the fused features.

[0005] In some embodiments, the backbone network includes: five convolution modules, a first cross-stage partial connection layer and three multi-scale entanglement learning modules, the convolution module is used to extract basic features, the first cross-stage partial connection layer is used to enhance feature fusion, and the multi-scale entanglement learning module is used to extract multi-scale local features and frequency features, and fuse them to obtain initial fusion features, wherein the five convolution modules include: a first convolution module, a second convolution module, a third convolution module, a fourth convolution module and a fifth convolution module, and the three multi-scale entanglement learning modules include: a first multi-scale entanglement learning module, a second multi-scale entanglement learning module and a third multi-scale entanglement learning module, the input of the first convolution module is image information, and the input of the second convolution module is The output of the first convolution module, the input of the first cross-stage partially connected layer is the input of the second convolution module, the output of the first cross-stage partially connected layer is the input of the third convolution module, the input of the first multi-scale entangled learning module is the output of the third convolution module, the output of the first multi-scale entangled learning module is the input of the neck network and the fourth convolution module, the output of the fourth convolution module is the input of the second multi-scale entangled learning module, the output of the second multi-scale entangled learning module is the input of the neck network and the fifth convolution module, the output of the fifth convolution module is the input of the third multi-scale entangled learning module, and the output of the third multi-scale entangled learning module is the input of the neck network.

[0006] In some embodiments, the multi-scale entanglement learning module includes: a multi-scale local feature extraction unit, a frequency operation unit and an output unit, the input of the multi-scale local feature extraction unit is feature information, the multi-scale local feature extraction unit is used to extract multi-scale features based on parallel convolution layers, and splice and fuse the multi-scale features, and input the spliced ​​multi-scale features into the frequency operation unit and the output unit, the frequency operation unit is used to perform Fourier transform on the spliced ​​multi-scale features to obtain frequency domain features, and perform point convolution processing on the frequency domain features and then perform inverse Fourier transform to obtain frequency domain enhanced features, and output the frequency domain enhanced features to the output unit, the output unit is used to perform residual connection fusion on the frequency domain enhanced features and the spliced ​​multi-scale features, and adjust the channel weights to generate initial fusion features, and output the initial fusion features to the neck network.

[0007] In some embodiments, the neck network includes: multiple upsampling modules, multiple connection modules, multiple cross-stage partial connection layers, multiple convolution modules and multiple global adaptive attention mechanism modules, the upsampling module is used to increase the size of the feature map, the connection module is used to perform a splicing operation and merge multiple feature maps along the channel dimension, the cross-stage partial connection layer is used to enhance feature fusion, the global adaptive attention mechanism module is used to fuse the global information of the channel and space, and output the fused features to the detection head network, the multiple upsampling modules include: a first upsampling module and a second upsampling module, the multiple connection modules include: a first connection module, a second connection module, a third A connection module and a fourth connection module, multiple cross-stage partially connected layers include: a second cross-stage partially connected layer, a third cross-stage partially connected layer, a fourth cross-stage partially connected layer and a fifth cross-stage partially connected layer, multiple convolution modules include: a sixth convolution module and a seventh convolution module, multiple global adaptive attention mechanism modules include: a first global adaptive attention mechanism module, a second global adaptive attention mechanism module and a third global adaptive attention mechanism module, wherein the input of the first connection module is the output of the second multi-scale entangled learning module and the first upsampling module, the input of the first upsampling module is the output of the third multi-scale entangled learning module, and The input of the second cross-stage partially connected layer is the output of the first connection module, the output of the second cross-stage partially connected layer is the input of the second upsampling module and the fourth connection module, the input of the second connection module is the output of the second upsampling module and the first multi-scale entangled learning module, the output of the second connection module is the input of the third cross-stage partially connected layer, the output of the third cross-stage partially connected layer is the input of the first global adaptive attention mechanism module and the seventh convolution module, the output of the first global adaptive attention mechanism module is the input of the detection head network, the output of the seventh convolution module is the input of the fourth connection module, the output of the fourth connection module is the input of the fourth cross-stage partially connected layer, the output of the fourth cross-stage partially connected layer is the input of the second global adaptive attention mechanism module and the sixth convolution module, the output of the second global adaptive attention mechanism module is the input of the detection head network, the output of the sixth convolution module is the input of the third connection module, the output of the third connection module is the input of the fifth cross-stage partially connected layer, the output of the fifth cross-stage partially connected layer is the input of the third global adaptive attention mechanism module, and the output of the third global adaptive attention mechanism module is the input of the detection head network.

[0008] In some embodiments, the global adaptive attention mechanism module includes: a multi-scale channel attention mechanism unit and a multi-headed axial mixed self-attention mechanism unit, the multi-scale channel attention mechanism unit is used to input a feature map and output a channel-weighted feature map, the input of the multi-headed axial mixed self-attention mechanism unit is a channel-weighted feature map, the output of the multi-headed axial mixed self-attention mechanism unit is a fused feature after spatial attention weighting, and the fused feature is output to the detection head network.

[0009] In some embodiments, the multi-scale channel attention mechanism unit is used to extract features through multiple parallel convolutions, and to splice the extracted features to obtain spliced ​​feature information, and to perform global average pooling and maximum pooling on the spliced ​​feature information, and to generate channel weights through a multi-layer perceptron, and to perform point multiplication based on the channel weights and the feature information input into the multi-scale channel attention mechanism unit to obtain channel-weighted feature information, and to output the channel-weighted feature information to the multi-headed axial hybrid self-attention mechanism unit, and the multi-headed axial hybrid self-attention mechanism unit is used to perform axial feature decomposition on the channel-weighted feature information to obtain each axial feature, to generate a query matrix, a key matrix and a value matrix based on each axial feature, and to calculate an attention score based on the query matrix and the key matrix, and to multiply the attention features of each axis based on the attention score and the value matrix, and to multiply the attention features of each axis, and to generate spatial weights based on the multiplied attention features, and to perform point multiplication of the spatial weights and the channel-weighted feature information to obtain the fused feature.

[0010] In some embodiments, the regression loss function of the lane segmentation network includes:

[0011] in, , is a fixed constant, , , , is the square of the Euclidean distance between the center point of the real box and the predicted box, is the square of the diagonal length of the rectangle surrounding the real box and the predicted box, is the width and height of the real frame, is the width and height of the predicted box, and IoU is the intersection over union of the predicted box and the true box. The coordinate value of the prediction box in the t direction, is the coordinate value of the prediction box in the t direction, is the coordinate value of the real frame in the t direction, where the t direction is the x-axis direction or the y-axis direction, and c tIt is the width of the minimum enclosing rectangle of the predicted box and the real box in the x-axis direction or the height in the y-axis direction.

[0012] In a second aspect, an embodiment of the present application provides a lane line segmentation device, comprising: An acquisition module, used to acquire image information of a road, wherein the image information includes: lane lines; A segmentation module is used to input the image information into a pre-established lane line segmentation network to obtain the lane lines in the image information. The lane line segmentation network includes: a backbone network, a neck network and a detection head network. The input of the backbone network is the image information. The backbone network is used to extract multi-scale local features and frequency features in the image information, and fuse the local features with the frequency features to obtain initial fused features, and output the initial fused features to the neck network. The neck network is used to fuse the initial fused features to obtain fused features, and output the fused features to the detection head network. The detection head network is used to output the lane lines in the image information based on the fused features.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method provided in the first aspect when executing the computer program.

[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method provided in the first aspect is implemented.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it is at least used to implement a method as described in any one of the first aspect or the third aspect.

[0016] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The lane line segmentation method provided in the embodiment of the present application obtains image information of a road, wherein the image information includes: lane lines; the image information is input into a pre-established lane line segmentation network to obtain the lane lines in the image information, the lane line segmentation network includes: a backbone network, a neck network and a detection head network, the input of the backbone network is the image information, the backbone network is used to extract multi-scale local features and frequency features in the image information, and fuse the local features with the frequency features to obtain initial fused features, and output the initial fused features to the neck network, the neck network is used to fuse the initial fused features to obtain fused features, and output the fused features to the detection head network, the detection head network is used to output the lane lines in the image information based on the fused features, which can improve the segmentation accuracy of the lane lines.

[0017] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0019] Figure 1 A schematic diagram of a flow chart of a lane line segmentation method provided by the present application; Figure 2 A schematic diagram of the number of road lane line categories provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a lane segmentation network provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of a multi-scale entanglement learning module provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of a global adaptive attention mechanism module provided in an embodiment of the present application; Figure 6 A schematic diagram of the structure of a lane line segmentation device provided in an embodiment of the present application; Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0021] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0022] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0023] As used in the specification of this application and the appended claims, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrases "if it is determined" or "if it is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce detected" or "in response to detecting" depending on the context.

[0024] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0025] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. appearing in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0026] Before introducing the embodiments of the present application, a brief introduction to the related technologies and problems in the related technologies is given: At present, with the rapid development of intelligent driving and advanced driver assistance systems, lane segmentation, as an important part of road scene understanding, has received extensive attention. The accuracy and robustness of lane segmentation directly affect the path planning and decision-making control of the vehicle. However, due to the complex shape characteristics of lanes, dynamic lighting changes, and the influence of interference objects in the scene (such as shadows, marker line wear, occlusions, etc.), traditional methods have limitations in dealing with diverse scenes. In recent years, lane segmentation methods based on deep learning have made significant progress. Through convolutional neural networks (CNNs), attention mechanisms, and specific task optimization, results have been achieved in improving the accuracy and robustness of lane segmentation. However, these methods still face the following challenges: 1. Insufficient shape feature capture: Existing methods have limited ability to capture complex shapes of lanes, especially the segmentation effect of lanes with large curvature or discontinuous lanes is poor. 2. Insufficient fusion of global information and local information: In lane segmentation, it is necessary to capture global road layout information and pay attention to local detail features, which puts higher requirements on the design of the model. 3. Loss function optimization problem: The traditional IoU (Intersection over Union) loss does not perform well in the segmentation of small objects or dense areas, resulting in inaccurate edge details of lane line prediction.

[0027] Based on the technical problems of related technologies, the embodiment of the present application provides a lane line segmentation method that can be applied to electronic devices. The electronic devices may include: mobile phones, tablet computers, vehicle-mounted devices, wearable devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. The embodiment of the present application does not impose any restrictions on the specific types of electronic devices.

[0028] The present application provides a lane line segmentation method. Figure 1 The present application provides a flow chart of a lane line segmentation method, such as Figure 1 As shown, the method includes: Step S101, obtaining image information of a road, wherein the image information includes lane lines.

[0029] In the embodiment of the present application, the image information of the road is the original road data obtained by a vehicle-mounted camera, lidar or high-precision map, including elements such as lane lines, curbs, and vehicles.

[0030] Step S102, input the image information into a pre-established lane line segmentation network to obtain the lane lines in the image information. The lane line segmentation network includes: a backbone network, a neck network and a detection head network. The input of the backbone network is the image information. The backbone network is used to extract multi-scale local features and frequency features in the image information, and fuse the local features and frequency features to obtain initial fused features, and output the initial fused features to the neck network. The neck network is used to fuse the initial fused features to obtain fused features, and output the fused features to the detection head network. The detection head network is used to output the lane lines in the image information based on the fused features.

[0031] In the embodiment of the present application, the lane line segmentation network is an end-to-end deep learning model that inputs image information and outputs pixel-level lane line masks (such as white solid lines, yellow dashed lines, etc.). The lane line segmentation network can be represented as: AE-LaneNet network. The backbone network fuses multi-scale local features and frequency features to solve the problem of discontinuous segmentation of curved / intermittent lane lines.

[0032] In the embodiment of the present application, the lane line segmentation network can be trained by the data set. For example, traffic markings and road signs on urban roads can be collected and analyzed. All pictures of the data set are from real vehicle scenes, including nearly 20,000 high-resolution annotated images, each with a resolution of 8 million pixels. High-resolution images can provide more detailed lane line information, especially in subtle lane markings, different road conditions and complex traffic environments, and can capture more details. The data set covers 28 different types of lane line annotations, which cover common lane line types, such as white solid lines, yellow solid lines, double white solid lines, etc., and also include some special marking types, such as variable lane lines, curbs, speed bumps, etc. This diversity ensures that the lane line segmentation network can perform accurate segmentation under various road conditions, not only for highways, but also for urban streets and complex road sections. The data set uses a high-quality manual annotation method to ensure that the annotation of each image is accurate and can fully reflect the shape and changes of lane lines in the real world. This provides a solid foundation for the subsequent lane line segmentation network. In order to clearly show the common characteristics and distribution of various lane lines, detailed data statistics were conducted for each category and the results were presented in the form of charts. Figure 2 A schematic diagram of the number of road lane line categories provided in an embodiment of the present application, such as Figure 2 As shown, the figure includes the quantity corresponding to each category.

[0033] After obtaining the dataset, the dataset can be cleaned and its annotated labels can be converted into the format required for model training. The dataset can be divided into training set and verification set in a ratio of 8:2. The lane segmentation network can be trained with the training set, and the lane segmentation network can be verified with the verification set to ensure that the lane segmentation network can also show good generalization ability on unseen data.

[0034] In order to enhance the robustness and adaptability of the lane segmentation network, the data can be enhanced before being input into the lane segmentation network. The data enhancement includes but is not limited to randomly adjusting the image size, adding random noise, and simulating different weather conditions to simulate various scenarios that may be encountered in the real world. The enhanced images are then input into the network so that the lane segmentation network can stably learn under a wider range of conditions.

[0035] In the process of training the lane segmentation network, the gradient descent method can be used to iteratively update the model parameters and find the optimal solution by minimizing the loss function. As the training progresses, the model gradually learns how to make accurate predictions in various tasks. After the training is completed, the final model file is saved, which contains the network structure information and the optimized weight parameters to ensure that it can be directly loaded and used in future application deployment.

[0036] In the embodiment of the present application, by pre-establishing a lane line segmentation network, after obtaining image information, it can be input into the lane line segmentation network, thereby realizing lane line segmentation.

[0037] The lane line segmentation method provided in the embodiment of the present application obtains image information of the road, wherein the image information includes: lane lines; inputs the image information into a pre-established lane line segmentation network to obtain the lane lines in the image information, the lane line segmentation network includes: a backbone network, a neck network and a detection head network, the input of the backbone network is the image information, the backbone network is used to extract multi-scale local features and frequency features in the image information, and fuse the local features and frequency features to obtain initial fused features, and output the initial fused features to the neck network, the neck network is used to fuse the initial fused features to obtain fused features, and output the fused features to the detection head network, the detection head network is used to output the lane lines in the image information based on the fused features, which can improve the segmentation accuracy of the lane lines.

[0038] In some embodiments, the backbone network includes: five convolution modules, a first cross-stage partial connection layer and three multi-scale entanglement learning modules (MSEL), the convolution module is used to extract basic features, the first cross-stage partial connection layer is used to enhance feature fusion, and the multi-scale entanglement learning module is used to extract multi-scale local features and frequency features, and fuse them to obtain initial fused features. The five convolution modules include: a first convolution module, a second convolution module, a third convolution module, a fourth convolution module and a fifth convolution module. The three multi-scale entanglement learning modules include: a first multi-scale entanglement learning module, a second multi-scale entanglement learning module and a third multi-scale entanglement learning module. The input of the first convolution module is image information, the input of the second convolution module is the output of the first convolution module, the input of the first cross-stage partial connection layer is the input of the second convolution module, the output of the first cross-stage partial connection layer is the input of the third convolution module, the input of the first multi-scale entanglement learning module is the output of the third convolution module, the output of the first multi-scale entanglement learning module is the input of the neck network and the fourth convolution module, the output of the fourth convolution module is the input of the second multi-scale entanglement learning module, the output of the second multi-scale entanglement learning module is the input of the neck network and the fifth convolution module, the output of the fifth convolution module is the input of the third multi-scale entanglement learning module, and the output of the third multi-scale entanglement learning module is the input of the neck network.

[0039] In the embodiment of the present application, Figure 3 A schematic diagram of the structure of a lane segmentation network provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the convolution module is represented by ConvModule, and the convolution module is used for basic feature extraction and downsampling. The cross-stage partial connection layer is represented by CSPLayer. The first cross-stage partial connection layer can reduce computational redundancy and enhance gradient flow. The first cross-stage partial connection layer can split the features into two parts, and only half of them are involved in the calculation. The multi-scale entanglement learning module is represented by MSEL. The first convolution module, the second convolution module, the third convolution module, the fourth convolution module and the fifth convolution module are respectively: the first ConvModule, the second ConvModule, the third ConvModule, the fourth ConvModule and the fifth ConvModule. The first cross-stage partial connection layer is the first CSPLayer. The first multi-scale entanglement learning module, the second multi-scale entanglement learning module and the third multi-scale entanglement learning module are respectively: the first MSEL, the second MSEL and the third MSEL.

[0040] In some embodiments, the multi-scale entanglement learning module includes: a multi-scale local feature extraction unit, a frequency operation unit and an output unit. The input of the multi-scale local feature extraction unit is feature information. The multi-scale local feature extraction unit is used to extract multi-scale features based on parallel convolution layers, and splice and fuse the multi-scale features, and input the spliced ​​multi-scale features into the frequency operation unit and the output unit. The frequency operation unit is used to perform Fourier transform on the spliced ​​multi-scale features to obtain frequency domain features, and perform point convolution processing on the frequency domain features and then perform inverse Fourier transform to obtain frequency domain enhanced features, and output the frequency domain enhanced features to the output unit. The output unit is used to perform residual connection fusion on the frequency domain enhanced features and the spliced ​​multi-scale features, and adjust the channel weights to generate initial fused features, and output the initial fused features to the neck network.

[0041] In the embodiment of the present application, Figure 4 A schematic diagram of the structure of a multi-scale entanglement learning module provided in an embodiment of the present application is shown in FIG. Figure 4 As shown in the figure, the multi-scale local feature extraction unit consists of a step size of 2 The convolution and four parallel operations have a stride of 1 and a size of , , In the embodiment of the present application, 5×5 convolution is used to capture the local continuity of lane lines, 7×7 convolution is used to perceive the curvature change in the medium range, and 9×9 convolution is used to model the regularity of the interval of long-distance dashed lines.

[0042] In the embodiment of the present application, the frequency domain operation is performed in sequence with a step size of 2. Convolution, Fourier transform (FT), step size 1 It is connected with the inverse Fourier transform (IFT), and the output units include: residual, which is connected and outputted through the residual.

[0043] In the embodiment of the present application, not only more complex features are captured by large convolution kernels and features under different receptive fields are integrated, but also the influence of noise is reduced, the feature extraction capability can be more effectively improved, and there is more advantage in lane line recognition.

[0044] Assume that the input features are , the output of the multi-scale local feature unit is , the output of the output unit is , then the MSEL module can be expressed as: For the multi-scale local feature extraction unit, the multi-scale feature extraction is expressed as Splicing fusion means: The convolution results of different scales (kernel size is 5, 7, 9) are concatenated and fused to form multi-scale features. .

[0045] For the frequency operation unit, the frequency domain features obtained by Fourier transform can be expressed as: , which represents the Perform a two-dimensional Fourier transform to obtain frequency domain features ; Then it is processed by the following formula: , which represents the linear transformation of the real part of the frequency domain feature (weight , Bias ) and convolution operations (weights , Bias ),get , The inverse Fourier transform can be expressed as: , which represents the Perform a two-dimensional inverse Fourier transform to restore the time domain features .

[0046] For the output unit, the residual connection fusion can be expressed as: After obtaining the fusion features of the residual connection, the formula This formula represents the convolution operation on the real part of the fusion result (weight , Bias ), and get the output .

[0047] In the embodiment of the present application, convolution is good at extracting local features, but the perception of global features requires a deeper network or a larger receptive field. The Fourier transform can directly obtain the global frequency information of the image in the frequency domain. The MSEL module can help the model better understand the image content and accurately segment the target through its ability to fuse local and global features; the Fourier transform converts the convolution operation from the spatial domain to the frequency domain, turning the convolution operation into a point multiplication, which complements the smoothness of the convolution extracted features, making the network more adaptable to interference and greatly reducing the amount of calculation. In the case of segmentation of multiple targets and complex lane objects, the MSEL module can significantly improve the accuracy and robustness of segmentation.

[0048] In some embodiments, the neck network includes: the neck network includes: a plurality of upsampling modules, a plurality of connection modules, a plurality of cross-stage partial connection layers, a plurality of convolution modules and a plurality of global adaptive attention mechanism modules (GAAM, Global adaptive attention mechanism), an upsampling module is used to increase the size of the feature map, a connection module is used to perform a splicing operation and merge multiple feature maps along the channel dimension, a cross-stage partial connection layer is used to enhance feature fusion, a global adaptive attention mechanism module is used to fuse the global information of the channel and space, and output the fused features to the detection head network, multiple upsampling modules include: a first upsampling module and a second upsampling module, multiple connection modules include: a first connection module, a second connection module, a third connection module and a fourth connection module, multiple cross-stage partial connection layers include: a second cross-stage partial connection layer, a third cross-stage partial connection layer, a fourth cross-stage partial connection layer and a fifth cross-stage partial connection layer, multiple convolution modules include: a sixth convolution module and a seventh convolution module, multiple global adaptive attention mechanism modules include: a first global adaptive attention mechanism module, a second global adaptive attention mechanism module and a third global adaptive attention mechanism module, wherein the input of the first connection module is the output of the second multi-scale entanglement learning module and the first upsampling module, the input of the first upsampling module is the output of the third multi-scale entanglement learning module, and the second cross-stage partial connection module is used to perform a splicing operation and merge multiple feature maps along the channel dimension, the cross-stage partial connection layer is used to enhance feature fusion, and the global adaptive attention mechanism module is used to fuse the global information of the channel and space, and output the fused features to the detection head network, multiple upsampling modules include: a first connection module, a second connection module, a third connection module and a fourth connection module, multiple cross-stage partial connection layers include: a second cross-stage partial connection layer, a third cross-stage partial connection layer, a fourth cross-stage partial connection layer and a fifth cross-stage partial connection layer, multiple convolution modules include: a sixth convolution module and a seventh convolution module, multiple global adaptive attention mechanism modules include: a first global adaptive attention mechanism module, a second global adaptive attention mechanism module and a third global adaptive attention mechanism module, wherein the input of the first connection module is the output of the second multi-scale entanglement learning module and The input of the partial connection layer is the output of the first connection module, the output of the second cross-stage partial connection layer is the input of the second upsampling module and the fourth connection module, the input of the second connection module is the output of the second upsampling module and the first multi-scale entanglement learning module, the output of the second connection module is the input of the third cross-stage partial connection layer, the output of the third cross-stage partial connection layer is the input of the first global adaptive attention mechanism module and the seventh convolution module, the output of the first global adaptive attention mechanism module is the input of the detection head network, the output of the seventh convolution module is the input of the fourth connection module, the output of the fourth connection module is the input of the fourth cross-stage partial connection layer, the output of the fourth cross-stage partial connection layer is the input of the second global adaptive attention mechanism module and the sixth convolution module, the output of the second global adaptive attention mechanism module is the input of the detection head network, the output of the sixth convolution module is the input of the third connection module, the output of the third connection module is the input of the fifth cross-stage partial connection layer, the output of the fifth cross-stage partial connection layer is the input of the third global adaptive attention mechanism module, and the output of the third global adaptive attention mechanism module is the input of the detection head network.

[0049] Continue to see Figure 2, the second cross-stage partially connected layer is represented as the second CSPLayer, the third cross-stage partially connected layer is represented as the third CSPLayer, the seventh convolution module is represented as the seventh ConvModule, and the sixth convolution module is represented as the sixth ConvModule. The fourth cross-stage partially connected layer is represented as the fourth CSPLaye, and the fifth cross-stage partially connected layer is represented as the fifth CSPLayer. The first global adaptive attention mechanism module, the second global adaptive attention mechanism module, and the third global adaptive attention mechanism module are respectively represented as: the first GAAM, the second GAAM, and the third GAAM. The connection module can be represented as a Concat module.

[0050] In an embodiment of the present application, the connection module is used for multi-level feature fusion: deep semantic features (from upsampling) are combined with shallow geometric features (from MSEL), and the cross-stage partial connection layers can replace standard convolutions with deep separable convolutions to reduce the amount of calculation.

[0051] In some embodiments, the global adaptive attention mechanism module includes: a multi-scale channel attention mechanism unit and a multi-headed axial mixed self-attention mechanism unit. The multi-scale channel attention mechanism unit is used to input a feature map and output a channel-weighted feature map. The input of the multi-headed axial mixed self-attention mechanism unit is a channel-weighted feature map. The output of the multi-headed axial mixed self-attention mechanism unit is a fused feature after spatial attention weighting, and the fused feature is output to the detection head network.

[0052] In an embodiment of the present application, a multi-scale channel attention mechanism unit is used to extract features through multiple parallel convolutions, and to splice the extracted features to obtain spliced ​​feature information, and to perform global average pooling and maximum pooling on the spliced ​​feature information, and to generate channel weights through a multi-layer perceptron, and to perform point multiplication based on the channel weights and the feature information input into the multi-scale channel attention mechanism unit to obtain channel-weighted feature information, and to output the channel-weighted feature information to a multi-headed axial hybrid self-attention mechanism unit, and the multi-headed axial hybrid self-attention mechanism unit is used to perform axial feature decomposition on the channel-weighted feature information to obtain each axial feature, and to generate a query matrix, a key matrix and a value matrix based on each axial feature, and to calculate an attention score based on the query matrix and the key matrix, and to multiply the attention features of each axis based on the attention score and the value matrix, and to multiply the attention features of each axis, and to generate spatial weights based on the multiplied attention features, and to perform point multiplication of the spatial weights and the channel-weighted feature information to obtain a fused feature.

[0053] Figure 5 A schematic diagram of the structure of a global adaptive attention mechanism module provided in an embodiment of the present application is shown in FIG. Figure 5As shown in Figure 1, the global adaptive attention mechanism module includes: a multi-scale channel attention mechanism (Multi scale channel attention mechanism, MSCA) and a multi-head axial self-attention mechanism (Multi head axial self-attention mechanism, MASA). The MSCA module in turn includes a plurality of convolution kernels with a step size of 1 and a The convolution is performed by splicing, channel transformation, pooling and other operations to obtain features under different receptive fields, and the channel information of average pooling and maximum pooling is integrated. Finally, the MLP fully connected layer is used to generate The channel weight of size is multiplied with the input feature to generate a new feature map with channel weights; the multi-head axial hybrid self-attention mechanism includes average pooling in the X-axis and Y-axis directions, multiple convolution kernels with a step size of 1 and a The convolution kernel generates three learnable matrices: query, key, and value. The attention score is calculated through the query matrix and the key matrix to represent the correlation with other positions in the feature map. Then, the attention features of each axis are generated by taking the weighted dot product of the value matrix. Finally, the attention features of the X-axis and the Y-axis are multiplied and generated through softmax. The spatial attention weights of size are multiplied by the input to obtain the final output matrix.

[0054] Assume that the input features are , the MSCA output is , the MASA output is , the extracted features are concatenated to obtain the concatenated feature information which can be expressed by the following formula: , extract multi-scale features through convolution operations at different levels, and splice the original features with the features of each scale to enhance feature diversity, represents the concatenation of features along the channel dimension, , , : Represents the convolution kernel weights at different levels, which are used to extract multi-scale features.

[0055] After obtaining the spliced ​​feature information, it is processed by the following formula: , this formula is used to perform convolution, batch normalization and nonlinear activation on the concatenated features to generate high-order fusion features ,in represents batch normalization, : represents the convolution kernel weight, which is used to fuse the concatenated features. is the activation function.

[0056] Global average pooling and maximum pooling can be expressed as: , in this formula is the global average pooling, It is global maximum pooling, which captures channel importance by combining the global information of average and maximum pooling.

[0057] The calculation of channel attention weight can be expressed as: ,in, , is the weight of the fully connected layer (used for dimensionality reduction and dimensionality increase); , is the bias term, is the channel attention weight, which is generated through a fully connected layer and nonlinear transformation to represent the importance of each channel.

[0058] The dot multiplication of the channel weight and the feature information input into the multi-scale channel attention mechanism unit can be expressed as: , Represents channel-by-channel dot product, which uses attention weights to adjust the input features and enhance the response of important channels.

[0059] Height directional feature extraction: ; The query matrix is ​​represented as: ; The bond matrix is ​​represented as: ; The value matrix is ​​represented as: ; Generate a high degree of directional attention representation as: ,This formula captures highly directional spatial dependencies through the self-attention mechanism and generates attention-weighted features.

[0060] Similarly, the width direction attention can be expressed as: ; Multiplying the attention features of each axis can be expressed as: ; The generated spatial weight can be expressed as follows: is the total number of attention heads, Softmax: normalizes along the spatial dimension to generate a probability distribution, Global spatial attention weight,This formula generates global spatial weights by fusing multi-head attention and normalizing it to identify important areas in the image.

[0061] The dot product of the spatial weight and the channel-weighted feature information can be expressed by the following formula: , which uses global spatial weights to adjust features, enhance the response of key areas (such as lane lines), and suppress irrelevant background.

[0062] In the embodiment of the present application, the global adaptive attention mechanism module (GAAM) has significant advantages in the recognition of lane lines, and can efficiently fuse the global information of channels and spaces in the feature extraction process. The features of lane lines at different perspectives may have scale differences. In the MSCA module, the rich features under different receptive fields are effectively captured through multi-scale channel feature extraction operations, and the average pooling and maximum pooling strategies are combined to further enhance the expression ability of channel dimension information. In addition, lane lines are usually ductile in the image along the longitudinal and transverse directions. The MASA module uses a multi-head axial self-attention mechanism to model spatial dependencies along the X-axis and Y-axis respectively, and realizes a comprehensive perception of the position correlation in the feature map. Through the query, key, value matrix learning and attention score calculation, the global correlation between different positions of the feature map is accurately captured, and the model is improved to learn multiple feature representations in different subspaces. This weight mechanism is particularly suitable for processing complex road conditions with multiple lane lines or different types of lane lines, and enhances the generalization ability of the model. Finally, the channel and spatial weights are fused into a new feature representation to improve the model's representation ability and robustness for lane lines.

[0063] In some embodiments, the regression loss function of the lane segmentation network includes:

[0064] in, , is a fixed constant, , , , is the square of the Euclidean distance between the center point of the real box and the predicted box, is the square of the diagonal length of the rectangle surrounding the real box and the predicted box, is the width and height of the real frame, is the width and height of the predicted box, and IoU is the intersection over union of the predicted box and the true box. is the coordinate value of the prediction box in the t direction, is the coordinate value of the ground truth in the t direction, where the t direction is the x-axis direction or the y-axis direction, and c t It is the width of the minimum enclosing rectangle of the predicted box and the real box in the x-axis direction or the height in the y-axis direction. In the embodiment of the present application, the detection head network includes: a detection head and a segmentation head. The detection head is used for the object detection task (Object Detection) and outputs the bounding box (Bounding Box) and category label of the object in the image. The segmentation head is used for the semantic segmentation task (Semantic Segmentation), outputs the category label of each pixel, and generates a pixel-level mask (Mask).

[0065] The method provided in the embodiment of the present application designs the AE-LaneNet network by solving the problems of multi-scale feature fusion, detail feature identification and adjacent target distinction in the lane segmentation task. By introducing the multi-scale entangled learning module (MSEL), the global adaptive attention mechanism module (GAAM) and the distance-penalized IoU loss (DPIoU), the lane segmentation task is comprehensively optimized from three levels: feature extraction, feature fusion and loss design. The multi-scale entangled learning module makes full use of the association between features of different scales, significantly improving the network's ability to capture lane details in complex scenes; the global adaptive attention mechanism module combines multi-scale channel attention with shared spatial self-attention to enhance the network's global modeling capabilities for lane features; the distance-penalized IoU loss effectively solves the confusion problem of adjacent lane annotations by introducing directional distance constraints, thereby improving positioning accuracy.

[0066] In the embodiment of the present application, the multi-scale entanglement learning module adopts Fourier convolution hybrid technology, which can effectively identify and segment lane lines with high curvature, and can maintain high-precision segmentation effects under different lane line shapes (such as curves, intersections, etc.), effectively solving the limitations of traditional methods in complex scenes. The global adaptive attention mechanism module is committed to realizing the deep fusion of global road layout and local detail features in the model. The module adopts multi-scale feature extraction and self-attention mechanism, which can capture local detail features while maintaining global context information. Through the adaptive multi-level information fusion method, more accurate lane line segmentation can be achieved in complex scenes, ensuring that the model can understand the overall structure of the road and grasp the local lane line features in detail to cope with difficult-to-segment areas such as complex road conditions and intersections. The regression loss function introduces additional distance penalties, which can better handle the details of the lane line, distinguish adjacent lane lines that are too close, especially in areas with dense lane lines, and improve the accuracy of lane line segmentation.

[0067] The lane segmentation network provided in the embodiment of the present application can not only efficiently adapt to complex road conditions, but also has stronger robustness and generalization capabilities, providing a more efficient and accurate technical solution for practical application scenarios such as intelligent driving.

[0068] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0069] According to the aforementioned embodiments, the embodiments of the present application provide a lane line segmentation device, and the modules included in the device and the units included in the modules can be implemented by a processor in a computer device; of course, they can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU, Central Processing Unit), a microprocessor (MPU, Microprocessor Unit), a digital signal processor (DSP, Digital Signal Processing) or a field programmable gate array (FPGA, Field Programmable Gate Array), etc.

[0070] The present application provides a lane segmentation device. Figure 6 A schematic diagram of the structure of a lane segmentation device provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the lane line segmentation device 600 includes: The acquisition module 601 is used to acquire image information of the road, wherein the image information includes: lane lines; The segmentation module 602 is used to input the image information into a pre-established lane line segmentation network to obtain the lane lines in the image information. The lane line segmentation network includes: a backbone network, a neck network and a detection head network. The input of the backbone network is the image information. The backbone network is used to extract multi-scale local features and frequency features in the image information, and fuse the local features with the frequency features to obtain initial fused features, and output the initial fused features to the neck network. The neck network is used to fuse the initial fused features to obtain fused features, and output the fused features to the detection head network. The detection head network is used to output the lane lines in the image information based on the fused features.

[0071] In some embodiments, the backbone network includes: five convolution modules, a first cross-stage partial connection layer and three multi-scale entanglement learning modules, the convolution module is used to extract basic features, the first cross-stage partial connection layer is used to enhance feature fusion, and the multi-scale entanglement learning module is used to extract multi-scale local features and frequency features, and fuse them to obtain initial fusion features, wherein the five convolution modules include: a first convolution module, a second convolution module, a third convolution module, a fourth convolution module and a fifth convolution module, and the three multi-scale entanglement learning modules include: a first multi-scale entanglement learning module, a second multi-scale entanglement learning module and a third multi-scale entanglement learning module, the input of the first convolution module is image information, and the input of the second convolution module is The output of the first convolution module, the input of the first cross-stage partially connected layer is the input of the second convolution module, the output of the first cross-stage partially connected layer is the input of the third convolution module, the input of the first multi-scale entangled learning module is the output of the third convolution module, the output of the first multi-scale entangled learning module is the input of the neck network and the fourth convolution module, the output of the fourth convolution module is the input of the second multi-scale entangled learning module, the output of the second multi-scale entangled learning module is the input of the neck network and the fifth convolution module, the output of the fifth convolution module is the input of the third multi-scale entangled learning module, and the output of the third multi-scale entangled learning module is the input of the neck network.

[0072] In some embodiments, the multi-scale entanglement learning module includes: a multi-scale local feature extraction unit, a frequency operation unit and an output unit, the input of the multi-scale local feature extraction unit is feature information, the multi-scale local feature extraction unit is used to extract multi-scale features based on parallel convolution layers, and splice and fuse the multi-scale features, and input the spliced ​​multi-scale features into the frequency operation unit and the output unit, the frequency operation unit is used to perform Fourier transform on the spliced ​​multi-scale features to obtain frequency domain features, and perform point convolution processing on the frequency domain features and then perform inverse Fourier transform to obtain frequency domain enhanced features, and output the frequency domain enhanced features to the output unit, the output unit is used to perform residual connection fusion on the frequency domain enhanced features and the spliced ​​multi-scale features, and adjust the channel weights to generate initial fusion features, and output the initial fusion features to the neck network.

[0073] In some embodiments, the neck network includes: multiple upsampling modules, multiple connection modules, multiple cross-stage partial connection layers, multiple convolution modules and multiple global adaptive attention mechanism modules, the upsampling module is used to increase the size of the feature map, the connection module is used to perform a splicing operation and merge multiple feature maps along the channel dimension, the cross-stage partial connection layer is used to enhance feature fusion, the global adaptive attention mechanism module is used to fuse the global information of the channel and space, and output the fused features to the detection head network, the multiple upsampling modules include: a first upsampling module and a second upsampling module, the multiple connection modules include: a first connection module, a second connection module, a third A connection module and a fourth connection module, multiple cross-stage partially connected layers include: a second cross-stage partially connected layer, a third cross-stage partially connected layer, a fourth cross-stage partially connected layer and a fifth cross-stage partially connected layer, multiple convolution modules include: a sixth convolution module and a seventh convolution module, multiple global adaptive attention mechanism modules include: a first global adaptive attention mechanism module, a second global adaptive attention mechanism module and a third global adaptive attention mechanism module, wherein the input of the first connection module is the output of the second multi-scale entangled learning module and the first upsampling module, the input of the first upsampling module is the output of the third multi-scale entangled learning module, and The input of the second cross-stage partially connected layer is the output of the first connection module, the output of the second cross-stage partially connected layer is the input of the second upsampling module and the fourth connection module, the input of the second connection module is the output of the second upsampling module and the first multi-scale entangled learning module, the output of the second connection module is the input of the third cross-stage partially connected layer, the output of the third cross-stage partially connected layer is the input of the first global adaptive attention mechanism module and the seventh convolution module, the output of the first global adaptive attention mechanism module is the input of the detection head network, the output of the seventh convolution module is the input of the fourth connection module, the output of the fourth connection module is the input of the fourth cross-stage partially connected layer, the output of the fourth cross-stage partially connected layer is the input of the second global adaptive attention mechanism module and the sixth convolution module, the output of the second global adaptive attention mechanism module is the input of the detection head network, the output of the sixth convolution module is the input of the third connection module, the output of the third connection module is the input of the fifth cross-stage partially connected layer, the output of the fifth cross-stage partially connected layer is the input of the third global adaptive attention mechanism module, and the output of the third global adaptive attention mechanism module is the input of the detection head network.

[0074] In some embodiments, the global adaptive attention mechanism module includes: a multi-scale channel attention mechanism unit and a multi-headed axial mixed self-attention mechanism unit, the multi-scale channel attention mechanism unit is used to input a feature map and output a channel-weighted feature map, the input of the multi-headed axial mixed self-attention mechanism unit is a channel-weighted feature map, the output of the multi-headed axial mixed self-attention mechanism unit is a fused feature after spatial attention weighting, and the fused feature is output to the detection head network.

[0075] In some embodiments, the multi-scale channel attention mechanism unit is used to extract features through multiple parallel convolutions, and to splice the extracted features to obtain spliced ​​feature information, and to perform global average pooling and maximum pooling on the spliced ​​feature information, and to generate channel weights through a multi-layer perceptron, and to perform point multiplication based on the channel weights and the feature information input into the multi-scale channel attention mechanism unit to obtain channel-weighted feature information, and to output the channel-weighted feature information to the multi-headed axial hybrid self-attention mechanism unit, and the multi-headed axial hybrid self-attention mechanism unit is used to perform axial feature decomposition on the channel-weighted feature information to obtain each axial feature, to generate a query matrix, a key matrix and a value matrix based on each axial feature, and to calculate an attention score based on the query matrix and the key matrix, and to multiply the attention features of each axis based on the attention score and the value matrix, and to multiply the attention features of each axis, and to generate spatial weights based on the multiplied attention features, and to perform point multiplication of the spatial weights and the channel-weighted feature information to obtain the fused feature.

[0076] In some embodiments, the regression loss function of the lane segmentation network includes:

[0077] in, , is a fixed constant, , , , is the square of the Euclidean distance between the center point of the real box and the predicted box, is the square of the diagonal length of the rectangle surrounding the real box and the predicted box, is the width and height of the real frame, is the width and height of the predicted box, and IoU is the intersection over union of the predicted box and the true box. is the coordinate value of the prediction box in the t direction, is the coordinate value of the real frame in the t direction, where the t direction is the x-axis direction or the y-axis direction, and c t It is the width of the minimum enclosing rectangle of the predicted box and the real box in the x-axis direction or the height in the y-axis direction.

[0078] in addition, Figure 6 The lane line segmentation shown may be a software unit, a hardware unit, or a combination of software and hardware units built into existing electronic devices, or may be integrated into electronic devices as independent accessories, or may exist as independent terminal devices.

[0079] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0080] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0081] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 7 As shown, the electronic device 3 of this embodiment may include: at least one processor 30 ( Figure 7 Only one processor 30 is shown in the figure), a memory 31, and a computer program 32 stored in the memory 31 and executable on at least one processor 30. When the processor 30 executes the computer program 32, the steps in any of the above-mentioned method embodiments are implemented; or, when the processor 30 executes the computer program 32, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0082] Exemplarily, the computer program 32 may be divided into one or more modules / units, one or more modules / units are stored in the memory 31, and are executed by the processor 30 to complete the present application. One or more modules / units may be a series of computer program 32 instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3.

[0083] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program 32. When the computer program 32 is executed by the processor 30, the steps in the above-mentioned method embodiments can be implemented.

[0084] An embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0085] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. According to this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program 32. The computer program 32 can be stored in a computer-readable storage medium. When the computer program 32 is executed by the processor 30, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program 32 includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the terminal, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0086] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0087] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0088] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0089] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0090] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

[0091] The relevant user personal information that may be involved in the various embodiments of this application is strictly in accordance with the requirements of laws and regulations, following the principles of legality, legitimacy and necessity, based on the reasonable purposes of business scenarios, to process the personal information that users actively provide during the use of products / services or generated due to the use of products / services, as well as the personal information obtained with the user's authorization.

[0092] The user personal information processed by the applicant will vary depending on the specific product / service scenario, and shall be based on the specific scenario in which the user uses the product / service, which may involve the user's account information, device information, driving information, vehicle information or other related information. The applicant will treat the user's personal information and its processing with a high degree of diligence.

[0093] The Applicant attaches great importance to the security of user personal information and has adopted reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.

Claims

1. A lane line segmentation method, characterized in that: include: Acquire image information of a road, wherein the image information includes: lane lines; The image information is input into a pre-established lane line segmentation network to obtain the lane lines in the image information. The lane line segmentation network includes: a backbone network, a neck network and a detection head network. The input of the backbone network is the image information. The backbone network is used to extract multi-scale local features and frequency features in the image information, and fuse the local features with the frequency features to obtain initial fused features, and output the initial fused features to the neck network. The neck network is used to fuse the initial fused features to obtain fused features, and output the fused features to the detection head network. The detection head network is used to output the lane lines in the image information based on the fused features.

2. The method according to claim 1, characterized in that The backbone network includes: five convolution modules, a first cross-stage partial connection layer and three multi-scale entanglement learning modules, the convolution module is used to extract basic features, the first cross-stage partial connection layer is used to enhance feature fusion, the multi-scale entanglement learning module is used to extract multi-scale local features and frequency features, and fuse them to obtain initial fusion features, wherein the five convolution modules include: a first convolution module, a second convolution module, a third convolution module, a fourth convolution module and a fifth convolution module, the three multi-scale entanglement learning modules include: a first multi-scale entanglement learning module, a second multi-scale entanglement learning module and a third multi-scale entanglement learning module, the input of the first convolution module is image information, the input of the second convolution module is the first The output of the convolution module, the input of the first cross-stage partially connected layer is the input of the second convolution module, the output of the first cross-stage partially connected layer is the input of the third convolution module, the input of the first multi-scale entangled learning module is the output of the third convolution module, the output of the first multi-scale entangled learning module is the input of the neck network and the fourth convolution module, the output of the fourth convolution module is the input of the second multi-scale entangled learning module, the output of the second multi-scale entangled learning module is the input of the neck network and the fifth convolution module, the output of the fifth convolution module is the input of the third multi-scale entangled learning module, and the output of the third multi-scale entangled learning module is the input of the neck network.

3. The method according to claim 2, characterized in that The multi-scale entanglement learning module includes: a multi-scale local feature extraction unit, a frequency operation unit and an output unit. The input of the multi-scale local feature extraction unit is feature information. The multi-scale local feature extraction unit is used to extract multi-scale features based on parallel convolution layers, and splice and fuse the multi-scale features, and input the spliced ​​multi-scale features into the frequency operation unit and the output unit. The frequency operation unit is used to perform Fourier transform on the spliced ​​multi-scale features to obtain frequency domain features, and perform point convolution processing on the frequency domain features and then perform inverse Fourier transform to obtain frequency domain enhanced features, and output the frequency domain enhanced features to the output unit. The output unit is used to perform residual connection fusion on the frequency domain enhanced features and the spliced ​​multi-scale features, and adjust the channel weights to generate initial fused features, and output the initial fused features to the neck network.

4. The method according to claim 2, characterized in that: The neck network includes: multiple upsampling modules, multiple connection modules, multiple cross-stage partial connection layers, multiple convolution modules and multiple global adaptive attention mechanism modules, the upsampling module is used to increase the size of the feature map, the connection module is used to perform a splicing operation and merge multiple feature maps along the channel dimension, the cross-stage partial connection layer is used to enhance feature fusion, the global adaptive attention mechanism module is used to fuse the global information of the channel and space, and output the fused features to the detection head network, the multiple upsampling modules include: a first upsampling module and a second upsampling module, the multiple connection modules include: a first connection module, a second connection module, a third connection module and a third connection module The four-connection module, the multiple cross-stage partially connected layers include: the second cross-stage partially connected layer, the third cross-stage partially connected layer, the fourth cross-stage partially connected layer and the fifth cross-stage partially connected layer, the multiple convolution modules include: the sixth convolution module and the seventh convolution module, the multiple global adaptive attention mechanism modules include: the first global adaptive attention mechanism module, the second global adaptive attention mechanism module and the third global adaptive attention mechanism module, wherein the input of the first connection module is the output of the second multi-scale entangled learning module and the first upsampling module, the input of the first upsampling module is the output of the third multi-scale entangled learning module, and the second cross The input of the stage partial connection layer is the output of the first connection module, the output of the second cross-stage partial connection layer is the input of the second upsampling module and the fourth connection module, the input of the second connection module is the output of the second upsampling module and the first multi-scale entanglement learning module, the output of the second connection module is the input of the third cross-stage partial connection layer, the output of the third cross-stage partial connection layer is the input of the first global adaptive attention mechanism module and the seventh convolution module, the output of the first global adaptive attention mechanism module is the input of the detection head network, the output of the seventh convolution module is the input of the fourth connection module, the output of the fourth connection module is the input of the fourth cross-stage partial connection layer, the output of the fourth cross-stage partial connection layer is the input of the second global adaptive attention mechanism module and the sixth convolution module, the output of the second global adaptive attention mechanism module is the input of the detection head network, the output of the sixth convolution module is the input of the third connection module, the output of the third connection module is the input of the fifth cross-stage partial connection layer, the output of the fifth cross-stage partial connection layer is the input of the third global adaptive attention mechanism module, and the output of the third global adaptive attention mechanism module is the input of the detection head network.

5. The method according to claim 4, characterized in that The global adaptive attention mechanism module includes: a multi-scale channel attention mechanism unit and a multi-head axial hybrid self-attention mechanism unit. The multi-scale channel attention mechanism unit is used to input a feature map and output a channel-weighted feature map. The input of the multi-head axial hybrid self-attention mechanism unit is a channel-weighted feature map. The output of the multi-head axial hybrid self-attention mechanism unit is a fused feature after spatial attention weighting, and the fused feature is output to the detection head network.

6. The method according to claim 5, characterized in that The multi-scale channel attention mechanism unit is used to extract features through multiple parallel convolutions, and to obtain spliced ​​feature information by splicing the extracted features, and to generate channel weights through a multi-layer perceptron after global average pooling and maximum pooling of the spliced ​​feature information, and to obtain channel-weighted feature information by performing point multiplication based on the channel weights and the feature information input into the multi-scale channel attention mechanism unit, and to output the channel-weighted feature information to the multi-headed axial hybrid self-attention mechanism unit, and the multi-headed axial hybrid self-attention mechanism unit is used to perform axial feature decomposition on the channel-weighted feature information to obtain each axial feature, to generate a query matrix, a key matrix and a value matrix based on each axial feature, and to calculate an attention score based on the query matrix and the key matrix, and to multiply the attention features of each axis based on the attention score and the value matrix, and to multiply the attention features of each axis, and to generate spatial weights based on the multiplied attention features, and to perform point multiplication of the spatial weights and the channel-weighted feature information to obtain the fused feature.

7. The method according to claim 1, characterized in that The regression loss function of the lane segmentation network includes: in, , is a fixed constant, , , , is the square of the Euclidean distance between the center point of the real box and the predicted box, is the square of the diagonal length of the rectangle surrounding the real box and the predicted box, is the width and height of the real frame, is the width and height of the predicted box, and IoU is the intersection over union of the predicted box and the true box. is the coordinate value of the prediction box in the t direction, is the coordinate value of the real frame in the t direction, where the t direction is the x-axis direction or the y-axis direction, and c t It is the width of the minimum enclosing rectangle of the predicted box and the real box in the x-axis direction or the height in the y-axis direction.

8. A lane line segmentation device, characterized in that: include: An acquisition module, used to acquire image information of a road, wherein the image information includes: lane lines; A segmentation module is used to input the image information into a pre-established lane line segmentation network to obtain the lane lines in the image information. The lane line segmentation network includes: a backbone network, a neck network and a detection head network. The input of the backbone network is the image information. The backbone network is used to extract multi-scale local features and frequency features in the image information, and fuse the local features with the frequency features to obtain initial fused features, and output the initial fused features to the neck network. The neck network is used to fuse the initial fused features to obtain fused features, and output the fused features to the detection head network. The detection head network is used to output the lane lines in the image information based on the fused features.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Lane line detection method based on starting point guidance

    CN118552924A

  • Lane line detection method with row-column anchor adaptive selection capability

    CN118942067A

  • Lane line recognition method, electronic device and storage medium

    US20240071102A1

Cited By

  • Street lamp brightness lack detection method and street lamp brightness lack detection device based on deep learning

    CN120472381A

  • A method and device for detecting dim light of a street lamp based on deep learning

    CN120472381B

  • Retina artery and vein blood vessel image segmentation method and device

    CN121074956A