Automatic driving vehicle lane line identification method
Through the Attention-Unet model and feature transfer learning, the problems of low recognition accuracy, poor stability and high cost in the recognition of autonomous driving lane line are solved, and high-precision and low-cost lane line recognition are achieved.
Patent Information
- Application Number
- CN202510392370.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing lane line recognition technology of autonomous driving vehicles has problems such as low recognition accuracy, poor stability and insufficient robustness, and requires a large number of data sets to be trained, resulting in high costs.
The Attention-Unet model is combined with the feature transfer learning method, and the public data set is used as the source domain data and the algorithm to arrange the scene's data set as the target domain data. The inter-domain difference is minimized in the feature space through the attention mechanism and the maximum mean difference algorithm, model training is carried out to reduce the data set demand.
It improves the accuracy and stability of lane line recognition, enhances the robustness of the model, reduces the cost of data set production, and is suitable for a variety of scenarios.
Smart Images

Figure CN120339984A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and specifically relates to a lane line recognition method for autonomous driving vehicles. Background Art
[0002] The positioning ability of autonomous driving vehicles is the key to intelligent technology, which can ensure accurate positioning in various traffic environments. At present, expensive positioning devices (such as millimeter-wave lidar) are gradually entering the field of autonomous driving. Due to the need for a large amount of dataset support and complex algorithms, the cost is high, which limits the popularization of the technology. The existing technologies are as follows: 1. The traditional lane line recognition algorithm based on traditional OpenCV and machine learning has poor processing ability for complex environments, resulting in poor stability and recognition accuracy, and there are situations of misrecognition or non-recognition, which cannot well meet the safety requirements of autonomous driving; 2. The lane line recognition algorithm based on single deep learning lacks the generalization ability of demand scenarios, resulting in poor robustness; 3. The lane line recognition algorithm based on deep learning requires a large amount of dataset support. For most application scenarios, it is difficult to obtain a large amount of dataset for training, and the cost of obtaining a large amount of dataset is relatively high, which affects the effect of the deep learning algorithm in the corresponding practical scenarios. Summary of the Invention
[0003] The purpose of the present invention is to provide a lane line recognition method for autonomous driving vehicles with high stability and recognition accuracy, good performance of the model in the required application scenarios, improved robustness, and the use of a small amount of dataset in the application scenarios to achieve high recognition accuracy, while reducing the cost of dataset production.
[0004] A lane line recognition method for autonomous driving vehicles according to the present invention is characterized by including the following steps:
[0005] 1) Obtain the dataset:
[0006] Take a large number of public datasets as source domain data, and collect the datasets of the algorithm deployment application scenarios as target domain datasets;
[0007] 2) Data preprocessing:
[0008] Process the collected target domain data into the source domain data format, and perform lane line annotation on the target domain data.
[0009] 3) Build a training model:
[0010] First, select the deep learning network Unet as the basic model, and add the attention mechanism Attention module to jointly construct the Attention-Unet model as the training model;
[0011] The Attention Mechanism model consists of a multi-layer perceptron model structure composed of a query matrix (Query), a key, and a weighted average. First, the processed image X passes through the attention mechanism unit. X is first added to the gate vector gi of the attention mechanism, passed through σ1, i.e., the ReLU activation function, multiplied by the weights of the attention mechanism, and then added to the bias, that is Subsequently, the attention coefficient α is obtained through σ2, i.e., the Sigmoid activation function, that is Then the output α of the attention mechanism is obtained i *X;
[0012] After adding the Attention module, the extended path of the Unet neural network is jointly obtained by the contraction path, the value of the previous path, and the attention mechanism. The update formula for its next layer is:
[0013]
[0014] 4) Transfer learning:
[0015] Perform H-space feature mapping on the source domain dataset and the target domain dataset. In the mapping space, use the Maximum Mean Difference (MMD) algorithm to define the data difference between the two domains and minimize the difference between the two datasets;
[0016] 5) Model training:
[0017] Send the constructed dataset into the Attention-Unet neural network for training. After adding transfer learning, the loss of the neural network is updated as:
[0018] Loss = MSELoss + λMMD(D S , D T )
[0019] By backpropagating this loss, continuously reduce MSELoss and MMD to reach the minimum value of model fitting and the gap between the source domain and the target domain, and finally obtain a lane line recognition model suitable for the required scenario.
[0020] For the Maximum Mean Difference, MMD algorithm, the domain adaptation metric used is the maximum mean difference, which is used to measure the difference between the data distributions of the two domains. The data samples of the two domains are mapped to a new feature space H, called the Reproducing Kernel Hilbert Space (RKHS), through a kernel function, and then the Euclidean distance is used to calculate the feature distance between the two domains in the new feature space.
[0021] Compared with the prior art, the present invention has obvious beneficial effects. From the above technical solutions, it can be seen that the present invention selects the deep learning network Attention-Unet as the training model, takes a large number of public data sets as the source domain data, collects the data sets of the application scenarios arranged by a certain algorithm as the target domain data sets, and adopts the feature-based transfer learning method to project the features of the source domain and the target domain into a specific space, obtain the feature representations of the two domains in the space and measure the distribution difference between them. Finally, minimize the distribution difference between the two data, encode the instances in the space and send them into the model for training, to make up for the problem that the target domain data set is too small to establish a lane line recognition model with high fitting degree and accuracy, so as to ensure the accuracy of model recognition, enhance the stability of the system. At the same time, there is no need to make a large number of data sets in the application scenarios, reducing the cost of the solution. There is no need to collect too many data sets of the required application scenarios, and there is no need to perform a large amount of data set annotation work, thus further saving costs. It improves the robustness and generalization ability of the model, can be applied to a large number of lane line recognition scenarios, not limited to specific scenarios. It realizes high stability and recognition accuracy, the model has good performance ability in the required application scenarios, improves the robustness, and uses a small amount of data sets in the application scenarios to achieve high recognition accuracy, while reducing the cost of data set production. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the flowchart of the present invention;
[0023] Figure 2 is the structural diagram of the attention mechanism of the present invention;
[0024] Figure 3 is the symmetric encoding and decoding structural diagram based on the fully convolutional network of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Embodiment 1:
[0026] A method for identifying lane lines of an autonomous driving vehicle according to the present invention is characterized in that it includes the following steps:
[0027] 1) Obtain data sets:
[0028] Take a large number of public data sets as the source domain data, and collect the data sets of the application scenarios arranged by the algorithm as the target domain data sets;
[0029] 2) Data preprocessing:
[0030] Process the collected target domain data into the source domain data format, and perform lane line annotation on the target domain data.
[0031] 3) Build a training model;
[0032] First, select the deep learning network Unet as the basic model, and add the Attention mechanism module to jointly construct the Attention-Unet model as the training model;
[0033] The Attention Mechanism is a data processing method in machine learning, and its structure is as follows Figure 2 As shown, where; Input xi xi: This is the input feature map, usually a multi-dimensional tensor, representing the input data of the model at a certain time step or a certain position;
[0034] gate gi: This is a gating signal used to control the activation degree of the attention mechanism. It can be a signal from other parts of the model or a feature related to the input;
[0035] W i : 1×1×1 This is a 1×1 convolutional layer used to perform a linear transformation on the input feature map xixi. This type of convolutional layer is usually used for feature fusion between channels without changing the spatial dimension;
[0036] W g : 1×1×1 This is a similar 1×1 convolutional layer used to perform a linear transformation on the gating signal gigi;
[0037] The plus sign in the figure indicates matrix addition of the feature maps after transformation by Wi and Wg. This step combines the information of the input feature and the gating signal;
[0038] ReLU: Rectified Linear Unit activation function, used to introduce non-linearity. The ReLU function is defined as f(x) = max(0, x), which can enhance the expressive ability of the model;
[0039] Another ψ: 1×1×11×1 convolutional layer, used to further process the feature map after ReLU activation;
[0040] Sigmoid: The Sigmoid function maps the values of the feature map to between (0, 1)(0, 1) to generate an attention weight map;
[0041] The Sigmoid function is defined as Its output can be interpreted as the attention weight at each position;
[0042] Resampler: Spatial resampling to align the feature dimensions. This step may be used to adjust the size of the attention weight map to match the size of the input feature map xi. This step is necessary in some cases, especially when the sizes of the input and output feature maps are inconsistent;
[0043] The multiplication symbol in the figure represents matrix multiplication of the attention weight map after being processed by the Resampler and the input feature map \(x_i\). This step applies the attention weights to the input features, enabling the model to focus on the important parts of the input;
[0044] Output \(x_o\) o : This is the output of the attention unit, representing the feature map after being processed by the attention mechanism. The output feature map \(x_o\) o retains the important information in the input feature map \(x_i\) and suppresses the unimportant parts.
[0045] Figure 3 It is divided into two paths. The downward path is the encoding process of the Unet encoder, and the upward path is the decoding process of the Unet decoder.
[0046] Among them, \(F\) n represents the \(n\)-th layer of convolution, \(H\) n and \(W\) n represent the data height and width of the \(n\)-th layer, and \(D\) n represents the number of data channels of the \(n\)-th layer. \(H\), \(W\), and \(D\) together describe the spatial resolution of the input data.
[0047] Encoding process:
[0048] After the data enters the Unet neural network, it first passes through 2 layers of 3x3 convolution and the ReLU activation function to extract local features of the data through convolution;
[0049] Subsequently, it passes through a 2x2 max pooling layer to focus on global semantics, and the spatial resolution \((H, W)\) is halved while the depth \((D)\) gradually increases.
[0050] Through the above two steps, one encoding is completed, and the entire encoding process undergoes 3 times of the above encoding operations;
[0051] Decoding process:
[0052] At the bottom, the data first passes through 2 layers of 3x3 convolution and the ReLU activation function to extract local features of the data through convolution;
[0053] Subsequently, it passes through a 2x2 upsampling layer with the same matrix size as the max pooling layer to restore the data spatial resolution;
[0054] Finally, the data is connected to the deep features extracted by the attention mechanism through the Concatenation layer. Among them, the attention unit is to improve the semantic segmentation accuracy, and the role of the Concatenation layer is to supplement detailed information;
[0055] One decoding is completed through the above three steps. The entire decoding process corresponds to the encoding process and undergoes 3 times of the above decoding operations.
[0056] Output layer:
[0057] Finally, features are extracted through a 2-layer 3x3 convolution and ReLU activation function to obtain the output of the Unet neural network. Among them, N c is the data category of the Unet output, and c is the number of categories.
[0058] The attention mechanism is widely used in various types of machine learning tasks such as natural language processing, image recognition, and speech recognition. The degree of attention (importance) of the attention mechanism to different information is reflected by weights. The attention mechanism can be regarded as a multi-layer perceptron composed of a query matrix (Query), a key (key), and a weighted average. Its formula is:
[0059]
[0060] First, the processed image X will pass through the attention mechanism unit. X is first added to the gate vector gi of the attention mechanism, passed through σ1, that is, the ReLU activation function, and then multiplied by the weights of the attention mechanism and added with the bias, that is Subsequently, the attention coefficient α is obtained through σ2, that is, the Sigmoid activation function, that is Then the output of the attention mechanism α is obtained i *X;
[0061] U-Net is one of the earlier algorithms for semantic segmentation using fully convolutional networks. It is an image segmentation network with a symmetric encoder-decoder structure based on a fully convolutional network, which extracts and restores image features through the downsampling and upsampling processes. The UNet network consists of two parts, the contracting path and the expanding path;
[0062] After adding the Attention module, the expanding path of the Unet neural network is jointly obtained by the contracting path, the value of the previous path, and the attention mechanism. The update formula for its next layer is:
[0063]
[0064] 4) Transfer learning;
[0065] The source domain dataset and the target domain dataset are subjected to H-space feature mapping. In the mapping space, the maximum mean difference is adopted, and the domain adaptation metric method used is the maximum mean difference (MMD), which is used to measure the difference between the data distributions of the two domains. MMD first maps the data samples of the two domains to a new feature space H through a kernel function, which is called the Reproducing Kernel Hilbert Space (RKHS). Then, the Euclidean distance is used to calculate the feature distance between the two domains in the new feature space. The formula is expressed as follows:
[0066]
[0067] Among them, and represent the datasets of the source domain and the target domain respectively, m and n represent the total amounts of data in the source domain and the target domain respectively, φ(·) is the way to map the data samples to the RKHS feature space, is the norm of the RKHS. The above formula is expanded in the form of matrix multiplication as follows:
[0068]
[0069] Finally, the above formula is transformed by using the kernel function k(u,v) = φ(u)φ(v) to obtain the following formula:
[0070]
[0071] 4) Model training;
[0072] The constructed dataset is sent into the Attention-Unet neural network for training. After the addition of transfer learning, the loss of the neural network is updated as:
[0073] Loss = MSELoss + λMMD(D S ,D T )
[0074] By backpropagating this loss, continuously reducing MSELoss and MMD, minimizing this loss to narrow the distance between the source domain features and the target domain features, and at the same time making the model reach fitting, finally obtaining a lane line recognition model applicable to the required scenario.
[0075] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A lane line recognition method for an autonomous vehicle, characterized in that; It includes the following steps: 1) Obtain the dataset: Take a large number of public datasets as the source domain data, and collect the datasets of the application scenarios of the acquisition algorithm as the target domain datasets; 2) Data preprocessing: Process the collected target domain data into the source domain data format, and perform lane line annotation on the target domain data; 3) Build a training model: First, select the deep learning network Unet as the basic model, and add the attention mechanism Attention module to jointly construct the Attention-Unet model as the training model; The Attention Mechanism model consists of a multi-layer perceptron model structure composed of a query matrix (Query), a key, and a weighted average; first, the processed image X passes through the attention mechanism unit. X is first added to the gate vector gi of the attention mechanism, passed through σ1, i.e., the ReLU activation function, multiplied by the weights of the attention mechanism, and then added to the bias, that is , and then the attention coefficient α is obtained through σ2, i.e., the Sigmoid activation function, that is , then the output α of the attention mechanism is obtained i *X; After adding the Attention module, the extended path of the Unet neural network is jointly obtained by the contraction path, the value of the previous path, and the attention mechanism. The update formula for its next layer is: 4) Transfer learning; Perform H-space feature mapping on the source domain dataset and the target domain dataset. In the mapping space, use the Maximum Mean Difference (MMD) algorithm to define the data differences between the two domains, and minimize the differences between the two datasets; 5) Model training; Send the constructed dataset into the Attention-Unet neural network for training. After adding transfer learning, the loss of the neural network is updated as: By backpropagating this loss, continuously reduce the MSELoss and MMD, reach the minimum value of the model fitting and the gap between the source domain and the target domain, and finally obtain a lane line recognition model suitable for the required scenario.
2. The lane line recognition method for an autonomous vehicle according to claim 1, wherein; The Maximum Mean Difference, MMD algorithm uses the domain adaptation metric of maximum mean difference to measure the difference between the data distributions of the two domains. The data samples of the two domains are mapped to a new feature space H, called the Reproducing Kernel Hilbert Space (RKHS), through a kernel function, and then the Euclidean distance is used to calculate the feature distance between the two domains in the new feature space.