Lane line prediction model training method and device and computer equipment

By acquiring the lane line images of continuous frames for feature fusion and difference determination, the problems of low training efficiency and insufficient performance of the existing lane line prediction model are solved, and efficient and accurate lane line detection is achieved.

CN120279516AActive Publication Date: 2025-07-08XIAN OUYE SEMICONDUCTOR CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510350148.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing lane line prediction model requires manual labeling of the truth value of each sequence, which is time-consuming and labor-intensive, and does not fully utilize inter-frame correlation information, resulting in low training efficiency and insufficient performance.

Method used

By acquiring the lane line images of continuous frames, extracting the initial feature map and performing feature fusion, using the homography matrix to simulate vehicle movement, determining the difference in marked and predicted lane lines, model training is performed based on the difference, and a lane line prediction model is generated.

Benefits of technology

It improves the accuracy and robustness of lane line detection, reduces manual labeling workload, improves training efficiency and saves costs, while maintaining high detection accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279516A_ABST
    Figure CN120279516A_ABST
Patent Text Reader

Abstract

The invention relates to a lane line prediction model training method and device and computer equipment. The method comprises the following steps: obtaining a first lane line image and a second lane line image, wherein the second lane line image is a previous frame image of the first lane line image; extracting an initial feature map from the first lane line image, wherein the first lane line image and the second lane line image both comprise marked lane lines; performing feature fusion on the initial feature map of the first lane line image and the marked lane line in the second lane line image to obtain a fused feature map; image generation is carried out based on the fusion feature map, a lane line prediction image is obtained, and the lane line prediction image comprises a prediction lane line; and determining the difference between the marked lane line and the predicted lane line, and performing model training based on the difference to obtain a lane line prediction model. The method can improve the training efficiency and effect and save the training cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a method and apparatus for training a lane line prediction model, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Lane lines are the lines marked on the road surface and are used to guide drivers to maintain the correct driving route. For autonomous vehicles, lane line prediction helps the vehicle identify the road boundaries and lane positions, thereby realizing functions such as automatic steering and path planning. This not only improves the driving comfort but also enhances the driving safety.

[0003] However, the current lane line prediction models usually require manual annotation of the ground truth (i.e., the true lane line positions) for each sequence before feature extraction and fusion can be performed. This method is not only time-consuming and laborious, but also, since the lane lines are static, the same information may appear repeatedly between multiple frames, further increasing the workload. In addition, the existing methods have deficiencies in dealing with the inter-frame motion relationships and do not make full use of the inter-frame correlation information to improve the model performance. Summary of the Invention

[0004] Based on this, it is necessary to provide a method and apparatus for training a lane line prediction model, a computer device, a computer-readable storage medium, and a computer program product that can improve the training efficiency and effect and save the training cost for the above technical problems.

[0005] In a first aspect, the present application provides a method for training a lane line prediction model, including:

[0006] Obtaining a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image;

[0007] Extracting an initial feature map from the first lane line image, where both the first lane line image and the second lane line image include annotated lane lines;

[0008] Performing feature fusion on the initial feature map of the first lane line image and the annotated lane lines in the second lane line image to obtain a fused feature map;

[0009] Generating an image based on the fused feature map to obtain a lane line prediction image, where the lane line prediction image includes predicted lane lines;

[0010] Determining the difference between the annotated lane lines and the predicted lane lines, and training a model based on the difference to obtain a lane line prediction model.

[0011] In one embodiment, obtaining the first lane line image and the second lane line image includes:

[0012] Obtaining an initial frame lane line image, where the initial frame lane line image includes labeled lane lines; determining a homography transformation matrix based on the annotation content of the labeled lane lines; transforming the labeled lane lines in the initial frame lane line image through the homography transformation matrix to obtain a sequence of lane line images transformed by simulating vehicle movement, where the sequence of lane line images includes a plurality of consecutive frame lane line images after the initial frame; selecting the first lane line image and the second lane line image from the plurality of consecutive frame lane line images.

[0013] In one embodiment, the labeled lane lines include a labeled angle and a plurality of labeled points; the predicted lane lines include a predicted angle and a plurality of predicted points;

[0014] Determining the difference between the labeled lane lines and the predicted lane lines includes:

[0015] Based on the positions of the plurality of labeled points in the first lane line image and the positions of the plurality of predicted points in the lane line prediction image, determining the position difference between the labeled lane lines and the predicted lane lines; based on the labeled angle and the predicted angle, determining the angle difference between the labeled lane lines and the predicted lane lines.

[0016] In one embodiment, based on the positions of the plurality of labeled points in the first lane line image and the positions of the plurality of predicted points in the lane line prediction image, determining the position difference between the labeled lane lines and the predicted lane lines includes:

[0017] For each labeled point in the first lane line image, searching for a target predicted point that matches the labeled point among the plurality of predicted points in the lane line prediction image; determining the horizontal distance in the horizontal direction and the vertical distance in the vertical direction between the labeled point and the target predicted point; based on the horizontal distance and the vertical distance, determining the Euclidean distance between the labeled point and the target predicted point; using the Euclidean distance between the labeled point and the target predicted point as the position difference between the labeled lane lines and the predicted lane lines.

[0018] In one embodiment, training a lane line prediction model based on the difference includes:

[0019] Determine a target annotation point among the multiple annotation points based on the Euclidean distance between the targeted annotation point and the target prediction point and a preset distance threshold; adjust the parameters of the model based on the number of the target annotation points until the number of the target annotation points is less than a preset annotation point value, so as to obtain a lane line prediction model.

[0020] In one embodiment, feature fusion of the initial feature map of the first lane line image and the annotated lane lines in the second lane line image is performed to obtain a fusion feature map, including:

[0021] Perform a convolution operation on the initial feature map of the first lane line image by using a first convolutional kernel to obtain a first feature map; perform a convolution operation on the annotated lane lines in the second lane line image by using a second convolutional kernel to obtain a second feature map; perform a linear combination on the first feature map and the second feature map to obtain a combined feature map; compress the feature values of the combined feature map through a prediction activation function to obtain an output feature map, where the feature values in the output feature map are used to represent the information retention ratio at each position of the annotated lane lines in the second lane line image; generate a fusion feature map based on the output feature map, the initial feature map, and the annotated lane lines in the second lane line image.

[0022] In a second aspect, the present application further provides a lane line prediction model training device, including:

[0023] An obtaining module, configured to obtain a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image;

[0024] An extraction module, configured to extract an initial feature map from the first lane line image, where both the first lane line image and the second lane line image include annotated lane lines;

[0025] A fusion module, configured to perform feature fusion on the initial feature map of the first lane line image and the annotated lane lines in the second lane line image to obtain a fusion feature map;

[0026] A generation module, configured to perform image generation based on the fusion feature map to obtain a lane line prediction image, where the lane line prediction image includes predicted lane lines;

[0027] A training module, configured to determine the difference between the annotated lane lines and the predicted lane lines, and perform model training based on the difference to obtain a lane line prediction model.

[0028] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0029] Obtain a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image;

[0030] Extract an initial feature map from the first lane line image. Both the first lane line image and the second lane line image include labeled lane lines;

[0031] Perform feature fusion on the initial feature map of the first lane line image and the labeled lane lines in the second lane line image to obtain a fused feature map;

[0032] Perform image generation based on the fused feature map to obtain a lane line prediction image, where the lane line prediction image includes predicted lane lines;

[0033] Determine the difference between the labeled lane lines and the predicted lane lines, and perform model training based on the difference to obtain a lane line prediction model.

[0034] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0035] Obtain a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image;

[0036] Extract an initial feature map from the first lane line image. Both the first lane line image and the second lane line image include labeled lane lines;

[0037] Perform feature fusion on the initial feature map of the first lane line image and the labeled lane lines in the second lane line image to obtain a fused feature map;

[0038] Perform image generation based on the fused feature map to obtain a lane line prediction image, where the lane line prediction image includes predicted lane lines;

[0039] Determine the difference between the labeled lane lines and the predicted lane lines, and perform model training based on the difference to obtain a lane line prediction model.

[0040] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0041] Obtain a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image;

[0042] Extract an initial feature map from the first lane line image. Both the first lane line image and the second lane line image include labeled lane lines.

[0043] Fuse the initial feature map of the first lane line image and the labeled lane lines in the second lane line image to obtain a fused feature map.

[0044] Generate an image based on the fused feature map to obtain a lane line prediction image, which includes predicted lane lines.

[0045] Determine the difference between the labeled lane lines and the predicted lane lines, and perform model training based on the difference to obtain a lane line prediction model.

[0046] The above-mentioned lane line prediction model training method, device, computer device, computer-readable storage medium, and computer program product obtain a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image. By acquiring the image data of consecutive frames (i.e., the previous frame and the current frame), it provides a basis for subsequent feature extraction and fusion. Utilizing the information in the time dimension can better capture the changes during the vehicle movement, thereby improving the accuracy and robustness of lane line detection.

[0047] Extract an initial feature map from the first lane line image. Both the first lane line image and the second lane line image include labeled lane lines. By extracting features from a single-frame image, the key features of the lane lines can be identified, which is the basis for subsequent feature fusion and helps to more accurately understand the position of the lane lines in the three-dimensional space.

[0048] Fuse the initial feature map of the first lane line image and the labeled lane lines in the second lane line image to obtain a fused feature map. Combining the features of two frames of images enables the model to not only consider the information within a single frame but also incorporate the changes in the time series. The advantage of this is that it can effectively overcome the occlusion or interference problems that may occur during single-frame analysis and can fully utilize the correlation information between frames to improve the model performance.

[0049] Generate an image based on the fused feature map to obtain a lane line prediction image, which includes predicted lane lines. Generating a prediction image through the fused feature map can, to a certain extent, simulate the real lane line distribution. This method is more flexible and efficient compared to traditional rule-based post-processing algorithms and can maintain a high detection accuracy especially in complex environments.

[0050] Determine the difference between the labeled lane lines and the predicted lane lines, and train the model based on the difference to obtain a lane line prediction model. By comparing the difference between the prediction result and the actual ground truth and adjusting the model parameters accordingly, the performance of the model can be continuously optimized. In addition, the time-consuming and laborious manual annotation process in traditional methods is avoided, greatly reducing the cost and improving the efficiency at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.

[0052] Figure 1 FIG. is an application environment diagram of the lane line prediction model training method in an embodiment;

[0053] Figure 2 FIG. is a schematic flowchart of the lane line prediction model training method in an embodiment;

[0054] Figure 3 FIG. is a schematic plan view of the initial frame lane line image in an embodiment;

[0055] Figure 4 FIG. is an architecture diagram of the lane line prediction model training in an embodiment;

[0056] Figure 5 FIG. is a structural block diagram of the lane line prediction model training device in an embodiment;

[0057] Figure 6 FIG. is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0059] The lane line prediction model training method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through a model. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other model servers. The terminal 102 is used to generate a lane line prediction model training request and send the lane line prediction model training request to the server 104, so that the server 104 determines the difference between the labeled lane lines and the predicted lane lines, and performs model training based on the difference to obtain the lane line prediction model. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0060] In an exemplary embodiment, as Figure 2 shown, a lane line prediction model training method is provided. Taking the method applied to Figure 1 the server 104 in

[0061] Step 202, obtain a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image.

[0062] Among them, the first lane line image refers to the frame image currently being processed. In this application, it refers to the frame image used to extract the initial feature map. This frame image contains road information, including but not limited to the position and shape of the lane lines.

[0063] The second lane line image refers to the frame image immediately preceding the first lane line image, that is, the previous frame image. This frame image also contains road information and has been labeled with lane lines.

[0064] Specifically, first, manual annotation is performed on a single-frame image. By connecting points, a complete annotated lane line is obtained. A continuous homography transformation is performed on the single-frame image to simulate the perspective change effect caused by the vehicle's movement at different time points, thereby generating a set of continuous time-series images. The simulated time-series images include lane line images of multiple consecutive frames. It can be understood that the lane line images of multiple consecutive frames in the time-series images also include annotated lane lines.

[0065] Then, in the lane line images of multiple consecutive frames in the time-series images, the latter frame of two consecutive frames is selected as the "first lane line image", which represents the latest road scene to be analyzed. This image contains the latest road information and possible lane line annotations. The former frame of the above two frames is defined as the "second lane line image". This image represents the road condition at the time point immediately before the "first lane line image" and also contains corresponding lane line annotations.

[0066] Step 204, extract an initial feature map from the first lane line image. Both the first lane line image and the second lane line image include annotated lane lines.

[0067] Among them, the initial feature map refers to the feature representation extracted from the first lane line image (current frame).

[0068] The annotated lane line refers to the lane line that has been manually or automatically algorithmically annotated in the second lane line image (previous frame). The annotation content in the annotated lane line includes but is not limited to the starting point, ending point, angle of the annotated lane line, and the distance between key points, etc. It can be understood that these annotation contents are used to fuse with the feature map extracted from the current frame (first lane line image) to enhance the accuracy of lane line detection in the current frame using the information of the historical frame. For example, when performing feature fusion, the displacement difference of the lane line between two adjacent frames is considered as part of the supervision signal to help optimize the model training process.

[0069] Specifically, first, the lane line images of multiple consecutive frames in the time-series images are respectively sent to the feature extraction unit for feature extraction to obtain the initial feature map corresponding to each lane line image. Then, the initial feature maps corresponding to each lane line image are all sent to the data pool for storage. Then, further, the initial feature map corresponding to the first lane line image is obtained from the data pool. It can be understood that the data pool, as an intermediate storage area, stores the feature information extracted from the images at each time point. The feature extraction unit processes the input lane line image using a deep learning model to identify and extract the key features that help identify the lane line. These features can be edges, textures, shapes, etc., which together constitute the description of the position and shape of the lane line.

[0070] Step 206: Feature fusion is performed on the initial feature map of the first lane line image and the labeled lane lines in the second lane line image to obtain a fused feature map.

[0071] The fused feature map refers to a new feature representation generated by combining the initial feature map of the current frame (the first lane line image) with the labeled lane line information in the previous frame (the second lane line image). Understandably, this feature fusion process aims to enhance the accuracy and robustness of lane line detection by leveraging information in the time dimension.

[0072] Specifically, the most recent feature extraction result from the data pool (i.e., the initial feature map of the first lane line image) and the detection result of the previous frame (i.e., the labeled lane lines in the second lane line image) are retrieved and fed into a convolutional gated recurrent unit for feature fusion. Understandably, the convolutional gated recurrent unit is a type of recurrent neural network structure specifically designed to process sequential data (time series images in this application). During the processing, the convolutional gated recurrent unit combines the features of the current frame with the detection result of the previous frame to perform a feature fusion operation and obtain a new fused feature map.

[0073] Step 208: Image generation is performed based on the fused feature map to obtain a lane line prediction image, which includes predicted lane lines.

[0074] The lane line prediction image refers to an image generated based on the fused feature map. In this lane line prediction image, the lane line prediction model predicts the positions of the lane lines in the current frame according to the learned patterns and rules. In other words, it is a new image inferred by the lane line prediction model based on the input data (i.e., the fused feature map obtained by fusing the initial feature map of the first lane line image and the labeled lane line information in the second lane line image). This image aims to accurately reflect the positions of the lane lines in the actual scene as much as possible.

[0075] The predicted lane lines refer to the positions and shapes of the lane lines identified in the "lane line prediction image". It is the result obtained by the lane line prediction model after analyzing the input data and represents the best guess of the lane line prediction model for the positions of the lane lines in the current frame. The predicted lane lines usually appear as lines or curves in the lane line prediction image, indicating the path that the vehicle should follow.

[0076] Specifically, the fused feature map obtained in the above steps is used as the input and fed into a pre-trained deep learning model, namely the lane line prediction model. This lane line prediction model has been fully trained and can predict the positions and shapes of the lane lines based on the input feature map.

[0077] Based on the fused feature map, a lane line prediction model is used to generate a new image containing the predicted lane lines, i.e., the lane line prediction image. In this process, the lane line prediction model analyzes the information in the fused feature map and marks the parts in the image that it believes are most likely to be lane lines.

[0078] In the generated lane line prediction image, the predicted lane lines appear in the form of lines or curves, representing the positions of the lane lines inferred by the lane line prediction model. These predicted lane lines reflect the best estimate of the lane line prediction model for the position of the lane lines in the current frame.

[0079] Step 210, determine the difference between the labeled lane lines and the predicted lane lines, and based on the difference, perform model training to obtain the lane line prediction model.

[0080] Among them, the difference refers to the gap between the labeled lane lines (i.e., the true or expected lane line positions, determined based on manual annotation or other precise methods) and the lane lines predicted by the model (i.e., the estimated lane line positions obtained after processing the image by the current model). Optionally, this gap is quantified by a defined loss function, and the loss function measures the error between the model output and the actual target.

[0081] The lane line prediction model refers to a deep learning model that can identify and predict the positions of lane lines from the input lane line images after training. The lane line prediction model uses feature extraction technology to obtain an initial feature map from the first lane line image and fuses it with the labeled lane line information in the second lane line image to generate a fused feature map. Based on this fused feature map, the lane line prediction model can generate a lane line prediction image, which contains the positions of the lane lines inferred by the lane line prediction model. By continuously comparing the lane line prediction results with the true lane line positions and optimizing the parameters of the lane line prediction model based on the difference, a more accurate and robust lane line prediction model is finally obtained. It can be understood that the lane line prediction model is the core of the entire temporal supervision modeling method, and its performance directly affects the accuracy and stability of the lane line detection in the autonomous driving system.

[0082] Specifically, refer to Figure 3, first, compare the predicted lane lines in the lane line prediction image with the corresponding annotated lane lines. Use a pre - defined loss function to quantify the difference between the predicted lane lines and the annotated lane lines. Understandably, the loss function is a metric used to measure the error between the output of the lane line prediction model and the real situation. The differences include, but are not limited to, the position deviation of the lane lines, shape mismatches, etc. For example, calculate the Euclidean distance between the key points of the predicted lane lines (such as the starting point, ending point, angle, key points, where the key points are the points on the line connecting the starting point and the ending point) and the corresponding points of the annotated lane lines as part of the difference.

[0083] Then, according to the difference (loss value) calculated in the previous step, use the backpropagation algorithm to update the parameters of the lane line prediction model. The purpose of this step is to minimize the loss function so that the lane line prediction model can more accurately predict the position of the lane lines. Understandably, this process involves multiple iterations, and each iteration will use a new data set or different batches of data from the same data set to optimize the model parameters in the lane line prediction model. To further improve the model performance, the above process needs to be continuously repeated, that is, starting from obtaining a new batch of two consecutive frame lane line images from the time - series images, through feature extraction, feature fusion, lane line prediction until calculating the difference and adjusting the model parameters of the lane line prediction model.

[0084] As the number of training rounds increases, the lane line prediction model gradually learns how to more accurately identify lane lines in different scenarios, and can maintain a high detection accuracy and robustness even in the presence of occlusion or interference. After multiple rounds of training, when the lane line prediction model can achieve satisfactory performance metrics on the validation set or test set, it can be considered that the training is completed, and the final lane line prediction model is obtained. This lane line prediction model can be used to process the road images captured by the vehicle's front - facing camera in real - time and accurately predict the position of the lane lines, providing important support for the autonomous driving system.

[0085] In one embodiment, obtain an initial frame lane line image, where the initial frame lane line image includes annotated lane lines; determine a homography transformation matrix based on the annotation content of the annotated lane lines; transform the annotated lane lines in the initial frame lane line image through the homography transformation matrix to obtain a sequence of lane line images that are transformed by simulating vehicle movement, and the sequence of lane line images includes multiple consecutive frame lane line images after the initial frame; select a first lane line image and a second lane line image from the multiple consecutive frame lane line images.

[0086] Among them, the initial frame lane line image refers to the first road image containing lane lines obtained from the vehicle camera. Understandably, the initial frame lane line image will serve as the basis for subsequent transformations.

[0087] The labeled lane lines refer to the positions of the lane lines marked in the lane line image of the initial frame. Optionally, this labeling is done manually to indicate information such as the exact position and shape of the lane lines.

[0088] The annotation content refers to various information recorded for the labeled lane lines, such as the starting point, ending point, angle, and distances between key points, etc. Understandably, these detailed information are crucial for determining how to apply the homography transformation matrix.

[0089] The homography transformation matrix, as a mathematical tool, is used to simulate the perspective change effect caused by vehicle movement. Through this homography transformation matrix, the lane lines in the initial frame can be transformed into lane lines in a series of consecutive frames, thus generating a series of images that seemingly capture the vehicle while it is moving.

[0090] The lane line sequence images refer to a set of consecutive time - series images obtained by transforming the labeled lane lines in the initial frame lane line image using the homography transformation matrix. This set of lane line sequence images simulates the lane line situations seen from different perspectives due to vehicle movement at different time points, including multiple consecutive frame lane line images after the initial frame.

[0091] Specifically, first, a road image is obtained from a vehicle front - facing camera or other sensors as the initial frame. Optionally, the initial frame should contain clearly visible and labeled lane lines. The annotation content includes, but is not limited to, information such as the starting point, ending point, angle, and distances between key points.

[0092] Then, according to the specific annotation content (such as the starting point, ending point, angle, etc.) of the labeled lane lines in the initial frame, a homography transformation matrix for simulating the vehicle movement effect is calculated. The homography transformation matrix will be used to generate a series of consecutive frames that can mimic the different lane line positions seen by the vehicle during driving due to perspective changes.

[0093] Next, the calculated homography transformation matrix is used to transform the labeled lane lines in the initial frame to generate a set of consecutive time - series images. These generated images represent the lane line views at different time points when the vehicle is moving, that is, the lane line sequence images. Each sequence image contains the position change of the lane lines, reflecting the dynamic process of the vehicle driving on the road.

[0094] Finally, two adjacent frames are selected from the generated multiple consecutive frame lane line images as the first lane line image (the current frame) and the second lane line image (the previous frame). Understandably, these two images will be used for subsequent feature extraction and fusion processing to train the lane line prediction model.

[0095] For the purpose of automatically generating a series of consecutive frames from a single initial frame using mathematical transformation techniques and selecting specific frames for further analysis. This method not only reduces the workload of manual annotation but also improves the diversity and accuracy of the training data for the lane line detection model.

[0096] In one embodiment, based on the positions of multiple annotation points in the first lane line image and the positions of multiple prediction points in the lane line prediction image, determine the position difference between the annotated lane line and the predicted lane line; based on the annotation angle and the prediction angle, determine the angle difference between the annotated lane line and the predicted lane line.

[0097] Herein, the annotation point refers to the specific position point manually or automatically marked in the initial frame lane line image. Optionally, these annotation points are distributed along the actual lane line, including but not limited to the starting point, ending point, and key feature points of the lane line, etc., for accurately describing the position and shape of the lane line.

[0098] The prediction point refers to the position point on the lane line identified in the lane line prediction image generated by the lane line prediction model..

[0099] The position difference is the spatial distance difference between the annotation point (the position of the real lane line) and the prediction point (the position of the lane line predicted by the lane line prediction model). It can be understood that calculating this difference is to evaluate the accuracy of the lane line prediction model in predicting the position of the lane line. The smaller the position difference, the more accurate the lane line prediction model predicts.

[0100] The annotation angle refers to the angle value determined when annotating the lane line in the initial frame lane line image. This angle can be the angle of the lane line relative to a certain reference direction (such as the vehicle driving direction or the direction of the image coordinate system), used to describe the directionality of the lane line.

[0101] The prediction angle refers to the lane line angle calculated in the lane line prediction image based on the output of the lane line prediction model. It reflects the prediction result of the lane line prediction model on the directionality of the lane line.

[0102] The angle difference is the difference between the annotation angle and the prediction angle. Calculating the angle difference is to measure the accuracy of the lane line prediction model in predicting the direction of the lane line. A smaller angle difference means that the lane line prediction model can better capture the direction characteristics of the lane line.

[0103] Specifically, first analyze the labeled lane lines in the initial frame lane line image. Specifically, it includes: the labeled angle, that is, determining the angle of the lane line relative to a certain reference direction (such as the vehicle driving direction or the direction of the image coordinate system), and multiple labeled points, that is, identifying and recording the position information of multiple key position points on the lane line. These points include but are not limited to the starting point, the ending point, and other feature points.

[0104] Then, in the generated lane line prediction image, analyze the lane lines predicted by the lane line prediction model. Specifically, it includes: the predicted angle, that is, calculating the angle of the lane line according to the prediction result of the lane line prediction model, and multiple predicted points, that is, identifying and recording the position information of the key position points corresponding to the actual labeled points on the predicted lane line.

[0105] Next, compare the position differences between the labeled points and the predicted points. Specifically, for each labeled point, find the corresponding predicted point in the lane line prediction image. And calculate the spatial distance between each labeled point and its corresponding predicted point, that is, the position difference. Optionally, it can be achieved by measurement methods such as the Euclidean distance. Integrate the position differences of all points to evaluate the position accuracy of the entire lane line.

[0106] Finally, compare the differences between the labeled angle and the predicted angle. Specifically, directly compare the labeled angle of the labeled lane line with the predicted angle of the predicted lane line, and calculate the difference between the two to determine the angle difference. It can be understood that this difference reflects the accuracy of the lane line prediction model in predicting the direction of the lane line.

[0107] Since determining the position difference and the angle difference provides specific error metrics for model training. These metrics can be used as part of the loss function to guide the backpropagation algorithm to adjust the model parameters to minimize the gap between the predicted value and the true value, thereby improving the model performance. And by accurately measuring the position and direction differences of the lane lines and improving the model accordingly, it is possible to reduce incorrect decisions caused by inaccurate lane line detection, such as incorrect steering or lane deviation, which is crucial for improving the safety and stability of the autonomous driving system.

[0108] In one embodiment, for each labeled point in the first lane line image, among the multiple predicted points in the lane line prediction image, find the target predicted point that matches the targeted labeled point; determine the horizontal distance in the horizontal direction and the vertical distance in the vertical direction between the targeted labeled point and the target predicted point; based on the horizontal distance and the vertical distance, determine the Euclidean distance between the targeted labeled point and the target predicted point; use the Euclidean distance between the targeted labeled point and the target predicted point as the position difference between the labeled lane line and the predicted lane line

[0109] Among them, the target prediction point refers to the most matching prediction point found in the lane line prediction image for each annotation point in the first lane line image. Optionally, a distance metric or similarity comparison algorithm is used to determine which prediction point is closest to a given annotation point.

[0110] The horizontal direction refers to the X-axis direction in the image coordinate system. In this application, the horizontal direction is used to describe the relative positional relationship in the horizontal direction between the annotation point and the target prediction point.

[0111] The vertical direction refers to the Y-axis direction in the image coordinate system. In this application, the vertical direction is used to describe the relative positional relationship in the vertical direction between the annotation point and the target prediction point.

[0112] The horizontal distance refers to the difference in the X-axis coordinates between a specific annotation point and the corresponding target prediction point in the image coordinate system. That is, the horizontal distance when two points are in the same row but different columns.

[0113] The vertical distance refers to the difference in the Y-axis coordinates between a specific annotation point and the corresponding target prediction point in the image coordinate system. That is, the vertical distance when two points are in the same column but different rows.

[0114] The Euclidean distance refers to the straight-line distance between two points calculated based on the Pythagorean theorem. In this application, the actual spatial distance between the annotation point and its corresponding target prediction point is calculated by combining the horizontal distance and the vertical distance.

[0115] Specifically, first, for each annotation point in the first lane line image, find the most matching target prediction point in the lane line prediction image. Optionally, by comparing the similarity or distance between each annotation point and all prediction points to determine which prediction point is closest to the annotation point. This step is the basis for ensuring the accuracy of subsequent calculations.

[0116] Then, for each pair of found annotation points and target prediction points, calculate their horizontal distance in the horizontal direction (X-axis) and vertical distance in the vertical direction (Y-axis). Specifically, it includes: horizontal distance = |X coordinate of the target prediction point - X coordinate of the annotation point|; vertical distance = |Y coordinate of the target prediction point - Y coordinate of the annotation point|. It can be understood that these distance values reflect the offsets of the two points in their respective coordinate axis directions, providing the necessary parameters for the next calculation of the Euclidean distance.

[0117] Next, use the horizontal distance and vertical distance obtained in the above steps to calculate the actual spatial distance between the annotation point and the target prediction point through the Euclidean distance formula. It can be understood that the Euclidean distance provides an intuitive way to quantify the straight-line distance between two points. It comprehensively considers the deviations in the X-axis and Y-axis directions and can more comprehensively reflect the differences between the two.

[0118] Finally, the Euclidean distance between each pair of labeled points and target prediction points is used as the position difference between the labeled lane line and the predicted lane line at that point. This position difference can be used to evaluate the accuracy of the model in predicting the lane line position. For example, by analyzing the position differences of multiple points, the excellent parts and areas that need improvement of the model can be identified, thereby guiding the further optimization of the model.

[0119] Since using a detailed comparison between multiple labeled points and prediction points, rather than a single overall metric, provides a more comprehensive evaluation of the model performance. This helps to discover local anomalies and ensure that the model can maintain good performance under various different conditions. And taking the Euclidean distance as part of the loss function can be directly used to guide the adjustment of parameters during the model training process, making the training objective more clear, that is, to minimize the spatial distance difference between the actual lane line and the predicted lane line.

[0120] In one embodiment, based on the Euclidean distance between the targeted labeled points and the target prediction points and a preset distance threshold, target labeled points are determined among multiple labeled points; based on the number of target labeled points, the parameters of the model are adjusted until the number of target labeled points is less than the preset labeled point value, and a lane line prediction model is obtained.

[0121] Among them, the preset distance threshold refers to a preset value used to determine whether a labeled point is considered a "target labeled point". Specifically, if the Euclidean distance between a labeled point and its corresponding target prediction point exceeds this threshold, then this labeled point may be regarded as an example where the lane line prediction model fails to accurately predict. This threshold helps to define the acceptable error range of the prediction result of the lane line prediction model.

[0122] Target labeled points refer to those labeled points whose Euclidean distance from their corresponding target prediction points exceeds the preset distance threshold. It can be understood that target labeled points are those position points where the lane line prediction model predicts inaccurately. Identifying these target labeled points helps to locate the areas where the lane line prediction model needs to be improved.

[0123] The preset labeled point value refers to a threshold of the maximum allowed number of mislabeled points preset. During the training process of the lane line prediction model, if the number of target labeled points (i.e., the number of inaccurately predicted points) is less than or equal to this value, then it is considered that the lane line prediction model has reached the expected performance standard. This value is an important metric for measuring the performance of the lane line prediction model and reflects the requirement for the overall prediction accuracy of the lane line prediction model.

[0124] The lane line prediction model refers to the final model after training, which can predict the position and shape of lane lines based on the input image. In this application, "obtaining the lane line prediction model" means continuously adjusting the model parameters to reduce the number of target annotation points until the preset standard is met (i.e., the number of target annotation points is less than the preset annotation point value), thereby obtaining a model with performance meeting the expected requirements.

[0125] Specifically, first, for each annotation point in the first lane line image, calculate the Euclidean distance between it and the corresponding target prediction point. Use a preset distance threshold to filter out those annotation points whose Euclidean distance exceeds the threshold as target annotation points. These target annotation points represent the positions where the lane line prediction model predicts inaccurately. It can be understood that the selection of target annotation points is based on their large deviation from the prediction points, indicating that the lane line prediction model performs poorly at these positions and needs to be improved.

[0126] Then count the number of all determined target annotation points. The purpose is to understand the degree of prediction error of the current lane line prediction model. The number of target annotation points directly reflects the accuracy of the lane line prediction model in the lane line detection task.

[0127] Next, adjust the model parameters according to the number of target annotation points. If the number of target annotation points exceeds the preset annotation point value (i.e., the maximum allowable number of mislabeled points), then the performance needs to be optimized by adjusting the model parameters, including but not limited to updating weights, changing the network architecture, or adjusting hyperparameters, etc. The purpose is to reduce the prediction error of the lane line prediction model so that more annotation points can be correctly predicted.

[0128] Finally, repeat the above steps until the number of target annotation points drops below the preset annotation point value, and an optimized lane line prediction model is obtained. At this time, the lane line prediction model reaches the expected performance level and can accurately predict the position and shape of lane lines.

[0129] Since the Euclidean distance between the annotation points and the target prediction points is calculated and a preset distance threshold is used to filter out the target annotation points, the positions where the model has large prediction deviations can be accurately identified. This method enables developers to improve the model targeted rather than making blind global adjustments. And adjusting the model parameters based on the number of target annotation points until the number of target annotation points is lower than the preset annotation point value ensures that the model is considered trained only when it can accurately predict lane lines in most cases. This can significantly reduce the false prediction rate of the model and improve the overall detection accuracy.

[0130] In one embodiment, a first convolutional kernel is applied to perform a convolution operation on the initial feature map of the first lane line image to obtain a first feature map; a second convolutional kernel is applied to perform a convolution operation on the labeled lane lines in the second lane line image to obtain a second feature map; the first feature map and the second feature map are linearly combined to obtain a combined feature map; the eigenvalues of the combined feature map are compressed through a prediction activation function to obtain an output feature map, and the eigenvalues in the output feature map are used to represent the information retention ratio at each position of the labeled lane lines in the second lane line image; based on the output feature map, the initial feature map, and the labeled lane lines in the second lane line image, a fused feature map is generated.

[0131] Among them, the first convolutional kernel refers to a filter used to perform a convolution operation on the initial feature map of the first lane line image. Through the application of this convolutional kernel, specific feature information can be extracted from the initial feature map to form the first feature map.

[0132] The first feature map refers to the result obtained by applying the first convolutional kernel to the initial feature map of the first lane line image. It contains the feature representations extracted from the original image and helps with the subsequent feature fusion process.

[0133] The second convolutional kernel refers to another filter used to perform a convolution operation on the labeled lane lines in the second lane line image. Its purpose is to extract useful feature information from the labeled lane lines to generate the second feature map.

[0134] The second feature map refers to the result obtained by applying the second convolutional kernel to the labeled lane lines in the second lane line image. It reflects the important feature information of the labeled lane lines.

[0135] The combined feature map refers to the new feature map obtained by linearly combining (such as weighted summation) the first feature map and the second feature map. The purpose of this step is to integrate the feature information from the two images and provide a richer data basis for subsequent processing.

[0136] The prediction activation function, as a non-linear function, is used in neural networks to increase the expressiveness of the model. In this application, the prediction activation function is used to compress or adjust the eigenvalues in the combined feature map to better reflect the information retention ratio at each position.

[0137] The output feature map refers to the combined feature map after being processed by the prediction activation function. The eigenvalues in it represent the information retention ratio at each position in the second lane line image, that is, the degree to which the information at that position is retained during the fusion process.

[0138] The eigenvalue refers to the value of each pixel point in the feature map, representing the feature intensity in the corresponding area or the probability of the existence of a certain feature. In this application, after being processed by the prediction activation function, the eigenvalue can be understood as the importance degree or retention ratio of the information at each position.

[0139] The information retention ratio refers to the ratio of the information at each position being retained relative to the original input (such as the labeled lane lines in the second lane line image) during the fusion process. Understandably, a higher ratio means more original information is retained in the final fused feature map.

[0140] Specifically, first, apply the first convolutional kernel to perform a convolutional operation on the initial feature map of the first lane line image. Through this operation, key feature information is extracted from the original image to obtain the first feature map. And use the second convolutional kernel to perform a convolutional operation on the labeled lane lines in the second lane line image to obtain the second feature map, which is a feature representation extracted based on the labeled lane lines in the second lane line image.

[0141] Then linearly combine (such as weighted summation) the obtained first feature map and the second feature map to generate a combined feature map. This combined feature map integrates the key feature information from the two images, providing a rich data basis for subsequent processing.

[0142] Next, compress or adjust the eigenvalues in the combined feature map through the prediction activation function to obtain the output feature map. After the eigenvalues in the combined feature map are adjusted, they are used to characterize the information retention ratio at each position in the second lane line image, that is, the degree to which the information at that position is retained during the fusion process.

[0143] Finally, based on the obtained output feature map, the initial feature map, and the labeled lane lines in the second lane line image, comprehensively generate the final fused feature map. Complete the feature fusion process. The finally generated fused feature map integrates information from different frames, enhancing the model's detection ability for lane lines and improving accuracy and robustness.

[0144] Since the final fused feature map is generated based on the output feature map, the initial feature map, and the labeled lane lines, this process integrates multi-source information, making the detection result of the lane line more accurate and reliable. Especially for modeling the displacement difference between consecutive frames, it helps to capture dynamic changes and further improves the overall performance of lane line detection. And fusing the information of different frames (i.e., the first lane line image and the second lane line image) can effectively address the problem of insufficient information in a single-frame image caused by occlusion, lighting changes, or other interference factors. This temporal information supplement enhances the robustness and stability of the model in complex environments.

[0145] In one embodiment, refer toFigure 4 , the lane line prediction model training method mainly includes three steps, which are respectively:

[0146] The first step: data generation, mainly including: taking a single-frame image as input, using homography transformation to generate a time series of lane line pictures from the annotated single-frame image, and outputting an image sequence.

[0147] The second step: feature extraction, mainly including: taking the image sequence as input, using a feature extraction unit to extract features from the image sequence to obtain a feature map. And adding the obtained feature map to the data pool.

[0148] The third step: temporal supervision, mainly including: taking the feature map as input, using a convolutional gated recurrent unit to process feature fusion and optimize the temporal supervision process. And adjusting the model parameters through a loss function to minimize the prediction error (i.e., the key point distance), and outputting a trained lane line prediction model.

[0149] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0150] Based on the same inventive concept, the embodiments of the present application also provide a lane line prediction model training device for implementing the above-mentioned lane line prediction model training method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the lane line prediction model training device provided below can refer to the limitations on the lane line prediction model training method in the above text, and will not be repeated here.

[0151] In an exemplary embodiment, as Figure 5 shown, a lane line prediction model training device 500 is provided, including: an acquisition module 502, an extraction module 504, a fusion module 506, a generation module 508, and a training module 510, where:

[0152] The acquisition module 502 is used to acquire a first lane line image and a second lane line image, and the second lane line image is the previous frame image of the first lane line image;

[0153] An extraction module 504, configured to extract an initial feature map from a first lane line image, where both the first lane line image and the second lane line image include labeled lane lines;

[0154] A fusion module 506, configured to perform feature fusion on the initial feature map of the first lane line image and the labeled lane lines in the second lane line image to obtain a fused feature map;

[0155] A generation module 508, configured to perform image generation based on the fused feature map to obtain a lane line prediction image, where the lane line prediction image includes predicted lane lines;

[0156] A training module 510, configured to determine the difference between the labeled lane lines and the predicted lane lines, and perform model training based on the difference to obtain a lane line prediction model.

[0157] In one embodiment, an acquisition module 502 is configured to acquire an initial frame lane line image, where the initial frame lane line image includes labeled lane lines; determine a homography transformation matrix based on the annotation content of the labeled lane lines; perform transformation on the labeled lane lines in the initial frame lane line image through the homography transformation matrix to obtain a lane line sequence image in which the lane lines are transformed by simulating vehicle movement, where the lane line sequence image includes a plurality of consecutive frame lane line images after the initial frame; select a first lane line image and a second lane line image from the plurality of consecutive frame lane line images.

[0158] In one embodiment, the training module 510 is configured to determine the position difference between the labeled lane lines and the predicted lane lines based on the positions of a plurality of labeled points in the first lane line image and the positions of a plurality of predicted points in the lane line prediction image; determine the angle difference between the labeled lane lines and the predicted lane lines based on the labeled angle and the predicted angle.

[0159] In one embodiment, the training module 510 is configured to, for each labeled point in the first lane line image, find a target predicted point that matches the targeted labeled point among the plurality of predicted points in the lane line prediction image; determine the horizontal distance in the horizontal direction and the vertical distance in the vertical direction between the targeted labeled point and the target predicted point; determine the Euclidean distance between the targeted labeled point and the target predicted point based on the horizontal distance and the vertical distance; use the Euclidean distance between the targeted labeled point and the target predicted point as the position difference between the labeled lane lines and the predicted lane lines.

[0160] In one embodiment, the training module 510 is configured to determine a target labeled point among the plurality of labeled points based on the Euclidean distance between the targeted labeled point and the target predicted point and a preset distance threshold; adjust the parameters of the model based on the number of target labeled points until the number of target labeled points is less than a preset labeled point value to obtain a lane line prediction model.

[0161] In one embodiment, the fusion module 506 is configured to perform a convolution operation on the initial feature map of the first lane line image by applying a first convolution kernel to obtain a first feature map; perform a convolution operation on the labeled lane lines in the second lane line image by applying a second convolution kernel to obtain a second feature map; perform a linear combination on the first feature map and the second feature map to obtain a combined feature map; compress the feature values of the combined feature map through a prediction activation function to obtain an output feature map, where the feature values in the output feature map are used to represent the information retention ratio at each position of the labeled lane lines in the second lane line image; and generate a fused feature map based on the output feature map, the initial feature map, and the labeled lane lines in the second lane line image.

[0162] Each module in the above lane line prediction model training device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0163] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to the training of the lane line prediction model. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a model connection. When the computer program is executed by the processor, it implements a lane line prediction model training method.

[0164] Those skilled in the art can understand that Figure 6 the structure shown in

[0165] In one embodiment, a computer device is further provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0166] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0167] In one embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0168] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0169] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0170] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.

[0171] The above embodiments only represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for training a lane line prediction model, characterized in that The method includes: Obtaining a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image; Extracting an initial feature map from the first lane line image, where both the first lane line image and the second lane line image include labeled lane lines; Performing feature fusion on the initial feature map of the first lane line image and the labeled lane lines in the second lane line image to obtain a fused feature map; Performing image generation based on the fused feature map to obtain a lane line prediction image, where the lane line prediction image includes predicted lane lines; Determining the difference between the labeled lane lines and the predicted lane lines, and training a model based on the difference to obtain a lane line prediction model.

2. The method according to claim 1, characterized in that The obtaining of the first lane line image and the second lane line image includes: Obtaining an initial frame lane line image, where the initial frame lane line image includes labeled lane lines; Determining a homography transformation matrix based on the annotation content of the labeled lane lines; Transforming the labeled lane lines in the initial frame lane line image through the homography transformation matrix to obtain a lane line sequence image transformed by simulating vehicle movement, where the lane line sequence image includes a plurality of consecutive frame lane line images after the initial frame; Selecting the first lane line image and the second lane line image from the plurality of consecutive frame lane line images.

3. The method according to claim 1, characterized in that The labeled lane lines include a labeled angle and a plurality of labeled points; the predicted lane lines include a predicted angle and a plurality of predicted points; The determining of the difference between the labeled lane lines and the predicted lane lines includes: Determining the position difference between the labeled lane lines and the predicted lane lines based on the positions of the plurality of labeled points in the first lane line image and the positions of the plurality of predicted points in the lane line prediction image; Determining the angle difference between the labeled lane lines and the predicted lane lines based on the labeled angle and the predicted angle.

4. The method according to claim 3, wherein The determining of the position difference between the labeled lane lines and the predicted lane lines based on the positions of the plurality of labeled points in the first lane line image and the positions of the plurality of predicted points in the lane line prediction image includes: For each labeled point in the first lane line image, searching for a target predicted point that matches the labeled point among the plurality of predicted points in the lane line prediction image; Determining the horizontal distance in the horizontal direction and the vertical distance in the vertical direction between the labeled point and the target predicted point; Determining the Euclidean distance between the labeled point and the target predicted point based on the horizontal distance and the vertical distance; Taking the Euclidean distance between the labeled point and the target predicted point as the position difference between the labeled lane lines and the predicted lane lines.

5. The method according to claim 4, characterized in that, The training of the model based on the difference to obtain a lane line prediction model includes: Determining target labeled points among the plurality of labeled points based on the Euclidean distance between the labeled point and the target predicted point and a preset distance threshold; Adjust the parameters of the model based on the number of the target annotation points until the number of the target annotation points is less than a preset annotation point value, so as to obtain a lane line prediction model.

6. The method according to any one of claims 1 to 5, characterized in that The step of fusing the initial feature map of the first lane line image and the annotated lane lines in the second lane line image to obtain a fused feature map includes: Performing a convolution operation on the initial feature map of the first lane line image by using a first convolutional kernel to obtain a first feature map; Performing a convolution operation on the annotated lane lines in the second lane line image by using a second convolutional kernel to obtain a second feature map; Performing a linear combination on the first feature map and the second feature map to obtain a combined feature map; Compressing the feature values of the combined feature map by using a prediction activation function to obtain an output feature map, where the feature values in the output feature map are used to represent the information retention ratio at each position of the annotated lane lines in the second lane line image; Generating a fused feature map based on the output feature map, the initial feature map, and the annotated lane lines in the second lane line image.

7. A lane line prediction model training device, characterized in that The device includes: An obtaining module, configured to obtain a first lane line image and a second lane line image, where the second lane line image is the previous frame image of the first lane line image; An extracting module, configured to extract an initial feature map from the first lane line image, where both the first lane line image and the second lane line image include annotated lane lines; A fusing module, configured to fuse the initial feature map of the first lane line image and the annotated lane lines in the second lane line image to obtain a fused feature map; A generating module, configured to generate an image based on the fused feature map to obtain a lane line prediction image, where the lane line prediction image includes predicted lane lines; A training module, configured to determine the difference between the annotated lane lines and the predicted lane lines, and perform model training based on the difference to obtain a lane line prediction model.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Angle-adaptive vehicle-mounted aerial view system and implementation method thereof

    CN109367483A

  • Lane line detection method and system based on homography transformation and feature window

    CN110414385A

  • Special scene lane line detection method and device using video time sequence information

    CN116612417A

  • Lane line display method and device, electronic equipment, vehicle and medium

    CN116753978A

  • Three-dimensional lane line detection model training method and device, and three-dimensional lane line detection method and device

    CN117274933A