Lane line marking method, device, equipment, medium and product

By acquiring a set of target road data pairs, a trained neural network model is used to identify point cloud data and image information, generating high-precision lane line labeling results. This solves the problems of low accuracy and low efficiency in existing lane line detection methods, and achieves efficient lane line labeling.

CN121963114APending Publication Date: 2026-05-01HUIZHOU DESAY SV AUTOMOTIVE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUIZHOU DESAY SV AUTOMOTIVE
Filing Date
2024-10-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing lane line detection methods, the lane line ground truth values ​​obtained by simple projection have low accuracy and are not smooth, making them difficult to apply to complex scenarios and requiring manual correction. Furthermore, the lane line ground truth values ​​generated by existing methods are of poor quality and difficult to use in actual production.

Method used

By acquiring a set of target road data pairs, a trained neural network model is used to identify point cloud data and image information to generate lane line labeling results. This process includes acquiring target road point clouds and images, iteratively training the neural network model using a training sample set, and generating target lane line labeling results.

Benefits of technology

It improves the accuracy and efficiency of lane line annotation results by directly recognizing point cloud data and image information based on the trained target model, generating high-precision lane line annotation results, shortening the tool chain, and improving annotation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963114A_ABST
    Figure CN121963114A_ABST
Patent Text Reader

Abstract

The invention discloses a lane line marking method, device and equipment, a medium and a product. The method comprises the steps that a target road data pair set is acquired, and target road data pairs are target road point cloud and target road images under the same timestamp; each target road data pair is input into a target model to obtain a target lane line marking result, the target model is obtained by iteratively training a neural network model through a training sample set, and the training sample set comprises a historical road data pair set and a lane line marking result corresponding to the historical road data pair set. Through the technical scheme of the invention, the accuracy and the marking efficiency of the lane line marking result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of lane line recognition technology, and in particular to a lane line marking method, device, equipment, medium and product. Background Technology

[0002] Lane detection is a key technology in the field of autonomous driving, enabling self-driving cars to identify lane categories and location information, thereby ensuring safe driving within the lane. Currently, mainstream lane detection methods primarily generate ground truth lane values ​​using visual lane detection models to determine the two-dimensional coordinates of each point on the lane line in the image coordinate system. Then, based on the projection matrix between the camera and LiDAR, the three-dimensional coordinates of each point on the lane line are determined. However, the ground truth lane values ​​obtained through simple projection have low accuracy and are not smooth, requiring manual correction. Furthermore, the ground truth values ​​generated are poor for complex scenes, making them unsuitable for practical production. Summary of the Invention

[0003] This invention provides a lane marking method, apparatus, device, medium, and product to improve the accuracy and efficiency of lane marking results.

[0004] According to one aspect of the present invention, a lane marking method is provided, comprising:

[0005] Obtain a set of target road data pairs, where each target road data pair consists of a target road point cloud and a target road image at the same timestamp;

[0006] Each target road data pair is input into the target model to obtain the target lane line labeling result. The target model is obtained by iteratively training a neural network model using a training sample set, which includes a set of historical road data pairs and the lane line labeling results corresponding to the set of historical road data pairs.

[0007] According to another aspect of the present invention, a lane marking device is provided, the device comprising:

[0008] The acquisition module is used to acquire a set of target road data pairs, wherein the target road data pairs are the target road point cloud and the target road image at the same timestamp;

[0009] The input module is used to input each target road data pair into the target model to obtain the target lane line labeling result. The target model is obtained by iteratively training a neural network model through a training sample set, which includes: a set of historical road data pairs and the lane line labeling results corresponding to the set of historical road data pairs.

[0010] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0011] At least one processor; and

[0012] A memory communicatively connected to the at least one processor; wherein,

[0013] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the lane marking method according to any embodiment of the present invention.

[0014] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the lane marking method according to any embodiment of the present invention.

[0015] According to another aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the lane marking method described in any embodiment of the present invention.

[0016] This invention provides a training sample set by acquiring a set of historical road data pairs and the corresponding lane line annotation results. A target model is obtained by iteratively training a neural network model using this training sample set. Then, a target road data pair set is acquired, consisting of multiple pairs of target road point clouds and target road images at the same time stamp. These target road data pairs are input into the target model to obtain the target lane line annotation results. This invention allows for direct recognition of point cloud data and image information based on the trained target model to obtain lane line annotation results. Compared to existing methods that rely solely on simple projection to obtain lane line ground truth, this approach improves the accuracy and efficiency of lane line annotation.

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a lane marking method according to an embodiment of the present invention;

[0020] Figure 2 This is a schematic diagram of the structure of a lane marking device according to an embodiment of the present invention;

[0021] Figure 3 This is a schematic diagram of the structure of an electronic device that implements the lane marking method of this invention. Detailed Implementation

[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and their derivatives, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0025] Example 1

[0026] Figure 1 This is a flowchart of a lane marking method according to an embodiment of the present invention. This embodiment is applicable to lane marking situations. The method can be executed by the lane marking device in this embodiment of the present invention, which can be implemented in software and / or hardware, such as... Figure 1 As shown, the method specifically includes the following steps:

[0027] S101. Obtain the target road data set.

[0028] It should be noted that the target road data pair set can be a collection of road point cloud data and road image pairs collected by the vehicle's onboard LiDAR and onboard camera respectively during vehicle operation. In this embodiment, the target road data pair set includes multiple target road data pairs. Specifically, each target road data pair consists of a target road point cloud and a target road image at the same timestamp. The target road point cloud can be road point cloud data collected by the LiDAR, and the target road image can be a road image captured by the camera.

[0029] Specifically, the vehicle can be equipped with LiDAR and cameras. During the driving process, road data, including point cloud data and images, can be collected by LiDAR and cameras respectively. A time synchronization algorithm is used to synchronize each frame of LiDAR point cloud and image to form multiple data pairs for subsequent lane line detection and labeling.

[0030] S102. Input each target road data pair into the target model to obtain the target lane line annotation results.

[0031] In this embodiment, the target model can be a model used for lane line detection and annotation based on point cloud data and images. Preferably, the target model can be a trained neural network model, and the target model can operate offline.

[0032] The target model is obtained by iteratively training a neural network model using a training sample set.

[0033] The training sample set includes: a set of historical road data pairs and the lane line annotation results corresponding to the set of historical road data pairs.

[0034] It should be noted that the historical road data set can be a collection of road point cloud data and road image pairs collected by the vehicle's onboard LiDAR and onboard camera during the vehicle's operation within a historical period. In this embodiment, the historical road data set may include multiple historical road data pairs. The lane line annotation result corresponding to the historical road data set can be lane lines annotated based on the road point cloud data and road image pairs in the historical road data set.

[0035] Specifically, a training sample set is formed by acquiring a set of historical road data pairs and the lane line annotation results corresponding to the set of historical road data pairs. The target model is obtained by iteratively training the neural network model using the training sample set. Then, each target road data pair is input into the target model to obtain the target lane line annotation result.

[0036] This invention provides a training sample set by acquiring a set of historical road data pairs and the corresponding lane line annotation results. A target model is obtained by iteratively training a neural network model using this training sample set. Then, a target road data pair set is acquired, consisting of multiple pairs of target road point clouds and target road images at the same time stamp. These target road data pairs are input into the target model to obtain the target lane line annotation results. This invention allows for direct recognition of point cloud data and image information based on the trained target model to obtain lane line annotation results. Compared to existing methods that rely solely on simple projection to obtain lane line ground truth, this approach improves the accuracy and efficiency of lane line annotation.

[0037] Optionally, the neural network model is iteratively trained using a training sample set, including:

[0038] Obtain a set of historical road data pairs.

[0039] Among them, the historical road data pairs are historical road point clouds and historical road images at the same timestamp.

[0040] Among them, historical road point cloud can be road point cloud data collected by lidar within a historical time period, and target road image can be road image captured by camera within a historical time period.

[0041] Specifically, a dedicated road data collection vehicle can be equipped with LiDAR and a camera. While the vehicle is in motion, road data, including point cloud data and images, can be collected by the LiDAR and camera. A time synchronization algorithm is used to synchronize each frame of the LiDAR point cloud and image, forming multiple data pairs for subsequent training of neural network models.

[0042] The pre-lane line marking results are determined based on historical road data.

[0043] It should be noted that the lane line annotation result can be the result of simply projecting the point cloud data onto the image and then performing lane line detection and annotation.

[0044] Specifically, the historical road point cloud at the same timestamp is projected onto the historical road image for lane line detection and annotation.

[0045] Determine the image domain classification result corresponding to each frame of historical road image.

[0046] In this embodiment, the image domain classification result can be the result obtained by classifying objects in historical road images according to different locations, different types, etc.

[0047] Specifically, each frame of the historical road image is segmented and identified to obtain the image domain classification result corresponding to each frame. For example, the segmentation and identification results may include the sky, cars, cyclists, vegetation, lane lines, roadside objects, etc.

[0048] The information map set is determined based on the image domain classification results corresponding to each historical road point cloud and each frame of historical road image.

[0049] It should be explained that an information map set can be a collection of multiple information maps obtained by projecting point cloud data onto the image domain classification results and the original image respectively, obtaining relevant parameter information to color the point cloud data.

[0050] Specifically, the historical road point cloud can be projected onto the image domain classification results and the historical road image at the same timestamp, respectively. Then, relevant parameter information is obtained from the image, and the historical road point cloud is colored according to each parameter to obtain an information map set.

[0051] The pre-lane line annotation results are corrected based on the infographic set, and the true values ​​of the lane lines are generated.

[0052] It should be noted that the correction operation can be used to correct the results of lane line detection and annotation after only projecting point cloud data onto an image, so as to obtain more accurate lane line detection and annotation results.

[0053] Among them, the true value of the lane line can be the value of the three-dimensional position coordinate information of the lane line obtained from the lane line annotation result after correction.

[0054] Specifically, the infographics in the infographic collection are imported into the annotation tool, and the advantages of each layer are utilized to correct the pre-lane line annotation results and generate the true values ​​of the lane lines.

[0055] The neural network model is trained iteratively based on the true values ​​of lane lines and the set of infographics.

[0056] Specifically, the ground truth lane line data is divided into slices of the required length. These slices are then used to train a neural network model, which can directly output an infographic. The lane line annotation results from the trained target model can then be directly used, shortening the toolpath and improving the accuracy and efficiency of lane line ground truth recognition.

[0057] Optionally, an information map set is determined based on the image domain classification results corresponding to each historical road point cloud and each frame of historical road image, including:

[0058] The static point cloud corresponding to each historical road image is determined based on the image domain classification results corresponding to each historical road image and each historical road image frame.

[0059] In this embodiment, the static point cloud can be the image domain classification result corresponding to the historical road image at the same timestamp, after projecting the historical road point cloud onto the image domain classification result corresponding to the historical road image at the same timestamp, and filtering out dynamic targets (such as moving cars, cyclists, pedestrians, etc.) in the point cloud.

[0060] Specifically, based on the calibration parameters of the lidar and camera (preferably, the calibration parameters can be, for example, a projection matrix), a single-frame lidar point cloud is projected onto the image domain classification result, dynamic targets in the point cloud are filtered out, and a single-frame static point cloud is generated.

[0061] The attribute information of each static point cloud includes: first coordinate information, second coordinate information, third coordinate information, reflection intensity information, first color information, second color information, and third color information.

[0062] In this embodiment, the first coordinate information, the second coordinate information, and the third coordinate information can each be one of the three-dimensional coordinate information of each static point cloud. For example, the first coordinate information can be the x-coordinate, the second coordinate information can be the y-coordinate, and the third coordinate information can be the z-coordinate. The reflection intensity information can be the reflection intensity of each static point cloud, which can be represented by 'i' in this embodiment. It should be noted that the first color information, the second color information, and the third color information can each be the RGB information corresponding to each static point cloud in the image domain classification result (RGB color mode is an industry color standard that obtains various colors by varying the red (R), green (G), and blue (B) color channels and superimposing them; RGB represents the colors of the red, green, and blue channels). For example, the first color information can be represented as MR, the second color information as MG, and the third color information as MB.

[0063] In practical operation, based on the calibration parameters of the LiDAR and camera (preferably, the calibration parameters can be, for example, a projection matrix), a single frame of LiDAR point cloud is projected onto the image domain classification result. Dynamic targets in the point cloud are filtered out, generating a single frame of static point cloud. Each point cloud contains seven types of parameter information: x, y, z, i, MR, MG, and MB. Among them, x, y, and z are the three-dimensional coordinates of each point cloud, i is the reflection intensity of each point cloud, and MR, MG, and MB are the RGB information obtained from the image domain classification result.

[0064] Project the static point cloud corresponding to each frame of historical road image onto each frame of historical road image.

[0065] Specifically, based on the calibration parameters of the lidar and camera (preferably, the calibration parameters can be, for example, a projection matrix), a single frame of lidar point cloud is projected onto the original image, i.e., the historical road image.

[0066] Obtain the fourth, fifth, and sixth color information corresponding to each static point cloud in each frame of historical road image.

[0067] It should be noted that the fourth, fifth, and sixth color information can be the RGB information corresponding to each static point cloud in the historical road image. For example, the fourth color information can be represented as R, the fifth color information can be represented as G, and the sixth color information can be represented as B.

[0068] Specifically, the static point cloud is projected onto the original image to obtain the corresponding RGB information of the static point cloud in the original image. By combining the image domain classification results with the information obtained from the original image, ten dimensions of information—x, y, z, i, MR, MG, MB, R, G, and B—can be obtained.

[0069] The current vehicle pose information is obtained, and the static point cloud corresponding to each frame of historical road image is stitched together based on the current vehicle pose information to obtain at least one local point cloud map.

[0070] The current vehicle can be a road data collection vehicle currently equipped with LiDAR and a camera. The pose information of the current vehicle can be obtained by performing image recognition on the images acquired by the camera.

[0071] In this embodiment, the stitching operation can be an operation of stitching static point clouds into a continuous, time-sequential point cloud map based on the current vehicle's pose information. Specifically, the local point cloud map can be a point cloud map formed by stitching together static point clouds of a certain length (which can be preset according to actual conditions; this embodiment does not limit this, for example, it could be 150 meters) based on the current vehicle's pose information.

[0072] For example, a static point cloud is stitched together based on the current vehicle's pose information, and a local point cloud map is generated every 150 meters.

[0073] A reflection intensity map is generated based on the reflection intensity information and the local point cloud map.

[0074] Among them, the reflection intensity map can be an information map generated after coloring the local point cloud map according to the reflection intensity information i.

[0075] In actual operation, different reflection intensity information i can be set to different colors, or different reflection intensity information i in different intervals can be set to different colors. The local point cloud map is colored according to different reflection intensity information i to generate a reflection intensity map.

[0076] A height map is generated based on the third coordinate information and the local point cloud map.

[0077] Among them, the height map can be an information map generated by coloring a local point cloud map according to the third coordinate information z.

[0078] In practice, different colors can be set for different third coordinate information z, or different colors can be set for the third coordinate information z of different intervals. The local point cloud map is colored according to the different third coordinate information z to generate a height map.

[0079] A semantic map is generated based on the first color information, the second color information, the third color information, and the local point cloud map.

[0080] The semantic graph can be an information graph generated by coloring a local point cloud map according to the first color information MR, the second color information MG, and the third color information MB.

[0081] In practice, different colors can be set for the first color information MR, the second color information MG, and the third color information MB, or different colors can be set for the first color information MR, the second color information MG, and the third color information MB in different intervals. The local point cloud map is colored according to the different first color information MR, the second color information MG, and the third color information MB to generate a semantic map.

[0082] An RGB image is generated based on the fourth, fifth, and sixth color information, as well as the local point cloud map.

[0083] The RGB image can be an information image generated by coloring a local point cloud map according to the fourth color information R, the fifth color information G, and the sixth color information B.

[0084] In actual operation, different colors can be set for the fourth color information R, the fifth color information G, and the sixth color information B, or different colors can be set for the fourth color information R, the fifth color information G, and the sixth color information B in different intervals. The local point cloud map is colored according to the different fourth color information R, the fifth color information G, and the sixth color information B to generate an RGB image.

[0085] The information graph set is determined based on the reflection intensity map, height map, semantic map, and RGB map.

[0086] Specifically, it consists of an information graph set composed of a reflection intensity map, a height map, a semantic map, and an RGB map.

[0087] Optionally, the neural network model is iteratively trained based on the lane line ground truth and infographic set, including:

[0088] Establish a neural network model.

[0089] Obtain the pose of the end accumulation point corresponding to each local point cloud map.

[0090] It should be explained that the pose information of the final accumulation point can be the pose information of the last point cloud corresponding to each local point cloud map when the local point cloud map is generated.

[0091] For example, a static point cloud is stitched together based on the current vehicle's pose information. Every 150 meters accumulated, a local point cloud map is generated, and the pose information of the end accumulation point is saved.

[0092] The lane line ground truth is divided based on the pose information of the end accumulation point corresponding to each local point cloud map, resulting in at least one slice of data.

[0093] It should be noted that the segmentation operation can be an operation of dividing the lane line ground truth into segments according to the pose information of the end accumulation point. Specifically, the slice data can be several lane line ground truth segments obtained by dividing the lane line ground truth based on the pose information of the end accumulation point corresponding to each local point cloud map.

[0094] Specifically, based on the pose information of the end accumulation point, the true value of the lane line is divided into several slices of data of the required length.

[0095] Each slice of data is input into a neural network model to obtain the predicted lane line labeling results.

[0096] It should be noted that the predicted lane line labeling results can be the lane line labeling results output by the neural network model after making predictions based on the input slice data.

[0097] The predicted lane marking results include: predicted reflection intensity map, predicted height map, predicted semantic map, and predicted RGB map.

[0098] Specifically, these slice data are used to train a neural network model. Each slice data is input into the neural network model, and the neural network model can directly output a predicted reflection intensity map, a predicted height map, a predicted semantic map, and a predicted RGB map.

[0099] The parameters of the neural network model are trained based on the objective function formed by the lane line labeling results and historical road data corresponding to the lane line labeling results.

[0100] The objective function can be a loss function, and the parameters can be weights or other parameters of the neural network model.

[0101] Return to the previous step and execute the operation of inputting each slice of data into the neural network model to obtain the predicted lane line labeling results, until the target model is obtained.

[0102] Optionally, the pre-lane line marking results are determined from the set based on historical road data, including:

[0103] The detection of each historical road image is performed based on a preset model group, and the image domain detection result corresponding to each historical road image is obtained.

[0104] In this embodiment, the preset model group can be a model group used to identify lane lines, curbs, stop lines and road signs for each frame of historical road images.

[0105] The preset model group includes: preset lane line detection model, preset curb detection model, preset stop line detection model, and preset road sign detection model.

[0106] In actual operation, the preset lane line detection model, preset curb detection model, preset stop line detection model, and preset road sign detection model can all be pre-trained neural network models. This embodiment does not limit the training process of these models, and historical road images can be directly input into each pre-trained model for detection.

[0107] In this embodiment, the image domain detection result can be the result obtained after detecting and classifying lane lines, curbs, stop lines, and road signs in historical road images.

[0108] Specifically, each frame of historical road image is processed using a preset lane line detection model, a preset curb detection model, a preset stop line detection model, and a preset road sign detection model to obtain the image domain detection results corresponding to lane lines, curbs, stop lines, and road signs, respectively.

[0109] Based on the first segmentation model, the historical road point cloud is identified to obtain a set of road point clouds.

[0110] In this embodiment, the first segmentation model can be a ground segmentation model, used to identify and segment the road surface point cloud from each frame of point cloud.

[0111] The road point cloud set can be a dense road point cloud generated by encrypting the road surface point cloud. This embodiment does not limit the specific encryption algorithm.

[0112] Specifically, a ground segmentation model is used to process the point cloud of each historical road frame to obtain the road surface point cloud, and a dense road point cloud is generated through an encryption algorithm.

[0113] The pre-lane line annotation results are determined based on the point cloud in the road point cloud set and the image domain detection results corresponding to each frame of historical road image.

[0114] Specifically, based on the calibration parameters of the lidar and camera (preferably, the calibration parameters can be, for example, a projection matrix), the dense road point cloud is projected onto the image domain detection results to obtain the three-dimensional coordinates of lane lines, road edges, stop lines, and road signs, and generate pre-lane line annotation results.

[0115] Optionally, determine the image domain classification result corresponding to each frame of historical road image, including:

[0116] The second segmentation model is used to identify each frame of historical road image, and the image domain classification result corresponding to each frame of historical road image is obtained.

[0117] In this embodiment, the second segmentation model can be a panoramic segmentation model, used to identify and segment different objects such as sky, cars, cyclists, vegetation, lane lines, and curbs from each frame of historical road images.

[0118] Specifically, a panoramic segmentation model is used to process each frame of historical road image to obtain the image domain classification result corresponding to each frame of historical road image. For example, the result may include sky, cars, cyclists, vegetation, lane lines, roadside, etc.

[0119] Existing technologies generate lane line ground truth values ​​that are only discrete data from a single frame, lacking temporal sequence and making them inconvenient for downstream tracking and control modules. The technical solution of this invention first uses deep learning models such as lane line detection, curb detection, stop line detection, and road sign detection to output 2D lane line detection results. The point cloud is then projected onto the image to generate 3D lane lines. The lane lines are then stitched together based on the vehicle's pose information to generate 4D lane lines. The pre-annotation results are further optimized in an annotation tool to generate 4D lane line ground truth values. The generated 4D lane lines are then used to train a neural network model. This offline training of the neural network model to generate pre-annotated labels improves the accuracy and efficiency of the annotation results.

[0120] As an exemplary description of an embodiment of the present invention, the lane line marking method can be described as follows:

[0121] S1. By using a data acquisition vehicle (i.e., a road data acquisition vehicle) equipped with LiDAR and a camera, road data is collected, and a time synchronization algorithm is used to synchronize each frame of LiDAR point cloud and image.

[0122] S2. Use the trained lane line detection model, curb detection model, stop line detection model, and road sign detection model to process each frame of the image to obtain the image domain detection results of lane lines, curbs, stop lines, and road signs.

[0123] S3. Use the ground segmentation model to process each frame of point cloud to obtain road surface point cloud, and generate dense road point cloud through encryption algorithm.

[0124] S4. Based on the calibration parameters of the lidar and camera, project the dense road point cloud in S3 onto the image domain detection results in S2 to obtain the three-dimensional coordinates of lane lines, curbs, stop lines and road signs, and generate pre-lane line annotation results.

[0125] S5. Use the panoramic segmentation model to process each frame of the image and obtain the classification results of the image domain, including the sky, cars, cyclists, vegetation, lane lines, curbs, etc.

[0126] S6. Based on the calibration parameters of the LiDAR and camera, firstly, the single-frame LiDAR point cloud is projected onto the image domain classification result to filter out dynamic targets in the point cloud, generating a single-frame static point cloud. Each point contains seven parameters: x, y, z, i, MR, MG, and MB. x, y, and z represent the three-dimensional coordinates of each point, i represents the reflection intensity of each point, and MR, MG, and MB are the RGB information obtained from the image domain classification result. Next, the static point cloud is projected onto the original image to obtain the corresponding RGB information of the static point cloud in the original image. Finally, the image domain classification result and the information obtained from the original image are combined to obtain ten dimensions of information: x, y, z, i, MR, MG, MB, R, G, and B.

[0127] S7. Based on the pose information of the data acquisition vehicle, stitch together a static point cloud. Every 150 meters accumulated, generate a local point cloud map and save the pose of the end accumulation point.

[0128] S8. Color the local point cloud map according to i to generate a reflection intensity map. Color the local point cloud map according to z to generate a height map. Color the local point cloud map according to MR, MG, and MB to generate a semantic map. Color the local point cloud map according to R, G, and B to generate an RGB map.

[0129] S9. Import the reflection intensity map, height map, semantic map, and RGB map into the annotation tool. Utilize the advantages of each of the four layers to correct the pre-lane line annotation results and generate the true values ​​of the lane lines.

[0130] S10. Based on the pose information of the end accumulation point, divide the lane line ground truth into slices of the required length. Use these slices to train the offline large model (i.e., the target model). The offline large model directly outputs the reflection intensity map, height map, semantic map, and RGB map from S8.

[0131] S11. Repeat steps S1 to S10 until the accuracy of offline large model inference is better than the result of S8. Then end steps S2 to S8 and switch to using the pre-labeled results of offline large model to shorten the tool chain and improve the accuracy and efficiency of lane line ground truth.

[0132] The technical solution of this invention consists of two stages. The first stage uses models from multiple image domains to obtain rich semantic and geometric information, and stitches the point cloud into a local map based on the vehicle's pose information. Compared to single-frame results, temporal data can solve the problem of lane line breaks caused by category jumps and vehicle occlusion. Existing methods mostly use lane line detection models to obtain image domain detection results, and then obtain the 3D coordinates of the lane lines through projected laser point clouds. Lane lines are prone to distortion and truncation. The first stage of this invention combines the detection results and semantic information of multiple models, utilizing the temporal nature of the data to generate lane line ground truth values ​​with better continuity and higher accuracy. The second stage first uses the results of the first stage to generate high-precision 4D lane line ground truth values ​​through annotation tools, and then uses these 4D lane line ground truth values ​​to train an offline large-scale model. The offline large-scale model can directly output 4D lane line ground truth values. The lane line ground truth values ​​generated in this way are far more accurate than those generated in the first stage, and are also more efficient. When the accuracy of the offline large model exceeds the accuracy of the first stage, switch to generating pre-lane line annotation results using the offline large model, and continuously optimize the offline large model using accumulated data. In the end, only a small number of pre-annotation results need to be corrected in the annotation tool, which can significantly improve the generation efficiency and accuracy of lane line ground truth values.

[0133] Example 2

[0134] Figure 2 This is a schematic diagram of a lane marking device according to an embodiment of the present invention. This embodiment is applicable to lane marking applications. The device can be implemented using software and / or hardware, and can be integrated into any device that provides lane marking functionality, such as… Figure 2 As shown, the lane marking device specifically includes: an acquisition module 201 and an input module 202.

[0135] The acquisition module 201 is used to acquire a set of target road data pairs, wherein the target road data pairs are the target road point cloud and the target road image at the same timestamp;

[0136] The input module 202 is used to input each of the target road data pairs into the target model to obtain the target lane line labeling results. The target model is obtained by iteratively training a neural network model through a training sample set. The training sample set includes: a set of historical road data pairs and the lane line labeling results corresponding to the set of historical road data pairs.

[0137] Optionally, the input module 202 includes:

[0138] The acquisition unit is used to acquire a set of historical road data pairs, wherein the historical road data pairs are historical road point clouds and historical road images at the same timestamp;

[0139] The first determining unit is used to determine the pre-lane line marking result of the set based on the historical road data;

[0140] The second determining unit is used to determine the image domain classification result corresponding to each frame of the historical road image;

[0141] The third determining unit is used to determine an information map set based on the image domain classification results corresponding to each historical road point cloud and each frame of the historical road image;

[0142] The correction unit is used to correct the pre-lane line marking results based on the information map set and generate lane line true values.

[0143] The training unit is used to iteratively train a neural network model based on the true values ​​of the lane lines and the set of information graphs.

[0144] Optionally, the third determining unit is specifically used for:

[0145] The static point cloud corresponding to each historical road image is determined based on the image domain classification results corresponding to each historical road image and each historical road image frame. The attribute information of each static point cloud includes: first coordinate information, second coordinate information, third coordinate information, reflection intensity information, first color information, second color information, and third color information.

[0146] Project the static point cloud corresponding to each frame of historical road image onto each frame of historical road image;

[0147] Obtain the fourth, fifth, and sixth color information corresponding to each static point cloud in each frame of historical road image;

[0148] The pose information of the current vehicle is obtained, and the static point cloud corresponding to each frame of historical road image is stitched together based on the pose information of the current vehicle to obtain at least one local point cloud map.

[0149] A reflection intensity map is generated based on the reflection intensity information and the local point cloud map;

[0150] A height map is generated based on the third coordinate information and the local point cloud map;

[0151] A semantic map is generated based on the first color information, the second color information, the third color information, and the local point cloud map;

[0152] An RGB image is generated based on the fourth color information, the fifth color information, the sixth color information, and the local point cloud map;

[0153] The information graph set is determined based on the reflection intensity map, the height map, the semantic map, and the RGB map.

[0154] Optionally, the training unit is specifically used for:

[0155] Establish a neural network model;

[0156] Obtain the pose information of the end accumulation point corresponding to each of the local point cloud maps;

[0157] The lane line ground truth is divided according to the pose information of the end accumulation point corresponding to each local point cloud map to obtain at least one slice data;

[0158] Each slice of data is input into the neural network model to obtain the predicted lane line labeling results, wherein the predicted lane line labeling results include: predicted reflection intensity map, predicted height map, predicted semantic map and predicted RGB map;

[0159] The parameters of the neural network model are trained based on the objective function formed by the lane line marking results and the historical road data corresponding to the lane line marking results.

[0160] Return to the operation of inputting each slice of data into the neural network model to obtain the predicted lane line labeling results, until the target model is obtained.

[0161] Optionally, the first determining unit is specifically used for:

[0162] Each frame of historical road image is detected based on a preset model group to obtain the image domain detection result corresponding to each frame of historical road image. The preset model group includes: a preset lane line detection model, a preset curb detection model, a preset stop line detection model, and a preset road sign detection model.

[0163] Based on the first segmentation model, the historical road point cloud is identified to obtain a set of road point clouds;

[0164] The pre-lane line labeling result is determined based on the point cloud in the road point cloud set and the image domain detection result corresponding to each frame of historical road image.

[0165] Optionally, the second determining unit is specifically used for:

[0166] The second segmentation model is used to identify each frame of historical road image, and the image domain classification result corresponding to each frame of historical road image is obtained.

[0167] The above-mentioned products can execute the lane marking method provided in any embodiment of the present invention, and have the corresponding functional modules and beneficial effects of the execution method.

[0168] Example 3

[0169] Figure 3A schematic diagram of an electronic device 30 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0170] like Figure 3 As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory (ROM) 32 or a random access memory (RAM) 33, communicatively connected to the at least one processor 31. The memory stores computer programs executable by the at least one processor. The processor 31 can perform various appropriate actions and processes based on the computer program stored in the ROM 32 or loaded from storage unit 38 into the RAM 33. The RAM 33 can also store various programs and data required for the operation of the electronic device 30. The processor 31, ROM 32, and RAM 33 are interconnected via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.

[0171] Multiple components in electronic device 30 are connected to I / O interface 35, including: input unit 36, such as keyboard, mouse, etc.; output unit 37, such as various types of monitors, speakers, etc.; storage unit 38, such as disk, optical disk, etc.; and communication unit 39, such as network card, modem, wireless transceiver, etc. Communication unit 39 allows electronic device 30 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0172] Processor 31 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 31 performs the various methods and processes described above, such as the lane marking method:

[0173] Obtain a set of target road data pairs, where each target road data pair consists of a target road point cloud and a target road image at the same timestamp;

[0174] Each target road data pair is input into the target model to obtain the target lane line labeling result. The target model is obtained by iteratively training a neural network model using a training sample set, which includes a set of historical road data pairs and the lane line labeling results corresponding to the set of historical road data pairs.

[0175] In some embodiments, the lane marking method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 38. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 30 via ROM 32 and / or communication unit 39. When the computer program is loaded into RAM 33 and executed by processor 31, one or more steps of the lane marking method described above may be performed. Alternatively, in other embodiments, processor 31 may be configured to perform the lane marking method by any other suitable means (e.g., by means of firmware).

[0176] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0177] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0178] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0179] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0180] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0181] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0182] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the lane marking method of any embodiment of the present invention.

[0183] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0184] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0185] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for marking lane lines, characterized in that, include: Obtain a set of target road data pairs, where each target road data pair consists of a target road point cloud and a target road image at the same timestamp; Each target road data pair is input into the target model to obtain the target lane line labeling result. The target model is obtained by iteratively training a neural network model using a training sample set, which includes a set of historical road data pairs and the lane line labeling results corresponding to the set of historical road data pairs.

2. The method according to claim 1, characterized in that, The neural network model is trained iteratively using a training sample set, including: Obtain a set of historical road data pairs, where each historical road data pair consists of a historical road point cloud and a historical road image at the same timestamp; The pre-lane line marking results are determined based on the historical road data. Determine the image domain classification result corresponding to each frame of the historical road image; An information map set is determined based on the image domain classification results corresponding to each historical road point cloud and each frame of the historical road image; The pre-lane line annotation results are corrected based on the information map set, and lane line true values ​​are generated; The neural network model is iteratively trained based on the true values ​​of the lane lines and the set of information graphs.

3. The method according to claim 2, characterized in that, An information map set is determined based on the image domain classification results corresponding to each historical road point cloud and each frame of the historical road image, including: The static point cloud corresponding to each historical road image is determined based on the image domain classification results corresponding to each historical road image and each historical road image frame. The attribute information of each static point cloud includes: first coordinate information, second coordinate information, third coordinate information, reflection intensity information, first color information, second color information, and third color information. Project the static point cloud corresponding to each frame of historical road image onto each frame of historical road image; Obtain the fourth, fifth, and sixth color information corresponding to each static point cloud in each frame of historical road image; The pose information of the current vehicle is obtained, and the static point cloud corresponding to each frame of historical road image is stitched together based on the pose information of the current vehicle to obtain at least one local point cloud map. A reflection intensity map is generated based on the reflection intensity information and the local point cloud map; A height map is generated based on the third coordinate information and the local point cloud map; A semantic map is generated based on the first color information, the second color information, the third color information, and the local point cloud map; An RGB image is generated based on the fourth color information, the fifth color information, the sixth color information, and the local point cloud map; The information graph set is determined based on the reflection intensity map, the height map, the semantic map, and the RGB map.

4. The method according to claim 3, characterized in that, Iteratively train a neural network model based on the true values ​​of the lane lines and the set of information maps, including: Establish a neural network model; Obtain the pose information of the end accumulation point corresponding to each of the local point cloud maps; The lane line ground truth is divided according to the pose information of the end accumulation point corresponding to each local point cloud map to obtain at least one slice data; Each slice of data is input into the neural network model to obtain the predicted lane line labeling results, wherein the predicted lane line labeling results include: predicted reflection intensity map, predicted height map, predicted semantic map and predicted RGB map; The parameters of the neural network model are trained based on the objective function formed by the lane line marking results and the historical road data corresponding to the lane line marking results. Return to the operation of inputting each slice of data into the neural network model to obtain the predicted lane line labeling results, until the target model is obtained.

5. The method according to claim 2, characterized in that, The pre-lane line marking results are determined based on the historical road data, including: Each frame of historical road image is detected based on a preset model group to obtain the image domain detection result corresponding to each frame of historical road image. The preset model group includes: a preset lane line detection model, a preset curb detection model, a preset stop line detection model, and a preset road sign detection model. Based on the first segmentation model, the historical road point cloud is identified to obtain a set of road point clouds; The pre-lane line labeling result is determined based on the point cloud in the road point cloud set and the image domain detection result corresponding to each frame of historical road image.

6. The method according to claim 2, characterized in that, Determine the image domain classification result corresponding to each frame of the historical road image, including: The second segmentation model is used to identify each frame of historical road image, and the image domain classification result corresponding to each frame of historical road image is obtained.

7. A lane marking device, characterized in that, include: The acquisition module is used to acquire a set of target road data pairs, wherein the target road data pairs are the target road point cloud and the target road image at the same timestamp; The input module is used to input each target road data pair into the target model to obtain the target lane line labeling result. The target model is obtained by iteratively training a neural network model through a training sample set, which includes: a set of historical road data pairs and the lane line labeling results corresponding to the set of historical road data pairs.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the lane marking method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the lane marking method according to any one of claims 1-6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the lane marking method according to any one of claims 1-6.