Automatic driving high-value sample data mining method

Through the high-value sample data mining method of autonomous driving, multimodal large models and point cloud data are used for scene understanding and clustering, and high-value data in the autonomous driving system is screened and labeled, which solves the problem of unbalanced data training and improves the stability and robustness of the model.

CN119964113APending Publication Date: 2025-05-09BEIJING TIEMUNIU INTELLIGENT MACHINERY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510067137.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively screen and label high-value sample data in autonomous driving systems, resulting in uneven data training and affecting the stability and robustness of the model.

Method used

The high-value sample data mining method of autonomous driving is used to understand global and local scenes through multimodal large models, unsupervised clustering is carried out in combination with point cloud data, distinguish known and unknown obstacle types, select badcase and hardcase for special annotation, and use the marked data for model training.

Benefits of technology

It improves the balance and customization of data in the autonomous driving system, enhances the stability and robustness of the model, and can learn and understand complex driving scenarios more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964113A_ABST
    Figure CN119964113A_ABST
Patent Text Reader

Abstract

The invention provides an automatic driving high-value sample data mining method. The automatic driving high-value sample data mining method comprises the following steps: acquiring data acquired by a vehicle end, performing data analysis and data cleaning, and then storing the data in a data platform; training an automatic driving multi-mode large model based on the data medium table and the scene understanding data set to carry out global scene understanding to obtain global scene understanding information, and storing the global scene understanding information as a global information label in a search server; training a target detection model based on the data middle table and the target detection data set to obtain local known category information and a target detection result, and storing the local known category information as a local known information label in a search server; performing point cloud unsupervised clustering based on the Euclidean distance between points in the point cloud data of the data middle station to obtain a clustering result list; and projecting a point cloud clustering result into an image, matching a projection result on the image with a target detection result to obtain a local unknown category information tag, and storing the local unknown category information tag into a search server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and in particular relates to a method for mining high-value sample data for autonomous driving. Background Art

[0002] Data collection is the starting point for autonomous driving model training. It can rely on sensor technology to collect data through road collection vehicles, mass-produced vehicles, and data contributions from car owners. It can also use multimodal large model technology to understand the collected data in the scene, and combine visual data and point cloud data to distinguish between known and unknown obstacle categories on the road, and solve specific data needs and corner case problems.

[0003] In the field of intelligent driving, data is the source of life. Without data, those complex algorithms and models are like dried-up rivers, unable to nourish the fruits of intelligence. The autonomous driving model can only learn driving behavior and understanding of the environment through driving video clips, so it is difficult to give the content that humans want the model to learn to the data and let the model learn this prior knowledge during training. Because each video clip of human driving actually contains a wealth of driving behavior, it is not easy for the model to understand a certain abstract prior knowledge in these video clips (such as turning left to go straight).

[0004] By training the model with a large amount of data, the autonomous driving system can recognize and predict various driving scenarios. The input of high-quality data directly determines the accuracy and reliability of the model output. This data not only needs to cover various road conditions, weather changes and traffic conditions, but also ensure the accuracy and diversity of its annotations.

[0005] From the data dimension, massive and high-quality data is becoming a "scarce commodity" in the autonomous driving industry. Generally, the lidar algorithm needs at least hundreds of thousands of frames of data training to meet the performance requirements of autonomous driving. Monocular cameras have higher requirements and require millions of frames of training data. To achieve high-value and more intelligent autonomous driving levels, massive, diverse, and high-quality data is the primary prerequisite.

[0006] Often, most of the collected data is worthless data. This type of data is real road scene data, but its distribution is biased and cannot meet the needs of balanced and customized autonomous driving training data. Summary of the invention

[0007] The purpose of the present invention is to solve the problems existing in the prior art and propose a method for mining high-value sample data for autonomous driving, which can replace manual screening and improve work efficiency.

[0008] In order to achieve the above object, the present invention adopts the following technical scheme.

[0009] The method for mining high-value sample data for autonomous driving comprises:

[0010] Obtain the data collected by the vehicle, perform data analysis and data cleaning, and then save it to the data center;

[0011] Based on the data center and the scene understanding data set, a multimodal large model of autonomous driving is trained to perform global scene understanding to obtain global scene understanding information, and the global scene understanding information is saved as a global information tag to the search server;

[0012] Training a target detection model based on the data center and the target detection data set to obtain local known category information and target detection results, and saving the local known category information as a local known information label to the search server;

[0013] Performing unsupervised point cloud clustering based on the Euclidean distance between points in the point cloud data of the data center to obtain a clustering result list;

[0014] The point cloud clustering result is projected into the image, the projection result on the image is matched with the target detection result, and the local unknown category information label is obtained and saved in the search server.

[0015] Furthermore, the acquisition of data collected by the vehicle end and the data parsing and cleaning and then saving to the data center include:

[0016] Collect vehicle-side data based on vehicle-mounted systems and sensors, and convert the collected vehicle-side data into the same format;

[0017] Identify missing values ​​in vehicle-side data through data cleaning, and remove or complete missing values ​​based on the missing mechanism;

[0018] Deviations in sensor data or errors in vehicle-side data entry can be identified and corrected through a calibration process or by comparing with other data sources.

[0019] Furthermore, the clustering result list obtained by performing unsupervised clustering of point cloud based on the Euclidean distance between points in the point cloud data of the data center includes:

[0020] If the distance between any two points in the point cloud data is less than the set threshold, the two points will be clustered into one category.

[0021] Furthermore, the clustering result list obtained by performing unsupervised clustering of point cloud based on the Euclidean distance between points in the point cloud data of the data center includes:

[0022] Create an empty result list to store clustering results;

[0023] Initialize the flag array to mark whether the point has been processed;

[0024] Build a spatial index structure, traverse a single point in the point cloud data, and start clustering if the point has not been processed;

[0025] Select a single unprocessed point as the starting point, find all neighboring points within the set threshold distance, mark them as processed and add them to the current cluster until all neighboring points are added to the cluster;

[0026] Continue to traverse the unprocessed points until all points in the point cloud data are processed, and return the clustering result list.

[0027] Furthermore, projecting the point cloud clustering results into the image includes: projecting the lidar data onto the image through the internal and external parameter relationship between the camera and the lidar, and after the image and the lidar are synchronized in time, a single image corresponds to the laser point cloud at the same moment.

[0028] Furthermore, it also includes:

[0029] The point cloud data and image data of the LiDAR are collected for calibration to obtain the internal and external parameters of the camera and LiDAR.

[0030] Further, including,

[0031] Calculate the camera intrinsic parameters through several chessboard images and store the camera intrinsic parameters;

[0032] Collect images and point cloud data from the calibration room, and project the point cloud into the image according to the parameter adjustment parameters until the point cloud and image coincide to obtain the external parameter parameters.

[0033] Furthermore, matching the projection result on the image with the target detection result to obtain the local unknown category information label and saving it to the search server includes: on the image, taking the target whose target detection result matches the point cloud clustering result as the local known obstacle type, and taking the target whose point cloud clustering result does not match the target detection result as the unknown obstacle type.

[0034] In order to achieve the above-mentioned purpose, the present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a program running on the processor, and when the processor runs the program, the steps of the method for mining high-value sample data for autonomous driving as described above are executed.

[0035] In order to achieve the above-mentioned objectives, the present invention also provides a computer-readable storage medium on which computer instructions are stored, and when the computer instructions are executed, the steps of the method for mining high-value sample data for autonomous driving as described above are executed.

[0036] The present invention proposes a method for mining high-value sample data for autonomous driving, which has the following beneficial effects:

[0037] The present invention targets the massive amount of collected data samples for autonomous driving, introduces a large multimodal model to perform scene understanding, distinguishes known obstacle types from unknown obstacle types, selects badcases and hardcases in the autonomous driving system, and then performs special labeling. The labeled data is added to the autonomous driving model training, and the autonomous driving model is continuously iterated to improve the stability and robustness of the autonomous driving system.

[0038] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be apparent from the description, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0040] Figure 1 This is a flow chart of a method for mining high-value sample data for autonomous driving according to the present invention;

[0041] Figure 2 The figure is a schematic diagram of the overall process of a method for mining high-value sample data for autonomous driving according to the present invention. DETAILED DESCRIPTION

[0042] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0043] Example 1

[0044] Figure 1 The flowchart of the method for mining high-value sample data for autonomous driving according to the present invention is as follows: Figure 1 , the high-value sample data mining method for autonomous driving of the present invention is described in detail.

[0045] In step 101, the data collected by the vehicle is obtained and the data is analyzed and cleaned before being saved to the data center.

[0046] Optionally, the data collected on the vehicle side usually comes from various sensors and vehicle systems, and is usually stored in the form of BAG packages. Before the data is stored in the database, data parsing must be performed first, and the data needs to be converted into a unified format (formatted) for further processing. The timestamps need to be in a unified format to synchronize data from different sources, and sensor readings may need to be converted into the same dimension or unit system. Vehicle-side data may be missing due to technical problems or signal interference. Data cleaning needs to identify these missing values ​​and eliminate or complete them according to the missing mechanism. After completing the above operations, import the collected data into the data center to complete the storage.

[0047] In step 102, a large multimodal model of autonomous driving is trained based on the data center and the scene understanding data set to perform global scene understanding to obtain global scene understanding information, which is saved as a global information tag in the search server.

[0048] Optionally, the multimodal large model trained based on a massive public data set of images and texts has a deep understanding of the real world, especially for global autonomous driving scenarios, such as weather conditions (sunny, cloudy, rainy, snowy, foggy), road conditions (congested), and lighting conditions (strong light, weak light, backlight, front light). It has strong recognition capabilities.

[0049] In this embodiment, the public data set includes scene understanding data sets annotated for autonomous driving, as well as common COCO, VG, SBU, CC3M, CC12M and 115M LAION400M images. If a specific data set (can be a small amount) is needed for fine-tuning the model, a portion of it can be annotated, such as a few hundred images.

[0050] Optionally, the multimodal large model structure adopts the structure of BLIP2.

[0051] In this embodiment, the image encoder and the large language model are frozen. As a pre-training method, Q-Former plays a vital role. The left and right sides are the two stages of pre-training. The first stage is dedicated to improving the model's representation learning ability so that the query token and the text token can be aligned and the most relevant visual features of the text can be extracted. The second stage is dedicated to improving the model's vision-to-language generation learning ability so that the query token can be understood by the large language model.

[0052] Optionally, output the global scene understanding information, save it as label information, and store it in a retrieval database. For example, use the label as the key and the image name and address as the value. By searching the key value, you can find the value and find the data that meets our requirements.

[0053] In step 103, the target detection model is trained using real-world driving data and an autonomous driving target detection dataset to obtain local known category information and target detection results, and the local known category information is saved as a local known information label in the search server.

[0054] Optionally, a target detection model is trained based on real-world driving data, and the categories include common obstacle types for autonomous driving, such as cars, buses, trucks, pedestrians, cyclists, motorcycles, traffic lights, etc.

[0055] Optionally, the dataset includes an object detection dataset annotated for autonomous driving, as well as common datasets for autonomous driving such as COCO.

[0056] In this embodiment, the target detection dataset can be a public dataset, such as the COCO dataset, or a self-annotated dataset.

[0057] Optionally, the model structure can adopt a yolo series model or a DETR series model.

[0058] Optionally, the output local known categories are saved to a label database.

[0059] In step 104, unsupervised point cloud clustering is performed based on the Euclidean distance between points in the point cloud data to generate a point cloud clustering result list.

[0060] Optionally, point cloud Euclidean clustering is a method for clustering based on the Euclidean distance between points in point cloud data. As long as the distance between any points is less than a set threshold, the two points are clustered into one category.

[0061] Optionally, the general steps of Euclidean clustering algorithm programming are as follows:

[0062] (1) Initialization:

[0063] Create an empty result list to store the clustering results.

[0064] Initialize a flag array to mark whether the point has been processed.

[0065] (2) Constructing the search structure:

[0066] Build a spatial index structure (such as a KD tree) to quickly find the neighborhood of each point.

[0067] (3) Traversing the point cloud:

[0068] Iterate over every point in the point cloud.

[0069] If the point has not been processed yet, a new clustering is started.

[0070] (4) Regional growth:

[0071] Choose an unprocessed point as the starting point.

[0072] Find all neighboring points within a given distance threshold and mark them as processed.

[0073] Add these neighboring points to the current cluster.

[0074] Repeat the above process for these neighboring points until no new points are added to the cluster.

[0075] (5) Save the clustering results:

[0076] Add the current cluster to the result list.

[0077] (6) Repeat steps (3) to (5):

[0078] Continue to traverse the unprocessed points and repeat the above steps until all points are processed.

[0079] (7) Output results:

[0080] Returns a list of all clustering results.

[0081] In step 105, point cloud and visual projection are performed to obtain local known and unknown obstacle category labels.

[0082] Optionally, after the visual and lidar calibration, the lidar data is projected onto the image through the internal and external parameter relationship of the camera and lidar. After the image and radar are synchronized in time, each image corresponds to the laser point cloud at the same moment.

[0083] In this embodiment, the point cloud data and image data of the laser radar are first collected and calibrated to obtain the internal and external parameters of the camera and the laser radar. The camera's internal parameter calibration mainly uses Zhang's calibration method, which is one of the common camera calibration methods. It calculates the camera's internal parameters through multiple checkerboard images. The commonly used tool is camera-calibration in ROS, and finally a yaml file is generated to store the camera's internal parameter parameters. Collect pictures and point cloud data in the calibration room, and then project the point cloud into the image according to the manual adjustment parameters until the point cloud and the image coincide to obtain the final external parameter parameters.

[0084] Optionally, for targets on the image where the target detection model and the point cloud clustering results match, they are considered to be locally known obstacle types, and the point cloud clustering results outside the target box are considered to be unknown obstacle types.

[0085] Optionally, output local known and unknown obstacle category labels.

[0086] In step 106, the global information tags and local information tags and their corresponding image and point cloud data addresses are saved in the search server.

[0087] In this embodiment, ES (Elasticsearch, search server) is an open source distributed search and analysis engine built on Apache Lucene. It provides a powerful full-text search engine that can handle large-scale data and supports real-time search and analysis. Its distributed architecture allows users to distribute data on multiple nodes to achieve high availability and scalability. It is usually used in scenarios such as log analysis, full-text search, and real-time data screening.

[0088] In this embodiment, the global information labels and local information labels, and their corresponding image and point cloud data addresses are stored in ES, and global and local information matching queries, precision queries, range queries, and Boolean queries are performed as needed to improve the efficiency of data mining.

[0089] The present invention proposes a method for mining high-value sample data for autonomous driving. For the massive collected data samples of autonomous driving, a large multimodal model is introduced to understand the scene, distinguish known obstacle types from unknown obstacle types, select badcases (abnormal events) and hardcases (difficult to handle events) in the autonomous driving system, and then perform special labeling. The labeled data is added to the autonomous driving model training, and the autonomous driving model is continuously iterated to improve the stability and robustness of the autonomous driving system.

[0090] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a program running on the processor, and the processor executes the steps of the above-mentioned method for mining high-value sample data for autonomous driving when running the program.

[0091] The present invention also provides a computer-readable storage medium having computer instructions stored thereon. When the computer instructions are executed, the steps of the above-mentioned method for mining high-value sample data for autonomous driving are executed. The method for mining high-value sample data for autonomous driving has been introduced in the aforementioned part and will not be described in detail here.

[0092] Those skilled in the art can understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions recorded in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for mining high-value sample data for autonomous driving, characterized in that: include: Obtain the data collected by the vehicle, perform data analysis and data cleaning, and then save it to the data center; Based on the data center and the scene understanding data set, a multimodal large model of autonomous driving is trained to perform global scene understanding to obtain global scene understanding information, and the global scene understanding information is saved as a global information tag to the search server; Training a target detection model based on the data center and the target detection data set to obtain local known category information and target detection results, and saving the local known category information as a local known information label to the search server; Performing unsupervised point cloud clustering based on the Euclidean distance between points in the point cloud data of the data center to obtain a clustering result list; The point cloud clustering result is projected into the image, the projection result on the image is matched with the target detection result, and the local unknown category information label is obtained and saved in the search server.

2. The method for mining high-value sample data for autonomous driving according to claim 1, characterized in that: The acquisition of data collected by the vehicle end and the data parsing and cleaning and then saving to the data center include: Collect vehicle-side data based on vehicle-mounted systems and sensors, and convert the collected vehicle-side data into the same format; Identify missing values ​​in vehicle-side data through data cleaning, and remove or complete missing values ​​based on the missing mechanism; Deviations in sensor data or errors in vehicle-side data entry can be identified and corrected through a calibration process or by comparing with other data sources.

3. The method for mining high-value sample data for autonomous driving according to claim 1, characterized in that: The clustering result list obtained by performing unsupervised clustering of point cloud based on the Euclidean distance between points in the point cloud data of the data center includes: If the distance between any two points in the point cloud data is less than the set threshold, the two points will be clustered into one category.

4. The method for mining high-value sample data for autonomous driving according to claim 1 or 3, characterized in that: The clustering result list obtained by performing unsupervised clustering of point cloud based on the Euclidean distance between points in the point cloud data of the data center includes: Create an empty result list to store clustering results; Initialize the flag array to mark whether the point has been processed; Build a spatial index structure, traverse a single point in the point cloud data, and start clustering if the point has not been processed; Select a single unprocessed point as the starting point, find all neighboring points within the set threshold distance, mark them as processed and add them to the current cluster until all neighboring points are added to the cluster; Continue to traverse the unprocessed points until all points in the point cloud data are processed, and return the clustering result list.

5. The method for mining high-value sample data for autonomous driving according to claim 1, characterized in that: The projecting of the point cloud clustering results into the image includes: projecting the laser radar data onto the image through the internal and external parameter relationship between the camera and the laser radar, and after the image and the laser radar are synchronized in time, a single image corresponds to the laser point cloud at the same moment.

6. The method for mining high-value sample data for autonomous driving according to claim 5, characterized in that: Also includes, The point cloud data and image data of the LiDAR are collected for calibration to obtain the internal and external parameters of the camera and LiDAR.

7. The method for mining high-value sample data for autonomous driving according to claim 6, characterized in that: include, Calculate the camera intrinsic parameters through several chessboard images and store the camera intrinsic parameters; Collect images and point cloud data from the calibration room, and project the point cloud into the image according to the parameter adjustment parameters until the point cloud and image coincide to obtain the external parameter parameters.

8. The method for mining high-value sample data for autonomous driving according to claim 1 or 5, characterized in that: The matching of the projection result on the image with the target detection result to obtain the local unknown category information label and save it to the search server includes: on the image, taking the target whose target detection result matches the point cloud clustering result as the local known obstacle type, and taking the target whose point cloud clustering result does not match the target detection result as the unknown obstacle type.

9. An electronic device, characterized in that: It includes a memory and a processor, wherein the memory stores a program running on the processor, and when the processor runs the program, it executes a method for mining high-value sample data for autonomous driving as described in any one of claims 1-8.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed, a method for mining high-value sample data for autonomous driving as described in any one of claims 1 to 8 is executed.