Automatic driving data set management system and method based on document database

By adopting a hierarchical data architecture and hybrid database based on document databases in the autonomous driving dataset management system, combining multi-threading technology and Bayesian estimation algorithm, the inlet and retrieval speed problems in autonomous driving dataset management and storage are solved, and efficient data processing and storage efficiency are achieved.

CN120123294APending Publication Date: 2025-06-10SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510194187.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently manage and store autonomous driving data sets, especially when processing binary and metadata information, and traditional databases have shortcomings in inbound and retrieval speeds.

Method used

The hierarchical data architecture and hybrid database based on document database are adopted to store data through MongoDB, and HDF5 files are written using improved formats and algorithms during the data acquisition stage. Combining multi-threading technology and Bayesian estimation algorithms, the data retrieval and database entry process is optimized.

Benefits of technology

It significantly improves data processing and storage efficiency, reduces database inlet time and file extraction time, and is about twice as fast as the database inlet speed and retrieval speed compared with the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123294A_ABST
    Figure CN120123294A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving data set management system and method based on a document database, and the system comprises a data collection platform, a data processing platform and a data product platform, and the data collection platform obtains original data through an embedded database, a data collection vehicle and a multi-mode sensor, and outputs the original data to the data processing platform; the data processing platform performs data synchronization on the original data and outputs an obtained data set with a truth value to a data product platform; the data product platform displays and retrieves the processed data set while previewing and downloading the structure and content of the data set. According to the method, through the MongoDB-based hierarchical data architecture and the hybrid database, the HDF5 file is written in by adopting an improved format and algorithm in a data acquisition stage, and the database storage time and the file extraction time are remarkably reduced through an improved data retrieval and high-speed storage method, so that the data processing and storage efficiency is improved. The system is particularly suitable for storing and managing a large amount of sensor data and status data generated in an automatic driving process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of autonomous driving, specifically a management system and method for autonomous driving data sets based on a document database. Background Art

[0002] The defects existing in the formats of current data sets are mainly in two aspects. First, as the size of autonomous driving data sets continues to increase, in many cases, researchers only want to conduct targeted training in some scenarios to improve the generalization of the model. In this way, it is inefficient to download and train an entire data set completely each time. The industry needs a platform that can have a database of data sets for customized retrieval and download. Second, current data sets are all built based on offline platforms, which are not convenient for large-scale management and storage. The industry needs to design a data set storage database that can be deployed in the cloud to manage a large number of data sets, and at the same time have efficient data extraction and warehousing capabilities. Summary of the Invention

[0003] In view of the deficiencies that it is difficult for the prior art to simultaneously process binary and metadata (MetaData) information generated during the autonomous driving process, and a large amount of hardware resources will be consumed for format conversion and storage after preprocessing, the present invention proposes a management system and method for autonomous driving data sets based on a document database. Through a hierarchical data architecture and a hybrid database based on MongoDB, an improved format and algorithm are used to write HDF5 files during the data collection stage, and the database warehousing time and file extraction time are significantly reduced through improved data retrieval and high-speed warehousing methods, thereby improving data processing and storage efficiency. This system is particularly suitable for storing and managing a large amount of sensor data and status data generated during the autonomous driving process.

[0004] The present invention is realized through the following technical solutions:

[0005] The present invention relates to a management system for autonomous driving data sets based on a document database, including: a data collection platform, a data processing platform, and a data product platform, wherein: the data collection platform obtains raw data through an embedded database, a data collection vehicle, and multi-modal sensors and outputs it to the data processing platform; the data processing platform synchronizes the raw data and outputs the data set with ground truth to the data product platform; the data product platform displays and retrieves the processed data sets, and at the same time previews and downloads the structure and content of the data sets.

[0006] The original data described above includes: binary data in YUV format of the camera serial port, message packets broadcast by the lidar to the local area network, PCD format files generated after drive unpacking sent by the millimeter-wave radar to the host, and GPS standard packets sent through the local area network after GPS communicates with satellites.

[0007] The data processing platform includes: a data preprocessing module, a visualization module, a data processing pipeline module, a data evaluation module, and a dataset service module. Among them: the data preprocessing module converts the data into a format suitable for model training and then outputs it to the visualization module and the data processing pipeline module respectively for storage; the visualization module performs Three.js visualization processing of HTML based on the collected sensor metadata information to obtain results such as 3D stereoscopic point clouds, GPS coordinate trajectories, and camera stitched BEVs; the data processing pipeline module performs labeling and object detection based on the collected sensor metadata and the metadata information generated by the preprocessing module to obtain an h5 file containing labels and bounding boxes and outputs it to the data evaluation module; the data evaluation module tests and evaluates the h5 file, generates a complete dataset with ground truth, and then outputs it to the data product platform.

[0008] The data collection platform includes: a data collection vehicle (Vehicle) and a cloud server (Cloud Server). Among them: the data collection vehicle collects the original data through a data collection board (Development Board), packs it into an HDF5 file, and communicates with the Cloud Server through the MQTT protocol in the local area network to upload the collected original data to the cloud server; the cloud server classifies the received HDF5 files and stores them in MongoDB. Technical Effects

[0009] The present invention aims to solve the problems of the data format of the autonomous driving dataset and the storage and collection speed of the database. After the database is built, indoor data collection simulation tests are carried out in this application, and the entire test process passes. In addition, this application saves the HDF5 files for a set of comparative experiments, which mainly compares the storage speed and retrieval speed of the database built in this application with the currently popular Influx DB. Compared with the prior art, the feasibility of the present invention is verified through experiments, and the storage speed and retrieval are improved compared with other traditional database architectures. Brief Description of the Drawings

[0010] Figure 1 It is a schematic structural diagram of the system of the present invention;

[0011] Figure 2 It is a schematic diagram of system module interaction;

[0012] Figure 3 It is a schematic diagram of the HDF5 file structure;

[0013] Figure 4 It is a comparison chart of database retrieval efficiency and visualization;

[0014] Figure 5 It is a screenshot of the web side of the data visualization platform. Specific implementation manners

[0015] As Figure 1 shown, this embodiment relates to an autonomous driving dataset management system based on a document database, including: a data collection platform, a data processing platform, and a data product platform, where: the data collection platform collects multi-modal sensor raw data from a data acquisition vehicle through external devices and an embedded database and outputs it to the data processing platform; the data processing platform sequentially performs data collection, packing and compression, processes sensor data and performs data synchronization, and processes and stores rich sensor information and feedback on the collected raw data, and outputs the obtained dataset with Ground Truth to the data product platform; the data product is used to display and retrieve the processed dataset, and at the same time preview and download the structure and content of the dataset.

[0016] As Figure 2 shown, this embodiment is a method for managing an autonomous driving dataset based on a document database based on the above system, including:

[0017] Step 1. In the data acquisition stage, use the HDF5 file as shown in Figure 3 to temporarily pack and store the sensor data of each frame.

[0018] Step 2. Use multiple threads to write to the HDF5 file, specifically including:

[0019] Thread 1: Open the serial port and synchronize the time data to ensure that all data has a consistent timestamp.

[0020] Thread 2: Write the MetaData of the dataset, including data such as acquisition time, acquisition location, and vehicle used. At the same time, according to the sensor-related data passed in, establish a data space for the camera, lidar, millimeter wave radar, and combined inertial navigation (GNSS).

[0021] Thread 3-11: A total of 9 threads are created, and the threads are allocated according to the incoming data volume. The lidar has 3 threads, GNSS uses 1 thread, Camera uses 4 threads, and the millimeter-wave radar uses 1 thread to write data. Among them, the camera stores matrix data in the original YUV format, and the lidar and millimeter-wave radar store point cloud data, such as

[0022] Thread 12: Create an 8G memory space to cache the unprocessed data.

[0023] Step 3. In the data processing stage, use MongoDB to save the file paths of the original data set, and use the HDF5 file to save the MetaData to form a MongoDB+HDF5 database and perform real-time preview on the original data set through the preloading and extraction algorithms based on the Bayesian estimation algorithm, specifically including:

[0024] 3.1 Initialization: Set the data index set available for preview as X, and let the preview duration set of the preview data be T. Construct Gaussian sampling points from X and T, and establish a Gaussian model represents the predicted preview data index, represents the predicted preview data duration.

[0025] 3.2 Establish a Bayesian joint probability model Among them: p(x,τ|θ) represents the joint probability density function of the data index x and the data preview duration τ under the condition of θ, and the parameter θ is determined by the mean μ x , μ τ , variance σ x , σ τ . And are the marginal probability density functions of the corresponding data index and data preview market.

[0026] 3.3 Calculate the Gaussian-Gamma prior function: Among them: the hyperparameter m 0 represents the mean parameter of the normal distribution function reflecting the mean μ x , the hyperparameter κ 0 represents the variance parameter of the normal distribution of the mean. Correspondingly, the hyperparameters a 0 , b 0 represent the mean and variance of the inverse gamma distribution of the variance σ x . The hyperparameters m' 0 κ 0 ' and a 0 , b 0 are the same by analogy.

[0027] 3.4 When the user performs a preview, the data will be pre-loaded and extracted.

[0028] The extraction satisfies: the predicted index model according to the index set X and the duration T and Extract the set composed of the 1-σ points at the peak of the Gaussian model distribution in the index set X and the duration T and According to the designed sensitivity policy η (representing the minimum duration that needs to be pre-loaded, and those less than this value will not be pre-loaded), allocate multiple threads to perform the pre-loading function Load the preview content (video cache) into the memory and wait for the user to click on the corresponding data index.

[0029] 3.5 After the user finishes clicking the index - performing data preview - and preview ends, record the data index x clicked this time k+1 and the preview duration τ k+1 And perform model update, specifically including:

[0030] a) Update the two predicted Gaussian models using the Bayesian hierarchical framework: Among them: the introduced hyperparameters m 0 , κ 0 and a 0 , b 0 The updates of all are called the python multi-point fitting distribution algorithm.

[0031] b) Similarly, update the preview time: Among them: similarly, the updates of the hyperparameters are all called the multi-point fitting probability distribution algorithm.

[0032] Step 4. In the dataset visualization stage, when receiving a retrieval or visualization command, start multiple threads to process different sensor data and MetaData data respectively. And use another thread to perform pre-loading according to the timestamp of the retrieved dataset, and finally preview the dataset images through the web viewer to achieve the preview of the original dataset.

[0033] After specific actual experiments, under the specific environmental settings of having local original sensor data packets, locally started MongoDB and MQTT protocol network environments, and locally started HTML visualization dataset platforms, the above-mentioned data storage method and timestamp retrieval preloading method are run with the HDF5 file parameters of a 10s sample sensor data packet (including 100x3 lidar data, 240x8 camera data, 100 GPS data, 300x8 millimeter-wave radar data). Data is collected for data storage operations, and the data storage speed, sequential reading speed (reading into memory, the same as below), random reading speed, reverse reading speed, and the time of the timestamp retrieval and reading process of the entire database (TimeStamp Search and Load) are detected respectively, as shown in Table 1.

[0034] Table 1 Search method MongoDB Patent method Sequential reading 0.1719 0.0359 Random reading 0.0334 0.0229 Reverse reading 0.0809 0.0280 Data writing 1.2891 0.9595 Timestamp retrieval 0.2053 0.0589

[0035] The detection method for the time of timestamp retrieval and reading in Table 1 is the time from the visualization API of the corresponding Web segment to the successful loading of the visualization content. The detection environment is a local area network to eliminate the influence of the network.

[0036] Compared with the prior art, the present invention uses multi-thread technology to accelerate the writing and sequential reading of the database, reducing the data storage time, which is about twice as fast as the widely used Influx DB at present, as shown in Table 2.

[0037] Table 2

[0038] In summary, the present invention models the user's click preview behavior by introducing a two-layer Bayesian network, predicts the user's preview habits, obtains an adaptive loading index range and preloading time, while accelerating the preview opening speed of large files such as videos, and also avoids overloading the system, resulting in a speed improvement of nearly 3 times compared to MondoDB in the timestamp retrieval and preview stages.

[0039] The above specific implementation can be locally adjusted by those skilled in the art in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the constraints of the present invention.

Claims

1. An autonomous driving data set management system based on a document database, characterized in that: include: Data collection platform, data processing platform and data product platform, where: the data collection platform obtains raw data through embedded database, data collection vehicle and multimodal sensor and outputs it to the data processing platform; the data processing platform synchronizes the raw data and outputs the obtained data set with true value to the data product platform; the data product platform displays and retrieves the processed data set, and previews and downloads the structure and content of the data set; The raw data includes: binary data in YUV format from the camera serial port, message packets broadcast by the laser radar to the local area network, PCD format files generated by the driver unpacking and sent by the millimeter wave radar to the host, and GPS standard packets sent through the local area network after GPS communicates with the satellite.

2. The autonomous driving data set management system based on a document database according to claim 1, characterized in that: The data processing platform includes: a data preprocessing module, a visualization module, a data processing pipeline module, a data evaluation module and a data set service module, wherein: the data preprocessing module converts the data into a format suitable for model training and outputs it to the visualization module and the data processing pipeline module respectively for storage; the visualization module performs HTML Three.js visualization processing based on the collected sensor metadata information to obtain 3D point cloud, GPS coordinate trajectory, and camera stitching BEV results; The data processing pipeline module performs labeling and target recognition based on the collected sensor metadata and metadata information generated by the preprocessing module, obtains the h5 file containing the labeling and detection box, and outputs it to the data evaluation module; the data evaluation module tests and evaluates the h5 file, generates a complete data set with true values, and outputs it to the data product platform.

3. The autonomous driving data set management system based on a document database according to claim 1, characterized in that: The data collection platform includes: a data collection vehicle and a cloud server, wherein: the data collection vehicle collects raw data through a data collection board and packages it into an HDF5 file, communicates with the Cloud Server through the MQTT protocol on the local area network, and uploads the collected raw data to the cloud server; the cloud server classifies the received HDF5 files and stores them in MongoDB.

4. A method for managing an autonomous driving data set based on the system according to any one of claims 1 to 3, characterized in that: include: Step 1: During the data collection phase, the HDF5 file shown in FIG3 is used on the data collection vehicle to temporarily package and store each frame of sensor data; Step 2: Use multiple threads to write HDF5 files; Step 3: In the data processing stage, use MongoDB to save the file path of the original data set, use HDF5 files to save MetaData, store it in the database to form a MongoDB+HDF5 database, and use the preloading and extraction algorithm based on the Bayesian estimation algorithm to preview the original data set in real time; Step 4: In the dataset visualization stage, when a retrieval or visualization command is received, multiple threads are started to process different sensor data and MetaData data respectively, and another thread is used to preload according to the timestamp of the retrieved dataset. Finally, the dataset is previewed through the web viewer to realize the preview of the original dataset.

5. The method for managing an autonomous driving dataset according to claim 4, wherein: The multiple threads specifically include: Thread 1: Open the serial port, synchronize time data, and ensure that all data has consistent timestamps; Thread 2: Write the MetaData of the dataset, including the collection time, collection location, vehicle used, and other data. At the same time, according to the incoming sensor-related data, establish the data space (DataSpace) of the camera (Camera), lidar (Lidar), millimeter-wave radar (Radar) and combined inertial navigation (GNSS); Threads 3-11: Threads are allocated according to the amount of data passed in. LiDAR has 3 threads, GNSS uses 1 thread, Camera uses 4 threads, and millimeter-wave radar uses 1 thread to write data. The camera stores matrix data in the original YUV format, while LiDAR and millimeter-wave radar store point cloud data. Thread 12: Create an 8G memory space to cache unprocessed data.

6. The method for managing an autonomous driving dataset according to claim 4, wherein: The step 3 specifically includes: 3.1 Initialization: Set the index set of data available for preview to X, set the preview duration set of the previewed data to T, use X and T to form Gaussian sampling points, and establish a Gaussian model Represents the index of the preview data for the prediction. Represents the duration of the predicted preview data; 3.2 Establishing Bayesian Joint Probability Model Where: p(x,τ|θ) represents the joint probability density function of data index x and data preview duration τ under the condition of θ, and the parameter θ is the mean μ x ,μ τ , variance σ x ,σ τ Decide, and is the marginal probability density function of the corresponding data index and data preview market; 3.3 Calculate the Gaussian-Gamma prior function: Where: The hyperparameter m0 represents the mean μ of the reaction pair x The mean parameter of the normal distribution function, the hyperparameter κ0 represents the variance parameter of the normal distribution with respect to the mean, and correspondingly, the hyperparameters a0 and b0 represent the variance σ x The mean and variance of the inverse gamma distribution of the preview time are similar to the designed hyperparameters m'0κ0' and a0, b0; 3.4 When the user previews, the data will be preloaded and extracted; 3.5 When the user clicks the index, previews the data, and then records the data index x of this click. k+1 and the preview duration τ k+1 And update the model.

7. The method for managing an autonomous driving dataset according to claim 6, wherein: The extraction satisfies: According to the index set X and the duration T, the predicted index model and Extract the Gaussian model distribution in the index set X and the set of 1-σ points of the peak of duration T and According to the designed sensitivity strategy η (representing the minimum length of preloading, and no preloading is performed if it is less than this value), multiple threads are allocated to preload functions. Load the previewed content (video cache) into memory and wait for the user to click the corresponding data index.

8. The method for managing an autonomous driving dataset according to claim 6, wherein: The step 3.5 specifically includes: a) The two predicted Gaussian models are updated by Bayesian hierarchical framework: Among them: the updates of the introduced hyperparameters m0, κ0, a0, b0 all call the python multi-point fitting distribution algorithm; b) Similarly, update the preview time: Among them: Similarly, the update of hyperparameters all calls the multi-point fitting probability distribution algorithm.