Data set construction method and device, equipment and storage medium
Through pre-set data set production configuration files and automated processing processes, candidate data are obtained from intelligent driving data and compliance checks are carried out to build intelligent driving data sets of different purposes, solving the problem of building high-quality data sets in the existing technology, and achieving rapid construction and flexible adjustment of data sets.
Patent Information
- Application Number
- CN202510395124.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult for the prior art to efficiently build intelligent driving data sets that meet the needs and have high-quality characteristics from massive data.
Through the pre-set data set production configuration files, candidate data is obtained from the intelligent driving data, and after compliance checks are performed, different intelligent driving data sets are built using the data tags and file paths.
It realizes the rapid construction of intelligent driving data sets under custom conditions, supports the construction of any version of the basic data set, has the ability to expand functions, and can adapt to technological development and demand changes.
Smart Images

Figure CN120336397A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent driving technology, and particularly to a method, apparatus, device, and storage medium for constructing a data set. Background Art
[0002] In the field of intelligent driving, due to the extremely complex traffic conditions and diverse scenarios, a large number of models have been introduced in various fields (such as perception, mapping, planning and control, etc.) to solve problems. To achieve better results, the training of models requires a vast amount of data. For different problems, different models need to be developed, which requires quickly constructing a data set related to such problems. Even when training the same model, sometimes in order to improve the performance of the model in a specific scenario, data related to that specific scenario needs to be added.
[0003] Generally, the data for intelligent driving is obtained by the data collection vehicle or uploaded by users of mass-produced models already on the market. A series of operations such as data parsing, data extraction, data analysis, data label mining, data key obstacle mining, and data preprocessing are performed on the original data to obtain the input data required by the model. However, at the practical operation level, how to efficiently construct an intelligent driving data set that meets the requirements and has high-quality characteristics from a vast amount of data is still a major problem faced by the industry. Summary of the Invention
[0004] The main purpose of this application is to provide a method, apparatus, device, and storage medium for constructing a data set, aiming to solve the technical problem of how to efficiently construct an intelligent driving data set that meets the requirements and has high-quality characteristics from a vast amount of data.
[0005] To achieve the above purpose, this application proposes a method for constructing a data set, the method including:
[0006] Obtain candidate data from intelligent driving data according to a preset data set production configuration file;
[0007] Construct different intelligent driving data sets using data labels and the file paths of the candidate data.
[0008] In an embodiment, the step of obtaining candidate data from intelligent driving data according to a preset data set production configuration file includes:
[0009] Extract or filter the intelligent driving data according to a preset data set production configuration file to obtain candidate data;
[0010] Perform compliance checks on the candidate data to obtain the candidate data.
[0011] In one embodiment, the intelligent driving data includes production vehicle data. The step of extracting or filtering the intelligent driving data according to a pre-set data set production profile to obtain candidate data includes:
[0012] Configure the storage location of the production vehicle data according to the data production method;
[0013] Extract the production vehicle data that meets the data time range from the storage location;
[0014] Extract or filter the production vehicle data that meets the data time range according to the event type and vehicle information to obtain candidate data.
[0015] In one embodiment, the step of extracting or filtering the production vehicle data that meets the data time range according to the event type and vehicle information to obtain candidate data includes:
[0016] Check whether the event type associated with the production vehicle data that meets the data time range is in the preset event type list;
[0017] If the event type is not in the preset event type list, filter the production vehicle data corresponding to the event type;
[0018] If the event type is in the preset event type list, extract the production vehicle data corresponding to the event type as the initial candidate data;
[0019] Check whether the vehicle information associated with the initial candidate data is in the preset vehicle number list;
[0020] If the vehicle information is not in the preset vehicle number list, filter the production vehicle data corresponding to the vehicle information;
[0021] If the vehicle information is in the preset vehicle list, extract the production vehicle data corresponding to the vehicle information as the final candidate data.
[0022] In one embodiment, the intelligent driving data includes collection vehicle data. The collection vehicle data is stored in the form of collection vehicle data packages. The step of extracting or filtering the intelligent driving data according to a pre-set data set production profile to obtain candidate data further includes:
[0023] Query the database to obtain the name list of the collection vehicle data packages within the data time range and the vehicle numbers corresponding to the name list;
[0024] Extract or filter the preset vehicle number list according to the vehicle number to obtain candidate data; or
[0025] Extract or filter the data of the acquisition vehicle according to the bucket name to obtain candidate data.
[0026] In one embodiment, the step of extracting or filtering the preset vehicle number list according to the vehicle number to obtain candidate data includes:
[0027] Check whether the vehicle number is in the preset vehicle number list;
[0028] If the vehicle number is not in the preset vehicle number list, filter the data of the acquisition vehicle corresponding to the vehicle number;
[0029] If the vehicle number is in the preset vehicle number list, extract the data of the acquisition vehicle corresponding to the vehicle number as the final candidate data.
[0030] In one embodiment, before the step of obtaining candidate data from the intelligent driving data according to the preset data set production configuration file, the following steps are further included:
[0031] Define the data set production configuration file, which includes multiple configuration subtasks. The configuration subtasks include data set production methods, data set compliance check configurations, data time intervals, data labels, event types, vehicle filtering information, and data sources. The data sources include mass-produced vehicle data and acquisition vehicle data.
[0032] In one embodiment, the step of performing a compliance check on the candidate data to obtain candidate data includes:
[0033] Based on the data set compliance check configuration, perform an integrity check on the candidate data;
[0034] When the candidate data passes the integrity check, perform a compliance check on the map file, lane change file, and label file corresponding to the candidate data to obtain candidate data.
[0035] In one embodiment, the data label includes a scene-level label and a target-level label. The step of constructing different intelligent driving data sets using the data label and the file path of the candidate data includes:
[0036] Use the scene-level label and the target-level label to group the candidate data in different dimensions;
[0037] In the grouping in different dimensions, construct a data set according to the file path of the candidate data to obtain different intelligent driving data sets.
[0038] In addition, to achieve the above object, the present application also proposes a data set construction device, which includes:
[0039] A processing module, configured to generate a configuration file according to a preset data set and obtain candidate data from intelligent driving data;
[0040] A construction module, configured to use data tags and file paths of the candidate data to construct different intelligent driving data sets.
[0041] In addition, to achieve the above object, the present application further provides a data set construction device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the data set construction method as described above.
[0042] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the data set construction method as described above are implemented.
[0043] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the data set construction method as described above are implemented.
[0044] An embodiment of the present application provides a data set construction method, device, device, and storage medium. The method includes generating a configuration file according to a preset data set, obtaining candidate data from intelligent driving data; using data tags and file paths of the candidate data to construct different intelligent driving data sets. This solution can quickly construct an intelligent driving data set under custom conditions through a preset configuration file and an automated processing process. And it supports constructing any version of the basic data set to meet different requirements. At the same time, it has the ability to expand functions and can be adjusted and optimized as technology develops and requirements change. Description of the Drawings
[0045] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application, and are used together with the description to explain the principles of the present application.
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the data set construction method of the present application;
[0048] Figure 2 Schematic diagram of the candidate data acquisition process provided in the first embodiment of this application;
[0049] Figure 3 Schematic diagram of the process provided in the second embodiment of the dataset construction method of this application;
[0050] Figure 4 Schematic diagram of the relationship between the configuration file and the configuration subtasks provided in the second embodiment of this application;
[0051] Figure 5 Schematic diagram of the parameter configuration of the configuration subtasks provided in the second embodiment of this application;
[0052] Figure 6 Schematic diagram of the brief process of the dataset construction method provided in the first and second embodiments of this application;
[0053] Figure 7 Schematic diagram of the module structure of the dataset construction device in the embodiment of this application;
[0054] Figure 8 Schematic diagram of the device structure of the hardware operating environment involved in the dataset construction method in the embodiment of this application.
[0055] The realization of the purpose, functional characteristics and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0056] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0057] In order to better understand the technical solutions of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0058] The main solution of the embodiment of this application is: according to the preset dataset production configuration file, extract or filter the obtained intelligent driving data to obtain candidate data; based on the preset dataset production configuration file, perform compliance checks on the candidate data to obtain candidate data; use the data tags and the file paths of the candidate data to construct different intelligent driving datasets.
[0059] Generally, the data of intelligent driving is obtained by the acquisition vehicle or uploaded by the users of the mass-produced models already on the market. A series of operations such as data parsing, data extraction, data analysis, data label mining, data key obstacle mining, and data preprocessing are performed on the original data to obtain the input data required by the model. However, at the practical operation level, how to efficiently construct an intelligent driving dataset that meets the requirements and has high-quality characteristics from a large amount of data is still a major problem faced by the industry.
[0060] This application provides a solution. According to a pre-set dataset production configuration file, intelligent driving data obtained is extracted or filtered to obtain candidate data; based on the pre-set dataset production configuration file, compliance checks are performed on the candidate data to obtain candidate data; data tags and the file paths of the candidate data are used to construct different intelligent driving datasets. This solution can significantly improve the performance and reliability of the intelligent driving system by constructing diverse high-quality datasets, laying a solid foundation for the future development of autonomous driving technology.
[0061] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions. Hereinafter, taking a personal computer as an example, this embodiment and the following embodiments will be described.
[0062] Based on this, the embodiments of this application provide a method for constructing a dataset, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the dataset construction method of this application.
[0063] In this embodiment, the dataset construction method includes steps S10 to S20:
[0064] Step S10, obtain candidate data from the intelligent driving data according to a pre-set dataset production configuration file;
[0065] It should be noted that the dataset production configuration file is a document or a set of rules for guiding how to extract useful information from the originally collected data to form a dataset suitable for a specific purpose (such as model training). The dataset production configuration file includes, but is not limited to, data source definitions (production vehicle data, collected vehicle data), vehicle filtering information, generation methods (online or offline), validity verification, concurrency, data version, data time interval, tags, and event types specific to production vehicles and other parameter configurations.
[0066] Candidate data is the part of the original data that is considered to meet the research or application requirements after preliminary screening and may be further processed (such as labeled) and finally added to the formal training dataset.
[0067] It can be understood that since the original intelligent driving data may contain a large amount of noise, irrelevant information, or data points that do not meet the research requirements, performing step S10 can avoid directly using low-quality or data that does not meet the current task requirements to participate in model training, thereby improving the model training efficiency and the performance of the final model.
[0068] In an alternative embodiment, step S10 may further include steps S11 to S12:
[0069] Step S11, extracting or filtering the intelligent driving data according to a pre-set data set production profile to obtain candidate data;
[0070] Optionally, please refer to Figure 2 , Figure 2 for the schematic diagram of the candidate data acquisition process. In this embodiment, according to different data sources and data generation methods, the data storage methods are also different, and the methods for obtaining candidate data will also be different.
[0071] In a feasible embodiment, step S10 may include steps A11 to A13:
[0072] Step A11, configuring the storage location of the production vehicle data according to the data production method;
[0073] For the production vehicle data, according to the data generation method, such as whether the production vehicle data is generated online or offline, configure the storage path of the production vehicle data.
[0074] Specifically, for the production vehicle data generated online, the object storage service provided by a cloud service provider is usually used. And set up an upload mechanism to ensure the real-time transmission of the production vehicle data; define the production vehicle data retention policy, such as archiving old data regularly; implement strict access control measures, etc.
[0075] For the production vehicle data generated offline, the vehicle first uses a local storage device (such as an in-vehicle hard disk). When conditions permit, the data is synchronized to the cloud or data center through the network. And ensure there is enough local storage space; set up data compression and encryption mechanisms to save bandwidth and protect privacy; establish a reliable data synchronization mechanism to support the resume function.
[0076] Through the above steps, determine the most suitable storage scheme according to the data generation method (online or offline) to ensure that the data can be saved and accessed efficiently and securely.
[0077] Step A12, extracting the production vehicle data that meets the data time interval from the storage location according to the data time interval;
[0078] Specifically, first determine the production vehicle data of the specific date and time range to be extracted, and then use the database query language or API interface to filter out the production vehicle data that does not meet the determined specific date and time range of extraction according to the timestamp field, and export the qualified data to a temporary folder or other easily accessible locations for subsequent operations.
[0079] Step A13: Extract or filter the production vehicle data within the data time range according to the event type and vehicle information to obtain candidate data.
[0080] It should be noted that the event type refers to the classification related to specific situations or behaviors that occur during the operation of the vehicle. Vehicle information refers to a series of attributes and status data related to the vehicle itself. Among them, vehicle information includes but is not limited to vehicle identifiers, specific models and configurations of the vehicle, production dates, version numbers, etc.
[0081] Specifically, first, determine a list that includes all event types of interest. Check each intelligent driving data record that meets the time range, and compare the event type in each record with the preset event type list. If the event type in each record is not in the preset event type list, then filter out this production vehicle data. If the event type in each record is in the preset event type, then retain this production vehicle data as the initial candidate data.
[0082] Then, determine a vehicle list that includes all key vehicle identifiers. Check the obtained initial candidate data, and compare the vehicle identifier in each initial candidate data with the preset vehicle list. If the vehicle identifier in the initial candidate data is not in the preset vehicle list, then filter out this initial candidate data. If the vehicle identifier in the initial candidate data is in the preset vehicle list, then retain this initial candidate data as the final candidate data.
[0083] In another feasible implementation manner, step S10 may include steps B11 to B13:
[0084] Step B11: According to the data time range, query the database to obtain the name list of the collected vehicle data packets within the data time range and the vehicle numbers corresponding to the name list;
[0085] It should be noted that for the collected vehicle data, the collection work of the collected vehicle data usually lasts for a long time. Therefore, the storage method of the collected vehicle data is stored in units of large packets. And in order to facilitate the processing and analysis of the collected vehicle data and quickly obtain segment data, the large packets are usually cut into several small packets at equal intervals for storage.
[0086] Specifically, first, judge the production method of the collected vehicle data. If it is the online production method, then according to the time range of the data, query the database through the query language to obtain all the large packet names (large packet IDs) and the corresponding vehicle numbers in this time range.
[0087] Step B12: Extract or filter the preset vehicle number list according to the vehicle number to obtain candidate data;
[0088] Specifically, first, determine a vehicle number list that includes all the vehicle numbers of interest, and then compare the vehicle numbers corresponding to the large packets obtained in the above steps with the vehicle number list. If it is detected that the vehicle number is in the vehicle number list, retain the collected vehicle data corresponding to the vehicle number and use it as candidate data. If it is detected that the vehicle number is not in the vehicle number list, filter the collected vehicle data corresponding to the vehicle number.
[0089] Step B13: Extract or filter the collected vehicle data through the bucket name to obtain candidate data.
[0090] Since the offline data storage of the collection vehicle is generally bucketed by vehicle number and date, if the production mode of the collection vehicle data is the offline production mode, directly filter or extract the data through the bucket name (vehicle number, date) of the bucket. Specifically, first, determine the name corresponding to the bucket to which each piece of collected vehicle data is assigned. If the bucket name is in the preset bucket name list, retain this piece of data as candidate data. If it is not in the preset list, ignore (filter out) this piece of data.
[0091] Through the above steps, data that meets the requirements of time and vehicle number can be effectively extracted from the massive collected vehicle data, and data that meets the preset conditions can also be effectively extracted from the offline stored collected vehicle data, providing a high-quality data basis for subsequent data analysis and model training.
[0092] Step S12: Perform compliance checks on the candidate data to obtain candidate data.
[0093] Since there may be file anomalies and other situations during the processes of online large-scale data production, transmission, storage, etc. And the data sources are diverse, and some data may not meet the requirements of the required data set. Therefore, in this embodiment, performing compliance checks on the candidate data can effectively avoid the occurrence of the above problems.
[0094] In this embodiment, the compliance check is carried out according to the corresponding check items of the compliance check configuration in the pre-set data set production configuration file, which is mainly divided into data integrity check, data anomaly check, and existence check of necessary files.
[0095] Specifically, during the data production, transmission, and storage processes, due to anomalies occurring in different situations, the data files of candidate data are ultimately saved incompletely, resulting in anomalies when reading the files later. Therefore, integrity checks are performed on the candidate data, including but not limited to checking the size of the data files of the candidate data, comparing the actual size of the data files with the expected size in the record (e.g., obtained through metadata or log files); verifying whether the headers and tails of the files contain correct identifiers or end markers; verifying whether the internal structure of the files conforms to expectations; verifying whether the data in the files is consistent and free of obvious errors; verifying whether there are duplicate data files, etc. If it is found through the checks that the data files of the candidate data are incomplete, the candidate data is discarded, and at the same time, the incomplete count in the result statistics dictionary file is incremented by 1.
[0096] Map data is important information for the intelligent driving dataset. The abnormal absence of map data will have an adverse impact on model training. In this embodiment, the map data check in the candidate data is used as a basic check when constructing the dataset to ensure the validity of the map data. Among them, the map compliance check mainly includes:
[0097] (1) There is map data missing after a certain number of starting frames (allowing map data to be missing in the starting several frames): Read the first N frames of the data file, check whether the map data in these frames is missing, and starting from the (N + 1)-th frame, check whether the map data exists in each frame.
[0098] (2) Existence check of the map file: Obtain the path of the map file and use the file system API to check whether the map file exists.
[0099] (3) Integrity check of the map file: Calculate the checksum and hash value of the map file, and compare the calculated checksum and hash value with the expected checksum and hash value. If they are consistent, it means that the content of the map file is complete and not damaged.
[0100] At the same time, record the abnormal situations in the result statistics dictionary and discard the data with anomalies.
[0101] The special behaviors of candidate data are generally the key focus objects. The information on vehicle lane changes in candidate data is stored in the lane change file. To ensure the validity of the lane change file, the check of the lane change file is usually used as a basic check. Among them, the lane change file compliance check mainly includes:
[0102] (1) Existence check of the lane change file: Obtain the path of the lane change file and use the file system API to check whether the lane change file exists.
[0103] (2) Lane change file integrity check: Calculate the checksum and hash value of the lane change file, and compare the calculated checksum and hash value with the expected checksum and hash value. If they are the same, it means that the content of the lane change file is complete and not damaged.
[0104] (3) Lane change data compliance check: Read the data in the lane change file, check whether the data format is correct (for example, timestamp, vehicle ID, lane information, etc.), and then check the logical consistency of the data (whether the timestamps are continuous and whether the lane changes are reasonable).
[0105] Ensure that each frame of data has corresponding lane change information, and at the same time record abnormal situations in the result statistics dictionary, and discard candidate data with abnormalities.
[0106] Data labels depict different dimensions of data, are also important bases for screening data, and play an important role in aspects such as data balance. Therefore, in this embodiment, the inspection of the label file is used as the basic inspection. The compliance inspection of the label file mainly includes: (1) file existence inspection, (2) file integrity inspection, and at the same time record abnormal situations in the result statistics dictionary, and discard data with abnormalities. The specific inspection methods are similar to those of the above-mentioned map file and lane change file.
[0107] In addition, due to the complex and diverse data sources, there may be unreasonable outliers in the data values. These outliers will have an important impact on the model. In this embodiment, a basic inspection will be carried out on whether the data output file contains outliers when constructing the dataset to ensure the validity of the candidate data. Focus on checking for situations such as none values, infinite values, and abnormally large values in the candidate data. At the same time, record abnormal situations in the result statistics dictionary, and discard candidate data with abnormalities.
[0108] Step S20, use the data label and the file path of the candidate data to construct different intelligent driving datasets.
[0109] It should be noted that data labels are meta-information added to data, used to describe the content, features, or attributes of data, including but not limited to scene-level labels and object-level labels. Among them, scene-level labels are overall descriptions of the entire data segment (such as a video or a series of images), usually used to represent specific events or environmental conditions that occur in this data segment, and object labels refer to the annotation of specific objects or targets in the data segment, usually used to describe detailed information such as the positions, categories, and states of these objects.
[0110] In a feasible embodiment method, step S20 may further include steps S21 to S22:
[0111] Step S21: Group the candidate data in different dimensions by using the scenario-level label and the target-level label.
[0112] Specifically, after the compliance check of the candidate data, the desired data is further obtained through the label information. By specifying the combination of a single data label or multiple data labels, the candidate data passing the compliance check is traversed and filtered according to the specified label conditions, and the candidate data meeting the label conditions is respectively placed into different groups. Through the combination of different dimensions of labels, data that meets different scenario requirements can be obtained to construct data sets for different purposes.
[0113] Step S22: In the groups in different dimensions, construct data sets according to the file paths of the candidate data to obtain different intelligent driving data sets.
[0114] Subsequently, obtain the file paths of the data for constructing the data sets from the data in different groups in step S21, and generate intelligent driving data sets according to the combination of these file paths.
[0115] Through the above steps, the scenario-level label and the target-level label are effectively used to group the candidate data in different dimensions, and different intelligent driving data sets for different purposes are constructed according to these groups.
[0116] Through the method of the above embodiment, first, candidate data is obtained from intelligent driving data according to a preset data set production configuration file; data labels and the file paths of the candidate data are used to construct different intelligent driving data sets. This method can quickly construct intelligent driving data sets under custom conditions through the preset configuration file and automated processing flow. And it supports constructing basic data sets of any version to meet different requirements. At the same time, it has the ability of function expansion and can be adjusted and optimized with the development of technology and the change of requirements.
[0117] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , before step S10, the data set construction method further includes step S01:
[0118] Step S01: Define a data set production configuration file, which contains multiple configuration subtasks. The configuration subtasks include data set production methods, data set compliance check configurations, data time intervals, data labels, event types, vehicle information, and data sources. The data sources include production vehicle data and collected vehicle data.
[0119] Specifically, first, set a configuration file for the production of intelligent driving data sets. This configuration file mainly configures the requirements for data set production, such asFigure 4 As shown Figure 4 in the figure, it is a schematic diagram of the relationship between the configuration file and the configuration subtask. As can be seen from Figure 4, the configuration file can contain multiple configuration subtasks. When configuring multiple subtasks, each configuration subtask specifies a complete data set production requirement.
[0120] Among them, as Figure 5 shown Figure 5 in the figure, it is a schematic diagram of the parameter configuration of the configuration subtask. The configuration subtask mainly includes: (1) The data set production method, online production or offline production; (2) Data compliance check configuration, including items that need to check the data, data integrity, data outliers, file necessity, etc.; (3) Data source, select to collect data from the collection vehicle or production vehicle data; (4) Concurrency of data processing, in order to make full use of the machine resources, set a reasonable number of concurrent processes; (5) Data version, because to meet different requirements, data production may correspond to multiple versions. Here, it mainly specifies the data version used when constructing the data set; (6) Data time interval, mainly specifies the start time and end time of the data selected when constructing the data set; (7) Vehicle filtering information, mainly specifies vehicle-related information, used to select or filter specific vehicle data; (8) Data label, the label mainly describes the category of data from different dimensions. According to different needs, the label can be combined to obtain a combined label, which can be used to select or filter specific data; (9) Production vehicle configuration, when the data source is selected as production vehicle data, the configuration event type can also be selected. Since the production vehicle data is very large, according to different return purposes, the production vehicle data is classified according to the event type.
[0121] And for each configuration subtask, it is necessary to initialize the result statistics dictionary. Among them, the result statistics dictionary is used to record the statistics of the abnormal reasons and valid data of each data during the data set construction process. Since data production supports a multi-concurrent production mode, here, according to different numbers of concurrent processes, for scenarios where the number of concurrent processes is greater than 1, it is necessary to configure lock parameters to avoid abnormal data statistics.
[0122] Through the method of the above embodiment, a data set production configuration file can be defined, which can systematically manage and control all links of data set generation, ensure the quality and consistency of the data set, and provide reliable data support for the research and development and testing of the intelligent driving system.
[0123] Exemplarily, in order to help understand the implementation process of the data set production method obtained by combining the above Embodiment 1 and Embodiment 2 in this embodiment, please refer to Figure 6 , Figure 6 which provides a brief flow schematic diagram of a data set production method. Specifically:
[0124] First, set the configuration file for generating the intelligent driving dataset. This configuration file mainly sets the requirements for dataset generation and includes multiple configuration subtasks. Among them, the parameter configuration in each configuration subtask mainly includes the dataset production method, data compliance check configuration, data source, concurrency, data version, data time interval, vehicle filtering information (vehicle-related information), data labels, production vehicle data configuration (including event types), etc.
[0125] Then, initialize the result statistics dictionary for counting the abnormal reasons and valid data of each data during the dataset construction process.
[0126] Subsequently, extract candidate data according to the data source and production method. This embodiment takes production vehicle data and collected vehicle data as examples. For production vehicle data, first extract the data that meets the time interval. Secondly, select or filter the data according to the event type configuration. When the event type does not exist in the preset event type list, it is default not to detect and filter the data corresponding to this event type. When the event type exists in the preset event type list, extract the data corresponding to this event type. Finally, select or filter the data according to the vehicle configuration. When the vehicle configuration does not exist in the preset vehicle number list, it is default not to detect and filter the data corresponding to this vehicle configuration. When the vehicle configuration exists in the preset vehicle number list, extract the data corresponding to this vehicle configuration as candidate data.
[0127] For collected vehicle data, it usually exists in the form of large package storage. If the generation method is online production, according to the data time interval, obtain the ID list corresponding to all large package names in this time interval by querying the database. Further query according to the large package name to determine the collected vehicle number of this large package, and select or filter the data according to the vehicle configuration. When the vehicle configuration does not exist in the preset vehicle number list, it is default not to detect and filter the data corresponding to this vehicle configuration. When the vehicle configuration exists in the preset vehicle number list, extract the data corresponding to this vehicle configuration as candidate data. The offline data storage of collected vehicles is generally bucketed by vehicle number and date. If the production method is offline production, directly filter or extract the data through the bucket name (vehicle number, date).
[0128] Next, perform compliance checks on the candidate data according to the corresponding check items in the compliance check configuration in the sub-configuration, mainly including data integrity check, data anomaly check, and check for the existence of necessary files. Among them, the check for the existence of necessary files mainly includes compliance checks on map files, lane change files, label files, etc.
[0129] After data compliance checking, the desired data is further obtained through tag information. The tag file stores scene-level tags and target-level tags. A single tag can be specified, or the acquisition or discard of data can be determined through tag combinations. By combining different dimensions of tags, data that meets the requirements of different scenarios can be obtained to construct datasets for different purposes.
[0130] Finally, the file paths of the data for constructing the dataset are extracted and combined to generate the dataset.
[0131] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the dataset construction method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0132] This application also provides a dataset construction device. Please refer to Figure 7 , the dataset construction device includes:
[0133] A processing module 10, configured to obtain candidate data from intelligent driving data according to a preset dataset production configuration file;
[0134] A construction module 20, configured to construct different intelligent driving datasets by using data tags and the file paths of the candidate data.
[0135] The dataset construction device provided by this application adopts the dataset construction method in the above embodiment, and can solve the technical problem of how to efficiently construct an intelligent driving dataset that meets the requirements and has high-quality characteristics from a large amount of data. Compared with the prior art, the beneficial effects of the dataset construction device provided by this application are the same as those of the dataset construction method provided by the above embodiment, and other technical features in the dataset construction device are the same as the features disclosed in the method of the above embodiment, which will not be elaborated here.
[0136] This application provides a dataset construction device. The dataset construction device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the dataset construction method in the first embodiment above.
[0137] Next, refer to Figure 8, which shows a schematic structural diagram of a data set construction device suitable for implementing the embodiments of the present application. The data set construction device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), PMPs (Portable Media Player), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The shown data set construction device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0138] As Figure 8 shown, the data set construction device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the data set construction device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the data set construction device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a data set construction device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0139] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0140] The data set construction device provided by the present application adopts the data set construction method in the above embodiments, and can solve the technical problem of how to efficiently construct an intelligent driving data set that meets the requirements and has high-quality characteristics from a large amount of data. Compared with the prior art, the beneficial effects of the data set construction device provided by the present application are the same as those of the data set construction method provided by the above embodiments, and other technical features in the data set construction device are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.
[0141] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0142] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0143] The present application provides a computer-readable storage medium, which has computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the data set construction method in the above embodiments.
[0144] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0145] The above computer-readable storage medium may be included in the data set construction device; or it may exist separately without being assembled into the data set construction device.
[0146] The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed by the data set construction device, the data set construction device is caused to: extract or filter the intelligent driving data according to a pre-set data set production profile to obtain candidate data; and construct different intelligent driving data sets using data tags and the file paths of the candidate data.
[0147] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0149] The modules described in the embodiments of this application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0150] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned data set construction method, and can solve the technical problem of how to efficiently construct an intelligent driving data set that meets the requirements and has high-quality characteristics from massive data. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the data set construction method provided in the above embodiments, and will not be elaborated here.
[0151] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-described data set construction method.
[0152] The computer program product provided by the present application can solve the technical problem of how to efficiently construct an intelligent driving data set that meets the requirements and has high-quality characteristics from massive data. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the data set construction method provided by the above embodiments, and will not be elaborated here.
[0153] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A method for constructing a data set, characterized in that, The method includes: Generating a configuration file according to a preset data set, and obtaining candidate data from the intelligent driving data; Using data tags and the file paths of the candidate data to construct different intelligent driving data sets.
2. The method according to claim 1, characterized in that, The step of obtaining candidate data from the intelligent driving data according to a preset data set production configuration file includes: Extracting or filtering the intelligent driving data according to a preset data set production configuration file to obtain candidate data; Performing compliance checks on the candidate data to obtain candidate data.
3. The method according to claim 2, wherein The intelligent driving data includes production vehicle data, and the step of extracting or filtering the intelligent driving data according to a preset data set production configuration file to obtain candidate data includes: Configuring the storage location of the production vehicle data according to the data production method; Extracting the production vehicle data within the data time range from the storage location; Extracting or filtering the production vehicle data within the data time range according to the event type and vehicle information to obtain candidate data.
4. The method according to claim 3, characterized in that, The step of extracting or filtering the production vehicle data within the data time range according to the event type and vehicle information to obtain candidate data includes: Checking whether the event type associated with the production vehicle data within the data time range is in a preset event type list; If the event type is not in the preset event type list, filtering the production vehicle data corresponding to the event type; If the event type is in the preset event type list, extracting the production vehicle data corresponding to the event type as initial candidate data; Checking whether the vehicle information associated with the initial candidate data is in a preset vehicle number list; If the vehicle information is not in the preset vehicle number list, filtering the production vehicle data corresponding to the vehicle information; If the vehicle information is in the preset vehicle list, extracting the production vehicle data corresponding to the vehicle information as final candidate data.
5. The method according to claim 2, wherein The intelligent driving data includes collection vehicle data, and the collection vehicle data is stored in the form of collection vehicle data packages. The step of extracting or filtering the intelligent driving data according to a preset data set production configuration file to obtain candidate data further includes: Querying a database to obtain a name list of collection vehicle data packages within the data time range and the vehicle numbers corresponding to the name list; Extracting or filtering a preset vehicle number list according to the vehicle numbers to obtain candidate data; or Extracting or filtering the collection vehicle data through the bucket name to obtain candidate data.
6. The method according to claim 5, wherein The step of extracting or filtering a preset vehicle number list according to the vehicle numbers to obtain candidate data includes: Checking whether the vehicle number is in a preset vehicle number list; If the vehicle number is not in the preset vehicle number list, filtering the collection vehicle data corresponding to the vehicle number; If the vehicle number is in the preset vehicle number list, extracting the collection vehicle data corresponding to the vehicle number as final candidate data.
7. The method according to any one of claims 2 to 6, characterized in that Before the step of obtaining candidate data from the intelligent driving data according to a preset data set production configuration file, it further includes: Define a dataset production configuration file, which contains multiple configuration subtasks. The configuration subtasks include dataset production methods, data set compliance check configurations, data time intervals, data labels, event types, vehicle filtering information, and data sources. The data sources include mass-produced vehicle data and collected vehicle data.
8. The method according to claim 2, characterized in that The steps of performing a compliance check on the candidate data and obtaining the candidate data include: Based on the data set compliance check configuration, perform an integrity check on the candidate data; When the candidate data passes the integrity check, perform a compliance check on the map file, lane change file, and label file corresponding to the candidate data to obtain the candidate data.
9. The method according to claim 1, wherein The data labels include scenario-level labels and target-level labels. The steps of using the data labels and the file paths of the candidate data to construct different intelligent driving data sets include: Use the scenario-level labels and the target-level labels to group the candidate data in different dimensions; In the groupings in different dimensions, construct data sets according to the file paths of the candidate data to obtain different intelligent driving data sets.
10. A dataset construction device, characterized in that, The device includes: A processing module for obtaining candidate data from intelligent driving data according to a preset dataset production configuration file; A construction module for using data labels and the file paths of the candidate data to construct different intelligent driving data sets.
11. A data set construction device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the dataset construction method according to any one of claims 1 to 9.
12. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the dataset construction method according to any one of claims 1 to 9.