Data management method and system and electronic equipment

By classifying and storing intelligent driving vehicle data and pre-configured workflow task processing, and associating data batch numbers in structured data storage and search engines, the problems of slow data screening speed and low data flow efficiency are solved, and more efficient data management and screening are achieved.

CN120020746APending Publication Date: 2025-05-20SAIC MOTOR
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311544350.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

At this stage, the screening speed of intelligent driving vehicle data is low and the data flow efficiency is not high, resulting in low data management efficiency.

Method used

By storing multiple data collected by the target vehicle during intelligent driving to different areas of the object storage module according to the data type, and processing the data using the pre-configured target data storage workflow task, the processed data is recorded in the structured data storage and search engine, and the data in the search engine and the data in the structured data storage are associated based on the data batch number.

Benefits of technology

It improves the speed of filtering out target data and data flow efficiency, reduces the time consumption of users who need to go to the local workstation to operate, avoids inefficiency problems caused by manual screening, and supports multiple people to process data at the same time, improving the overall processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020746A_ABST
    Figure CN120020746A_ABST
Patent Text Reader

Abstract

The invention provides a data management method and system and electronic equipment. Firstly, multiple pieces of data collected by a target vehicle in the intelligent driving process are stored in different areas of an object storage module according to data types. And processing the multiple pieces of data by utilizing a pre-configured target data storage workflow task. And recording the processed data in a structured data storage library and a search engine, and associating the data in the search engine with the data in the structured data storage library based on the data batch number. Through cloud storage and cloud computing, a one-stop data service from data storage to screening is provided, the data extraction speed is increased, the data circulation process is simplified, time consumption caused by the fact that a user needs to go to a local workstation for operation can be avoided, and the user experience is improved. And the problem of low screening speed caused by manual screening of the data by a local workstation is also avoided, multiple persons are supported to process the data at the same time, and the screening speed and the processing efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving data processing, and in particular, to a data management method, system, and electronic device. Background Art

[0002] With the development of technology, intelligent driving vehicles have emerged. During the driving process of intelligent driving vehicles, a lot of data will be generated, such as unstructured data generated by various types of sensors and structured data generated by the computing center. These data can realize the intelligent driving function and the iterative update of the intelligent driving system. Intelligent driving vehicles generate data in the order of terabytes every day. This data includes scenario data, computing data, control data, vehicle performance data, driver operation data, fault data, etc. By processing the data, the performance, advantages, and defects of the intelligent driving system can be objectively evaluated, so as to iteratively update the intelligent driving system.

[0003] At present, when the vehicle journey ends, the data on the large-capacity hard disk installed on the vehicle needs to be first copied to a mobile hard disk, and then the data in the mobile hard disk is uploaded to the local workstation, where the data is stored and screened. However, currently, the screening speed of filtering out the target data is relatively low, and the efficiency of data transfer is not high. Summary of the Invention

[0004] This application provides a data management method, system, and electronic device for improving the screening speed of filtering out target data and the data transfer efficiency.

[0005] In a first aspect, an embodiment of this application provides a data management method, and the method includes:

[0006] Obtain a plurality of data collected during the intelligent driving of a target vehicle, and store the plurality of data in different areas of an object storage module according to the data type, where the data type includes structured data and unstructured data;

[0007] Process the plurality of data by using a pre-configured target data warehousing workflow task;

[0008] Record the processed plurality of data in a structured data repository and a search engine, and associate the data in the search engine with the data in the structured data repository based on the data batch number;

[0009] When the data batch number is input in the search engine, obtain the target data corresponding to the data batch number.

[0010] Optionally, the configuration method of the target data warehousing workflow task includes:

[0011] Store the custom data processing model in the operator repository of the object storage module, and record the basic information of the custom data processing model in the operator table of the structured repository;

[0012] In response to obtaining a target operator from the operator repository, configure the target data ingestion workflow task using the target operator. The configuration of the target data ingestion workflow task includes recording workflow project information in the workflow project table of the structured repository and recording workflow task information in the workflow task table of the structured storage. The workflow project information includes the workflow belonging project and project allocation resource information, and the workflow task information includes workflow node information, dynamic parameter information, and system configuration parameter information.

[0013] Optionally, the processing of the multiple data using the pre-configured target data ingestion workflow task includes:

[0014] Filter out the target data ingestion workflow task from the workflow project table and the workflow task table; in response to obtaining pre-set input parameters, trigger the target data ingestion workflow task to process the multiple data, and record the time information of the target data ingestion workflow task, the parameter values of the adjusted parameters, and the job link in the workflow job table.

[0015] Optionally, in response to obtaining pre-set input parameters, trigger the target data ingestion workflow task to process the multiple data, including:

[0016] In response to obtaining pre-set input parameters, based on the distributed task processing framework caused by the automated container operation open-source platform and the open-source container local workflow, trigger the data ingestion workflow task to process the multiple data.

[0017] Optionally, the recording of the processed multiple data in the structured data repository and the search engine, and associating the data in the search engine with the data in the structured data repository based on the data batch number, includes:

[0018] Store the relative position information, file download link, and attribute information of the processed unstructured data in a dataset with a pre-set data structure; the data structure creates an index in the search engine in advance, and the index name of the index is generated according to a pre-set format;

[0019] Insert the metadata in the dataset into the search engine, and record the structure information of the dataset in the structured data repository, so that the data in the dataset has a mapping relationship with the index information in the search engine; the index information includes the index name and the metadata.

[0020] Optionally, the method further includes:

[0021] In response to receiving the identity identifier of the data set, obtaining the data set and the index information of the search engine;

[0022] According to the index information, locating the relative position information of the data;

[0023] According to the relative position information, obtaining the target data.

[0024] Optionally, the method further includes:

[0025] Traversing the metadata in the search engine and the file download link corresponding to the metadata;

[0026] Obtaining the data to be downloaded from the storage location corresponding to the file download link, and downloading the data to be downloaded.

[0027] Optionally, the unstructured data includes vehicle signal data, hardware parameters, picture data, video data, and point cloud data.

[0028] In a second aspect, the present application provides a data management system, the system includes:

[0029] A data upload module, configured to obtain a plurality of data collected during the intelligent driving process of the target vehicle, and store the plurality of data in different areas of the object storage module according to the data type, where the data type includes structured data and unstructured data;

[0030] A workflow management module, configured to process the plurality of data by using a pre-configured target data warehousing workflow task;

[0031] A data set management module, configured to record the processed plurality of data in a structured data repository and a search engine, and associate the data in the search engine with the data in the structured data repository based on a data batch number; when the data batch number is input in the search engine, obtaining the target data corresponding to the data batch number.

[0032] In a third aspect, the present application provides an electronic device, the electronic device includes a memory and a processor, and the memory is coupled to the processor;

[0033] The memory stores a program, and when the program is executed by the processor, the electronic device is caused to execute the method according to any one of the first aspect.

[0034] Beneficial effects:

[0035] The present application provides a data management method, system and electronic device. When executing the method, first, multiple data collected during the intelligent driving of a target vehicle are stored in different areas of an object storage module according to the data type. The multiple data are processed using a pre-configured target data warehousing workflow task. The processed multiple data are recorded in a structured data repository and a search engine, and the data in the search engine and the data in the structured data repository are associated based on a data batch number. In this way, a user can, on a client side, input the data batch number based on the search engine to filter out the target data. It can be seen that in the embodiments of the present application, by classifying and storing data and inputting the data batch number on the search engine, the target data can be filtered out. Through cloud storage and cloud computing, a one-stop data service from data storage to filtering is provided, accelerating the data extraction speed and streamlining the data flow process. This method can avoid the time consumption caused by the user having to go to a local workstation for operation, and also avoid the problem of slow filtering speed caused by manual filtering of data by the local workstation, and supports multiple people to process data simultaneously, improving the filtering speed and processing efficiency. Description of the Drawings

[0036] Figure 1 It is a flowchart of a data management method provided by an embodiment of the present application;

[0037] Figure 2 It is a data management system provided by an embodiment of the present application;

[0038] Figure 3 It is a flowchart of the implementation of data upload and data warehousing provided by an embodiment of the present application;

[0039] Figure 4 It is a flowchart of the implementation of data warehousing workflow configuration provided by an embodiment of the present application;

[0040] Figure 5 It is a flowchart of a method for generating a data set provided by an embodiment of the present application. Detailed Embodiments

[0041] Terms such as "first", "second" and "third" in the specification, claims and drawings of the present application are used to distinguish different objects, rather than to limit a specific order.

[0042] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0043] As described above, the current method for processing intelligent driving data is as follows: A large-capacity hard disk is installed in the vehicle, and the large-capacity hard disk is used to collect the data generated by the target vehicle during intelligent driving. The data includes structured data, such as data batch number, target vehicle model, collection time, vehicle driving road type, collection weather, and collection city. The data also includes unstructured data, such as the vehicle signal data, hardware parameters, and content data of the target vehicle. Among them, the content data is the collected picture data, video data, and point cloud data.

[0044] The data collected by the large-capacity hard disk on the target vehicle first needs to be copied to a mobile hard disk, and then, using the mobile hard disk, the user uploads the collected data to the local workstation, where the data is managed, such as stored, parsed, and screened. After analysis by the inventor, it is found that: at present, when managing data at the local workstation, the data screening speed is relatively low.

[0045] Further analysis reveals that the reason for the low screening speed is that the data needs to be parsed at the local workstation, and after parsing, the target data is obtained through manual screening. This results in a low screening speed at present.

[0046] In view of the above problems, the present application provides a data management method. By classifying and storing the data, entering the data batch number in the search engine can screen out the target data. Through cloud storage and cloud computing, a one-stop data service from data storage to screening is provided, accelerating the data extraction speed and streamlining the data flow process. This method can avoid the time consumption caused by the user having to go to the local workstation for operations, and also avoid the problem of slow screening speed caused by manual screening of data at the local workstation, and supports multiple people to process the data simultaneously, improving the screening speed and processing efficiency.

[0047] The data management method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0048] See Figure 1 , which is a flowchart of a data management method provided by an embodiment of the present application. This method is applied to the client. The method specifically includes the following processes:

[0049] S11: Obtain a plurality of data collected by the target vehicle during intelligent driving.

[0050] It can be understood that a large-capacity hard disk is installed on the target vehicle. When the target vehicle is in the process of intelligent driving, multiple data are collected by sensors, such as image data obtained by camera photographing and point cloud data of lidar. The sensors transmit the multiple collected data to the large-capacity hard disk. The large-capacity hard disk is inserted into the upper disk drive, triggering the upload of the multiple collected data to the background server for data management. This method does not require the user to manually confirm the upload command, and the user experience is good.

[0051] The multiple data include structured data and unstructured data. Structured data refers to vehicle-related metadata, and metadata refers to collection conditions, such as collection time, collection weather, collection city, collection duration, file size, and collection mileage, etc. Structured data also includes data batch numbers, target vehicle models, etc. Unstructured data refers to vehicle signal data, hardware parameters, image data, video data, and point cloud data.

[0052] S12: Store the multiple data in different areas of the object storage module according to the data type.

[0053] The multiple data are classified and stored in the object storage module according to the data type. The data type includes two types: structured data and unstructured data. Specifically, the structured data is stored in area A, and the unstructured data is stored in area B. Among them, area A and area B are different areas in the object storage module. This standardized storage will not cause storage conflicts, and the data in different areas can be processed separately, improving the data processing speed.

[0054] The object storage module is the corresponding memory in the cloud, which can be a private cloud or a public cloud.

[0055] The multiple obtained data do not need to be stored in the local workstation through a mobile hard disk. It only needs to store the multiple data in different areas of the cloud. Therefore, it improves the problem of poor user experience of storing data in the local workstation through a mobile hard disk.

[0056] Optionally, in the embodiment of the present application, every time the data collected during the intelligent driving of the target vehicle is obtained, a corresponding unique data batch number will be generated. For example, the data batch number V013-2013-0912-V1 represents the first batch of data obtained on September 12, 2013, and the data batch number V013-2013-0912-V2 represents the second batch of data obtained on September 12, 2013. The data batch number is used to associate the unstructured data and metadata in the object storage module. In this way, the user can input the data batch number in the corresponding user interface of the client, and the data corresponding to the data batch number can be displayed in the user interface. The operation is simple and the usability is good.

[0057] Optionally, after multiple data are stored in the object storage module, the object storage module automatically generates the storage address of each data according to a preset rule. According to the storage address, a file download link can be generated. Optionally, it can be presented on the client through a preset third interface. The user can perform a download operation corresponding to the file download link on the client to download the data corresponding to the file download link. Therefore, the present application has good usability and a good user experience.

[0058] S13: Process multiple data using a pre-configured target data warehousing workflow task.

[0059] The data warehousing workflow task refers to the workflow and the process information of the workflow required to store data in the destination repository. In the embodiments of the present application, the data warehousing workflow task can be configured through the workflow project name, such as the project to which the data warehousing workflow task belongs, the workflow name, the business type, the operator, and the data processing stage. Among them, the business type includes forward view, surround view, and panoramic view, the operator includes algorithm models such as data pre-annotation, video data playback extraction, surround view stitching, and image feature extraction, and the data processing stages include preprocessing, data processing, and label sending, etc.

[0060] Optionally, store the custom data processing model in the operator warehouse of the object storage, and record the basic information of the custom data processing model in the operator table in the structured repository. The basic information includes resource requirement information, container image information, execution command entry and path information, etc.

[0061] At the same time, obtain the target operator from the operator warehouse, that is, the data processing model, and configure the target data warehousing workflow task using the target operator. Specifically, record the workflow project information, such as the workflow attribution project and project allocated resources, etc., in the workflow project table, and record the workflow task information, such as workflow node information, dynamic parameter information, and system configuration parameter information, in the workflow task table.

[0062] Among them, both the workflow project table and the workflow task table are located in the structured data repository. The user screens out the target data warehousing workflow task from the workflow project table and the workflow task table through a preset port, and processes the data using the target data warehousing workflow task.

[0063] The embodiments of the present application can process the data warehousing workflow task in the following manner, specifically including:

[0064] Step 1: Screen out the target data warehousing workflow task from the workflow project table and the workflow task table.

[0065] In the embodiment of the present application, the user can manually screen the data warehousing workflow tasks, or set the corresponding relationship between the trigger conditions and the data warehousing workflow tasks. When the trigger conditions are met, the data warehousing workflow tasks are automatically screened out.

[0066] Step 2: Trigger the data warehousing workflow task according to the set input parameters to process multiple data.

[0067] In the embodiment of the present application, the user directly inputs parameters in the user interface of the client. For example, the input parameter is the storage path of the file in the object storage (origin_data / *** / *** / ***_data batch number), so that different types of data can be processed.

[0068] In the embodiment of the present application, an open-source platform for automated container operations and a distributed task processing framework caused by the open-source container local workflow are built. For example, it is deployed based on the Kubernetes cluster and the Argo distributed task processing framework.

[0069] In the embodiment of the present application, the deployment of the Kubernetes cluster and the Argo distributed task processing can make full use of the resources of multiple underlying computers and can improve convenience and reusability.

[0070] Then, according to the input parameters, the data warehousing workflow task is automatically triggered. Optionally, after the data warehousing workflow task is triggered, the Kubernetes cluster will automatically allocate resources such as memory, disk, and network, reference the operators in the operator repository, process multiple data, and record the time information, input parameters, and job links of the data warehousing workflow task in the workflow job table. The user can monitor the data warehousing workflow task through the workflow job table of the client.

[0071] S1: Record the processed multiple data in the structured data repository and the search engine, and associate the data in the search engine with the data in the structured data repository based on the data batch number.

[0072] In the embodiment of the present application, the processed multiple data are respectively recorded in the structured data repository and the search engine. Optionally, the following are recorded in the structured data repository: structured data, the relative storage location information of the unstructured data in the object storage module, the file download link, attribute data, such as the timestamp for generating the data set, the data batch number, the frame number, and the data set identity identifier, etc. The metadata is stored in the search engine.

[0073] Metadata includes the collected data, which is a type of structured data used to describe the data. Metadata includes the collection time, collection city, collection mileage, collection duration, collection weather, and collection signal information. Inserting metadata into the search engine facilitates directly obtaining the target data through the metadata.

[0074] To establish the association between the data in the search engine and the data in the structured data repository, the data batch number is used for the association. That is, the data in the structured data repository is associated with the data batch number, and the data in the search engine is associated with the data batch number.

[0075] In an optional implementation, data management can be performed through a data set. Specifically:

[0076] Step 1: Store the relative position information, file download link, and attribute information of the processed unstructured data in a data set with a preset data structure.

[0077] It can be understood that multiple data and file download links are added to the corresponding positions of the data structure to generate the corresponding data set.

[0078] In the embodiment of the present application, the data structure creates an index in the search engine in advance, and the index name of the index is generated according to a preset format. For example, a corresponding index is created in ElasticSearch, and the index name is generated according to a preset format. For example, the preset format is "Project Number_X.Dataset Number_X", and the index name "Project Number_X.Dataset Number_X" is directly generated. In this way, there is a mapping relationship between the data set and the index name of the search engine. The user can directly input the index name in the search engine to obtain the corresponding data set.

[0079] Step 2: Insert the metadata in the data set into the search engine, and record the structure information of the data set in the structured data repository.

[0080] At this time, there is a mapping relationship between the data in the data set and the index information of the search engine. Among them, the index information includes the index name and the metadata.

[0081] S15: Input the data batch number in the search engine to filter out the target data corresponding to the data batch number.

[0082] The user can directly input the data batch number in the search engine to query the corresponding data set and the index information of the search engine, thereby locating the position of the data and achieving the purpose of quickly finding the data.

[0083] The present application provides a data management method. First, multiple data collected by a target vehicle during the intelligent driving process are stored in different regions of an object storage module according to the data type. The multiple data are processed by using a pre-configured data warehousing workflow task to generate a data set based on a file download link for the processed data. The metadata of the data set is inserted into a search engine, and the data batch number is associated with the metadata in the search engine and the unstructured data in the object storage module. Through cloud storage and cloud computing, a one-stop data service from data storage to screening is provided, which accelerates the data extraction speed and streamlines the data transfer process. This method can avoid the time consumption caused by the user having to go to a local workstation for operations and also avoid the problem of slow screening speed caused by manual screening of data by the local workstation. Moreover, it supports multiple people to process data simultaneously, improving the screening speed and processing efficiency.

[0084] The above data management method will be described in detail below.

[0085] See Figure 2 , which is a data management system provided by an embodiment of the present application, including a client and a server. The client displays a user interface, and the user can operate on the user interface of the client. Specifically, it includes:

[0086] A data upload module 201, configured to obtain multiple data collected by a target vehicle during the intelligent driving process after uploading the data on a large-capacity hard disk, and store them in different regions of an object storage module according to the data type.

[0087] Optionally, the user can query the upload record from the data upload module 201 according to the hard disk serial number, upload time, data type, and upload status.

[0088] A workflow management module 202, configured to process the multiple data by using a pre-configured target data warehousing workflow task. Optionally, the workflow management module 202 includes a workflow creation sub-module, a workflow management sub-module, an operator management sub-module, and a project management sub-module. The user can create a data warehousing workflow task in the workflow creation sub-module to process the multiple collected data according to the workflow project name, workflow name, business type, operator, and data processing node. The workflow management sub-module is used to view the logs of each workflow task and screen data according to the data batch number and the like during the workflow creation process. The operator management sub-module is used to manage and query the algorithm models required for data processing. The project management sub-module is used to manage and query each project team or project name.

[0089] The dataset management module 203 is used to process multiple processed data records in the structured data repository and the search engine, and associate the data in the search engine with the data in the structured data repository based on the data batch number; when the data batch number is input in the search engine, the target data corresponding to the data batch number is obtained.

[0090] The server includes:

[0091] The object storage module 204 is used to store the structured data (such as the collection time, the collection city, etc.) and unstructured data (such as pictures, point clouds, and videos, etc.) collected by the vehicle in a classified manner, and automatically generate the storage address of each data according to the preset address rule.

[0092] Optionally, the object storage module includes an operator warehouse for providing operator components required for data processing (such as data annotation, video data playback, surround view stitching, file format conversion, video frame extraction, data pre-classification, etc.).

[0093] The structured data repository 205 is used to store the underlying structure and business data of each module of the construction platform. The unstructured data repository 206 stores various types of data such as videos, point clouds, pictures, and automotive signal data in accordance with the agreed path format.

[0094] The cluster scheduling module 207 is used to schedule and allocate cluster resources (such as computing, storage, and network, etc.) for the workflow tasks created by the user, so as to manage data using the same interface.

[0095] See Figures 3 to 5 , which is a schematic diagram for implementing the data management method provided by the embodiment of the present application.

[0096] See Figure 3 , which is a flowchart for implementing a data upload and data warehousing provided by the embodiment of the present application.

[0097] 101: Insert a large-capacity hard disk into the disk loader to upload the large-capacity hard disk data.

[0098] 102: The data upload module obtains the structured data and unstructured data collected by the target vehicle during the intelligent driving process through the disk loader, uploads the structured data and unstructured data to the object storage module, and automatically generates a storage address.

[0099] In the embodiment of the present application, the storage address can be generated according to "main tab page / metadata field 1 / metadata field 2 / metadata field 3... / data batch number" to form a standardized storage and avoid conflicts caused by data duplication.

[0100] 103: The workflow management module screens data for the data storage workflow task, processes multiple data for storage, and records the processed data in the structured data repository and the search engine database. The data in the structured data repository and the search engine database are associated through the data batch number.

[0101] 104: The user inputs query information on the client side to locate the data.

[0102] When the input query information is metadata information, use the API1 interface for querying all metadata information to search for the metadata information and quickly locate the data. When the input query information is the data batch number, use the API2 interface for file data associated with the data batch number to obtain the unstructured data and metadata information in the structured data of the object storage module associated with the data batch number. When the input query information is the file download link, call the interface for generating the file download link by the object storage and the API3 interface for the file hierarchy interface to locate the data to be downloaded and perform data download.

[0103] In the embodiment of the present application, through the preset interfaces and various query conditions, the target data can be located, and the screening speed of the target data can be improved.

[0104] See Figure 4 , which is the implementation flowchart of a data storage workflow configuration provided by the embodiment of the present application.

[0105] 105: Use the object storage module to store the custom data processing model in the operator repository, and record the basic information of the operator, such as resource requirement information, container image information, and execution command storage and path information, etc. in the operator table in the structured data repository.

[0106] 106: The workflow management module manages the data storage workflow task. To better manage the retrieved data storage workflow, the user calls the API4 interface for reading / writing and updating the operator table to reference the operator configuration data in the operator table for the data storage workflow task. And record the workflow project information, such as the workflow attribution project and project allocated resources, etc. in the workflow project table, and record the workflow task information, such as workflow node information, dynamic parameter information, and system configuration parameter information, in the workflow task table. Among them, the workflow project table and the workflow task table are stored in the structured data repository.

[0107] 107: The user calls the interface API5 for reading / writing and updating the workflow task table and the interface API6 for reading / writing and updating the workflow job table to view the workflow project table and the workflow task table. And trigger the data storage workflow task through the input parameters to call the cluster computing resources to process the data, and record the time information of the data storage workflow task, the parameter values of the adjusted parameters, and the job link of the data storage workflow task in the workflow job table.

[0108] 108: The user calls the API7 interface to query the job status of the data storage workflow task.

[0109] See Figure 5 , which is a flowchart of a method for generating a data set provided by an embodiment of the present application.

[0110] 109: Data set preprocessing.

[0111] Preset the data structure of the data set. For example, the data structure includes the relative storage location, the downloadable link, the timestamp for generating the data set, and the data batch number from top to bottom.

[0112] To better manage and query the data in the later stage, create a corresponding index in ElasticSearch for the preset data structure. The index name is generated according to the preset format, such as project number.dataset number. For example, the index name is "project number_X.dataset number_X".

[0113] 110: Add the data processed by the data storage workflow task to the data set.

[0114] According to the data storage workflow task, process multiple data collected during the intelligent driving of the target vehicle, and fill the processed data into the above data structure to generate a data set.

[0115] Optionally, if the multiple collected data includes unstructured data, store the unstructured data in the object storage module and insert the metadata corresponding to the unstructured data into ElasticSearch. At this time, there is a mapping relationship between the data set and the index of ElasticSearch.

[0116] 111: Data set and data query.

[0117] The user can enter the identity identifier of the data set in the user interface of the client. According to the identity identifier of the data set, the corresponding data set and the index information of ElasticSearch can be queried. According to the data set and the index information of ElasticSearch, the relative position information of the data can be located, and the data can be quickly queried according to the relative position information.

[0118] 112: Dataset Update and Annotation Screening.

[0119] Using the native interface of ElasticSearch, the index name and the identity identifier of the dataset, the updatable data tags are obtained. When the user previews the pictures in the dataset, the data screening tags can be updated, and the screened data can be exported for further processing later.

[0120] 113: Dataset Download.

[0121] Traverse the metadata information of ElasticSearch in the dataset to obtain the download link of the object storage. When the user operates the download link, the dataset to be downloaded can be obtained based on the location information of the storage block.

[0122] 114: The metadata in the dataset is stored in ElasticSearch.

[0123] The metadata includes the location information of the object storage and the file download link, and also includes attribute data such as the data batch number, the timestamp for generating the dataset, the frame number, the identity identifier, etc.

[0124] 115: The unstructured data in the dataset is stored in the object storage module.

[0125] 116: Establish the index interaction between the dataset and the interfaces for inserting data, updating data, downloading data, and querying data.

[0126] Through the index creation of ElasticSearch itself, the indexes of the interfaces for inserting data, updating data, downloading data, and querying data are obtained, and the index interaction between the dataset and the interfaces for inserting data, updating data, downloading data, and querying data is established. In this way, the user can retrieve, screen, and manage the unstructured data on the user interface of the client.

[0127] In summary, the embodiments of the present application provide a one-stop data service from data storage to screening and task processing through cloud storage + cloud computing, accelerating the data extraction speed, streamlining the data flow process, and making it easier for users to obtain the desired high-value data.

[0128] The technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.

[0129] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A data management method, characterized in that: The method comprises: Acquire multiple data collected by the target vehicle during the intelligent driving process, and store the multiple data in different areas of the object storage module according to data types, wherein the data types include structured data and unstructured data; Processing the plurality of data using a pre-configured target data warehousing workflow task; Recording the processed data in a structured data repository and a search engine, and associating the data in the search engine with the data in the structured data repository based on a data batch number; When the data batch number is input into the search engine, target data corresponding to the data batch number is obtained.

2. The method according to claim 1, characterized in that: The configuration method of the target data storage workflow task includes: Storing the custom data processing model in the operator warehouse of the object storage module, and recording basic information of the custom data processing model in the operator table in the structured repository; In response to obtaining a target operator from the operator warehouse, the target data warehousing workflow task is configured using the target operator, wherein the configuration of the target data warehousing workflow task includes recording workflow project information in a workflow project table in the structured storage repository, and recording workflow task information in a workflow task table in the structured storage; the workflow project information includes workflow belonging project and project allocation resource information, and the workflow task information includes workflow node information, dynamic parameter information and system configuration parameter information.

3. The method according to claim 2, characterized in that: The processing of the plurality of data using a pre-configured target data storage workflow task includes: The target data warehousing workflow task is filtered out from the workflow project table and the workflow task table; in response to obtaining the preset input parameters, the target data warehousing workflow task is triggered to process the multiple data, and the time information, parameter values ​​of the adjustment parameters and the job link of the target data warehousing workflow task are recorded in the workflow job table.

4. The method according to claim 3, characterized in that: In response to obtaining the preset input parameters, triggering the target data storage workflow task to process the multiple data, including: In response to obtaining the preset input parameters, based on the distributed task processing framework caused by the automated container operation open source platform and the open source container local workflow, a data warehousing workflow task is triggered to process the multiple data.

5. The method according to claim 1, characterized in that: The step of recording the processed data in a structured data repository and a search engine, and associating the data in the search engine with the data in the structured data repository based on a data batch number, comprises: The relative position information, file download link and attribute information of the processed unstructured data are stored in a data set of a preset data structure; the data structure is pre-indexed in the search engine, and the index name of the index is generated according to a preset format; The metadata in the data set is inserted into the search engine, and the structural information of the data set is recorded in the structured data repository, so that the data in the data set has a mapping relationship with the index information of the search engine; the index information includes the index name and the metadata.

6. The method according to claim 5, characterized in that: The method further comprises: In response to receiving the identity of the data set, obtaining the data set and index information of the search engine; According to the index information, relative position information of the positioning data is obtained; Target data is acquired according to the relative position information.

7. The method according to claim 5, characterized in that: The method further comprises: Traversing the metadata in the search engine and the file download links corresponding to the metadata; The data to be downloaded is obtained from the storage location corresponding to the file download link, and the data to be downloaded is downloaded.

8. The method according to any one of claims 1 to 7, characterized in that: The unstructured data includes vehicle signal data, hardware parameters, image data, video data, and point cloud data.

9. A data management system, characterized in that: The system comprises: A data uploading module is used to obtain a plurality of data collected by the target vehicle during the intelligent driving process, and store the plurality of data in different areas of the object storage module according to data types, wherein the data types include structured data and unstructured data; A workflow management module, used for processing the plurality of data using a pre-configured target data storage workflow task; A data set management module is used to record the processed multiple data in a structured data repository and a search engine, and to associate the data in the search engine and the data in the structured data repository based on a data batch number; when the data batch number is entered in the search engine, the target data corresponding to the data batch number is obtained.

10. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is coupled to the processor; The memory stores a program, and when the program is executed by the processor, the electronic device executes the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Vehicle multi-modal data management method and system, electronic equipment and storage medium

    CN120540603A