Driving data processing method, device, electronic device and readable storage medium
By collecting and processing a variety of driving data, identifying and labeling long-tail data, and optimizing the decision-making model of the autonomous driving system, the problem of long-tail data not being effectively utilized in existing technologies is solved, and the system's decision-making ability and safety in rare situations are improved.
Patent Information
- Application Number
- CN202411072840.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-08-06
AI Technical Summary
When faced with rare situations, the decision-making ability of existing autonomous driving systems based on long-tail data is limited, and existing data processing technologies fail to effectively tap the value of long-tail data.
By collecting vehicle status data, environmental perception data, driver behavior data and communication data, coarse-grained screening, spatiotemporal alignment, fine-grained screening and data standardization are performed to extract the optimal feature data set, determine the weight based on correlation, and annotate long-tail data to optimize the decision model.
It improves the decision-making ability and safety of autonomous driving systems in complex and rare scenarios, enhances the efficiency and accuracy of data processing, and ensures the generalization and adaptability of decision-making models.
Smart Images

Figure CN119028045B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of autonomous driving data mining, and in particular to a method for processing driving data, a device for processing driving data, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of autonomous driving technology, more and more vehicles are being equipped with autonomous driving systems. During the autonomous driving process, a large amount of driving data is generated, including abnormal situations such as vehicle failures and sensor malfunctions. This data, often referred to as "long-tail data," although less frequent, has a significant impact on the safety and stability of autonomous driving systems. However, due to the low frequency of long-tail data, the collected data may contain high noise and inaccurate labeling. These issues often affect the quality of long-tail data and limit the vehicle's decision-making capabilities.
[0003] The collection and processing of long-tail data is crucial for the development of autonomous driving technology. Existing autonomous driving systems often overlook the importance of this data, limiting the vehicle's decision-making capabilities in rare situations. Furthermore, existing data processing technologies may not effectively address these challenges, resulting in the inability to fully tap the value of long-tail data.
[0004] Specifically, in the existing long-tail data mining process, long-tail data is usually obtained based on perception data, and perception data is often concentrated on common driving scenarios and normal operations, which results in a very small proportion of long-tail data in the overall data set. This deviation makes the long-tail data insufficiently weighted in model training, making it difficult to effectively influence the vehicle's decision-making model. Summary of the Invention
[0005] An embodiment of the present invention provides a method, device, electronic device, and computer-readable storage medium for processing driving data to solve the problem in the prior art that the decision-making ability of the autonomous driving system is limited due to the lack of long-tail data collection and processing that is easily ignored.
[0006] An embodiment of the present invention discloses a method for processing driving data, the method comprising:
[0007] Collecting driving data of the vehicle through vehicle detection equipment; the driving data includes: vehicle status data, environmental perception data, driver behavior data and communication data;
[0008] Processing the driving data to obtain pre-processed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data standardization;
[0009] Performing feature processing on the preprocessed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening;
[0010] Obtaining, based on the driving data and the optimal feature data set, a plurality of data results and weights corresponding to the plurality of data results; wherein the weights are determined based on correlations between the plurality of data results;
[0011] Obtaining long-tail data according to the plurality of data results and the weights respectively corresponding to the plurality of data results;
[0012] Labeling the long-tail data according to the characteristics of the long-tail data to obtain labeled long-tail data;
[0013] The decision model of the vehicle is optimized according to the labeled long-tail data; the decision model is used to guide the vehicle to make a decision.
[0014] Optionally, the processing the driving data to obtain pre-processed data includes:
[0015] Performing the coarse-grained screening on the driving data according to a data quality threshold to obtain first data;
[0016] Performing the spatiotemporal alignment on the spatiotemporal information of the first data to obtain second data; the spatiotemporal information includes time information and space information;
[0017] Performing the fine-grained screening on the second data to obtain third data;
[0018] The third data is subjected to the data standardization process to obtain preprocessed data; the data standardization process is used to convert the third data into a target data format.
[0019] Optionally, performing feature processing on the preprocessed data to obtain an optimal feature data set includes:
[0020] Performing feature extraction on the preprocessed data to obtain feature data sets corresponding to multiple feature types;
[0021] The feature screening is performed on the feature data sets respectively corresponding to the multiple feature types to obtain the optimal feature data set, where the optimal feature data set includes the optimal feature data respectively corresponding to the multiple feature types.
[0022] Optionally, obtaining a plurality of data results and weights corresponding to the plurality of data results according to the driving data and the optimal feature data set includes:
[0023] Constructing a long-tail data recognition model based on the optimal feature data set;
[0024] Inputting the driving data into the long-tail data recognition model for data recognition to obtain a plurality of data results;
[0025] Constructing the plurality of data results into a graph structure; wherein the graph nodes in the graph structure are the data results;
[0026] Determine a correlation coefficient between adjacent graph nodes in the graph structure, and determine weights corresponding to the plurality of data results based on the correlation coefficient.
[0027] Optionally, obtaining long-tail data according to the plurality of data results and the weights respectively corresponding to the plurality of data results includes:
[0028] Multiplying the data result by the weight corresponding to the data result to obtain weighted data corresponding to the data result;
[0029] The weighted data corresponding to the plurality of data results are aggregated to obtain the long-tail data.
[0030] Optionally, after obtaining the long-tail data according to the plurality of data results and the weights respectively corresponding to the plurality of data results, the method further includes:
[0031] Determine the similarity between the plurality of long-tail data, and determine the association relationship between the plurality of long-tail data based on the similarity between the plurality of long-tail data;
[0032] Analyze the similarities between the plurality of long-tail data to determine a plurality of data groups;
[0033] Determining a data pattern corresponding to the data group; the data pattern is used to characterize a pattern corresponding to long-tail data that appears repeatedly in the data group;
[0034] The long-tail data recognition model is optimized and updated according to the long-tail data, the association relationship between the long-tail data, the data group and the data pattern.
[0035] Optionally, optimizing the vehicle decision model according to the labeled long-tail data includes:
[0036] Determining, in the annotated long-tail data, annotated feature data corresponding to the annotated information based on the annotated information in the annotated long-tail data; the annotated information includes driving scenes, driving difficulty, and driving risk factors;
[0037] Training and adjusting the decision model according to the labeled feature data;
[0038] Evaluate the decision model obtained through training and adjustment to obtain an evaluation result;
[0039] The decision model obtained through training and adjustment is optimized according to the evaluation results, so that the decision model outputs corresponding decisions when dealing with different driving scenarios, driving difficulties and driving risk factors.
[0040] Optionally, after labeling the long-tail data according to the characteristics of the long-tail data to obtain labeled long-tail data, the method further includes:
[0041] The annotated long-tail data is stored in a target database; the annotated long-tail data is displayed through a visual interface;
[0042] Searching the target database according to the keywords corresponding to the annotated long-tail data;
[0043] and / or,
[0044] The target database is searched based on the conditional keywords corresponding to the annotated long-tail data; the conditional keywords are obtained by conditionally combining multiple keywords.
[0045] An embodiment of the present invention further discloses a driving data processing device, the device comprising:
[0046] An acquisition module is used to collect driving data of the vehicle through the vehicle's detection equipment; the driving data includes: vehicle status data, environmental perception data, driver behavior data and communication data;
[0047] A data processing module, configured to process the driving data to obtain pre-processed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data standardization;
[0048] A feature processing module is used to perform feature processing on the preprocessed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening;
[0049] a weighting module, configured to obtain, based on the driving data and the optimal feature data set, a plurality of data results and weights corresponding to the plurality of data results; wherein the weights are determined based on correlations between the plurality of data results;
[0050] A long-tail data module, configured to obtain long-tail data based on the plurality of data results and the weights corresponding to the plurality of data results;
[0051] a labeling module, configured to label the long-tail data according to characteristics of the long-tail data to obtain labeled long-tail data;
[0052] An optimization module is used to optimize the decision model of the vehicle based on the labeled long-tail data; the decision model is used to guide the vehicle to make decisions.
[0053] An embodiment of the present invention further discloses an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0054] The memory is used to store computer programs;
[0055] The processor is configured to implement the method described in the embodiment of the present invention when executing the program stored in the memory.
[0056] An embodiment of the present invention further discloses a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the embodiment of the present invention.
[0057] An embodiment of the present invention further discloses a computer-readable storage medium having instructions stored thereon. When executed by one or more processors, the processors are enabled to execute the method according to the embodiment of the present invention.
[0058] The embodiments of the present invention include the following advantages:
[0059] In an embodiment of the present invention, driving data of the vehicle is collected by a detection device of the vehicle; the driving data includes: vehicle status data, environmental perception data, driver behavior data and communication data; data processing is performed on the driving data to obtain preprocessed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening and data normalization processing; feature processing is performed on the preprocessed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening; based on the driving data and the optimal feature data set, multiple data results and weights corresponding to the multiple data results are obtained; wherein the weights are determined based on the correlation between the multiple data results; long-tail data is obtained based on the multiple data results and the weights corresponding to the multiple data results; the long-tail data is labeled according to the characteristics of the long-tail data to obtain labeled long-tail data; the decision model of the vehicle is optimized based on the labeled long-tail data; the decision model is used to guide the vehicle to make decisions. The embodiments of the present invention use vehicle detection equipment to collect a variety of driving data, including vehicle status data, environmental perception data, driver behavior data, and communication data. This not only expands the scope of data collection but also increases data diversity. This extensive data collection method helps capture more rare long-tail data, thereby providing a richer information source for subsequent decision model optimization. During the data processing and feature processing stages, coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data normalization processes ensure data quality and consistency. Feature extraction and feature screening can extract the most valuable features for decision models from large amounts of data, which not only improves data processing efficiency but also enhances the generalization and adaptability of decision models. Traditional autonomous driving systems often ignore the importance of long-tail data. However, the embodiments of the present invention, through refined data processing and feature processing, can effectively identify and utilize this long-tail data, enabling autonomous driving systems to better respond to rare situations and reduce accident risks. Furthermore, through automated data collection and annotation processes, data processing efficiency and accuracy are improved. Decision models are optimized using long-tail data, enhancing the decision-making capabilities of autonomous driving systems in complex and rare scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flowchart of a method for processing driving data provided in an embodiment of the present invention;
[0061] Figure 2 This is a structural block diagram of a driving data processing device provided in an embodiment of the present invention;
[0062] Figure 3 is a schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention;
[0063] Figure 4 is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0065] Reference Figure 1 , shows a flowchart of a method for processing driving data provided in an embodiment of the present invention, which may specifically include the following steps:
[0066] In an embodiment of the present invention, a method for processing driving data provided in an embodiment of the present invention is implemented through an efficient long-tail data mining and processing system. The system specifically includes a data collection module, a data preprocessing module, a long-tail data identification module, a data annotation module, a data storage and retrieval module, a model optimization module, and a user interface module. The details are as follows:
[0067] 1) Data Collection Module: This module collects comprehensive driving data in real time through the vehicle's sensor network and communication system, including vehicle status data, environmental perception data, driver behavior data, and communication data. This comprehensive data collection method helps capture more rare long-tail data, providing a rich data source for subsequent data processing and model training.
[0068] 2) Data preprocessing module: Through coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data standardization, the quality and availability of long-tail data are effectively improved. These data processing steps ensure the consistency and accuracy of the data, laying a solid foundation for subsequent feature processing and model training.
[0069] 3) Long-tail data identification module: Utilizes deep learning models (such as CNN and LSTM) for feature extraction and obtains optimal feature datasets through feature screening. These optimal feature datasets contain the most relevant and influential features and can more accurately reflect the key information of long-tail data, thereby improving the model's prediction accuracy and performance.
[0070] 4) Data Labeling Module: Label long-tail data manually or automatically to clarify the characteristics and importance of the data. The labeled data serves as high-quality training samples for subsequent training and optimization of decision-making models, improving the model's generalization ability and robustness.
[0071] 5) Data storage and retrieval module: Stores labeled long-tail data in a high-performance database and provides keyword and conditional keyword-based retrieval capabilities. This efficient retrieval method greatly improves the efficiency of data retrieval and provides strong support for subsequent data analysis and model training.
[0072] 6) Model Optimization Module: This module uses the mined long-tail data to optimize the autonomous driving system's decision-making model. By incorporating these rare cases into the training set, the decision-making model's accuracy and robustness in handling similar situations can be improved. The optimization process uses advanced machine learning algorithms (such as deep reinforcement learning) combined with the characteristics and annotation information of the long-tail data to iteratively train and adjust the decision-making model.
[0073] 7) User Interface Module: Provides an intuitive and user-friendly interface to display the mined long-tail data and its annotation information; users can view, analyze, and process this data through the interface to further understand the performance of the autonomous driving system in rare situations and optimization directions.
[0074] Step 101: Collecting driving data of the vehicle through vehicle detection equipment; the driving data includes: vehicle status data, environment perception data, driver behavior data and communication data;
[0075] In this embodiment of the present invention, a data collection module collects driving data from the autonomous vehicle's sensors and systems. Specifically, the data collection module collects driving data from the autonomous vehicle in real time through a network of onboard sensors (such as GPS, radar, cameras, and inertial measurement units) and communication systems (such as vehicle-to-everything (V2X) technology).
[0076] In the existing long-tail data acquisition process, in the existing long-tail data mining process, long-tail data is usually obtained based on perception data. However, perception data is often concentrated on common driving scenarios and normal operations, which results in a very small proportion of long-tail data in the overall data set. This deviation makes the long-tail data insufficiently weighted in model training, making it difficult to effectively influence the vehicle's decision-making model. Therefore, in order to more effectively mine and utilize long-tail data, and thereby improve the safety and stability of the autonomous driving system, the embodiment of the present invention collects driving data such as vehicle status data, environmental perception data, driver behavior data, and communication data through the vehicle's detection equipment, so that the collected data is more comprehensive, covering not only the vehicle's own operating status, but also the vehicle's environmental information, the driver's behavior pattern, and communication data between vehicles, etc. This diversity and comprehensiveness makes the data set richer, which helps to capture more rare long-tail data in the future.
[0077] The data collection module has advanced data capture capabilities to ensure that all types of data generated during autonomous driving, regardless of size, frequency, and type, are fully recorded. Furthermore, the data collection module also features a data cache function, temporarily storing data when data transmission is limited or interrupted, and retransmitting it when the network is restored.
[0078] Step 102: Process the driving data to obtain pre-processed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data normalization.
[0079] In existing technologies, due to the low frequency of long-tail data, the collected data may contain problems such as high noise. These problems often affect the quality of long-tail data and limit the vehicle's decision-making ability. Therefore, in the embodiments of the present invention, driving data is processed through four processing steps: coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data normalization, to address the problem of low quality of existing collected long-tail data.
[0080] Specifically, problems such as large noise and many outliers that may exist in long-tail data can be solved by coarse-grained screening to remove obviously erroneous or irrelevant data, such as outliers caused by sensor failures and data missing due to communication interruptions. This solves the data quality problem and ensures the accuracy and effectiveness of subsequent processing steps. Since data collected by different sensors may have time delays and spatial position differences, the timestamps and spatial coordinates of the data are unified by spatiotemporal alignment, which solves the data synchronization problem, enables data from different sources to be accurately associated and analyzed, and improves the comprehensive utilization value of the data. On the basis of coarse-grained screening, fine-grained screening is used to further eliminate data that does not meet specific standards or conditions, such as behavioral data that does not conform to driving logic and noise data in environmental perception, which solves the data accuracy problem and ensures that every data in the data set is highly relevant and available. Since data collected by different sensors and systems often have different dimensions and ranges, data standardization is used to convert the data into a unified format and range, eliminating the dimensional differences between the data and facilitating subsequent feature extraction and model training.
[0081] The embodiments of the present invention effectively solve the quality problem of long-tail data through coarse-grained screening, spatiotemporal alignment, fine-grained screening and data standardization processing, and improve the quality and availability of long-tail data sets.
[0082] Step 103: performing feature processing on the pre-processed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening;
[0083] In practical applications, raw preprocessed data may contain a large number of dimensions, many of which contribute little to the decision model. High-dimensional data can also lead to overly complex models, increasing computational costs and the risk of overfitting. In embodiments of the present invention, feature extraction and feature screening can effectively reduce data dimensionality, improving data processing efficiency and model training speed.
[0084] The embodiments of the present invention extract valuable features from preprocessed data for decision-making models through feature extraction, transforming large amounts of preprocessed data into a more concise and representative feature set. This not only reduces the complexity of data processing and simplifies the model structure, but also improves model training efficiency and prediction accuracy. Subsequently, feature screening is used to further select the most relevant and influential features from the extracted feature set to construct an optimal feature dataset. This removes redundant and irrelevant features, reduces the data dimensionality, reduces the complexity of the model, and avoids the risk of overfitting.
[0085] Specifically, the long-tail data recognition module first uses a deep learning model (such as a convolutional neural network (CNN) or a long short-term memory (LSTM) network) to extract features from the preprocessed data to obtain feature data sets corresponding to multiple feature types. In the embodiments of the present invention, there is no restriction on the deep learning model used for feature extraction. Next, feature selection is implemented using a feature selection algorithm specific to long-tail data. The most representative optimal feature data is selected from the feature data sets corresponding to each feature type. Subsequently, the optimal feature data sets corresponding to each feature type are aggregated to obtain the optimal feature data set.
[0086] Since long-tail data occurs infrequently, feature extraction and feature screening can extract valuable features from these rare data, enhancing the model's ability to process long-tail data and improving the system's decision-making ability and security when facing rare situations.
[0087] Step 104: obtaining a plurality of data results and weights corresponding to the plurality of data results based on the driving data and the optimal feature data set; wherein the weights are determined based on the correlation between the plurality of data results;
[0088] In an embodiment of the present invention, long-tail data is identified through a long-tail data identification module, which is the core component of the system. First, in the long-tail data identification module, a long-tail data identification model based on a support vector machine (SVM) or random forest is constructed based on the optimal feature data set. This model mainly uses a machine learning algorithm to classify and identify the optimal feature data in the optimal feature data set, thereby accurately identifying long-tail data from a large amount of data in the subsequent long-tail data identification process.
[0089] The driving data is then fed into the long-tail data recognition model for data identification, generating multiple data results. These data results represent different types of long-tail data identified by the model. These results are then constructed into a graph structure, with graph nodes representing the data results. The system determines the correlation coefficients between adjacent graph nodes in the graph structure and, based on these correlation coefficients, assigns weights to each data result. These weights reflect the relevance and importance of the data results.
[0090] The efficient mining and processing system for long-tail data of autonomous driving in the embodiment of the present invention quantifies the correlation between data results and assigns corresponding weights to each data result, so that the data can be used more reasonably in the subsequent decision-making process, which can effectively improve the safety and reliability of the autonomous driving system. The determination of weights improves the intelligence level of data processing and improves data processing efficiency, enabling the system to more effectively utilize long-tail data and optimize model performance. The system has broad application prospects and market value.
[0091] Step 105: Obtain long-tail data according to the plurality of data results and the weights corresponding to the plurality of data results;
[0092] In an embodiment of the present invention, the system uses weights to quantify the importance of each data result, thereby screening out the long-tail data that is most valuable to the decision-making model from numerous data results. Specifically, each data result is multiplied by its corresponding weight to obtain weighted data corresponding to each data result, so as to ensure that in subsequent data aggregation, data results with high importance have a greater impact. The weighted data corresponding to multiple data results are aggregated to obtain the final long-tail data.
[0093] By combining data results with weights, the embodiment of the present invention enables the system to identify and extract long-tail data that has a significant impact on the autonomous driving system. These data usually involve rare but critical driving scenarios, such as extreme weather conditions, complex traffic conditions, vehicle failures, etc., which solves the problem of ignoring rare events in traditional data processing and provides more comprehensive data support for vehicle decision-making model training and optimization.
[0094] The efficient mining and processing system for long-tail data of autonomous driving in the embodiment of the present invention achieves efficient identification and extraction of long-tail data by comprehensively considering all data results and their weights, thereby ensuring the accuracy and representativeness of the long-tail data.
[0095] Step 106: label the long-tail data according to the characteristics of the long-tail data to obtain labeled long-tail data;
[0096] In practical applications, the labeling of long-tail data requires a combination of professional knowledge and algorithms to ensure the accuracy and consistency of the labeling. However, the long-tail data collected through existing technologies often have inaccurate labeling, resulting in low labeling quality for the long-tail data, which affects the subsequent data processing and model training effects.
[0097] In an embodiment of the present invention, long-tail data is annotated using a data annotation module. Specifically, the data annotation module manually or automatically annotates the identified long-tail data. The annotation process includes clarifying the characteristics and importance of the data, such as the type of driving scenario, difficulty level, and risk factors.
[0098] Annotation can be performed manually or automatically. For long-tail data that is difficult to automatically identify, the module provides manual annotation, typically performed by a team of experts to ensure data accuracy and reliability. Automatic annotation may involve the use of machine learning algorithms or rule engines to automatically identify and annotate data. These algorithms can identify characteristics of the data based on predefined rules or patterns and annotate accordingly.
[0099] The efficient mining and processing system for long-tail data of autonomous driving in the embodiment of the present invention uses the labeled data as high-quality training samples for the training and optimization of subsequent decision models.
[0100] Step 107: Optimize the decision model of the vehicle based on the labeled long-tail data; the decision model is used to guide the vehicle to make decisions.
[0101] In an embodiment of the present invention, the decision model of the autonomous driving system is optimized through a model optimization module. Specifically, the model optimization module uses the mined long-tail data to optimize the decision model of the autonomous driving system. By incorporating these rare cases (long-tail data) into the training set, the accuracy and robustness of the decision model in handling similar situations can be improved. The optimization process uses advanced machine learning algorithms (such as deep reinforcement learning) and combines the characteristics and annotation information of long-tail data to iteratively train and adjust the decision model. The optimized decision model can better cope with complex and rare driving scenarios and improve the safety and reliability of the autonomous driving system.
[0102] In an embodiment of the present invention, driving data of the vehicle is collected by a detection device of the vehicle; the driving data includes: vehicle status data, environmental perception data, driver behavior data and communication data; data processing is performed on the driving data to obtain preprocessed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening and data normalization processing; feature processing is performed on the preprocessed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening; based on the driving data and the optimal feature data set, multiple data results and weights corresponding to the multiple data results are obtained; wherein the weights are determined based on the correlation between the multiple data results; long-tail data is obtained based on the multiple data results and the weights corresponding to the multiple data results; the long-tail data is labeled according to the characteristics of the long-tail data to obtain labeled long-tail data; the decision model of the vehicle is optimized based on the labeled long-tail data; the decision model is used to guide the vehicle to make decisions. The embodiments of the present invention use vehicle detection equipment to collect a variety of driving data, including vehicle status data, environmental perception data, driver behavior data, and communication data. This not only expands the scope of data collection but also increases data diversity. This extensive data collection method helps capture more rare long-tail data, thereby providing a richer information source for subsequent decision model optimization. During the data processing and feature processing stages, coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data normalization processes ensure data quality and consistency. Feature extraction and feature screening can extract the most valuable features for decision models from large amounts of data, which not only improves data processing efficiency but also enhances the generalization and adaptability of decision models. Traditional autonomous driving systems often ignore the importance of long-tail data. However, the embodiments of the present invention, through refined data processing and feature processing, can effectively identify and utilize this long-tail data, enabling autonomous driving systems to better respond to rare situations and reduce accident risks. Furthermore, through automated data collection and annotation processes, data processing efficiency and accuracy are improved. Decision models are optimized using long-tail data, enhancing the decision-making capabilities of autonomous driving systems in complex and rare scenarios.
[0103] In one embodiment of the present invention, in step 102, processing the driving data to obtain pre-processed data includes:
[0104] Performing the coarse-grained screening on the driving data according to a data quality threshold to obtain first data;
[0105] Performing the spatiotemporal alignment on the spatiotemporal information of the first data to obtain second data; the spatiotemporal information includes time information and space information;
[0106] Performing the fine-grained screening on the second data to obtain third data;
[0107] The third data is subjected to the data standardization process to obtain preprocessed data; the data standardization process is used to convert the third data into a target data format.
[0108] In the embodiment of the present invention, the driving data is preprocessed by the data preprocessing module, which generally needs to be implemented through four stages: coarse-grained screening, spatiotemporal alignment, fine-grained screening, and standardization.
[0109] Specifically, during the coarse-grained screening phase, the data preprocessing module first cleans the collected raw data. Using set data quality thresholds, it quickly removes obvious substandard data, such as noise, invalid data, and erroneous data. This simple data quality threshold allows for rapid processing, effectively reducing data volume and improving subsequent processing efficiency.
[0110] Spatiotemporal alignment: Advanced sensor data calibration techniques are used to ensure temporal and spatial synchronization between sensors, guaranteeing data accuracy and consistency. Performing spatiotemporal alignment after coarse-grained filtering ensures alignment is performed on a relatively clean dataset, reducing errors and complexity during the alignment process.
[0111] Fine-grained screening: Noise filtering algorithms (such as Kalman filters) are used to further remove outliers and interference from the data. This typically involves more complex algorithms and calculations, resulting in relatively slow processing speeds. Fine-grained screening, performed after coarse-grained screening and spatiotemporal alignment, ensures that the algorithm operates on a higher-quality dataset, improving the accuracy and effectiveness of the screening.
[0112] Standardization: Data standardization is implemented, converting all data to a common measurement range, laying the foundation for subsequent data mining and analysis. This common measurement range can be understood as a data format that the system can recognize and process, also known as the target data format. Standardization after fine-grained screening ensures that standardization is performed on a high-quality dataset that has undergone multiple rounds of screening, improving the efficiency and effectiveness of standardization.
[0113] Coarse-grained filtering can quickly remove large amounts of unqualified data, reducing the data volume and saving time and computing resources in subsequent processing steps. Directly performing fine-grained filtering can waste computing resources and time on large amounts of unqualified data. The combined use of coarse-grained and fine-grained filtering can simplify the complexity of subsequent processing steps (such as spatiotemporal alignment and normalization), ensuring that they can be performed on relatively clean and consistent datasets, improving processing efficiency and effectiveness.
[0114] In one embodiment of the present invention, performing feature processing on the preprocessed data to obtain an optimal feature data set includes:
[0115] Performing feature extraction on the preprocessed data to obtain feature data sets corresponding to multiple feature types;
[0116] The feature screening is performed on the feature data sets respectively corresponding to the multiple feature types to obtain the optimal feature data set; the optimal feature data set includes the optimal feature data respectively corresponding to the multiple feature types.
[0117] In existing technologies, feature extraction often only focuses on some common features, while ignoring other features that may have an important impact on decision-making. This incomplete feature extraction method limits the performance and generalization ability of the model. When screening features, there is often a lack of refined screening methods, resulting in some redundant or irrelevant features being retained, which increases the complexity and computational cost of the model. It may also introduce noise and affect the accuracy of the model.
[0118] Currently, due to deficiencies in feature extraction and feature screening, model training often lacks sufficient high-quality feature support, limiting the model's generalization and robustness when faced with complex and rare situations. Therefore, in an embodiment of the present invention, feature extraction is used to extract the most valuable features for the decision model from a large amount of preprocessed data, and feature screening is used to remove redundant and irrelevant features, thereby constructing an efficient and streamlined optimal feature dataset.
[0119] 1) Feature Extraction: First, a deep learning model (such as CNN and LSTM) is used to extract feature datasets for each feature type from the preprocessed driving data to capture the complex patterns and relationships in the data. The extracted feature types may include but are not limited to: vehicle speed change patterns, acceleration distribution, directional stability, and environmental perception information (such as obstacle distance and traffic sign recognition).
[0120] These extracted features are classified into different feature types according to their properties and uses, such as dynamic features (speed, acceleration), static features (position, direction), environmental features (obstacles, traffic signs), etc.
[0121] 2) Feature Screening: Within each feature type, we use specific feature selection algorithms (such as principal component analysis (PCA), correlation analysis, and information gain) to evaluate and select the most representative and relevant features. This helps identify the features most critical for recognizing long-tail data, thereby reducing the dimensionality of the feature space and improving the model's processing efficiency and accuracy. This step yields the optimal feature data for each feature type, ultimately forming the optimal feature dataset.
[0122] The optimal feature dataset in this embodiment of the present invention contains the most relevant and influential feature data. This feature data more accurately reflects key information, thereby improving the model's predictive accuracy and performance. This optimal feature dataset can be used to train the model later, making it more adaptable to various driving scenarios, including rare ones, thereby enhancing the model's robustness and reliability.
[0123] Through feature extraction and screening, the embodiments of the present invention ensure that the features used in the recognition model are the most representative and relevant, thereby improving the recognition accuracy of long-tail data. The optimal feature dataset in the embodiments of the present invention can provide more accurate and reliable information support for the decision-making of the autonomous driving system, thereby optimizing the vehicle's decision-making process. This enables the efficient mining and processing of autonomous driving long-tail data systems to make more reasonable and safer decisions in complex and rare situations.
[0124] In one embodiment of the present invention, step 104 of obtaining a plurality of data results and weights corresponding to the plurality of data results based on the driving data and the optimal feature data set includes:
[0125] Constructing a long-tail data recognition model based on the optimal feature data set;
[0126] Inputting the driving data into the long-tail data recognition model for data recognition to obtain a plurality of data results;
[0127] Constructing the plurality of data results into a graph structure; wherein the graph nodes in the graph structure are the data results;
[0128] Determine a correlation coefficient between adjacent graph nodes in the graph structure, and determine weights corresponding to the plurality of data results based on the correlation coefficient.
[0129] In an embodiment of the present invention, a long-tail data recognition model is constructed based on the optimal feature dataset. The long-tail data recognition model can be based on machine learning algorithms such as support vector machines (SVMs) or random forests. These algorithms can process high-dimensional data and have good generalization capabilities, making them suitable for identifying long-tail data from massive amounts of data. Subsequently, the driving data is input into the long-tail data recognition model, which analyzes and processes the input driving data to produce multiple data results. These results represent different types of long-tail data identified by the model, reflecting rare situations that the vehicle may encounter during driving.
[0130] To better understand and analyze the relationships between these data results, multiple data results are constructed into a graph structure, where the graph nodes represent the data results. By constructing a graph structure, the system can intuitively display the associations and patterns between the data results, facilitating subsequent analysis and processing. The correlation coefficients between adjacent graph nodes in the graph structure are determined, and the weights corresponding to each data result are determined based on these correlation coefficients. These weights reflect the relevance and importance between the data results, helping the system to more rationally utilize this data in the subsequent decision-making process. Furthermore, by quantifying the relationships between the data results, the system can assign corresponding weights to each data result, thereby more effectively utilizing long-tail data in the decision-making process.
[0131] The long-tail data recognition model constructed in this embodiment of the present invention is based on an optimal feature dataset and can better generalize to unseen data, improving the model's ability to handle novel situations. By determining the graph structure and correlation coefficient, the system can more effectively process and analyze data results, improving data processing efficiency. The identified data results and their weighting provide a high-quality data foundation for subsequent long-tail data annotation and model optimization, helping to enhance the decision-making support capabilities of autonomous driving systems.
[0132] In one embodiment of the present invention, step 105 of obtaining long-tail data according to the plurality of data results and the weights corresponding to the plurality of data results includes:
[0133] Multiplying the data result by the weight corresponding to the data result to obtain weighted data corresponding to the data result;
[0134] The weighted data corresponding to the plurality of data results are aggregated to obtain the long-tail data.
[0135] Long-tail data refers to data that occurs less frequently but is of significant value. This data is often overlooked in traditional data processing methods. By mining and processing long-tail data, the present invention enables autonomous driving systems to better respond to rare situations and reduce accident risks.
[0136] In an embodiment of the present invention, each data result is first multiplied by its corresponding weight to obtain weighted data corresponding to each data result, thereby adjusting the influence of the data result according to the importance of the data result (represented by the weight). Weights are generally determined based on the relevance and importance between data results. Data results with higher weights will have a greater impact in subsequent data aggregation. Subsequently, the weighted data corresponding to multiple data results are aggregated, and all data results and their weights are comprehensively considered to obtain the final long-tail data.
[0137] Data aggregation can adopt a variety of methods, such as weighted average, weighted summation or other statistical methods. The specific method depends on the nature of the data and the application scenario, and is not limited in the embodiments of the present invention.
[0138] In one embodiment of the present invention, after obtaining the long-tail data in step 105 based on the plurality of data results and the weights corresponding to the plurality of data results, the method further includes:
[0139] Determine the similarity between the plurality of long-tail data, and determine the association relationship between the plurality of long-tail data based on the similarity between the plurality of long-tail data;
[0140] Analyze the similarities between the plurality of long-tail data to determine a plurality of data groups;
[0141] Determining a data pattern corresponding to the data group; the data pattern is used to characterize a pattern corresponding to long-tail data that appears repeatedly in the data group;
[0142] The long-tail data recognition model is optimized and updated according to the long-tail data, the association relationship between the long-tail data, the data group and the data pattern.
[0143] In an embodiment of the present invention, in order to continuously improve the accuracy and generalization ability of the recognition model, incremental learning technology is also used in the long-tail data recognition module to enable the long-tail data recognition model to continuously learn and adapt to new long-tail data.
[0144] In addition, similarity metrics (such as cosine similarity, Euclidean distance, etc.) are used to calculate the similarity between multiple long-tail data. Based on the similarity, association rule mining algorithms (such as Apriori or FP-Growth) are used to discover the association relationships between long-tail data, revealing the potential connections between the data, and cluster analysis (such as K-means or DBSCAN) is used to discover hidden patterns and groups (data patterns and data groups) in long-tail data, providing strong support for subsequent model optimization.
[0145] Specifically, association rule mining algorithms are used to discover associations between long-tail data. These algorithms can identify frequently occurring item sets and rules within a dataset, thereby revealing potential connections between long-tail data. For example, in the field of autonomous driving, suppose the long-tail data includes "sudden traffic jams" and "foggy weather." Association rule mining can reveal that these two events often occur together, forming an association rule: "If a sudden traffic jam occurs, heavy fog is likely to occur at the same time."
[0146] Cluster analysis is an unsupervised learning method that uses similarity between long-tail data to discover data groups and data patterns in long-tail data. It groups objects in a dataset based on similarity to form different clusters. Each cluster represents a data group. Objects within a cluster have high similarity, while objects between different clusters have low similarity. The data patterns corresponding to each data group are determined. These patterns characterize the common characteristics of long-tail data that recur in the group. For example, in the field of autonomous driving, suppose the long-tail data includes the performance of "tire burst" under different driving conditions. Through cluster analysis, it can be found that the performance patterns of "tire burst" are different under high-speed and low-speed driving conditions, forming two different clusters, representing tire burst under high-speed driving conditions and tire burst under low-speed driving conditions respectively.
[0147] Optimizing and updating the long-tail data recognition model based on long-tail data, the correlation between long-tail data, data groups and data patterns can improve the accuracy and generalization ability of the long-tail data recognition model, so that it can better adapt to new long-tail data.
[0148] In one embodiment of the present invention, step 107, optimizing the vehicle decision model based on the labeled long-tail data, includes:
[0149] Determining, in the annotated long-tail data, annotated feature data corresponding to the annotated information based on the annotated information in the annotated long-tail data; the annotated information includes driving scenes, driving difficulty, and driving risk factors;
[0150] Training and adjusting the decision model according to the labeled feature data;
[0151] Evaluate the decision model obtained through training and adjustment to obtain an evaluation result;
[0152] The decision model obtained through training and adjustment is optimized according to the evaluation results, so that the decision model outputs corresponding decisions when dealing with different driving scenarios, driving difficulties and driving risk factors.
[0153] In an embodiment of the present invention, the decision model of the autonomous driving system is optimized through a model optimization module. Specifically, the model optimization module uses the mined long-tail data to optimize the decision model of the autonomous driving system. By incorporating these rare cases (long-tail data) into the training set, the accuracy and robustness of the decision model in handling similar situations can be improved. The optimization process uses advanced machine learning algorithms (such as deep reinforcement learning) and combines the characteristics and annotation information of long-tail data to iteratively train and adjust the decision model. The optimized decision model can better cope with complex and rare driving scenarios and improve the safety and reliability of the autonomous driving system.
[0154] In an embodiment of the present invention, first, based on the annotation information in the annotated long-tail data, the annotated feature data corresponding to the annotated information is determined. The annotated information may include driving scenarios (such as driving at night, rainy and snowy weather, etc.), driving difficulty (such as easy, medium, difficult, etc.), and driving risk factors (such as high-speed driving, emergency braking, etc.). These annotated information provides key supervisory signals for the training of the decision model. Subsequently, the decision model is trained and adjusted based on the annotated feature data. In this training and adjustment process, the model is usually trained using a machine learning algorithm (such as deep reinforcement learning) so that it can output corresponding decisions based on the input feature data.
[0155] After model training and tuning, the resulting decision-making model needs to be evaluated to obtain evaluation results. This evaluation typically involves using a validation set or a test set to test the model's performance. Evaluation metrics may include accuracy, recall, and F1 score. The evaluation results reflect the model's performance in various driving scenarios, driving difficulties, and driving risk factors.
[0156] The decision model is trained and adjusted based on the evaluation results. If the evaluation results show that the model performs poorly in certain situations, the decision model can be optimized by adjusting model parameters, increasing training data, and improving feature engineering. The optimized decision model can better handle complex and rare driving scenarios, improving the safety and reliability of the autonomous driving system.
[0157] In practical applications, data annotation is usually performed on image data. However, the embodiments of the present invention can perform data annotation not only on image data, but also on other types of driving data, such as vehicle status data, environmental perception data, etc.
[0158] In an embodiment of the present invention, during the training and adjustment process of the decision model, labeled information of various driving scenarios, driving difficulty and driving risk factors is combined, so that the model can better adapt to various complex driving environments, enhance the adaptability and robustness of the model, and by combining machine learning algorithms and long-tail data, the safety and reliability of the autonomous driving system can be further improved.
[0159] In one embodiment of the present invention, after the long-tail data is labeled according to the characteristics of the long-tail data in step 106 to obtain labeled long-tail data, the method further includes:
[0160] The annotated long-tail data is stored in a target database; the annotated long-tail data is displayed through a visual interface;
[0161] Searching the target database according to the keywords corresponding to the annotated long-tail data;
[0162] and / or,
[0163] The target database is searched based on the conditional keywords corresponding to the annotated long-tail data; the conditional keywords are obtained by conditionally combining multiple keywords.
[0164] In the field of autonomous driving, long-tail data often involves rare but critical driving scenarios. The collection and management of this data is complex, and traditional data management methods often struggle to effectively process and utilize it. Labeling long-tail data can clarify information such as its category, status, and importance, improving its quality and usability. However, efficient management and retrieval of this labeled data remains a challenge that has yet to be fully addressed in existing technologies.
[0165] The embodiment of the present invention achieves efficient management and utilization of the annotated long-tail data by storing the annotated long-tail data in a target database and providing a search function based on keywords and conditional keywords.
[0166] In this embodiment of the present invention, the storage and retrieval of annotated long-tail data is implemented through a data storage and retrieval module. Specifically, the data storage and retrieval module is responsible for storing the annotated long-tail data in a high-performance database and providing efficient retrieval capabilities. The database design fully considers data security and scalability to ensure the stability and reliability of the annotated long-tail data during storage and retrieval.
[0167] In this embodiment of the present invention, the user interface module displays annotated long-tail data. Specifically, the module provides an intuitive and user-friendly interface that displays mined long-tail data and its annotated information. Users can view, analyze, and process this data through the interface, gaining a deeper understanding of the autonomous driving system's performance in rare situations and optimizing it. The interface design fully considers user habits and needs, offering a variety of data display methods and interactive features to facilitate efficient data management and decision-making analysis.
[0168] In addition, based on the data storage and retrieval module and the user interface module, users can quickly retrieve the required long-tail data through keywords, condition combinations, etc., providing strong support for subsequent data analysis and model training. For example, users can quickly find relevant long-tail data by entering "foggy weather". This retrieval method is simple and direct, and is suitable for quickly locating specific types of data. Users can also combine the keywords "foggy weather" and "highway" to retrieve long-tail data under foggy weather on highways. This retrieval method is more flexible and can meet more complex data retrieval needs.
[0169] The embodiment of the present invention uses a retrieval function based on keywords and conditional keywords, so the system can quickly locate and obtain the required annotated long-tail data. This efficient retrieval method greatly improves the efficiency of data retrieval and saves time and resources. The efficient retrieval function provides more accurate and reliable information support for the decision support of the autonomous driving system. Based on the retrieval results, the system can quickly obtain relevant long-tail data, optimize the decision-making process, and improve the accuracy and safety of decisions. By effectively utilizing the annotated long-tail data, the system can better respond to potential risks and challenges, thereby improving the performance and reliability of the autonomous driving system. The introduction of the retrieval function enables the system to more accurately identify and process various driving scenarios, improving the overall performance of the system.
[0170] The embodiments of the present invention facilitate users to view, analyze and process long-tail data by providing an intuitive user interface, providing strong support for the development and optimization of autonomous driving systems.
[0171] In an embodiment of the present invention, driving data of the vehicle is collected by a detection device of the vehicle; the driving data includes: vehicle status data, environmental perception data, driver behavior data and communication data; data processing is performed on the driving data to obtain preprocessed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening and data normalization processing; feature processing is performed on the preprocessed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening; based on the driving data and the optimal feature data set, multiple data results and weights corresponding to the multiple data results are obtained; wherein the weights are determined based on the correlation between the multiple data results; long-tail data is obtained based on the multiple data results and the weights corresponding to the multiple data results; the long-tail data is labeled according to the characteristics of the long-tail data to obtain labeled long-tail data; the decision model of the vehicle is optimized based on the labeled long-tail data; the decision model is used to guide the vehicle to make decisions. The embodiments of the present invention use vehicle detection equipment to collect a variety of driving data, including vehicle status data, environmental perception data, driver behavior data, and communication data. This not only expands the scope of data collection but also increases data diversity. This extensive data collection method helps capture more rare long-tail data, thereby providing a richer information source for subsequent decision model optimization. During the data processing and feature processing stages, coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data normalization processes ensure data quality and consistency. Feature extraction and feature screening can extract the most valuable features for decision models from large amounts of data, which not only improves data processing efficiency but also enhances the generalization and adaptability of decision models. Traditional autonomous driving systems often ignore the importance of long-tail data. However, the embodiments of the present invention, through refined data processing and feature processing, can effectively identify and utilize this long-tail data, enabling autonomous driving systems to better respond to rare situations and reduce accident risks. Through automated data collection and labeling processes, the efficiency and accuracy of data processing are improved, and long-tail data is used to optimize decision models, improving the decision-making capabilities of autonomous driving systems in complex and rare scenarios.
[0172] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0173] Reference Figure 2, shows a structural block diagram of a driving data processing device provided in an embodiment of the present invention, which may specifically include the following modules:
[0174] The acquisition module 201 is used to acquire the driving data of the vehicle through the vehicle detection equipment;
[0175] The data processing module 202 is used to process the driving data to obtain pre-processed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening and data normalization processing;
[0176] The feature processing module 203 is used to perform feature processing on the pre-processed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening;
[0177] A weighting module 204 is configured to obtain, based on the driving data and the optimal feature data set, a plurality of data results and weights corresponding to the plurality of data results; wherein the weights are determined based on correlations between the plurality of data results;
[0178] A long-tail data module 205 is configured to obtain long-tail data based on the plurality of data results and the weights corresponding to the plurality of data results;
[0179] A labeling module 206 is configured to label the long-tail data according to characteristics of the long-tail data to obtain labeled long-tail data;
[0180] The optimization module 207 is used to optimize the decision model of the vehicle according to the labeled long-tail data; the decision model is used to guide the vehicle to make decisions.
[0181] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0182] In addition, an embodiment of the present invention also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned driving data processing method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0183] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the various processes of the aforementioned driving data processing method embodiment are implemented, and the same technical effects are achieved. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0184] An embodiment of the present invention also provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the various processes of the above-mentioned driving data processing method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0185] Figure 3 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0186] The electronic device 300 includes but is not limited to: a radio frequency unit 301, a network module 302, an audio output unit 303, an input unit 304, a sensor 305, a display unit 306, a user input unit 307, an interface unit 308, a memory 309, a processor 310, and a power supply 311. It will be understood by those skilled in the art that Figure 3 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or may combine certain components or arrange the components differently. In the embodiments of the present invention, the electronic device includes but is not limited to a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle terminal, a wearable device, and a pedometer.
[0187] It should be understood that in this embodiment of the present invention, the RF unit 301 can be used to receive and transmit signals during information transmission or calls. Specifically, it receives downlink data from the base station and transmits it to the processor 310 for processing; in addition, it transmits uplink data to the base station. Typically, the RF unit 301 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and the like. Furthermore, the RF unit 301 can communicate with the network and other devices via a wireless communication system.
[0188] The electronic device provides users with wireless broadband Internet access through the network module 302, such as helping users to send and receive emails, browse web pages, and access streaming media.
[0189] The audio output unit 303 can convert audio data received by the RF unit 301 or the network module 302 or stored in the memory 309 into an audio signal and output it as sound. In addition, the audio output unit 303 can also provide audio output related to a specific function performed by the electronic device 300 (for example, a call signal reception sound, a message reception sound, etc.). The audio output unit 303 includes a speaker, a buzzer, a receiver, etc.
[0190] The input unit 304 is used to receive audio or video signals. The input unit 304 may include a graphics processing unit (GPU) 3041 and a microphone 3042. The GPU 3041 processes image data of still pictures or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 306. The image frames processed by the GPU 3041 can be stored in the memory 309 (or other storage medium) or transmitted via the RF unit 301 or the network module 302. The microphone 3042 can receive sound and process such sound into audio data. In the case of a telephone call mode, the processed audio data can be converted into a format that can be sent to a mobile communication base station via the RF unit 301 for output.
[0191] The electronic device 300 also includes at least one sensor 305, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 3061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 3061 and / or the backlight when the electronic device 300 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used to identify the posture of the electronic device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; the sensor 305 can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be repeated here.
[0192] The display unit 306 is used to display information input by the user or information provided to the user. The display unit 306 may include a display panel 3061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0193] The user input unit 307 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the electronic device. Specifically, the user input unit 307 includes a touch panel 3071 and other input devices 3072. The touch panel 3071, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using any suitable object or accessory such as a finger, stylus, etc. on or near the touch panel 3071). The touch panel 3071 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction and detects the signal caused by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 310, which receives and executes the command sent by the processor 310. In addition, the touch panel 3071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 3071, the user input unit 307 may also include other input devices 3072. Specifically, other input devices 3072 may include but are not limited to a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which are not described in detail here.
[0194] Furthermore, the touch panel 3071 may be overlaid on the display panel 3061. When the touch panel 3071 detects a touch operation on or near it, it transmits the information to the processor 310 to determine the type of touch event. Subsequently, the processor 310 provides corresponding visual output on the display panel 3061 according to the type of touch event. Figure 3 In the figure, the touch panel 3071 and the display panel 3061 are two independent components to realize the input and output functions of the electronic device. However, in some embodiments, the touch panel 3071 and the display panel 3061 can be integrated to realize the input and output functions of the electronic device, which is not limited here.
[0195] The interface unit 308 is an interface for connecting external devices to the electronic device 300. For example, the external devices may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 308 may be used to receive input (e.g., data information, power, etc.) from the external device and transmit the received input to one or more elements within the electronic device 300, or may be used to transmit data between the electronic device 300 and the external device.
[0196] Memory 309 can be used to store software programs and various data. Memory 309 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, memory 309 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0197] The processor 310 is the control center of the electronic device. It connects all parts of the electronic device using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 309 and accessing data stored in the memory 309, it performs various functions of the electronic device and processes data, thereby providing overall monitoring of the electronic device. The processor 310 may include one or more processing units; preferably, the processor 310 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 310.
[0198] The electronic device 300 may also include a power supply 311 (such as a battery) to supply power to each component. Preferably, the power supply 311 may be logically connected to the processor 310 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.
[0199] In addition, the electronic device 300 includes some functional modules not shown, which will not be described here.
[0200] like Figure 4 As shown, in another embodiment provided by the present invention, a computer-readable storage medium 401 is also provided, in which instructions are stored. When the computer-readable storage medium 401 is run on a computer, the computer executes the method for processing driving data described in the above embodiment.
[0201] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0202] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0203] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
[0204] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the embodiments of the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0205] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0206] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0207] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0208] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0209] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, ROM, RAM, a magnetic disk, or an optical disk.
[0210] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for processing driving data, characterized in that: The method comprises: Collecting driving data of the vehicle through vehicle detection equipment; the driving data includes: vehicle status data, environmental perception data, driver behavior data and communication data; Processing the driving data to obtain pre-processed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data standardization; Performing feature processing on the preprocessed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening; Obtaining, based on the driving data and the optimal feature data set, a plurality of data results and weights corresponding to the plurality of data results; wherein the weights are determined based on correlations between the plurality of data results; Obtaining long-tail data according to the plurality of data results and the weights respectively corresponding to the plurality of data results; Labeling the long-tail data according to the characteristics of the long-tail data to obtain labeled long-tail data; The decision model of the vehicle is optimized according to the labeled long-tail data; the decision model is used to guide the vehicle to make a decision.
2. The method according to claim 1, characterized in that The processing of the driving data to obtain pre-processed data includes: Performing the coarse-grained screening on the driving data according to a data quality threshold to obtain first data; Performing the spatiotemporal alignment on the spatiotemporal information of the first data to obtain second data; the spatiotemporal information includes time information and space information; Performing the fine-grained screening on the second data to obtain third data; The third data is subjected to the data standardization process to obtain preprocessed data; the data standardization process is used to convert the third data into a target data format.
3. The method according to claim 1, characterized in that The performing feature processing on the preprocessed data to obtain an optimal feature data set includes: Performing feature extraction on the preprocessed data to obtain feature data sets corresponding to multiple feature types; The feature screening is performed on the feature data sets respectively corresponding to the multiple feature types to obtain the optimal feature data set; the optimal feature data set includes the optimal feature data respectively corresponding to the multiple feature types.
4. The method according to claim 1, wherein The step of obtaining a plurality of data results and weights corresponding to the plurality of data results based on the driving data and the optimal feature data set includes: Constructing a long-tail data recognition model based on the optimal feature data set; Inputting the driving data into the long-tail data recognition model for data recognition to obtain a plurality of data results; Constructing the plurality of data results into a graph structure; wherein the graph nodes in the graph structure are the data results; Determine a correlation coefficient between adjacent graph nodes in the graph structure, and determine weights corresponding to the plurality of data results based on the correlation coefficient.
5. The method according to claim 1, wherein Obtaining long-tail data according to the plurality of data results and the weights respectively corresponding to the plurality of data results includes: Multiplying the data result by the weight corresponding to the data result to obtain weighted data corresponding to the data result; The weighted data corresponding to the plurality of data results are aggregated to obtain the long-tail data.
6. The method according to claim 1, characterized in that Optimizing the vehicle decision model according to the labeled long-tail data includes: Determining, in the annotated long-tail data, annotated feature data corresponding to the annotated information based on the annotated information in the annotated long-tail data; the annotated information includes driving scenes, driving difficulty, and driving risk factors; Training and adjusting the decision model according to the labeled feature data; Evaluate the decision model obtained through training and adjustment to obtain an evaluation result; The decision model obtained through training and adjustment is optimized according to the evaluation results, so that the decision model outputs corresponding decisions when dealing with different driving scenarios, driving difficulties and driving risk factors.
7. The method according to claim 1, characterized in that After labeling the long-tail data according to the characteristics of the long-tail data to obtain labeled long-tail data, the method further includes: The annotated long-tail data is stored in a target database; the annotated long-tail data is displayed through a visual interface; Searching the target database according to the keywords corresponding to the annotated long-tail data; and / or, The target database is searched based on the conditional keywords corresponding to the annotated long-tail data; the conditional keywords are obtained by conditionally combining multiple keywords.
8. A driving data processing device, characterized in that: The device comprises: An acquisition module is used to collect driving data of the vehicle through the vehicle's detection equipment; the driving data includes: vehicle status data, environmental perception data, driver behavior data and communication data; A data processing module, configured to process the driving data to obtain pre-processed data; the data processing includes coarse-grained screening, spatiotemporal alignment, fine-grained screening, and data standardization; A feature processing module is used to perform feature processing on the preprocessed data to obtain an optimal feature data set; the feature processing includes feature extraction and feature screening; a weighting module, configured to obtain, based on the driving data and the optimal feature data set, a plurality of data results and weights corresponding to the plurality of data results; wherein the weights are determined based on correlations between the plurality of data results; A long-tail data module, configured to obtain long-tail data based on the plurality of data results and the weights corresponding to the plurality of data results; a labeling module, configured to label the long-tail data according to characteristics of the long-tail data to obtain labeled long-tail data; An optimization module is used to optimize the decision model of the vehicle based on the labeled long-tail data; the decision model is used to guide the vehicle to make decisions.
9. An electronic device, characterized in that: comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is configured to implement the method according to any one of claims 1 to 7 when executing a program stored in the memory.
10. A computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Optimization method and system for long-tail data recognition model
CN116310590A
Data processing method and system for autonomous vehicle
CN118289029A