A data processing method, device and electronic equipment

CN116415192BActive Publication Date: 2026-08-21北京亮道智能汽车技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310237322.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-06
Publication Date
2026-08-21
Estimated Expiration
2043-03-06

AI Technical Summary

Technical Problem

但是,在自动驾驶中,随着驾驶场景的多样化,所采集的驾驶场景数据已从多年之前的几万公里,增长到现在的几百万公里甚至更多,再利用全量数据进行训练和测试不但耗费时间长、效率低,还可能因为计算资源无法承受而导致设备瘫痪等问题,以至于不可能再利用全量数据进行训练和测试

Benefits of technology

[0047]This invention provides a data processing method, apparatus, and electronic device. The method acquires data collected by a target sensor, preprocesses the data to obtain data to be classified, and the data to be classified includes parameters under different dimensions corresponding to different scenarios. For the data to be classified, multiple dimension combinations are determined from the dimensions corresponding to each scenario. For each dimension combination, the accuracy of the dimension combination is determined based on the vehicle behavior data corresponding to each scenario included in the dimension combination. Based on the accuracy of each dimension combination, a target dimension combination is determined, and a driving scenario data model is constructed based on the data contained in the target dimension combination. By combining data from multiple scenarios with different dimensions for parameters corresponding to different scenarios, and further using the vehicle behavior data corresponding to each scenario included in the dimension combination as the evaluation basis for the accuracy of the dimension combination, the accuracy of different dimension combinations is determined. Then, the optimal scenario dimension combination data is selected based on the accuracy, and data modeling is performed using the selected optimal scenario dimension combination data. This achieves efficient and automatic data classification, solves the problems caused by existing modeling based on experience or manual methods, improves data classification efficiency, and is applicable to data classification in big data scenarios, demonstrating high applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116415192B_ABST
    Figure CN116415192B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method, device and electronic equipment, and relate to the technical field of data processing. The method comprises: acquiring data collected by a target sensor, the target sensor comprising a camera on a data collection vehicle; preprocessing the data collected by the target sensor to obtain to-be-classified data, the to-be-classified data comprising parameters in different dimensions corresponding to different scenes, the data in a scene comprising data collected by the target sensor within a preset time length, and the dimension being used to represent a component of a driving scene; determining, for the to-be-classified data, a plurality of dimension combinations from the dimensions corresponding to each scene; determining, for each dimension combination, an accuracy of the dimension combination according to primary vehicle behavior data corresponding to each scene included in the dimension combination; determining a target dimension combination based on the accuracy of each dimension combination, and constructing a driving scene data model according to data contained in the target dimension combination, thereby improving the classification efficiency of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data processing method, apparatus, and electronic device. Background Technology

[0002] With the development of artificial intelligence and autonomous driving technologies, driving scenarios are becoming increasingly diverse, leading to a surge in data across these scenarios. In autonomous driving technology, identifying, classifying, and predicting targets in the driving environment of autonomous vehicles has become a crucial technology. Target identification, classification, and prediction are achieved through the collection and analysis of large amounts of real-world driving data, followed by training and testing with different types of extracted data.

[0003] For data collected in the driving environment of autonomous vehicles, related technologies utilize the full dataset for training and testing to achieve target identification, classification, or prediction. However, in autonomous driving, with the diversification of driving scenarios, the collected driving scenario data has increased from tens of thousands of kilometers years ago to millions of kilometers or even more today. Utilizing the full dataset for training and testing is not only time-consuming and inefficient, but may also lead to equipment failure due to insufficient computing resources, making it impossible to use the full dataset for training and testing.

[0004] Based on this, some technologies involve manually defining scenario types and then classifying driving scenario data to extract different types of data with high usage value for training and testing. However, due to the large volume of driving scenario data and the great diversity of scenarios, manually classifying driving scenario data or discovering new types requires a lot of manpower and time, the data classification efficiency is low, and it is not suitable for big data processing scenarios. Summary of the Invention

[0005] The purpose of this invention is to provide a data processing method, apparatus, and electronic device to improve data classification efficiency. The specific technical solution is as follows:

[0006] In a first aspect, embodiments of the present invention provide a data processing method, the method comprising:

[0007] Acquire data collected by a target sensor, wherein the target sensor includes a camera on the data acquisition vehicle;

[0008] The data collected by the target sensor is preprocessed to obtain data to be classified; wherein, the data to be classified includes: parameters under different dimensions corresponding to different scenarios, the data under a scenario includes the data collected by the target sensor within a preset time period, and the dimension is used to represent the components of the driving scenario;

[0009] For the data to be classified, multiple dimension combinations are determined from the dimensions corresponding to each scenario;

[0010] For each dimension combination, the accuracy of that dimension combination is determined based on the main vehicle behavior data corresponding to each scenario included in that dimension combination.

[0011] Based on the accuracy of each dimension combination, a target dimension combination is determined, and a driving scenario data model is constructed based on the data contained in the target dimension combination.

[0012] In one possible implementation, determining multiple dimension combinations from the dimensions corresponding to each scenario for the data to be classified includes:

[0013] Determine the dimensions corresponding to each scenario to obtain all dimensions;

[0014] Based on all the dimensions, a preset number of dimensions are combined to determine multiple dimension combinations.

[0015] In one possible implementation, determining the accuracy of each dimension combination based on the vehicle behavior data corresponding to each scenario included in the dimension combination includes:

[0016] For each dimension combination, the parameters contained in each scenario under that dimension combination are clustered to obtain multiple classes for that dimension combination;

[0017] For each class, the accuracy of the class is determined based on the main vehicle behavior data corresponding to each scenario included in that class.

[0018] The accuracy of the dimension combination is determined based on the accuracy of each class in the dimension combination.

[0019] In one possible implementation, for each dimension combination, the parameters contained in each scenario under that dimension combination are clustered to obtain multiple classes for that dimension combination, including:

[0020] For each dimension combination, a clustering algorithm corresponding to the data category of that dimension combination is used to cluster the parameters contained in each scenario under that dimension combination, resulting in multiple classes for that dimension combination.

[0021] In one possible implementation, after obtaining multiple classes of the dimension combination, the method further includes:

[0022] For multiple classes in this dimension combination, determine whether there are any classes that contain fewer than a preset threshold number of scenarios;

[0023] If there is a class containing fewer than a preset threshold number of scenarios, delete the class containing fewer than the preset threshold number of scenarios and the parameters contained in each scenario under that class, and re-cluster the parameters contained in each scenario under that dimension combination to obtain multiple classes for that dimension combination.

[0024] In one possible implementation, determining the accuracy of each class based on the main vehicle behavior data corresponding to each scenario included in that class includes:

[0025] For each class, obtain the number of the first scenarios that contain the preset main vehicle behavior in each scenario under that class;

[0026] For each class, obtain the number of second scenarios under that class that do not contain the preset master vehicle behavior;

[0027] The accuracy of the class is determined based on the first number of scenarios and the second number of scenarios.

[0028] In one possible implementation, the preset driver behavior includes: emergency braking of the driver.

[0029] In one possible implementation, determining the accuracy of each class based on the main vehicle behavior data corresponding to each scenario included in that class includes:

[0030] For each class, obtain the main vehicle's sequential behavior characteristics that occur in the time sequence within each scenario of that class.

[0031] The percentage corresponding to the most frequently occurring main vehicle series behavior feature in this category is determined as the accuracy of this category.

[0032] In one possible implementation, determining the accuracy of each class based on the main vehicle behavior data corresponding to each scenario included in that class includes:

[0033] For each class, the accuracy of the class is determined based on the preset master vehicle behavior and master vehicle serial behavior characteristics corresponding to each scenario included in the class.

[0034] In one possible implementation, the method further includes:

[0035] The autonomous driving classification model is trained using the data contained in the driving scenario data model.

[0036] In a second aspect, embodiments of the present invention provide a data processing apparatus, the apparatus comprising:

[0037] The data acquisition module is used to acquire data collected by a target sensor, the target sensor including a camera on the data acquisition vehicle;

[0038] The preprocessing module is used to preprocess the data collected by the target sensor to obtain data to be classified; wherein, the data to be classified includes: parameters under different dimensions corresponding to different scenarios, the data under a scenario includes the data collected by the target sensor within a preset time period, and the dimension is used to represent the components of the driving scenario;

[0039] The combination determination module is used to determine multiple dimension combinations from the dimensions corresponding to each scenario for the data to be classified.

[0040] The accuracy determination module is used to determine the accuracy of each dimension combination based on the main vehicle behavior data corresponding to each scenario included in that dimension combination.

[0041] The data processing module is used to determine the target dimension combination based on the accuracy of each dimension combination, and to construct a driving scenario data model based on the data contained in the target dimension combination.

[0042] Thirdly, embodiments of the present invention provide an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0043] Memory, used to store computer programs;

[0044] A processor, when executing a program stored in memory, implements any of the data processing method steps described above.

[0045] Fourthly, embodiments of the present invention also provide a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the data processing method steps described above.

[0046] Beneficial effects of the embodiments of the present invention:

[0047] This invention provides a data processing method, apparatus, and electronic device. The method acquires data collected by a target sensor, preprocesses the data to obtain data to be classified, and the data to be classified includes parameters under different dimensions corresponding to different scenarios. For the data to be classified, multiple dimension combinations are determined from the dimensions corresponding to each scenario. For each dimension combination, the accuracy of the dimension combination is determined based on the vehicle behavior data corresponding to each scenario included in the dimension combination. Based on the accuracy of each dimension combination, a target dimension combination is determined, and a driving scenario data model is constructed based on the data contained in the target dimension combination. By combining data from multiple scenarios with different dimensions for parameters corresponding to different scenarios, and further using the vehicle behavior data corresponding to each scenario included in the dimension combination as the evaluation basis for the accuracy of the dimension combination, the accuracy of different dimension combinations is determined. Then, the optimal scenario dimension combination data is selected based on the accuracy, and data modeling is performed using the selected optimal scenario dimension combination data. This achieves efficient and automatic data classification, solves the problems caused by existing modeling based on experience or manual methods, improves data classification efficiency, and is applicable to data classification in big data scenarios, demonstrating high applicability.

[0048] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0050] Figure 1 A flowchart illustrating a data processing method provided in an embodiment of the present invention;

[0051] Figure 2 This is a schematic diagram illustrating the dimensional display in different scenarios provided by an embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram illustrating a dimensional combination provided in an embodiment of the present invention;

[0053] Figure 4 A flowchart illustrating another data processing method provided in an embodiment of the present invention;

[0054] Figure 5 This is a schematic diagram of the structure of a data processing apparatus provided in an embodiment of the present invention;

[0055] Figure 6This is a schematic diagram of the structure of an electronic device used to implement the data processing method of the embodiments of the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of the present invention.

[0057] With the development of artificial intelligence and autonomous driving technology, autonomous driving research has entered the stage of automatic training and testing based on massive scene data. Training and testing require a large amount of scene data for support. To improve data classification efficiency for big data in driving scenarios, this invention provides a data processing method. This method involves acquiring data collected by a target sensor, preprocessing the data to obtain data to be classified, which includes parameters under different dimensions corresponding to different scenarios. For the data to be classified, multiple dimension combinations are determined from the dimensions corresponding to each scenario. For each dimension combination, the accuracy is determined based on the vehicle behavior data corresponding to each scenario included in that dimension combination. Based on the accuracy of each dimension combination, a target dimension combination is determined, and a driving scenario data model is constructed based on the data contained in the target dimension combination.

[0058] In this embodiment of the invention, for parameters under different dimensions corresponding to different scenarios, data from multiple scenarios with different dimensions can be combined. Furthermore, the vehicle behavior data corresponding to each scenario included in the dimension combination is used as the evaluation basis for the accuracy of the dimension combination to determine the accuracy of different dimension combinations. Then, the optimal scenario dimension combination data is selected based on the accuracy, and data modeling is performed using the selected optimal scenario dimension combination data. This achieves efficient and automatic data classification, solves the problems caused by existing modeling based on experience or manual methods, improves the efficiency of data classification, and is applicable to data classification in big data scenarios, with high applicability.

[0059] This invention provides a data processing method, apparatus, and electronic device applicable to scenarios such as the processing of big data in driving scenarios. For big data in driving scenarios, it achieves efficient data classification and extraction, facilitating the subsequent training and testing of autonomous driving classification models, detection models, or recognition models based on the obtained classified data.

[0060] See Figure 1 , Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of the present invention, which includes the following steps:

[0061] S101, acquire data collected by the target sensor.

[0062] In this embodiment of the invention, data in a driving scenario is collected by target sensors on a data acquisition vehicle. These data acquisition vehicles may be equipped with sensors including, but not limited to, cameras and lidar. Target sensors include: cameras, lidar, etc., on the data acquisition vehicle.

[0063] In one example, the data acquired by the target sensor is the data collected by the target sensor in the driving environment by the data acquisition vehicle. The data acquisition vehicle can be an autonomous vehicle equipped with an autonomous driving system, or it can be a vehicle without an autonomous driving system driving in an autonomous driving test environment.

[0064] S102, preprocess the data collected by the target sensor to obtain the data to be classified.

[0065] The data to be classified includes parameters under different dimensions corresponding to different scenarios. Data for a single scenario includes data collected by the target sensor within a preset duration. Dimensions represent the components of a driving scenario, such as environmental factors like road type, weather, and the status of surrounding vehicles; or the driving status of the main vehicle or target vehicle, such as speed, acceleration, straight-line driving, and turning. Different dimensions are combined to form a driving scenario. The preset duration can be set according to needs, such as 5 minutes, 10 minutes, or 30 minutes. Parameters under different dimensions represent the specific state values ​​of that dimension. For example, when the dimension is road, the parameters can be highway, urban road, etc.; when the dimension is weather, the parameters can be sunny, foggy, rainy, snowy, etc.; when the dimension is the maximum speed of the main vehicle, the parameter is the maximum speed value of the main vehicle within that time period.

[0066] In one example, data collected by sensors such as cameras and LiDAR on the data acquisition vehicle is synchronized in time, and the synchronized data is integrated to obtain fused data, which may be images or point cloud data. Further, the fused data undergoes object labeling, scene segmentation, and perception recognition. Object labeling may involve labeling pedestrians, vehicles, and buildings within the fused data; scene segmentation may involve dividing the labeled data according to a preset time interval; and perception recognition may involve perceiving and recognizing the behavior of target objects in the labeled data to generate parameters for different dimensions corresponding to different scenes, thus obtaining the data to be classified.

[0067] For example, data in a scenario could be data collected by the target sensor within a time period corresponding to a start time of 17:00 and an end time of 17:30, or data collected by the target sensor within a time period corresponding to a start time of 9:00 and an end time of 9:10, and so on.

[0068] For example, such as Figure 2 As shown, the types included in the dimension tree corresponding to the dimension in this embodiment of the invention may include: object type, target vehicle behavior, object interaction, object occlusion, environmental background, and road shape, etc. The object types can include: cars, trucks, vans, bicycles, motorcycles, and pedestrians, etc.; target vehicle behaviors can include: lateral control (such as going straight, merging, turning left or right, U-turns, lane fine-tuning, etc.) and longitudinal control (such as acceleration / deceleration, cruise control, and parking, etc.); object interactions can include: cutting in, cutting out, and following, etc. Cutting in means the vehicle in front of the data acquisition vehicle cuts in, and cutting out means the vehicle in front of the data acquisition vehicle cuts out (or changes lanes); object occlusion means the blind spot of the data acquisition vehicle, such as the data acquisition vehicle not being able to collect data on the slope while going uphill, or the data acquisition vehicle not being able to collect data on the back of the curve, etc.; environmental background can include: weather (such as cloudy, sunny, light rain, heavy snow, partly cloudy, foggy, etc.) and environment (such as open space, trees, low buildings, skyscrapers, and tunnels, etc.); road shape can include: shape (such as straight lines, curves, etc.) and lanes (such as unchanged, narrowing, widening, and ramps, etc.), etc.

[0069] For example, the parameters for different dimensions corresponding to different scenarios are shown in Table 1 below:

[0070] Table 1. Parameters under different dimensions for different scenarios.

[0071]

[0072] This embodiment of the invention merely illustrates the parameters under different dimensions corresponding to different scenarios using Table 1 above. The number of scenarios and dimensions in Table 1 does not constitute a specific limitation on this embodiment of the invention. For example, dimension A represents the maximum speed of the main vehicle, dimension B represents the average speed of the main vehicle, ..., dimension G represents whether the main vehicle makes a U-turn. Scenario 1 represents 8:00-8:10. Correspondingly, parameter 1A represents the maximum speed of the main vehicle during the time period 8:00-8:10, parameter 1B represents the average speed of the main vehicle during the time period 8:00-8:10, ..., parameter 1G represents whether the main vehicle makes a U-turn during the time period 8:00-8:10, and so on. Accordingly, a parameter under one scenario corresponds to one data point, and one data point corresponds to parameters under multiple dimensions in one scenario.

[0073] S103, for the data to be classified, determine multiple dimension combinations from the dimensions corresponding to each scenario.

[0074] In one example, the dimensions corresponding to each scenario in the data to be classified are determined, and all dimensions corresponding to the data to be classified are counted. Then, the counted dimensions are arbitrarily combined to determine multiple dimension combinations.

[0075] Assuming that all the data to be classified involves 3 dimensions (dimensions A, B, and C), then we can use permutation and combination to combine the dimensions to obtain multiple dimension combinations: (dimensions A and B), (dimensions A and C), (dimensions C and B), and (dimensions A, B, and C).

[0076] S104. For each dimension combination, determine the accuracy of the dimension combination based on the main vehicle behavior data corresponding to each scenario included in the dimension combination.

[0077] Each dimension combination contains state data (parameters) of the target objects in each scenario under that dimension combination, such as driving data of the main vehicle and the target vehicle. In one embodiment of the present invention, the accuracy of each dimension combination is evaluated based on the similarity of the main vehicle's behavior. The similarity of the main vehicle's behavior can be calculated based on the main vehicle's behavior data, such as whether the main vehicle accelerates, decelerates, or brakes.

[0078] In one example, for each dimension combination, which contains parameters corresponding to the target object in different scenarios, the parameters contained in each scenario under the dimension combination are classified or clustered. Then, based on the main vehicle behavior data corresponding to each scenario included in the dimension combination, the accuracy of the classification or clustering of the dimension combination is calculated. The accuracy of the classification or clustering is used to measure the accuracy of the scenario classification under the dimension combination. The accuracy of the classification or clustering of the dimension combination is determined as the accuracy of the dimension combination.

[0079] For example, as shown in Table 1 above, a dimension combination is a combination of two dimensions (such as dimensions A and B, A and C, A and D, ..., F and G). Dimension combination A and B contains all parameters of dimensions A and B in different scenarios (such as scenarios 1-6), and dimension combination A and C contains all parameters of dimensions A and C in different scenarios, and so on. For each dimension combination, the parameters contained in each scenario under that dimension combination are classified or clustered. Then, based on the consistency of the main vehicle behavior data corresponding to each scenario contained in that dimension combination, the accuracy of the classification or clustering of that dimension combination is calculated, and the accuracy of the classification or clustering of that dimension combination is determined as the accuracy of that dimension combination. Among them, the parameters in different scenarios contained in the dimension combination are the driving parameters in each scenario of that dimension combination.

[0080] S105. Based on the accuracy of each dimension combination, determine the target dimension combination, and construct a driving scenario data model based on the data contained in the target dimension combination.

[0081] High accuracy of dimension combination indicates high accuracy of scene classification under the corresponding dimension combination. Given the accuracy of each dimension combination, the dimension combination with the highest accuracy can be selected as the target dimension combination, or the first set number of dimension combinations with the highest accuracy in descending order can be selected as the target dimension combination. The selected target dimension combination is used as the optimal scene dimension combination data to participate in the construction of the driving scene data model.

[0082] In one example, the data contained in the determined target dimension combination includes at least: dimension combination name, accuracy, and data under different clusters of the dimension combination.

[0083] In this embodiment of the invention, for parameters under different dimensions corresponding to different scenarios, data from multiple scenarios with different dimensions can be combined. Furthermore, the vehicle behavior data corresponding to each scenario included in the dimension combination is used as the evaluation basis for the accuracy of the dimension combination to determine the accuracy of different dimension combinations. Then, the optimal scenario dimension combination data is selected based on the accuracy, and data modeling is performed using the selected optimal scenario dimension combination data. This achieves efficient and automatic data classification, solves the problems caused by existing modeling based on experience or manual methods, improves the efficiency of data classification, and is applicable to data classification in big data scenarios, with high applicability.

[0084] In one possible implementation, step S102, which preprocesses the data collected by the target sensor to obtain the data to be classified, may include:

[0085] The data collected by the target sensor is processed by format conversion and time synchronization to obtain parsed data; then the parsed data is labeled with objects to obtain labeled data; finally, the labeled data is divided into scenes and subjected to perception recognition to obtain parameters in different dimensions corresponding to different scenes.

[0086] In one example, the target sensors include cameras and LiDAR on the data acquisition vehicle. The cameras and LiDAR are used as the main sensors. Centered on the data acquisition vehicle, the data format of the data collected by the camera is converted to the data format of the data collected by the LiDAR. The data collected by the camera and the data collected by the LiDAR are time-stamped and synchronized to achieve data integration and obtain parsed data (i.e., integrated data with format conversion and time synchronization). This parsed data can be, for example, images or point cloud data.

[0087] The data collected by the target sensor is format converted and time synchronized. The resulting integrated data can improve the accuracy of data analysis. Furthermore, if one sensor fails, data collected by the other sensor can be used for data analysis, thus improving the reliability of data analysis.

[0088] Furthermore, the obtained parsed data can be labeled with objects using annotation tools or manual annotation methods to obtain labeled data. For example, objects such as pedestrians, vehicles, buildings, and static landmarks contained in the parsed data can be labeled.

[0089] In one example, a target recognition model can be used to perceive and identify objects in labeled data, and preset rules can be used to segment the labeled data into scenes to obtain parameters in different dimensions corresponding to different scenes. Perception and scene segmentation can be performed simultaneously or asynchronously. The target recognition model can be trained based on sample labeled data and sample recognition results. The preset rules can segment scenes according to a set duration, such as 5 minutes, 10 minutes, or 30 minutes, etc.

[0090] In this embodiment of the invention, the data collected by the target sensor is format converted and time synchronized. The resulting integrated data can improve the accuracy of data analysis. Furthermore, when one sensor fails, the data collected by the other sensor can be used for data analysis, which improves the reliability of data analysis. The parsed data is further annotated with objects, divided into scenes, and recognized by perception to obtain parameters in different dimensions corresponding to different scenes. This allows for the extraction of the optimal combination of scene dimensions from the parameters in different dimensions corresponding to different scenes for data modeling.

[0091] In one possible implementation, step S103 above, which determines all dimension combinations based on dimensions corresponding to different scenarios for the data to be classified, may include:

[0092] Determine the dimensions corresponding to each scenario to obtain all dimensions;

[0093] Based on all dimensions, a preset number of dimensions are combined to determine multiple dimension combinations.

[0094] In this embodiment of the invention, the dimensions corresponding to each scenario in the data to be classified are determined, and all dimensions corresponding to the data to be classified are counted to obtain all dimensions. Then, a preset number of dimensions are combined to obtain multiple dimension combinations. The preset number can be set according to actual needs or experience, such as 2-5, 2-7, 2-9, etc.

[0095] In one example, combining a preset number of dimensions for all determined dimensions can be done by permuting and combining the preset number of dimensions to obtain all possible combinations of the corresponding number of dimensions. For instance, a two-dimensional combination could be a combination of the target car's speed and the distance between the target car and the host car, while a three-dimensional combination could be a combination of the target car's acceleration, the density of surrounding cars, and the visibility of the target car, and so on.

[0096] For example, such as Figure 3 As shown, all determined dimensions include dimensions: A, B, C, ..., X, Y, Z; combinations of two dimensions include A, B, A, C, ..., Y, Z, etc.; combinations of three dimensions include A, B, C, A, C, D, ..., X, Y, Z, etc.; and combinations of seven dimensions include A, B, C, D, E, F, G, A, B, C, D, E, F, H, ..., T, U, V, W, X, Y, Z, etc.

[0097] The more dimensions in a dimension combination, the higher the accuracy of classification for that dimension combination will be. Correspondingly, the amount of computation will also be greater. However, in actual data processing, after the number of dimensions in a dimension combination reaches a certain number, further increasing the number of dimensions in the dimension combination will not increase the accuracy of the dimension combination. In order to reduce the amount of computation, in this embodiment of the invention, a preset number of dimensions are combined to determine all dimension combinations corresponding to the preset number of dimensions.

[0098] In one possible implementation, the process of determining the accuracy of each dimension combination based on the vehicle behavior data corresponding to each scenario included in the dimension combination, as described in step S104, may include:

[0099] For each dimension combination, the parameters contained in each scenario under that dimension combination are clustered to obtain multiple classes for that dimension combination;

[0100] For each class, the accuracy of the class is determined based on the main vehicle behavior data corresponding to each scenario included in that class.

[0101] The accuracy of the dimension combination is determined based on the accuracy of each class in the dimension combination.

[0102] In one example, for each dimension combination, a target clustering algorithm can be used to cluster the parameters contained in each scenario under that dimension combination, resulting in multiple classes for that dimension combination. Target clustering algorithms can include, but are not limited to: K-means (k-means clustering algorithm), DBSCAN (Density-Based Spatial Clustering of Applications with Noise), ISODATA (Iterative Selforganizing Data Analysis Techniques Algorithm), and BIRCH (Balanced Iterative Reducing and Clustering using Hierarchies), etc.

[0103] For example, a combination of certain dimensions (including scenarios 1 to 20) may be clustered into two categories. The clustering results show that scenarios 1 to 10 are classified into the first category, and scenarios 11 to 20 are classified into the second category.

[0104] For each class in this dimension combination, the consistency of the main vehicle behavior data corresponding to each scenario included in the class is calculated. The degree of consistency of the main vehicle behavior data corresponding to each scenario included in the class is determined as the accuracy of the class. Then, the accuracy of the dimension combination is calculated by combining the accuracy of each class in the dimension combination, so as to evaluate the clustering effect of the parameters included in each scenario under the dimension combination.

[0105] In one example, the accuracy of the dimension combination can be determined by the weighted value, summation, or average of the accuracy of each class in the dimension combination.

[0106] In this embodiment of the invention, for each dimension combination, the parameters contained in each scenario under the dimension combination are clustered using a target clustering algorithm to obtain multiple classes. The consistency of the main vehicle behavior data corresponding to each scenario included in each class is used as the evaluation benchmark to evaluate the accuracy of clustering of each class in the dimension combination. Furthermore, the accuracy of various clustering is combined to evaluate the clustering effect of the parameters contained in each scenario under the dimension combination.

[0107] In one possible implementation, since the clustering algorithms applicable to different data types may not be entirely consistent, the above-mentioned clustering of the parameters contained in each scenario under each dimension combination to obtain multiple classes for that dimension combination can include: for each dimension combination, using a clustering algorithm corresponding to the data category of that dimension combination to cluster the parameters contained in each scenario under that dimension combination to obtain multiple classes for that dimension combination. Clustering algorithms can include K-means, DBSCAN, ISODATA, and BIRCH, etc.

[0108] In one example, for each dimension combination, different clustering algorithms are used to cluster the parameters included in each scenario under that dimension combination, resulting in multiple clusters for that dimension combination under different clustering algorithms. For each cluster under different clustering algorithms, the accuracy of that cluster is determined based on the vehicle behavior data corresponding to each scenario included in that cluster, resulting in the accuracy of each cluster under different clustering algorithms. Then, for each clustering algorithm, the weighted value, sum, or average of the accuracy of each cluster under that clustering algorithm is determined, resulting in the accuracy of that dimension combination under that clustering algorithm. The clustering algorithm with the highest accuracy is determined as the clustering algorithm corresponding to the data category of that dimension combination. Then, the parameters included in each scenario under that dimension combination are clustered using the clustering algorithm corresponding to the data category of that dimension combination, resulting in multiple clusters for that dimension combination.

[0109] Accordingly, the data included in the target dimension combination determined above shall include at least: dimension combination name, clustering algorithm name, number of clusters, accuracy, and data under different clusters of the dimension combination.

[0110] When the clustering results show that a certain class contains too few scenes, it may affect the accuracy of the clustering. Therefore, in one possible implementation, after obtaining multiple classes of this dimension combination, the following may also be included:

[0111] For multiple classes in this dimension combination, determine whether there are any classes that contain fewer than a preset threshold number of scenarios;

[0112] If there is a class containing fewer than a preset threshold number of scenarios, delete the class containing fewer than the preset threshold number of scenarios and the parameters contained in each scenario under that class, and re-cluster the parameters contained in each scenario under that dimension combination to obtain multiple classes for that dimension combination.

[0113] In this embodiment of the invention, for each dimension combination, the parameters contained in each scene under that dimension combination are clustered to obtain multiple classes for that dimension combination. Then, it is further determined whether there are any classes containing fewer than a preset threshold number of scenes. If such a class exists, it indicates that it belongs to noise in machine learning clustering. At this point, the classes containing fewer than the preset threshold number of scenes and the parameters contained in each scene under that class are deleted to remove noise. The parameters contained in each scene under that dimension combination are then re-clustered to obtain multiple classes for that dimension combination. The preset threshold number can be set based on the number of scenes contained in the dimension combination; for example, the preset threshold number can be 0.5%, 1%, or 2% of the number of scenes contained in the dimension combination, etc.

[0114] To evaluate the effectiveness of driving scenario modeling and to determine the merits of dimension combinations and clustering algorithm parameters, this embodiment of the invention calculates the accuracy of each class in the dimension combination, and then determines the accuracy of the dimension combination by using the accuracy of each class in the dimension combination.

[0115] In one possible implementation, the ability to cluster scenarios with similar main vehicle behaviors into the same class is used as a criterion for evaluating the clustering effect. The similarity of main vehicle behaviors can be calculated based on the main vehicle behavior data.

[0116] In one possible implementation, the process of determining the accuracy of each class based on the main vehicle behavior data corresponding to each scenario included in that class may include:

[0117] For each class, obtain the number of the first scenarios that contain the preset main vehicle behavior in each scenario under that class;

[0118] For each class, obtain the number of second scenarios under that class that do not contain the preset main vehicle behavior;

[0119] The accuracy of the class is determined based on the number of the first scene and the number of the second scene.

[0120] In one possible implementation, the preset behavior of the main vehicle includes: emergency braking of the main vehicle.

[0121] In one example, for each class, the acceleration at the trigger time of the main vehicle state is obtained from the parameters of each scene within that class. The trigger time of the main vehicle state can be the time point corresponding to a change in the main vehicle speed, etc. For each scene in that class, it is determined whether the acceleration at the trigger time of the main vehicle state is greater than a preset acceleration threshold. If so, the scene is determined to contain a sudden braking of the main vehicle; otherwise, it is determined not to contain a sudden braking of the main vehicle. The number of scenes containing a sudden braking of the main vehicle within that class is taken as the first scene count, and the number of scenes not containing a sudden braking of the main vehicle within that class is taken as the second scene count. Further, the proportion of the first scene count in that class is taken as the first proportion, and the proportion of the second scene count in that class is taken as the second proportion. The sum of the first and second proportions is 1. The higher of the first and second proportions is determined as the accuracy of that class. The preset acceleration threshold can be set according to actual conditions; for example, the preset acceleration threshold could be -2.5 m / s², -5 m / s², or -8 m / s², etc.

[0122] In this embodiment of the invention, the presence of sudden braking by the main vehicle in each scenario within each category of the dimension combination is used as the criterion for judging the similarity of the main vehicle's behavior. The accuracy of each category in the dimension combination is calculated, and the weighted value, summation value, or average value of the accuracy of each category in the dimension combination is further determined as the accuracy of the dimension combination to evaluate the clustering effect of the dimension combination. Of course, other main vehicle behaviors, such as rapid acceleration and lane changing, can also be used as a representative of a main vehicle's behavior.

[0123] In one possible implementation, the process of determining the accuracy of each class based on the main vehicle behavior data corresponding to each scenario included in that class may include:

[0124] For each class, obtain the main vehicle's sequential behavior characteristics that occur in the time sequence within each scenario of that class.

[0125] The percentage corresponding to the most frequently occurring main vehicle series behavior feature in this category is determined as the accuracy of this category.

[0126] In one example, for each class, the vehicle behavior data for each scenario within that class is obtained. This vehicle behavior data can include acceleration, deceleration, cruising, or stopping. For each scenario within that class, the vehicle behaviors occurring in chronological order within the corresponding time period are concatenated to obtain the sequential vehicle behavior features occurring in chronological order within that scenario. Further, the most frequently occurring sequential vehicle behavior feature in that class is selected as the target behavior feature, and the proportion of the target behavior feature in that class is calculated. This proportion is determined as the accuracy of that class; that is, the proportion corresponding to the most frequently occurring sequential vehicle behavior feature in that class is determined as the accuracy of that class. Identical sequential vehicle behavior features indicate consistent vehicle behavior.

[0127] For example, acceleration is represented by the field `acc`, deceleration by the field `dec`, cruising by the field `acc_dec`, and parking by the field `park`. Suppose that machine learning clustering methods are used to cluster the parameters of a single-dimensional combination of 1000 scenarios, resulting in 10 classes. For instance, one class might contain 100 scenarios. Correspondingly, each scenario can obtain a concatenated field representing the primary vehicle's behavior, resulting in a primary vehicle's sequential behavior feature. Typically, the primary vehicle's sequential behavior features within a class are not entirely identical. For example, in a class containing 100 scenarios, 98 primary vehicle sequential behavior features are `dec_acc`, and 2 are `dec_park`. In this case, the primary vehicle sequential behavior feature `dec_acc` is determined as the target behavior feature for that class, and the percentage (98%) of the target behavior feature `dec_acc` in that class is determined as the accuracy of that class. For example, in a class containing 100 scenarios, 80 main vehicle series behavior features are dec_park, 10 main vehicle series behavior features are acc_dec_park, and 10 main vehicle series behavior features are dec_acc. Then, the main vehicle series behavior feature dec_park is determined as the target behavior feature of this class, and the proportion (80%) of the target behavior feature dec_park in this class is determined as the accuracy of this class, etc.

[0128] In this embodiment of the invention, the consistency of the main vehicle's serial behavior features corresponding to the main vehicle's behavior data in each scenario is used as the basis for judging the similarity of the main vehicle's behavior. The accuracy of each class in the dimension combination is calculated, and the weighted value, summation value or average value of the accuracy of each class in the dimension combination is further determined as the accuracy of the dimension combination in order to evaluate the clustering effect of the dimension combination.

[0129] In one possible implementation, the process of determining the accuracy of each class based on the main vehicle behavior data corresponding to each scenario included in that class may include:

[0130] For each class, the accuracy of the class is determined based on the preset master vehicle behavior and master vehicle serial behavior characteristics corresponding to each scenario included in the class.

[0131] In one example, the preset main vehicle behavior includes emergency braking. For each class, the acceleration at the trigger time of the main vehicle state is obtained from the parameters of each scenario within that class. For each scenario within that class, it is determined whether the acceleration at the trigger time of the main vehicle state is greater than a preset acceleration threshold to determine whether the scenario includes or does not include emergency braking. The higher percentage between the proportion of scenarios including emergency braking and the proportion of scenarios not including emergency braking is determined as the first accuracy value for that class. Furthermore, for each class, the main vehicle behavior data for each scenario within that class is obtained. For each scenario within that class, the main vehicle behaviors occurring in chronological order within the corresponding time period are concatenated to obtain the main vehicle sequential behavior features occurring in chronological order within the corresponding time period. The percentage corresponding to the most frequent main vehicle sequential behavior feature in the class is determined as the second accuracy value for that class. Finally, the weighted value, summation, or average of the first and second accuracy values ​​for that class is determined as the accuracy of that class.

[0132] In this embodiment of the invention, the similarity of the main vehicle behavior is calculated from two perspectives: whether there is a preset main vehicle behavior in each scenario and whether the main vehicle's serial behavior features are consistent. This measures the accuracy of the combination of each dimension and achieves a multi-dimensional evaluation of the clustering effect of the combination of each dimension.

[0133] In one possible implementation, the above method may further include:

[0134] Use the data contained in the driving scenario data model to train an autonomous driving classification model.

[0135] In this embodiment of the invention, after determining the target dimension combination and constructing a driving scenario data model based on the data contained in the target dimension combination, the data contained in the driving scenario data model can be further used to train autonomous driving classification models, target recognition models, etc.

[0136] In one example, as the amount of data collected by the target sensor increases, the data contained in the target dimension combination can be directly selected for the required data classification, scene classification, and training of autonomous driving classification models.

[0137] For example, such as Figure 4 As shown, another data processing method provided by an embodiment of the present invention may include the following steps:

[0138] S401, Acquire data collected by the target sensor, which includes: a camera on the data acquisition vehicle;

[0139] S402 performs format conversion and time synchronization processing on the data collected by the target sensor to obtain parsed data;

[0140] S403, Perform object annotation on the parsed data to obtain the annotated data;

[0141] S404 performs scene segmentation and perception recognition on the labeled data to obtain parameters in different dimensions corresponding to different scenes;

[0142] S405, determine multiple dimension combinations from the dimensions corresponding to each scenario;

[0143] Based on the dimensions corresponding to each scenario, all dimensions are first counted to obtain the total dimensions. Based on the total dimensions, a preset number of dimensions are combined to determine multiple dimension combinations.

[0144] S406, For each dimension combination, use the clustering algorithm corresponding to the data category of that dimension combination to cluster the parameters contained in each scenario under that dimension combination, and obtain multiple classes for that dimension combination;

[0145] S407, for each class, determine the accuracy of the class based on the main vehicle behavior data corresponding to each scenario included in the class;

[0146] S408, Determine the accuracy of the dimension combination based on the accuracy of each class in the dimension combination;

[0147] S409 determines the target dimension combination based on the accuracy of each dimension combination, and constructs a driving scenario data model based on the data contained in the target dimension combination.

[0148] In one example, the data contained in the determined target dimension combination includes at least: dimension combination name, clustering algorithm name, number of clusters, accuracy, and data under different clusters of the dimension combination.

[0149] In this embodiment of the invention, for parameters under different dimensions corresponding to different scenarios, data from multiple scenarios with different dimensions can be combined. Furthermore, the vehicle behavior data corresponding to each scenario included in the dimension combination is used as the evaluation basis for the accuracy of the dimension combination to determine the accuracy of different dimension combinations. Then, the optimal scenario dimension combination data is selected based on the accuracy, and data modeling is performed using the selected optimal scenario dimension combination data. This achieves efficient and automatic data classification, solves the problems caused by existing modeling based on experience or manual methods, improves the efficiency of data classification, and is applicable to data classification in big data scenarios, with high applicability.

[0150] This invention also provides a data processing apparatus, such as... Figure 5 As shown, the device includes:

[0151] Data acquisition module 501 is used to acquire data collected by a target sensor, including a camera on the data acquisition vehicle.

[0152] The preprocessing module 502 is used to preprocess the data collected by the target sensor to obtain the data to be classified; wherein, the data to be classified includes: parameters under different dimensions corresponding to different scenarios, the data under a scenario includes the data collected by the target sensor within a preset time period, and the dimension is used to represent the components of the driving scenario;

[0153] The combination determination module 503 is used to determine multiple dimension combinations from the dimensions corresponding to each scenario for the data to be classified.

[0154] The accuracy determination module 504 is used to determine the accuracy of each dimension combination based on the main vehicle behavior data corresponding to each scenario included in the dimension combination.

[0155] The data processing module 505 is used to determine the target dimension combination based on the accuracy of each dimension combination, and to construct a driving scenario data model based on the data contained in the target dimension combination.

[0156] In this embodiment of the invention, for parameters under different dimensions corresponding to different scenarios, data from multiple scenarios with different dimensions can be combined. Furthermore, the vehicle behavior data corresponding to each scenario included in the dimension combination is used as the evaluation basis for the accuracy of the dimension combination to determine the accuracy of different dimension combinations. Then, the optimal scenario dimension combination data is selected based on the accuracy, and data modeling is performed using the selected optimal scenario dimension combination data. This achieves efficient and automatic data classification, solves the problems caused by existing modeling based on experience or manual methods, improves the efficiency of data classification, and is applicable to data classification in big data scenarios, with high applicability.

[0157] In one possible implementation, the combination determination module 503 described above is specifically used for:

[0158] Determine the dimensions corresponding to each scenario to obtain all dimensions;

[0159] Based on all dimensions, a preset number of dimensions are combined to determine multiple dimension combinations.

[0160] In one possible implementation, the accuracy determination module 504 includes:

[0161] The clustering submodule is used to cluster the parameters contained in each scenario under each dimension combination to obtain multiple classes for that dimension combination.

[0162] The first determination submodule is used to determine the accuracy of each class based on the main vehicle behavior data corresponding to each scenario included in that class.

[0163] The second determining submodule is used to determine the accuracy of the dimension combination based on the accuracy of each class in the dimension combination.

[0164] In one possible implementation, the clustering submodule is specifically used to cluster the parameters contained in each scenario under each dimension combination using a clustering algorithm corresponding to the data category of that dimension combination, thereby obtaining multiple classes for that dimension combination.

[0165] In one possible implementation, the above-described apparatus further includes a repeat clustering submodule.

[0166] The repeated clustering submodule is used to determine whether there are classes containing fewer than a preset threshold number of scenes for multiple classes of this dimension combination; if there are classes containing fewer than the preset threshold number of scenes, the classes containing fewer than the preset threshold number of scenes and the parameters contained in each scene under the class are deleted, and the parameters contained in each scene under the dimension combination are re-clustered to obtain multiple classes of this dimension combination.

[0167] In one possible implementation, the first determining submodule is specifically used for:

[0168] For each class, obtain the number of the first scenarios that contain the preset main vehicle behavior in each scenario under that class;

[0169] For each class, obtain the number of second scenarios under that class that do not contain the preset main vehicle behavior;

[0170] The accuracy of the class is determined based on the number of the first scene and the number of the second scene.

[0171] In one possible implementation, the aforementioned preset vehicle behavior includes: emergency braking of the vehicle.

[0172] In one possible implementation, the first determining submodule is specifically used for:

[0173] For each class, obtain the main vehicle serial behavior characteristics that occur in sequence within the time period corresponding to each scenario under that class;

[0174] The accuracy of a class is determined by the proportion of the most frequently occurring main vehicle series behavior feature.

[0175] In one possible implementation, the first determining submodule described above is specifically used for:

[0176] For each class, the accuracy of the class is determined based on the preset master vehicle behavior and master vehicle serial behavior characteristics corresponding to each scenario included in the class.

[0177] In one possible implementation, the above-described apparatus further includes:

[0178] The model training module is used to train an autonomous driving classification model using the data contained in the driving scenario data model.

[0179] This invention also provides an electronic device, such as... Figure 6 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.

[0180] Memory 603 is used to store computer programs;

[0181] When the processor 601 executes the program stored in the memory 603, it implements the steps of any of the data processing methods described above to achieve the same technical effect.

[0182] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0183] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0184] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0185] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0186] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and the computer program, when executed by a processor, implements the steps of any of the above methods.

[0187] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the steps of any of the methods described in the above embodiments.

[0188] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0189] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0190] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device / electronic device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0191] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A data processing method, characterized in that, The method includes: Acquire data collected by a target sensor, wherein the target sensor includes a camera on the data acquisition vehicle; The data collected by the target sensor is preprocessed to obtain data to be classified; wherein, the data to be classified includes: parameters under different dimensions corresponding to different scenarios, the data under a scenario includes the data collected by the target sensor within a preset time period, and the dimension is used to represent the components of the driving scenario; For the data to be classified, multiple dimension combinations are determined from the dimensions corresponding to each scenario; For each dimension combination, the accuracy of that dimension combination is determined based on the main vehicle behavior data corresponding to each scenario included in that dimension combination. Based on the accuracy of each dimension combination, a target dimension combination is determined, and a driving scenario data model is constructed based on the data contained in the target dimension combination. Specifically, for each dimension combination, determining the accuracy of that dimension combination based on the vehicle behavior data corresponding to each scenario included in that dimension combination includes: For each dimension combination, the parameters contained in each scenario under that dimension combination are clustered to obtain multiple classes for that dimension combination; For each class, the accuracy of the class is determined based on the main vehicle behavior data corresponding to each scenario included in the class; the accuracy of the class is determined based on the consistency of the main vehicle behavior data corresponding to each scenario included in the class. The accuracy of the dimension combination is determined based on the accuracy of each class in the dimension combination.

2. The method according to claim 1, characterized in that, For the data to be classified, multiple dimension combinations are determined from the dimensions corresponding to each scenario, including: Determine the dimensions corresponding to each scenario to obtain all dimensions; Based on all the dimensions, a preset number of dimensions are combined to determine multiple dimension combinations.

3. The method according to claim 1, characterized in that, For each dimension combination, the parameters contained in each scenario under that dimension combination are clustered to obtain multiple classes for that dimension combination, including: For each dimension combination, a clustering algorithm corresponding to the data category of that dimension combination is used to cluster the parameters contained in each scenario under that dimension combination, resulting in multiple classes for that dimension combination.

4. The method according to claim 1, characterized in that, After obtaining multiple classes of this dimension combination, the method further includes: For multiple classes in this dimension combination, determine whether there are any classes that contain fewer than a preset threshold number of scenarios; If there is a class containing fewer than a preset threshold number of scenarios, delete the class containing fewer than the preset threshold number of scenarios and the parameters contained in each scenario under that class, and re-cluster the parameters contained in each scenario under that dimension combination to obtain multiple classes for that dimension combination.

5. The method according to claim 1 or 3, characterized in that, For each class, the accuracy of that class is determined based on the main vehicle behavior data corresponding to each scenario included in that class, including: For each class, obtain the number of the first scenarios that contain the preset main vehicle behavior in each scenario under that class; For each class, obtain the number of second scenarios under that class that do not contain the preset master vehicle behavior; The accuracy of the class is determined based on the first number of scenarios and the second number of scenarios.

6. The method according to claim 5, characterized in that, The preset main vehicle behavior includes: emergency braking of the main vehicle.

7. The method according to claim 1 or 3, characterized in that, For each class, the accuracy of that class is determined based on the main vehicle behavior data corresponding to each scenario included in that class, including: For each class, obtain the main vehicle's sequential behavior characteristics that occur in the time sequence within each scenario of that class. The percentage corresponding to the most frequently occurring main vehicle series behavior feature in this category is determined as the accuracy of this category.

8. The method according to claim 1 or 3, characterized in that, For each class, the accuracy of that class is determined based on the main vehicle behavior data corresponding to each scenario included in that class, including: For each class, the accuracy of the class is determined based on the preset master vehicle behavior and master vehicle serial behavior characteristics corresponding to each scenario included in the class.

9. A data processing apparatus, characterized in that, The device includes: The data acquisition module is used to acquire data collected by a target sensor, the target sensor including a camera on the data acquisition vehicle; The preprocessing module is used to preprocess the data collected by the target sensor to obtain data to be classified; wherein, the data to be classified includes: parameters under different dimensions corresponding to different scenarios, the data under a scenario includes the data collected by the target sensor within a preset time period, and the dimension is used to represent the components of the driving scenario; The combination determination module is used to determine multiple dimension combinations from the dimensions corresponding to each scenario for the data to be classified. The accuracy determination module is used to determine the accuracy of each dimension combination based on the main vehicle behavior data corresponding to each scenario included in that dimension combination. The data processing module is used to determine the target dimension combination based on the accuracy of each dimension combination, and to construct a driving scenario data model based on the data contained in the target dimension combination. The accuracy determination module includes: The clustering submodule is used to cluster the parameters contained in each scenario under each dimension combination to obtain multiple classes for that dimension combination. The first determination submodule is used to determine the accuracy of each class based on the main vehicle behavior data corresponding to each scenario included in that class; the accuracy of the class is determined based on the consistency of the main vehicle behavior data corresponding to each scenario included in that class. The second determining submodule is used to determine the accuracy of the dimension combination based on the accuracy of each class in the dimension combination.

10. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Information processing method and device, electronic equipment and storage medium

    CN110826616A

  • Method for building virtual scenario library for autonomous vehicle

    US20210197851A1