A data sampling and model training method, device, equipment and storage medium
By clustering and splitting the geographical locations to be sampled, combining regional area and random sampling, the problem of insufficient balance in data sampling distribution is solved, and the training effect of the autonomous driving technology algorithm model is significantly improved.
Patent Information
- Application Number
- CN202210356430.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-04-06
AI Technical Summary
In the prior art, the data sampling distribution is poor, resulting in poor training of neural network models involved in autonomous driving technology algorithms.
By initially classifying all sampling points of the geographical location to be sampled, multiple initial clusters are obtained and further split into multiple subclusters. Calculate the area of the area surrounded by each sampling point of each subcluster, determine the total amount of data to be sampled within each subcluster, and determine the final sampling result through multiple random sampling and distribution equality calculations.
It effectively expands the coverage of data sampling, improves the distribution balance of collected data samples, and thus improves the training effect of neural network models involved in autonomous driving technology algorithms.
Smart Images

Figure CN114692769B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data sampling, and in particular, to a data sampling and model training method, device, equipment, and storage medium. Background Art
[0002] With the development of science and technology, autonomous driving technology has also developed rapidly. In the field of autonomous driving technology, there are many autonomous driving technology algorithms, and sometimes deep neural network models are needed to implement these autonomous driving technology algorithms. In order to better implement some of the autonomous driving technology algorithms involved in autonomous driving technology, it is necessary to collect a large number of data samples of geographical locations in a large range as training data to train the neural network models involved in some of the autonomous driving technology algorithms.
[0003] For example, when training an identification model that can determine whether a vehicle needs to change lanes based on the states of obstacles such as vehicles and pedestrians around the driving environment, it may be necessary to collect different data samples for training the identification model according to different driving environment scenarios. Because in different geographical locations, the driving environment scenarios are different: for example, in some places, the traffic flow is greater, in some places, there may be large trucks, and in some places, there may be more pedestrians. Therefore, if the data sampling range can cover a larger geographical location range when collecting the data samples required for training data, more abundant data samples of driving environment scenarios can be collected. The richer the driving environment scenarios of the data samples, the better the training effect of the identification model can be improved.
[0004] In the actual application process, for the data samples required for some neural network models involved in some autonomous driving technology algorithms, the data samples collected are often sparse because the number of vehicles in the collected data samples is limited and the driving environment scenarios are not rich enough. The lack of richness in the driving environment scenarios of the collected data samples leads to poor training effects of some neural network models involved in autonomous driving technology algorithms.
[0005] Therefore, how to expand the coverage range of data collection and improve the distribution balance of the collected data samples has been a problem that people have been concerned about. Summary of the Invention
[0006] This application aims to at least solve one of the above technical defects. In view of this, this application provides a data sampling and model training method, device, equipment, and storage medium for solving the technical defect of poor distribution balance in data sampling in the prior art.
[0007] A data sampling method includes:
[0008] Preliminarily classify all sampling points of the geographical location to be sampled to obtain multiple initial clusters;
[0009] Split each of the initial clusters to obtain multiple sub-clusters;
[0010] Calculate the area of the region enclosed by the sampling points of each sub-cluster;
[0011] Based on the area of the region enclosed by the sampling points of each sub-cluster, determine the total amount of data to be sampled within each sub-cluster;
[0012] Based on the total amount of data to be sampled in each sub-cluster, perform multiple random samplings on each sub-cluster;
[0013] Calculate the distribution balance of each random sampling of each sub-cluster, and determine the sampling result with the most balanced distribution as the final sampling result of each sub-cluster.
[0014] Preferably, the step of preliminarily classifying all sampling points of the geographical location to be sampled to obtain multiple initial clusters includes:
[0015] Randomly determine a sampling point in the geographical location to be sampled as the first sampling point;
[0016] Classify the sampling points whose straight-line distance from the first sampling point is less than a preset first distance and the first sampling point into one initial cluster;
[0017] Classify the sampling points outside the initial cluster whose straight-line distance from each sampling point in the initial cluster is less than the preset first distance into the initial cluster until the straight-line distance between each sampling point outside the initial cluster and each sampling point inside the initial cluster is equal to or greater than the preset first distance;
[0018] Randomly select a sampling point from the sampling points outside the initial cluster as the first sampling point, and return to execute the operation of classifying the sampling points whose straight-line distance from the first sampling point is less than the preset first distance and the first sampling point into one initial cluster until all sampling points in the geographical location to be sampled are classified into each initial cluster.
[0019] Preferably, the step of classifying the sampling points outside the initial cluster whose straight-line distance from each sampling point in the initial cluster is less than the preset first distance into the initial cluster until the straight-line distance between each sampling point outside the initial cluster and each sampling point inside the initial cluster is equal to or greater than the preset first distance includes:
[0020] Calculate the straight-line distances between each sampling point in the initial cluster and each sampling point outside the initial cluster respectively;
[0021] Determine whether there are sampling points outside the initial cluster whose straight-line distances to each sampling point within the initial cluster are less than a preset first distance;
[0022] If there are sampling points outside the initial cluster whose straight-line distances to each sampling point within the initial cluster are less than the preset first distance, then classify the sampling points outside the initial cluster and whose straight-line distances to each sampling point of the initial cluster are less than the preset first distance into the initial cluster until the straight-line distances between each sampling point outside the initial cluster and each sampling point within the initial cluster are all equal to or greater than the preset first distance.
[0023] Preferably, the step of splitting each initial cluster respectively to obtain a plurality of sub-clusters includes:
[0024] Determine whether each initial cluster meets a preset splitting condition;
[0025] If an initial cluster meets the preset splitting condition, then use the initial cluster that meets the preset splitting condition as the cluster to be split;
[0026] Split the cluster to be split to obtain a plurality of sub-clusters;
[0027] Determine whether the split sub-clusters meet the preset splitting condition;
[0028] If the split sub-clusters meet the preset splitting condition, then use the sub-clusters that meet the preset splitting condition as the clusters to be split, and return to execute the operation of splitting the clusters to be split to obtain a plurality of sub-clusters until each split sub-cluster does not meet the preset splitting condition.
[0029] Preferably, the preset splitting condition includes:
[0030] If the maximum value of the straight-line distances between any two sampling points within each cluster is equal to or greater than a preset second distance, then it is considered that the cluster meets the preset splitting condition.
[0031] Preferably, the step of splitting the cluster to be split to obtain a plurality of sub-clusters includes:
[0032] Calculate the straight-line distances between any two sampling points within each cluster to be split respectively;
[0033] Determine the two sampling points with the maximum straight-line distance between any two sampling points within each cluster to be split as target sampling points;
[0034] Connect a straight line between the two target sampling points within each cluster to be split;
[0035] Calculate a target ratio between a straight-line distance between the target sampling points and a preset second distance;
[0036] Determine a plurality of central points on a straight line connecting two target sampling points according to the target ratio, such that a distance between two adjacent central points on the straight line connecting the two target sampling points is less than or equal to the preset second distance;
[0037] Draw a perpendicular line at each central point on the straight line connecting the target sampling points, and split the clustering to be split with each perpendicular line as a dividing line to obtain a plurality of sub-clusterings.
[0038] Preferably, determining the total amount of data to be sampled in each sub-clustering based on the area of the region surrounded by the sampling points of each sub-clustering includes:
[0039] Calculate the straight-line distance between any two sampling points within each sub-clustering;
[0040] Calculate the area of a circle with the maximum value of the straight-line distances between any two sampling points within each sub-clustering as the diameter, as the area of the region surrounded by the sampling points of each sub-clustering;
[0041] Accumulate the areas of the regions surrounded by the sampling points of each sub-clustering to obtain the total area of the geographical location to be sampled;
[0042] Calculate the ratio of the area of the region surrounded by the sampling points of each sub-clustering to the total area of the geographical location to be sampled;
[0043] Determine the total amount of data to be sampled in each sub-clustering based on the ratio of the area of the region surrounded by the sampling points of each sub-clustering to the total area of the geographical location to be sampled and a preset total amount of sampling data.
[0044] Preferably, calculating the distribution balance of each random sampling of each sub-clustering and determining the most balanced random sampling result as the final sampling result of each sub-clustering includes:
[0045] Calculate the straight-line distance between any two sampling points among the sampling points of each random sampling of each sub-clustering;
[0046] Calculate the sum of the straight-line distances between any two sampling points among the sampling points of each random sampling of each sub-clustering;
[0047] Take the random sampling result with the maximum sum of the straight-line distances between any two sampling points among the sampling points of each random sampling as the final sampling result of each sub-clustering.
[0048] A model training method includes:
[0049] Use the data collected by any of the data sampling methods described above as training data;
[0050] Use the training data to train a deep neural network model to obtain a trained unmanned vehicle control model, which is used to process one or more of the perception, prediction, and decision-making tasks of an unmanned vehicle.
[0051] A data sampling device, comprising:
[0052] A clustering unit for preliminarily classifying all sampling points of a geographical location to be sampled to obtain a plurality of initial clusters;
[0053] A splitting unit for splitting each initial cluster respectively to obtain a plurality of sub-clusters;
[0054] An area calculation unit for calculating the area of the region enclosed by the sampling points of each sub-cluster;
[0055] A data determination unit for determining the total amount of data to be sampled in each sub-cluster based on the area of the region enclosed by the sampling points of each sub-cluster;
[0056] A sampling unit for performing multiple random samplings on each sub-cluster based on the total amount of data to be sampled in each sub-cluster;
[0057] An equilibrium calculation unit for calculating the distribution equilibrium of each random sampling of each sub-cluster, and determining the most evenly distributed sampling result as the final sampling result of each sub-cluster.
[0058] A data sampling device, comprising: one or more processors, and a memory;
[0059] The memory stores computer-readable instructions, which, when executed by the one or more processors, implement the steps of the data sampling method described in any of the above and the model training method described above.
[0060] A storage medium stores computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to implement the steps of the data sampling method described in any of the above and the model training method described above.
[0061] As can be seen from the above technical solutions, the embodiments of the present application can preliminarily classify all sampling points of the geographical location to be sampled, and through the preliminary classification, the sampling range of the geographical location to be sampled can be initially divided. After obtaining multiple initial clusters, each initial cluster can be split to obtain multiple sub-clusters; splitting the multiple initial clusters can better further subdivide the sampling range with a larger sampling point density, which can help improve the distribution balance of the sampling data. After obtaining multiple sub-clusters, the area of the region enclosed by each sampling point in each sub-cluster can be calculated; based on the area of the region enclosed by each sampling point in each sub-cluster, the total amount of sampled data in each sub-cluster can be determined; finally, based on the total amount of data to be sampled in each sub-cluster, multiple random samplings can be performed on each sub-cluster; the distribution balance of each random sampling in each sub-cluster is calculated respectively, and the sampling result with the most balanced distribution is determined as the final sampling result of each sub-cluster.
[0062] After the clustering result of the geographical location to be sampled in the embodiments of the present application is split and refined into multiple sub-clusters, according to the area of the region enclosed by each sampling point in each sub-cluster, the total amount of sampled data in each sub-cluster can be determined, effectively expanding the coverage range of data sampling. Secondly, multiple random samplings are performed for each sub-cluster, the distribution balance of each random sampling in each sub-cluster is calculated, and the sampling result with the most balanced distribution is determined as the final sampling result of each sub-cluster. This effectively improves the distribution balance of the sampled data in each sub-cluster and the overall distribution balance of the total amount of sampled data. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0064] Figure 1 It is a flowchart of a method for implementing data sampling provided by the embodiments of the present application;
[0065] Figure 2 It is a schematic structural diagram of a data sampling device exemplified by the embodiments of the present application;
[0066] Figure 3 It is a hardware structure block diagram of a data sampling device disclosed by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0068] Next, in conjunction with Figure 1 , the process of the data sampling method given in the embodiments of the present application will be introduced. As Figure 1 shown, this process may include the following steps:
[0069] Step S101, preliminarily classify all sampling points of the geographical location to be sampled to obtain multiple initial clusters.
[0070] Specifically, after determining the geographical location of the data to be sampled, all sampling points of the geographical location to be sampled can be preliminarily classified to divide the geographical location to be sampled into multiple initial clusters.
[0071] Dividing each sampling point of the geographical location to be sampled into multiple initial clusters can ensure that the sampling data is collected according to the balance of the geographical distribution, avoiding missing the areas where the sampling points are relatively sparse; it helps to improve the balance of the sampling data distribution in the area surrounded by each sampling point of each initial cluster.
[0072] Step S102, split each initial cluster to obtain multiple sub-clusters.
[0073] Specifically, after preliminarily classifying each sampling point of the geographical location to be sampled to obtain multiple initial clusters, since the sampling point density of some initial clusters is relatively large, if data is directly collected in the area with a large sampling point density, it is possible to collect more data in the initial cluster with a large sampling point density. However, since the total amount of sample data to be collected is already determined, this will result in a reduction in the sampling quantity at other geographical locations where the sampling points are relatively sparse or direct omission, which is not conducive to ensuring the balance of the sampling data distribution. Therefore, each initial cluster obtained from the initial classification can be further split into smaller sub-clusters before collecting data.
[0074] Step S103, calculate the area of the region surrounded by each sampling point of each sub-cluster.
[0075] Specifically, the area enclosed by each sampling point of each sub-cluster after splitting is a relatively small area unit with respect to the geographical location to be sampled. If data can be collected in each sub-cluster, the distribution balance of the data collected in each sub-cluster is obviously more balanced than that of directly collecting data in the initial cluster with a larger geographical location. However, the number of sampling points of each sub-cluster after splitting is different, so the area of the region enclosed by each sampling point of each sub-cluster will also be different. Since the total amount of data to be collected is already determined, although the geographical location to be sampled has been divided into sub-clusters with approximately the same area and data is collected at the sampling points of each sub-cluster, the amount of sampled data of each sub-cluster will still be different. To further improve the distribution balance of the data of each sub-cluster, it is necessary to determine the total amount of data to be sampled at the sampling points of each sub-cluster. Therefore, the area of the region enclosed by each sampling point of each sub-cluster after splitting can be further calculated, so that the total amount of data to be sampled at the sampling points of each sub-cluster can be determined based on the area of the region enclosed by each sampling point of each sub-cluster after splitting.
[0076] Step S104: Determine the total amount of data to be sampled in each sub-cluster based on the area of the region enclosed by each sampling point of each sub-cluster.
[0077] Specifically, after determining the area of the region enclosed by each sampling point of each sub-cluster, the total amount of data to be sampled at the sampling points of each sub-cluster can be determined according to the proportion of the area of the region enclosed by each sampling point of each sub-cluster to the total area of the geographical location to be sampled.
[0078] Step S105: Conduct multiple random samplings for each sub-cluster based on the total amount of data to be sampled in each sub-cluster.
[0079] Specifically, after determining the total amount of data required at the sampling points of each sub-cluster, if the sampling points of each sub-cluster only collect data once, it may lead to inaccurate data collection due to some uncertain factors. To further improve the accuracy and distribution balance of the collected data, multiple random samplings can be conducted for the sampling points of each sub-cluster, so as to ensure that the distribution of the data collected at the sampling points of each sub-cluster is more balanced and the accuracy is higher.
[0080] Step S106: Calculate the distribution balance of each random sampling of each sub-cluster, and determine the sampling result with the most balanced distribution as the final sampling result of each sub-cluster.
[0081] Specifically, since the total amount of data to be collected for each sub-cluster has been determined, multiple random samplings are performed for each sub-cluster to improve the accuracy and distribution balance of the sampling for each sub-cluster. For the results of multiple random samplings for each sub-cluster, the distribution balance of each random sampling for each sub-cluster can be calculated, and the sampling result with the most balanced distribution is determined as the final sampling result for each sub-cluster.
[0082] As can be seen from the above technical solution, in the embodiment of the present application, the clustering result of the geographical location to be sampled can be split and refined into multiple sub-clusters, and then, according to the area of the region enclosed by each sampling point of each sub-cluster, the total amount of sampling data for each sub-cluster is determined, effectively expanding the coverage range of data sampling. Secondly, multiple random samplings are performed for each sub-cluster, the distribution balance of each random sampling for each sub-cluster is calculated, and the sampling result with the most balanced distribution is determined as the final sampling result for each sub-cluster, effectively improving the distribution balance of the sampling data for each sub-cluster and the overall distribution balance of the total amount of sampling data.
[0083] As introduced above, by initially classifying all sampling points of the geographical location to be sampled, multiple initial clusters can be obtained. The following introduces this process, which can include the following steps:
[0084] Step S11, randomly determine a sampling point in the geographical location to be sampled as the first sampling point.
[0085] Specifically, after determining the range of the geographical location to be sampled, a sampling point can be randomly determined in the geographical location to be sampled as the first sampling point, so as to perform the initial classification based on the first sampling point.
[0086] Step S12, classify the sampling points whose straight-line distance from the first sampling point is less than a preset first distance and the first sampling point into an initial cluster.
[0087] Specifically, after determining the first sampling point in the geographical location to be sampled, it can be further determined whether the straight-line distance between other sampling points in the geographical location to be sampled and the first sampling point is less than the preset first distance; if there are sampling points among other sampling points in the geographical location to be sampled whose straight-line distance from the first sampling point is less than the preset first distance, then the sampling points among other sampling points in the geographical location to be sampled whose straight-line distance from the first sampling point is less than the preset first distance and the first sampling point are classified into an initial cluster; if the straight-line distance between other sampling points in the geographical location to be sampled and the first sampling point is not less than the preset first distance, then the first sampling point is classified into an initial cluster.
[0088] Step S13: Incorporate all the sampling points that are outside the initial cluster and whose straight-line distances from each sampling point in the initial cluster are less than a preset first distance into the initial cluster until the straight-line distances between all the sampling points outside the initial cluster and all the sampling points inside the initial cluster are equal to or greater than the preset first distance.
[0089] Specifically, after determining the sampling points of the initial cluster, it can be further determined whether there are sampling points outside the initial cluster whose straight-line distances from each sampling point in the initial cluster are less than the preset first distance. If there are sampling points outside the initial cluster whose straight-line distances from each sampling point in the initial cluster are less than the preset first distance, then all the sampling points that are outside the initial cluster and whose straight-line distances from each sampling point in the initial cluster are less than the preset first distance can be incorporated into the initial cluster until the straight-line distances between all the sampling points outside the initial cluster and all the sampling points inside the initial cluster are equal to or greater than the preset first distance.
[0090] Step S14: Randomly select a sampling point from all the sampling points outside the initial cluster as the first sampling point, and then return to perform the operation of incorporating the sampling points whose straight-line distances from the first sampling point are less than the preset first distance and the first sampling point into an initial cluster until all the sampling points in the geographical location to be sampled are incorporated into each initial cluster.
[0091] Specifically, if it is determined that there are no sampling points outside the initial cluster whose straight-line distances from each sampling point in the initial cluster are less than the preset first distance, then a sampling point can be randomly selected from all the sampling points outside the initial cluster as the first sampling point, and then return to perform the operation of incorporating the sampling points whose straight-line distances from the first sampling point are less than the preset first distance and the first sampling point into an initial cluster, and re-primarily classify all the sampling points in the geographical location to be sampled that are not inside the initial cluster until all the sampling points in the geographical location to be sampled are incorporated into each initial cluster.
[0092] It can be seen from the above technical solution that the embodiments of the present application can primarily classify all the sampling points in the geographical location to be sampled to obtain multiple initial clusters, which can ensure that the sampling data is collected according to the balance of the geographical location distribution, avoiding missing the areas where the sampling points are relatively sparse; and it helps to improve the balance of the sampling data distribution in the area surrounded by each sampling point in each initial cluster.
[0093] As can be seen from the technical solutions introduced above, in the embodiments of the present application, sampling points outside the initial cluster and having a straight-line distance less than a preset first distance from each sampling point of the initial cluster can be classified into the initial cluster until the straight-line distance between each sampling point outside the initial cluster and each sampling point inside the initial cluster is equal to or greater than the preset first distance. Next, this process will be introduced, and this process may include the following steps:
[0094] Step S21: Calculate the straight-line distances between each sampling point of the initial cluster and each sampling point outside the initial cluster respectively.
[0095] Specifically, after determining each sampling point of the initial cluster, the straight-line distances between each sampling point of the initial cluster and other sampling points outside the initial cluster can be calculated respectively to determine whether there are still sampling points outside the initial cluster that can be classified into the initial cluster.
[0096] Step S22: Determine whether there is a sampling point outside the initial cluster whose straight-line distance from each sampling point inside the initial cluster is less than the preset first distance.
[0097] Specifically, after determining the straight-line distances between each sampling point outside the initial cluster and each sampling point inside the initial cluster, it can be further determined whether there is a sampling point outside the initial cluster whose straight-line distance from each sampling point inside the initial cluster is less than the preset first distance. If there is a sampling point outside the initial cluster whose straight-line distance from each sampling point inside the initial cluster is less than the preset first distance, step S23 can be executed; if there is no sampling point outside the initial cluster whose straight-line distance from each sampling point inside the initial cluster is less than the preset first distance, the preliminary classification of other sampling points outside the initial cluster can be started until the straight-line distance between each sampling point outside the initial cluster and each sampling point inside the initial cluster is equal to or greater than the preset first distance.
[0098] Step S23: Classify sampling points outside the initial cluster and having a straight-line distance less than the preset first distance from each sampling point of the initial cluster into the initial cluster until the straight-line distance between each sampling point outside the initial cluster and each sampling point inside the initial cluster is equal to or greater than the preset first distance.
[0099] Specifically, if it is determined that there are sampling points outside the initial cluster whose straight-line distances from each sampling point within the initial cluster are less than a preset first distance, then the sampling points outside the initial cluster and whose straight-line distances from each sampling point of the initial cluster are less than the preset first distance can be classified into the initial cluster until the straight-line distances between each sampling point outside the initial cluster and each sampling point within the initial cluster are all equal to or greater than the preset first distance.
[0100] As can be seen from the above technical solution, the embodiments of the present application can classify the sampling points outside the initial cluster and whose straight-line distances from each sampling point of the initial cluster are less than the preset first distance into the initial cluster until the straight-line distances between each sampling point outside the initial cluster and each sampling point within the initial cluster are all equal to or greater than the preset first distance. All sampling points of the geographical location to be sampled can be initially classified to obtain multiple initial clusters, which can ensure the balanced acquisition of sampling data according to the geographical location distribution, and avoid missing the areas with relatively sparse sampling point distributions; it helps to improve the balanced distribution of sampling data in the area surrounded by each sampling point of each initial cluster.
[0101] As can be known from the above introduction, the embodiments of the present application can split each initial cluster respectively to obtain multiple sub-clusters. Next, this process will be introduced, and this process can include the following steps:
[0102] Step S31, determine whether each initial cluster meets the preset splitting condition.
[0103] Specifically, as can be known from the above introduction, since the sampling point density of some initial clusters is relatively large, if data is directly collected in the area with a large sampling point density, it is possible to collect a larger quantity in the initial cluster with a large sampling point density. However, since the total amount of sample data to be collected is already determined, this will lead to a reduction in the sampling quantity at other geographical locations with relatively sparse sampling points or even direct omission, which is not conducive to ensuring the balanced distribution of sampling data. Therefore, it is necessary to split each initial cluster. After obtaining multiple initial clusters, it is possible to further determine whether each initial cluster meets the preset splitting condition. If each initial cluster meets the preset splitting condition, then step S32 can be executed.
[0104] Among them, if the maximum value of the straight-line distances between any two sampling points within each cluster is equal to or greater than a preset second distance, then it is considered that the cluster meets the preset splitting condition.
[0105] Among them, the preset second distance can be set to be greater than the preset first distance.
[0106] Step S32, use the initial cluster that meets the preset splitting condition as the cluster to be split.
[0107] Specifically, if the initial cluster satisfies the preset splitting condition, it means that the maximum value of the straight-line distance between any two sampling points within the initial cluster that satisfies the preset splitting condition is equal to or greater than the preset second distance. Then, the initial cluster that satisfies the preset splitting condition can be used as the cluster to be split.
[0108] Step S33: Split the cluster to be split to obtain multiple sub-clusters.
[0109] Specifically, after determining the cluster to be split, the cluster to be split can be split so that each initial cluster can be split into multiple sub-clusters.
[0110] Step S34: Determine whether the split sub-clusters satisfy the preset splitting condition.
[0111] Specifically, after splitting the initial cluster that satisfies the preset splitting condition, there may still be sub-clusters that satisfy the preset splitting condition among the multiple split sub-clusters. Therefore, it is still necessary to continue to determine whether the split sub-clusters satisfy the preset splitting condition. If the split sub-clusters satisfy the preset splitting condition, then step S35 can be executed.
[0112] Step S35: Use the sub-cluster that satisfies the preset splitting condition as the cluster to be split, and return to execute the operation of splitting the cluster to be split to obtain multiple sub-clusters until each split sub-cluster does not satisfy the preset splitting condition.
[0113] Specifically, if the split sub-clusters satisfy the preset splitting condition, then the sub-clusters that satisfy the preset splitting condition can continue to be used as the clusters to be split, and return to execute the operation of splitting the cluster to be split to obtain multiple sub-clusters until each split sub-cluster does not satisfy the preset splitting condition. This can ensure that the straight-line distance between any two sampling points in each split sub-cluster is less than the preset second distance.
[0114] As can be seen from the above technical solution, the embodiments of the present application can split each initial cluster respectively. After splitting the initial cluster to obtain multiple sub-clusters, further determine whether the split sub-clusters satisfy the preset splitting condition, so as to continue to split the sub-clusters that satisfy the preset splitting condition until each split sub-cluster does not satisfy the preset splitting condition. This can ensure that the straight-line distance between any two sampling points in each split sub-cluster is less than the preset second distance. To ensure that the area of the region surrounded by the sampling points of each cluster is controlled within a relatively reasonable range, which is beneficial to ensuring the distribution balance of the sampling data.
[0115] As can be seen from the above introduction, the embodiments of the present application can split the to-be-split cluster to obtain multiple sub-clusters. Next, this process will be introduced, and this process may include the following steps:
[0116] Step S41: Calculate the straight-line distance between any two sampling points in each to-be-split cluster respectively.
[0117] Specifically, after determining the to-be-split cluster, the straight-line distance between any two sampling points in each to-be-split cluster can be calculated respectively, so as to determine how to split the to-be-split cluster.
[0118] Step S42: Determine the two sampling points with the maximum straight-line distance between any two sampling points in each to-be-split cluster as the target sampling points.
[0119] Specifically, after determining the straight-line distance between any two sampling points in each to-be-split cluster, the two sampling points with the maximum straight-line distance between any two sampling points in each to-be-split cluster can be further determined as the target sampling points.
[0120] Step S43: Connect a straight line between the two target sampling points in each to-be-split cluster.
[0121] Specifically, after determining the straight-line distance between the two target sampling points in each to-be-split cluster, a straight line can be connected between the two target sampling points in each to-be-split cluster.
[0122] Step S44: Calculate the target ratio between the straight-line distance between the target sampling points and a preset second distance.
[0123] Specifically, after connecting a straight line between the two target sampling points in each to-be-split cluster, the target ratio between the straight-line distance between the target sampling points and a preset second distance can be further determined, so as to determine the split center point for the to-be-split cluster based on the target ratio.
[0124] Step S45: Determine multiple center points on the straight line connected by the two target sampling points according to the target ratio, so that the distance between two adjacent center points on the straight line connected by the two target sampling points is less than or equal to the preset second distance.
[0125] Specifically, after determining the target ratio, multiple center points can be determined on the straight line connected by the two target sampling points according to the target ratio, so that the distance between two adjacent center points on the straight line connected by the two target sampling points is less than or equal to the preset second distance.
[0126] Step S46: At each center point on the straight line connecting the target sampling points, draw a perpendicular line. Use each perpendicular line as a dividing line to split the clustering to be split, obtaining multiple sub-clusterings.
[0127] Specifically, after determining the split center points of the clustering to be split, at each center point on the straight line connecting the target sampling points, draw a perpendicular line. Use each perpendicular line as a dividing line to split the clustering to be split, obtaining multiple sub-clusterings.
[0128] As can be seen from the above technical solution, the embodiment of the present application can split the clustering to be split, obtaining multiple sub-clusterings, so as to ensure that the area of the region surrounded by the sampling points of each clustering is controlled within a relatively reasonable range, which is conducive to ensuring the distribution balance of the sampling data.
[0129] As can be known from the above introduction, the embodiment of the present application can determine the total amount of data to be sampled within each sub-clustering based on the area of the region surrounded by the sampling points of each sub-clustering. Next, this process will be introduced, and this process can include the following steps:
[0130] Step S51: Calculate the straight-line distance between any two sampling points within each sub-clustering.
[0131] Specifically, after splitting each clustering, each split sub-clustering can be used as a smallest sampling unit, and the total amount of data to be sampled for each sub-clustering can be further determined. To better determine the total amount of data to be sampled for each sub-clustering, among them, to determine the total amount of data to be sampled for each sub-clustering, it can be determined by the area of the region surrounded by the sampling points of each sub-clustering. Therefore, the straight-line distance between any two sampling points within each sub-clustering can be calculated, so that it can be used to calculate the area of the region surrounded by the sampling points of each sub-clustering.
[0132] Step S52: Calculate the area of a circle with the maximum value of the straight-line distance between any two sampling points within each sub-clustering as the diameter, and use it as the area of the region surrounded by the sampling points of each sub-clustering.
[0133] Specifically, after determining the straight-line distance between any two sampling points within each sub-clustering, to facilitate calculating the area of the region surrounded by the sampling points of each sub-clustering, the region surrounded by the sampling points of each sub-clustering can be regarded as an approximate circle. Therefore, the area of a circle with the maximum value of the straight-line distance between any two sampling points within each sub-clustering as the diameter can be calculated, and used as the area of the region surrounded by the sampling points of each sub-clustering.
[0134] Step S53: Accumulate the areas of the regions enclosed by the sampling points of each sub-cluster to obtain the total area of the geographical location to be sampled.
[0135] Specifically, after determining the areas of the regions enclosed by the sampling points of each sub-cluster, the areas of the regions enclosed by the sampling points of each sub-cluster can be accumulated to obtain the total area of the geographical location to be sampled. So as to determine the proportion of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled.
[0136] Step S54: Calculate the ratio of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled.
[0137] Specifically, after determining the total area of the geographical location to be sampled, the ratio of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled can be calculated. So as to determine the total amount of sampled data within each sub-cluster according to the ratio of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled and the preset total amount of sampled data.
[0138] Step S55: Based on the ratio of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled and the preset total amount of sampled data, determine the total amount of sampled data within each sub-cluster.
[0139] Specifically, after determining the ratio of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled, the total amount of sampled data within each sub-cluster can be determined according to the preset total amount of sampled data.
[0140] Among them, the total amount of sampled data within each sub-cluster is equal to the preset total amount of sampled data multiplied by the ratio of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled.
[0141] From the above technical solutions, it can be seen that the embodiments of the present application can determine the total amount of data to be sampled within each sub-cluster based on the area of the region enclosed by the sampling points of each sub-cluster, which can ensure that the sampled data is collected according to the balance of the geographical location distribution, avoiding missing the regions where the sampling points are sparsely distributed; and helps to improve the balance of the sampled data distribution in the regions enclosed by the sampling points of each initial cluster.
[0142] As can be seen from the above introduction, the embodiments of the present application can calculate the distribution balance of each random sampling of each sub-cluster, and determine the sampling result with the most balanced distribution as the final sampling result of each sub-cluster. Next, this process will be introduced, and this process may include the following steps:
[0143] Step S61: Calculate the straight-line distance between any two sampling points among all the sampling points of each random sampling of each sub-cluster.
[0144] Specifically, after determining the total amount of data to be sampled for each sub-cluster, in order to better improve the accuracy of data sampling and the distribution balance of data, and to determine the sampling data with the best distribution balance among the sampling data of each sub-cluster. The distribution balance of each random sampling result of each sub-cluster can be calculated. Therefore, the straight-line distance between any two sampling points among all the sampling points of each random sampling of each sub-cluster can be calculated, so as to be used to calculate the distribution balance of each sampling result.
[0145] Step S62: Calculate the sum of the straight-line distances between any two sampling points among all the sampling points of each random sampling of each sub-cluster.
[0146] Specifically, after determining the straight-line distance between any two sampling points among all the sampling points of each random sampling of each sub-cluster, the sum of the straight-line distances between any two sampling points among all the sampling points of each random sampling of each sub-cluster can be calculated. So as to determine the distribution balance of each random sampling result.
[0147] Step S63: Take the random sampling result with the maximum sum of the straight-line distances between any two sampling points among all the sampling points of each random sampling as the final sampling result of each sub-cluster.
[0148] Specifically, the larger the sum of the straight-line distances between any two sampling points among all the sampling points of each random sampling of each sub-cluster, it indicates that under the condition of collecting the same amount of data, the larger the sum of the straight-line distances between any two sampling points among all the sampling points of the random sampling, it indicates that the distribution range of the randomly sampled data is wider and the data distribution balance is better. Therefore, after determining the sum of the straight-line distances between any two sampling points among all the sampling points of each random sampling of each sub-cluster, the random sampling result with the maximum sum of the straight-line distances between any two sampling points among all the sampling points of each random sampling can be taken as the final sampling result of each sub-cluster.
[0149] As can be seen from the above technical solution, the embodiment of the present application can calculate the distribution balance of each random sampling of each sub-cluster, and determine the sampling result with the most balanced distribution as the final sampling result of each sub-cluster. It can ensure that the data of each random sampling of each cluster is sampled according to the distribution balance of the geographical location, which helps to improve the distribution balance of the sampling data in the area surrounded by the sampling points of each initial cluster.
[0150] In the technical field of autonomous driving, there are many autonomous driving technology algorithms, and to implement these autonomous driving technology algorithms, sometimes neural network models are needed. To better implement some of the autonomous driving technology algorithms involved in autonomous driving, a large number of geographical location data samples in a wide range need to be collected as training data to train the deep neural network models involved in some of the autonomous driving technology algorithms.
[0151] As can be seen from the data sampling scheme introduced above, using the data sampling method introduced above to collect data can improve the distribution balance of the collected data. In practical applications, to improve the accuracy of the deep neural network model involved in training the autonomous driving technology algorithm, the data samples used as training data need to be distributed over a wider geographical location range. Therefore, the data sampling method introduced above can be used to collect the training data required for training the deep neural network model involved in the autonomous driving technology algorithm.
[0152] For example, the data sampling method introduced above can be used to collect the training data required for training the unmanned vehicle control model. The trained unmanned vehicle control model can be used to handle one or more of the tasks such as perceiving the surrounding driving environment state of the unmanned vehicle, predicting the vehicle driving trajectory, or making decisions on vehicle lane changes.
[0153] For example, when a prediction model that can determine whether a vehicle needs to change lanes based on the states of surrounding vehicles, pedestrians, and other obstacles needs to be trained. Since the vehicle environment scenarios are different in different geographical locations. For example, in some places, the traffic flow may be relatively large, in some places, there may be large trucks, and in some places, the pedestrian flow may be relatively large. Therefore, if the collected training data can cover a larger geographical location range, the training data can have more abundant scenario data. The richer the scenarios of the training data, the better the training effect of the model can be improved.
[0154] As can be seen from the above introduction, the embodiments of the present application can use the data collected by any of the data sampling methods introduced above as training data to train the deep neural network model involved in the autonomous driving technology algorithm, so as to make the training effect of the model better, so that it can better handle one or more of the perception, prediction, and decision-making tasks of the unmanned vehicle.
[0155] Next, the data sampling device and model training device provided by the embodiments of the present application will be described. The data sampling device and model training device described below can be correspondingly referred to the data sampling method and model training method described above.
[0156] See Figure 2 , Figure 2 which is a schematic structural diagram of a data sampling device disclosed in an embodiment of the present application.
[0157] As shown in Figure 2 , the data sampling device may include:
[0158] A clustering unit 101, configured to preliminarily classify all sampling points of the geographical location to be sampled, and obtain a plurality of initial clusters;
[0159] A splitting unit 102, configured to split each initial cluster respectively to obtain a plurality of sub-clusters;
[0160] An area calculation unit 103, configured to calculate the area of the region enclosed by each sampling point of each sub-cluster;
[0161] A data determination unit 104, configured to determine the total amount of data to be sampled in each sub-cluster based on the area of the region enclosed by each sampling point of each sub-cluster;
[0162] A sampling unit 105, configured to perform multiple random samplings on each sub-cluster based on the total amount of data to be sampled in each sub-cluster;
[0163] A balance calculation unit 106, configured to calculate the distribution balance of each random sampling of each sub-cluster, and determine the most balanced sampling result as the final sampling result of each sub-cluster.
[0164] As can be seen from the above introduction, the device according to the embodiment of the present application can use the clustering unit 101 to initially classify each sampling point of the geographical location to be sampled. After using the splitting unit 102 to split and refine the clustering result into multiple sub-clusters, the area calculation unit 103 is used to calculate the area of the region enclosed by each sampling point of each sub-cluster, so as to use the data determination unit 104 to determine the total amount of sampling data for each sub-cluster, effectively expanding the coverage range of data sampling. Secondly, after determining the total amount of data sampled within each sub-cluster, the sampling unit 105 can be used to perform multiple random samplings for each sub-cluster, and the balance calculation unit 106 can be used to calculate the distribution balance of each random sampling of each sub-cluster, and determine the most balanced sampling result as the final sampling result of each sub-cluster. This effectively improves the distribution balance of the sampling data of each sub-cluster and improves the overall distribution balance of the total amount of sampling data.
[0165] Further optionally, the above-mentioned clustering unit 101 may include:
[0166] A first sampling point determination unit, configured to randomly determine a sampling point as the first sampling point in the geographical location to be sampled;
[0167] A classification unit, which is used to classify the sampling points whose straight-line distance from the first sampling point is less than a preset first distance and the first sampling point into an initial cluster; and classify the sampling points outside the initial cluster and whose straight-line distance from each sampling point in the initial cluster is less than the preset first distance into the initial cluster until the straight-line distance between each sampling point outside the initial cluster and each sampling point inside the initial cluster is equal to or greater than the preset first distance;
[0168] A second sampling point determination unit, which is used to randomly select a sampling point from each sampling point outside the initial cluster as the first sampling point, and return to execute the operation of classifying the sampling points whose straight-line distance from the first sampling point is less than the preset first distance and the first sampling point into an initial cluster until all sampling points in the geographical location to be sampled are classified into each initial cluster.
[0169] Further optionally, the above classification unit may include:
[0170] A first calculation unit, which is used to calculate the straight-line distance between each sampling point in the initial cluster and each sampling point outside the initial cluster respectively;
[0171] A first judgment unit, which is used to judge whether there is a sampling point outside the initial cluster whose straight-line distance from each sampling point in the initial cluster is less than the preset first distance;
[0172] A classification sub-unit, which is used to, when the judgment result of the first judgment unit is yes, classify the sampling points outside the initial cluster and whose straight-line distance from each sampling point in the initial cluster is less than the preset first distance into the initial cluster until the straight-line distance between each sampling point outside the initial cluster and each sampling point inside the initial cluster is equal to or greater than the preset first distance.
[0173] Further optionally, the above splitting unit 102 may include:
[0174] A second judgment unit, which is used to judge whether each initial cluster meets the preset splitting condition;
[0175] A first splitting sub-unit, which is used to, when the execution result of the second judgment unit is yes, use the initial cluster that meets the preset splitting condition as the cluster to be split;
[0176] A second splitting sub-unit, which is used to split the cluster to be split to obtain a plurality of sub-clusters;
[0177] A third judgment unit, which is used to judge whether the split sub-clusters meet the preset splitting condition;
[0178] A third splitting subunit, configured to, when the execution result of the third determination unit is yes, use a sub-cluster that meets a preset splitting condition as a cluster to be split, and return to perform an operation of splitting the cluster to be split to obtain multiple sub-clusters until each sub-cluster after splitting does not meet the preset splitting condition.
[0179] Further optionally, the second splitting subunit may include:
[0180] A second calculation unit, configured to calculate the straight-line distance between any two sampling points in each cluster to be split;
[0181] A target sampling point determination unit, configured to determine two sampling points with the maximum straight-line distance between any two sampling points in each cluster to be split as target sampling points;
[0182] A straight-line connection unit, configured to connect a straight line between the two target sampling points in each cluster to be split;
[0183] A third calculation unit, configured to calculate a target ratio between the straight-line distance between the target sampling points and a preset second distance;
[0184] A center point determination unit, configured to determine multiple center points on the straight line connected by the two target sampling points according to the target ratio, so that the distance between two adjacent center points on the straight line connected by the two target sampling points is less than or equal to the preset second distance;
[0185] A fourth splitting subunit, configured to draw a perpendicular line at each center point on the straight line connected by the target sampling points, and use each perpendicular line as a dividing line to split the cluster to be split to obtain multiple sub-clusters.
[0186] Further optionally, the above-mentioned balance calculation unit 106 may include:
[0187] A fourth calculation unit, configured to calculate the straight-line distance between any two sampling points among the sampling points of each random sampling of each sub-cluster;
[0188] A fifth calculation unit, configured to calculate the sum of the straight-line distances between any two sampling points among the sampling points of each random sampling of each sub-cluster;
[0189] A sampling result determination unit, configured to use the result of a random sampling with the maximum sum of the straight-line distances between any two sampling points among the sampling points of each random sampling as the final sampling result of each sub-cluster.
[0190] Wherein, the specific processing flow of each unit included in the above-mentioned data sampling device may refer to the relevant introduction in the previous part of the data sampling method, and will not be elaborated here.
[0191] A model training device may include:
[0192] A data acquisition unit, configured to use the data collected by any of the data sampling methods described above as training data;
[0193] A training unit, configured to train a deep neural network model using the training data obtained by the data acquisition unit, and obtain a trained unmanned vehicle control model, where the unmanned vehicle control model is used to process one or more of the perception, prediction, and decision-making tasks of an unmanned vehicle.
[0194] The data sampling device provided in the embodiments of the present application can be applied to data sampling devices, and the model training device can be applied to model training devices, such as terminals: mobile phones, computers, etc. Optionally, Figure 3 The hardware structure block diagram of the data sampling device or the model training device is shown. Referring to Figure 3 , the hardware structure of the data sampling device or the model training device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4.
[0195] In the embodiments of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4.
[0196] The processor 1 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application, etc.;
[0197] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;
[0198] Wherein, the memory stores a program, and the processor can call the program stored in the memory, and the program is used to: implement each processing flow in the foregoing terminal data sampling scheme and model training scheme.
[0199] The embodiments of the present application also provide a readable storage medium, which can store a program suitable for being executed by a processor, and the program is used to: implement each processing flow in the foregoing terminal data sampling scheme and model training scheme.
[0200] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0201] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the various embodiments, reference may be made to each other.
[0202] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. The various embodiments can be combined with each other. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model training method, characterized in that, Including: Using the data collected by the data sampling method as training data; Training a deep neural network model using the training data to obtain a trained unmanned vehicle control model, where the unmanned vehicle control model is used to handle one or more of the perception, prediction, and decision-making tasks of an unmanned vehicle; The data sampling method includes: Preliminarily classifying all sampling points at the geographical location to be sampled to obtain multiple initial clusters; Splitting each initial cluster respectively to obtain multiple sub-clusters; Calculating the area of the region enclosed by the sampling points of each sub-cluster; Based on the area of the region enclosed by the sampling points of each sub-cluster, determining the total amount of data to be sampled within each sub-cluster; Based on the total amount of data to be sampled for each sub-cluster, performing multiple random samplings on each sub-cluster; Calculating the distribution balance of each random sampling of each sub-cluster, and determining the most balanced sampling result as the final sampling result of each sub-cluster; including: calculating the straight-line distance between any two sampling points among the sampling points of each random sampling of each sub-cluster; calculating the sum of the straight-line distances between any two sampling points among the sampling points of each random sampling of each sub-cluster; taking the random sampling result with the maximum sum of the straight-line distances between any two sampling points among the sampling points of each random sampling as the final sampling result of each sub-cluster.
2. The model training method according to claim 1, characterized in that, The preliminary classification of all sampling points at the geographical location to be sampled to obtain multiple initial clusters includes: Randomly determining a sampling point in the geographical location to be sampled as the first sampling point; Grouping the sampling points whose straight-line distance from the first sampling point is less than a preset first distance and the first sampling point into one initial cluster; Grouping the sampling points outside the initial cluster whose straight-line distance from the sampling points of the initial cluster is less than the preset first distance into the initial cluster until the straight-line distance between the sampling points outside the initial cluster and the sampling points inside the initial cluster is equal to or greater than the preset first distance; Randomly selecting a sampling point from the sampling points outside the initial cluster as the first sampling point, and returning to perform the operation of grouping the sampling points whose straight-line distance from the first sampling point is less than the preset first distance and the first sampling point into one initial cluster until all sampling points in the geographical location to be sampled are grouped into each initial cluster.
3. The model training method according to claim 2, characterized in that, The grouping of the sampling points outside the initial cluster whose straight-line distance from the sampling points of the initial cluster is less than the preset first distance into the initial cluster until the straight-line distance between the sampling points outside the initial cluster and the sampling points inside the initial cluster is equal to or greater than the preset first distance includes: Calculating the straight-line distance between each sampling point of the initial cluster and the sampling points outside the initial cluster respectively; Judging whether there are sampling points outside the initial cluster whose straight-line distance from the sampling points inside the initial cluster is less than the preset first distance; If there are sampling points outside the initial cluster whose straight-line distances from each sampling point inside the initial cluster are less than a preset first distance, then include all the sampling points outside the initial cluster and with straight-line distances less than the preset first distance from each sampling point of the initial cluster into the initial cluster until the straight-line distances between each sampling point outside the initial cluster and each sampling point inside the initial cluster are all equal to or greater than the preset first distance.
4. The model training method according to claim 1, characterized in that, The splitting of each initial cluster into multiple sub-clusters includes: Judging whether each initial cluster meets the preset splitting condition; If an initial cluster meets the preset splitting condition, then use the initial cluster that meets the preset splitting condition as the cluster to be split; Split the cluster to be split to obtain multiple sub-clusters; Judging whether the split sub-clusters meet the preset splitting condition; If the split sub-clusters meet the preset splitting condition, then use the sub-clusters that meet the preset splitting condition as the clusters to be split, and return to perform the operation of splitting the clusters to be split to obtain multiple sub-clusters until each split sub-cluster does not meet the preset splitting condition.
5. The model training method according to claim 4, characterized in that, The preset splitting condition includes: If the maximum value of the straight-line distances between any two sampling points in each cluster is equal to or greater than the preset second distance, then it is considered that the cluster meets the preset splitting condition.
6. The model training method according to claim 5, characterized in that, The splitting of the cluster to be split to obtain multiple sub-clusters includes: Calculate the straight-line distances between any two sampling points in each cluster to be split respectively; Determine the two sampling points with the maximum straight-line distance between any two sampling points in each cluster to be split as the target sampling points; Connect a straight line between the two target sampling points in each cluster to be split; Calculate the target ratio between the straight-line distance between the target sampling points and the preset second distance; According to the target ratio, determine multiple center points on the straight line connected by the two target sampling points, so that the distance between two adjacent center points on the straight line connected by the two target sampling points is less than or equal to the preset second distance; Draw a perpendicular line at each center point on the straight line connected by the target sampling points, and use each perpendicular line as a dividing line to split the cluster to be split to obtain multiple sub-clusters.
7. The model training method according to any one of claims 1-6, characterized in that Based on the area of the region enclosed by the sampling points of each sub-cluster, determining the total amount of data to be sampled in each sub-cluster includes: Calculate the straight-line distances between any two sampling points in each sub-cluster; Calculate the area of the circle with the maximum value of the straight-line distances between any two sampling points in each sub-cluster as the diameter, and use it as the area of the region enclosed by the sampling points of each sub-cluster; Accumulate the areas of the regions enclosed by the sampling points of each sub-cluster to obtain the total area of the geographical location to be sampled; Calculate the ratio of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled; Based on the ratio of the area of the region enclosed by the sampling points of each sub-cluster to the total area of the geographical location to be sampled, and the preset total amount of sampling data, determine the total amount of data to be sampled in each sub-cluster.
8. A model training device, characterized in that including: A data acquisition unit for using the data collected by a preset data sampling method as training data; wherein, the preset data sampling method includes: preliminarily classifying all sampling points of the geographical location to be sampled to obtain a plurality of initial clusters; splitting each of the initial clusters to obtain a plurality of sub-clusters; calculating the area of the region enclosed by the sampling points of each sub-cluster; determining the total amount of data to be sampled in each sub-cluster based on the area of the region enclosed by the sampling points of each sub-cluster; performing multiple random samplings on each sub-cluster based on the total amount of data to be sampled in each sub-cluster; calculating the distribution balance of each random sampling of each sub-cluster, and determining the most balanced sampling result as the final sampling result of each sub-cluster; including: calculating the straight-line distance between any two sampling points among the sampling points of each random sampling of each sub-cluster; calculating the sum of the straight-line distances between any two sampling points among the sampling points of each random sampling of each sub-cluster; taking the random sampling result with the maximum sum of the straight-line distances between any two sampling points among the sampling points of each random sampling as the final sampling result of each sub-cluster. A training unit for training a deep neural network model using the training data obtained by the data acquisition unit to obtain a trained unmanned vehicle control model, where the unmanned vehicle control model is used to process one or more of the perception, prediction, and decision-making tasks of an unmanned vehicle.
9. A data sampling device, characterized in that Including: One or more processors, and a memory; Computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the one or more processors, the steps of the model training method according to any one of claims 1 to 7 are implemented.
10. A storage medium, characterized in that: Computer-readable instructions are stored in the storage medium, and when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to implement the steps of the model training method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Sampling point distribution method and device based on landscape heterogeneity and electronic equipment
CN113344105A