Core sample data set selection method for automatic driving control strategy learning

By building a reference system based on the key features of autonomous driving control and performing grid sampling, the problem of selecting core data sets for learning autonomous driving control strategy in the existing technology is solved, learning efficiency and model performance are improved, and more efficient data utilization and generalization capabilities are achieved.

CN120067681APending Publication Date: 2025-05-30PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510121070.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively select core data sets for autonomous driving control strategy learning, resulting in low efficiency and poor performance in model training, especially when faced with high-dimensional multimodal data and complex interactive scenarios.

Method used

By building a reference system based on key features of autonomous driving control and rasterizing it, data samples are projected into the grid, and using a grid-based sampling method for equalization, a core sample data set for self-driving control strategy learning is constructed.

Benefits of technology

It improves the learning efficiency and model performance of autonomous driving control strategies, suppresses the long-tail problem in the data center, and builds a data subset with higher distribution and higher information density, reducing data storage needs and improving the model's generalization ability on key data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067681A_ABST
    Figure CN120067681A_ABST
Patent Text Reader

Abstract

The invention provides a core sample data set selection method for automatic driving control strategy learning. The method comprises the following steps: constructing a reference system by using key features of automatic driving control, and rasterizing the reference system to obtain a plurality of grids; projecting data samples of various automatic driving scenes into grids, wherein the data samples in the same grid have common characteristics; and carrying out balanced sampling on the data samples in each grid by adopting a grid-based sampling method to obtain a core sample data set for automatic driving control strategy learning. The method can improve the learning efficiency and model performance of the automatic driving control strategy. According to the method, the long tail problem generally existing in an automatic driving data set can be inhibited, and data subsets with higher distribution and higher information density are constructed from mass data, so that the data storage is reduced, the training efficiency and the generalization ability of the model on key data are improved, and the method has very high application value.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for selecting a core sample data set for autonomous driving control strategy learning, characterized in that: include: A reference system is constructed using key features of autonomous driving control, and the reference system is gridded to obtain a plurality of grids; Project the data samples of various autonomous driving scenarios into the grid. The data samples in the same grid have common features. A grid-based sampling method is used to evenly sample the data samples in each grid to obtain the core sample data set for learning autonomous driving control strategies.

2. The method according to claim 1, characterized in that The reference system is constructed by utilizing the key features of the autonomous driving control, and the reference system is gridded to obtain a plurality of grids, including: Key features related to driving control are selected using driving experience knowledge, a reference system is constructed using the key features, and the reference system is rasterized to obtain a plurality of grids.

3. The method according to claim 2, characterized in that The reference system is constructed by utilizing the key features of the autonomous driving control, and the reference system is gridded to obtain a plurality of grids, including: A specific driving behavior completed by a vehicle is defined as a driving task. The driving data is divided according to the driving tasks. The core data sets are constructed according to the driving tasks. The starting point of the driving task is taken as the origin, the radial distance s along the reference trajectory is taken as the horizontal axis, and the lateral deviation distance d perpendicular to the reference trajectory is taken as the vertical axis. The radial distance s and the deviation distance d are taken as the key features of the automatic driving control. The reference system is constructed using the key features. For reference frame The coordinates are normalized to obtain the coordinates where s max and d max The data samples representing the driving task are in the reference frame The maximum value of the mid-coordinate.

4. The method according to claim 3, characterized in that The data samples of various autonomous driving scenarios are projected into the grid. The data samples in the same grid have common characteristics, including: According to the data sample z i The vehicle position (x i ,y i ), calculate the data sample z i In the normalized reference frame The coordinates of the reference system Divide into W×H grid regions according to fixed intervals, and obtain a series of grids E={E w,h }, project the data samples of various autonomous driving scenarios into each grid, each grid contains a sub-dataset Subdataset E w,h The data samples in the i In the normalized reference frame The coordinates of the data sample z i Mapped to the corresponding grid E w,h Inside.

5. The method according to claim 4, characterized in that The grid-based sampling method is used to perform balanced sampling on the data samples in each grid to obtain a core sample data set for learning the autonomous driving control strategy, including: After projecting the data samples of various autonomous driving scenarios into each grid, the data samples in the grid with redundant data are downsampled, and the data samples in the grid with small data volume are oversampled. The specific processing process includes: Given the gridded data set E = {E w,h }, the size of the core sample data set S is N S ; Request: core sample data set S 1) set S1=φ,S2=φ,S remain =φ 2) Let N occ is the number of non-empty grids, is the average number of samples per grid; 3) The first step of sampling: a) For each grid E w,h , set the number of samples b) From each grid E w,h Randomly select n w,h The sample is placed in S1, and the remaining samples are placed in S remain 4) The second step is sampling: a) For S remain Each sample Setting weights b) From S remain Medium weighted sampling N S - / S1 / sample, put into S2 5) Core sample data set S = S1 ∪ S2.