Core sample data set selection method for automatic driving control strategy learning
By building a reference system based on the key features of autonomous driving control and performing grid sampling, the problem of selecting core data sets for learning autonomous driving control strategy in the existing technology is solved, learning efficiency and model performance are improved, and more efficient data utilization and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510121070.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively select core data sets for autonomous driving control strategy learning, resulting in low efficiency and poor performance in model training, especially when faced with high-dimensional multimodal data and complex interactive scenarios.
By building a reference system based on key features of autonomous driving control and rasterizing it, data samples are projected into the grid, and using a grid-based sampling method for equalization, a core sample data set for self-driving control strategy learning is constructed.
It improves the learning efficiency and model performance of autonomous driving control strategies, suppresses the long-tail problem in the data center, and builds a data subset with higher distribution and higher information density, reducing data storage needs and improving the model's generalization ability on key data.
Smart Images

Figure CN120067681A_ABST
Abstract
Claims
1. A method for selecting a core sample data set for autonomous driving control strategy learning, characterized in that: include: A reference system is constructed using key features of autonomous driving control, and the reference system is gridded to obtain a plurality of grids; Project the data samples of various autonomous driving scenarios into the grid. The data samples in the same grid have common features. A grid-based sampling method is used to evenly sample the data samples in each grid to obtain the core sample data set for learning autonomous driving control strategies.
2. The method according to claim 1, characterized in that The reference system is constructed by utilizing the key features of the autonomous driving control, and the reference system is gridded to obtain a plurality of grids, including: Key features related to driving control are selected using driving experience knowledge, a reference system is constructed using the key features, and the reference system is rasterized to obtain a plurality of grids.
3. The method according to claim 2, characterized in that The reference system is constructed by utilizing the key features of the autonomous driving control, and the reference system is gridded to obtain a plurality of grids, including: A specific driving behavior completed by a vehicle is defined as a driving task. The driving data is divided according to the driving tasks. The core data sets are constructed according to the driving tasks. The starting point of the driving task is taken as the origin, the radial distance s along the reference trajectory is taken as the horizontal axis, and the lateral deviation distance d perpendicular to the reference trajectory is taken as the vertical axis. The radial distance s and the deviation distance d are taken as the key features of the automatic driving control. The reference system is constructed using the key features. For reference frame The coordinates are normalized to obtain the coordinates where s max and d max The data samples representing the driving task are in the reference frame The maximum value of the mid-coordinate.
4. The method according to claim 3, characterized in that The data samples of various autonomous driving scenarios are projected into the grid. The data samples in the same grid have common characteristics, including: According to the data sample z i The vehicle position (x i ,y i ), calculate the data sample z i In the normalized reference frame The coordinates of the reference system Divide into W×H grid regions according to fixed intervals, and obtain a series of grids E={E w,h }, project the data samples of various autonomous driving scenarios into each grid, each grid contains a sub-dataset Subdataset E w,h The data samples in the i In the normalized reference frame The coordinates of the data sample z i Mapped to the corresponding grid E w,h Inside.
5. The method according to claim 4, characterized in that The grid-based sampling method is used to perform balanced sampling on the data samples in each grid to obtain a core sample data set for learning the autonomous driving control strategy, including: After projecting the data samples of various autonomous driving scenarios into each grid, the data samples in the grid with redundant data are downsampled, and the data samples in the grid with small data volume are oversampled. The specific processing process includes: Given the gridded data set E = {E w,h }, the size of the core sample data set S is N S ; Request: core sample data set S 1) set S1=φ,S2=φ,S remain =φ 2) Let N occ is the number of non-empty grids, is the average number of samples per grid; 3) The first step of sampling: a) For each grid E w,h , set the number of samples b) From each grid E w,h Randomly select n w,h The sample is placed in S1, and the remaining samples are placed in S remain 4) The second step is sampling: a) For S remain Each sample Setting weights b) From S remain Medium weighted sampling N S - / S1 / sample, put into S2 5) Core sample data set S = S1 ∪ S2.