A Feature Selection Method for Autonomous Driving Simulation Testing

By building a robust discriminant feature space and sparse constraint mechanism, combined with unsupervised algorithm, the dimensional disaster problem of high-dimensional scene feature space in virtual simulation test of autonomous driving systems is solved, and efficient feature selection is achieved without prior labeling data, which improves the efficiency and accuracy of autonomous driving tests.

CN120030315BActive Publication Date: 2025-07-01SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510497546.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-01
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The virtual simulation testing technology of existing autonomous driving systems faces high computing resource consumption and scenario generation redundancy problems in complex dynamic environments. The traditional exhaustive testing is inefficient, deep learning lacks interpretability, and the existing feature selection methods fail to effectively solve the dimensional disaster problem of high-dimensional scene feature space.

Method used

By building a robust discriminant feature space, combining sparse constraints and polarized discrete scoring mechanisms, unsupervised algorithms quickly identify and select features of key test scenarios without prior labeling of data, and using technologies such as K-means clustering, sparse weight matrix and augmented Lagrangian method for feature selection.

Benefits of technology

It realizes efficient identification and selection of features most relevant to key test scenarios without prior labeling data, improves the efficiency and accuracy of autonomous driving tests and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030315B_ABST
    Figure CN120030315B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for feature selection in autonomous driving simulation testing, which relates to the technical field of autonomous driving system testing, and includes the steps: T1. Collect an unlabeled test set of the autonomous driving system, denoted as the first test set; preprocess the first test set to generate a second test set; construct a projection matrix that satisfies orthogonal constraints; construct a sparse weight matrix; T2. Obtain the initial pseudo-labels of the second test set, and divide the driving scene features into safe scenes or critical scenes; T3. Use the projection matrix to project the second test set into the discriminant feature space; T4. Apply sparse weights to the projection matrix; construct an objective optimization function for the discriminant feature space; solve the objective optimization function to obtain a discrete feature selection matrix; T5. Generate a final scene feature subset according to the non-zero row indices of the feature selection matrix. By constructing a robust discriminant feature space and combining sparse constraints with a polarization discrete scoring mechanism, the efficient identification and division of critical scene features are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving system testing, and in particular to a method for feature selection in autonomous driving simulation testing. Background Art

[0002] In the process of large-scale application of autonomous driving systems (ADS), the verification of safety and reliability in complex dynamic environments has become a key bottleneck - existing virtual simulation testing technologies need to cover a large number of extreme scenarios to evaluate the system behavior boundary, but their scenario generation faces multi-dimensional dynamic coupling challenges: on the one hand, the co-evolution of the kinematic characteristics of ADS, road topology, multi-agent interaction rules, and environmental parameters (such as lighting, meteorology) significantly increases the complexity of the scenario space; on the other hand, traditional exhaustive testing causes exponential consumption of computing resources due to the need to traverse all potential feature combinations, resulting in the test efficiency and cost being difficult to meet the actual engineering requirements. There is an urgent need to construct a scenario generation optimization method with high coverage and low redundancy to break through the verification efficiency boundary.

[0003] Current low-cost autonomous driving test research focuses on test case optimization technologies, including test suite minimization, test case selection, and priority ranking. This technology mainly reduces the scale of the test set through heuristic strategies, but fails to solve the curse of dimensionality problem brought by the high-dimensional scenario feature space. Although deep learning shows potential advantages in the field of feature dimensionality reduction, its inherent black-box characteristics lead to the lack of feature interpretability, and the training paradigm that relies on a large amount of labeled data is resource-intensive. In contrast to deep learning, traditional feature selection methods have the advantage of physical interpretability and are more suitable for the scenario of autonomous driving test tasks. However, existing research has not established a feature selection framework for the specific characteristics of autonomous driving scenarios. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the present invention provides a method for feature selection in autonomous driving simulation testing, which can efficiently identify and select the features most relevant to key test scenarios through rapid iterative optimization of the algorithm without prior labeled data.

[0005] The technical solution of the present invention is realized as follows:

[0006] A method for feature selection in autonomous driving simulation testing includes the following steps:

[0007] T1. Collect an unlabeled test set of the autonomous driving system, denoted as the first test set; the first test set contains d driving scenario features; each driving scenario feature corresponds to n test cases;

[0008] Preprocess the first test set to generate a second test set;

[0009] Construct a projection matrix that satisfies the orthogonal constraint; construct a sparse weight matrix;

[0010] T2. Obtain the initial pseudo-labels of the second test set, and divide the driving scenario features into safe scenarios or critical scenarios;

[0011] T3. Use the projection matrix to project the second test set into an r-dimensional discriminant feature space, ;

[0012] T4. Use the sparse weight matrix to impose sparse weights on the projection matrix; construct the objective optimization function of the discriminant feature space; solve the objective optimization function to obtain a discrete feature selection matrix; the feature selection matrix is the optimal matrix for scene definition feature dimensionality reduction.

[0013] T5. Generate the final scene feature subset according to the non-zero row indices of the feature selection matrix. The final scene feature subset is used for subsequent test case prioritization or simulation tests.

[0014] As a further optimization of the above solution, the preprocessing is to perform centering processing on the first test set to obtain the second test set, that is:

[0015] ;

[0016] where, represents the first test set, , R is the set of real numbers, represents the set of d×n-dimensional real matrices; represents the second test set; represents an n×1 all-ones column vector, represents an n×n matrix, and each element of the matrix is .

[0017] By matrix centering processing, the mean value of the second test set can be adjusted to zero, which is convenient for subsequent statistical analysis.

[0018] As a further optimization of the above solution, in T2, use a clustering algorithm to obtain the initial pseudo-labels of the second test set; the pseudo-labels are denoted as , and ; where, 0 corresponds to the safe scenario, and 1 corresponds to the critical scenario.

[0019] Furthermore, the clustering algorithm uses the K-means clustering algorithm.

[0020] As a further optimization of the above solution, in T3, use the objective function to construct the discriminant feature space, and the objective function is expressed as:

[0021] ;

[0022] wherein, represents the Frobenius norm constraint; represents the projection matrix, and , R is the set of real numbers, represents the set of d×n dimensional real matrices; is the clustering centroid matrix of; represents the identity matrix, represents the orthogonal constraint imposed on .

[0023] The Frobenius norm is a kind of matrix norm used to measure the overall size of a matrix. It is defined as the square root of the sum of the squares of the absolute values of all elements of the matrix.

[0024] To increase the discriminability of the feature space so that the low-dimensional feature space can effectively distinguish key scenarios, a constrained trace ratio objective is used to minimize the distance between projection points for test cases belonging to the same class and maximize the distance between test cases belonging to different classes, thereby learning a discriminative feature space. The constrained trace ratio objective is a special form of matrix expression of linear discriminant analysis.

[0025] As a further optimization of the above solution, the sparse weight matrix includes two, denoted as and , and ;

[0026] The i-th diagonal element of is denoted as : ;

[0027] The i-th diagonal element of is denoted as : ;

[0028] wherein, is the i-th column element of, is the i-th column element of; is a preset constant; is a sufficiently small constant used to avoid overfitting;

[0029] After applying the sparse weight, the objective function is expressed as:

[0030] ;

[0031] Among them, denotes the clustering centroid matrix of

[0032] Sparsity is imposed on the constrained trace ratio objective through a sparse weight matrix to enhance the robustness of the constructed discriminative feature space to abnormal test cases.

[0033] As a further optimization of the above scheme, an norm constraint is added to the objective function to obtain the objective optimization function, denoted as:

[0034] ;

[0035] Among them, denotes the norm constraint.

[0036] The norm constraint is a hybrid norm constraint method that combines structured sparsity and feature selection ability. Its purpose is to make the norm of each row as small as possible, with as many 0 elements as possible in the row.

[0037] Imposing an norm constraint on the projection matrix to impose feature sparsity is to learn discrete feature scores, polarize the scores of each scenario feature, so as to form a clearer and more well-defined decision boundary, making the objective optimization function applicable to the feature selection task.

[0038] As a further optimization of the above scheme, in T4, the augmented Lagrangian method and the alternating direction method of multipliers are used to solve the objective optimization function to obtain the feature selection matrix.

[0039] The augmented Lagrangian method (ALM) is an optimization technique that combines the Lagrangian multiplier method and the penalty function method, mainly used to solve constrained optimization problems. Its core idea is to transform the original constrained problem into an unconstrained optimization problem by introducing Lagrangian multipliers and quadratic penalty terms, while improving the convergence speed and numerical stability.

[0040] The basic idea of the alternating direction method of multipliers (ADMM) is to decompose a complex optimization problem into multiple sub-problems that can be solved in parallel, and alternately optimize the primal variables and dual variables to achieve efficient convergence.

[0041] The augmented Lagrangian and the alternating direction method of multipliers are used to solve the objective optimization function, that is, fix other variables and optimize a specified variable alone, and finally achieve the convergence of the result through alternating optimization to obtain the optimal discrete feature selection matrix. Alternating optimization is a process of continuous alternating update until the stopping iteration condition is reached. The stopping iteration condition is that the objective value is less than the set threshold or the number of iterations reaches the maximum number of iterations, and the optimal parameters are obtained, that is, the result convergence is completed.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] The present invention proposes a key scenario feature selection method for an unsupervised and parameter-free autonomous driving system (ADS). By constructing a robust discriminant feature space and combining sparse constraints with a polarization discrete scoring mechanism, accurate separation of safety critical scenarios is achieved. Specifically, the -norm is used as the distance metric in the spatial dimension to enhance the geometric separability of the scene safety domain and the risk domain. At the same time, the -norm is introduced to impose a sparsity constraint on the test cases, effectively suppressing the interference of abnormal samples. On this basis, a discrete scoring mechanism with a polarization effect is designed. By strengthening the weight difference of highly discriminative features, the decision boundary is driven to shrink towards the low-density region, realizing the rapid decoupling and efficient partitioning of safety scenarios and key scenario features without prior labeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic diagram of the logical processing of a feature selection method for autonomous driving simulation testing provided by an embodiment of the present invention;

[0045] Figure 2 is a schematic flowchart of a feature selection method for autonomous driving simulation testing provided by an embodiment of the present invention;

[0046] Figure 3 is a performance table of a feature selection method for autonomous driving simulation testing provided by an embodiment of the present invention under different classifier models and different numbers of features;

[0047] Figure 4 is a violin data view of different algorithm comparison tests provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0049] As Figure 1 , Figure 2 shown, this embodiment provides a method for feature selection in autonomous driving simulation testing, including the following steps:

[0050] T1. Collect an unlabeled test set of the autonomous driving system from the simulation platform, that is, the autonomous driving simulation test set, denoted as the first test set , , R is the set of real numbers, represents the set of d×n-dimensional real matrices; the first test set contains d driving scenario features; each driving scenario feature corresponds to n test cases. Preprocess the first test set, specifically centering processing, to generate the second test set ; the processing process is expressed as:

[0051] ;

[0052] where represents an all-ones column vector of n×1, represents an n×n matrix, and each element of the matrix is . Through matrix centering processing, the mean value of the second test set can be adjusted to zero, facilitating subsequent statistical analysis.

[0053] Randomly generate a projection matrix that satisfies the orthogonal constraint.

[0054] Construct a sparse weight matrix, including two column-sparse weight matrices and , and ; The i-th diagonal element of is expressed as ;

[0055] The i-th diagonal element of is expressed as ;

[0056] where is the i-th column element of is The i-th column element of is a preset constant; is a sufficiently small constant used to avoid overfitting.

[0057] T2. Use the K-means clustering algorithm to obtain the initial pseudo-labels of the second test set; the pseudo-labels are denoted as , and ; among them, 0 corresponds to the safe scenario, and 1 corresponds to the critical scenario.

[0058] T3. Use the projection matrix to project the second test set onto the r-dimensional discriminant feature space, ; in this embodiment, the discriminant feature space is constructed using the objective function, and the objective function is expressed as:

[0059] ;

[0060] Among them, represents the Frobenius norm constraint; represents the projection matrix, and ; is the clustering centroid matrix of represents the identity matrix, represents the orthogonal constraint imposed on

[0061] The Frobenius norm is a norm of a matrix used to measure the overall size of the matrix. It is defined as the square root of the sum of the squares of the absolute values of the elements of the matrix.

[0062] To increase the discriminability of the feature space so that the low-dimensional feature space can effectively distinguish critical scenarios, use the constrained trace ratio objective to minimize the distance between projection points among test cases belonging to the same class and maximize the distance between test cases belonging to different classes, thereby learning the discriminant feature space. The constrained trace ratio objective is a special form of matrix expression of linear discriminant analysis.

[0063] T4. Use the sparse weight matrix to impose sparse weights on the projection matrix; after imposing the sparse weights, the objective function is expressed as:

[0064] ;

[0065] Among them, represents the clustering centroid matrix of

[0066] Impose sparsity on the constrained trace ratio objective through the sparse weight matrix to enhance the robustness of the constructed discriminant feature space to abnormal test cases, that is, robust trace ratio analysis.

[0067] Construct the objective optimization function for the discriminant feature space; in this embodiment, add norm constraint to the objective function, and obtain the objective optimization function, which is expressed as:

[0068] ;

[0069] Among them, represents norm constraint.

[0070] The norm constraint is a mixed norm constraint method that combines structural sparsity and feature selection ability. Its purpose is to make the norm of each row as small as possible, and there are as many 0 elements as possible within the row.

[0071] Imposing norm constraint on the projection matrix to impose feature sparsity is the discrete scoring mechanism of the present invention. It is to learn discrete feature scores, polarize the scores of each scenario feature, so as to form a clearer and more well-defined decision boundary, make the objective optimization function applicable to the feature selection task, and obtain a discrete and robust discriminant feature space.

[0072] Solve the objective optimization function to obtain a discrete feature selection matrix; the feature selection matrix is the optimal matrix for scene definition feature dimension reduction. In this embodiment, the augmented Lagrangian method and the alternating direction method of multipliers are used to solve the objective optimization function, and the feature selection matrix is obtained.

[0073] The augmented Lagrangian method (ALM) is an optimization technique that combines the Lagrangian multiplier method and the penalty function method, mainly used to solve constrained optimization problems. Its core idea is to transform the original constrained problem into an unconstrained optimization problem by introducing Lagrangian multipliers and quadratic penalty terms, while improving the convergence speed and numerical stability.

[0074] The basic idea of the alternating direction method of multipliers (ADMM) is to decompose a complex optimization problem into multiple sub-problems that can be solved in parallel, and alternately optimize the primal variables and dual variables to achieve efficient convergence.

[0075] The augmented Lagrangian and alternating direction multiplier method are used to solve the objective optimization function, that is, fix other variables and optimize a specified variable alone, and finally achieve the convergence of the result through alternating optimization to obtain the optimal discrete feature selection matrix. Alternating optimization is a process of continuous alternating update until the stopping iteration condition is reached. The stopping iteration condition is that the objective value is less than the set threshold or the number of iterations reaches the maximum number of iterations, and the optimal parameters are obtained, that is, the result convergence is completed.

[0076] The target problem is decomposed into a pseudo-label matrix , a clustering centroid matrix , a projection matrix , a sparse weight and five sub-problems and optimize them alternately. The specific implementation steps are as follows:

[0077] Perform K-means clustering on the surrogate variable to obtain the optimal pseudo-label matrix ; calculate the clustering centroid of the surrogate variable ; dynamically update the sparse weight diagonal matrix and according to the projection vector and the clustering centroid ; solve the projection matrix through the discrete sparse subspace optimization algorithm to screen out non-zero scenario features.

[0078] T5. Generate the final scenario feature subset according to the non-zero row indices of the feature selection matrix . The final scenario feature subset is used for subsequent test case prioritization or simulation test.

[0079] The following is a further description of the technical effects of the present invention in this embodiment by combining comparative experiments and visualization experiments:

[0080] 1. Experimental conditions and settings

[0081] The Baidu Apollo 5.0 platform was evaluated using the LGSVL simulator. The scenarios included four different roads on the San Francisco map. These scenarios were designed to accurately simulate various driving situations, such as turning, lane changing, stopping at signals, interacting with pedestrians or obstacles, overtaking, and parking. In general, this experiment utilized 25,000 unique test scenarios, each scenario containing 65 features, of which 10,847 were classified as critical scenarios (collision or violation of safety distance), and 14,153 were classified as safe scenarios. Feature effectiveness is defined as the ability to handle data collinearity and the ability to distinguish between safe and unsafe scenarios. In this embodiment, critical scenarios refer to unsafe scenarios.

[0082] To evaluate the sufficiency of the selected features, four machine learning models were experimentally trained in the present invention: Random Forest (RF), k-Nearest Neighbor (KNN), Decision Tree (DT), and Multi-Layer Perceptron (MLP). Then, the trained models were used to predict the results of previously unobserved driving scenarios without running them in a simulator. All experiments were repeated 10 times, and the average results were recorded as performance metrics. For comparison, the experiments in the present invention were compared with Laplacian Score (LS) and Fast Sparse Discriminant K-means method (FSDK), and Random Feature Selection (denoted as Random) was used as a baseline. The experimental results are as Figure 3 shown.

[0083] 2. Comparative Experiments and Analysis

[0084] As can be seen from Figure 3 , the four applied models continuously showed good performance in classification. It is worth noting that when only 15 features were selected, the classification results obtained in the present invention were comparable to those using all 65 features, emphasizing its effectiveness in selecting the most discriminative features. In contrast, when only 5 features were selected, the classification performance of all methods decreased, which may be because 5 features were not sufficient to fully represent the scenarios in the simulation.

[0085] 3. Visualization Experiments and Analysis

[0086] In the experiments of the present invention, different algorithms (LS, FSDK) were used to select the top 4 features from the test set, and the label distributions of these features were shown using violin plots, as Figure 4 shown. The features selected by the present invention (Ours) are superior to those selected by LS because the corresponding label distributions show clearer differences. Compared with FSDK, although there is some overlap in the selected features, the present invention still shows superior performance. Although the features selected by FSDK also show obvious differences in label distribution, the overall quality of the features selected by the present invention is better.

[0087] This indicates that the present invention can learn a more discriminative feature space, thereby obtaining features that can more effectively distinguish safe and unsafe scenarios. According to the disclosure and teachings of the above specification, those skilled in the art to which the present invention pertains can also make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the protection scope of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for convenience of description and do not constitute any limitation to the present invention.

Claims

1. A feature selection method for autonomous driving simulation test, characterized in that: The following steps are involved: T1, collect an unlabeled test set of the autonomous driving system, recorded as the first test set; the first test set includes d driving scene features; each of the driving scene features corresponds to n test cases; Preprocessing the first test set to generate a second test set; Construct a projection matrix that satisfies orthogonal constraints; Construct a sparse weight matrix; The sparse weight matrix includes two, which are respectively expressed as and ,and ; T2. Use a clustering algorithm to obtain the initial pseudo-labels of the second test set, and divide the driving scene features into safety scenes or key scenes; the pseudo-labels are recorded as ,and ; Wherein, 0 corresponds to the safety scenario, and 1 corresponds to the critical scenario; T3, constructing a discriminant feature space using the objective function; projecting the second test set to the r-dimensional discriminant feature space using the projection matrix, ; The objective function is expressed as: ; in, represents the Frobenius norm constraint; represents the projection matrix, and , R is the set of real numbers, represents the set of d×n dimensional real matrices; for The cluster centroid matrix of represents the identity matrix, Indicated in Orthogonality constraints imposed on ; The i-th diagonal element of is represented by : ; The i-th diagonal element of is represented by : ; in, for The i-th column element of for The i-th column element of ; is a preset constant; T4. Apply sparse weights to the projection matrix using the sparse weight matrix; after applying the sparse weights, the objective function is expressed as: ; in, express The cluster centroid matrix of Adding to the objective function Norm constraint, the target optimization function is obtained, which is expressed as: ; in, express Norm constraints; Solving the objective optimization function to obtain a discrete feature selection matrix; T5. Generate a final scene feature subset according to the non-zero row indexes of the feature selection matrix.

2. The method for selecting features for an autonomous driving simulation test according to claim 1, characterized in that: The preprocessing is to perform centralization processing on the first test set to obtain the second test set, that is: ; in, represents the first test set, , R is the set of real numbers, represents the set of d×n dimensional real matrices; represents the second test set; represents an n×1 column vector of all 1s, represents an n×n matrix, and each element of the matrix is .

3. The method for selecting features for an autonomous driving simulation test according to claim 1, characterized in that: In T4, the objective optimization function is solved using the augmented Lagrangian method and the alternating direction multiplier method to obtain the feature selection matrix.

Citation Information

Patent Citations

  • URMFFS (unsupervised regularization matrix factorization feature selection) method

    CN107203787A

  • Vehicle lane changing trajectory prediction method based on random forest and improved Informer model

    CN117807413A