Automatic driving simulation test feature selection method

By building a robust discriminant feature space in the test of autonomous driving system and combining sparse constraints and polarized discrete scoring mechanisms, the problem of difficulty in efficiently identifying features related to key test scenarios in the existing technology is solved, and the efficiency and accuracy of autonomous driving system testing is achieved.

CN120030315AActive Publication Date: 2025-05-23SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202510497546.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

It is difficult for existing autonomous driving system testing technology to efficiently identify and select the features most relevant to key test scenarios, making it difficult to meet actual engineering needs for testing efficiency and cost.

Method used

Through an autonomous driving simulation test feature selection method, the algorithm is used to quickly iteratively optimize, and a robust discriminant feature space is built without prior labeling data, combining sparse constraints and polarized discrete scoring mechanisms to achieve accurate separation of safe critical scenarios.

Benefits of technology

实现了在无需先验标注数据条件下,安全场景与关键场景特征的快速解耦与高效划分,提高了自动驾驶系统测试的效率和准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030315A_ABST
    Figure CN120030315A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving simulation test feature selection method, and relates to the technical field of automatic driving system testing, and the method comprises the steps: T1, collecting a label-free test set of an automatic driving system, and recording the label-free test set as a first test set; preprocessing the first test set to generate a second test set; constructing a projection matrix meeting orthogonal constraints; constructing a sparse weight matrix; t2, acquiring an initial pseudo tag of the second test set, and dividing the driving scene features into a safety scene or a key scene; t3, projecting the second test set to the discriminant feature space by using the projection matrix; t4, sparse weight is applied to the projection matrix; constructing a target optimization function for discriminating the feature space; solving the target optimization function to obtain a discrete feature selection matrix; and T5, generating a final scene feature subset according to the non-zero line index of the feature selection matrix. By constructing a robust discrimination feature space and combining sparse constraint and a polarization discrete scoring mechanism, efficient recognition and division of key scene features are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving system testing, and in particular to a method for selecting features for autonomous driving simulation testing. Background Art

[0002] In the large-scale application process of autonomous driving systems (ADS), safety and reliability verification in complex dynamic environments has become a key bottleneck. Existing virtual simulation testing technology needs to cover a large number of extreme scenarios to evaluate the system behavior boundaries, but its scenario generation faces multi-dimensional dynamic coupling challenges: on the one hand, the kinematic characteristics of ADS, road topology, multi-agent interaction rules and the co-evolution of environmental parameters (such as lighting and weather) significantly increase the complexity of the scene space; on the other hand, traditional exhaustive testing requires traversing all potential feature combinations, which causes exponential computing resource consumption, resulting in test efficiency and cost that are difficult to meet actual engineering needs. It is urgent to build a high-coverage, low-redundancy scenario generation optimization method to break through the verification efficiency boundary.

[0003] Current research on low-cost autonomous driving testing focuses on test case optimization techniques, including test suite minimization, test case selection and prioritization. This technique mainly reduces the size of the test set through heuristic strategies, but fails to solve the dimensionality disaster problem caused by the high-dimensional scene feature space. Although deep learning has shown potential advantages in the field of feature dimensionality reduction, its inherent black box characteristics lead to a lack of feature interpretability, and the training paradigm that relies on a large amount of labeled data is resource-intensive. In contrast to deep learning, traditional feature selection methods have the advantage of physical interpretability, which is more suitable for autonomous driving test task scenarios. However, existing research has not yet established a feature selection framework for the specific characteristics of autonomous driving scenarios. Summary of the invention

[0004] In response to the problems existing in the prior art, the present invention provides a feature selection method for autonomous driving simulation testing. Through rapid iterative optimization of the algorithm, it can efficiently identify and select the features most relevant to key test scenarios without the need for prior labeled data.

[0005] The technical solution of the present invention is achieved in this way: A method for selecting features for an autonomous driving simulation test comprises the following steps: T1, collect an unlabeled test set of the autonomous driving system, recorded as the first test set; the first test set includes d driving scene features; each of the driving scene features corresponds to n test cases; Preprocessing the first test set to generate a second test set; Construct a projection matrix that satisfies orthogonal constraints; construct a sparse weight matrix; T2, obtaining the initial pseudo-labels of the second test set, and dividing the driving scene features into safety scenes or critical scenes; T3, using the projection matrix, projecting the second test set to an r-dimensional discriminant feature space, ; T4. Apply sparse weights to the projection matrix using the sparse weight matrix; construct a target optimization function for the discriminant feature space; solve the target optimization function to obtain a discrete feature selection matrix; the feature selection matrix is ​​the optimal matrix for scene definition feature dimensionality reduction.

[0006] T5. Generate a final scene feature subset according to the non-zero row index of the feature selection matrix. The final scene feature subset is used for subsequent test case prioritization or simulation testing.

[0007] As a further optimization of the above solution, the preprocessing is to perform centralization processing on the first test set to obtain the second test set, that is: ; in, represents the first test set, , R is the set of real numbers, represents the set of d×n dimensional real matrices; represents the second test set; represents an n×1 column vector of all 1s, represents an n×n matrix, and each element of the matrix is .

[0008] Through matrix centering, the mean of the second test set can be adjusted to zero, which is convenient for subsequent statistical analysis.

[0009] As a further optimization of the above solution, in T2, the initial pseudo-label of the second test set is obtained by using a clustering algorithm; the pseudo-label is recorded as ,and ; Among them, 0 corresponds to the safety scenario, and 1 corresponds to the critical scenario.

[0010] Furthermore, the clustering algorithm adopts K-means clustering algorithm.

[0011] As a further optimization of the above scheme, in T3, the discriminant feature space is constructed using an objective function, and the objective function is expressed as: ; in, represents the Frobenius norm constraint; represents the projection matrix, and , R is the set of real numbers, represents the set of d×n dimensional real matrices; for The cluster centroid matrix of represents the identity matrix, Indicated in Orthogonality constraints imposed on .

[0012] The Frobenius norm is a matrix norm that measures the overall size of the matrix. It is defined as the square root of the sum of the squares of the absolute values ​​of the matrix elements.

[0013] In order to increase the discriminability of the feature space so that the low-dimensional feature space can effectively distinguish key scenarios, the constrained trace ratio objective is used to minimize the distance between the projected points in the same class of test cases and maximize the distance between test cases in different classes, thereby learning the discriminative feature space. The constrained trace ratio objective is a special form of matrix expression of linear discriminant analysis.

[0014] As a further optimization of the above scheme, the sparse weight matrix includes two, which are respectively expressed as and ,and ; The i-th diagonal element of is represented by : ; The i-th diagonal element of is represented by : ; in, for The i-th column element of for The i-th column element of ; is a preset constant; is a sufficiently small constant to avoid overfitting; After applying sparse weights, the objective function is expressed as: ; in, express The cluster centroid matrix of .

[0015] Sparsity is imposed on the constrained trace ratio target through a sparse weight matrix to enhance the robustness of the constructed discriminative feature space to abnormal test cases.

[0016] As a further optimization of the above scheme, add norm constraint, the objective optimization function is obtained, which is expressed as: ; Among them, denotes the norm constraint.

[0017] The norm constraint is a mixed norm constraint method that combines structured sparsity and feature selection capabilities. Its purpose is to make the norm of each row as small as possible, with as many 0 elements as possible appearing within the row.

[0018] Applying the norm constraint on the projection matrix is to impose feature sparsity in order to learn discrete feature scores, polarize the scores of each scenario feature, form a clearer and more well-defined decision boundary, and make the objective optimization function applicable to the feature selection task.

[0019] As a further optimization of the above solution, in the T4, the augmented Lagrangian method and the alternating direction method of multipliers are used to solve the objective optimization function to obtain the feature selection matrix.

[0020] The augmented Lagrangian method (ALM) is an optimization technique that combines the Lagrangian multiplier method and the penalty function method, mainly used to solve constrained optimization problems. Its core idea is to transform the original constrained problem into an unconstrained optimization problem by introducing Lagrangian multipliers and quadratic penalty terms, while improving the convergence speed and numerical stability.

[0021] The basic idea of the alternating direction method of multipliers (ADMM) is to decompose a complex optimization problem into multiple sub-problems that can be solved in parallel, and alternately optimize the primal variables and dual variables to achieve efficient convergence.

[0022] Using the augmented Lagrangian and the alternating direction method of multipliers to solve the objective optimization function means fixing other variables and optimizing a specified variable alone, and finally achieving the convergence of the result through alternating optimization. Alternating optimization is a process of continuous alternating update until the stopping iteration condition is reached. The stopping iteration condition is that the objective value is less than the set threshold, or the number of iterations reaches the maximum number of iterations, obtaining the optimal parameters, that is, completing the convergence of the result.

[0023] Compared with the prior art, the present invention has the following beneficial effects: This paper proposes an unsupervised and parameter-free method for selecting key scene features of an automated driving system (ADS). By constructing a robust discriminant feature space and combining sparse constraints with a polarized discrete scoring mechanism, the accurate separation of safety critical scenes can be achieved. -norm is used as a distance metric in the spatial dimension to enhance the geometric separability of the scene safety domain and risk domain, while introducing The -norm imposes sparsity constraints on test cases, effectively suppressing the interference of abnormal samples; on this basis, a discrete scoring mechanism with a polarization effect is designed to drive the decision boundary to shrink to the low-density area by strengthening the weight difference of highly discriminative features, thereby achieving rapid decoupling and efficient division of safety scenarios and key scenario features without the need for prior labeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a logical processing diagram of a feature selection method for an autonomous driving simulation test provided by an embodiment of the present invention; Figure 2 It is a flowchart of a method for selecting features in an autonomous driving simulation test provided by an embodiment of the present invention; Figure 3 This is a performance table of a feature selection method for autonomous driving simulation testing provided by an embodiment of the present invention under different classifier models and different numbers of features; Figure 4 This is a violin data view of a comparative test of different algorithms provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the present invention more clear, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] like Figure 1 , Figure 2 As shown, this embodiment provides a method for selecting features for an autonomous driving simulation test, comprising the following steps: T1: Collect the unlabeled test set of the autonomous driving system from the simulation platform, that is, the autonomous driving simulation test set, recorded as the first test set , , R is the set of real numbers, represents a set of d×n dimensional real number matrices; the first test set contains d driving scene features; each driving scene feature corresponds to n test cases. The first test set is preprocessed, specifically centralized, to generate the second test set ; The processing process is expressed as: ; in, represents an n×1 column vector of all 1s, represents an n×n matrix, and each element of the matrix is Through matrix centering, the mean of the second test set can be adjusted to zero, which is convenient for subsequent statistical analysis.

[0027] Randomly generate a projection matrix that satisfies orthogonal constraints .

[0028] Construct a sparse weight matrix, including two column sparse weight matrices and ,and ; The i-th diagonal element of is represented by : ; The i-th diagonal element of is represented by : ; in, for The i-th column element of for The i-th column element of ; is a preset constant; is a sufficiently small constant to avoid overfitting.

[0029] T2, use K-means clustering algorithm to obtain the initial pseudo-label of the second test set; the pseudo-label is recorded as ,and ; Among them, 0 corresponds to a safe scenario and 1 corresponds to a critical scenario.

[0030] T3, using the projection matrix, project the second test set into the r-dimensional discriminant feature space, In this embodiment, the discriminant feature space is constructed using the objective function, and the objective function is expressed as: ; in, represents the Frobenius norm constraint; represents the projection matrix, and ; for The cluster centroid matrix of represents the identity matrix, Indicated in Orthogonality constraints imposed on .

[0031] The Frobenius norm is a matrix norm that measures the overall size of the matrix. It is defined as the square root of the sum of the squares of the absolute values ​​of the matrix elements.

[0032] In order to increase the discriminability of the feature space so that the low-dimensional feature space can effectively distinguish key scenarios, the constrained trace ratio objective is used to minimize the distance between the projected points in the same class of test cases and maximize the distance between test cases in different classes, thereby learning the discriminative feature space. The constrained trace ratio objective is a special form of matrix expression of linear discriminant analysis.

[0033] T4. Use the sparse weight matrix to apply sparse weights to the projection matrix. After applying the sparse weights, the objective function is expressed as: ; in, express The cluster centroid matrix of .

[0034] Sparsity is imposed on the constrained trace ratio target through a sparse weight matrix to enhance the robustness of the constructed discriminative feature space for abnormal test cases, namely, robust trace ratio analysis.

[0035] Construct the target optimization function of the discriminant feature space; in this embodiment, add Norm constraint, the target optimization function is obtained, which is expressed as: ; in, express Norm constraint.

[0036] The norm constraint is a hybrid norm constraint method that combines structured sparsity with feature selection capabilities. Its purpose is to make each row The norm is as small as possible, and there are as many 0 elements as possible in the row.

[0037] In the projection matrix Apply The norm constraint is used to impose feature sparsity, that is, the discrete scoring mechanism of the present invention, which is to learn discrete feature scores and polarize the scores of each scene feature to form a clearer and well-defined decision boundary, so that the objective optimization function is suitable for feature selection tasks and a discrete and robust discriminant feature space is obtained.

[0038] The objective optimization function is solved to obtain a discrete feature selection matrix; the feature selection matrix is ​​the optimal matrix for scene definition feature dimensionality reduction. In this embodiment, the objective optimization function is solved using the augmented Lagrangian method and the alternating direction multiplier method to obtain the feature selection matrix .

[0039] The Augmented Lagrangian Method (ALM) is an optimization technique that combines the Lagrangian multiplier method and the penalty function method. It is mainly used to solve optimization problems with constraints. Its core idea is to transform the original constrained problem into an unconstrained optimization problem by introducing Lagrangian multipliers and quadratic penalty terms, while improving the convergence speed and numerical stability.

[0040] The basic idea of ​​the Alternating Direction Method of Multipliers (ADMM) is to achieve efficient convergence by decomposing a complex optimization problem into multiple sub-problems that can be solved in parallel and alternately optimizing the primal variables and the dual variables.

[0041] The augmented Lagrangian and alternating direction multiplier methods are used to solve the target optimization function, that is, to optimize a specified variable alone while fixing other variables, and finally achieve convergence of the results through alternating optimization to obtain the optimal discrete feature selection matrix. Alternating optimization is a process of continuous alternating updates until the stopping iteration condition is reached, where the stopping iteration condition is that the target value is less than the set threshold, or the number of iterations reaches the maximum number of iterations, and the optimal parameters are obtained, that is, the result convergence is completed.

[0042] Solve the target problem into a pseudo-label matrix , cluster centroid matrix , projection matrix , sparse weight and Five sub-problems are optimized alternately. The specific implementation steps are: For proxy variables Perform K-means clustering to obtain the optimal pseudo-label matrix ; Calculate proxy variables The cluster centroid of ; According to the projection vector and cluster centroids Dynamically update the sparse weight diagonal matrix and ; Solve the projection matrix through discrete sparse subspace optimization algorithm , filter non-zero scene features.

[0043] T5. Select matrix based on features The non-zero row index of generates the final scenario feature subset. The final scenario feature subset is used for subsequent test case prioritization or simulation testing.

[0044] The following is a further description of the technical effects of the present invention in combination with comparative experiments and visualization experiments in this embodiment: 1. Experimental conditions and settings The Baidu Apollo 5.0 platform was evaluated using the LGSVL simulator. The scenarios consisted of four different roads on a map of San Francisco. The scenarios were designed to accurately simulate various driving situations, such as turning, changing lanes, stopping at signals, interacting with pedestrians or obstacles, overtaking, and parking. In total, this experiment utilized 25,000 unique test scenarios, each containing 65 features, of which 10,847 were classified as critical scenarios (collision or violation of safe distance) and 14,153 were classified as safe scenarios. Feature validity is defined as the ability to handle data collinearity and the ability to distinguish between safe and unsafe scenarios. In this example, critical scenarios refer to unsafe scenarios.

[0045] To evaluate the adequacy of the selected features, the present invention experiment trained four machine learning models: random forest (RF), k-nearest neighbor (KNN), decision tree (DT), and multi-layer perceptron (MLP). The trained models were then used to predict the results of previously unobserved driving scenarios without running them in the simulator. All experiments were repeated 10 times and the average results were recorded as performance indicators. For comparison, the present invention experiment was compared with the Laplace score (LS) and the fast sparse discriminant K-means method (FSDK), and random feature selection (denoted as Random) was used as a baseline. The experimental results are shown in Figure 2. Figure 3 shown.

[0046] 2. Comparative experiment and analysis from Figure 3 As can be seen in the , the 4 models applied consistently perform well in classification performance. It is noteworthy that when only 15 features are selected, the classification results obtained by the present invention are comparable to those using all 65 features, emphasizing its effectiveness in selecting the most discriminative features. In contrast, when only 5 features are selected, the classification performance of all methods decreases, which may be because 5 features are not enough to fully represent the scenario in the simulation.

[0047] 3. Visualization experiment and analysis The experiment of this invention uses different algorithms (LS, FSDK) to select the first 4 features from the test set, and uses violin plots to show the label distribution of these features, such as Figure 4As shown in Figure 2, the features selected by our method (Ours) are better than those selected by LS because the corresponding label distributions show clearer differences. Compared with FSDK, our method still shows superior performance despite some overlap in the selected features. Although the features selected by FSDK also show obvious differences in label distribution, the features selected by our method are of better overall quality.

[0048] This shows that the present invention is able to learn a more discriminative feature space, thereby obtaining features that can more effectively distinguish between safe and unsafe scenarios. According to the disclosure and teachings of the above specification, technicians in the field to which the present invention belongs can also change and modify the above implementation. Therefore, the present invention is not limited to the specific implementations disclosed and described above, and some modifications and changes to the present invention should also fall within the scope of protection of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for the convenience of explanation and do not constitute any limitation to the present invention.

Claims

1. A feature selection method for autonomous driving simulation test, characterized in that: The following steps are involved: T1, collect an unlabeled test set of the autonomous driving system, recorded as the first test set; the first test set includes d driving scene features; each of the driving scene features corresponds to n test cases; Preprocessing the first test set to generate a second test set; Construct a projection matrix that satisfies orthogonal constraints; Construct a sparse weight matrix; T2, obtaining the initial pseudo-labels of the second test set, and dividing the driving scene features into safety scenes or critical scenes; T3, using the projection matrix, projecting the second test set to an r-dimensional discriminant feature space, ; T4, applying sparse weights to the projection matrix using the sparse weight matrix; Constructing a target optimization function of the discriminant feature space; Solving the objective optimization function to obtain a discrete feature selection matrix; T5. Generate a final scene feature subset according to the non-zero row indexes of the feature selection matrix.

2. The method for selecting features for an autonomous driving simulation test according to claim 1, characterized in that: The preprocessing is to perform centralization processing on the first test set to obtain the second test set, that is: ; in, represents the first test set, , R is the set of real numbers, represents the set of d×n dimensional real matrices; represents the second test set; represents an n×1 column vector of all 1s, represents an n×n matrix, and each element of the matrix is .

3. The method for selecting features for an autonomous driving simulation test according to claim 1, characterized in that: In T2, the clustering algorithm is used to obtain the initial pseudo-label of the second test set; the pseudo-label is recorded as ,and ; Among them, 0 corresponds to the safety scenario, and 1 corresponds to the critical scenario.

4. The method for selecting features for an autonomous driving simulation test according to claim 3, characterized in that: In T3, the discriminant feature space is constructed using an objective function, and the objective function is expressed as: ; in, represents the Frobenius norm constraint; represents the projection matrix, and , R is the set of real numbers, represents the set of d×n dimensional real matrices; for The cluster centroid matrix of represents the identity matrix, Indicated in Orthogonality constraints imposed on .

5. The method for selecting features for an autonomous driving simulation test according to claim 4, characterized in that: The sparse weight matrix includes two, which are respectively expressed as and ,and ; The i-th diagonal element of is represented by : ; The i-th diagonal element of is represented by : ; in, for The i-th column element of for The i-th column element of ; is a preset constant; After applying sparse weights, the objective function is expressed as: ; in, express The cluster centroid matrix of .

6. The method for selecting features for an autonomous driving simulation test according to claim 5, characterized in that: Adding to the objective function norm constraint, the objective optimization function is obtained, which is expressed as: ; in, express Norm constraint.

7. The method for selecting features for an autonomous driving simulation test according to claim 1, characterized in that: In T4, the objective optimization function is solved using the augmented Lagrangian method and the alternating direction multiplier method to obtain the feature selection matrix.

Citation Information

Patent Citations

  • URMFFS (unsupervised regularization matrix factorization feature selection) method

    CN107203787A

  • Robust semi-supervised sparse feature selection method based on self-adjusting graph

    CN111652265A

  • Local adaptive feature selection method and device based on robust subspace representation

    CN113378926A

  • Vehicle lane changing trajectory prediction method based on random forest and improved Informer model

    CN117807413A

  • Image classification method based on low-rank regularization denoising discriminant regression

    CN117876744A

Cited By

  • Risk-aware automatic driving simulation test case sorting method

    CN121144211A